Acoustic signal processing device and acoustic signal processing method
Patent Information
- Application Number
- PCT/JP2026/011275
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-09-25
- Filing Date
- 2026-03-23
- Publication Date
- 2026-10-01
Smart Images

Figure JP2026011275_01102026_PF_FP_ABST
Abstract
Description
Acoustic signal processing apparatus and acoustic signal processing method
[0001] The present disclosure relates to an acoustic signal processing apparatus and the like.
[0002] Patent Document 1, Patent Document 2, and Non-Patent Document 1 disclose technologies related to acoustic signal processing.
[0003] Japanese National Publication of International Patent Application No. 2023-511862, Japanese Unexamined Patent Application Publication No. 2019-22049
[0004] ISO 23090-4:202#(X), ISO / IEC JTC 1 / SC 29 / WG 6, Date: 2024-07-19, Information technology-Coded representation of immersive media-Part 4: MPEG-I immersive audio, DIS stage
[0005] Further improvements to acoustic signal processing are desired. An object of the present disclosure is to improve acoustic signal processing.
[0006] An acoustic signal processing apparatus according to an aspect of the present disclosure includes a memory and a circuit capable of accessing the memory. In operation, the circuit prepares a transfer function for calculating a sound signal representing sound that reaches a listener in a virtual space from a sound source in the virtual space, and calculates the sound signal based on the transfer function. In preparing the transfer function, the circuit detects an amount of change in the relative positional relationship between the sound source and the listener in the virtual space, and updates the transfer function when the amount of change in the relative positional relationship exceeds a threshold range.
[0007] Note that these general or specific aspects may be implemented by a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, and may also be implemented by any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.
[0008] The present disclosure can contribute to improving acoustic signal processing.
[0009] This is a diagram showing an example of a 3D audio playback system in an embodiment. This is a block diagram showing an example of the configuration of an encoding device in an embodiment. This is a block diagram showing an example of the configuration of a decoding device in an embodiment. This is a block diagram showing another example of the configuration of an encoding device in an embodiment. This is a block diagram showing another example of the configuration of a decoding device in an embodiment. This is a block diagram showing an example of the configuration of a decoder in an embodiment. This is a block diagram showing another example of the configuration of a decoder in an embodiment. This is a diagram showing an example of the physical configuration of an acoustic signal processing device in an embodiment. This is a diagram showing an example of the physical configuration of an encoding device in an embodiment. This is a block diagram showing an example of the configuration of a rendering unit in an embodiment. This is a flowchart showing an example of the operation of an acoustic signal processing device in an embodiment. This is a block diagram showing an example of the configuration of a rendering unit for pipeline processing. This is a conceptual diagram showing a specific example of pipeline processing. This is a graph showing the computational load of each stage in pipeline processing. This is a block diagram showing an example of the configuration of a rendering unit for SESS-related processing. This is a diagram showing the correspondence between stages and threads. This is a conceptual diagram showing a sound source and a listener. This is a flowchart showing an example of a determination process for whether or not to update the sound image expansion filter. This is a conceptual diagram showing a process for detecting the shape of a sound source as seen from the listener's position. This is a conceptual diagram showing an example of a determination process in the first embodiment. This is a conceptual diagram showing an example of a determination process in a modified example of the first embodiment. This is a diagram showing an example of configuration information for a wearable profile. This figure shows an example of configuration information for a mobile profile. This figure shows an example of configuration information for a home profile. This figure shows an example of configuration information for a cloud profile. This is a conceptual diagram showing an example of the judgment process in the second embodiment. This is a conceptual diagram showing another example of the judgment process in the second embodiment. This is a conceptual diagram showing an example of the judgment process in the third embodiment. This is a conceptual diagram showing another example of the judgment process in the third embodiment. This is a conceptual diagram showing an example of the judgment process in the fourth embodiment. This is a conceptual diagram showing another example of the judgment process in the fourth embodiment. This is a conceptual diagram showing the relative positional relationship between the sound source and the listener using vectors. This is a conceptual diagram showing an example where the distance between the sound source and the listener does not change.This is a conceptual diagram showing an example where the direction connecting the sound source and the listener does not change. This is a conceptual diagram showing an example of a threshold set based on the arrangement of the HRTF group. This is a block diagram showing an example configuration of an acoustic signal processing device that calculates a sound signal indicating the sound reaching the listener. This is a flowchart showing an example of a determination process based on the relative positional relationship between the sound source and the listener. This is a block diagram showing a basic configuration example of an acoustic signal processing device in an embodiment. This is a flowchart showing a basic operation example of an acoustic signal processing device in an embodiment. This is a flowchart showing an operation example related to the preparation of the transfer function.
[0010] (Introduction) In recent years, technological development has been progressing in virtual reality (VR) and augmented reality (AR), which are virtual experiences from the user's perspective. With VR or AR, users can experience (i.e., become immersed in) a virtual space as if they were actually there. In particular, since the sense of immersion is enhanced by combining a three-dimensional visual experience with a three-dimensional auditory experience, technologies related to three-dimensional auditory experiences are also considered important in VR and AR.
[0011] Against this backdrop, the development of the MPEG-I standard is underway as a standard for the reproduction of three-dimensional sound in virtual spaces such as VR or AR, and Non-Patent Document 1 has already been published as a draft version of the international standard document. An international standard document is called IS (International Standard), and a draft version of an international standard document is called DIS (Draft International Standard).
[0012] For example, in a virtual space defined by the MPEG-I standard, it is possible to use a process to enlarge the sound image so that a sound image corresponding to the shape of the sound source is obtained from the sound signal (also expressed as an audio signal or acoustic signal) associated with the sound source having a shape. This technology is called SESS (Spatially Extended Sound Sources). The SESS technology is disclosed in Patent Document 1. By using SESS, the presence of the sound source is enhanced, and the virtual space is reproduced more realistically.
[0013] Furthermore, other sound processing may be performed in place of or in addition to SESS. This may result in a more realistic reproduction of the virtual space and an improved sense of immersion.
[0014] However, these acoustic processing tasks may be computationally intensive. For example, preparing the transfer function for calculating the sound signal that represents the sound reaching the listener from the sound source in a virtual space may involve computationally intensive processing.
[0015] Therefore, the acoustic signal processing device of Example 1 comprises a memory and a circuit that can access the memory, wherein the circuit prepares a transfer function for calculating an acoustic signal that represents sound reaching a listener in the virtual space from a sound source in the virtual space, calculates the acoustic signal based on the transfer function, and in preparing the transfer function, the circuit detects the amount of change in the relative positional relationship between the sound source and the listener in the virtual space, and updates the transfer function if the amount of change in the relative positional relationship exceeds a threshold range.
[0016] This can sometimes make it possible to appropriately control the updating of the transfer function in accordance with the amount of change in the relative positional relationship between the sound source and the listener. Therefore, it may be possible to update the transfer function in a timely manner while appropriately suppressing the processing load related to updating the transfer function.
[0017] Furthermore, the acoustic signal processing device of Example 2 may be the acoustic signal processing device of Example 1, wherein the sound source has a shape, the transfer function is determined based on the shape, and the circuit performs a sound image expansion preparation process which includes preparing the transfer function determined based on the shape, and a sound image expansion signal processing which includes calculating the sound signal based on the transfer function determined based on the shape, thereby expanding the sound image based on the shape.
[0018] This may make it possible to appropriately enlarge the sound image. Furthermore, in the sound image enlargement preparation process, it may be possible to appropriately control the update of the transfer function according to the amount of change in the relative positional relationship between the sound source and the listener. Therefore, it may be possible to appropriately suppress the load of the sound image enlargement preparation process.
[0019] Furthermore, the acoustic signal processing device of Example 3 may be the acoustic signal processing device of Example 1 or 2, wherein the relative positional relationship is represented by a vector connecting the sound source and the listener in the virtual space.
[0020] This can sometimes make it possible to appropriately represent the relative positional relationship between the sound source and the listener in the virtual space. Therefore, it may be possible to appropriately detect the amount of change in the relative positional relationship and appropriately control the updating of the transfer function according to the amount of change in the relative positional relationship.
[0021] Furthermore, the acoustic signal processing device of Example 4 may be any of the acoustic signal processing devices of Examples 1 to 3, wherein the amount of variation in the relative positional relationship includes the amount of variation in the direction connecting the sound source and the listener in the virtual space.
[0022] This can make it possible to appropriately control the update of the transfer function in accordance with the amount of change in the direction connecting the sound source and the listener in the virtual space. Therefore, it may be possible to appropriately suppress the processing load related to the update of the transfer function.
[0023] Furthermore, the acoustic signal processing device of Example 5 may be the acoustic signal processing device of Example 4, wherein the direction is represented by the direction of the vector connecting the sound source and the listener in the virtual space.
[0024] This can sometimes make it possible to appropriately represent the direction connecting the sound source and the listener in the virtual space. Therefore, it may be possible to appropriately detect the amount of directional variation and appropriately control the update of the transfer function according to the amount of directional variation.
[0025] Furthermore, the acoustic signal processing device of Example 6 may be the acoustic signal processing device of Example 4 or 5, wherein the amount of variation in the direction is expressed as the difference between the direction at a reference point and the direction at the point in time when the amount of variation in the relative positional relationship is detected.
[0026] This can make it possible to appropriately represent the amount of variation that affects hearing in terms of direction between the sound source and the listener in the virtual space. Therefore, it may be possible to appropriately detect the amount of variation related to direction and appropriately control the updating of the transfer function according to the amount of variation.
[0027] Furthermore, the acoustic signal processing device of Example 7 may be the acoustic signal processing device of Example 6, wherein the time point defined as the reference is the time point at which the transfer function was last updated.
[0028] This can sometimes make it possible to appropriately control the updating of the transfer function in accordance with the amount of change in the direction from the time the transfer function was last updated.
[0029] Furthermore, the acoustic signal processing device of Example 8 may be any of the acoustic signal processing devices of Examples 1 to 7, wherein the amount of variation in the relative positional relationship includes the amount of variation in the distance between the sound source and the listener in the virtual space.
[0030] This can make it possible to appropriately control the updating of the transfer function in response to variations in the distance between the sound source and the listener in the virtual space. Therefore, it may be possible to appropriately suppress the processing load related to updating the transfer function.
[0031] Furthermore, the acoustic signal processing device of Example 9 may be the acoustic signal processing device of Example 8, wherein the distance is represented by the magnitude of the vector connecting the sound source and the listener in the virtual space.
[0032] This can sometimes make it possible to accurately represent the distance between the sound source and the listener in the virtual space. Therefore, it may be possible to appropriately detect the amount of distance variation and appropriately control the update of the transfer function according to the amount of distance variation.
[0033] Furthermore, the acoustic signal processing device of Example 10 may be the acoustic signal processing device of Example 8 or 9, wherein the amount of distance variation is expressed as the ratio of the distance at a reference point to the distance at the point in time when the amount of variation in the relative positional relationship is detected.
[0034] This can sometimes make it possible to appropriately represent the amount of variation that affects hearing in relation to the distance between the sound source and the listener in the virtual space. Therefore, it may be possible to appropriately detect the amount of variation related to distance and appropriately control the updating of the transfer function according to the amount of variation.
[0035] Furthermore, the acoustic signal processing device of Example 11 may be the acoustic signal processing device of Example 10, wherein the time point defined as the reference is the time point at which the transfer function was last updated.
[0036] This can sometimes make it possible to appropriately control the updating of the transfer function in accordance with the amount of change in distance since the last update of the transfer function.
[0037] Furthermore, the acoustic signal processing device of Example 12 is an acoustic signal processing device of any of Examples 1 to 11, wherein (i) the amount of variation in the direction connecting the sound source and the listener in the virtual space, and (ii) the amount of variation in the distance between the sound source and the listener in the virtual space are each defined as the amount of variation in the relative positional relationship, and (i) the directional threshold range for the amount of variation in the direction, and (ii) the distance threshold range for the amount of variation in the distance are each defined as the threshold range, and the amount of variation in the relative positional relationship exceeds the threshold range if at least one of the conditions of the amount of variation in the direction exceeding the directional threshold range and the condition of the amount of variation in the distance exceeding the distance threshold range is satisfied, and the transfer function is updated when the amount of variation in the direction exceeds the directional threshold range, and the transfer function is updated when the amount of variation in the distance exceeds the distance threshold range.
[0038] This can make it possible to appropriately control the updating of the transfer function in response to variations in both direction and distance in the relative positional relationship between the sound source and the listener. Therefore, it may be possible to update the transfer function in a timely manner while appropriately suppressing the processing load related to the updating of the transfer function.
[0039] Furthermore, the acoustic signal processing device of Example 13 may be an acoustic signal processing device of any of Examples 1 to 12, wherein the threshold range is determined based on the auditory discrimination limit.
[0040] This can sometimes allow for appropriate control over the updating of the transfer function based on an appropriate threshold range determined by the auditory discrimination limit. Therefore, it may be possible to update the transfer function in a timely manner while appropriately suppressing the processing load related to the updating of the transfer function.
[0041] Furthermore, the acoustic signal processing device of Example 14 may be the acoustic signal processing device of Example 13, wherein the discrimination limit for determining the threshold range includes the auditory direction discrimination limit with respect to the direction of arrival of the sound.
[0042] This may in some cases make it possible to appropriately evaluate the amount of directional variation in the relative positional relationship between a sound source and a listener based on the auditory direction discrimination limit. Therefore, this may in some cases make it possible to appropriately control the update of the transfer function.
[0043] Also, the acoustic signal processing device according to Example 15 may be the acoustic signal processing device according to Example 13 or 14, wherein the discrimination limit for defining the threshold range includes the auditory loudness discrimination limit for the loudness of the sound.
[0044] This may in some cases make it possible to appropriately evaluate the amount of distance variation in the relative positional relationship between a sound source and a listener based on the auditory loudness discrimination limit. Therefore, this may in some cases make it possible to appropriately control the update of the transfer function.
[0045] Also, the acoustic signal processing device according to Example 16 may be the acoustic signal processing device according to any one of Examples 1 to 15, wherein the threshold range is adjusted based on a relaxation coefficient.
[0046] This may in some cases make it possible to adaptively adjust the threshold range for the amount of variation in the relative positional relationship between a sound source and a listener. Therefore, this may in some cases make it possible to appropriately adjust the update frequency of the transfer function.
[0047] Also, the acoustic signal processing device according to Example 17 may be the acoustic signal processing device according to Example 16, wherein the relaxation coefficient is determined based on computational resources of the acoustic signal processing device.
[0048] This may in some cases make it possible to adjust the threshold range for the amount of variation in the relative positional relationship between a sound source and a listener based on computational resources. Therefore, this may in some cases make it possible to appropriately adjust the update frequency of the transfer function based on computational resources.
[0049] Also, the acoustic signal processing device according to Example 18 may be the acoustic signal processing device according to Example 16 or 17, wherein the relaxation coefficient is defined as one of a plurality of setting parameters included in configuration information of motion in a virtual space.
[0050] This may allow the threshold range to be adjusted by the configuration parameters included in the configuration information. Therefore, it may be possible to appropriately adjust the frequency of transfer function updates by using the configuration parameters.
[0051] Furthermore, the acoustic signal processing method of Example 19 includes preparing a transfer function for calculating an acoustic signal that represents sound reaching a listener in a virtual space from a sound source in a virtual space, and calculating the acoustic signal based on the transfer function, wherein preparing the transfer function includes detecting the amount of change in the relative positional relationship between the sound source and the listener in the virtual space, and updating the transfer function if the amount of change in the relative positional relationship exceeds a threshold range.
[0052] This can make it possible to appropriately control the updating of the transfer function in accordance with the amount of change in the relative positional relationship between the sound source and the listener. Therefore, it may be possible to update the transfer function in a timely manner while appropriately suppressing the processing load related to updating the transfer function.
[0053] Furthermore, the program in Example 20 is a program that causes a computer to execute the acoustic signal processing method described in Example 19.
[0054] This may make it possible to execute the above-described acoustic signal processing method by program. Furthermore, this may allow for appropriate control of the transfer function update in response to changes in the relative positional relationship between the sound source and the listener. Therefore, it may be possible to update the transfer function in a timely manner while appropriately suppressing the processing load related to the transfer function update.
[0055] Furthermore, these comprehensive or specific embodiments may be implemented as systems, devices, methods, integrated circuits, computer programs, or non-temporary recording media such as computer-readable CD-ROMs, or as any combination of systems, devices, methods, integrated circuits, computer programs, and recording media.
[0056] (Embodiments) Embodiments will be described below with reference to the drawings. The embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, arrangement positions and connection forms of components, steps, and the order of steps shown in the following embodiments are examples and are not intended to limit the scope of the claims.
[0057] (Example of a 3D audio reproduction system) Figure 1 is a diagram showing an example of a 3D audio reproduction system. Specifically, Figure 1 shows a 3D audio reproduction system 1000, which is an example of a system to which the sound processing or decoding processing of this disclosure can be applied. 3D audio is also expressed as immersive audio. The 3D audio reproduction system 1000 includes an acoustic signal processing device 1001 and a sound presentation device 1002.
[0058] The acoustic signal processing device 1001 applies acoustic processing to the audio signal emitted by the virtual sound source to generate an acoustically processed audio signal that is presented to the listener. The audio signal is not limited to voice; it can be any audible sound. Acoustic processing is, for example, signal processing applied to an audio signal in order to reproduce one or more effects that the sound undergoes from the time it is generated at the sound source until it reaches the listener.
[0059] The acoustic signal processing device 1001 performs acoustic processing based on spatial information describing the factors that cause the above-described effects. The spatial information includes, for example, information indicating the positions of the sound source, the listener, and surrounding objects, information indicating the shape of the space, and parameters related to sound propagation. The acoustic signal processing device 1001 is, for example, a PC (Personal Computer), a smartphone, a tablet, or a game console.
[0060] The signal after acoustic processing is presented to the listener by the audio presentation device 1002. The audio presentation device 1002 is connected to the acoustic signal processing device 1001 via wireless or wired communication. The audio signal after acoustic processing generated by the acoustic signal processing device 1001 is transmitted to the audio presentation device 1002 via wireless or wired communication.
[0061] If the sound presentation device 1002 is composed of multiple devices, such as a device for the right ear and a device for the left ear, the multiple devices present sound synchronously through communication between the multiple devices, or through communication between each of the multiple devices and the acoustic signal processing device 1001. The sound presentation device 1002 is, for example, headphones, earphones, a head-mounted display worn on the listener's head, or a surround speaker system composed of multiple fixed speakers.
[0062] Furthermore, the 3D sound reproduction system 1000 may be used in combination with an image display device or a stereoscopic video display device that provides an XR (Extended Reality / Cross Reality) experience, including AR / VR, visually.
[0063] For example, the space handled by spatial information is a virtual space, and the positions of sound sources, listeners, and objects in that space are the virtual positions of virtual sound sources, virtual listeners, and virtual objects in that virtual space. This space can also be expressed as sound space. Furthermore, spatial information can also be expressed as sound space information.
[0064] Furthermore, although Figure 1 shows an example of a system configuration in which the acoustic signal processing device 1001 and the sound presentation device 1002 are separate devices, the stereophonic sound reproduction system 1000 to which the acoustic signal processing method or decoding method of this disclosure can be applied is not limited to the configuration shown in Figure 1. For example, the acoustic signal processing device 1001 may be included in the sound presentation device 1002, and the sound presentation device 1002 may perform both acoustic processing and sound presentation.
[0065] Furthermore, the acoustic signal processing device 1001 and the voice presentation device 1002 may share the task of performing the acoustic processing described in this disclosure. Alternatively, a server connected to the acoustic signal processing device 1001 or the voice presentation device 1002 via a network may perform part or all of the acoustic processing described in this disclosure.
[0066] Furthermore, the acoustic signal processing device 1001 may decode a bitstream generated by encoding at least a portion of the data between the audio signal and the spatial information used for acoustic processing, and then perform acoustic processing. Therefore, the acoustic signal processing device 1001 may also be referred to as a decoding device.
[0067] (Example of an encoding device) Figure 2A is a block diagram showing an example of the configuration of an encoding device. Specifically, Figure 2A shows the configuration of an encoding device 1100, which is an example of an encoding device according to the present disclosure.
[0068] The input data 1101 is data to be encoded, including spatial information and / or audio signals, which are input to the encoder 1102. Details of the spatial information will be explained later.
[0069] The encoder 1102 encodes the input data 1101 to generate encoded data 1103. The encoded data 1103 is, for example, a bitstream generated by the encoding process.
[0070] Memory 1104 stores the encoded data 1103. Memory 1104 may be, for example, a hard disk or an SSD (Solid State Drive), or other type of memory.
[0071] In the above description, a bitstream generated by the encoding process is given as an example of the encoded data 1103 stored in the memory 1104, but the encoded data 1103 may be data other than a bitstream. For example, the encoding device 1100 may convert the bitstream to a predetermined data format and store the converted data generated in the memory 1104. The converted data may be, for example, a file or multiplexed stream corresponding to one or more bitstreams.
[0072] Here, the file is a file having a file format such as ISOBMFF (ISO Base Media File Format). The encoded data 1103 may also be in the form of multiple packets generated by dividing the above bitstream or file.
[0073] For example, the bitstream generated by the encoder 1102 may be converted into data different from the bitstream. In this case, the encoding device 1100 may include a conversion unit (not shown) and perform the conversion process in the conversion unit, or it may perform the conversion process in a CPU (Central Processing Unit), which is an example of a processor described later.
[0074] (Example of a Decryption Device) Figure 2B is a block diagram showing an example of the configuration of a decryption device. Specifically, Figure 2B shows the configuration of a decryption device 1110, which is an example of a decryption device according to the present disclosure.
[0075] Memory 1114 stores, for example, the same data as the encoded data 1103 generated by the encoding device 1100. The stored data is read from memory 1114 and input to the decoder 1112 as input data 1113. Input data 1113 is, for example, the bitstream to be decoded. Memory 1114 may be, for example, a hard disk or SSD, or other type of memory.
[0076] The decoding device 1110 may not directly input the data read from the memory 1114 as input data 1113 to the decoder 1112, but may convert the read data and input the converted data as input data 1113 to the decoder 1112. The data before conversion may be, for example, multiplexed data containing one or more bitstreams. Here, the multiplexed data may be a file having a file format such as ISOBMFF.
[0077] Furthermore, the data before conversion may be multiple packets generated by splitting the bitstream or file as described above. Data different from the bitstream may be read from memory 1114 and converted into a bitstream. In this case, the decoding device 1110 may include a conversion unit (not shown) and perform the conversion process in the conversion unit, or it may be performed by a CPU, which is an example of a processor described later.
[0078] The decoder 1112 decodes the input data 1113 and generates an audio signal 1111 that represents the sound to be presented to the listener.
[0079] (Another Example of an Encoding Device) Figure 2C is a block diagram showing another example of the configuration of an encoding device. Specifically, Figure 2C shows the configuration of an encoding device 1120, which is another example of an encoding device of the present disclosure. In Figure 2C, the same reference numerals as in Figure 2A are used for the same components, and these components will not be described.
[0080] The encoding device 1100 stores the encoded data 1103 in the memory 1104. On the other hand, the encoding device 1120 differs from the encoding device 1100 in that it includes a transmission unit 1121 that transmits the encoded data 1103 to the outside.
[0081] The transmitting unit 1121 transmits a transmission signal 1122, generated based on encoded data 1103 or data converted from encoded data 1103 to another data format, to another device or server. The data used to generate the transmission signal 1122 is, for example, a bitstream, multiplexed data, file, or packet as described in the encoding device 1100.
[0082] (Another Example of a Decrypting Device) Figure 2D is a block diagram showing another example of the configuration of a decrypting device. Specifically, Figure 2D shows the configuration of a decrypting device 1130, which is another example of a decrypting device of the present disclosure. In Figure 2D, the same reference numerals as in Figure 2B are used for the same components, and these components will not be described.
[0083] The decoding device 1110 reads the input data 1113 from the memory 1114. On the other hand, the decoding device 1130 differs from the decoding device 1110 in that it includes a receiving unit 1131 that receives the input data 1113 from an external source.
[0084] The receiving unit 1131 receives the receiving signal 1132, acquires the received data, and outputs the input data 1113 that is input to the decoder 1112. The received data may be the same as the input data 1113 that is input to the decoder 1112, or it may be data in a different data format than the input data 1113.
[0085] If the data format of the received data differs from the data format of the input data 1113, the receiving unit 1131 may convert the received data to the input data 1113. Alternatively, a conversion unit or CPU (not shown) of the decoding device 1130 may convert the received data to the input data 1113. The received data may be, for example, a bitstream, multiplexed data, file, or packet as described in the encoding device 1120.
[0086] (Example of a Decoder) Figure 3A is a block diagram showing an example of a decoder configuration. Specifically, Figure 3A shows the configuration of decoder 1200, which is an example of decoder 1112 in Figure 2B or Figure 2D.
[0087] The input data 1113 is an encoded bitstream and includes encoded audio data, which is an encoded audio signal, and metadata used for acoustic processing.
[0088] The spatial information management unit 1201 acquires metadata contained in the input data 1113 and analyzes the metadata. The metadata includes information describing elements that act on sounds placed in the sound space. The spatial information management unit 1201 manages the spatial information used for sound processing obtained by analyzing the metadata and provides the spatial information to the rendering unit 1203.
[0089] In this disclosure, the information used for sound processing is referred to as spatial information, but other expressions may be used. For example, the information used for sound processing may be referred to as sound spatial information or scene information. Furthermore, if the information used for sound processing changes over time, the spatial information input to the rendering unit 1203 may be referred to as spatial state, sound spatial state, or scene state, etc.
[0090] Furthermore, spatial information may be managed for each sound space or each scene. For example, if multiple different rooms are each represented as virtual spaces, each of the multiple rooms may be managed as a separate, different scene. Also, even within the same space, spatial information may be managed as different scenes depending on the situation being represented.
[0091] Therefore, multiple spatial pieces of information may be managed for multiple sound spaces or multiple scenes. In managing multiple spatial pieces of information, identifiers that identify each of the multiple spatial pieces of information may be assigned to the spatial pieces of information.
[0092] Spatial information data may be included in a bitstream, which is an example of input data 1113. Alternatively, the bitstream may include a spatial information identifier, and the spatial information data may be obtained from an information source other than the bitstream. Specifically, if the bitstream contains only a spatial information identifier, in rendering, the spatial information data stored in the device's memory or an external server may be obtained as input data 1113 using the spatial information identifier.
[0093] Furthermore, the information managed by the spatial information management unit 1201 is not limited to the information contained in the bitstream. For example, the input data 1113 may include data that is not included in the bitstream, such as data indicating the characteristics and structure of the space obtained from software or a server that provides VR or AR.
[0094] Furthermore, the input data 1113 may include data indicating the characteristics and location of the listener or object. In addition, the input data 1113 may include information about the listener's location acquired by sensors equipped in a terminal including a decoding device (1110, 1130), or it may include information indicating the location of the terminal estimated based on the information acquired by the sensors.
[0095] In other words, the spatial information management unit 1201 may communicate with an external system or server to acquire spatial information and listener locations (i.e., listening positions). Alternatively, the spatial information management unit 1201 may acquire clock synchronization information from an external system and perform a process to synchronize it with the clock of the rendering unit 1203.
[0096] Furthermore, the space described above may be a virtually created space, i.e., a VR space, or it may be a real space or a virtual space corresponding to a real space, i.e., an AR space or an MR (Mixed Reality) space. The virtual space may also be expressed as a sound field or sound space. Furthermore, the position information described above may be information such as coordinate values indicating a position within space, information indicating a relative position with respect to a predetermined reference position, or information indicating the movement or acceleration of a position within space.
[0097] The audio data decoder 1202 decodes the encoded audio data contained in the input data 1113 to obtain an audio signal.
[0098] The encoded audio data acquired by the 3D audio playback system 1000 is a bitstream encoded in a predetermined format, such as MPEG-H 3D Audio (ISO / IEC 23008-3). Note that MPEG-H 3D Audio is merely one example of an encoding method that can be used to generate the encoded audio data included in the bitstream. The encoded audio data may be a bitstream encoded using another encoding method.
[0099] For example, the encoding scheme may be a lossy codec such as MP3 (MPEG-1 Audio Layer-3), AAC (Advanced Audio Coding), WMA (Windows Media Audio), AC3 (Audio Codec-3), or Vorbis. Alternatively, the encoding scheme may be a lossless codec such as ALAC (Apple Lossless Audio Codec) or FLAC (Free Lossless Audio Codec).
[0100] Alternatively, any encoding scheme other than those described above may be used. For example, PCM (pulse code modulation) data may be a type of encoded audio data. In this case, the decoding process may be, for example, a process of converting an N-bit binary number into a number format that the rendering unit 1203 can process (e.g., floating-point format) if the number of quantization bits of the PCM data is N.
[0101] The rendering unit 1203 acquires the audio signal and spatial information, applies acoustic processing to the audio signal using the spatial information, and outputs the audio signal after acoustic processing (audio signal 1111).
[0102] Before rendering begins, the spatial information management unit 1201 reads the metadata of the input signal, detects rendering items such as objects and sounds defined in the spatial information, and transmits them to the rendering unit 1203. After rendering begins, the spatial information management unit 1201 grasps the changes in spatial information and listener positions over time, and updates and manages the spatial information. Then, the spatial information management unit 1201 transmits the updated spatial information to the rendering unit 1203.
[0103] The rendering unit 1203 generates and outputs an audio signal with added acoustic processing based on the audio signal included in the input data 1113 and the spatial information received from the spatial information management unit 1201.
[0104] The spatial information update process and the audio signal output process with added acoustic processing may be executed in the same thread. Alternatively, the spatial information management unit 1201 and the rendering unit 1203 may each allocate processing to independent threads. If the spatial information management unit 1201 and the rendering unit 1203 execute the spatial information update process and the audio signal output process with added acoustic processing in different threads, the thread startup frequency may be set individually, or the processing may be executed in parallel.
[0105] When the spatial information management unit 1201 and the rendering unit 1203 execute processing in different, independent threads, it is possible to preferentially allocate computing resources to the rendering unit 1203. This makes it possible to safely perform sound output processing where even the slightest delay is unacceptable, for example, where a delay of one sample (0.02 msec) would cause a popping noise.
[0106] In this case, the allocation of computing resources to the spatial information management unit 1201 is limited. However, since updating spatial information is a low-frequency process compared to the output processing of audio signals (for example, updating the direction of the listener's face), it does not necessarily have to be instantaneous like the output processing of audio signals. Therefore, even if the allocation of computing resources is limited, it does not have a significant impact on the acoustic quality.
[0107] Spatial information updates may be performed periodically at predetermined times or intervals, or when predetermined conditions are met. Furthermore, spatial information updates may be performed manually by the listener or the sound space manager, or triggered by changes in an external system.
[0108] For example, spatial information may be updated if a listener operates a controller, causing their avatar's position to instantly warp or the time to instantly advance or rewind. Alternatively, spatial information may be updated if the virtual space administrator suddenly alters the environment of the space. In these cases, a thread for updating the spatial information managed by the spatial information management unit 1201 may be started not only periodically but also as a one-off interrupt.
[0109] The information update thread, which performs spatial information update processing, is responsible for tasks such as updating the position or orientation of the listener's avatar placed in the virtual space based on the position or orientation of the VR goggles worn by the listener, and updating the positions of objects moving within the virtual space. Such processing is handled within a processing thread that is activated at a relatively low frequency of several tens of Hz.
[0110] The process of updating information indicating the properties of the direct sound may be performed in such an infrequently occurring processing thread. This is because the frequency of changes in the properties of the direct sound is lower than the frequency of changes in audio processing frames for audio output. This makes it possible to relatively reduce the computational load of the process.
[0111] Furthermore, updating information at an unnecessarily rapid pace carries the risk of generating pulsative noise. This risk can be avoided by updating information at a lower frequency.
[0112] Figure 3B is a block diagram showing another example of the decoder configuration. Specifically, Figure 3B shows the configuration of decoder 1210, which is another example of decoder 1112 in Figure 2B or Figure 2D.
[0113] Figure 3B differs from Figure 3A in that the input data 1113 includes an unencoded audio signal rather than encoded audio data. The input data 1113 includes a bitstream containing metadata and an audio signal.
[0114] The spatial information management unit 1211 is the same as the spatial information management unit 1201 in Figure 3A, so its explanation is omitted. Note that the spatial information management unit 1201 in the following explanation may be replaced with the spatial information management unit 1211.
[0115] The rendering unit 1213 is the same as the rendering unit 1203 in Figure 3A, so its description is omitted. Note that the rendering unit 1203 in the following description may be replaced with the rendering unit 1213.
[0116] The decoders 1112, 1200, and 1210 may also be described as acoustic processing units that perform acoustic processing. Furthermore, the decoding devices 1110 and 1130 may also be acoustic signal processing units 1001, and may be described as acoustic processing units.
[0117] (Physical Configuration of the Acoustic Signal Processing Device) Figure 4 shows an example of the physical configuration of the acoustic signal processing device 1001. Note that the acoustic signal processing device 1001 in Figure 4 may be the decoding device 1110 in Figure 2B or the decoding device 1130 in Figure 2D. The multiple components shown in Figure 2B or Figure 2D may be implemented by the multiple components shown in Figure 4. Furthermore, some of the configuration described here may be provided in the sound presentation device 1002.
[0118] The acoustic signal processing device 1001 shown in Figure 4 comprises a processor 1402, a memory 1404, a communication interface (IF) 1403, a sensor 1405, and a speaker 1401.
[0119] The processor 1402 is, for example, a CPU, a DSP (Digital Signal Processor), or a GPU (Graphics Processing Unit). The CPU, DSP, or GPU may perform the acoustic processing or decoding processing of this disclosure by executing a program stored in memory 1404. Alternatively, the processor 1402 may be, for example, a circuit that performs information processing. The processor 1402 may also be a dedicated circuit that performs signal processing on an audio signal, including the acoustic processing of this disclosure.
[0120] The memory 1404 is composed of, for example, RAM (Random Access Memory) or ROM (Read Only Memory). The memory 1404 may also include a magnetic recording medium such as a hard disk or a semiconductor memory such as an SSD. The memory 1404 may also be an internal memory built into the CPU or GPU. The memory 1404 may also store spatial information managed by the spatial information management unit 1201. It may also store threshold data as described later.
[0121] In memory 1404, the value is undefined when the acoustic signal processing device 1001 is started, so a desired value is set during initialization. Although not shown in Figures 2B, 2D, 3A, and 3B, for example, the decoder 1200 or 1210 may have such an initialization unit built in.
[0122] The communication interface 1403 is a communication module that supports communication methods such as Bluetooth® or WIGIG®. The acoustic signal processing device 1001 communicates with other communication devices via the communication interface 1403 and obtains the bitstream to be decoded. The obtained bitstream is stored in the memory 1404, for example.
[0123] The communication IF1403 consists, for example, of a signal processing circuit and an antenna corresponding to the communication method. The communication method is not limited to Bluetooth® and WIGIG®, but may also be LTE (Long Term Evolution), NR (New Radio), or Wi-Fi®, etc.
[0124] Furthermore, the communication method is not limited to the wireless communication methods described above. The communication method may also be a wired communication method such as Ethernet®, USB (Universal Serial Bus), or HDMI® (High-Definition Multimedia Interface).
[0125] Sensor 1405 performs sensing to estimate the listener's position and orientation. Specifically, sensor 1405 estimates the listener's position and / or orientation based on the detection result of one or more of the position, orientation, movement, velocity, angular velocity, and acceleration of part or whole of the body, and generates position / or orientation information indicating the listener's position and / or orientation.
[0126] Furthermore, an external device to the acoustic signal processing device 1001 may be equipped with a sensor 1405. The part of the body may be the listener's head, etc. The position / orientation information may be information indicating the listener's position and / or orientation in real space, or it may be information indicating the displacement of the listener's position and / or orientation relative to the listener's position and / or orientation at a predetermined time. In addition, the position / orientation information may be information indicating the relative position and / or orientation with respect to the stereophonic sound reproduction system 1000 or an external device equipped with a sensor 1405.
[0127] Sensor 1405 is, for example, an imaging device such as a camera or a distance measuring device such as LiDAR (Laser Imaging Detection and Ranging). Sensor 1405 may detect the movement of the listener's head by imaging the movement of the listener's head and processing the captured image. Alternatively, a device that performs position estimation using wireless communication in an arbitrary frequency band such as millimeter waves may be used as sensor 1405.
[0128] Furthermore, the acoustic signal processing device 1001 may acquire position information from an external device equipped with a sensor 1405 via a communication IF 1403. In this case, the acoustic signal processing device 1001 does not need to include the sensor 1405. Here, the external device is, for example, the audio presentation device 1002 described in Figure 1, or a stereoscopic image playback device worn on the listener's head. In this case, the sensor 1405 is configured by combining various sensors, such as a gyro sensor and an acceleration sensor.
[0129] The sensor 1405 may, for example, detect the angular velocity of rotation with at least one of the three mutually orthogonal axes in the sound space as the axis of rotation, or it may detect the acceleration of displacement with at least one of the three axes as the direction of displacement.
[0130] The sensor 1405 may, for example, detect the amount of movement of the listener's head, such as a rotational amount with at least one of the three mutually orthogonal axes in the sound space as the axis of rotation, or it may detect a displacement amount with at least one of the three axes as the direction of displacement. Specifically, the sensor 1405 detects the listener's position in 6DoF (x, y, z) and angle (yaw, pitch, roll). The sensor 1405 is composed of a combination of various sensors used for motion detection, such as a gyroscope and an accelerometer.
[0131] The sensor 1405 may be implemented by a camera or GPS (Global Positioning System) receiver for detecting the listener's position. Position information obtained by performing self-position estimation using LiDAR or the like as the sensor 1405 may also be used. For example, if the 3D sound reproduction system 1000 is implemented by a smartphone, the sensor 1405 will be built into the smartphone.
[0132] Furthermore, the sensor 1405 may include a temperature sensor such as a thermocouple for detecting the temperature of the acoustic signal processing device 1001. The sensor 1405 may also include a sensor for detecting the remaining charge of the battery provided by the acoustic signal processing device 1001, or a battery connected to the acoustic signal processing device 1001.
[0133] The speaker 1401 includes, for example, a diaphragm, a drive mechanism such as a magnet or voice coil, and an amplifier, and presents the processed audio signal as sound to the listener. The speaker 1401 operates the drive mechanism in response to the audio signal (more specifically, a waveform signal showing the waveform of the sound) amplified via the amplifier, and the drive mechanism vibrates the diaphragm. In this way, the diaphragm vibrating in response to the audio signal generates sound waves, which propagate through the air and are transmitted to the listener's ears, allowing the listener to perceive the sound.
[0134] Although an example was given in which the acoustic signal processing device 1001 is equipped with a speaker 1401 and presents the processed audio signal via the speaker 1401, the means for presenting the audio signal is not limited to the above configuration.
[0135] For example, the processed audio signal may be output to an external audio presentation device 1002 connected via a communication module. The communication performed by the communication module may be wired or wireless. Another example is that the audio signal processing device 1001 may have a terminal for outputting an analog audio signal, and the audio signal may be presented through an earphone or the like by connecting a cable to the terminal.
[0136] In the above case, the audio presentation device 1002 may be headphones, earphones, a head-mounted display, a neck speaker, or a wearable speaker, etc., that are attached to the listener's head or a part of their body. Alternatively, the audio presentation device 1002 may be a surround speaker, etc., composed of multiple fixed speakers. The audio presentation device 1002 may also reproduce an audio signal.
[0137] (Physical configuration of the encoding device) Figure 5 shows an example of the physical configuration of the encoding device. The encoding device 1500 in Figure 5 may be the encoding device 1100 in Figure 2A or the encoding device 1120 in Figure 2C, and the multiple components shown in Figure 2A or Figure 2C may be implemented by the multiple components shown in Figure 5.
[0138] The encoding device 1500 in Figure 5 comprises a processor 1501, a memory 1503, and a communication IF 1502.
[0139] The processor 1501 is, for example, a CPU, DSP, or GPU. The CPU, DSP, or GPU may perform the encoding process of this disclosure by executing a program stored in the memory 1503. Alternatively, the processor 1501 may be, for example, a circuit that performs information processing. The processor 1501 may also be a dedicated circuit that performs signal processing on an audio signal, including the encoding process of this disclosure.
[0140] The memory 1503 is composed of, for example, RAM or ROM. The memory 1503 may also include a magnetic recording medium such as a hard disk or a semiconductor memory such as an SSD. Alternatively, the memory 1503 may be internal memory incorporated into the CPU or GPU.
[0141] The communication IF 1502 is a communication module that supports communication methods such as Bluetooth® or WIGIG®. The encoding device 1500 communicates with other communication devices via the communication IF 1502, for example, and transmits the encoded bitstream.
[0142] The communication IF1502 consists, for example, of a signal processing circuit and an antenna corresponding to the communication method. The communication method is not limited to Bluetooth® and WIGIG®, but may also be LTE, NR, or Wi-Fi®, etc. Furthermore, the communication method is not limited to wireless communication. The communication method may also be a wired communication method such as Ethernet®, USB, or HDMI®.
[0143] (Rendering Unit Configuration) Figure 6 is a block diagram showing an example of the rendering unit configuration. Specifically, Figure 6 shows a detailed example of the rendering unit 1300 corresponding to the rendering units 1203 and 1213 in Figures 3A and 3B.
[0144] The rendering unit 1300 consists of an analysis unit 1301, a selection unit 1302, and a playback unit 1303, and adds acoustic processing to the sound data included in the input signal and outputs it.
[0145] The input signal may consist, for example, of spatial information, sensor information, and sound data. The input signal may also include a bitstream consisting of sound data and metadata (control information), in which case the metadata may include spatial information.
[0146] Spatial information is information about the sound space (three-dimensional sound field) created by the 3D sound reproduction system 1000, and consists of information about objects included in the sound space and information about the listener. Objects include sound source objects that emit sound and act as sound sources, and non-sounding objects that do not emit sound. Sound source objects can also be simply referred to as sound sources.
[0147] Non-sounding objects act as obstacles that reflect sound emitted by a sound source object, but sound source objects can also act as obstacles that reflect sound emitted by another sound source object. Obstacle objects may also be referred to as reflection objects.
[0148] Information that is assigned to both sound source objects and non-sounding objects includes positional information, shape information, and the rate of sound attenuation when the object reflects sound.
[0149] Position information is represented by coordinate values on three axes in Euclidean space, for example, the X, Y, and Z axes, but it does not necessarily have to be three-dimensional information. For example, position information may be two-dimensional information represented by coordinate values on two axes, the X and Y axes. The position information of an object is determined by the representative position of the shape represented by a mesh or voxel.
[0150] Shape information may include information about the surface material.
[0151] The attenuation rate may be expressed as a real number between 0 and 1, or as a negative decibel value. In real space, sound volume is not amplified by reflection, so the attenuation rate is set to a negative decibel value. However, for example, to create an eerie atmosphere in an unreal space, an attenuation rate of 1 or greater, i.e., a positive decibel value, may be deliberately set.
[0152] Furthermore, the attenuation rate may be set to a different value for each frequency band that constitutes multiple frequency bands, or a value may be set independently for each frequency band. Also, if the attenuation rate is set for each type of material on the object surface, the corresponding attenuation rate value may be used based on information about the surface material.
[0153] Furthermore, the spatial information may include information indicating whether or not the object belongs to a living organism, and information indicating whether or not the object is a moving object. If the object is a moving object, the position indicated by the position information may move over time. In this case, information about the changed position or the amount of change is transmitted to the rendering unit 1300.
[0154] Information regarding a sound source object includes, in addition to information commonly assigned to both sound source and non-sounding objects, sound data and information necessary for radiating the sound data into the sound space. The sound data is data that indicates information such as the frequency and intensity of the sound, and represents the sound perceived by the listener.
[0155] The audio data is typically a PCM signal, but it may also be data compressed using an encoding method such as MP3. In that case, the signal needs to be decoded at least before it arrives at the playback unit 1303, so the rendering unit 1300 may include a decoding unit (not shown). Alternatively, the signal may be decoded by the audio data decoder 1202.
[0156] A single sound source object may have one sound data set, or it may have multiple sound data set. Furthermore, identification information may be attached to each sound data, and information about the sound source object may include this identification information.
[0157] The information necessary to radiate sound data into sound space may include, for example, information on the reference volume used as a reference in the playback of sound data, information indicating the properties (also called characteristics) of the sound data, information on the position of the sound source object, and information on the orientation of the sound source object (i.e., information on the directivity of the sound emitted by the sound source object).
[0158] The reference volume information may be, for example, the effective value of the amplitude of the sound data at the sound source position when the sound data is radiated into the sound space, and may be expressed as a floating-point value in decibels (dB).
[0159] For example, if the reference volume is 0 dB, it may indicate that the sound is radiated into the sound space from the position indicated by the information regarding the location of the sound source object, at the same volume without increasing or decreasing the volume of the signal level indicated by the sound data. Alternatively, if the reference volume is -6 dB, it may indicate that the volume of the signal level indicated by the sound data is reduced by approximately half, and the sound is radiated into the sound space from the position indicated by the information regarding the location of the sound source object.
[0160] Reference volume information may be assigned to each audio data file, or it may be assigned to multiple audio data files collectively.
[0161] Information indicating the nature of sound data may, for example, be information regarding the volume of the sound source, and may also be information indicating the time-series fluctuations of the volume of the sound source.
[0162] For example, if the sound space is a virtual conference room and the sound source is a speaker, the volume will transition intermittently over short periods of time. That is, there will be alternating periods of sound and silence. If the sound space is a concert hall and the sound source is a performer, the volume will be maintained for a certain duration. If the sound space is a battlefield and the sound source is an explosive, the volume of the explosion will be loud for just a moment, and then remain silent or quiet.
[0163] Thus, the volume information of a sound source may include not only information about the loudness of the sound, but also information about the transition of the loudness. Such information may be used as information that indicates the properties of the sound data.
[0164] Transition information may be represented by data showing the frequency characteristics in time series. Transition information may be represented by data showing the duration of the sounded interval. Transition information may be represented by data showing the time series of the duration of the sounded interval and the duration of the silent interval. Transition information may be represented by data that enumerates multiple pairs in time series the duration during which the amplitude of the sound signal can be considered stationary (considered to be approximately constant) and the amplitude value of the signal during that period.
[0165] The transition information may be represented by data showing the duration for which the frequency characteristics of the sound signal can be considered stationary. The transition information may also be represented by data listing multiple pairs of durations for which the frequency characteristics of the sound signal can be considered stationary, and the frequency characteristics during those durations, in a time series. The transition information may also be represented, for example, in the form of data showing the general shape of a spectrogram.
[0166] Furthermore, the volume used as the reference for the frequency characteristics described above may be the reference volume described above. The reference volume information and the information indicating the properties of the sound data may be used in the calculation process of the volume of direct sound or reflected sound perceived by the listener, or they may be used in the selection process of whether or not to make the sound perceived by the listener.
[0167] Note that the reflected sound in this embodiment is an example of indirect sound. Indirect sound may be reflected sound, diffracted sound, or the like. In this embodiment, a reflected sound, which is an example of indirect sound, is used for explanation, but the same processing is performed even if indirect sound is used instead of reflected sound.
[0168] Information regarding the orientation of a sound source object (orientation information) is typically represented by yaw, pitch, and roll. Alternatively, the roll rotation may be omitted, and the orientation information of the sound source object may be represented by azimuth (yaw) and elevation (pitch). The orientation information of the sound source object may change over time, and if it changes, it is transmitted to the rendering unit 1300.
[0169] Information about the listener is information about the listener's position and orientation in the sound space. Position information is represented by the position along the XYZ axes in Euclidean space, but it does not necessarily have to be three-dimensional information; it may also be two-dimensional information. Information about the listener's orientation is typically represented by yaw, pitch, and roll. Alternatively, the rotation of roll may be omitted, and the listener's orientation information may be represented by azimuth (yaw) and elevation (pitch).
[0170] The listener's location and orientation information may change over time, and if it does, it is transmitted to the rendering unit 1300.
[0171] The sensor information includes the amount of rotation or displacement detected by the sensor 1405 worn by the listener, as well as the listener's position and orientation. The sensor information is transmitted to the rendering unit 1300, which updates the listener's position and orientation information based on the sensor information. The sensor information may also include position information obtained by a mobile terminal performing self-position estimation using GPS, a camera, or LiDAR, for example.
[0172] Furthermore, information acquired from an external source via a communication module, rather than from the sensor 1405, may be detected as sensor information. Information indicating the temperature of the acoustic signal processing device 1001 and information indicating the remaining battery level may be acquired from the sensor 1405. In addition, the computing resources (CPU capacity, memory resources, or PC performance, etc.) of the acoustic signal processing device 1001 or the voice presentation device 1002 may be acquired in real time.
[0173] The analysis unit 1301 analyzes the audio signal included in the input signal and the spatial information received from the spatial information management unit 1201, and calculates the information necessary for generating direct sound and reflected sound in the playback unit 1303, as well as the information necessary for selecting whether or not to generate reflected sound.
[0174] The information necessary for generating direct and reflected sound includes, for example, values relating to the path, time, and volume of the direct and reflected sound to the listening position, respectively. The values relating to the path, time, and volume of the direct and reflected sound to the listening position are, for example, values indicating the path, time, and volume of the direct and reflected sound to the listening position, respectively.
[0175] The information necessary for selecting the reflected sound to be output is information that shows the relationship between the direct sound and the reflected sound, such as a value relating to the time difference between the direct sound and the reflected sound, and a value relating to the volume ratio of the direct sound and the reflected sound at the listening position. The value relating to the time difference between the direct sound and the reflected sound, and the value relating to the volume ratio of the direct sound and the reflected sound at the listening position are, for example, a value indicating the time difference between the direct sound and the reflected sound, and a value indicating the volume ratio of the direct sound and the reflected sound at the listening position, respectively.
[0176] Furthermore, when volume is expressed in decibels on a logarithmic scale (i.e., when volume is expressed in the decibel domain), it goes without saying that the volume ratio of two signals is expressed as the difference in decibel values. Specifically, the volume ratio of two signals may be the difference when the amplitude values of each signal are expressed in the decibel domain. This value may be calculated based on energy values or power values, etc. Also, in the decibel domain, this difference may be called the gain difference or simply the gain difference.
[0177] In other words, the volume ratio in this disclosure is substantially the ratio of the signal amplitudes, and may be expressed as Sound volume ratio, Volume ratio, Amplitude ratio, Sound level ratio, Sound intensity ratio, or Gain ratio, etc. Furthermore, when the unit of volume is decibels, it goes without saying that the volume ratio in this disclosure can be rephrased as the volume difference.
[0178] In this disclosure, "volume ratio" typically means the gain difference when the volumes of two sounds are expressed in decibels, and in the examples of the embodiments, the threshold data is also typically defined by a gain difference expressed in the decibel domain. However, the volume ratio is not limited to a gain difference in the decibel domain. When a volume ratio expressed in a domain other than the decibel domain is used, the threshold data defined in the decibel domain may be converted to the unit of the calculated volume ratio and used. Alternatively, threshold data defined in each unit may be stored in memory in advance.
[0179] In other words, it is clear that the algorithm in this disclosure can be applied to solving the problem in this disclosure even if a ratio of, for example, an energy value or a power value is used instead of a volume ratio.
[0180] The time difference between the arrival of a direct sound and a reflected sound is, for example, the time difference between the arrival time of the direct sound and the arrival time of the reflected sound. For simplicity, the time difference between the arrival of a direct sound and a reflected sound is sometimes written as the time difference between the direct sound and the reflected sound. The time difference between the direct sound and the reflected sound may also be the time difference between when the direct sound and the reflected sound arrive at the listening position, the time difference between when the direct sound finishes being produced and when the reflected sound arrives at the listening position, or the time difference between when the direct sound finishes being produced and when the reflected sound arrives at the listening position.
[0181] The selection unit 1302 uses the information calculated by the analysis unit 1301 and threshold data to select whether or not the playback unit 1303 should generate a reflected sound. In other words, the selection unit 1302 determines whether or not to select a reflected sound as the reflected sound to be generated. To put it another way, the selection unit 1302 selects which of the multiple reflected sounds the playback unit 1303 should generate.
[0182] Threshold data can be represented, for example, as a graph with the time difference between direct sound and reflected sound on the horizontal axis and the volume ratio between direct sound and reflected sound on the vertical axis, where the threshold is the boundary between whether the reflected sound is perceived or not. Threshold data may be expressed as an approximation formula with the time difference between direct sound and reflected sound as a variable, or as an array with the time difference between direct sound and reflected sound as an index and corresponding thresholds.
[0183] The selection unit 1302 selects to generate a reflected sound if, for example, the volume ratio between the volume of the direct sound at arrival and the volume of the reflected sound at arrival, based on the time difference between the arrival time of the direct sound and the arrival time of the reflected sound, is greater than a threshold set by referring to threshold data. Note that the volume at arrival refers to the volume when the sound arrives at the listening position.
[0184] The time difference between the arrival time of the direct sound and the arrival time of the reflected sound is, in other words, the difference in the time it takes for the direct sound and the reflected sound to arrive at the listening position, respectively. Alternatively, the time difference between the end of the direct sound's production and the arrival time of the reflected sound at the listening position may be used as the time difference between the direct sound and the reflected sound. In that case, threshold data different from the threshold data determined using the time difference between the arrival time of the direct sound and the arrival time of the reflected sound may be used, or common threshold data may be used.
[0185] The threshold data may be obtained from the memory 1404 of the acoustic signal processing device 1001, or it may be obtained from an external storage device via a communication module.
[0186] The playback unit 1303 combines the direct sound audio signal with the reflected sound audio signal that the selection unit 1302 has selected to generate.
[0187] Specifically, the playback unit 1303 processes the input audio signal to generate direct sound based on the direct sound arrival time and volume information at the time of direct sound arrival calculated by the analysis unit 1301. The playback unit 1303 also processes the input audio signal to generate reflected sound based on the reflected sound arrival time and volume information for the reflected sound selected by the selection unit 1302. Finally, the playback unit 1303 synthesizes the generated direct sound and reflected sound and outputs it.
[0188] The playback unit 1303 may include a dispersion filter. The dispersion filter is a filter that disperses (varies) the phase of sound and can simulate sound diffusion. The dispersion filter may have a signal amplification factor set in advance, or the playback unit 1303 may set the signal amplification factor.
[0189] Furthermore, the playback unit 1303 may perform volume compensation processing based on the reflected sound that was not selected.
[0190] (Example of Rendering Unit Operation) Figure 7 is a flowchart showing an example of the operation of the acoustic signal processing unit 1001. Figure 7 mainly shows the processing performed by the rendering unit 1300 of the acoustic signal processing unit 1001.
[0191] In the input signal analysis process (S101 in Figure 7), the analysis unit 1301 analyzes the input signal input to the acoustic signal processing device 1001 to detect direct sound and reflected sound that may occur in the sound space. The reflected sound detected here is a candidate for reflected sound to be selected by the selection unit 1302 as the reflected sound to be ultimately generated by the playback unit 1303. The analysis unit 1301 also analyzes the input signal to calculate the information necessary for generating direct sound and reflected sound, and the information necessary for selecting the reflected sound to be generated.
[0192] First, the characteristics of both the direct sound and the reflected sound are calculated. Specifically, the arrival time and volume of each sound upon arrival at the listener are calculated. If multiple objects exist in the sound space as reflective objects, the characteristics of the reflected sound are calculated for each of the multiple objects.
[0193] The direct sound arrival time (td) is calculated based on the direct sound arrival path (pd). The direct sound arrival path (pd) is the path connecting the location information S (xs, ys, zs) of the sound source object and the location information A1 (xa, ya, za) of the listener. The direct sound arrival time (td) is obtained by dividing the length of the path connecting the location information S (xs, ys, zs) and the location information A1 (xa, ya, za) by the speed of sound (approximately 340 m / s).
[0194] For example, the path length (X) can be calculated as ((xs - xa)^2 + (ys - ya)^2 + (zs - za)^2)^0.5. Sound volume decreases inversely with distance. Therefore, if the sound volume at the location information S(xs, ys, zs) of the sound source object is N and the unit distance is U, the sound volume at direct sound arrival (ld) can be calculated as ld = N * U / X.
[0195] The volume N at the sound source location may be the reference volume as explained earlier.
[0196] The arrival time of the reflected sound (tr) is calculated based on the reflected sound arrival path (pr). The reflected sound arrival path (pr) is the path connecting the position of the reflected sound image to the position information A1 (xa, ya, za).
[0197] Furthermore, the position of the sound image of reflected sound may be derived using methods such as the "image method" or "ray tracing method," or any other method for deriving the sound image position may be used. The image method is a technique that simulates a sound image by assuming that a mirror image exists on the wall surface of a room at a position symmetrical to the sound source relative to the wall surface, and that sound waves are radiated from the position of that mirror image. The ray tracing method is a technique that simulates an image (sound image) observed at a certain point by tracing waves that propagate linearly, such as light rays or sound rays.
[0198] For example, a sound image of the reflected sound is formed symmetrically with respect to the sound source position, separated by a wall. By determining the position of the reflected sound image along the x, y, and z axes, the arrival time of the reflected sound can be determined in the same way as calculating the arrival time of the direct sound.
[0199] The arrival time of the reflected sound (tr) is obtained by dividing the length (Y) of the path connecting the position of the reflected sound image and the position information A1 (xa, ya, za) by the speed of sound (approximately 340 m / s). Sound volume decreases inversely with distance. Therefore, if the sound volume at the sound source is N, the unit distance is U, and the rate of sound volume attenuation in reflection is G, the sound volume at the arrival of the reflected sound (lr) can be calculated as lr = N * G * U / Y.
[0200] As explained earlier, the attenuation rate G may be expressed as a real number between 0 and 1, or as a negative decibel value. In this case, the overall volume of the signal is attenuated by the amount of G. The attenuation rate may also be set for each frequency band that makes up multiple frequency bands. In this case, the analysis unit 1301 multiplies each frequency component of the signal by the specified attenuation rate. Furthermore, in order to reduce the amount of computation, the analysis unit 1301 may use a representative value or average value of multiple attenuation rates for multiple frequency bands as the overall attenuation rate and attenuate the overall volume of the signal by that amount.
[0201] Next, in the reflected sound selection process (S102 in Figure 7), the selection unit 1302 selects the reflected sound generated by the playback unit 1303 based on the information calculated by the analysis unit 1301. In other words, the selection unit 1302 selects whether or not to generate a reflected sound. The selection of whether or not to generate a reflected sound may be a selection of whether or not to play the direct sound and the reflected sound separately, or a selection of whether or not to combine the direct sound and the reflected sound.
[0202] Next, in the direct sound and reflected sound generation process (S103 in Figure 7), the playback unit 1303 generates and combines the audio signal of the direct sound and the audio signal of the reflected sound selected as the target reflected sound by the selection unit 1302. The playback unit 1303 may also perform volume compensation processing based on the reflected sounds that were not selected.
[0203] The direct sound audio signal is generated by applying the direct sound arrival time (td) and the direct sound arrival volume (ld), calculated by the analysis unit 1301, to the sound data of the sound source object included in the input information. Specifically, the sound data is delayed by the direct sound arrival time (td) and then multiplied by the direct sound arrival volume (ld). The process of delaying the sound data is a process of moving the position of the sound data forward or backward on the time axis. For example, a process for delaying sound data without degrading sound quality, such as the one disclosed in Patent Document 2, may be applied.
[0204] The reflected sound signal is generated, similar to the direct sound, by applying the reflected sound arrival time (tr) and the reflected sound arrival volume (lr), calculated by the analysis unit 1301, to the sound data of the sound source object.
[0205] However, the volume of the reflected sound upon arrival (lr) in the generation of reflected sound differs from the volume of the direct sound upon arrival (ld) in that the attenuation rate G for volume in reflection is applied to the value. G may be an attenuation rate applied uniformly to the entire frequency band. Alternatively, the reflectance rate may be defined for each predetermined frequency band in order to reflect the bias in frequency components caused by reflection. In that case, the process of applying the volume of the reflected sound upon arrival (lr) may be carried out as a frequency equalizer process, which is a process of multiplying by the attenuation rate for each band.
[0206] (Examples of direct and reflected sound) For example, direct sound is sound that has not been reflected by a reflective object, and reflected sound is sound that has been reflected by a reflective object. Direct sound may be sound that arrives at the listener without being reflected by a reflective object from the sound source, and reflected sound may be sound that arrives at the listener after being reflected by a reflective object from the sound source.
[0207] Furthermore, direct sound and reflected sound are not limited to sounds that have reached the listener, but may also be sounds that have not yet reached the listener. For example, direct sound may be the sound emitted from the sound source, or in other words, the sound of the sound source itself.
[0208] For example, if an obstacle object is present between the sound source object and the listener, direct sound may not reach the listener. In this case, the sound emitted from the sound source object, passing through the obstacle object and reaching the listener may be considered direct sound. The sound emitted from the sound source object, diffracted by the obstacle object and reaching the listener may be considered reflected sound.
[0209] Furthermore, the two sounds compared in the selection process are not limited to direct and reflected sounds based on a single sound source. For example, sound selection may be performed by comparing two reflected sounds based on a single sound source. In this case, the direct sound in this disclosure may be interpreted as the sound that reaches the listener first, and the reflected sound in this disclosure may be interpreted as the sound that reaches the listener later.
[0210] (Example of pipeline processing in the rendering unit) The processing performed in the analysis unit 1301, selection unit 1302, and playback unit 1303 described above may be performed as pipeline processing.
[0211] Figure 8 is a block diagram showing an example configuration for the rendering unit 1300 to perform pipeline processing.
[0212] The rendering unit 1300 in Figure 8 includes a reverberation processing unit 1311, an early reflection processing unit 1312, a distance attenuation processing unit 1313, a selection unit 1314, a generation unit 1315, and a binaural processing unit 1316. These multiple components may be composed of the multiple components of the rendering unit 1300 shown in Figure 6, or at least some of the multiple components of the acoustic signal processing device 1001 shown in Figure 4.
[0213] Pipelining refers to the process of dividing the process of applying sound effects into multiple steps and executing each step sequentially. Each of these steps may involve, for example, signal processing of the audio signal or the generation of parameters used in the signal processing.
[0214] The rendering unit 1300 may perform reverberation processing, early reflection processing, distance attenuation processing, and binaural processing as pipeline processing. However, these processing are just examples, and pipeline processing may include other processing or may omit some of the processing. For example, pipeline processing may include diffraction processing and occlusion processing. Also, for example, reverberation processing may be omitted if it is not necessary.
[0215] Furthermore, each process may be referred to as a stage. Also, the resulting audio signals, such as reflected sound, may be referred to as rendering items. The number of stages in a pipeline process, and their order, are not limited to the example shown in Figure 8.
[0216] Here, the parameters used in the selection process (arrival paths, arrival times, and volume ratios for direct and reflected sounds) are calculated in one of several stages for generating rendering items. In other words, the parameters used for selecting reflected sounds are calculated as part of the pipeline process for generating rendering items. Note that not all stages are performed in the rendering unit 1300. For example, some stages may be omitted or performed outside of the rendering unit 1300.
[0217] This section describes reverberation processing, early reflection processing, distance attenuation processing, selection processing, generation processing, and binaural processing, which may be included as stages in pipeline processing. In each stage, metadata contained in the input signal may be analyzed to calculate parameters used for generating reflected sound.
[0218] In reverberation processing, the reverberation processing unit 1311 generates an audio signal indicating reverberation, or parameters used to generate the audio signal. Reverberation is the sound that arrives to the listener as reverberation after the direct sound. For example, reverberation is the sound that arrives to the listener relatively late (for example, around 100 ms from the arrival of the direct sound) after the initial reflections described later have arrived to the listener, and after undergoing many more reflections (for example, several dozen times) than the initial reflections.
[0219] The reverberation processing unit 1311 refers to the audio signal and spatial information included in the input signal and calculates the reverberation sound using a predetermined function that has been prepared in advance as a function for generating reverberation sound.
[0220] The reverberation processing unit 1311 may generate reverberation sound by applying a known reverberation generation method to the audio signal included in the input signal. An example of a known reverberation generation method is the Schroeder method, but known reverberation generation methods are not limited to the Schroeder method. Furthermore, in applying a known reverberation generation method, the reverberation processing unit 1311 uses the shape and acoustic characteristics of the sound reproduction space indicated by the spatial information. This allows the reverberation processing unit 1311 to calculate parameters for generating reverberation sound.
[0221] In the initial reflection processing, the initial reflection processing unit 1312 calculates parameters for generating initial reflections based on spatial information. Initial reflections are reflected sounds that arrive at the listener relatively early after the direct sound from the sound source object has arrived (for example, within a few tens of milliseconds from the arrival of the direct sound), after undergoing one or more reflections.
[0222] The initial reflection processing unit 1312 calculates the path of the reflected sound that travels from the sound source object, reflects off reflection objects, and reaches the listener, for example, by referring to the audio signal and metadata. For example, the shape of the three-dimensional sound field (space), the size of the three-dimensional sound field, the position of reflection objects such as structures, and the reflectivity of the reflection objects may be used in the path calculation.
[0223] Furthermore, the initial reflection processing unit 1312 may also calculate the path of the direct sound. This path information may be used by the initial reflection processing unit 1312 as a parameter for generating the initial reflected sound, or it may be used by the selection unit 1314 as a parameter for selecting the reflected sound.
[0224] In distance attenuation processing, the distance attenuation processing unit 1313 calculates the volume of direct sound and reflected sound reaching the listener based on the length of the paths of the direct sound and reflected sound. The volume of direct sound and reflected sound reaching the listener is attenuated in proportion to the distance of the path to the listener (inversely proportional to the distance) relative to the volume of the sound source. Therefore, the distance attenuation processing unit 1313 can calculate the volume of the direct sound by dividing the volume of the sound source by the length of the direct sound path, and can calculate the volume of the reflected sound by dividing the volume of the sound source by the length of the reflected sound path.
[0225] In the selection process, the selection unit 1314 selects the reflected sound to be generated based on parameters calculated before the selection process. Any selection method of this disclosure may be used to select the reflected sound to be generated.
[0226] The selection process may be performed on all reflected sounds, or, as mentioned above, it may be performed only on reflected sounds with high evaluation values based on the evaluation process. In other words, reflected sounds with low evaluation values may be determined to be unselected without even needing to perform the selection process. For example, reflected sounds with very low volume may be considered to have a low evaluation value and may be determined to be unselected.
[0227] Alternatively, for example, a selection process may be performed on all reflected sounds. Then, the evaluation value of the selected reflected sounds is determined, and reflected sounds with low evaluation values may be re-determined as unselected.
[0228] The selection process and the evaluation process may be executed independently or in combination. When the selection process and the evaluation process are executed together, either process may be executed first.
[0229] In the generation process, the generation unit 1315 generates direct sound and reflected sound. For example, the generation unit 1315 generates direct sound from the audio signal included in the input signal based on the arrival time and volume at arrival of the direct sound. In addition, for the reflected sound selected in the selection process, the generation unit 1315 generates reflected sound from the audio signal included in the input signal based on the arrival time and volume at arrival of the reflected sound.
[0230] In binaural processing, the binaural processing unit 1316 performs signal processing so that the direct sound audio signal is perceived as sound arriving to the listener from the direction of the sound source object. Furthermore, the binaural processing unit 1316 performs signal processing so that the reflected sound selected by the selection unit 1314 is perceived as sound arriving to the listener from the reflecting object.
[0231] For example, the binaural processing unit 1316 performs a process to apply the HRIR DB to the listener so that the sound arrives from the location of the sound source object or the location of the obstacle object, based on the listener's position and orientation in the sound space.
[0232] HRIR (Head-Related Impulse Responses) is the response characteristic when a single impulse is generated. Specifically, HRIR is the response characteristic obtained by converting the head-related transfer function, which represents the changes in sound caused by peripheral objects including the ear, head, and shoulders, from a frequency domain representation to a time domain representation using a Fourier transform. The HRIR DB is a database that contains this kind of information.
[0233] Furthermore, the position and orientation of the listener in the sound space are, for example, the position and orientation of a virtual listener in a virtual sound space. The position and orientation of the virtual listener in the virtual sound space may change in accordance with the movement of the listener's head. Alternatively, the position and orientation of the virtual listener in the virtual sound space may be determined based on information acquired from sensor 1405.
[0234] The programs, spatial information, HRIR DB, threshold data, or other parameters used in the above processing are obtained from the memory 1404 provided in the acoustic signal processing device 1001 or from outside the acoustic signal processing device 1001.
[0235] Furthermore, the pipeline processing may include other processing. The rendering unit 1300 may also include processing units (not shown) for performing other processing included in the pipeline processing. For example, the rendering unit 1300 may include a diffraction processing unit and an occlusion processing unit.
[0236] The diffraction processing unit performs a process to generate an audio signal that includes diffracted sound caused by obstacle objects between the listener and the sound source object in a three-dimensional sound field (space). Diffracted sound is sound that travels from the sound source object to the listener by bending around obstacle objects when such obstacle objects are present between the sound source object and the listener.
[0237] The diffraction processing unit, for example, refers to the audio signal and metadata to calculate the path of diffracted sound arriving from the sound source object, bypassing obstacle objects, and reaching the listener, and generates diffracted sound based on that path. In calculating the path, the positions of the sound source object, listener, and obstacle objects in the three-dimensional sound field (space), as well as the shape and size of the obstacle objects, may be used.
[0238] The occlusion processing unit generates an audio signal representing sound leaking through an obstacle object, based on spatial information and information such as the material of the obstacle object, when a sound source object exists on the other side of an obstacle object.
[0239] (Threads) The MPEG-I Impressive Audio standard specifies that the audio rendering process is divided into two independent workflows, and that these two independent workflows are executed on different threads.
[0240] The two independent workflows are called the "Control Workflow" and the "Rendering Workflow." The thread on which the Control Workflow is executed is called the Update thread (also referred to as the Information Update thread). The thread on which the Rendering Workflow is executed is called the Audio thread (also referred to as the Audio thread).
[0241] Note that while multiple workflows are executed by multiple threads within a process, the actions performed by each thread according to its workflow are sometimes also referred to as a process. Furthermore, the act of a thread executing a process is sometimes expressed as "the thread executing the process."
[0242] The spatial information management unit 1201 of this disclosure primarily executes a portion of the metadata update processing handled by the Scene Controller, as defined in the control workflow, as an information update thread. On the other hand, the rendering unit 1203 primarily executes the sound processing handled by each stage of the Renderer Pipeline, as defined in the rendering workflow, as an audio thread.
[0243] Note that the spatial information management unit 1201 may be replaced with the spatial information management unit 1211. Also, the rendering unit 1203 may be replaced with the rendering unit 1213 or with the rendering unit 1300.
[0244] (Update Thread) For example, in a control workflow, the scene controller creates a separate thread as the Update thread to update and interpolate information based on time and position. The update routines within the Update thread are executed at a relatively low rate of several tens of Hz or about 100 Hz. The Update thread is responsible for preparing and updating metadata for acoustic rendering. Specifically, the Update thread executes processing to update spatial information managed by the spatial information management unit 1201.
[0245] For example, the Update thread updates the position and orientation of the listener's avatar in the virtual space based on the position and orientation of the VR goggles worn by the listener. The Update thread also updates the position of objects moving within the virtual space. Furthermore, for SESS, the Update thread generates ray hit information through geometry information analysis and generates adjustment parameters for interaural perceptual cues as parameters for the sound image expansion filter based on the ray hit information.
[0246] Interaural sensory cues include IACC (Interaural Cross-Correlation), IAPD (Interaural Phase Difference), or IALD (Interaural Level Difference).
[0247] (Audio Thread) On the other hand, the processBlock routines, which are the audio processing blocks for each stage in the rendering pipeline, are executed at a fixed audio frame rate in the host system's audio callback thread, i.e., the "Audio thread". For example, a fixed audio frame rate is approximately 375 Hz when a block of 128 samples is processed at a sample rate of 48 kHz. This Audio thread is mainly handled by the rendering unit 1203.
[0248] (Example of spatial information update operation) Spatial information and parameters of the sound image expansion filter are updated by the spatial information management unit 1201 in the information update thread, and used for acoustic processing by the rendering unit 1203 in the audio thread. Data exchange between these threads is performed through a shared memory area (e.g., a buffer or queue) that enables the audio thread to asynchronously read acoustic processing instruction data output as a processing result from the information update thread.
[0249] Specifically, the information update thread prepares sound processing instructions that include setting information such as filter coefficients, algorithm parameters, and references to input / output signal buffers for sound processing in the rendering unit 1203, and writes them to the shared memory area. The audio thread periodically reads the latest sound processing instructions stored in this shared memory area without blocking, and performs processing of the audio stream based on them.
[0250] This allows the information update thread and the audio thread to operate independently at different frequencies, while still ensuring consistent audio processing based on the latest metadata at all times.
[0251] When the spatial information management unit 1201 and the rendering unit 1203 execute processing in different threads, the startup frequency of each thread may be set individually, or the processing may be executed in parallel.
[0252] When the spatial information management unit 1201 and the rendering unit 1203 execute processing in different, independent threads, it is possible to preferentially allocate computing resources to the rendering unit 1203. This makes it possible to safely perform sound output processing where even the slightest delay is unacceptable, for example, where a delay of one sample (0.02 msec) would cause a popping noise.
[0253] In this case, the spatial information management unit 1201 may be limited in the allocation of computing resources. However, since updating spatial information is a low-frequency process compared to the output processing of audio signals (for example, updating the direction of the listener's face), it does not necessarily have to be done instantaneously like the output processing of audio signals, and therefore does not significantly affect the acoustic quality. Consequently, limitations on the allocation of computing resources do not cause problems.
[0254] This is because the frequency of changes in the properties of direct sound is lower than the frequency of audio processing frames for audio output. This low-frequency processing makes it possible to relatively reduce the computational load of the processing. Furthermore, updating information at an unnecessarily fast frequency carries the risk of generating pulsating noise. Updating information at a low frequency makes it possible to avoid this risk.
[0255] Spatial information updates may be performed periodically at predetermined times or intervals, or when predetermined conditions are met. Furthermore, spatial information updates may be performed manually by the listener or the sound space manager, or triggered by changes in an external system. Spatial information updates may also include the following three types of updates:
[0256] (1) Timed Scene Update This update is performed automatically at predetermined times or intervals. This update is used to handle scene changes that occur over time, such as the movement of the sun in the virtual space, the transition between day and night, or events that occur at a specific time (e.g., 10 minutes after the start of the game) (e.g., the appearance of a specific character). The rendering system periodically updates this spatial information based on its internal time.
[0257] (2) Voxel sub-scene update This update is a type that updates only the spatial information of a specific voxel region when the virtual space is represented as a collection of small cubes called voxels. A voxel is also called a volume pixel. A voxel region is also called a subscene. This update is used, for example, in a vast virtual city, when a user, who is the listener, enters a specific area, such as the interior of a particular building, to update only the spatial information related to the acoustic environment, such as the reverberation characteristics and the placement of obstacles inside that building.
[0258] This makes it possible to efficiently perform detailed updates of spatial information by focusing on specific sub-regions (subscenes) that the user is interested in or has moved to.
[0259] (3) Triggered Scene Update This update is triggered manually or automatically in accordance with a specific event or condition provided from an external source. For example, spatial information may be updated when a listener operates a controller and their avatar's position is instantaneously warped (also called "teleport") or time is instantly advanced or reversed. Alternatively, spatial information may be updated when the administrator of the virtual space suddenly performs an action that changes the environment of the space.
[0260] The update mechanisms defined by these three update types are used selectively according to their respective characteristics to accommodate diverse scenarios in virtual space and provide a real-time, immersive audio experience.
[0261] For example, these processes may be performed when the virtual space is created (when the software is created), when processing of the virtual space begins (when the software is started or rendering begins), or when an information update thread periodically occurs during virtual space processing. Furthermore, when the virtual space is created, it may be when the virtual space is constructed before the start of sound processing, when the virtual space information (spatial information) is acquired, or when the software is acquired.
[0262] Note that "thread" is merely one example of a term that refers to a unit of execution of a series of instructions that perform a predetermined role on a processor, and may be replaced with other terms such as module, subroutine, process, procedure, task, step, or workflow.
[0263] (Example of bitstream structure) A bitstream includes, for example, an audio signal and metadata. The audio signal is sound data that represents sound, and indicates information such as the frequency and intensity of the sound. The metadata includes spatial information about the sound space, which is the space of the sound field.
[0264] For example, spatial information is information about the space in which a listener is located when hearing a sound based on an audio signal. Specifically, spatial information is information about a predetermined location (localization position) in sound space (e.g., a three-dimensional sound field) that is used to localize a sound image to that location, that is, to allow the listener to perceive a sound coming from a direction corresponding to that predetermined location. Spatial information includes, for example, sound source object information and location information indicating the listener's position.
[0265] Sound source object information is information about a sound source object that generates sound based on an audio signal. In other words, sound source object information is information about an object (sound source object) that reproduces an audio signal, and is information about a virtual sound source object placed in a virtual sound space. Here, the virtual sound space may correspond to the real space where the sound-generating object is located, and the sound source object in the virtual sound space may correspond to the sound-generating object in the real space.
[0266] The sound source object information may indicate the position of the sound source object placed in the sound space, the orientation of the sound source object, the directivity of the sound emitted by the sound source object, whether or not the sound source object belongs to a living organism, and whether or not the sound source object is a moving object. For example, an audio signal is associated with one or more sound source objects indicated by the sound source object information.
[0267] A bitstream has a data structure that consists of, for example, metadata (control information) and audio signals.
[0268] The audio signal and metadata may be contained in a single bitstream or separately in multiple bitstreams. Furthermore, the audio signal and metadata may be contained in a single file or separately in multiple files.
[0269] A bitstream may exist for each audio source, or for each playback time. Even if a bitstream exists for each playback time, multiple bitstreams may be processed in parallel simultaneously.
[0270] Metadata may be assigned to each bitstream, or it may be assigned to multiple bitstreams collectively as information for controlling multiple bitstreams. In this case, multiple bitstreams may share metadata. Metadata may also be assigned per playback time.
[0271] If multiple bitstreams or files exist, one or more bitstreams or files may contain information indicating the associated bitstream or file. Alternatively, each of all bitstreams or each of all files may contain information indicating the associated bitstream or file.
[0272] Here, the associated bitstream or associated file refers to, for example, a bitstream or file that may be used simultaneously during audio processing. It may also include a bitstream or file that collectively describes information indicating the associated bitstream or associated file.
[0273] Here, the information indicating the associated bitstream or associated file may be, for example, an identifier indicating the associated bitstream or associated file. Alternatively, the information indicating the associated bitstream or associated file may be, for example, a file name, URL (Uniform Resource Locator), or URI (Uniform Resource Identifier) indicating the associated bitstream or associated file.
[0274] In this case, the acquisition unit identifies and acquires the related bitstream or related file based on information indicating the related bitstream or related file. Furthermore, the bitstream or file may contain information indicating the related bitstream or related file, and another bitstream or another file may also contain information indicating the related bitstream or related file.
[0275] Here, the file containing information indicating the associated bitstream or associated file may be a control file, such as a manifest file used for content distribution.
[0276] Furthermore, all or some of the metadata may be obtained from sources other than the audio signal bitstream. For example, either the metadata for controlling the sound or the metadata for controlling the video may be obtained from a source other than the bitstream, or both types of metadata may be obtained from a source other than the bitstream.
[0277] Furthermore, metadata for controlling the video may be included in the bitstream acquired by the stereophonic sound playback system 1000. In this case, the stereophonic sound playback system 1000 may output metadata for controlling the video to a display device that displays an image, or a stereophonic video playback device that plays stereophonic video.
[0278] (Examples of information included in metadata) Metadata may also be information used to describe a scene represented in a sound space. Here, "scene" refers to the collection of all elements that represent the three-dimensional image and sound events in the sound space modeled by the 3D sound reproduction system 1000 using metadata.
[0279] In other words, metadata may include not only information for controlling sound processing, but also information for controlling video processing. Metadata may include only one of these, or both.
[0280] The 3D audio playback system 1000 generates virtual sound effects by performing acoustic processing on the audio signal using metadata included in the bitstream and additionally acquired interactive listener location information. Among the acoustic effects, early reflection processing, obstacle processing, diffraction processing, blockage processing, and reverberation processing may be performed, or other acoustic processing may be performed using metadata. For example, acoustic effects such as distance attenuation, localization, or Doppler effect may be added.
[0281] Furthermore, metadata may include information on switching all or some of the sound effects on or off, or priority information for multiple sound effect processing.
[0282] Furthermore, as an example, metadata includes information about the sound space, including sound source objects and obstacle objects, and information about localization locations for localizing a sound image to a predetermined position within the sound space (i.e., making the listener perceive sound arriving from a predetermined direction).
[0283] Here, an obstacle object is an object that can affect the sound perceived by the listener by, for example, blocking or reflecting sound before the sound emitted by the sound source object reaches the listener. Obstacle objects may include not only stationary objects but also moving objects such as animals or machines. Animals may also be people, etc.
[0284] Furthermore, if multiple sound source objects exist in a sound space, other sound source objects can become obstacle objects for any given sound source object. In other words, non-sound-emitting objects such as building materials or inanimate objects, as well as sound-emitting sound source objects, can become obstacle objects.
[0285] Metadata includes information representing all or part of the shape of the sound space, the shape and location of obstacle objects in the sound space, the shape and location of sound source objects in the sound space, and the position and orientation of the listener in the sound space.
[0286] The sound space may be either a closed space or an open space. The metadata may also include information representing the reflectivity of obstacle objects that can reflect sound within the sound space. For example, floors, walls, or ceilings that constitute the boundary of the sound space may also constitute obstacle objects.
[0287] The reflectance is the ratio of the energy of the reflected sound to the energy of the incident sound, and may be set for each frequency band of the sound. Of course, the reflectance may also be set uniformly, regardless of the frequency band of the sound. In the case of an open sound space, for example, parameters such as a uniformly set attenuation rate, diffracted sound, and early reflections may be used.
[0288] Metadata may include information other than reflectivity as parameters for obstacle objects or sound source objects. For example, metadata may include information about the material of objects as parameters for both sound source objects and non-sounding objects. Specifically, metadata may include information such as diffusion, transmittance, and sound absorption.
[0289] Information regarding a sound source object may include volume, radiation characteristics (directivity), playback conditions, the number and type of sound sources in a single object, and information indicating the sound source area within the object. Playback conditions may specify, for example, whether the sound is continuously playing or triggered by an event. The sound source area within an object may be determined by the relative relationship between the listener's position and the object's position, or it may be determined using the object as a reference.
[0290] For example, if the sound source area is defined by the relative relationship between the listener's position and the object's position, it is possible to make the listener perceive sound A coming from the right side of the object and sound B coming from the left side of the object, from the listener's perspective.
[0291] Furthermore, when the sound source area is defined using an object as a reference, it is possible to fix which part of the object emits which sound, using the object as a reference. For example, if the listener views the object from the front, it is possible to make the listener perceive high-pitched sounds from the right side of the object and low-pitched sounds from the left side. And if the listener views the object from behind, it is possible to make the listener perceive low-pitched sounds from the right side of the object and high-pitched sounds from the left side.
[0292] The metadata related to the space may include the time to the first reflection, the reverberation time, and the ratio of direct sound to diffused sound. When the ratio of direct sound to diffused sound is zero, it is possible to make the listener perceive only direct sound.
[0293] (Example of a sound source object) In the above example, the position information assigned to the sound source object indicates the position of the sound source object as a "point" in the virtual space. In other words, in the above example, the sound source is defined as a "point sound source".
[0294] On the other hand, a sound source in virtual space may be defined as an object having length, size, and shape, that is, not a point source, but a spatially extended sound source. In this case, the distance between the listener and the sound source, and the direction of the sound's arrival are not determined. Therefore, reflected sound originating from such a sound source may be limited to being selected by the selection unit 1302 without analysis by the analysis unit 1301, or regardless of the analysis results. This makes it possible to avoid the degradation of sound quality that may occur by not selecting reflected sound.
[0295] Alternatively, a representative point such as the center of gravity of the object may be determined, and the processing of the present disclosure may be applied assuming that the sound originates from that representative point. In this case, the threshold may be adjusted according to information on the spatial extension of the sound source.
[0296] (Specially Extended Sound Sources) Figure 9 is a conceptual diagram showing a specific example of pipeline processing. In the MPEG-I standard, sound signal processing is divided into multiple processes such as "reflection," "diffraction," "reverberation," "directivity control," and "distance attenuation." These multiple processes are then processed in a vertical order as pipeline processing. This simulates the propagation of sound in a virtual space.
[0297] Among these various processes is a process related to "SESS (Spatially Extended Sound Sources)". SESS corresponds to the process of expanding a sound image from an audio signal associated with a sound source that has a shape, so that a sound image corresponding to the shape is obtained. Here, shape includes size. In other words, a sound source that has a shape is a sound source that has a size corresponding to its shape, and a sound image corresponding to the shape is also a sound image corresponding to the size of the shape. Furthermore, the shape may be, for example, a mesh or a portal.
[0298] Figure 10 is a graph showing the computational load of each stage in pipeline processing. In other words, Figure 10 shows the computational load required for each process in pipeline processing.
[0299] In Figure 10, "DiscoverSESS" is a process for preparing for sound image expansion processing, and specifically, it is a process for understanding the state of the virtual space. "HomogeneousExtent," which is written as "Homogen.Extent" in Figure 10, consists of a process for constructing a sound image expansion filter based on the results of DiscoverSESS, and a process for executing the sound image expansion filter. The construction of the sound image expansion filter includes the generation and updating of the sound image expansion filter.
[0300] A sound image expansion filter is a filter used to expand the sound image of a sound source to the shape of the sound source. Specifically, a sound image expansion filter is a transfer function that calculates a sound signal from the sound source's sound signal that represents the sound so that the sound image of the sound source is expanded to the shape of the sound source and perceived by the listener. For example, a sound image expansion filter is an integrated transfer function obtained by integrating multiple transfer functions that calculate the sound reaching the listener's position from multiple positions included in the shape of the sound source. This integration can also be described as synthesis.
[0301] The sound image corresponding to the sound source is enlarged to the shape of the sound source by a sound image enlargement filter and perceived by the listener.
[0302] The processes for understanding the virtual space environment and constructing the sound image expansion filter are executed in routines on the Update thread. The process for executing the sound image expansion filter is performed in a routine called processBlock on the Audio thread. The processBlock is also called the sound processing block.
[0303] Figure 10 shows that the computational load of SESS-related processing is high, and in particular, the processing of the Update thread among the SESS-related processing appears to have a high computational load.
[0304] (Sound image augmentation preparation unit and sound image augmentation processing unit) Figure 11 is a block diagram showing an example configuration for the rendering unit 1300 to perform SESS-related processing. The rendering unit 1300 in Figure 11 corresponds to the rendering units 1203, 1213, and 1300 in Figures 3A, 3B, and 8.
[0305] In the example in Figure 11, compared to the example in Figure 6, a sound image expansion preparation unit 1341 and a sound image expansion processing unit 1342 are added, while the analysis unit 1301, selection unit 1302, and playback unit 1303 are omitted. However, some or all of the analysis unit 1301, selection unit 1302, and playback unit 1303 may be combined with the sound image expansion preparation unit 1341 and the sound image expansion processing unit 1342. The sound image expansion preparation unit 1341 may be included in the analysis unit 1301, and the sound image expansion processing unit 1342 may be included in the playback unit 1303.
[0306] Furthermore, in the example of Figure 11, compared to the example of Figure 8, the sound image expansion preparation unit 1341 and the sound image expansion processing unit 1342 are added, and several other processing units related to pipeline processing are omitted. However, some or all of the other processing units may be combined with the sound image expansion preparation unit 1341 and the sound image expansion processing unit 1342. The sound image expansion preparation unit 1341, the sound image expansion processing unit 1342, and some or all of the other processing units may perform pipeline processing.
[0307] For example, SESS rendering is performed primarily through the following two dedicated stages in a rendering pipeline as shown in the example in Figure 8. In the example in Figure 11, these two dedicated stages are implemented by the sound image expansion preparation unit 1341 and the sound image expansion processing unit 1342.
[0308] Specifically, the sound image expansion preparation unit 1341 performs operations corresponding to the DiscoverSESS stage. In other words, the sound image expansion preparation unit 1341 is responsible for the DiscoverSESS stage, which is a helper stage that supports the rendering of SESS. More specifically, the sound image expansion preparation unit 1341 casts a predetermined number of rays in all directions from the listener's position in the virtual 3D scene and stores ray hit information that indicates the result of the hit on the shape (i.e., extent) of each expanded sound source.
[0309] This process is executed in each update cycle when the associated scene object or listener position changes. Ray hit information is used in the subsequent sound image augmentation processing unit 1342.
[0310] Furthermore, similar processing may be used not only for detecting the shape of the extended sound source, but also for detecting shapes for other processing such as reverberation and occlusion. The ray hit information may be used not only by the sound image expansion processing unit 1342, but also by the reverberation processing unit 1311, the occlusion processing unit, or other processing units.
[0311] More specifically, the sound image expansion preparation unit 1341 mainly executes the following processes corresponding to the first and second stages as routines of the Update thread.
[0312] In the first stage, it is determined whether or not the sound image expansion filter needs to be updated. Specifically, the sound image expansion preparation unit 1341 detects changes in the listener's position and orientation and the sound source's position, and based on the detection results, it determines whether or not the sound image expansion filter needs to be updated.
[0313] In the second stage, ray hit information is generated. Specifically, if it is determined in the first stage that the sound image expansion filter needs to be updated, the sound image expansion preparation unit 1341 generates numerous rays from the listener's position toward the entire sphere of the virtual space in order to grasp the shape of the sound source (i.e., the shape of the sound image) as seen from the listener's position. The sound image expansion preparation unit 1341 then calculates the position where each ray hits the shape of the sound source.
[0314] Here, the shape of the sound source corresponds to its spatial extent. The position where each ray hits the shape of the sound source is generated as ray hit information and provided for subsequent processing.
[0315] The series of processes in the first and second stages are executed as a routine of the Update thread in the DiscoverSESS stage.
[0316] Furthermore, the sound image expansion processing unit 1342 performs operations corresponding to the Homogeneous Extension stage. For example, based on the information obtained by the sound image expansion preparation unit 1341, the sound image expansion processing unit 1342 performs the main processing to make the listener perceive the spatial extent of SESS. Specifically, the sound image expansion processing unit 1342 performs a series of processes corresponding to the following stages 2.5 to 4.
[0317] In step 2.5, a sound image expansion filter is generated. The sound image expansion filter corresponds to a transfer function and may be expressed as coefficients or parameters. For example, in the generation of the sound image expansion filter, parameters for adjusting binaural perceptual cues such as IACC, IAPD, or IALD are generated, thereby generating the sound image expansion filter. Here, the adjustment of binaural perceptual cues may be a synthesis of binaural perceptual cues.
[0318] Specifically, in the DiscoverSESS stage, parameters are generated based on the ray hit information generated by the sound image expansion preparation unit 1341 to adjust interaural perceptual cues such as IACC, IAPD, or IALD. These parameters contribute to the spatial extent of the perceived sound source or sound image and are used as filter coefficients for adjusting interaural perceptual cues, and as signal conversion coefficients for performing said adjustments.
[0319] By generating the coefficients or parameters described above, a sound image expansion filter corresponding to the transfer function is generated. Note that the transfer function may also be a collective term for such filters, coefficients, and parameters.
[0320] The generation process in stage 2.5 is mainly executed as a routine in the Update thread.
[0321] In the third stage, the sound image expansion filter is updated. Specifically, the sound image expansion filter generated in stage 2.5 is updated based on the detected ray hit information. The sound image expansion filter may also be updated by being regenerated based on the detected ray hit information. This prepares a sound image expansion filter that is optimal for the current scene state. This update process is mainly executed as a routine in the Update thread.
[0322] In the fourth stage, the sound image of the audio signal is expanded. Specifically, the sound image of the audio signal attached to the sound source is expanded using the sound image expansion filter updated in the third stage. This audibly represents the spatial extent of SESS. This signal processing is mainly executed as a processBlock routine in the Audio thread.
[0323] Figure 12 is a diagram illustrating the correspondence between stages and threads. In other words, Figure 12 shows the correspondence between stages and threads in a organized manner, based on the explanation described above.
[0324] The sound image expansion filter may be a "composite filter" of SESS as defined in the MPEG-I standard. Furthermore, the sound image expansion filter may include a series of processes in the SESS rendering process as defined in the MPEG-I standard that adjust interaural perceptual cues to allow the listener to perceive the spatial extent of the sound source. The sound image expansion filter may also be a collective term for this series of processes and the parameters used in them.
[0325] The process of applying an image expansion filter to an audio signal is primarily performed in the HomogeneousExtent stage of the rendering pipeline as defined in the MPEG-I standard.
[0326] The prerequisites for SESS-related processing are as follows: The listener can move while changing direction (6DoF). That is, the listener retains information about its position (LP) and orientation (LO). The sound source can also move. Furthermore, the sound source is not a point source. That is, the sound source retains information about its position (OP) and shape (Extent).
[0327] Figure 13 is a conceptual diagram showing a sound source and a listener. In Figure 13, the sound source is a piano, making it difficult to imagine a situation where the sound source is moving. However, the sound source could be a person, animal, or vehicle that moves while emitting sound. Therefore, positional information can also be assigned to the sound source. In this example, orientation information is not assigned to the sound source. However, since the sound source has a shape, orientation information may also be assigned to the sound source.
[0328] Furthermore, a representative position of the sound source may be used as the sound source position (OP). For example, if the sound source has a shape, a representative position of the sound source, such as the center position of the sound source, may be used as the sound source position (OP). Similarly, a representative position of the listener, such as the center position of the listener, may be used as the listener position (LP).
[0329] Figure 14 is a flowchart showing an example of a process for determining whether or not to update the sound image expansion filter. This determination process corresponds to the first stage of the SESS-related processing.
[0330] In this example, the sound image expansion preparation unit 1341 determines whether or not to update the sound image expansion filter according to three conditions. The first condition is that the listener's position (LP) has changed by a threshold (TLP) or more (S201). For example, the threshold (TLP) is 1 mm. The second condition is that the listener's orientation (LO) has changed by a threshold (TLO) or more (S202). For example, the threshold (TLO) is 1°. The third condition is that the sound source position (OP) has changed by a threshold (TOP) or more (S203). For example, the threshold (TOP) is 1 mm.
[0331] The sound image expansion preparation unit 1341 then determines to update the sound image expansion filter if at least one of the first, second, and third conditions is met (S205). On the other hand, the sound image expansion preparation unit 1341 determines not to update the sound image expansion filter if none of the first, second, and third conditions are met. In other words, in this case, the sound image expansion preparation unit 1341 determines to maintain the sound image expansion filter (S204).
[0332] If it is determined that the sound image expansion filter should be updated, the sound image expansion preparation unit 1341 updates the sound image expansion filter. If it is determined that the sound image expansion filter should be maintained, the sound image expansion preparation unit 1341 maintains the sound image expansion filter. The sound image expansion processing unit 1342 expands the sound image using the sound image expansion filter that has been updated or maintained by the sound image expansion preparation unit 1341.
[0333] Figure 15 is a conceptual diagram showing the process for detecting the shape of a sound source as seen from the listener's position. When a sound image expansion filter is generated or updated, the sound image expansion preparation unit 1341 performs processing to detect the shape of the sound source as seen from the listener's position. Specifically, the sound image expansion preparation unit 1341 emits, for example, 4096 rays from the listener's position. The sound image expansion preparation unit 1341 then detects the position where each ray hits the shape of the sound source. In this way, the sound image expansion preparation unit 1341 detects the shape of the sound source as seen from the listener's position.
[0334] Subsequently, the sound image expansion preparation unit 1341 generates or updates a sound image expansion filter based on the detected shape. That is, the sound image expansion preparation unit 1341 derives a transfer function obtained by synthesizing multiple transfer functions corresponding to multiple positions of the detected shape as a sound image expansion filter. As a result, the sound image expansion preparation unit 1341 generates or updates a sound image expansion filter. The sound image expansion processing unit 1342 expands the sound image of the sound signal using the sound image expansion filter derived in the sound image expansion preparation unit 1341.
[0335] As described above, SESS-related processing includes, for example, the process of constructing a sound image expansion filter and the process of executing the sound image expansion filter. The process of constructing a sound image expansion filter corresponds to the process of constructing a transfer function for calculating an audio signal that represents the sound reaching the listener from the sound source. The process of executing the sound image expansion filter corresponds to the process of calculating an audio signal that represents the sound reaching the listener from the sound source based on the transfer function.
[0336] The process of constructing the sound image expansion filter is performed based on the positions of the sound source and the listener. For example, a sound source has a unique shape. Therefore, the computational cost required to identify the unique shape of the sound source as seen from the listener's position is large. However, if both the sound source and the listener are stationary and not moving, the sound image expansion filter can be continuously used and does not need to be updated. Therefore, in such cases, the process of constructing the sound image expansion filter only needs to be performed once at initialization and does not need to be performed thereafter. This reduces the computational cost.
[0337] The process of executing the sound image expansion filter is also performed based on the positions of the sound source and the listener. However, even if both the sound source and the listener are stationary and not moving, the sound signal will continue to be played. Therefore, even in such cases, the process of executing the sound image expansion filter will be repeated at regular intervals.
[0338] In the example above, the process of constructing a sound image expansion filter is performed if at least one of the following conditions is met: the sound source position changes by 1 mm or more, the listener's position changes by 1 mm or more, or the listener's orientation changes by 1° or more. In other words, changes are detected with extremely sensitive sensitivity. As a result, in the example above, the process of constructing the sound image expansion filter is executed at an excessive frequency.
[0339] Therefore, several embodiments for reducing the frequency of processing to construct a sound image expansion filter while suppressing the degradation of the acoustic quality of the virtual space are shown below. Combinations of these embodiments may be used.
[0340] (First Embodiment) In this embodiment, the threshold can be set by the user or others. This may make it possible to adjust the balance between quality and computational load, which have a trade-off relationship.
[0341] For example, SESS parameters are updated when the sound source or listener moves in the virtual space. Therefore, a threshold (T) is defined to determine whether or not the sound source or listener has moved. If the movement of the sound source or listener exceeds the threshold (T), the SESS parameters are updated; otherwise, the parameters are maintained.
[0342] In this embodiment, the acoustic signal processing device 1001 includes a threshold setting unit that allows the user, administrator, or creator of the virtual space to set a threshold (T). The threshold setting unit may be part of a setting unit for setting configuration information that defines the operating conditions of the virtual space. This may make it possible to directly reduce the amount of computation at the expense of reducing the acoustic effects of the virtual space. Conversely, it may be possible to enhance the acoustic effects of the virtual space by using more computation.
[0343] For example, a threshold (T) is set by a user, administrator, or creator of the virtual space via the threshold setting unit. The acoustic signal processing device 1001 then decides whether to update or maintain the SESS parameters based on the threshold (T).
[0344] Figure 16 is a conceptual diagram showing an example of the judgment process in the first embodiment. When the threshold (T) is set small, the parameter is updated with only a slight movement. As a result, high-precision acoustic effects are obtained. When the threshold (T) is set large, the parameter is not updated unless there is a large movement. As a result, the computational load of the acoustic processing is reduced.
[0345] For example, the threshold (T) may be set in a format such as Configuration parameter (Table 139) in Non-Patent Document 1.
[0346] The threshold (T) may be determined according to the computing resources of the computer (specifically, the acoustic signal processing device 1001) that performs processing in the virtual space. For example, if computing resources are abundant, the threshold (T) will be set low. Conversely, if computing resources are poor, the threshold (T) will be set high.
[0347] The "abundance of computing resources" mentioned above is not necessarily expressed numerically. For example, terms such as "supercomputer" and "cloud" correspond to extremely large computing resources. Similarly, the term "PC" corresponds to large computing resources. Furthermore, the term "mobile terminal" corresponds to medium computing resources. And the term "CPU built into devices such as HMDs (Head-Mounted Displays)" corresponds to small computing resources. The "abundance of computing resources" may be expressed based on these descriptions.
[0348] Alternatively, the "abundance of computing resources" can be expressed by specifying the intended use. For example, the term "wearable use" indicates a small amount of computing resources. The term "mobile use" indicates a medium amount of computing resources. The term "home use" indicates a large amount of computing resources. The term "server use" indicates an extremely large amount of computing resources. The "abundance of computing resources" can be expressed based on these descriptions.
[0349] Furthermore, "abundance of computing resources" may include the abundance of battery capacity. In other words, the abundance of battery capacity may be considered in relation to "abundance of computing resources." Specifically, the threshold may be set higher as the battery capacity decreases. This can suppress the decrease in battery capacity when the battery capacity is low.
[0350] Furthermore, in a modified version of the first embodiment, the acoustic signal processing device 1001 may calculate a threshold (T) according to a reference threshold (C: criteria) and a relaxation coefficient (R: relaxation). Here, the relaxation coefficient (R) is a coefficient for relaxing the sensitivity to detect whether or not movement has occurred. The reference threshold (C) may also be predetermined.
[0351] Furthermore, the acoustic signal processing device 1001 includes a relaxation coefficient setting unit that allows the user, administrator, or creator of the virtual space to set the relaxation coefficient (R). The relaxation coefficient setting unit may be part of a setting unit for setting configuration information that defines the operating conditions of the virtual space. This may make it possible to directly reduce the amount of computation at the expense of reducing the acoustic effects of the virtual space. Conversely, it may be possible to enhance the acoustic effects of the virtual space by using more computation.
[0352] Specifying the relaxation coefficient is assumed to be more intuitively understandable than specifying the threshold itself. For example, the relaxation coefficient (R) is set by the user, administrator, or creator of the virtual space via the relaxation coefficient setting unit. The acoustic signal processing device 1001 then determines the threshold (T) according to the reference threshold (C) and the relaxation coefficient (R), and decides whether to update or maintain the SESS parameters based on the threshold (T). For example, the threshold (T) may be determined as T = C × R.
[0353] Furthermore, the relaxation coefficient (R) may be determined according to the computing resources of the computer (specifically, the acoustic signal processing device 1001) that performs processing in the virtual space. For example, if computing resources are abundant, the relaxation coefficient (R) will be set to a small value. Conversely, if computing resources are poor, the relaxation coefficient (R) will be set to a large value. In addition, "abundance of computing resources" may include the abundance of battery capacity. Specifically, the relaxation coefficient (R) may be set to a large value as the battery capacity decreases. This can suppress the decrease in battery capacity when the battery capacity is low.
[0354] Figure 17 is a conceptual diagram showing an example of the judgment process in a modified version of the first embodiment. When the relaxation coefficient is set small, the parameters are updated with only slight movement. As a result, high-precision acoustic effects are obtained. When the relaxation coefficient is set large, the parameters are not updated unless there is a large movement. As a result, the computational load of the acoustic processing is reduced.
[0355] For example, the relaxation coefficient (R) may be set in a format such as the Configuration parameter (Table 139) in Non-Patent Document 1.
[0356] The relaxation coefficient (R) may be determined according to the computing resources of the computer (specifically, the acoustic signal processing device 1001) that performs processing in the virtual space. For example, if computing resources are abundant, the relaxation coefficient (R) may be set to a small value. Conversely, if computing resources are poor, the relaxation coefficient (T) may be set to a large value.
[0357] The acoustic signal processing device 1001 of this embodiment calculates an acoustic signal indicating sound reaching a listener in a virtual space in which a sound source having a shape and a defined representative position (OP) and a listener exist. For example, the acoustic signal processing device 1001 of this embodiment comprises a first processing unit, a second processing unit, an acquisition unit, a detection unit, a flag management unit, an update unit, a signal generation unit, and a setting unit. These components may be implemented by circuits and memory, and the processing performed by these components may be performed by circuits.
[0358] The first processing unit performs sound image expansion preparation processing. In the sound image expansion preparation processing, M (M > 1) positions (Pm (1 ≤ m ≤ M)) are identified based on the shape associated with the sound source. A transfer function is also generated based on these M positions. The transfer function is a function for calculating an audio signal that indicates the sound reaching the listener's position. The second processing unit performs sound image expansion signal processing. In the sound image expansion signal processing, an audio signal is calculated based on the transfer function.
[0359] The acquisition unit acquires at least one piece of information from the representative location of the sound source, the location of the listener, and the orientation of the listener. The detection unit detects the amount of change in the acquired information over time. The flag management unit sets a flag to true if the amount of change exceeds a threshold (T). The update unit updates the transfer function by causing the first processing unit to perform sound image expansion preparation processing if the flag set by the flag management unit is true.
[0360] The signal generation unit periodically causes the second processing unit to perform sound image expansion signal processing, regardless of the flags in the flag management unit.
[0361] The threshold (T) is determined via a setting unit. The threshold (T) may be determined according to a predetermined reference threshold (C) and a changeable relaxation coefficient (R). The relaxation coefficient (R) may be determined via a setting unit. The threshold (T) may be set according to the computing resources available to the computing device (i.e., the acoustic signal processing device 1001) for generating the virtual space. The threshold (T) may be obtained from configuration information defining the operating conditions of the virtual space. The relaxation coefficient (R) may be obtained from configuration information defining the operating conditions of the virtual space.
[0362] A threshold (T) is set by the administrator or user of the virtual space via the settings unit. The threshold (T) may be determined according to the computing resources of the computer that performs the processing in the virtual space.
[0363] For example, if computing resources are abundant, the threshold (T) is set low. This allows sound image expansion preparation processing to be performed in response to even minute changes, and the transfer function to be updated. Conversely, if computing resources are scarce, the threshold (T) is set high. This reduces sensitivity to changes and decreases the frequency of sound image expansion preparation processing. Consequently, the amount of computation required for sound image expansion preparation processing is saved.
[0364] Alternatively, instead of a threshold (T), a relaxation coefficient (R) may be set according to the computing resources of the computer (i.e., the acoustic signal processing device 1001).
[0365] Furthermore, the relaxation coefficient (R) may be used to control the update frequency (i.e., the frequency at which judgment and updates are performed). For example, if the relaxation coefficient is 1.0, that is, if the reference threshold (C) is set to the threshold (T), the update frequency may be set to 50 Hz. If the relaxation coefficient (R) is greater than 1.0, the update frequency may be set to less than 50 Hz.
[0366] In other words, if the relaxation coefficient is large, the sensitivity to fluctuations in the sound source and listener in the virtual space may be reduced, and at the same time, the update frequency itself may be lowered. Alternatively, the relaxation coefficient (R) may not be used to set the threshold, but only to control the update frequency. For example, the threshold may be immutable, and only the update frequency of the transfer function may be changeable.
[0367] Conversely, a parameter indicating the update frequency may be set, and the relaxation coefficient may be determined according to the update frequency. For example, if the update frequency is set to 50 Hz, the relaxation coefficient may be determined to be 1.0, and if the update frequency is set to 10 Hz, the relaxation coefficient may be determined to be 5.0.
[0368] Figure 18A shows an example of configuration information for a wearable profile. When the profile assumed for the computer that reproduces the virtual space (specifically, the acoustic signal processing device 1001) is defined by a term such as "wearable profile" which has limited computing resources, the settings shown in the configuration information in Figure 18A are adopted.
[0369] Figure 18B shows an example of configuration information for a mobile profile. When the profile expected for the computer that reproduces the virtual space (specifically, the acoustic signal processing device 1001) is defined by a term such as "mobile profile" which indicates moderate computing resources, the settings shown in the configuration information in Figure 18B are adopted.
[0370] Figure 18C shows an example of home profile configuration information. When the profile expected for the computer that reproduces the virtual space (specifically, the acoustic signal processing device 1001) is defined by a term such as "home profile" which has large computing resources, the settings shown in the configuration information in Figure 18C are adopted.
[0371] Figure 18D shows an example of cloud profile configuration information. When the profile expected for the computer that reproduces the virtual space (specifically, the acoustic signal processing device 1001) is defined by a term such as "cloud profile" which indicates extremely large computing resources, the settings shown in the configuration information in Figure 18D are adopted.
[0372] The "configuration information defining the operating conditions of the virtual space" may, for example, be a list of information defining the sampling frequency and processing frame length of the sound signal in the virtual space. Alternatively, the "configuration information defining the operating conditions of the virtual space" may, for example, be a list of information defining the conditions of headphones, etc., for receiving the sound represented by the sound signal in the virtual space.
[0373] (Second Embodiment) In this embodiment, an appropriate threshold is automatically set according to the conditions of the virtual space. For example, the acoustic signal processing device 1001 in this embodiment determines the threshold (T) according to the level of attention the listener directs towards the sound source. This makes it possible to reduce the amount of computation without substantially impairing the acoustic effect of the virtual space in terms of auditory perception.
[0374] Figure 19 is a conceptual diagram showing an example of the judgment process in the second embodiment. In this example, the acoustic signal processing device 1001 sets the threshold (T) to a smaller value when the distance between the sound source and the listener is short than when the distance is long. As a result, the SESS parameters are updated in a sensitive response to small changes. This simulates acoustics that correspond to the everyday feeling that attention is paid to nearby sound sources, but less to distant sound sources.
[0375] In other words, the acoustic signal processing device 1001 reduces the threshold (T) when the distance between the sound source and the listener is short. As a result, even a slight movement of the sound source or listener is judged as "moved". On the other hand, the acoustic signal processing device 1001 increases the threshold (T) when the distance between the sound source and the listener is long. As a result, a movement that would be judged as "moved" when the distance between the sound source and the listener is short may be judged as "not moved" when the distance between the sound source and the listener is long.
[0376] Figure 20 is a conceptual diagram showing another example of the determination process in the second embodiment. In this example, the acoustic signal processing device 1001 sets the threshold (T) to a smaller value when the sound source is in front of the listener (or within the listener's field of vision) than when it is not. This allows the SESS parameters to be updated in a sensitive response to small changes. This simulates acoustics that correspond to the everyday sensation that attention is paid to visible sound sources, but not to invisible sound sources.
[0377] In other words, the acoustic signal processing device 1001 reduces the threshold (T) when the sound source is in front of the listener. This causes the device to determine that the sound source or listener has "moved" even if they have only moved slightly. On the other hand, the acoustic signal processing device 1001 increases the threshold (T) when the sound source is behind the listener. This allows the device to determine that the sound source or listener has "moved" if they are close to each other, but not if they are far apart.
[0378] The acoustic signal processing device 1001 of this embodiment calculates an acoustic signal indicating sound reaching a listener in a virtual space in which a sound source having a shape and a defined representative position (OP) and a listener exist. For example, the acoustic signal processing device 1001 of this embodiment comprises a first processing unit, a second processing unit, an acquisition unit, a detection unit, a flag management unit, an update unit, and a signal generation unit. These components may be implemented by circuits and memory, and the processing performed by these components may be performed by circuits.
[0379] The first processing unit performs sound image expansion preparation processing. In the sound image expansion preparation processing, M (M > 1) positions (Pm (1 ≤ m ≤ M)) are identified based on the shape associated with the sound source. A transfer function is also generated based on these M positions. The transfer function is a function for calculating an audio signal that indicates the sound reaching the listener's position. The second processing unit performs sound image expansion signal processing. In the sound image expansion signal processing, an audio signal is calculated based on the transfer function.
[0380] The acquisition unit acquires at least one piece of information from the representative location of the sound source, the location of the listener, and the orientation of the listener. The detection unit detects the amount of change in the acquired information over time. The flag management unit sets a flag to true if the amount of change exceeds a threshold (T). The update unit updates the transfer function by causing the first processing unit to perform sound image expansion preparation processing if the flag set by the flag management unit is true.
[0381] The signal generation unit periodically causes the second processing unit to perform sound image expansion signal processing, regardless of the flags in the flag management unit.
[0382] The threshold (T) is determined for each sound source, depending on the state of the sound source and the listener. The threshold (T) may also be controlled according to the distance between the sound source and the listener. For example, the longer the distance, the greater the threshold (T) may be. Alternatively, the threshold (T) may also be controlled according to the direction from the listener to the sound source. For example, the closer the direction from the listener to the sound source is to the listener's front, the smaller the threshold (T) may be. Here, the listener's front corresponds to the listener's orientation.
[0383] (Third Embodiment) In this embodiment, an appropriate relaxation coefficient is automatically set according to the conditions of the virtual space. For example, the acoustic signal processing device 1001 in this embodiment calculates a threshold (T) according to a reference threshold (C: criteria) and a relaxation coefficient (R: relaxation), similar to the modified embodiment of the first embodiment. The acoustic signal processing device 1001 may hold a reference threshold (C) or may hold a relaxation coefficient (R). Here, the relaxation coefficient (R) is a coefficient for relaxing the sensitivity to detect whether or not movement has occurred. The reference threshold (C) may also be predetermined.
[0384] Furthermore, for example, the acoustic signal processing device 1001 of this embodiment determines the relaxation coefficient (R) according to the level of attention the listener directs towards the sound source. This makes it possible to reduce the amount of computation without substantially impairing the acoustic effects of the virtual space from an auditory perspective.
[0385] Figure 21 is a conceptual diagram showing an example of the judgment process in the third embodiment. In this example, the acoustic signal processing device 1001 sets the relaxation coefficient (R) to a smaller value when the distance between the sound source and the listener is short than when the distance is long. As a result, the SESS parameters are updated in a sensitive response to small changes. This simulates acoustics that correspond to the everyday feeling that attention is paid to nearby sound sources, but less to distant sound sources.
[0386] In other words, the acoustic signal processing device 1001 reduces the relaxation coefficient (R) when the distance between the sound source and the listener is short. As a result, even a slight movement of the sound source or listener is judged as "movement." On the other hand, the acoustic signal processing device 1001 increases the relaxation coefficient (R) when the distance between the sound source and the listener is long. As a result, a movement that would be judged as "movement" of the sound source or listener when the distance between the sound source and the listener is short may be judged as "not moving" when the distance between the sound source and the listener is long.
[0387] Figure 22 is a conceptual diagram showing another example of the determination process in the third embodiment. In this example, the acoustic signal processing device 1001 sets the relaxation coefficient (R) to a smaller value when the sound source is in front of the listener (or within the listener's field of vision) than when it is not. This allows the SESS parameters to be updated in a sensitive response to small changes. This simulates acoustics that correspond to the everyday sensation that attention is paid to visible sound sources, but not to invisible sound sources.
[0388] In other words, the acoustic signal processing device 1001 reduces the relaxation coefficient (R) when the sound source is in front of the listener. This causes the device to determine that the sound source or listener has "moved" even if they only move slightly. On the other hand, the acoustic signal processing device 1001 increases the relaxation coefficient (R) when the sound source is behind the listener. This allows the device to determine that the sound source or listener has "moved" in the event of a movement that would otherwise be determined as such when the distance between the sound source and the listener is short, while determining that the sound source and listener have "not moved" when the distance between the sound source and the listener is far.
[0389] In this embodiment, the threshold (T) is determined according to a predetermined reference threshold (C) and a changeable relaxation coefficient (R).
[0390] Furthermore, the relaxation coefficient (R) is determined for each sound source, depending on the state of the sound source and the listener. The relaxation coefficient (R) may also be controlled according to the distance between the sound source and the listener. For example, the longer the distance, the larger the relaxation coefficient (R) may be. The relaxation coefficient (R) may also be controlled according to the direction from the listener to the sound source. For example, the closer the direction from the listener to the sound source is to the listener's front, the smaller the relaxation coefficient (R) may be. Here, the listener's front corresponds to the listener's orientation.
[0391] (Fourth Embodiment) In this embodiment, a threshold is set for the relative positional relationship between the sound source and the listener. For example, even if the sound source or the listener moves, if their relative positional relationship does not change, the acoustic effect related to SESS will not change. Therefore, the threshold (T) is defined as a threshold for the amount of change in the positional relationship between the sound source and the listener.
[0392] Specifically, the threshold (T) may be a threshold for the amount of change in direction from the listener to the sound source. Even if the distance the sound source travels is the same, the impact of the sound source's movement on the listener is greater when the sound source is moving closer to the listener than when the sound source is moving farther away. Therefore, considering everyday experience, it is reasonable to apply a threshold not to the distance traveled, but to the difference in direction from the listener to the sound source before and after the movement (i.e., the angle between the direction before and after the movement).
[0393] Alternatively, the threshold (T) may be a threshold for the ratio of the distance between the sound source and the listener before and after movement. For example, even if the distance the sound source travels is the same, the impact on the listener is greater when the sound source is moving closer to the listener than when it is moving farther away. Therefore, considering everyday experience, it is reasonable to apply a threshold not to the distance traveled, but to the ratio of the distance between the sound source and the listener before and after movement (i.e., the ratio of the distance before movement to the distance after movement).
[0394] This makes it possible to reduce the amount of computation without substantially compromising the perceived acoustic effects of the virtual space.
[0395] Figure 23 is a conceptual diagram showing an example of the determination process in the fourth embodiment. In this example, the acoustic signal processing device 1001 evaluates the movement based on the opening angle of the azimuth before and after the update.
[0396] Here, the angle of azimuth before and after the update is the angle between the direction from the listener to the sound source before the sound source or listener's position is updated, and the direction from the listener to the sound source after the sound source or listener's position is updated. The direction from the listener to the sound source before the update may be the direction from the listener to the sound source at the time the SESS parameters were last updated. The direction from the listener to the sound source after the update may be the direction from the listener to the sound source at the time the movement is evaluated.
[0397] As a result, for the same distance traveled, if the opening angle exceeds a threshold, it is determined that the sound source or listener has "moved," and if the opening angle does not exceed the threshold, it is determined that neither the sound source nor the listener has "moved."
[0398] Furthermore, the threshold for the boundary between "moved" and "not moved" may be determined as follows.
[0399] For example, the threshold may be determined based on the auditory discrimination limit for the direction of sound arrival. The auditory discrimination limit for the direction of sound arrival may be the auditory discrimination limit for changes in the direction of sound arrival. Specifically, the threshold may be a value corresponding to the boundary of whether or not it is possible to distinguish a change in the direction of sound arrival based on the auditory discrimination limit. Alternatively, the threshold may be determined within a range in which it is possible to distinguish a change in the direction of sound arrival based on the auditory discrimination limit.
[0400] Furthermore, for example, the threshold may be determined based on the spatial resolution of the virtual space being provided. Specifically, the threshold may be a value corresponding to the boundary of whether or not it is possible to represent changes in the direction of sound arrival based on the spatial resolution of the virtual space. Alternatively, the threshold may be determined within a range in which it is possible to represent changes in the direction of sound arrival based on the spatial resolution of the virtual space.
[0401] Furthermore, the spatial resolution mentioned above may also be the spatial resolution in SOFA (Spatially Oriented Format for Acoustics), which includes data for the virtual space being provided. Here, SOFA includes data such as HRTF (Head-Related Transfer Function) in the virtual space. HRTF is a transfer function used to simulate the direction of sound arrival.
[0402] Furthermore, the relaxation coefficient shown in the modified example of the first embodiment may be applied to the threshold value described above. In other words, the threshold value may be adjusted by the relaxation coefficient. For example, the threshold value may be adjusted by multiplying it by the relaxation coefficient.
[0403] Figure 24 is a conceptual diagram showing another example of the determination process in the fourth embodiment. In this example, the acoustic signal processing device 1001 evaluates the motion based on the ratio of the distance before and after the update.
[0404] Here, the ratio of distances before and after the update is the ratio of the distance between the sound source and the listener before the sound source or listener's position is updated to the distance between the sound source and the listener after the sound source or listener's position is updated. For example, in Figure 24, the distance between the sound source and the listener before the update is shown by a dotted line, and the distance between the sound source and the listener after the update is shown by a solid line.
[0405] The distance between the sound source and the listener before the update may be the distance between the sound source and the listener at the time the SESS parameters were last updated. The distance between the sound source and the listener after the update may be the distance between the sound source and the listener at the time the motion is evaluated.
[0406] Furthermore, the ratio of the distance before and after the update may be the ratio of the distance after the update to the distance before the update, or the ratio of the distance before the update to the distance after the update. Also, the ratio of the distance before and after the update may be the larger of the ratio of the distance after the update to the distance before the update and the ratio of the distance before the update to the distance after the update.
[0407] As a result, for the same distance traveled, if the ratio of distances exceeds a threshold, it is determined that the sound source or listener has "moved," and if the ratio of distances does not exceed the threshold, it is determined that neither the sound source nor the listener has "moved." The ratio of distances is the ratio of the distance between the sound source and the listener shown by the dotted line in Figure 24 to the distance between the sound source and the listener shown by the solid line. For example, the ratio of distances may be the ratio of the distance shown by the solid line to the distance shown by the dotted line in Figure 24.
[0408] Furthermore, the threshold for the boundary between "moved" and "not moved" may be determined as follows.
[0409] For example, the threshold may be determined based on the auditory discrimination limit for volume. The auditory discrimination limit for volume may be the discrimination limit for changes in volume. Specifically, the threshold may be a value corresponding to the boundary of whether or not it is possible to distinguish changes in volume based on the auditory discrimination limit. Alternatively, the threshold may be determined within a range in which it is possible to distinguish changes in volume based on the auditory discrimination limit.
[0410] Furthermore, the relaxation coefficient shown in the modified example of the first embodiment may be applied to the threshold value described above. In other words, the threshold value may be adjusted by the relaxation coefficient. For example, the threshold value may be adjusted by multiplying it by the relaxation coefficient.
[0411] The acoustic signal processing device 1001 of this embodiment calculates an acoustic signal indicating sound reaching a listener in a virtual space in which a sound source having a shape and a defined representative position (OP) and a listener exist. For example, the acoustic signal processing device 1001 of this embodiment comprises a first processing unit, a second processing unit, an acquisition unit, a detection unit, a flag management unit, an update unit, and a signal generation unit. These components may be implemented by circuits and memory, and the processing performed by these components may be performed by circuits.
[0412] The first processing unit performs sound image expansion preparation processing. In the sound image expansion preparation processing, M (M > 1) positions (Pm (1 ≤ m ≤ M)) are identified based on the shape associated with the sound source. A transfer function is also generated based on these M positions. The transfer function is a function for calculating an audio signal that indicates the sound reaching the listener's position. The second processing unit performs sound image expansion signal processing. In the sound image expansion signal processing, an audio signal is calculated based on the transfer function.
[0413] The acquisition unit acquires at least one piece of information from the representative location of the sound source, the location of the listener, and the orientation of the listener. The detection unit detects the amount of change in the acquired information over time. The flag management unit sets a flag to true if the amount of change exceeds a threshold (T). The update unit updates the transfer function by causing the first processing unit to perform sound image expansion preparation processing if the flag set by the flag management unit is true.
[0414] The signal generation unit periodically causes the second processing unit to perform sound image expansion signal processing, regardless of the flags in the flag management unit.
[0415] The threshold (T) is a threshold for the amount of variation in the relative positional relationship between the sound source and the listener. For example, the threshold (T) may be a threshold for the amount of variation in the direction from the listener to the sound source. Alternatively, for example, the threshold (T) may be a threshold for the amount of variation in the distance between the sound source and the listener.
[0416] Alternatively, for example, the threshold (T) may be defined as a pair of a directional threshold (Td) for the variation in direction from the listener to the sound source and a distance threshold (Tl) for the variation in distance between the sound source and the listener.
[0417] The directional threshold (Td) may be determined according to the auditory discrimination limit for the direction of sound arrival. For example, if the auditory discrimination limit for the direction of sound arrival is 2°, the directional threshold (Td) may be set as Td = 2°.
[0418] The distance threshold (Tl) may be determined according to the auditory discrimination limit for sound intensity (i.e., volume). For example, if the auditory discrimination limit for sound intensity is 0.3 dB, the distance threshold (Tl) may be set to Tl = 1.035 based on the following calculation.
[0419] For example, the volume of sound reaching the listener from a sound source decreases inversely proportional to the distance between the sound source and the listener. Therefore, if the distance at a reference time is A and the distance at the current time is B, the volume ratio (RT) is expressed as RT = A / B. If the auditory discrimination limit for volume is 0.3 dB, the distance threshold (Tl) corresponding to this discrimination limit can be determined as RT = Tl = 1.035, based on the relationship 0.3 dB = 20 × log 10 (RT).
[0420] The values above are illustrative, and the discrimination limit and threshold are not limited to the examples given. Other values may be used for the discrimination limit and threshold. Furthermore, thresholds may be determined without regard to the discrimination limit. For example, the direction threshold (Td) may be 1°. Also, for example, the distance threshold (Tl) may be 1.01.
[0421] The threshold (T) may be adjusted based on the relaxation coefficient (R). For example, the above discrimination limit may be defined as the reference threshold (C), and the threshold (T) may be determined using the aforementioned relaxation coefficient (R). Alternatively, two relaxation coefficients (Rd, Rl) may be applied to two thresholds (Td, Tl), respectively.
[0422] Here, the directional relaxation coefficient (Rd) with respect to the directional threshold (Td) may be set to a value smaller than the distance relaxation coefficient (Rl) with respect to the distance threshold (Tl). In other words, control may be performed so that the directional threshold (Td) is not relaxed more than the distance threshold (Tl). For example, the size of a sound image is controlled by IACC (interaural correlation), and IACC is more greatly affected by the direction of sound arrival than by volume. The effect of relaxation can be suppressed by controlling the directional threshold (Td) so that it is not relaxed.
[0423] In the above, the direction from the sound source to the listener may be used instead of the direction from the listener to the sound source. The direction from the listener to the sound source and the direction from the sound source to the listener can be collectively expressed as the direction connecting the sound source and the listener.
[0424] In a modified version of this embodiment, the relative positional relationship between the sound source and the listener may be represented by a vector connecting the sound source and the listener.
[0425] Figure 25 is a conceptual diagram showing the relative positional relationship between a sound source and a listener using vectors. As shown in Figure 25, the relative positional relationship between a sound source and a listener can be represented by a vector connecting the sound source and the listener. The direction connecting the sound source and the listener may be represented by the direction of the vector connecting the sound source and the listener. The distance between the sound source and the listener may also be represented by the magnitude of the vector connecting the sound source and the listener. Furthermore, the amount of change in the relative positional relationship between the sound source and the listener may be represented by the amount of change in the vector connecting the sound source and the listener.
[0426] The acoustic signal processing device 1001 may evaluate motion based on the amount of change in the vector before and after the update. Here, the amount of change in the vector before and after the update is the amount of change between the vector before the sound source or listener's position is updated and the vector after the sound source or listener's position is updated. The vector before the update may be the vector at the time when the SESS parameters were last updated. The vector after the update may be the vector at the time when motion is evaluated.
[0427] The amount of change in the vector before and after the update may be the ratio of the magnitudes of the vectors before and after the update, the difference in direction of the vectors before and after the update, or a combination of the ratio of the magnitudes of the vectors before and after the update and the difference in direction of the vectors before and after the update.
[0428] Furthermore, while Figure 25 shows a vector from the listener to the sound source as the vector connecting the sound source and the listener, a vector from the sound source to the listener may also be used.
[0429] The relative positional relationship between the sound source and the listener can be appropriately represented by a vector connecting the sound source and the listener. Based on this representation, the amount of change in the relative positional relationship between the sound source and the listener can be appropriately detected.
[0430] A modified acoustic signal processing device 1001 of this embodiment calculates an acoustic signal indicating sound reaching a listener in a virtual space where a sound source with a designated representative position (OP) and a listener exist. For example, the modified acoustic signal processing device 1001 of this embodiment comprises a first processing unit, a second processing unit, an acquisition unit, a detection unit, a flag management unit, an update unit, and a signal generation unit. These components may be implemented by circuits and memory, and the processing performed by these components may be performed by circuits.
[0431] The first processing unit performs sound image expansion preparation processing. In sound image expansion preparation processing, a transfer function is generated based on the position of the sound source. The transfer function is a function for calculating an audio signal that represents the sound reaching the listener's position from the position of the sound source. The second processing unit performs sound image expansion signal processing. In sound image expansion signal processing, an audio signal is calculated based on the transfer function.
[0432] The acquisition unit acquires information on the representative location of the sound source and the location of the listener. The detection unit detects the amount of change in the acquired information over time. The flag management unit sets a flag to true if the amount of change exceeds a threshold (T). The update unit updates the transfer function by causing the first processing unit to perform sound image expansion preparation processing if the flag set by the flag management unit is true.
[0433] The signal generation unit periodically causes the second processing unit to perform sound image expansion signal processing, regardless of the flags in the flag management unit.
[0434] Furthermore, the detection unit detects the magnitude and direction of the change in the vector connecting the representative position of the sound source (OP) and the listener's position (LP). For example, the detection unit detects the magnitude and direction of the change in the vector based on the vector VTa at time Ta and the vector VTb at time Tb.
[0435] The flag management unit also comprises a first comparison unit and a second comparison unit. The first comparison unit compares the magnitude variation of a vector with a threshold (Tl) based on the auditory volume discrimination limit to determine whether the magnitude variation of the vector exceeds the threshold (Tl). The second comparison unit compares the direction variation of a vector with a threshold (Td) based on the auditory direction discrimination limit to determine whether the direction variation of the vector exceeds the threshold (Td). The flag management unit sets a flag to true if at least one of the first and second comparison units determines that the variation exceeds the threshold.
[0436] Furthermore, in the modified version of this embodiment, both the magnitude and direction of the vector are used. That is, two thresholds are used: a threshold (Tl) based on the volume discrimination limit and a threshold (Td) based on the direction discrimination limit. The magnitude of the vector corresponds to the distance between the sound source and the listener, and the direction of the vector corresponds to the direction connecting the sound source and the listener.
[0437] Figure 26 is a conceptual diagram illustrating an example where the distance between the sound source and the listener does not change. For example, if the sound source moves along a circle or sphere with the listener's position as the center, the distance between the sound source and the listener does not change, and the volume of the sound reaching the listener remains the same, but the direction of arrival of the sound from the sound source to the listener changes. In such a case, if only a threshold (Tl) based on the volume discrimination limit is used, the sound source and the listener will be judged as "not moving."
[0438] In this modified embodiment, by using a threshold (Td) based on the direction discrimination limit in addition to a threshold (Tl) based on the volume discrimination limit, it is possible to appropriately determine that the sound source or the listener has "moved" even in the above-mentioned case.
[0439] Figure 27 is a conceptual diagram illustrating an example where the direction connecting the sound source and the listener does not change. For example, if the sound source approaches the listener in a straight line, the direction connecting the listener and the sound source does not change, but the volume of the sound reaching the listener from the sound source changes. In such a case, if only a threshold (Td) based on the direction discrimination limit is used, the sound source and the listener will be judged as "not moving."
[0440] In a modified version of this embodiment, by using a threshold (Tl) based on volume discrimination in addition to a threshold (Td) based on direction discrimination, it is possible to appropriately determine that the sound source or the listener has "moved" even in the above-mentioned case.
[0441] (Fifth Embodiment) In this embodiment, an appropriate threshold is automatically set according to the spatial resolution of the virtual space. This makes it possible to efficiently reduce the amount of computation.
[0442] The SESS process includes a sound image expansion preparation process (first process) and a sound image expansion signal processing (second process). The second process calculates an audio signal that represents the sound reaching the listener from the sound source. This process corresponds to simulating sound propagating from the sound source towards the listener, and can also be described as a process that propagates the audio signal towards the listener. This process is performed using a celestial spherical HRTF group (head-related transfer function group) to simulate sound localization.
[0443] The spatial resolution of the virtual space (more specifically, the resolution of the direction of sound arrival) corresponds to the density of multiple HRTFs in the celestial HRTF group. It is reasonable to control the aforementioned threshold (T) according to the density of multiple HRTFs in the celestial HRTF group. This point will be explained below.
[0444] Figure 28 is a conceptual diagram showing an example of a threshold set based on the arrangement of HRTF groups. In the HRTF group on the left side of Figure 28, multiple HRTFs are densely arranged. Each point is associated with an HRTF directed from that point to the listener. In the HRTF group on the right side of Figure 28, multiple HRTFs are loosely arranged. Each vertex is associated with an HRTF directed from that vertex to the listener.
[0445] For example, in SESS processing, by integrating multiple HRTFs corresponding to multiple positions according to the shape of the sound image, an integrated HRTF is calculated as a transfer function to calculate an audio signal that represents the sound reaching the listener's position from multiple positions in the shape of the sound image.
[0446] The dense arrangement of a plurality of HRTFs as shown on the left side of FIG. 28 means that the virtual space can be controlled with high spatial resolution. Therefore, in this arrangement, it is effective to detect motion with high sensitivity and frequently update the sound image expansion preparation process. On the other hand, the coarse arrangement of a plurality of HRTFs as shown on the right side of FIG. 28 means that the virtual space cannot be controlled with high spatial resolution. Therefore, in this arrangement, it is not effective to detect motion with high sensitivity and frequently update the sound image expansion preparation process.
[0447] Therefore, in the present embodiment, the lower the resolution is, the larger the above-mentioned threshold (T) is set according to the HRTF group. This can suppress unnecessary waste of calculation amount.
[0448] Furthermore, the mitigation coefficient shown in the modified example of the first aspect may be applied to the above threshold (T). That is, the above threshold may be adjusted by the mitigation coefficient. For example, the threshold may be adjusted by multiplying the above threshold by the mitigation coefficient. Alternatively, the above threshold (T) may be used as the reference threshold (C) of the third aspect.
[0449] Alternatively, the method for determining the threshold (T) in any of the plurality of aspects and the plurality of modified examples described above may be combined with the method for determining the threshold (T) of the present aspect. That is, the threshold (T) may be adjusted based on the spatial resolution of the HRTF group. Specifically, the denser the arrangement of the HRTF group is, the lower the threshold (T) may be adjusted, and the coarser the arrangement of the HRTF group is, the higher the threshold (T) may be adjusted.
[0450] The acoustic signal processing apparatus 1001 of the present aspect calculates an audio signal representing sound reaching a listener in a virtual space where the listener and a sound source having a shape and a defined representative position (OP) exist. For example, the acoustic signal processing apparatus 1001 of the present aspect includes a first processing unit, a second processing unit, an obtaining unit, a detecting unit, a flag management unit, an updating unit, and a signal generating unit. These constituent elements may be implemented by circuits and a memory, and the processing executed by these constituent elements may be executed by circuits.
[0451] The first processing unit performs sound image expansion preparation processing. In the sound image expansion preparation processing, M (M>1) positions (Pm (1≤m≤M)) resulting from the shape associated with the sound source are specified. Further, a transfer function is generated based on the M positions. The transfer function is a function for calculating a sound signal representing sound that reaches a listener's position. The second processing unit performs sound image expansion signal processing. In the sound image expansion signal processing, the sound signal is calculated based on the transfer function.
[0452] The acquisition unit acquires at least one piece of information among the representative position of the sound source, the position of the listener, and the orientation of the listener. The detection unit detects a temporal fluctuation amount of the acquired information. The flag management unit sets a flag to true when the fluctuation amount exceeds a threshold (T). The updating unit updates the transfer function by causing the first processing unit to perform the sound image expansion preparation processing when the flag of the flag management unit is true.
[0453] The signal generation unit causes the second processing unit to periodically perform the sound image expansion signal processing regardless of the flag of the flag management unit.
[0454] The threshold (T) is determined according to the resolution of the sound arrival direction in the virtual space. For example, the transfer function for calculating a sound signal representing sound that reaches the listener's position may be generated using a celestial-sphere HRTF group (head-related transfer function group) for simulating sound localization. Then, the threshold (T) may be determined according to the density of a plurality of HRTFs in the HRTF group. The denser the plurality of HRTFs in the HRTF group are, the smaller the threshold (T) may be.
[0455] The above HRTF may be replaced with HRIR. Also, in the present disclosure, HRTF and HRIR may be replaced with each other.
[0456] (Adjustment of Implementation Frequency or Computational Resources) For example, the acoustic signal processing apparatus 1001 reproduces a sound signal in a virtual space of 6DoF (Six Degrees of Freedom). In the 6DoF virtual space, at least the position of the listener can be changed by the listener. Therefore, the positional relationship between the sound source and the listener in the virtual space changes moment by moment. The acoustic signal processing apparatus 1001 reflects an acoustic effect corresponding to the positional relationship in the sound heard by the listener.
[0457] For example, the acoustic signal processing device 1001 comprises a first processing unit and a second processing unit. The first processing unit performs processing, including determining the positional relationship between the sound source and the listener, at a predetermined frequency in order to grasp the listener's positional information moment by moment. The second processing unit reproduces an audio signal for the listener to hear. The first processing unit is provided separately from the second processing unit.
[0458] The processing of the second processing unit is real-time. In other words, any interruption in the processing of the second processing unit will be immediately perceived as a quality defect. Therefore, the second processing unit must allocate the processing frequency or computing resources appropriately. On the other hand, the impact of the processing of the first processing unit is limited to the level of acoustic detail. Therefore, the frequency of execution or allocation of computing resources for the processing of the first processing unit may be left to the discretion of the creator, administrator, or user of the virtual space.
[0459] Furthermore, the acoustic signal processing device 1001 that realizes the virtual space is expected to be high-performance, such as a large-scale server or a supercomputer, depending on the application of the virtual space. Alternatively, it is expected to be built into a home PC, a personal mobile device (specifically a smartphone or tablet, etc.), or VR goggles that visualize the virtual space.
[0460] Regardless of the processing capacity of the acoustic signal processing device 1001 that realizes the virtual space, that is, regardless of what kind of device the acoustic signal processing device 1001 is, the second processing unit is allocated processing with an execution frequency or computing resources that satisfy real-time requirements.
[0461] On the other hand, an appropriate frequency of execution or computing resources may be allocated to the processing of the first processing unit according to the processing capacity of the acoustic signal processing unit 1001 that realizes the virtual space. For example, the acoustic signal processing unit 1001 that realizes the virtual space may appropriately adjust the frequency of execution or computing resources allocated to the processing of the first processing unit according to the capacity of the acoustic signal processing unit 1001.
[0462] Furthermore, within the virtual space, there may be sound sources that can move other than the listener (e.g., cars, airplanes, people, or animals). Therefore, even if the listener does not move, the positional information assigned to each sound source changes, causing the positional relationship between the sound source and the listener to change moment by moment. Thus, the amount of change in the positional relationship between the sound source and the listener depends heavily on the attributes of the sound sources that make up the virtual space. For example, the rate at which the positional relationship between the sound source and the listener changes will be significantly different when the sound source is a racing car compared to when the sound source is a cat.
[0463] Therefore, the execution frequency or computing resources allocated to the processing of the first processing unit may be appropriately adjusted according to the attributes of the sound sources constituting the virtual space. For example, a high execution frequency or large computing resources may be allocated to a fast-moving sound source, while a low execution frequency or small computing resources may be allocated to a slow-moving sound source or a sound source that does not move.
[0464] Furthermore, even when a racing car exists in a virtual space, the rate at which the positional relationship between the sound source and the listener changes differs significantly depending on whether the racing car is stationary, simply moving, or in a race. Therefore, the execution frequency or computing resources allocated to the processing of the first processing unit may be appropriately adjusted according to the state of the sound source constituting the virtual space.
[0465] Furthermore, the frequency of execution or computing resources allocated to the processing of the first processing unit regarding the positional relationship between the sound source and the listener in the 6DoF virtual space may be adjusted according to the characteristics of the virtual space or the performance of the acoustic signal processing unit 1001.
[0466] Furthermore, the execution frequency or computational resources allocated to the transfer function update process in the first processing unit may be appropriately adjusted according to an appropriate threshold.
[0467] For example, the frequency of execution or computing resources may be adjusted in conjunction with individually set thresholds. A smaller threshold may result in a higher execution frequency or larger computing resources being allocated, while a larger threshold may result in a lower execution frequency or smaller computing resources being allocated.
[0468] Alternatively, instead of individually evaluating the amount of change in the sound source's position and the amount of change in the listener's position, the transfer function update process may be performed at an appropriate timing by evaluating the amount of change in the relative positional relationship between the sound source and the listener. This allows the frequency of execution or the computational resources allocated to the transfer function update process in the first processing unit to be appropriately adjusted.
[0469] (Example of configuration of an acoustic signal processing device) Figure 29 is a block diagram showing an example of the configuration of an acoustic signal processing device 1001 that calculates an acoustic signal indicating sound reaching a listener. In this example, the acoustic signal processing device 1001 calculates an acoustic signal indicating sound reaching a listener in a virtual space where a sound source having a shape and a defined representative position (OP) and a listener exist.
[0470] Specifically, in this example, the acoustic signal processing device 1001 comprises a detection unit 1351, a determination unit 1352, an update unit 1353, a first processing unit 1354, and a second processing unit 1355. These components may be implemented by circuits and memory, and the processing performed by these components may be performed by circuits. For example, the detection unit 1351, the determination unit 1352, the update unit 1353, and the first processing unit 1354 correspond to the sound image expansion preparation unit 1341, and the second processing unit 1355 corresponds to the sound image expansion processing unit 1342.
[0471] The first processing unit 1354 performs sound image expansion preparation processing. In the sound image expansion preparation processing, M (M > 1) positions (Pm (1 ≤ m ≤ M)) are identified due to the shape associated with the sound source. A transfer function is also generated based on the M positions. The transfer function is a function for calculating an audio signal that indicates the sound reaching the listener's position. The second processing unit 1355 performs sound image expansion signal processing. In the sound image expansion signal processing, an audio signal is calculated based on the transfer function.
[0472] The detection unit 1351 detects the amount of change in the relative positional relationship between the sound source and the listener from the representative position information (OP) of the sound source and the position information (LP) of the listener. The determination unit 1352 determines whether the amount of change exceeds a threshold (T). If the amount of change exceeds the threshold (T), the update unit 1353 updates the transfer function by causing the first processing unit 1354 to perform sound image expansion preparation processing.
[0473] The threshold (T) is the threshold for the amount of change in the relative positional relationship between the sound source and the listener. For example, the threshold (T) may be a threshold determined according to the auditory discrimination limit for changes in the sound reaching the listener's position. Specifically, the threshold (T) may be a threshold (Td) determined according to the auditory discrimination limit for changes in the direction of arrival of the sound. Alternatively, the threshold (T) may be a threshold (Tl) determined according to the auditory discrimination limit for changes in volume.
[0474] Alternatively, each of the two thresholds (Td and Tl) may be used as threshold (T). In this case, the determination unit 1352 determines that the amount of variation exceeds threshold (T) even if it exceeds the threshold (Td) based on the directional discrimination limit, or even if it exceeds the threshold (Tl) based on the volume discrimination limit.
[0475] Furthermore, the first processing unit 1354 may perform preparation processing that is not limited to sound image expansion preparation processing. For example, such preparation processing may be a process of selecting the sound transmission characteristics from the SOFA of the HRTF from the representative position (OP) of the sound source to the listener's position information (LP). In this preparation processing, a transfer function is generated based on the representative position (OP) of the sound source.
[0476] Furthermore, the second processing unit 1355 may perform signal processing that is not limited to sound image expansion signal processing. For example, such signal processing may involve applying the HRTF selected from the SOFA to the sound signal of the sound source. In this signal processing, the sound signal is calculated based on a transfer function generated in a preparation process that is not limited to sound image expansion preparation processing. Then, if the amount of variation exceeds a threshold (T), the update unit 1353 updates the transfer function by having the first processing unit 1354 perform preparation processing.
[0477] Furthermore, the detection unit 1351 may detect the amount of change in the vector connecting the representative position (OP) and the listener's position (LP) as the amount of change in the relative positional relationship between the sound source and the listener. For example, the detection unit 1351 detects the amount of change in the magnitude and direction of the vector based on the vector VTa at time Ta and the vector VTb at time Tb as the amount of change in the relative positional relationship between the sound source and the listener. The determination unit 1352 may then compare the amount of change with a threshold (T) to determine whether the amount of change exceeds the threshold (T).
[0478] The determination unit 1352 may also compare the amount of change in the magnitude of the vector with a threshold (Tl) for the amount of change in the magnitude of the vector to determine whether the amount of change exceeds the threshold (Tl). Alternatively, the determination unit 1352 may also compare the amount of change in the direction of the vector with a threshold (Td) for the amount of change in the direction of the vector to determine whether the amount of change exceeds the threshold (Td).
[0479] The determination unit 1352 may include a first comparison unit that compares the amount of change in the magnitude of a vector with its threshold (Tl), and a second comparison unit that compares the amount of change in the direction of a vector with its threshold (Td). The determination unit 1352 may then determine that the amount of change in the vector, i.e., the amount of change in the positional relationship, exceeds the threshold (T) if the amount of change exceeds the threshold in the comparison of at least one of the first and second comparison units.
[0480] Furthermore, the threshold (T) for the variation in the vector may be determined based on the auditory discrimination limit. Specifically, the threshold (Tl) for the variation in the magnitude of the vector may be determined based on the auditory discrimination limit for volume. Also, the threshold (Td) for the variation in the direction of the vector may be determined based on the auditory discrimination limit for the direction of sound arrival.
[0481] In other words, the determination unit 1352 may compare the amount of vector variation with a threshold (T) based on the auditory discrimination limit to determine whether the amount of vector variation exceeds the threshold (T). Specifically, the determination unit 1352 may compare the amount of variation in vector magnitude with a threshold (Tl) based on the auditory loudness discrimination limit to determine whether the amount of variation exceeds the threshold (Tl). The determination unit 1352 may also compare the amount of variation in vector direction with a threshold (Td) based on the auditory direction discrimination limit to determine whether the amount of variation exceeds the threshold (Td).
[0482] Furthermore, the determination unit 1352 may include a first comparison section that compares the amount of variation in vector magnitude with the threshold (Tl) based on the auditory loudness discrimination limit, and a second comparison section that compares the amount of variation in vector direction with the threshold (Td) based on the auditory direction discrimination limit. Then, the determination unit 1352 may determine that the amount of variation exceeds the threshold when the amount of variation exceeds the threshold in at least one of the comparisons performed by the first comparison section and the second comparison section.
[0483] Note that a change in the distance between the sound source and the listener causes a change in the loudness of the sound that reaches the listener. Therefore, the auditory discrimination limit for loudness corresponds to the auditory discrimination limit for the distance between the sound source and the listener.
[0484] Furthermore, the threshold (Td) for the amount of variation in vector direction may be determined according to the auditory discrimination limit for the direction of sound arrival, as described above. For example, the auditory discrimination limit for the direction of sound arrival is 2°. In this case, the threshold (Td) may be set as Td=2°.
[0485] Furthermore, the threshold (Tl) for the amount of variation in vector magnitude may be determined according to the auditory discrimination limit for sound magnitude (i.e., loudness), as described above. For example, the auditory discrimination limit for sound magnitude is 0.3 dB. In this case, the threshold (Tl) may be set as Tl=1.035 based on the calculation below.
[0486] For example, the volume of sound reaching the listener from a sound source decreases inversely proportional to the distance between the sound source and the listener. Therefore, if the distance at a reference time is A and the distance at the current time is B, the volume ratio (RT) is expressed as RT = A / B. If the auditory discrimination limit for volume is 0.3 dB, the threshold (Tl) corresponding to this discrimination limit can be determined as RT = Tl = 1.035, based on the relationship 0.3 dB = 20 × log 10 (RT).
[0487] In other words, if A / B exceeds 1.035, it is determined that the amount of change exceeds the threshold. On the other hand, changes in volume include changes that increase the volume and changes that decrease the volume. For either of these changes, if the amount of change is large, it is better to update the transfer function. Therefore, it can be determined that the amount of change exceeds the threshold not only when A / B exceeds 1.035 above, but also when A / B exceeds 1 / 1.035 below.
[0488] A threshold can be replaced with a threshold range, which is the range defined by the threshold. In other words, the expression for a threshold and the expression for a threshold range are interchangeable.
[0489] Since volume is inversely proportional to distance, the threshold range for the volume ratio mentioned above corresponds to the threshold range for the distance ratio. Alternatively, in calculating the amount of variation corresponding to two distances at two different timings, the ratio of the larger distance to the smaller distance may be detected as the amount of variation. Alternatively, the larger of A / B and B / A may be detected as the amount of variation. This ensures that the amount of variation is appropriately represented.
[0490] The values above are illustrative, and the discrimination limit and threshold are not limited to the examples given. Other values may be used for the discrimination limit and threshold. Furthermore, the threshold may be determined without regard to the discrimination limit. For example, the direction threshold (Td) may be 1°. Also, for example, the distance threshold (Tl) may be 1.01 or 1 / 1.01.
[0491] The above thresholds (Td, Tl) may be adjusted by a relaxation coefficient (R) specified separately. The relaxation coefficient (R) may also be set using a separately provided setting unit. Furthermore, the relaxation coefficient (R) may be set according to the processing capability of the acoustic signal processing device 1001.
[0492] Furthermore, two relaxation coefficients (Rd, Rl) may be applied to each of the two thresholds (Td, Tl). Here, in order to suppress the relaxation of the threshold (Td) related to direction, which has a significant effect on sound localization, the relaxation coefficient related to direction (Rd) may be set smaller than the relaxation coefficient related to distance (volume) (Rl).
[0493] (Decision processing based on relative positional relationship) Figure 30 is a flowchart showing an example of a decision processing based on the relative positional relationship between the sound source and the listener. This decision processing determines whether or not to update the sound image expansion filter based on the relative positional relationship between the sound source and the listener.
[0494] For example, the sound image expansion preparation unit 1341 determines whether the relative positional relationship between the sound source and the listener has changed (S401). For example, the sound image expansion preparation unit 1341 determines whether the relative positional relationship has changed by a greater amount than a threshold. More specifically, the sound image expansion preparation unit 1341 may determine whether the amount of change in the relative positional relationship exceeds a threshold range.
[0495] Then, if the sound image expansion preparation unit 1341 determines that the relative positional relationship has changed (Yes in S401), it determines to update the sound image expansion filter (S403). On the other hand, if the sound image expansion preparation unit 1341 determines that the relative positional relationship has not changed (No in S401), it determines to maintain the sound image expansion filter (S402).
[0496] If it is determined that the sound image expansion filter should be updated, the sound image expansion preparation unit 1341 updates the sound image expansion filter. If it is determined that the sound image expansion filter should be maintained, the sound image expansion preparation unit 1341 maintains the sound image expansion filter. The sound image expansion processing unit 1342 expands the sound image using the sound image expansion filter that has been updated or maintained by the sound image expansion preparation unit 1341.
[0497] For example, the determination process shown in Figure 14 may be replaced with the determination process shown in Figure 30. Alternatively, the determination of changes in the listener's position (S201) and the determination of changes in the sound source's position (S203) in Figure 14 may be replaced with the determination of changes in the relative positional relationship between the sound source and the listener (S401) in Figure 30.
[0498] Furthermore, in the example shown in Figure 30, the decision to update the sound image expansion filter is made based on whether or not the relative positional relationship between the sound source and the listener has changed. The sound image expansion filter is a transfer function for calculating an audio signal that represents the sound reaching the listener's position from multiple positions on the shape of the sound source. More specifically, the sound image expansion filter is a transfer function that integrates multiple transfer functions for calculating an audio signal that represents the sound reaching the listener's position from multiple positions on the shape of the sound source.
[0499] In other words, in examples such as Figure 30, the decision of whether or not to update the transfer function corresponding to the sound image expansion filter is made based on whether or not the relative positional relationship between the sound source and the listener has changed. However, the decision of whether or not to update other transfer functions is not limited to whether or not to update the transfer function corresponding to the sound image expansion filter; the decision of whether or not to update other transfer functions may also be made based on whether or not the relative positional relationship between the sound source and the listener has changed.
[0500] In other words, for various transfer functions that are affected by the relative positional relationship between the sound source and the listener, the decision of whether or not to update the transfer function may be made based on whether or not the relative positional relationship between the sound source and the listener has changed. For example, the decision of whether or not to update transfer functions related to reverberation, early reflections, distance attenuation, or binaural processing may be made based on whether or not the relative positional relationship between the sound source and the listener has changed.
[0501] On the other hand, as shown in Figure 10, the computational cost required to update the transfer function corresponding to the sound image expansion filter is expected to be large. Therefore, the effect of determining whether or not to update the transfer function corresponding to the sound image expansion filter based on whether or not the relative positional relationship between the sound source and the listener has changed is significant.
[0502] (Representative Example) Figure 31 is a block diagram showing a basic configuration example of the acoustic signal processing device 1001 in an embodiment. In the example of Figure 31, the acoustic signal processing device 1001 includes a circuit 1601 and a memory 1602.
[0503] For example, the multiple components of the acoustic signal processing device 1001 shown in Figure 4 may be implemented by the circuit 1601 and memory 1602 shown in Figure 31. Furthermore, other components of the acoustic signal processing device 1001 or its components, such as the decoder 1200 or rendering unit 1300, may also be implemented by the circuit 1601 and memory 1602 shown in Figure 31.
[0504] Circuit 1601 is a dedicated or general-purpose circuit for performing information processing in the acoustic signal processing device 1001, and is a processor that can access the memory 1602. Circuit 1601 may be an electrical circuit. Alternatively, circuit 1601 may be a processor such as a CPU. Furthermore, circuit 1601 may be a collection of multiple circuits.
[0505] Memory 1602 is a dedicated or general-purpose memory for storing information in the acoustic signal processing device 1001. Memory 1602 may also be an electrical circuit. Memory 1602 may be connected to circuit 1601 or may be included in circuit 1601. Memory 1602 may be a collection of multiple memories. Memory 1602 may also be a magnetic disk or an optical disk, or may be described as storage or a recording medium. Memory 1602 may also be a non-volatile memory or a volatile memory.
[0506] For example, memory 1602 may store spatial information, sound signals, bitstreams, or other information. Furthermore, memory 1602 may store information for circuit 1601 to perform information processing, or it may store a program for circuit 1601 to perform information processing.
[0507] Figure 32 is a flowchart illustrating a basic example of operation of the acoustic signal processing device 1001 in an embodiment. For example, the circuit 1601 of the acoustic signal processing device 1001 shown in Figure 31 performs the processing shown in Figure 32 during operation. Specifically, first, the circuit 1601 prepares a transfer function for calculating an audio signal that represents the sound reaching a listener in a virtual space from a sound source in a virtual space (S501). The circuit 1601 then calculates the audio signal based on the transfer function (S502).
[0508] The circuit 1601 may repeat the process of preparing the transfer function (S501) and calculating the sound signal (S502). Alternatively, the circuit 1601 may repeat the process of preparing the transfer function (S501) and calculating the sound signal (S502) at different intervals.
[0509] Figure 33 is a flowchart illustrating an example of the operation related to the preparation of the transfer function. For example, the circuit 1601 of the acoustic signal processing device 1001 shown in Figure 31 performs the process shown in Figure 33 during the preparation of the transfer function (S501) in the operation example shown in Figure 32.
[0510] Specifically, in preparing the transfer function, circuit 1601 detects the amount of change in the relative positional relationship between the sound source and the listener in the virtual space (S511). Then, if the amount of change in the relative positional relationship exceeds a threshold range (Yes in S512), circuit 1601 updates the transfer function (S513).
[0511] This can sometimes make it possible to appropriately control the updating of the transfer function in accordance with the amount of change in the relative positional relationship between the sound source and the listener. Therefore, it may be possible to update the transfer function in a timely manner while appropriately suppressing the processing load related to updating the transfer function.
[0512] For example, the sound source may have a shape. The transfer function may be determined based on the shape. Circuit 1601 may enlarge the sound image based on the shape by performing a sound image enlargement preparation process which includes preparing a transfer function determined based on the shape, and a sound image enlargement signal processing which includes calculating a sound signal based on the transfer function determined based on the shape.
[0513] This may make it possible to appropriately enlarge the sound image. Furthermore, in the sound image enlargement preparation process, it may be possible to appropriately control the update of the transfer function according to the amount of change in the relative positional relationship between the sound source and the listener. Therefore, it may be possible to appropriately suppress the load of the sound image enlargement preparation process.
[0514] Furthermore, for example, the relative positional relationship may be represented by a vector connecting the sound source and the listener in the virtual space. This can sometimes make it possible to appropriately represent the relative positional relationship between the sound source and the listener in the virtual space. Therefore, it may be possible to appropriately detect the amount of change in the relative positional relationship and appropriately control the updating of the transfer function according to the amount of change in the relative positional relationship.
[0515] Furthermore, for example, the amount of change in the relative positional relationship may include the amount of change in the direction connecting the sound source and the listener in the virtual space. This may make it possible to appropriately control the updating of the transfer function according to the amount of change in the direction connecting the sound source and the listener in the virtual space. Therefore, it may be possible to appropriately suppress the processing load related to updating the transfer function.
[0516] Furthermore, for example, direction may be represented by the direction of the vector connecting the sound source and the listener in the virtual space. This may make it possible to appropriately represent the direction connecting the sound source and the listener in the virtual space. Therefore, it may be possible to appropriately detect the amount of change in direction and appropriately control the updating of the transfer function according to the amount of change in direction.
[0517] Furthermore, for example, the amount of directional variation may be expressed as the difference between the direction at a predetermined reference point and the direction at the point in time when the variation in the relative positional relationship is detected. This may make it possible to appropriately represent the amount of variation that affects hearing with respect to the direction connecting the sound source and the listener in the virtual space. Therefore, it may be possible to appropriately detect the amount of directional variation and appropriately control the updating of the transfer function according to the amount of variation.
[0518] Furthermore, for example, the time point set as the reference point may be the time when the transfer function was last updated. This may make it possible to appropriately control the updating of the transfer function according to the amount of change in the direction from the time the transfer function was last updated.
[0519] Furthermore, for example, the amount of change in relative positional relationship may include the amount of change in the distance between the sound source and the listener in the virtual space. This may make it possible to appropriately control the updating of the transfer function in accordance with the amount of change in the distance between the sound source and the listener in the virtual space. Therefore, it may be possible to appropriately suppress the processing load related to updating the transfer function.
[0520] Furthermore, distance may be represented, for example, by the magnitude of the vector connecting the sound source and the listener in the virtual space. This can sometimes make it possible to appropriately represent the distance between the sound source and the listener in the virtual space. Therefore, it may be possible to appropriately detect the amount of distance variation and appropriately control the updating of the transfer function according to the amount of variation.
[0521] Furthermore, for example, the amount of distance variation may be expressed as the ratio of the distance at a predetermined reference point to the distance at the point in time when the variation in the relative positional relationship is detected. This may make it possible to appropriately represent the amount of variation that affects hearing with respect to the distance between the sound source and the listener in the virtual space. Therefore, it may be possible to appropriately detect the amount of distance variation and appropriately control the updating of the transfer function according to the amount of variation.
[0522] Furthermore, for example, the time point set as the reference point may be the time when the transfer function was last updated. This may make it possible to appropriately control the updating of the transfer function in accordance with the amount of change in distance since the time the transfer function was last updated.
[0523] Furthermore, for example, (i) the amount of change in the direction connecting the sound source and the listener in the virtual space, and (ii) the amount of change in the distance between the sound source and the listener in the virtual space, may each be defined as an amount of change in the relative positional relationship. Also, (i) the directional threshold range for the amount of change in direction, and (ii) the distance threshold range for the amount of change in distance, may each be defined as a threshold range.
[0524] Furthermore, the fact that the amount of change in relative position exceeds the threshold range may also mean that at least one of the following conditions is met: the amount of change in direction exceeds the directional threshold range, or the amount of change in distance exceeds the distance threshold range. If the amount of change in direction exceeds the directional threshold range, the transfer function may be updated. Furthermore, if the amount of change in distance exceeds the distance threshold range, the transfer function may be updated.
[0525] This can make it possible to appropriately control the updating of the transfer function in response to variations in both direction and distance in the relative positional relationship between the sound source and the listener. Therefore, it may be possible to update the transfer function in a timely manner while appropriately suppressing the processing load related to the updating of the transfer function.
[0526] Furthermore, for example, the threshold range may be determined based on the auditory discrimination limit. This may allow for appropriate control of the transfer function update based on an appropriate threshold range determined based on the auditory discrimination limit. Therefore, it may be possible to update the transfer function in a timely manner while appropriately suppressing the processing load related to the transfer function update.
[0527] Furthermore, for example, the discrimination limit for determining the threshold range may include the auditory directional discrimination limit for the direction of sound arrival. This may make it possible to appropriately evaluate the amount of directional variation in the relative positional relationship between the sound source and the listener based on the auditory directional discrimination limit. Therefore, it may be possible to appropriately control the updating of the transfer function. Note that the directional discrimination limit may be 2°.
[0528] Furthermore, for example, the discrimination limit for determining the threshold range may include the auditory volume discrimination limit for sound volume. This may make it possible to appropriately evaluate the amount of distance variation in the relative positional relationship between the sound source and the listener based on the auditory volume discrimination limit. Therefore, it may be possible to appropriately control the updating of the transfer function. The volume discrimination limit may be 0.3 dB.
[0529] Furthermore, for example, the threshold range may be adjusted based on relaxation coefficients. This may allow for adaptive adjustment of the threshold range to the amount of variation in the relative positional relationship between the sound source and the listener. Consequently, it may be possible to appropriately adjust the frequency of transfer function updates.
[0530] Furthermore, for example, the relaxation coefficient may be determined based on the computational resources of the acoustic signal processing device 1001. This may make it possible to adjust the threshold range for the amount of variation in the relative positional relationship between the sound source and the listener based on the computational resources. Consequently, it may be possible to appropriately adjust the frequency of transfer function updates based on the computational resources.
[0531] Furthermore, for example, the relaxation coefficient may be defined as one of several setting parameters included in the configuration information for operation in the virtual space. This may allow the threshold range to be adjusted by the setting parameters included in the configuration information. Consequently, it may be possible to appropriately adjust the frequency of transfer function updates by the setting parameters.
[0532] Furthermore, some or all of the configuration examples and operation examples described using Figures 1 to 30 may be added to the configuration examples and operation examples described using Figures 31, 32, and 33, or modifications may be made in accordance with such some or all.
[0533] Furthermore, a sound source in a virtual space is a virtual sound source and can be described as a sound source object or simply an object. Also, a listener in a virtual space may be an avatar of the listener. And the listener's position in a virtual space may be the position of the avatar corresponding to the listener in that virtual space.
[0534] (Other examples) Although the embodiments of the acoustic signal processing device have been described above according to the embodiments, the embodiments of the acoustic signal processing device are not limited to these embodiments. Modifications that a person skilled in the art can conceive of may be made to the embodiments, and the multiple components in the embodiments may be combined arbitrarily.
[0535] For example, in the embodiment, a process performed by a specific component may be performed by another component instead. Also, the order of multiple processes may be changed, or multiple processes may be executed in parallel. Furthermore, the ordinal numbers such as the first and second used in the description may be rearranged, removed, or newly assigned as appropriate. These ordinal numbers do not necessarily correspond to a meaningful order and may be used to identify elements.
[0536] Furthermore, for example, the expression "at least one of the first element, the second element, and the third element" corresponds to the first element, the second element, the third element, or any combination thereof.
[0537] Furthermore, the acoustic signal processing method, which includes steps performed by each component of the acoustic signal processing device, may be performed by any device. In other words, this acoustic signal processing method may be performed by the acoustic signal processing device or by other devices.
[0538] For example, part or all of the acoustic signal processing method may be executed by a computer equipped with a processor, memory, and input / output circuits, etc. In this case, the acoustic signal processing method may be executed by the computer executing a program that causes the computer to execute the acoustic signal processing method.
[0539] For example, the above program causes a computer to perform an acoustic signal processing method which includes preparing a transfer function for calculating an acoustic signal that represents sound reaching a listener in a virtual space from a sound source in a virtual space, and calculating the acoustic signal based on the transfer function, wherein preparing the transfer function includes detecting the amount of change in the relative positional relationship between the sound source and the listener in the virtual space, and updating the transfer function if the amount of change in the relative positional relationship exceeds a threshold range.
[0540] Furthermore, the above program may be recorded on a non-temporary computer-readable recording medium such as a CD-ROM.
[0541] Furthermore, each component of the acoustic signal processing device may be composed of dedicated hardware, general-purpose hardware that executes the above-mentioned programs, or a combination of these. The general-purpose hardware may consist of memory on which the program is stored, and a general-purpose processor that reads and executes the program from the memory. Here, the memory may be semiconductor memory or a hard disk, and the general-purpose processor may be a CPU.
[0542] Furthermore, dedicated hardware may consist of memory and a dedicated processor, etc. For example, a dedicated processor may refer to memory and execute the above-described acoustic signal processing method.
[0543] Furthermore, each component of the acoustic signal processing device may be an electrical circuit. These electrical circuits may form a single electrical circuit as a whole, or they may be separate electrical circuits. Also, these electrical circuits may correspond to dedicated hardware, or they may correspond to general-purpose hardware that executes the above-mentioned programs, etc.
[0544] This disclosure is applicable to an acoustic signal processing device that calculates sound signals based on a transfer function, and can be used in a 3D sound reproduction system that provides 3D sound in a virtual space, etc.
[0545] 1000 3D sound reproduction system 1001 Sound signal processing device 1002 Voice presentation device 1100, 1120, 1500 Encoding device 1101, 1113 Input data 1102 Encoder 1103 Encoded data 1104, 1114, 1404, 1503, 1602 Memory 1110, 1130 Decoding device 1111 Audio signal 1112, 1200, 1210 Decoder 1121 Transmitting unit 1122 Transmitting signal 1131 Receiving unit 1132 Receiving signal 1201, 1211 Spatial information management unit 1202 Audio data decoder 1203, 1213, 1300 Rendering unit 1301 Analysis unit 1302, 1314 Selection unit 1303 Playback unit 1311 Reverberation processing unit 1312 Early reflection processing unit 1313 Distance attenuation processing unit 1315 Generation unit 1316 Binaural processing unit 1341 Sound image expansion preparation unit 1342 Sound image expansion processing unit 1351 Detection unit 1352 Determination unit 1353 Update unit 1354 First processing unit 1355 Second processing unit 1401 Speaker 1402, 1501 Processor 1403, 1502 Communication IF 1405 Sensor 1601 Circuit
Claims
1. An acoustic signal processing device comprising a memory and a circuit that can access the memory, wherein the circuit, in operation, prepares a transfer function for calculating an acoustic signal representing sound reaching a listener in a virtual space from a sound source in a virtual space, calculates the acoustic signal based on the transfer function, and in preparing the transfer function, detects a change in the relative positional relationship between the sound source and the listener in the virtual space, and updates the transfer function if the change in the relative positional relationship exceeds a threshold range.
2. The sound source has a shape, the transfer function is determined based on the shape, and the circuit performs sound image expansion preparation processing, which includes preparing the transfer function determined based on the shape, and sound image expansion signal processing, which includes calculating the sound signal based on the transfer function determined based on the shape, thereby expanding the sound image based on the shape, according to claim 1.
3. The acoustic signal processing apparatus according to claim 1 or 2, wherein the relative positional relationship is represented by a vector connecting the sound source and the listener in the virtual space.
4. The acoustic signal processing apparatus according to claim 1 or 2, wherein the amount of variation in the relative positional relationship includes the amount of variation in the direction connecting the sound source and the listener in the virtual space.
5. The acoustic signal processing apparatus according to claim 4, wherein the direction is represented by the direction of a vector connecting the sound source and the listener in the virtual space.
6. The acoustic signal processing apparatus according to claim 4, wherein the amount of variation in the direction is expressed as the difference between the direction at a reference point and the direction at the point in time when the amount of variation in the relative positional relationship is detected.
7. The acoustic signal processing apparatus according to claim 6, wherein the time point defined as the criterion is the time when the transfer function was last updated.
8. The acoustic signal processing apparatus according to claim 1 or 2, wherein the amount of variation in the relative positional relationship includes the amount of variation in the distance between the sound source and the listener in the virtual space.
9. The acoustic signal processing apparatus according to claim 8, wherein the distance is expressed by the magnitude of the vector connecting the sound source and the listener in the virtual space.
10. The acoustic signal processing apparatus according to claim 8, wherein the amount of variation in distance is expressed as the ratio of the distance at a reference point to the distance at the point in time when the amount of variation in the relative positional relationship is detected.
11. The acoustic signal processing apparatus according to claim 10, wherein the time point defined as the criterion is the time when the transfer function was last updated.
12. (i) The amount of variation in the direction connecting the sound source and the listener in the virtual space, and (ii) The amount of variation in the distance between the sound source and the listener in the virtual space, are each defined as the amount of variation in the relative positional relationship, (i) The directional threshold range for the amount of variation in the direction, and (ii) The distance threshold range for the amount of variation in the distance, are each defined as the threshold range, and the amount of variation in the relative positional relationship exceeds the threshold range means that at least one of the conditions of the amount of variation in the direction exceeding the directional threshold range and the condition of the amount of variation in the distance exceeding the distance threshold range is satisfied, and if the amount of variation in the direction exceeds the directional threshold range, the transfer function is updated, and if the amount of variation in the distance exceeds the distance threshold range, the transfer function is updated, the acoustic signal processing device according to claim 1 or 2.
13. The acoustic signal processing apparatus according to claim 1 or 2, wherein the threshold range is determined based on the auditory discrimination limit.
14. The acoustic signal processing apparatus according to claim 13, wherein the discrimination limit for determining the threshold range includes the auditory directional discrimination limit with respect to the direction of arrival of the sound.
15. The acoustic signal processing apparatus according to claim 13, wherein the discrimination limit for determining the threshold range includes the auditory volume discrimination limit for the volume of the sound.
16. The acoustic signal processing apparatus according to claim 1 or 2, wherein the threshold range is adjusted based on a relaxation coefficient.
17. The acoustic signal processing apparatus according to claim 16, wherein the relaxation coefficient is determined based on the computational resources of the acoustic signal processing apparatus.
18. The acoustic signal processing device according to claim 16, wherein the relaxation coefficient is defined as one of a plurality of setting parameters included in the configuration information of the operation in the virtual space.
19. An acoustic signal processing method comprising: preparing a transfer function for calculating an acoustic signal that represents sound reaching a listener in a virtual space from a sound source in a virtual space; and calculating the acoustic signal based on the transfer function, wherein preparing the transfer function includes detecting the amount of change in the relative positional relationship between the sound source and the listener in the virtual space; and updating the transfer function if the amount of change in the relative positional relationship exceeds a threshold range.
20. A program for causing a computer to execute the acoustic signal processing method described in claim 19.