Acoustic signal processing method, computer program, and acoustic signal processing device

JPWO2024084949A5Pending Publication Date: 2025-07-02
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024551436
Authority / Receiving Office
JP · JP
Patent Type
Applications
Filing Date
2025-03-28
Publication Date
2025-07-02

AI Technical Summary

Technical Problem

Existing acoustic signal processing methods fail to provide a realistic sense of presence by not accurately simulating the timing of aerodynamic sounds in virtual environments, leading to discomfort for listeners.

Method used

An acoustic signal processing method that acquires object information indicating changes causing wind and outputs aerodynamic sound data after a predetermined time, based on factors like wind speed and distance, to synchronize the timing of aerodynamic sounds with the listener's perception in virtual spaces.

Benefits of technology

This approach allows listeners to experience aerodynamic sounds at appropriate timings, enhancing realism and reducing discomfort by simulating the timing of wind-generated sounds similar to real-world scenarios.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

This acoustic signal processing method includes: an acquisition step for acquiring object information indicating a change in an object that causes wind, and indicating a prescribed timing relating to the change in the object; and an output step for, starting from the prescribed timing indicated by the acquired object information, outputting aerodynamic sound data indicating an aerodynamic sound due to the wind after a prescribed amount of time based on the change in the object.
Need to check novelty before this filing date? Find Prior Art

Description

Acoustic signal processing method, computer program, and acoustic signal processing device

[0001] The present disclosure relates to an acoustic signal processing method and the like.

[0002] Patent Document 1 discloses a technology related to a stereophonic calculation method, which is an acoustic signal processing method, in which the arrival time of sound at a listener (observer) is controlled to vary depending on the distance between the sound source and the listener and the speed of sound.

[0003] JP 2013-201577 A International Publication No. 2021 / 180938

[0004] However, with the technology disclosed in Patent Document 1, it may be difficult to give a sense of realism to the listener.

[0005] Therefore, an object of the present disclosure is to provide an acoustic signal processing method and the like that can give a sense of realism to listeners.

[0006] An acoustic signal processing method according to one aspect of the present disclosure includes an acquisition step of acquiring object information indicating a change in an object that causes wind and a predetermined timing related to the change in the object, and an output step of outputting aerodynamic sound data indicating aerodynamic sound caused by the wind a predetermined time period based on the change in the object from the predetermined timing indicated by the acquired object information.

[0007] Furthermore, a computer program according to one aspect of the present disclosure causes a computer to execute the above-described acoustic signal processing method.

[0008] Moreover, an acoustic signal processing device according to one aspect of the present disclosure includes an acquisition unit that acquires object information indicating a change in an object that causes wind and a predetermined timing related to the change in the object, and an output unit that outputs aerodynamic sound data indicating the aerodynamic sound caused by the wind a predetermined time after the predetermined timing indicated by the acquired object information based on the change in the object.

[0009] These comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0010] According to the acoustic signal processing method according to one aspect of the present disclosure, it is possible to give a sense of realism to the listener.

[0011] FIG. 1 is a diagram illustrating a stereophonic (immersive audio) reproduction system as an example of a system to which the acoustic processing or decoding processing of the present disclosure can be applied. FIG. 2 is a functional block diagram illustrating the configuration of an encoding device as an example of the encoding device of the present disclosure. FIG. 3 is a functional block diagram illustrating the configuration of a decoding device as an example of the decoding device of the present disclosure. FIG. 4 is a functional block diagram illustrating the configuration of an encoding device as another example of the encoding device of the present disclosure. FIG. 5 is a functional block diagram illustrating the configuration of a decoding device as another example of the decoding device of the present disclosure. FIG. 6 is a functional block diagram illustrating the configuration of a decoder as an example of the decoder in FIG. 3 or FIG. 5. FIG. 7 is a functional block diagram illustrating the configuration of a decoder as another example of the decoder in FIG. 3 or FIG. 5. FIG. 8 is a diagram illustrating an example of the physical configuration of an acoustic signal processing device. FIG. 9 is a diagram illustrating an example of the physical configuration of an encoding device. FIG. 10 is a block diagram illustrating the functional configuration of an acoustic signal processing device according to an embodiment. FIG. 11 is a flowchart of operation example 1 of the acoustic signal processing device according to the embodiment. FIG. 12 is a diagram illustrating a fan as an object and a listener according to operation example 1. FIG. 13A is a diagram illustrating the process of determining the predetermined time in step S40 shown in FIG. 11 . FIG. 13B is a diagram illustrating a detailed example of output of aerodynamic sound data according to an embodiment. FIG. 13C is a diagram illustrating another detailed example of output of aerodynamic sound data according to an embodiment. FIG. 14 is a flowchart of operation example 2 of the acoustic signal processing device according to an embodiment. FIG. 15 is a diagram showing an ambulance and a listener, which are objects according to operation example 2. FIG. 16 is a schematic diagram illustrating predetermined timing according to operation example 2. FIG. 17 is a flowchart illustrating details of step S35 according to operation example 2. FIG. 18 is a flowchart illustrating details of step S35 according to another first example of operation example 2. FIG. 19 is a diagram illustrating an example of functional blocks and steps for explaining a case where the rendering units of FIGS. 6 and 7 perform pipeline processing.

[0012] (Findings that Form the Basis of the Present Disclosure) Conventionally, an acoustic signal processing method is known in which the arrival time of sound at a listener in a virtual space is controlled.

[0013] Patent Literature 1 discloses a technology related to a stereophonic calculation method, which is an acoustic signal processing method. In this acoustic signal processing method, the arrival time of sound at a listener is controlled to change depending on the distance between the sound source and the listener and the speed of sound. More specifically, the arrival time is controlled to become longer as the distance increases and the slower the speed of sound. This allows the listener to recognize the distance between the object emitting the sound, i.e., the sound source, and themselves.

[0014] Such controlled sound is used in applications for reproducing three-dimensional sound in a space (virtual space) where a user (listener) is present, such as virtual reality (VR) or augmented reality (AR). Such controlled sound is particularly used in virtual spaces where information on the listener's 6 DoF (Degrees of Freedom) is sensed.

[0015] Incidentally, the sound reaching the listener disclosed in Patent Document 1 is the traveling sound of a vehicle (moving sound source), which is an object in VR or AR, and is a sound (such as an engine sound) emitted by the vehicle itself. However, in real space, for example, a vehicle creates wind when it moves. The wind created by this vehicle reaches the listener's ears, generating aerodynamic sound. This aerodynamic sound is generated, for example, according to the shape of the listener L's ear when wind generated by an object (for example, a vehicle) reaches the listener. Note that the object that creates wind is not limited to a traveling (moving) object such as the vehicle, but also includes an object that generates wind, such as an electric fan.

[0016] However, Patent Document 1 does not disclose how to allow a listener to hear aerodynamic sound. More specifically, Patent Document 1 does not disclose a technology for controlling the time it takes for aerodynamic sound to reach the listener when an object creates wind. With the technology disclosed in Patent Document 1, the listener cannot hear the aerodynamic sound at the appropriate timing, which causes the listener to feel uncomfortable and makes it difficult for the listener to experience a sense of realism. Therefore, there is a need for an audio signal processing method or the like that can provide the listener with a sense of realism.

[0017] Therefore, the acoustic signal processing method according to the first aspect of the present disclosure includes an acquisition step of acquiring object information indicating a change in an object that causes wind and a predetermined timing related to the change in the object, and an output step of outputting aerodynamic sound data indicating the aerodynamic sound caused by the wind a predetermined time after the predetermined timing indicated by the acquired object information that is based on the change in the object.

[0018] This makes it possible to output aerodynamic sound data at a timing when a predetermined time has elapsed from a predetermined timing. As a result, the listener can hear the aerodynamic sound at the appropriate timing, and the listener is less likely to feel uncomfortable and can obtain a sense of realism. In other words, an acoustic signal processing method that can give the listener a sense of realism is realized.

[0019] For example, an acoustic signal processing method according to a second aspect of the present disclosure is the acoustic signal processing method according to the first aspect, in which the object information indicates a change in the wind due to a change in the object and the predetermined timing is the timing of the change in the wind, and the acoustic signal processing method includes a determination step of determining the predetermined time based on the wind indicated by the acquired object information.

[0020] This allows aerodynamic sound data to be output when a predetermined time determined based on the wind has elapsed since the wind changed, allowing the listener to hear the aerodynamic sound at a more appropriate timing.

[0021] For example, an acoustic signal processing method according to a third aspect of the present disclosure is the acoustic signal processing method according to the second aspect, in which the change in wind indicated by the object information indicates a change in the wind speed, and in the determination step, the specified time is determined based on the wind speed.

[0022] This allows the predetermined time to be determined based on the wind speed, allowing the listener to hear the aerodynamic sound at a more appropriate timing.

[0023] Furthermore, for example, an acoustic signal processing method according to a fourth aspect of the present disclosure is the acoustic signal processing method according to the third aspect, in which the aerodynamic sound is sound generated at the changed wind speed.

[0024] This allows the aerodynamic sound heard by the listener in the virtual space to be closer to the aerodynamic sound heard by the listener in the real space.

[0025] Furthermore, for example, an acoustic signal processing method according to a fifth aspect of the present disclosure is the acoustic signal processing method according to the first aspect, in which the object information indicates the position of the object, and the acoustic signal processing method includes a determination step of determining the predetermined time based on the distance between the position of a listener of the aerodynamic sound and the position of the object indicated by the acquired object information.

[0026] As a result, the predetermined time is determined based on the distance, allowing the listener to hear the aerodynamic sound at a more appropriate timing.

[0027] Furthermore, for example, an acoustic signal processing method according to a sixth aspect of the present disclosure is the acoustic signal processing method according to the third or fourth aspect, wherein the object information indicates the position of the object, and the determining step determines the predetermined time based on the wind speed and the distance between the position of a listener of the aerodynamic sound and the position of the object indicated by the acquired object information.

[0028] This allows the predetermined time to be determined based on the wind speed and the distance, allowing the listener to hear the aerodynamic sound at a more appropriate timing.

[0029] Furthermore, for example, an acoustic signal processing method according to a seventh aspect of the present disclosure is an acoustic signal processing method according to any one of the first to sixth aspects, in which the object information indicates that the predetermined timing is a first timing for outputting sound data associated with the object, and in the output step, the aerodynamic sound data is output after the predetermined time from the first timing indicated by the acquired object information.

[0030] This allows, for example, when an object generates a sound, aerodynamic sound data to be output at a timing when a predetermined time has elapsed from the first timing at which the sound is output, allowing the listener to hear the aerodynamic sound at a more appropriate timing.

[0031] Furthermore, for example, an acoustic signal processing method according to an eighth aspect of the present disclosure is an acoustic signal processing method according to any one of the first to sixth aspects, wherein the object information indicates the position of the object and that the predetermined timing is a second timing at which the distance between the position of a listener of the aerodynamic sound and the position of the object becomes shorter than a predetermined distance, and the output step outputs the aerodynamic sound data after the predetermined time has elapsed since the second timing indicated by the acquired object information.

[0032] This allows the aerodynamic sound data to be output at the second timing when the distance becomes shorter than the predetermined distance, that is, at the timing when a predetermined time has elapsed since the second timing when the object approached the listener, allowing the listener to hear the aerodynamic sound at a more appropriate timing.

[0033] Furthermore, for example, an acoustic signal processing method according to a ninth aspect of the present disclosure is an acoustic signal processing method according to any one of the first to sixth aspects, wherein the object information indicates that the change in wind due to a change in the object is a change in the wind direction, and the predetermined timing is a third timing at which the change in wind direction occurred, and the output step outputs the aerodynamic sound data after the predetermined time from the third timing indicated by the acquired object information.

[0034] This allows the aerodynamic sound data to be output at a timing when a predetermined time has elapsed from the third timing when the change in wind direction occurred, allowing the listener to hear the aerodynamic sound at a more appropriate timing.

[0035] Furthermore, for example, an acoustic signal processing method according to a tenth aspect of the present disclosure is the acoustic signal processing method according to the sixth aspect, in which the object is an object that generates the sound and the wind indicated by sound data associated with the object, and the aerodynamic sound is aerodynamic sound that is generated when the wind generated by the object reaches the listener.

[0036] This allows an object such as an electric fan that generates sound and wind to be used, and aerodynamic sound caused by wind blowing out from the object to be realized.

[0037] Furthermore, for example, an acoustic signal processing method according to an eleventh aspect of the present disclosure is the acoustic signal processing method according to the tenth aspect, in which, when the distance is D, the distance from the position of the object at which the wind speed is S is U, and the predetermined time is t, t satisfies the following formula:

[0038] t={(D-U)^2} / {So×U×(log(D)-log(U))

[0039] As a result, in the determination step, the predetermined time can be determined to be the time from the predetermined timing until the wind generated by the object reaches the listener. Therefore, since the aerodynamic sound data can be output at a timing when this predetermined time has elapsed from the predetermined timing, the listener can hear the aerodynamic sound at a more appropriate timing.

[0040] Also, for example, an acoustic signal processing method according to a twelfth aspect of the present disclosure is the acoustic signal processing method according to the sixth aspect, in which the object is an object that generates the wind by moving the position of the object, and the aerodynamic sound is aerodynamic sound that is produced when the wind generated by the movement reaches the listener.

[0041] This allows a vehicle or the like that generates wind as it moves to be used as an object, and makes it possible to realize aerodynamic sound caused by the wind generated by that movement.

[0042] For example, an acoustic signal processing method according to a thirteenth aspect of the present disclosure is the acoustic signal processing method according to the twelfth aspect, in which the specified timing indicated by the object information is the timing at which the amount of change in the distance over time turns from negative to positive.

[0043] This allows the aerodynamic sound data to be output when a predetermined time has elapsed since the time when the listener's position and the object's position are closest to each other, allowing the listener to hear the aerodynamic sound at a more appropriate time.

[0044] Furthermore, for example, an acoustic signal processing method according to a fourteenth aspect of the present disclosure is the acoustic signal processing method according to the twelfth or thirteenth aspect, in which, when the distance is D, the distance from the position of the object at which the wind speed of the wind generated by the movement is So is U, and the predetermined time is t, t satisfies the following formula:

[0045] t={(D-U)^2} / {So×U×(log(D)-log(U))

[0046] As a result, in the determination step, the predetermined time can be determined to be the time from the predetermined timing until the wind generated by the object reaches the listener. Therefore, since the aerodynamic sound data can be output at a timing when this predetermined time has elapsed from the predetermined timing, the listener can hear the aerodynamic sound at a more appropriate timing.

[0047] Furthermore, for example, a computer program according to a fifteenth aspect of the present disclosure is a computer program for causing a computer to execute the acoustic signal processing method according to any one of the first to fourteenth aspects.

[0048] This allows the computer to execute the above-described acoustic signal processing method in accordance with the computer program.

[0049] Furthermore, for example, an acoustic signal processing device according to a sixteenth aspect of the present disclosure includes an acquisition unit that acquires object information indicating a change in an object that causes wind and a predetermined timing related to the change in the object, and an output unit that outputs aerodynamic sound data indicating the aerodynamic sound caused by the wind a predetermined time after the predetermined timing indicated by the acquired object information based on the change in the object.

[0050] This makes it possible to output aerodynamic sound data at a timing when a predetermined time has elapsed from a predetermined timing. As a result, the listener can hear the aerodynamic sound at the appropriate timing, and the listener is less likely to feel uncomfortable and can obtain a sense of realism. In other words, an acoustic signal processing device that can give the listener a sense of realism is realized.

[0051] Furthermore, these comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0052] Hereinafter, the embodiments will be specifically described with reference to the drawings.

[0053] The embodiments described below are all comprehensive or specific examples, and the numerical values, shapes, materials, components, the arrangement and connection of the components, steps, and the order of steps shown in the following embodiments are merely examples and are not intended to limit the scope of the claims.

[0054] In the following description, ordinal numbers such as "first" and "second" may be attached to elements. These ordinal numbers are attached to elements to identify the elements and do not necessarily correspond to a meaningful order. These ordinal numbers may be rearranged, newly added, or removed as appropriate.

[0055] Furthermore, each figure is a schematic diagram and is not necessarily an exact illustration. Therefore, the scales and the like do not necessarily match in each figure. In each figure, the same reference numerals are used to denote substantially the same components, and redundant explanations will be omitted or simplified.

[0056] In this specification, terms indicating relationships between elements such as verticality, and numerical ranges, are not expressions that only express a strict meaning, but also expressions that include a substantially equivalent range, for example, a difference of about a few percent.

[0057] (Embodiments) [Example of a device to which the acoustic processing technique or encoding / decoding technique of the present disclosure can be applied] <Stereophonic sound reproduction system> Fig. 1 is a diagram showing a stereophonic (immersive audio) reproduction system A0000, which is an example of a system to which the acoustic processing or decoding process of the present disclosure can be applied. The stereophonic sound reproduction system A0000 includes an acoustic signal processing device A0001 and an audio presentation device A0002.

[0058] The acoustic signal processing device A0001 performs acoustic processing on an audio signal emitted by a virtual sound source to generate an acoustically processed audio signal to be presented to a listener (i.e., a listener). The audio signal is not limited to voice, and may be any audible sound. The acoustic processing is, for example, signal processing performed on an audio signal to reproduce one or more sound-related effects that a sound generated from a sound source experiences between the time the sound is emitted and the time the listener hears the sound. The acoustic signal processing device A0001 performs the acoustic processing based on information describing factors that cause the above-mentioned sound-related effects. The spatial information includes, for example, information indicating the positions of the sound source, the listener, and surrounding objects, information indicating the shape of the space, parameters related to sound propagation, etc. The acoustic signal processing device A0001 is, for example, a PC (personal computer), a smartphone, a tablet, a game console, or the like.

[0059] The signal after acoustic processing is presented to a listener (user) from the audio presentation device A0002. The audio presentation device A0002 is connected to the audio signal processing device A0001 via wireless or wired communication. The audio signal after acoustic processing generated by the audio signal processing device A0001 is transmitted to the audio presentation device A0002 via wireless or wired communication. If the audio presentation device A0002 is composed of multiple devices, such as a device for the right ear and a device for the left ear, the multiple devices present sound in synchronization by communication between the multiple devices or between each of the multiple devices and the audio signal processing device A0001. The audio presentation device A0002 is, for example, headphones, earphones, or a head-mounted display worn on the listener's head, or a surround speaker composed of multiple fixed speakers.

[0060] The stereophonic sound reproduction system A0000 may be used in combination with an image presentation device or a stereoscopic video presentation device that provides an ER (Extended Reality) experience, including visual VR or AR.

[0061] 1 shows an example system configuration in which the acoustic signal processing device A0001 and the audio presentation device A0002 are separate devices, but the stereophonic sound reproduction system A0000 to which the acoustic signal processing method or decoding method of the present disclosure can be applied is not limited to the configuration shown in FIG. 1. For example, the acoustic signal processing device A0001 may be included in the audio presentation device A0002, which may perform both acoustic processing and sound presentation. Furthermore, the acoustic signal processing device A0001 and the audio presentation device A0002 may share the acoustic processing described in the present disclosure, or a server connected to the acoustic signal processing device A0001 or the audio presentation device A0002 via a network may perform part or all of the acoustic processing described in the present disclosure.

[0062] In the above description, the acoustic signal processing device A0001 is referred to as such, but if the acoustic signal processing device A0001 performs acoustic processing by decoding a bit stream generated by encoding at least a portion of the data of an audio signal or spatial information used for acoustic processing, the acoustic signal processing device A0001 may also be referred to as a decoding device.

[0063] <Example of Encoding Device> FIG. 2 is a functional block diagram showing the configuration of an encoding device A0100, which is an example of an encoding device according to the present disclosure.

[0064] The input data A0101 is data to be coded, including spatial information and / or an audio signal, and is input to the encoder A0102. Details of the spatial information will be described later.

[0065] The encoder A0102 encodes the input data A0101 to generate encoded data A0103. The encoded data A0103 is, for example, a bit stream generated by the encoding process.

[0066] The memory A 0104 stores the encoded data A 0103. The memory A 0104 may be, for example, a hard disk or a solid-state drive (SSD), or may be other memory.

[0067] In the above description, a bitstream generated by an encoding process is given as an example of the encoded data A0103 stored in the memory A0104. However, data other than a bitstream may also be used. For example, the encoding device A0100 may convert a bitstream into a predetermined data format and store the converted data in the memory A0104. The converted data may be, for example, a file or multiplexed stream storing one or more bitstreams. Here, the file may have a file format such as ISOBMFF (ISO Base Media File Format). The encoded data A0103 may also be in the form of multiple packets generated by dividing the bitstream or file. When converting the bitstream generated by the encoder A0102 into data other than the bitstream, the encoding device A0100 may include a conversion unit (not shown), or the conversion process may be performed by a CPU (Central Processing Unit).

[0068] <Example of Decoding Device> FIG. 3 is a functional block diagram showing the configuration of a decoding device A 0110, which is an example of a decoding device according to the present disclosure.

[0069] The memory A0114 stores, for example, the same data as the coded data A0103 generated by the coding device A0100. The memory A0114 reads out the stored data and inputs it as input data A0113 to the decoder A0112. The input data A0113 is, for example, a bitstream to be decoded. The memory A0114 may be, for example, a hard disk or an SSD, or may be some other memory.

[0070] Note that the decoding device A0110 may not directly use the data stored in the memory A0114 as the input data A0113, but may convert the read data and generate converted data as the input data A0113. The data before conversion may be, for example, multiplexed data storing one or more bitstreams. Here, the multiplexed data may be, for example, a file having a file format such as ISOBMFF. The data before conversion may also be in the form of multiple packets generated by dividing the bitstream or file. When converting data different from the bitstream read from the memory A0114 into a bitstream, the decoding device A0110 may be provided with a conversion unit (not shown), or the conversion process may be performed by a CPU.

[0071] The decoder A0112 decodes the input data A0113 to generate an audio signal A0111 that is presented to the listener.

[0072] <Another Example of Encoding Device> Fig. 4 is a functional block diagram showing the configuration of an encoding device A0120, which is another example of an encoding device of the present disclosure. In Fig. 4, components having the same functions as those in Fig. 2 are assigned the same reference numerals, and descriptions of these components will be omitted.

[0073] The encoding device A0120 differs from the encoding device A0100 in that the encoding device A0100 stores the encoded data A0103 in a memory A0104, whereas the encoding device A0120 includes a transmitting unit A0121 that transmits the encoded data A0103 to the outside.

[0074] The transmitter A0121 transmits a transmission signal A0122 to another device or a server based on the encoded data A0103 or data in another data format generated by converting the encoded data A0103. The data used to generate the transmission signal A0122 is, for example, the bit stream, multiplexed data, file, or packet described in the encoding device A0100.

[0075] <Another Example of Decoding Device> Fig. 5 is a functional block diagram showing the configuration of a decoding device A0130, which is another example of a decoding device of the present disclosure. In Fig. 5, components having the same functions as those in Fig. 3 are assigned the same reference numerals, and descriptions of these components will be omitted.

[0076] The decoding device A0130 differs from the decoding device A0110 in that the decoding device A0110 reads the input data A0113 from a memory A0114, whereas the decoding device A0130 includes a receiving unit A0131 that receives the input data A0113 from outside.

[0077] The receiving unit A0131 receives the received signal A0132, acquires the received data, and outputs the input data A0113 to be input to the decoder A0112. The received data may be the same as the input data A0113 to be input to the decoder A0112, or may be data in a different data format from the input data A0113. If the received data is data in a different data format from the input data A0113, the receiving unit A0131 may convert the received data into the input data A0113, or a conversion unit or CPU (not shown) included in the decoding device A0130 may convert the received data into the input data A0113. The received data is, for example, a bit stream, multiplexed data, a file, or a packet, as described in the encoding device A0120.

[0078] <Functional Description of Decoder> FIG. 6 is a functional block diagram showing the configuration of a decoder A0200, which is an example of the decoder A0112 in FIG. 3 or FIG.

[0079] The input data A0113 is an encoded bitstream, and includes encoded audio data, which is an encoded audio signal, and metadata used for acoustic processing.

[0080] The spatial information management unit A0201 acquires metadata included in the input data A0113 and analyzes the metadata. The metadata includes information describing elements that act on sounds arranged in a sound space. The spatial information management unit A0201 manages spatial information necessary for acoustic processing obtained by analyzing the metadata and provides the spatial information to the rendering unit A0203. Note that, although the information used for acoustic processing is referred to as spatial information in this disclosure, it may be referred to by other names. The information used for acoustic processing may be referred to, for example, as sound space information or as scene information. Furthermore, when the information used for acoustic processing changes over time, the spatial information input to the rendering unit A0203 may be referred to as a space state, a sound space state, a scene state, or the like.

[0081] Furthermore, spatial information may be managed for each sound space or for each scene. For example, when different rooms are represented as virtual spaces, the spatial information may be managed as scenes of different sound spaces for each room, or the spatial information may be managed as different scenes depending on the situation being represented, even for the same space. In managing the spatial information, an identifier for identifying each piece of spatial information may be assigned. The spatial information data may be included in a bitstream, which is one form of input data, or the bitstream may include an identifier for the spatial information, and the spatial information data may be acquired from a source other than the bitstream. When the bitstream includes only the identifier for the spatial information, the identifier for the spatial information may be used during rendering to acquire the spatial information data stored in the memory of the acoustic signal processing device A0001 or an external server as input data.

[0082] Note that the information managed by the spatial information management unit A0201 is not limited to information included in the bitstream. For example, the input data A0113 may include, as data not included in the bitstream, data indicating the characteristics or structure of a space acquired from a software application or server providing VR or AR. Furthermore, for example, the input data A0113 may include, as data not included in the bitstream, data indicating the characteristics or position of a listener or object. Furthermore, the input data A0113 may include, as information indicating the position of the listener, information acquired by a sensor provided in a terminal including a decoding device, or information indicating the position of the terminal estimated based on information acquired by the sensor. In other words, the spatial information management unit A0201 may communicate with an external system or server to acquire spatial information and the position of the listener. Furthermore, the spatial information management unit A0201 may acquire clock synchronization information from an external system and execute a process of synchronizing with the clock of the rendering unit A0203. Note that the space in the above description may be a virtually formed space, i.e., a VR space, or may be a real space (i.e., a physical space) or a virtual space corresponding to the real space, i.e., an AR or MR (Mixed Reality). The virtual space may also be called a sound field or a sound space. Furthermore, the information indicating a position in the above description may be information such as coordinate values ​​indicating a position within the space, information indicating a relative position with respect to a predetermined reference position, or information indicating the movement or acceleration of a position within the space.

[0083] The audio data decoder A0202 decodes the encoded audio data included in the input data A0113 to obtain an audio signal.

[0084] The encoded audio data acquired by the stereophonic sound reproduction system A0000 is a bitstream encoded in a predetermined format such as MPEG-H 3D Audio (ISO / IEC 23008-3). Note that MPEG-H 3D Audio is merely one example of an encoding method that can be used to generate the encoded audio data included in the bitstream, and the encoded audio data may be included in a bitstream encoded in another encoding method. For example, the encoding method used may be a lossy codec such as MP3 (MPEG-1 Audio Layer-3), AAC (Advanced Audio Coding), WMA (Windows Media Audio), AC3 (Audio Codec-3), or Vorbis, or a lossless codec such as ALAC (Apple Lossless Audio Codec) or FLAC (Free Lossless Audio Codec), or any other encoding method may be used. For example, PCM (pulse code modulation) data may be a type of encoded audio data. In this case, the decoding process may be, for example, a process of converting an N-bit binary number into a number format (e.g., floating-point format) that can be processed by the rendering unit A0203, where the number of quantization bits of the PCM data is N.

[0085] The rendering unit A0203 receives an audio signal and spatial information as input, performs acoustic processing on the audio signal using the spatial information, and outputs an audio signal A0111 after acoustic processing.

[0086] Before starting rendering, the spatial information management unit A0201 reads metadata of the input signal, detects rendering items such as objects or sounds defined in the spatial information, and sends them to the rendering unit A0203. After starting rendering, the spatial information management unit A0201 grasps changes over time in the spatial information and the position of the listener, and updates and manages the spatial information. The spatial information management unit A0201 then sends the updated spatial information to the rendering unit A0203. The rendering unit A0203 generates and outputs an audio signal to which acoustic processing has been applied based on the audio signal included in the input data A0113 and the spatial information received from the spatial information management unit A0201.

[0087] The spatial information update process and the audio signal output process with added acoustic processing may be executed in the same thread, or the spatial information management unit A0201 and the rendering unit A0203 may be assigned to independent threads. When the spatial information update process and the audio signal output process with added acoustic processing are executed in different threads, the thread startup frequency may be set individually, or the processes may be executed in parallel.

[0088] By having the spatial information management unit A0201 and the rendering unit A0203 execute their processes in different, independent threads, computational resources can be preferentially allocated to the rendering unit A0203. Therefore, in the case of sound output processing in which even the slightest delay cannot be tolerated, for example, sound output processing in which a delay of even one sample (0.02 msec) would cause a popping noise, can be safely performed. In this case, the allocation of computational resources to the spatial information management unit A0201 is limited. However, compared to audio signal output processing, updating spatial information is a less frequent process (e.g., a process such as updating the listener's facial orientation). Therefore, unlike audio signal output processing, updating spatial information does not necessarily require an instantaneous response, and therefore limiting the allocation of computational resources does not significantly affect the acoustic quality experienced by the listener.

[0089] The spatial information may be updated periodically at preset times or intervals, or when preset conditions are met. The spatial information may be updated manually by a listener or a sound space manager, or may be triggered by a change in an external system. For example, if a listener operates a controller to instantly warp the position of their avatar, instantly advance or reverse the time, or if a virtual space manager suddenly changes the environment of the venue, the thread in which the spatial information management unit A0201 is located may be activated as a one-off interrupt process in addition to being activated periodically.

[0090] The role of the information update thread that executes the spatial information update process is, for example, to update the position or orientation of the listener's avatar placed in the virtual space based on the position or orientation of the VR goggles worn by the listener, and to update the position of objects moving in the virtual space. These tasks are handled within a processing thread that runs relatively infrequently, on the order of several tens of Hz. Processing to reflect the properties of direct sound may be performed in such an infrequently occurring processing thread. This is because the properties of direct sound change less frequently than the frequency of audio processing frames for audio output. Doing so can relatively reduce the computational load of the process and also avoid the risk of pulsive noise occurring when information is updated at an unnecessarily fast frequency.

[0091] FIG. 7 is a functional block diagram showing the configuration of a decoder A0210, which is another example of the decoder A0112 in FIG. 3 or 5.

[0092] The decoder A0210 shown in Fig. 7 differs from the decoder A0200 shown in Fig. 6 in that the input data A0113 includes an unencoded audio signal rather than encoded audio data. The input data A0113 includes a bitstream including metadata and an audio signal.

[0093] The spatial information management unit A0211 is the same as the spatial information management unit A0201 in FIG. 6, and therefore a description thereof will be omitted.

[0094] The rendering unit A0213 is the same as the rendering unit A0203 in FIG. 6, and therefore a description thereof will be omitted.

[0095] 7 is referred to as a decoder A0210 in the above description, it may also be referred to as an audio processing unit that performs audio processing. Furthermore, a device including an audio processing unit may also be referred to as an audio processing device rather than a decoding device. Furthermore, the audio signal processing device A0001 may also be referred to as an audio processing device.

[0096] <Physical configuration of audio signal processing device> Fig. 8 is a diagram showing an example of the physical configuration of an audio signal processing device. Note that the audio signal processing device in Fig. 8 may be a decoding device. Furthermore, part of the configuration described here may be provided in an audio presentation device A0002. Furthermore, the audio signal processing device shown in Fig. 8 is an example of the audio signal processing device A0001 described above.

[0097] The acoustic signal processing device of FIG. 8 includes a processor, a memory, a communication IF, a sensor, and a speaker.

[0098] The processor may be, for example, a CPU (Central Processing Unit), a DSP (Digital Signal Processor), or a GPU (Graphics Processing Unit), and the CPU, DSP, or GPU may execute a program stored in a memory to perform the audio processing or decoding processing of the present disclosure. Alternatively, the processor may be a dedicated circuit that performs signal processing on an audio signal, including the audio processing of the present disclosure.

[0099] The memory may be configured, for example, by RAM (Random Access Memory) or ROM (Read Only Memory). The memory may also include a magnetic storage medium such as a hard disk or a semiconductor memory such as an SSD (Solid State Drive). The term "memory" may also refer to an internal memory built into a CPU or GPU.

[0100] The communication IF (Interface) is a communication module compatible with a communication method such as Bluetooth (registered trademark) or WIGIG (registered trademark). The acoustic signal processing device shown in Fig. 8 has a function of communicating with other communication devices via the communication IF and acquires a bitstream to be decoded. The acquired bitstream is stored in a memory, for example.

[0101] The communication module is composed of, for example, a signal processing circuit and an antenna corresponding to the communication method. In the above example, Bluetooth (registered trademark) or WIGIG (registered trademark) is used as the communication method, but other communication methods such as LTE (Long Term Evolution), NR (New Radio), or Wi-Fi (registered trademark) may also be supported. Furthermore, the communication IF may be a wired communication method such as Ethernet (registered trademark), USB (Universal Serial Bus), or HDMI (registered trademark) (High-Definition Multimedia Interface) instead of the wireless communication method described above.

[0102] The sensor performs sensing to estimate the position or orientation of the listener. Specifically, the sensor estimates the position and / or orientation of the listener based on one or more detection results of the position, orientation, movement, velocity, angular velocity, or acceleration of a part or the entire body of the listener, such as the head, and generates position information indicating the position and / or orientation of the listener. Note that the position information may be information indicating the position and / or orientation of the listener in real space, or information indicating a displacement of the position and / or orientation of the listener based on the position and / or orientation of the listener at a predetermined time. Furthermore, the position information may be information indicating the position and / or orientation relative to the stereophonic sound reproduction system A0000 or an external device equipped with the sensor.

[0103] The sensor may be, for example, an imaging device such as a camera or a ranging device such as LiDAR (Light Detection and Ranging), and may capture an image of the listener's head movement and detect the movement of the listener's head by processing the captured image. Alternatively, the sensor may be a device that performs position estimation using wireless signals of any frequency band, such as millimeter waves.

[0104] The audio signal processing device shown in Fig. 8 may acquire position information from an external device equipped with a sensor via a communication IF. In this case, the audio signal processing device may not include a sensor. Here, the external device may be, for example, the audio presentation device A0002 described in Fig. 1 or a 3D video playback device worn on the listener's head. In this case, the sensor may be a combination of various sensors, such as a gyro sensor and an acceleration sensor.

[0105] The sensor may, for example, detect the angular velocity of rotation around at least one of three mutually perpendicular axes in the sound space as the axis of rotation as the speed of movement of the listener's head, or may detect the acceleration of displacement with at least one of the three axes as the direction of displacement.

[0106] For example, the sensor may detect the amount of rotation about at least one of three mutually orthogonal axes in the sound space as the rotation axis, or the amount of displacement about at least one of the three axes as the displacement direction, as the amount of movement of the listener's head. Specifically, the sensor detects 6 DoF (position (x, y, z) and angle (yaw, pitch, roll)) as the position of the listener. The sensor is configured by combining various sensors used for detecting movement, such as a gyro sensor and an acceleration sensor.

[0107] The sensor may be any device capable of detecting the position of the listener, such as a camera or a GPS (Global Positioning System) receiver. Alternatively, the sensor may use location information obtained by performing self-position estimation using LiDAR (Laser Imaging Detection and Ranging). For example, when the audio signal reproduction system is implemented by a smartphone, the sensor is built into the smartphone.

[0108] The sensor may also include a temperature sensor such as a thermocouple that detects the temperature of the acoustic signal processing device shown in Figure 8, and a sensor that detects the remaining charge of a battery provided in or connected to the acoustic signal processing device.

[0109] A speaker has, for example, a diaphragm, a drive mechanism such as a magnet or a voice coil, and an amplifier, and presents an audio signal after acoustic processing to a listener as sound. The speaker operates the drive mechanism in response to the audio signal (more specifically, a waveform signal indicating the waveform of the sound) amplified by the amplifier, and the drive mechanism vibrates the diaphragm. In this way, the diaphragm vibrating in response to the audio signal generates sound waves, which propagate through the air and reach the listener's ears, causing the listener to perceive the sound.

[0110] Note that, although the description has been given here of an example in which the acoustic signal processing device shown in FIG. 8 includes a speaker and presents an audio signal after acoustic processing via the speaker, the means for presenting the audio signal is not limited to the above configuration. For example, the audio signal after acoustic processing may be output to an external audio presentation device A0002 connected via a communication module. Communication via the communication module may be wired or wireless. As another example, the acoustic signal processing device shown in FIG. 8 may include a terminal for outputting an analog audio signal, and a cable such as earphones may be connected to the terminal to present the audio signal from the earphones. In the above case, the audio signal is reproduced by headphones, earphones, a head-mounted display, a neck speaker, a wearable speaker, a surround speaker composed of multiple fixed speakers, or the like, which are worn on the head or part of the body of the listener, which is the audio presentation device A0002.

[0111] <Physical Configuration of Encoding Device> Fig. 9 is a diagram showing an example of the physical configuration of an encoding device. The encoding device shown in Fig. 9 is an example of the encoding devices A0100 and A0120 described above.

[0112] The encoding device of FIG. 9 includes a processor, a memory, and a communication IF.

[0113] The processor may be, for example, a CPU (Central Processing Unit) or a DSP (Digital Signal Processor), and the CPU or GPU may execute a program stored in a memory to perform the encoding process of the present disclosure. Alternatively, the processor may be a dedicated circuit that performs signal processing on an audio signal, including the encoding process of the present disclosure.

[0114] The memory may be configured, for example, by RAM (Random Access Memory) or ROM (Read Only Memory). The memory may also include a magnetic storage medium such as a hard disk or a semiconductor memory such as an SSD (Solid State Drive). The term "memory" may also refer to an internal memory built into a CPU or GPU.

[0115] The communication IF (Interface) is a communication module compatible with a communication method such as Bluetooth (registered trademark) or WIGIG (registered trademark). The encoding device has a function of communicating with other communication devices via the communication IF and transmits an encoded bitstream.

[0116] The communication module is composed of, for example, a signal processing circuit and an antenna corresponding to the communication method. In the above example, Bluetooth (registered trademark) or WIGIG (registered trademark) is used as the communication method, but other communication methods such as LTE (Long Term Evolution), NR (New Radio), or Wi-Fi (registered trademark) may also be supported. Furthermore, the communication IF may be a wired communication method such as Ethernet (registered trademark), USB (Universal Serial Bus), or HDMI (registered trademark) (High-Definition Multimedia Interface) instead of the wireless communication method described above.

[0117] [Configuration] The configuration of the acoustic signal processing device 100 according to the embodiment will now be described. Fig. 10 is a block diagram showing the functional configuration of the acoustic signal processing device 100 according to the present embodiment.

[0118] The acoustic signal processing device 100 according to this embodiment is a device for outputting aerodynamic sound data indicating aerodynamic sound caused by wind generated by an object in a virtual space (sound reproduction space). The acoustic signal processing device 100 according to this embodiment is a device that is used in various applications in virtual spaces, such as virtual reality or augmented reality (VR or AR), for example.

[0119] The object in the virtual space is included in the content (here, a video as an example) displayed on the display unit 300 that displays the content executed in the virtual space. The object is not particularly limited as long as it generates wind.

[0120] An object is, for example, a moving body that generates wind by moving its position. Moving bodies include, for example, objects that represent plants and animals, man-made objects, or natural objects. Examples of objects that represent man-made objects include vehicles, bicycles, and airplanes. Examples of objects that represent man-made objects include sports equipment such as baseball bats and tennis rackets, and furniture such as desks, chairs, and grandfather clocks. Note that, for example, an object may be at least one of an object that can move within the content and an object that can be moved, but is not limited to this.

[0121] Furthermore, for example, the object may be an object that can blow air, such as an electric fan, a circulator, a hand fan, or an air conditioner.

[0122] Aerodynamic sound according to this embodiment will now be described. Aerodynamic sound is sound that is generated when wind generated by an object reaches the ears of a listener in a virtual space.

[0123] If the object is an object that can blow air, such as an electric fan, the aerodynamic sound is the aerodynamic sound that is generated when the wind generated by the object reaches the listener. More specifically, the aerodynamic sound is the sound that is generated when the wind blown out from the electric fan reaches the listener, depending on the shape of the listener's ear, for example.

[0124] When the object is a moving body (e.g., a vehicle), the aerodynamic sound is generated when wind generated by the movement of the object reaches the listener, and more specifically, it is the sound that is generated when the wind reaches the listener, for example, depending on the shape of the listener's ear.

[0125] The object may also be an object that creates wind and generates sound. The sound generated by the object is a sound indicated by sound data associated with the object (hereinafter, sometimes referred to as object sound data). For example, if the object is a fan, the sound generated by the object is a motor sound generated by a motor possessed by the fan. For example, if the object is an ambulance, the sound generated by the object is a siren sound emitted by the ambulance.

[0126] In this embodiment, the object is an electric fan, which is an example of an object that can blow air.

[0127] The acoustic signal processing device 100 outputs aerodynamic sound data indicating aerodynamic sound in a virtual space to the headphones 200 .

[0128] Next, the headphones 200 will be described.

[0129] The headphones 200 are devices that play aerodynamic sound and are audio output devices that present the aerodynamic sound to a listener. More specifically, the headphones 200 play aerodynamic sound based on the aerodynamic sound data output by the acoustic signal processing device 100. This allows the listener to hear the aerodynamic sound. Note that instead of the headphones 200, other output channels such as speakers may be used.

[0130] As shown in FIG. 10, the headphones 200 include a head sensor unit 201 and an output unit 202 .

[0131] The head sensor unit 201 senses the position of the listener, which is determined by the horizontal coordinates and vertical height in the virtual space, and outputs second position information indicating the position of the listener of the aerodynamic sound in the virtual space to the acoustic signal processing device 100.

[0132] The head sensor unit 201 may sense 6 DoF information of the listener's head, and may be, for example, an inertial measurement unit (IMU), an accelerometer, a gyroscope, a magnetic sensor, or a combination thereof.

[0133] The output unit 202 is a device that reproduces the sound that reaches the listener in the sound reproduction space. More specifically, the output unit 202 reproduces the aerodynamic sound based on aerodynamic sound data that indicates the aerodynamic sound output from the acoustic signal processing device 100.

[0134] Furthermore, if the object is an electric fan, sound data indicating a motor sound is output from the acoustic signal processing device 100, and the output unit 202 reproduces the motor sound based on the output sound data. Similarly, if the object is an ambulance, sound data indicating a siren sound is output from the acoustic signal processing device 100, and the output unit 202 reproduces the siren sound based on the output sound data.

[0135] Next, the display unit 300 will be described.

[0136] The display unit 300 is a display device that displays content (video) including objects in a virtual space. The process by which the display unit 300 displays content will be described later. The display unit 300 is realized by a display panel such as a liquid crystal panel or an organic EL (Electro Luminescence) panel, for example.

[0137] Next, the acoustic signal processing device 100 shown in Fig. 10 will be described. In this embodiment, the acoustic signal processing device 100 outputs aerodynamic sound data to the headphones 200 after a predetermined time has elapsed from a predetermined timing.

[0138] As shown in FIG. 10, the acoustic signal processing device 100 includes an acquisition unit 110, a determination unit 120, an output unit 130, and a storage unit 140.

[0139] The acquisition unit 110 acquires object information. The object information is information indicating a change in an object that causes wind, a predetermined timing related to the change in the object, a change in the wind due to the change in the object, and a position of the object. Note that, hereinafter, the object information will be treated as information including first change information indicating a change in the object that causes wind, timing information indicating a predetermined timing related to the change in the object, second change information indicating a change in the wind due to the change in the object, and first position information indicating the position of the object.

[0140] If the object is an object that generates a sound, the object information includes sound data indicating the sound (object sound data). The object information may also include shape information indicating the shape of the object.

[0141] The acquisition unit 110 acquires second position information. As described above, the second position information is information that indicates the position of the listener in the virtual space. The acquisition unit 110 acquires aerodynamic sound data that indicates aerodynamic sound. The aerodynamic sound data is stored in the storage unit 140, and the acquisition unit 110 acquires the aerodynamic sound data stored in the storage unit 140.

[0142] The acquisition unit 110 may acquire the object information, second position information, and aerodynamic sound data from, for example, an input signal, or may acquire the object information, second position information, and aerodynamic sound data from other sources. The input signal will be described below. In the following, the object sound data and aerodynamic sound data may be collectively referred to as sound data.

[0143] The input signal may be composed of, for example, spatial information, sensor information, and sound data (audio signal). The above information and sound data may be included in a single input signal, or may be included in multiple separate signals. The input signal may include a bitstream composed of sound data and metadata (control information), in which case the metadata may include information identifying the spatial information and sound data.

[0144] The first change information, timing information, second change information, first position information, shape information, object sound data, second position information, and aerodynamic sound data described above may be included in the input signal. More specifically, the first change information, timing information, second change information, first position information, and shape information may be included in spatial information, and the second position information may be generated based on information acquired from sensor information. The sensor information may be acquired from the head sensor unit 201 or from another external device.

[0145] The spatial information is information about the sound space (three-dimensional sound field) created by the stereophonic sound reproduction system A0000, and is composed of information about objects included in the sound space and information about the listener. Objects include sound source objects that emit sound and act as sound sources, and non-sound-emitting objects that do not emit sound. Non-sound-emitting objects function as obstacle objects that reflect sounds emitted by sound source objects, but sound source objects may also function as obstacle objects that reflect sounds emitted by other sound source objects. Obstacle objects may also be called reflecting objects.

[0146] Information commonly assigned to sound source objects and non-sound-producing objects includes position information, shape information, and the rate of attenuation of the volume when the object reflects sound.

[0147] The position information is expressed as coordinate values ​​on three axes, for example, the X-axis, Y-axis, and Z-axis, in Euclidean space, but does not necessarily have to be three-dimensional information. The position information may also be two-dimensional information expressed as coordinate values ​​on two axes, for example, the X-axis and Y-axis. The position information of an object is determined by a representative position of a shape expressed by a mesh or voxels.

[0148] The shape information may include information about the surface material.

[0149] The attenuation rate may be expressed as a real number equal to or less than 1 or equal to or greater than 0, or may be expressed as a negative decibel value. Since sound volume is not amplified by reflection in real space, a negative decibel value is set for the attenuation rate. However, for example, to create an eerie feeling in an unreal space, an attenuation rate of equal to or greater than 1, i.e., a positive decibel value, may be set. Furthermore, different values ​​for the attenuation rate may be set for each frequency band constituting multiple frequency bands, or values ​​may be set independently for each frequency band. Furthermore, if an attenuation rate is set for each type of material on the object surface, a corresponding attenuation rate value may be used based on information about the surface material.

[0150] Furthermore, the information commonly assigned to the sound source object and the non-sound-producing object may include information indicating whether the object belongs to a living thing or whether the object is a moving object, etc. If the object is a moving object, the position information may move over time, and the changed position information or the amount of change is transmitted to the rendering units A0203 and A0213.

[0151] The information about the sound source object includes, in addition to the information commonly assigned to the sound source object and the non-sound generating object, object sound data and information necessary for radiating the object sound data into the sound space. The object sound data is data representing the sound perceived by the listener, including information about the frequency and intensity of the sound. The object sound data is typically a PCM signal, but may also be data compressed using an encoding method such as MP3. In this case, the signal must be decoded at least before reaching the generation unit (the generation unit 907 described later in FIG. 19 ). Therefore, the rendering units A0203 and A0213 may include a decoding unit (not shown). Alternatively, the signal may be decoded by the audio data decoder A0202.

[0152] At least one object sound data may be set for one sound source object, and multiple object sound data may be set for one sound source object. Furthermore, identification information for identifying each object sound data may be assigned, and the identification information for the object sound data may be stored as metadata as information about the sound source object.

[0153] The information necessary for radiating object sound data into the sound space may include, for example, information on the reference volume that serves as a reference when playing the object sound data, information on the position of the sound source object, information on the orientation of the sound source object, and information on the directionality of the sound emitted by the sound source object.

[0154] The reference volume information may be, for example, the effective value of the amplitude value of the object sound data at the sound source position when the object sound data is radiated into the sound space, and may be expressed as a floating-point decibel (dB) value. For example, when the reference volume is 0 dB, the reference volume information may indicate that sound is to be radiated into the sound space from the position indicated by the information regarding the position at the same volume without increasing or decreasing the volume of the signal level indicated by the object sound data. When the reference volume is −6 dB, the reference volume information may indicate that sound is to be radiated into the sound space from the position indicated by the information regarding the position with the volume of the signal level indicated by the object sound data reduced to about half. The reference volume information may be assigned to one object sound data or to multiple object sound data collectively.

[0155] The volume information included in the information necessary to radiate object sound data into a sound space may include, for example, information indicating time-series fluctuations in the volume of the sound source. For example, if the sound space is a virtual conference room and the sound source is a speaker, the volume transitions intermittently over a short period of time. Expressed more simply, this can be said to mean that sound portions and silent portions occur alternately. Furthermore, if the sound space is a concert hall and the sound source is a performer, the volume is maintained for a certain period of time. Furthermore, if the sound space is a battlefield and the sound source is an explosive, the volume of the explosion sound increases for only a moment and then remains silent. In this way, the volume information of the sound source includes not only information about the volume of the sound but also information about the transition of the volume of the sound, and such information may be used as information indicating the properties of the object sound data.

[0156] Here, the information on loudness transitions may be data indicating frequency characteristics in a time series. The information on loudness transitions may be data indicating the duration of a section in which sound is present. The information on loudness transitions may be data indicating a time series of the duration of a section in which sound is present and the duration of a section in which sound is absent. The information on loudness transitions may be data listing, in a time series, multiple sets of durations during which the amplitude of a sound signal can be considered steady (considered to be roughly constant) and data on the amplitude values ​​of the signal during those periods. The information on loudness transitions may be data indicating the durations during which the frequency characteristics of a sound signal can be considered steady. The information on loudness transitions may be data listing, in a time series, multiple sets of durations during which the frequency characteristics of a sound signal can be considered steady and data on the frequency characteristics during those periods. The data format of the information on loudness transitions may be, for example, data indicating the outline of a spectrogram. Furthermore, the volume that serves as a reference for the frequency characteristics may be the reference volume. The information on the reference volume and the information indicating the properties of the object sound data may be used to calculate the volume of the direct sound or reflected sound to be perceived by the listener, as well as in a selection process to select whether or not to perceive it.

[0157] Orientation information is typically expressed using yaw, pitch, and roll. Alternatively, the roll rotation may be omitted and the information may be expressed using azimuth (yaw) and elevation (pitch). Orientation information may change over time, and if it does change, it is transmitted to the rendering units A0203 and A0213.

[0158] The information about the listener is information about the position and orientation of the listener in sound space. The position information is expressed as positions on the X-, Y-, and Z-axes in Euclidean space, but does not necessarily have to be three-dimensional information and may be two-dimensional information. The information about orientation is typically expressed using yaw, pitch, and roll. Alternatively, the information about orientation may be expressed using azimuth (yaw) and elevation (pitch) without the roll rotation. The position information and orientation information may change over time, and if they change, they are transmitted to the rendering units A0203 and A0213.

[0159] The sensor information includes the amount of rotation or displacement detected by a sensor worn by the listener, as well as the position and orientation of the listener. The sensor information is transmitted to the rendering units A0203 and A0213, which update the position and orientation information of the listener based on the sensor information. The sensor information may be, for example, position information obtained by a mobile terminal performing self-position estimation using a GPS, a camera, or LiDAR (Laser Imaging Detection and Ranging). Information obtained from an external source other than a sensor via a communication module may also be detected as sensor information. Information indicating the temperature of the acoustic signal processing device 100 and information indicating the remaining battery capacity may be acquired from the sensor. Information indicating the computing resources (CPU capacity, memory resources, PC performance) of the acoustic signal processing device 100 or the audio presentation device A0002 may be acquired in real time as sensor information.

[0160] In the present embodiment, the acquiring unit 110 acquires the object information from the storage unit 140, but is not limited to this, and may acquire the object information from a device other than the acoustic signal processing device 100 (for example, a server device 500 such as a cloud server). Furthermore, the acquiring unit 110 acquires the second position information from the headphones 200 (more specifically, the head sensor unit 201), but is not limited to this.

[0161] Here, the information contained in the object information will be described.

[0162] First, the first change information will be described.

[0163] The first change information is information that indicates a change in the object that causes wind. In this embodiment, the change in the object means a change in the state of the object. In this example, since the object is an electric fan, the following examples of changes in the state of the object are given.

[0164] For example, a change in the state of an object is when an electric fan is switched between ON and OFF (hereinafter, this may be referred to as an "ON / OFF switch"). Another example of a change in the state of an object is when a switch that specifies the fan's wind speed is switched from low to high (hereinafter, this may be referred to as an "wind speed switch"). Another example of a change in the state of an object is when a switch that specifies the fan's oscillation is switched from no oscillation to oscillation (hereinafter, this may be referred to as an "wind direction switch").

[0165] The second change information will now be described.

[0166] The second change information is information that indicates a change in the wind due to a change in the object. The second change information indicates a change in the wind speed or a change in the wind direction (wind direction) as a change in the wind due to a change in the object. In this embodiment, the content of the information indicated by the second change information changes in accordance with the change in the state of the object indicated by the first change information.

[0167] When the change in the state of the object indicated by the first change information is an "ON / OFF change," the second change information indicates, for example, that the wind speed has changed from 0 m / s to V1 m / s (V1 > 0). Furthermore, when the change in the state of the object indicated by the first change information is a "wind speed change," the second change information indicates, for example, that the wind speed has changed from V2 m / s to V3 m / s (V3 > V2). Furthermore, when the change in the state of the object indicated by the first change information is a "wind direction change," the second change information indicates, for example, that the wind direction has changed from a constant state to a changing state. In this way, the second change information may be information that depends on the first change information.

[0168] Note that the wind speeds V1, V2, and V3 are, for example, the wind speeds at the position where the electric fan, which is an object, is placed.

[0169] Next, the timing information will be described.

[0170] The timing information is information that indicates a predetermined timing related to a change in an object. As described above, the acoustic signal processing device 100 outputs aerodynamic sound data to the headphones 200 a predetermined time after this predetermined timing. The predetermined timing indicates the start of the predetermined time for outputting the aerodynamic sound data.

[0171] The predetermined timing indicated by the timing information is the timing of a change in wind, more specifically, the timing of a change in wind due to a change in object. For example, the predetermined timing is the timing of a change in wind speed or direction due to a change in object.

[0172] Furthermore, a case where the predetermined timing is the timing when the wind speed changes will be described.

[0173] An example of a change in wind speed is when a fan, which is an object, is switched from OFF to ON. In this case, for example, the wind speed changes from 0 m / s to V1 m / s, and the predetermined timing is the timing at which the wind speed changes, that is, the timing at which the wind speed changes from 0 m / s to V1 m / s. When the fan is switched from OFF to ON, as described above, the fan generates a motor sound. Therefore, in this case, the predetermined timing is the timing at which the wind speed changes and also the timing (first timing) at which sound data (object sound data) associated with the fan, which is an object, is output. In other words, the acoustic signal processing device 100 (more specifically, the output unit 130) according to this embodiment outputs sound data (object sound data) associated with the fan at the predetermined timing (first timing). The timing information included in the object information indicates that the predetermined timing is the timing of a change in wind and also the first timing.

[0174] The predetermined timing may be a timing designated by an administrator of the acoustic signal processing device 100, for example.

[0175] The first location information will be further described.

[0176] As described above, the object in the virtual space is included in the content (video) displayed on the display unit 300, and in this embodiment is an electric fan.

[0177] The first position information is information indicating the position of the electric fan in the virtual space at a given time. Note that in the virtual space, for example, the electric fan may be moved by a user picking up the electric fan and moving it. Therefore, the acquisition unit 110 continuously acquires the first position information. The acquisition unit 110 acquires the first position information, for example, each time the space information is updated by the space information management units A0201 and A0211.

[0178] Furthermore, object sound data associated with an object and sound data including aerodynamic sound data will be described.

[0179] The sound data including the object sound data and aerodynamic sound data described in this specification may be a sound signal such as PCM (Pulse Code Modulation) data, but is not limited to this and may be any information that indicates the properties of the sound.

[0180] As an example, if a sound signal is a noise signal with a volume of X decibels, the sound data related to the sound signal may be the PCM data itself representing the sound signal, or may be data consisting of information indicating that the component is a noise signal and information indicating that the volume is X decibels. As another example, if a sound signal is a noise signal with a frequency component having a predetermined peak / dip characteristic, the sound data related to the sound data may be the PCM data itself representing the sound signal, or may be data consisting of information indicating that the component is a noise signal and information indicating the peak / dip of the frequency component.

[0181] In this specification, a sound signal based on sound data means PCM data representing the sound data.

[0182] As described above, the aerodynamic sound data is stored in advance in storage unit 140. Aerodynamic sound data is data obtained by collecting sounds that occur when wind reaches a human ear or a model that imitates a human ear. In this embodiment, the aerodynamic sound data is data obtained by collecting sounds that occur when wind reaches a model that imitates a human ear. A dummy head microphone or the like is used as the model that imitates a human ear, and the aerodynamic sound data is collected.

[0183] Furthermore, as described above, in this embodiment, the wind changes as the object changes. The aerodynamic sound is the aerodynamic sound caused by the wind before the change or the wind after the change. The aerodynamic sound may be the aerodynamic sound caused by the wind after the change, for example, the aerodynamic sound caused by the wind at the changed wind speed or the aerodynamic sound caused by the wind in the changed wind direction.

[0184] Next, the shape information will be described.

[0185] Shape information is information that indicates the shape of an object in virtual space. The shape information indicates the shape of the object, and more specifically, indicates the three-dimensional shape of the object as a rigid body. The shape of the object is indicated by, for example, a sphere, a rectangular parallelepiped, a cube, a polyhedron, a cone, a pyramid, a cylinder, a prism, or a combination thereof. Note that the shape information may be expressed as, for example, mesh data, or a set of multiple faces each consisting of a voxel, a three-dimensional point cloud, or vertices with three-dimensional coordinates.

[0186] The first change information includes object identification information for identifying an object. The timing information also includes object identification information, the second change information also includes object identification information, the first position information also includes object identification information, the object sound data also includes object identification information, and the shape information also includes object identification information.

[0187] Therefore, even if the acquisition unit 110 acquires the first change information, timing information, second change information, first position information, object sound data, and shape information separately, the object indicated by each of the first change information, timing information, second change information, first position information, object sound data, and shape information can be identified by referencing the object identification information included in each of the first change information, timing information, second change information, first position information, object sound data, and shape information. For example, in this example, it is easy to identify that the object indicated by each of the first change information, timing information, second change information, first position information, object sound data, and shape information is the same electric fan. In other words, by referencing the six pieces of object identification information, it becomes clear that the first change information, timing information, second change information, first position information, object sound data, and shape information acquired by the acquisition unit 110 are information related to an electric fan. Therefore, the first change information, the timing information, the second change information, the first position information, the object sound data, and the shape information are linked as information indicating the electric fan.

[0188] Next, the second location information will be described.

[0189] The listener may move in the virtual space. The second position information is information indicating the position in the virtual space where the listener is located at a given time. Since the listener may move in the virtual space, the acquisition unit 110 continuously acquires the second position information. The acquisition unit 110 acquires the second position information, for example, each time the spatial information is updated by the spatial information management units A0201 and A0211.

[0190] The first change information, timing information, second change information, first position information, shape information, object sound data, second position information, and aerodynamic sound data may be included in metadata, control information, or header information included in the input signal. When sound data including object sound data and aerodynamic sound data is a sound signal (PCM data), information identifying the sound signal may be included in the metadata, control information, or header information, or the sound signal may be included outside the metadata, control information, or header information. In other words, the acoustic signal processing device 100 (more specifically, the acquisition unit 110) may acquire metadata, control information, or header information included in the input signal and perform acoustic processing based on the metadata, control information, or header information. The acoustic signal processing device 100 (more specifically, the acquisition unit 110) may acquire the first change information, timing information, second change information, first position information, shape information, object sound data, second position information, and aerodynamic sound data, and the source of acquisition is not limited to the input signal. The sound data including the object sound data and the aerodynamic sound data and the metadata may be stored in one input signal, or may be stored separately in a plurality of input signals.

[0191] Furthermore, sound signals other than sound data including object sound data and aerodynamic sound data may be stored in the input signal as audio content information. The audio content information may be encoded using MPEG-H 3D Audio (ISO / IEC 23008-3) (hereinafter referred to as MPEG-H 3D Audio) or other encoding technology. The encoding technology is not limited to MPEG-H 3D Audio, and other well-known technologies may also be used. Information such as the above-mentioned first change information, timing information, second change information, first position information, shape information, object sound data, second position information, and aerodynamic sound data may also be subject to encoding processing.

[0192] In other words, the acoustic signal processing device 100 acquires sound signals and metadata contained in an encoded bitstream. Audio content information is acquired and decoded in the acoustic signal processing device 100. In this embodiment, the acoustic signal processing device 100 functions as a decoder (e.g., decoders A0200 and A0210) included in a decoding device (e.g., decoding devices A0110 and A0130), and more specifically, functions as rendering units A0203 and A0213 included in the decoder. Note that the term "audio content information" in the present disclosure is to be interpreted as meaning information including the sound signal itself, first change information, timing information, second change information, first position information, shape information, object sound data, second position information, and aerodynamic sound data, depending on the technical content.

[0193] The acquisition unit 110 outputs the acquired object information and second position information to the determination unit 120 and the output unit 130 .

[0194] The determination unit 120 determines the predetermined time based on the wind indicated by the object information acquired by the acquisition unit 110. That is, the determination unit 120 determines the predetermined time based on the wind caused by the object.

[0195] For example, the determination unit 120 determines the predetermined time based on the wind speed indicated by the second change information included in the acquired object information and the distance between the position of the listener and the position of the object. If the predetermined time is t seconds, then, for example, t>0 is satisfied, but this is not limited thereto, and the predetermined time may be, for example, 0.1 seconds or more and 5 seconds or less. The determination unit 120 can determine, for example, a time specified by an administrator of the acoustic signal processing device 100 as the predetermined time. Furthermore, the determination unit 120 calculates the distance as follows.

[0196] The determination unit 120 calculates the distance between the position of the listener and the position of the object based on the first position information included in the object information acquired by the acquisition unit 110 and the acquired second position information. As described above, the acquisition unit 110 acquires the first position information and the second position information in the virtual space each time the space information is updated by the space information management units A0201 and A0211. The determination unit 120 calculates the distance between the position of the listener and the position of the object in the virtual space based on the multiple pieces of first position information and the multiple pieces of second position information acquired each time the space information is updated.

[0197] The determination unit 120 determines the predetermined time and outputs it to the output unit 130 .

[0198] The output unit 130 outputs the aerodynamic sound data acquired by the acquisition unit 110 a predetermined time determined by the determination unit 120 from the predetermined timing indicated by the object information acquired by the acquisition unit 110. Here, the output unit 130 outputs the aerodynamic sound data to the headphones 200. This allows the headphones 200 to play back the aerodynamic sound indicated by the output aerodynamic sound data. In other words, the listener can hear the aerodynamic sound a predetermined time after the predetermined timing.

[0199] The storage unit 140 is a storage device that stores computer programs executed by the acquisition unit 110, the determination unit 120, and the output unit 130, object information, and aerodynamic sound data.

[0200] Here, the shape information according to the present embodiment will be explained again. The shape information is information used to generate an image of an object in a virtual space and is also information indicating the shape of the object (electric fan). In other words, the shape information is also information used to generate content (image) to be displayed on the display unit 300.

[0201] The acquisition unit 110 also outputs the acquired shape information to the display unit 300. The display unit 300 acquires the shape information output by the acquisition unit 110. The display unit 300 further acquires attribute information indicating attributes (such as color) of the object (electric fan) other than the shape in the virtual space. The display unit 300 may acquire the attribute information directly from a device other than the acoustic signal processing device 100 (the server device 500), or may acquire the attribute information from the acoustic signal processing device 100. The display unit 300 generates and displays content (video) based on the acquired shape information and attribute information.

[0202] Hereinafter, a first operational example of the acoustic signal processing method performed by the acoustic signal processing device 100 will be described.

[0203] 11 is a flowchart of an operation example 1 of the acoustic signal processing device 100 according to this embodiment. FIG. 12 is a diagram showing an electric fan F and a listener L, which are objects according to the operation example 1.

[0204] 11 , first, the acquisition unit 110 acquires object information (S10). As described above, the object information includes first change information indicating a change in the object causing the wind W, timing information indicating a predetermined timing for the change in the object, second change information indicating a change in the wind W due to the change in the object, and first position information indicating the position of the object. The object information also includes object sound data indicating a motor sound and shape information. Step S10 corresponds to the acquisition step.

[0205] Here, the second change information indicates a change in the wind W due to a change in the object, that is, a change in the wind speed of the wind W. The predetermined timing indicated by the timing information is the timing of the change in the wind W, more specifically, the timing of the change in the wind W due to a change in the object.

[0206] Next, the acquisition unit 110 acquires second position information indicating the position of the listener L in the virtual space from the headphones 200 (S20). Furthermore, the acquisition unit 110 acquires aerodynamic sound data indicating aerodynamic sound stored in the storage unit 140 (S30).

[0207] Next, the determination unit 120 determines the predetermined time based on the wind speed indicated by the second change information and the distance between the position of the listener L and the position of the object (electric fan F) (S40). This step S40 corresponds to the determination step.

[0208] Furthermore, the output unit 130 outputs sound data (object sound data) associated with the electric fan F at a predetermined timing (S50). Then, after a predetermined time has elapsed since the predetermined timing, the output unit 130 outputs aerodynamic sound data indicating aerodynamic sound caused by the wind W (S60). This step S60 corresponds to the output step.

[0209] Here, the predetermined timing and predetermined time in this operation example will be explained.

[0210] Here, the predetermined timing is the timing of a change in the wind W, that is, the timing of a change in the wind speed due to a change in the object. As an example, when the listener L is watching content in which an electric fan F is displayed on the display unit 300, the predetermined timing is the timing when the electric fan F is switched from OFF to ON.

[0211] In real space, the listener L hears the aerodynamic sound when the time has passed from the time when the electric fan F is switched from OFF to ON (i.e., the predetermined time) until the wind W generated by the electric fan F reaches the listener L. Therefore, the determination unit 120 may determine the predetermined time as the time from the predetermined time until the wind W generated by the electric fan F reaches the listener L.

[0212] FIG. 13A is a diagram illustrating the process of determining the predetermined time in step S40 shown in FIG.

[0213] The distance between the position of the listener L and the position of the object (electric fan F) is defined as D. More specifically, the distance between the position of the listener L's ear and the position of the object (electric fan F) is defined as D. Note that the distance D is calculated by the determination unit 120 based on the first position information included in the object information acquired by the acquisition unit 110 and the acquired second position information.

[0214] Let U be the distance from the position of the object (electric fan F) at which the wind speed of the wind W generated by the object, electric fan F, is So. Furthermore, the direction from the electric fan F toward the listener L is the x-axis direction, and the distance from the electric fan F in the x-axis direction is x. Because the wind speed V of the wind W is inversely proportional to the distance x, the wind speed V and the distance x satisfy the following equation.

[0215] V = So×(U / x)

[0216] The average wind speed up to a position at a distance D satisfies the following formula.

[0217]

[0218] t, which is the time (predetermined time) from when the fan F is switched from OFF to ON (i.e., the predetermined time) until the wind W generated by the object, the fan F, reaches the listener L, is the value obtained by dividing the distance by the average wind speed and satisfies the following formula.

[0219] t = {(D-U)^2} / {So×U×(log(D)-log(U))

[0220] In the above formula, "^" represents an operator for finding a power.

[0221] Then, as described above, in step S60, aerodynamic sound data is output at a timing when a predetermined time t has elapsed from the predetermined timing.

[0222] As a result, the listener L can hear the aerodynamic sound output from the headphones 200 at the timing when the fan F is switched from OFF to ON (i.e., the predetermined timing) and the time it takes for the wind W created by the fan F to reach the listener L (predetermined time t) has elapsed. Therefore, the listener L can hear the aerodynamic sound at the same timing as in real space, that is, at the appropriate timing, so the listener L is less likely to feel uncomfortable and can obtain a sense of presence.

[0223] Furthermore, in this operation example, the predetermined timing is the timing when the electric fan F is switched from OFF to ON, which is the first timing when the object sound data associated with the electric fan F, which is an object, is output.

[0224] It goes without saying that the above operation includes the following meaning. That is, this meaning means that "from a predetermined timing until a predetermined time t has elapsed, the aerodynamic sound indicated by the aerodynamic sound data is output so as to become a sound with an amplitude that can be perceived by the listener L." This is realized, for example, when outputting the aerodynamic sound data, by a filter with the predetermined time t as its time constant. Specifically, it may be done as follows.

[0225] Fig. 13B is a diagram illustrating a detailed example of the output of aerodynamic sound data according to this embodiment. Fig. 13C is a diagram illustrating another detailed example of the output of aerodynamic sound data according to this embodiment.

[0226] (a) of Fig. 13B is a diagram showing a trigger signal that indicates the ON / OFF change of the electric fan F. (a) of Fig. 13B shows a trigger signal that has a value of "0" when the electric fan F is OFF and a value of "1" when the electric fan F is ON. (b) of Fig. 13B is a diagram showing the trigger signal multiplied by a time constant t. In other words, the trigger signal is filtered through a low-pass filter whose time constant is a predetermined time t. (c) of Fig. 13B is a diagram showing aerodynamic sound data whose amplitude has been amplified in accordance with the magnitude of the output signal of the low-pass filter.

[0227] This makes it possible to very simply simulate the operation of outputting aerodynamic sound data after a predetermined time t has elapsed, and also automatically simulate the operation when the cause of the aerodynamic sound disappears (the operation when the electric fan F changes from ON to OFF).

[0228] Here, t does not necessarily have to be a value calculated accurately based on the following formula, but may be a value simply approximated so that t increases as the distance D increases.

[0229] t = {(D-U)^2} / {So×U×(log(D)-log(U))

[0230] In the above formula, "^" represents an operator for finding a power.

[0231] (a) of Fig. 13C, like (a) of Fig. 13B, is a diagram showing a trigger signal that indicates the ON / OFF change of the electric fan F. (b) of Fig. 13C, like (b) of Fig. 13B, is a diagram showing the above-mentioned trigger signal multiplied by a time constant t, and shows a trigger signal multiplied by a time constant t that is smaller than the time constant t in (b) of Fig. 13B. (c) of Fig. 13C is a diagram showing aerodynamic sound data controlled in accordance with the value of the trigger signal multiplied by the time constant t shown in (b) of Fig. 13C.

[0232] As described above, the predetermined timing is the timing when the electric fan F is switched from OFF to ON, and is the first timing when the object sound data associated with the electric fan F, which is an object, is output.

[0233] Therefore, by the processing of step S50, at the timing when the electric fan F is switched from OFF to ON, the listener L can hear the motor sound of the electric fan F output from the headphones 200. Furthermore, by the processing of step S60, at the timing when the time has passed since the listener L heard the motor sound, for the wind W caused by the electric fan F being switched from OFF to ON to reach the listener L, the listener L can hear the aerodynamic sound output from the headphones 200.

[0234] In real space, the motor sound reaches the listener L at the speed of sound and is heard by the listener L, and the aerodynamic sound is heard by the listener L when the wind W reaches the listener L. In real space, the speed of sound is generally faster than the wind speed, and in this operation example, as in real space, the listener L hears the motor sound first and then the aerodynamic sound. Therefore, the listener L can hear the motor sound (sound represented by sound data associated with the object) and the aerodynamic sound at the same timing as in real space, that is, at the appropriate timing, so the listener L is less likely to feel uncomfortable and can experience a sense of realism.

[0235] In operation example 1, the specified timing is the timing when the wind speed changes and the timing (first timing) when sound data (object sound data) associated with the object, the fan F, is output, but this is not limited to this.

[0236] For example, the object information may indicate a change in the direction of the wind W due to a change in the object (electric fan F). More specifically, the object information may indicate a change in the direction (wind direction) of the wind W as a change in the wind W due to a change in the object (electric fan F). This case is, for example, a case in which the change in the state of the object indicated by the first change information is a "wind direction change," and the second change information indicates that the wind direction has changed from a constant state to a changing state.

[0237] In this case, the timing information included in the object information indicates that the predetermined timing is the third timing at which a change in the direction of the wind W (wind direction) occurs.

[0238] In this way, when the wind direction of the electric fan F changes, the state of the wind W reaching the listener L changes, and therefore the aerodynamic sound heard by the listener L also changes. For this reason, in step S60 shown in Fig. 11, the output unit 130 may output aerodynamic sound data indicating the aerodynamic sound caused by the wind W a predetermined time after the third timing (predetermined timing) indicated by the object information.

[0239] Furthermore, the predetermined timing and predetermined time are not limited to those shown in Operation Example 1. The predetermined timing may be a timing (specified timing) specified by a user (for example, an administrator of the acoustic signal processing device 100), and the predetermined time may be a time (specified time) specified by the administrator. The determination unit 120 may determine the timing and time specified by the user as the predetermined timing and predetermined time. For example, the acoustic signal processing device 100 may include a reception unit, which receives the timing and time specified by the user, and the determination unit 120 may determine the timing and time received by the reception unit as the predetermined timing and predetermined time. In this case, the administrator specifies the specified timing and time so that the listener L can hear the aerodynamic sound at the same timing as in real space.

[0240] Even in this case, the listener L can hear the aerodynamic sound at the same timing as in the real space, that is, at the appropriate timing, so the listener L is less likely to feel uncomfortable and can get a sense of realism.

[0241] Furthermore, in the first operational example of the embodiment, the aerodynamic sound data is stored in advance in the storage unit 140, but this is not limited to this. For example, the determination unit 120 may generate the aerodynamic sound data. For example, the determination unit 120 may generate the aerodynamic sound data by acquiring a noise signal and processing the acquired noise signal with each of a plurality of band emphasis filters.

[0242] Furthermore, in the first operational example of the embodiment, the determiner 120 determined the predetermined time based on the wind speed indicated by the second change information and the distance between the position of the listener L and the position of the object (electric fan F), but this is not limited to this. For example, the object information may include first position information indicating the position of the object, and the determiner 120 may determine the predetermined time based on the distance between the position of the listener L of the aerodynamic sound and the position of the object indicated by the first position information included in the acquired object information. For example, it is preferable that a predetermined time corresponding to a reference distance is set, and that the predetermined time is determined so that the longer the distance between the position of the listener L of the aerodynamic sound and the position of the object is, the longer the predetermined time is, and so that the shorter the distance between the position of the listener L of the aerodynamic sound and the position of the object is, the shorter the predetermined time is.

[0243] (Modifications of the embodiment) Modifications of the embodiment will be described below, focusing on differences from the embodiment, and explanations of commonalities will be omitted or simplified.

[0244] [Configuration] In this modification, the acoustic signal processing device 100 according to the embodiment is used, but the object in the virtual space is different. The object according to this modification is a vehicle, which is a moving body. More specifically, the object is an ambulance. In this case, the aerodynamic sound is a sound that is generated when wind W, which is generated by the movement of the object's position, reaches the listener L. Furthermore, the object, the ambulance, is an object that generates sound, and generates a siren sound.

[0245] The object information according to this modification is information indicating a change in the object causing the wind W, a predetermined timing for the change in the object, a change in the wind W due to the change in the object, and a position of the object. As in the embodiment, the object information is handled as information including first change information indicating a change in the object causing the wind W, timing information indicating a predetermined timing for the change in the object, second change information indicating a change in the wind W due to the change in the object, and first position information indicating the position of the object.

[0246] The first change information is information that indicates a change in the object that causes the wind W, and in this modified example, the change in the object means a change in the position of the object.

[0247] The first location information is information indicating the location of the ambulance in the virtual space at a given time. Note that in the virtual space, the ambulance may travel and move as it is operated by a driver, for example. For this reason, the acquisition unit 110 continuously acquires the first location information.

[0248] The second change information is information that indicates a change in the wind W due to a change in the object. In the present embodiment, the content of the information indicated by the second change information changes according to a change in the position of the object indicated by the first change information.

[0249] For example, when the first change information indicates that the position of the object has changed, the second change information indicates that the wind speed of the wind W generated by the movement of the object has changed from a first predetermined value to a second predetermined value or that the wind direction has changed from a first predetermined direction to a second predetermined direction. Note that the first and second predetermined values ​​are, for example, the wind speed at the position where the ambulance is located, and the first and second predetermined directions are, for example, the wind direction at the position where the ambulance is located.

[0250] As a more specific example, a case will be described in which the first change information indicates that an ambulance approaches and then moves away from the listener L. In this case, the wind W generated by the movement of the ambulance blows strongly toward the listener L while the ambulance is approaching the listener L, and blows weakly toward the listener L while the ambulance is moving away from the listener L. Therefore, the wind speed of the wind W is high toward the listener L while the ambulance is approaching the listener L, and low toward the listener L while the ambulance is moving away from the listener L. In this way, the wind W (more specifically, the wind speed of the wind W) is changing.

[0251] In this modification, the speed of the wind W generated by the object, an ambulance, is considered to be the same as the moving speed of the ambulance, which is calculated by differentiating the position of the ambulance with respect to time in the virtual space based on the first position information.

[0252] Next, the timing information will be described.

[0253] The timing information is information indicating a predetermined timing related to a change in the object. The predetermined timing indicated by the timing information is the timing of a change in the wind W, more specifically, the timing of a change in the wind W due to a change in the position of the object. For example, the predetermined timing is the timing when the wind speed changes due to a change in the position of the object, such as the timing when an ambulance approaches the listener L and then moves away from the listener L. In this case, the predetermined timing is the timing when the amount of change in the distance between the position of the listener L and the position of the object in the virtual space changes from negative to positive over time. In other words, this predetermined timing is the timing when the object is closest to the listener L in the virtual space. Alternatively, for example, the predetermined timing may be the timing when the wind direction changes due to a change in the position of the object.

[0254] A second example of the operation of the acoustic signal processing method performed by the acoustic signal processing device 100 will now be described.

[0255] 14 is a flowchart of an operation example 2 of the acoustic signal processing device 100 according to this embodiment. FIG. 15 is a diagram showing an ambulance A and a listener L, which are objects according to the operation example 2.

[0256] 14, first, the acquisition unit 110 acquires object information (S10). As described above, the object information includes first change information indicating a change in the object causing the wind W, timing information indicating a predetermined timing for the change in the object, second change information indicating a change in the wind W due to the change in the object, and first position information indicating the position of the object. The object information also includes object sound data indicating a siren sound and shape information.

[0257] Here, the second change information indicates a change in the wind W due to a change in the object, that is, a change in the wind speed of the wind W. The predetermined timing indicated by the timing information is the timing of the change in the wind W, more specifically, the timing of the change in the wind W due to a change in the object.

[0258] Next, the acquisition unit 110 acquires second position information indicating the position of the listener L in the virtual space from the headphones 200 (S20). Furthermore, the acquisition unit 110 acquires aerodynamic sound data indicating aerodynamic sound stored in the storage unit 140 (S30).

[0259] Furthermore, the output unit 130 determines whether or not a predetermined timing has arrived (S35). If the predetermined timing has not arrived (No in step S35), the process of step S35 is repeated.

[0260] If it is the specified timing (Yes in step S35), the determination unit 120 determines the specified time based on the wind speed indicated by the second change information and the distance between the position of the listener L and the position of the object (ambulance A) (S40).

[0261] Then, the output unit 130 outputs aerodynamic sound data indicating aerodynamic sound caused by the wind W after a predetermined time has elapsed since the predetermined timing (S60).

[0262] The predetermined timing and the process of step S35 in this operation example will now be described in more detail.

[0263] In this operation example, the predetermined timing is the timing of a change in the wind W. More specifically, the predetermined timing is the timing when the wind speed changes due to a change in the position of the object, and is the timing when the amount of change in the distance between the position of the listener L and the position of the object in the virtual space turns from negative to positive over time.

[0264] FIG. 16 is a schematic diagram for explaining predetermined timing according to the second operation example.

[0265] Ambulance A moves in the order of (a), (b), and (c) shown in FIG. 16. Furthermore, it is assumed that the position of listener L remains constant while ambulance A moves from (a) to (c). While ambulance A moves from (a) to (b), the amount of change in the distance between the position of listener L and the position of the object in the virtual space is negative. While ambulance A moves from (b) to (c), the amount of change in the distance between the position of listener L and the position of the object in the virtual space is positive. Therefore, the timing at which the amount of change in distance changes from negative to positive is the timing at which ambulance A is at position (b) shown in FIG. 16.

[0266] Therefore, in step S35, the process shown in Fig. 17 is performed as follows: Fig. 17 is a flowchart for explaining the details of step S35 according to the second operation example.

[0267] After the process of step S30 is performed, the determination unit 120 determines whether the timing (predetermined timing) has arrived at which the amount of change in the distance between the position of the listener L and the position of the object (ambulance A) in the virtual space has changed from negative to positive (S35a). The determination unit 120 calculates the distance between the position of the listener L and the position of the object (ambulance A) and differentiates the calculated distance to calculate the amount of change in the distance. If the answer is Yes in step S35a, the process of step S40 is performed, and if the answer is No in step S35a, the process of step S35 is repeated.

[0268] Furthermore, the predetermined time period according to this operation example will be described in more detail.

[0269] In real space, the listener L hears the aerodynamic sound when the time has passed from the time when the amount of change in the distance between the position of the listener L and the position of the object turns from negative to positive until the wind W generated by the ambulance A reaches the listener L. As described above, the time when the amount of change in the distance turns from negative to positive is the time when the object is closest to the listener L, and is the predetermined time. Therefore, the determination unit 120 may determine the time from the predetermined time until the wind W generated by the ambulance A reaches the listener L as the predetermined time.

[0270] In this operation example, the predetermined time is determined based on the same concept as in Fig. 13A described in Operation Example 1. That is, as shown in Fig. 15, the distance between the position of listener L and the position of the object (ambulance A) is set to D, and more specifically, the distance between the position of ambulance A at the position (b) shown in Fig. 16 and the position of listener L is set to D.

[0271] Let U be the distance from the position of the object (ambulance A) at which the wind speed of wind W generated by the object ambulance A is So. Furthermore, let the direction from ambulance A toward listener L be the x-axis direction, and let x be the distance from ambulance A in the x-axis direction. Because the wind speed V of wind W is inversely proportional to the distance x, the wind speed V and the distance x satisfy the following equation.

[0272] V = So×(U / x)

[0273] The average wind speed up to a position at a distance D satisfies the following formula.

[0274]

[0275] The time (predetermined time) from the moment when the change in the distance between the position of the listener L and the position of the object turns from negative to positive (i.e., a predetermined moment) until the wind W generated by the object, ambulance A, reaches the listener L is t, which is the value obtained by dividing the distance by the average wind speed, and satisfies the following formula.

[0276] t = {(D-U)^2} / {So×U×(log(D)-log(U))

[0277] Then, as described above, in step S60, aerodynamic sound data is output at a timing when a predetermined time t has elapsed from the predetermined timing.

[0278] As a result, the listener L can hear the aerodynamic sound output from the headphones 200 at the timing when the amount of change in the distance between the position of the listener L and the position of the object turns from negative to positive (i.e., the predetermined timing) and the time it takes for the wind W created by the ambulance A to reach the listener L (predetermined time t) has elapsed. Therefore, the listener L can hear the aerodynamic sound at the same timing as in real space, that is, at the appropriate timing, so the listener L is less likely to feel uncomfortable and can obtain a sense of presence.

[0279] A further explanation is as follows. In real space, the listener L hears the aerodynamic sound after a vehicle such as ambulance A comes closest to the listener L. Therefore, in virtual space, if the listener L hears the aerodynamic sound before the ambulance A comes closest to the listener L, the listener L will feel uncomfortable. In operation example 2, the predetermined timing is the timing when the amount of change in the distance between the position of the listener L and the position of the object turns from negative to positive (in other words, the timing when the object comes closest to the listener L). This allows the listener L to hear the aerodynamic sound after the object, a vehicle such as ambulance A, comes closest to the listener L. In other words, the listener L can hear the aerodynamic sound at the appropriate timing, so the listener L is less likely to feel uncomfortable and can obtain a sense of realism.

[0280] The ambulance A is a sound-generating object that generates a siren sound. As shown in Fig. 16, when the position of the ambulance A changes, that is, when the ambulance A moves, the output unit 130 may output an object sound signal indicating a siren sound so that the listener L hears a siren sound accompanied by the Doppler effect.

[0281] In the above-described operation example 2, the predetermined timing is the timing when the amount of change in the distance between the position of the listener L and the position of the object changes from negative to positive, but this is not limited to this. For example, in another first example of operation example 2, the predetermined timing may be the timing (second timing) when the distance between the position of the listener L and the position of the object becomes shorter than the predetermined distance. The predetermined distance is, for example, several meters to several tens of meters, and is a distance that indicates that the distance between the position of the listener L and the position of the object has become sufficiently close. The predetermined distance may be a value designated by, for example, an administrator of the acoustic signal processing device 100.

[0282] In this case, in step S35, the process shown in Fig. 18 is performed as follows: Fig. 18 is a flowchart illustrating the details of step S35 according to another first example of the second operation example.

[0283] After the process of step S30 is performed, the determination unit 120 determines whether or not the timing (second timing) has arrived at which the distance between the position of the listener L and the position of the object (ambulance A) in the virtual space has become shorter than a predetermined distance (S35b). As described above, if the answer is Yes in step S35b, the process of step S40 is performed, and if the answer is No in step S35b, the process of step S35 is repeated.

[0284] In this way, in the other first example of operation example 2, the listener L can hear the aerodynamic sound output from the headphones 200 at the time when the time has passed from the second time when the distance between the position of the listener L and the position of the object (ambulance A) becomes sufficiently close to the time when the wind W created by the ambulance A reaches the listener L.

[0285] Furthermore, a second example of another operation example 2 will be described. In this second example of another operation example 2, in step S35, the processes of both steps S35a and S35b shown in Figures 17 and 18 are performed. If both steps S35a and S35b are Yes, the process of step S40 is performed, and if at least one of steps S35a and S35b is No, the process of step S35 is repeated. The process shown in this second example of another operation example 2 may be performed.

[0286] Next, the pipeline processing will be described.

[0287] Some or all of the processing performed by the above-described acoustic signal processing device 100 may be performed as part of pipeline processing such as that described in Patent Document 2, for example. Fig. 19 is a functional block diagram and a diagram showing an example of steps for explaining a case in which the rendering units A0203 and A0213 in Figs. 6 and 7 perform pipeline processing. In the explanation using Fig. 19, a rendering unit 900, which is an example of the rendering units A0203 and A0213 in Figs. 6 and 7, will be used.

[0288] Pipeline processing refers to dividing the process for adding sound effects into multiple processes and executing each process one by one in sequence. Each of the divided processes performs, for example, signal processing on an audio signal or generation of parameters used in signal processing.

[0289] The rendering unit 900 in this embodiment includes, as pipeline processing, processes that perform, for example, reverberation effects, early reflection processing, distance attenuation effects, binaural processing, and the like. However, the above processes are merely examples, and other processes may be included, or some processes may not be included. For example, the rendering unit 900 may include diffraction processing or occlusion processing as pipeline processing, or may omit reverberation processing if it is not necessary. Furthermore, each process may be represented as a stage, and audio signals such as reflected sound generated as a result of each process may be represented as rendering items. The order of each stage in the pipeline processing and the stages included in the pipeline processing are not limited to the example shown in FIG. 19 .

[0290] It should be noted that not all of the stages shown in FIG. 19 may be included in the rendering unit 900, and some stages may be omitted, or other stages may be present in addition to the rendering unit 900.

[0291] As an example of pipeline processing, we will explain the processes performed in each of the following: reverberation processing, early reflection processing, distance attenuation processing, selection processing, generation processing, and binaural processing. Each process analyzes metadata included in the input signal and calculates the parameters necessary to generate reflected sounds.

[0292] 19, the rendering unit 900 includes a reverberation processing unit 901, an early reflection processing unit 902, a distance attenuation processing unit 903, a selection unit 904, a calculation unit 906, a generation unit 907, and a binaural processing unit 905. Here, an example will be described in which the reverberation processing unit 901 performs a reverberation processing step, the early reflection processing unit 902 performs an initial reflection processing step, the distance attenuation processing unit 903 performs a distance attenuation processing step, the selection unit 904 performs a selection processing step, and the binaural processing unit 905 performs a binaural processing step.

[0293] In the reverberation processing step, the reverberation processor 901 generates an audio signal indicating reverberant sound or parameters required to generate an audio signal. The reverberant sound is a sound that includes reverberant sound that reaches the listener as reverberation after the direct sound. As an example, the reverberant sound is a reverberant sound that reaches the listener after having been reflected more times (e.g., several tens of times) than the initial reflection sound, at a relatively later stage (e.g., about a hundred and several tens of milliseconds after the direct sound arrives) after the initial reflection sound described below reaches the listener. The reverberation processor 901 refers to the audio signal and spatial information included in the input signal and performs calculations using a predetermined function prepared in advance for generating reverberant sound.

[0294] The reverberation processor 901 may generate reverberation by applying a known reverberation generation method to the sound signal. One example of a known reverberation generation method is the Schroeder method, but the method is not limited to this. When applying the known reverberation generation process, the reverberation processor 901 uses the shape and acoustic characteristics of the sound reproduction space indicated by the spatial information. This allows the reverberation processor 901 to calculate parameters for generating an audio signal indicating reverberation.

[0295] In the early reflection processing step, the early reflection processing unit 902 calculates parameters for generating early reflection sounds based on spatial information. Early reflection sounds are reflected sounds that reach the listener after one or more reflections at a relatively early stage (e.g., approximately several tens of milliseconds after the direct sound arrives) after the direct sound arrives from the sound source object to the listener. The early reflection processing unit 902, for example, refers to the sound signal and metadata, and calculates the path (path length) of the reflected sound that travels from the sound source object to the listener after reflecting off the object, using the shape and size of the three-dimensional sound field (space), the positions of objects such as structures, and the reflectance of the objects. The early reflection processing unit 902 may also calculate the path (path length) of the direct sound. Information indicating the path may be used as a parameter for generating early reflection sounds, and may also be used as a parameter for the selection process of the reflected sound by the selection unit 904.

[0296] In the distance attenuation processing step, a distance attenuation processing unit 903 calculates the volume of sound reaching the listener based on the difference between the path length of the direct sound and the path length of the reflected sound calculated by the early reflection processing unit 902. The volume of sound reaching the listener attenuates in proportion to the distance to the listener (inversely proportional to the distance) relative to the volume of the sound source, so the volume of the direct sound can be obtained by dividing the volume of the sound source by the length of the path of the direct sound, and the volume of the reflected sound can be calculated by dividing the volume of the sound source by the length of the path of the reflected sound.

[0297] In the selection process step, the selection unit 904 selects a sound to be generated. The selection process may be performed based on parameters calculated in the previous steps.

[0298] When the selection process is performed as part of the pipeline process, sounds not selected in the selection process may not be subjected to subsequent processing in the pipeline process. By not performing subsequent processing on the unselected sounds, it is possible to reduce the computational load on the acoustic signal processing device 100 compared to when it is decided not to perform binaural processing on the unselected sounds.

[0299] Furthermore, when the selection processing described in this embodiment is executed as part of pipeline processing, if the order of the selection processing is set to be an earlier order among the orders of multiple processes in the pipeline processing, more of the processing after the selection processing can be omitted, thereby making it possible to reduce the amount of calculations even more. For example, if the selection processing is executed in an order earlier than the processing of the calculation unit 906 and the generation unit 907, it is possible to omit processing for aerodynamic sounds related to objects that have been determined not to be selected, making it possible to further reduce the amount of calculations in the acoustic signal processing device 100.

[0300] Additionally, parameters calculated during a part of the pipeline process that generates the rendering items may be used by the selection unit 904 or the calculation unit 906 .

[0301] In the binaural processing step, the binaural processing unit 905 performs signal processing on the audio signal of the direct sound so that the sound is perceived as reaching the listener from the direction of the sound source object. Furthermore, the binaural processing unit 905 performs signal processing so that the reflected sound is perceived as reaching the listener from an obstacle object involved in the reflection. Based on the coordinates and orientation of the listener in the sound space (i.e., the position and orientation of the listening point), a process of applying a head-related impulse response (HRIR) database (data base) is performed so that the sound reaches the listener from the position of the sound source object or the position of the obstacle object. Note that the position and direction of the listening point may change in accordance with, for example, the movement of the listener's head. Information indicating the position of the listener may also be acquired from a sensor.

[0302] The programs used for pipeline processing and binaural processing, spatial information required for acoustic processing, the HRIR DB, and other parameters such as threshold data are acquired from memory provided in the acoustic signal processing device 100 or from an external source. HRIR (Head-Related Impulse Responses) are response characteristics when a single impulse is generated. In other words, HRIR is a response characteristic converted from a frequency domain representation to a time domain representation by Fourier transforming a head-related transfer function (HRTF), which represents, as a transfer function, changes in sound caused by surrounding objects including the auricle, the human head, and shoulders. The HRIR DB is a database containing such information.

[0303] As an example of pipeline processing, the rendering unit 900 may include processing units (not shown), such as a diffraction processing unit or an occlusion processing unit.

[0304] The diffraction processing unit executes processing to generate an audio signal representing a sound including diffracted sound caused by an obstacle between the listener and the sound source object in a three-dimensional sound field (space). When there is an obstacle between the sound source object and the listener, the diffracted sound is sound that travels from the sound source object to the listener by going around the obstacle.

[0305] The diffraction processing unit, for example, refers to the sound signal and metadata, and uses the position of the sound source object in the three-dimensional sound field (space), the position of the listener, and the position, shape, and size of obstacles to calculate a path from the sound source object to the listener, bypassing obstacles, and generates diffracted sound based on that path.

[0306] The occlusion processing unit generates an audio signal that can be heard when a sound source object is located behind an obstacle object, based on the spatial information acquired in any of the steps and information such as the material of the obstacle object.

[0307] In the above embodiment, the position information assigned to the sound source object is defined as a "point" in the virtual space, and the details of the invention have been described assuming that the sound source is a so-called "point sound source." On the other hand, as a method of defining a sound source in a virtual space, a spatially extended sound source, rather than a point sound source, may be defined as an object having length, size, shape, etc. In such cases, the distance between the listener and the sound source or the direction of sound arrival is not determined, so the reflected sound resulting from this may be limited to "selection" by the selection unit 904 without analysis, or regardless of the analysis results. This is because it is possible to avoid deterioration in sound quality that may occur when the reflected sound is not selected. Alternatively, a representative point, such as the center of gravity of the object, may be defined, and the processing of the present disclosure may be applied assuming that the sound is generated from that representative point. In this case, the threshold may be adjusted according to information about the spatial extension of the sound source before applying the processing of the present disclosure.

[0308] Next, an example of the structure of a bitstream will be described.

[0309] The bitstream includes, for example, an audio signal and metadata. The audio signal is sound data that represents sound, indicating information such as the frequency and intensity of the sound. The spatial information included in the metadata is information about the space in which a listener who listens to sound based on the audio signal is located. Specifically, the spatial information is information about a predetermined position (localization position) when a sound image of the sound is localized at a predetermined position in a sound space (e.g., in a three-dimensional sound field), that is, when the listener perceives the sound as arriving from a predetermined direction. The spatial information includes, for example, sound source object information and position information indicating the position of the listener.

[0310] The sound source object information is information about an object that generates a sound based on an audio signal, that is, that indicates an object that plays the audio signal, and is information about a virtual object (sound source object) that is placed in a sound space, which is a virtual space that corresponds to the real space in which the object is placed. The sound source object information includes, for example, information indicating the position of the sound source object placed in the sound space, information about the orientation of the sound source object, information about the directionality of the sound emitted by the sound source object, information indicating whether the sound source object belongs to a living thing, and information indicating whether the sound source object is a moving object. For example, the audio signal corresponds to one or more sound source objects indicated by the sound source object information.

[0311] As an example of the data structure of a bitstream, the bitstream is made up of, for example, metadata (control information) and an audio signal.

[0312] The audio signal and metadata may be stored in a single bitstream or in separate bitstreams, and similarly, the audio signal and metadata may be stored in a single file or in separate files.

[0313] A bitstream may exist for each sound source or for each playback time. When a bitstream exists for each playback time, multiple bitstreams may be processed in parallel at the same time.

[0314] The metadata may be attached to each bitstream, or may be attached collectively as information for controlling a plurality of bitstreams, or may be attached to each playback time.

[0315] When the audio signal and metadata are stored separately in multiple bitstreams or multiple files, the audio signal and metadata may be included in information indicating other bitstreams or files related to one or some of the bitstreams or files, or the audio signal and metadata may be included in information indicating other bitstreams or files related to each of all of the bitstreams or files. Here, the related bitstreams or files are, for example, bitstreams or files that may be used simultaneously during audio processing. Furthermore, the related bitstreams or files may include a bitstream or file that collectively describes information indicating other related bitstreams or files. Here, the information indicating other related bitstreams or files may be, for example, an identifier indicating the other bitstream, a file name indicating the other file, a URL (Uniform Resource Locator), or a URI (Uniform Resource Identifier). In this case, the acquisition unit 110 identifies or acquires the bitstream or file based on the information indicating the other related bitstreams or files. Furthermore, a bitstream may contain information indicating other related bitstreams, and may also contain information indicating a bitstream or file related to another bitstream or file, where the file containing information indicating the related bitstream or file may be, for example, a control file such as a manifest file used in content distribution.

[0316] Note that all or some of the metadata may be acquired from sources other than the bitstream of the audio signal. For example, either the metadata controlling the audio or the metadata controlling the video may be acquired from sources other than the bitstream, or both may be acquired from sources other than the bitstream. Furthermore, when the metadata controlling the video is included in the bitstream acquired by the audio signal reproduction system, the audio signal reproduction system may have a function of outputting the metadata that can be used to control the video to a display device that displays images or a 3D video reproduction device that reproduces 3D video.

[0317] Furthermore, an example of information contained in the metadata will be described.

[0318] The metadata may be information used to describe a scene represented in sound space. Here, a scene is a term that refers to the collection of all elements representing three-dimensional images and acoustic events in sound space that are modeled in an audio signal reproduction system using metadata. In other words, the metadata here may include not only information that controls sound processing, but also information that controls video processing. Of course, the metadata may include information that controls only one of sound processing and video processing, or information used to control both.

[0319] The audio signal reproduction system generates virtual sound effects by performing sound processing on the audio signal using metadata included in the bitstream and additionally acquired interactive listener position information. Here, a case where early reflection processing, obstacle processing, diffraction processing, blocking processing, and reverberation processing are performed among the sound effects will be described, but other sound processing may also be performed using the metadata. For example, the audio signal reproduction system may add sound effects such as distance attenuation effect, localization, and Doppler effect. Furthermore, information for switching all or some of the sound effects on or off, and priority information may also be added as metadata.

[0320] As an example, the encoded metadata includes information about a sound space including a sound source object and an obstacle object, and information about the localization position when the sound image of the sound is localized at a predetermined position in the sound space (i.e., perceived as sound arriving from a predetermined direction). Here, an obstacle object is an object that can affect the sound perceived by a listener by, for example, blocking or reflecting the sound emitted by the sound source object before it reaches the listener. Obstacle objects may include not only stationary objects but also animals such as people or moving objects such as machines. Furthermore, when multiple sound source objects exist in a sound space, other sound source objects can be obstacle objects for any sound source object. Non-sounding objects, such as building materials or inanimate objects, that do not emit sound, as well as sound source objects that emit sound, can be obstacle objects.

[0321] The metadata includes all or part of the information representing the shape of the sound space, shape and position information of obstacle objects present in the sound space, shape and position information of sound source objects present in the sound space, and the position and orientation of the listener in the sound space.

[0322] The sound space may be either a closed space or an open space. The metadata also includes information indicating the reflectance of structures that can reflect sound in the sound space, such as floors, walls, and ceilings, and the reflectance of obstacle objects that exist in the sound space. Here, the reflectance is the ratio of the energy of reflected sound to incident sound, and is set for each frequency band of sound. Of course, the reflectance may be set uniformly regardless of the frequency band of sound. When the sound space is an open space, parameters such as a uniform attenuation rate, diffracted sound, and early reflected sound may be used.

[0323] In the above description, reflectance was mentioned as a parameter related to an obstacle object or a sound source object included in the metadata, but information other than reflectance may also be included. For example, information other than reflectance may include information about the material of the object as metadata related to both the sound source object and the non-sound-producing object. Specifically, information other than reflectance may include parameters such as diffusion rate, transmittance, and sound absorption rate.

[0324] Information about a sound source object may include information such as volume, radiation characteristics (directivity), playback conditions, the number and type of sound sources emitted from a single object, and information specifying the sound source area of ​​the object. The playback conditions may, for example, determine whether the sound is a continuous sound or an event-triggered sound. The sound source area of ​​the object may be determined relative to the position of the listener and the position of the object, or may be determined based on the object. When the sound source area of ​​the object is determined relative to the position of the listener and the position of the object, the surface of the object the listener is looking at can be used as the reference, allowing the listener to perceive sound C as emanating from the right side of the object and sound E as emanating from the left side of the object as viewed from the listener. When the sound source area of ​​the object is determined based on the object, the sound emitted from which area of ​​the object can be fixed regardless of the direction the listener is looking. For example, when viewed from the front, the listener can perceive a high-pitched sound coming from the right side and a low-pitched sound coming from the left side. In this case, if the listener goes behind the object, the listener can perceive a low-pitched sound coming from the right side and a high-pitched sound coming from the left side as viewed from the back.

[0325] Spatial metadata can include time to early reflections, reverberation time, direct to diffuse ratio, etc. If the direct to diffuse ratio is zero, only direct sound will be perceived by the listener.

[0326] (Effects, etc.) The acoustic signal processing method according to the embodiment includes an acquisition step of acquiring object information indicating a change in an object that causes wind W and a predetermined timing related to the change in the object, and an output step of outputting aerodynamic sound data indicating aerodynamic sound caused by wind W a predetermined time after the predetermined timing indicated by the acquired object information that is based on the change in the object.

[0327] This makes it possible to output aerodynamic sound data at a timing when a predetermined time has elapsed from a predetermined timing. Therefore, the listener L can hear the aerodynamic sound at the appropriate timing, and the listener L is less likely to feel uncomfortable and can obtain a sense of realism. In other words, an acoustic signal processing method that can give the listener L a sense of realism is realized.

[0328] For example, as shown in the first operational example, the predetermined timing is, for example, the timing of a change in the wind W, and the predetermined time is, for example, the time it takes for the wind W generated by the fan F to reach the listener L.

[0329] For example, as shown in the operation example 2, the predetermined timing is, for example, the timing of a change in the wind W, and the predetermined time is, for example, the time it takes for the wind W generated by the ambulance A to reach the listener L.

[0330] In the cases shown in Operation Examples 1 and 2, the listener L can hear the aerodynamic sound at the same timing as in real space, that is, at the appropriate timing, so the listener L is less likely to feel uncomfortable and can obtain a sense of realism. In this way, the acoustic signal processing method according to the embodiment can provide the listener L with a sense of realism.

[0331] Furthermore, for example, the predetermined timing may be a timing (designated timing) designated by the user, and the time designated by the user may be the predetermined time. In this case, the user may designate the designated timing and time, and the designated timing and time may be set as the predetermined timing and predetermined time, so that the listener L can hear the aerodynamic sound at the same timing as in real space. Even in this case, the listener L can hear the aerodynamic sound at the same timing as in real space, that is, at the appropriate timing, so the listener L is less likely to feel uncomfortable, and the listener L can obtain a sense of presence.

[0332] In addition, in the acoustic signal processing method according to the embodiment, the object information indicates a change in the wind W due to a change in the object, and the predetermined timing indicates the timing of the change in the wind W. The acoustic signal processing method includes a determination step of determining the predetermined time based on the wind W indicated by the acquired object information.

[0333] This allows aerodynamic sound data to be output when a predetermined time determined based on the wind W has elapsed since the wind W changed, allowing the listener L to hear the aerodynamic sound at a more appropriate timing.

[0334] In the sound signal processing method according to the embodiment, the change in the wind W indicated by the object information indicates a change in the wind speed of the wind W, and in the determining step, the predetermined time is determined based on the wind speed.

[0335] This allows the predetermined time to be determined based on the wind speed, allowing the listener L to hear the aerodynamic sound at a more appropriate timing.

[0336] In the acoustic signal processing method according to the embodiment, the aerodynamic sound is a sound generated at the changed wind speed.

[0337] This allows the aerodynamic sound heard by the listener L in the virtual space to be closer to the aerodynamic sound heard by the listener L in the real space.

[0338] In addition, in the acoustic signal processing method according to the embodiment, the object information indicates the position of the object. The acoustic signal processing method includes a determination step of determining a predetermined time based on the distance between the position of a listener L of the aerodynamic sound and the position of the object indicated by the acquired object information.

[0339] As a result, the predetermined time is determined based on the distance, and the listener L can hear the aerodynamic sound at a more appropriate timing.

[0340] In the acoustic signal processing method according to the embodiment, the object information indicates the position of the object. In the determining step, the predetermined time is determined based on the wind speed and the distance between the position of the listener L of the aerodynamic sound and the position of the object indicated by the acquired object information.

[0341] As a result, the predetermined time is determined based on the wind speed and the distance, allowing the listener L to hear the aerodynamic sound at a more appropriate timing.

[0342] In the acoustic signal processing method according to the embodiment, the object information indicates that the predetermined timing is a first timing for outputting sound data associated with the object, and in the output step, the aerodynamic sound data is output a predetermined time after the first timing indicated by the acquired object information.

[0343] This allows, for example, when an object generates a sound, the aerodynamic sound data to be output at a timing when a predetermined time has elapsed from the first timing at which the sound is output, so that the listener L can hear the aerodynamic sound at a more appropriate timing.

[0344] For example, as shown in Operation Example 1, if the object is an electric fan F that generates a motor sound, the predetermined timing is, for example, the timing when the electric fan F is switched from OFF to ON. When a time (predetermined time) has elapsed from this predetermined timing for the wind W generated by the electric fan F to reach the listener L, the listener L can hear the aerodynamic sound output from the headphones 200. Therefore, the listener L can hear the aerodynamic sound at the same timing as in real space, that is, at the appropriate timing, so the listener L is less likely to feel uncomfortable and can obtain a sense of realism. In this way, the acoustic signal processing method according to the embodiment can provide the listener L with a sense of realism.

[0345] In addition, in the acoustic signal processing method according to the modified example of the embodiment, the object information indicates the position of the object, and the predetermined timing is a second timing at which the distance between the position of the listener L of the aerodynamic sound and the position of the object becomes shorter than the predetermined distance. In the output step, the aerodynamic sound data is output a predetermined time after the second timing indicated by the acquired object information.

[0346] This allows aerodynamic sound data to be output at the second timing when the distance becomes shorter than the predetermined distance, that is, at the timing when a predetermined time has elapsed since the second timing when the object approached the listener L, so that the listener L can hear the aerodynamic sound at a more appropriate timing.

[0347] For example, as shown in Operation Example 2, the predetermined timing is, for example, the timing when the amount of change in the distance between the position of the listener L and the position of the object changes from negative to positive. When the time (predetermined time) for the wind W created by the ambulance A to reach the listener L has elapsed from this predetermined timing, the listener L can hear the aerodynamic sound output from the headphones 200. Therefore, the listener L can hear the aerodynamic sound at the same timing as in real space, that is, at the appropriate timing, so the listener L is less likely to feel uncomfortable and can obtain a sense of realism. In this way, the acoustic signal processing method according to the modified example of the embodiment can provide the listener L with a sense of realism.

[0348] In the acoustic signal processing method according to the embodiment, the object information indicates that the change in the wind W due to the change in the object is a change in the direction of the wind W, and the predetermined timing is a third timing at which the change in direction of the wind W occurred. In the output step, the aerodynamic sound data is output a predetermined time after the third timing indicated by the acquired object information.

[0349] This allows aerodynamic sound data to be output at a timing when a predetermined time has elapsed from the third timing when the change in direction of the wind W occurs, so that the listener L can hear the aerodynamic sound at a more appropriate timing.

[0350] In addition, in the acoustic signal processing method according to the embodiment, the object is an object that generates a sound and wind W indicated by sound data associated with the object, and the aerodynamic sound is aerodynamic sound that is generated when the wind W generated by the object reaches the listener L.

[0351] This allows an object such as an electric fan F that generates sound and wind W to be created, and aerodynamic sound caused by the wind W blown out from the object can be realized.

[0352] In the acoustic signal processing method according to the embodiment, the distance is defined as D, and the distance from the object position at which the wind speed is So is defined as U. When the predetermined time is defined as t, t satisfies the following formula.

[0353] t={(D-U)^2} / {So×U×(log(D)-log(U))

[0354] As a result, in the determination step, the predetermined time can be determined as the time from the predetermined timing until the wind W generated by the object reaches the listener L. Therefore, since the aerodynamic sound data can be output at a timing when such a predetermined time has elapsed from the predetermined timing, the listener L can hear the aerodynamic sound at a more appropriate timing.

[0355] For example, as shown in Operation Example 1, in the determining step, the time at which the wind W generated by the fan F reaches the listener L can be determined as the predetermined time. Therefore, the listener L can hear the aerodynamic sound at the same timing as in real space, that is, at the appropriate timing, so the listener L is less likely to feel uncomfortable and can obtain a sense of realism. In this way, the acoustic signal processing method according to the embodiment can provide the listener L with a sense of realism.

[0356] In addition, in the acoustic signal processing method according to the modified example of the embodiment, the object is an object that generates wind W by moving its position, and the aerodynamic sound is aerodynamic sound that is generated when the wind W generated by the movement reaches a listener L.

[0357] This allows a vehicle or the like that generates wind W as it moves to be used as an object, and makes it possible to realize aerodynamic sound caused by the wind W generated by the movement.

[0358] In the acoustic signal processing method according to the modified example of the embodiment, the predetermined timing indicated by the object information is the timing when the amount of change in distance over time turns from negative to positive.

[0359] This allows aerodynamic sound data to be output at a predetermined time after the position of the listener L and the position of the object are closest, allowing the listener L to hear the aerodynamic sound at a more appropriate time.

[0360] In the acoustic signal processing method according to the modified example of the embodiment, the distance is defined as D, and the distance from the position of the object at which the wind speed of the wind W generated by the movement becomes So is defined as U. When the predetermined time is defined as t, t satisfies the following formula.

[0361] t={(D-U)^2} / {So×U×(log(D)-log(U))

[0362] As a result, in the determination step, the predetermined time can be determined as the time from the predetermined timing until the wind W generated by the object reaches the listener L. Therefore, since the aerodynamic sound data can be output at a timing when such a predetermined time has elapsed from the predetermined timing, the listener L can hear the aerodynamic sound at a more appropriate timing.

[0363] For example, as shown in Operation Example 2, in the determining step, the time at which the wind W generated by the ambulance A reaches the listener L can be determined as the predetermined time. Therefore, the listener L can hear the aerodynamic sound at the same timing as in real space, that is, at the appropriate timing, so the listener L is less likely to feel uncomfortable and can obtain a sense of realism. In this way, the acoustic signal processing method according to the embodiment can provide the listener L with a sense of realism.

[0364] A computer program according to the embodiment is a computer program for causing a computer to execute the above-described acoustic signal processing method.

[0365] This allows the computer to execute the above-described acoustic signal processing method in accordance with the computer program.

[0366] Furthermore, the acoustic signal processing device 100 according to the embodiment includes an acquisition unit 110 that acquires object information indicating a change in an object that causes wind W and a predetermined timing related to the change in the object, and an output unit 130 that outputs aerodynamic sound data indicating aerodynamic sound caused by the wind W a predetermined time after the predetermined timing indicated by the acquired object information, which is based on the change in the object.

[0367] This makes it possible to output aerodynamic sound data at a timing when a predetermined time has elapsed from a predetermined timing. Therefore, the listener L can hear the aerodynamic sound at the appropriate timing, and the listener L is less likely to feel uncomfortable and can obtain a sense of realism. In other words, an acoustic signal processing device 100 that can give the listener L a sense of realism is realized.

[0368] (Other Embodiments) While the acoustic signal processing method and acoustic signal processing device according to aspects of the present disclosure have been described above based on embodiments and modifications, the present disclosure is not limited to these embodiments and modifications. For example, other embodiments realized by arbitrarily combining the components described in this specification or excluding some of the components may also be considered as embodiments of the present disclosure. Furthermore, modifications obtained by applying various modifications that would occur to a person skilled in the art to the above embodiments and modifications without departing from the spirit of the present disclosure, i.e., the meaning indicated by the wording of the claims, are also included in the present disclosure.

[0369] In the above embodiment, an example was shown in which the object was an electric fan F, but the object is not limited to this. Here, an object that generates wind W will be exemplified.

[0370] The object that generates the wind W may be, for example, an object into which the wind W blows, such as a window or a door. In an example in which a listener L is inside a building in a virtual space and wind W is blowing outside the building, when the window or door opens, the wind W blows into the building, causing the listener L to hear aerodynamic sound. In this example, the timing when the window or door opens corresponds to the predetermined timing, and the wind W is generated at the position of the window or door, and the technology of the present disclosure can be applied.

[0371] The object that generates the wind W may be, for example, an object from which the wind W blows out, such as a vent or an exhaust hole. For wind W blowing out from a vent or an exhaust hole, it is meaningless in the virtual space to precisely define the location where the wind W is generated, and the technology of the present disclosure can be applied by assuming that the wind W is generated at the exit position of the vent or exhaust hole. In this case, the predetermined timing can be determined by an administrator of the virtual space or an administrator of the acoustic signal processing device 100. For example, a receiving unit included in the acoustic signal processing device 100 may receive a timing specified by the administrator, and the determination unit 120 may determine the timing received by the receiving unit as the predetermined timing.

[0372] The following embodiments may also be included within the scope of one or more aspects of the present disclosure.

[0373] (1) Some of the components constituting the above-mentioned audio signal processing device may be a computer system comprising a microprocessor, ROM, RAM, hard disk unit, display unit, keyboard, mouse, etc. A computer program is stored in the RAM or hard disk unit. The microprocessor operates in accordance with the computer program to achieve its functions. Here, the computer program is composed of a combination of multiple instruction codes that indicate commands to a computer to achieve a predetermined function.

[0374] (2) Some of the components constituting the above-described audio signal processing device may be configured as a single system LSI (Large Scale Integration). A system LSI is an ultra-multifunctional LSI manufactured by integrating multiple components on a single chip, and specifically, is a computer system configured to include a microprocessor, ROM, RAM, etc. A computer program is stored in the RAM. The system LSI achieves its functions by the microprocessor operating in accordance with the computer program.

[0375] (3) Some of the components constituting the above-mentioned audio signal processing device may be configured as an IC card or a standalone module that can be attached to each device. The IC card or the module may be a computer system configured with a microprocessor, ROM, RAM, etc. The IC card or the module may include the above-mentioned ultra-multifunctional LSI. The IC card or the module achieves its functions when the microprocessor operates according to a computer program. The IC card or the module may be tamper-resistant.

[0376] (4) Furthermore, some of the components constituting the above-described acoustic signal processing device may be the computer program or the digital signal recorded on a computer-readable recording medium, such as a flexible disk, hard disk, CD-ROM, MO, DVD, DVD-ROM, DVD-RAM, BD (Blu-ray (registered trademark) Disc), semiconductor memory, etc. Alternatively, the components may be digital signals recorded on such recording media.

[0377] Furthermore, some of the components constituting the above-mentioned acoustic signal processing device may transmit the computer program or the digital signal via a telecommunications line, a wireless or wired communication line, a network such as the Internet, data broadcasting, etc.

[0378] (5) The present disclosure may be embodied as the methods described above, a computer program that implements these methods on a computer, or a digital signal that includes the computer program.

[0379] (6) The present disclosure may also be a computer system having a microprocessor and a memory, the memory storing the computer program, and the microprocessor operating in accordance with the computer program.

[0380] (7) The program or the digital signal may also be implemented by another independent computer system by recording it on the recording medium and transferring it, or by transferring the program or the digital signal via the network or the like.

[0381] The present disclosure is applicable to an acoustic signal processing method and an acoustic signal processing device, and is particularly applicable to an acoustic system and the like.

[0382] 100 Acoustic signal processing device 110 Acquisition unit 120 Decision unit 130 Output unit 140 Storage unit 200 Headphones 201 Head sensor unit 202 Output unit 300 Display unit 900 Rendering unit 901 Reverberation processing unit 902 Early reflection processing unit 903 Distance attenuation processing unit 904 Selection unit 905 Binaural processing unit 906 Calculation unit 907 Generation unit A Ambulance A0000 Stereophonic sound reproduction system A0001 Acoustic signal processing device A0002 Audio presentation device A0100 Encoding device A0101 Input data A0102 Encoder A0103 Encoded data A0104 Memory A0110 Decoding device A0111 Audio signal A0112 Decoder A0113 Input data A0114 Memory A0120 Encoding device A0121 Transmitting unit A0122 Transmitted signal A0130 Decoding device A0131 Receiving unit A0132 Received signal A0200 Decoder A0201 Spatial information management unit A0202 Audio data decoder A0203 Rendering unit A0210 Decoder A0211 Spatial information management unit A0213 Rendering unit F Fan L Listener

Claims

1. an acquisition step of acquiring object information indicating a change in an object causing wind and a predetermined timing related to the change in the object; an output step of outputting aerodynamic sound data indicating aerodynamic sound caused by the wind after a predetermined time based on a change in the object from the predetermined timing indicated by the acquired object information; Including, Acoustic signal processing method.

2. The object information is A change in the wind due to a change in the object; the predetermined timing is a timing of a change in the wind, The acoustic signal processing method includes a determination step of determining the predetermined time based on the wind indicated by the acquired object information.

2. The method of claim 1.

3. the change in the wind indicated by the object information indicates a change in the wind speed, In the determination step, the predetermined time is determined based on the wind speed.

3. The method of claim 2.

4. The aerodynamic sound is a sound generated at the changed wind speed. The method of processing an acoustic signal according to claim 3.

5. the object information indicates a position of the object; The acoustic signal processing method includes a determination step of determining the predetermined time based on a distance between a position of a listener of the aerodynamic sound and a position of the object indicated by the acquired object information.

2. The method of claim 1.

6. the object information indicates a position of the object; In the determining step, the predetermined time is determined based on the wind speed and a distance between a position of a listener of the aerodynamic sound and a position of the object indicated by the acquired object information. The method of processing an acoustic signal according to claim 3.

7. the object information indicates that the predetermined timing is a first timing for outputting sound data associated with the object, In the output step, the aerodynamic sound data is output after the predetermined time from the first timing indicated by the acquired object information.

2. The method of claim 1.

8. The object information is The position of the object; the predetermined timing is a second timing at which a distance between a position of a listener of the aerodynamic sound and a position of the object becomes shorter than a predetermined distance, In the output step, the aerodynamic sound data is output after the predetermined time from the second timing indicated by the acquired object information.

2. The method of claim 1.

9. The object information is The change in the wind due to the change in the object is a change in the direction of the wind; the predetermined timing is a third timing at which the change in wind direction occurs, In the output step, the aerodynamic sound data is output after the predetermined time from a third timing indicated by the acquired object information.

2. The method of claim 1.

10. the object is an object that generates the sound indicated by sound data associated with the object and the wind, The aerodynamic sound is generated when the wind generated by the object reaches the listener.

7. The method of claim 6.

11. The distance is D, Let U be the distance from the object position at which the wind speed S is reached, When the predetermined time is t, the t satisfies the following formula: t={(D-U)^2} / {So×U×(log(D)-log(U))} The method of processing an acoustic signal according to claim 10.

12. the object generates the wind by moving a position of the object, The aerodynamic sound is generated when the wind generated by the movement reaches the listener.

7. The method of claim 6.

13. the predetermined timing indicated by the object information is a timing at which the amount of change in the distance over time turns from negative to positive; The method of processing an acoustic signal according to claim 12.

14. The distance is D, The distance from the position of the object at which the wind speed of the wind generated by the movement becomes So is defined as U, When the predetermined time is t, the t satisfies the following formula: t={(D-U)^2} / {So×U×(log(D)-log(U))} The method of processing an acoustic signal according to claim 12.

15. A computer program for causing a computer to execute the acoustic signal processing method according to any one of claims 1 to 14.

16. an acquisition unit that acquires object information indicating a change in an object that causes wind and a predetermined timing related to the change in the object; an output unit that outputs aerodynamic sound data indicating aerodynamic sound caused by the wind after a predetermined time based on a change in the object from the predetermined timing indicated by the acquired object information; Equipped with Acoustic signal processing device.