Acoustic signal processing method, acoustic signal processing apparatus, and computer program

JPWO2024084950A5Pending Publication Date: 2025-07-03
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024551437
Authority / Receiving Office
JP · JP
Patent Type
Applications
Filing Date
2025-04-10
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing acoustic signal processing technologies struggle to provide a realistic sense of presence in virtual spaces, particularly in reproducing wind sounds with fluctuations in real space, leading to discomfort for listeners.

Method used

An acoustic signal processing method that acquires sound data indicating a waveform of a reference sound, processes it to change frequency components, phase, and amplitude values based on simulation information that simulates fluctuations in natural phenomena, such as wind speed, and outputs the processed sound data to create a more realistic wind sound experience.

Benefits of technology

The method effectively enhances the sense of realism for listeners by introducing fluctuations in frequency components, phase, and amplitude values, reducing discomfort and improving the immersion in virtual environments.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

An acoustic signal processing method comprising: an acquisition step for acquiring sound data indicating the waveform of a reference sound; a processing step for processing the sound data so as to change at least one of a frequency component, the phase, and the amplitude value of the waveform on the basis of mimicry information that mimics fluctuations in a natural phenomenon; and an output step for outputting the processed sound data.
Need to check novelty before this filing date? Find Prior Art

Description

Acoustic signal processing method, computer program, and acoustic signal processing device

[0001] The present disclosure relates to an acoustic signal processing method and the like.

[0002] Furthermore, Patent Document 1 discloses a technology for outputting images and sounds in order to create a realistic virtual space, and also discloses a technology for changing the sound of the wind in accordance with changes in the strength of the wind in the virtual space.

[0003] JP 1998-2151162 A International Publication No. 2021 / 180938

[0004] Yoshinori Dobashi, 2 others, Real-time rendering of aerodynamic sound using sound textures based on computational fluid dynamics, ACM Transactions on Graphics, Vol. 22, No. 3, p732-740

[0005] However, with the technology disclosed in Patent Document 1, it may be difficult to give a sense of realism to the listener.

[0006] Therefore, an object of the present disclosure is to provide an acoustic signal processing method and the like that can give a sense of realism to listeners.

[0007] An acoustic signal processing method according to one aspect of the present disclosure includes an acquisition step of acquiring sound data indicating a waveform of a reference sound, a processing step of processing the sound data so as to change at least one of a frequency component, a phase, and an amplitude value of the waveform based on simulation information that simulates fluctuations in a natural phenomenon, and an output step of outputting the processed sound data.

[0008] Furthermore, a computer program according to one aspect of the present disclosure causes a computer to execute the above-described acoustic signal processing method.

[0009] In addition, an acoustic signal processing device according to one aspect of the present disclosure includes an acquisition unit that acquires sound data indicating the waveform of a reference sound, a processing unit that processes the sound data so as to change at least one of the frequency component, phase, and amplitude value of the waveform based on simulation information that simulates fluctuations in a natural phenomenon, and an output unit that outputs the processed sound data.

[0010] These comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0011] According to the acoustic signal processing method according to one aspect of the present disclosure, it is possible to give a sense of realism to the listener.

[0012] FIG. 1 is a diagram showing a stereophonic (immersive audio) reproduction system as an example of a system to which the acoustic processing or decoding processing of the present disclosure can be applied. FIG. 2 is a functional block diagram showing the configuration of an encoding device as an example of the encoding device of the present disclosure. FIG. 3 is a functional block diagram showing the configuration of a decoding device as an example of the decoding device of the present disclosure. FIG. 4 is a functional block diagram showing the configuration of an encoding device as another example of the encoding device of the present disclosure. FIG. 5 is a functional block diagram showing the configuration of a decoding device as another example of the decoding device of the present disclosure. FIG. 6 is a functional block diagram showing the configuration of a decoder as an example of the decoder in FIG. 3 or FIG. 5. FIG. 7 is a functional block diagram showing the configuration of a decoder as another example of the decoder in FIG. 3 or FIG. 5. FIG. 8 is a diagram showing an example of the physical configuration of an acoustic signal processing device. FIG. 9 is a diagram showing an example of the physical configuration of an encoding device. FIG. 10 is a block diagram showing the functional configuration of an acoustic signal processing device according to embodiment 1. FIG. 11 is a diagram showing a fan and a listener as examples of objects according to embodiment 1. FIG. 12 is a diagram showing sound data according to embodiment 1. FIG. 13 is a diagram showing an example of a smooth function according to embodiment 1. FIG. 14 is a flowchart of a first operational example of the acoustic signal processing device according to the first embodiment. FIG. 15 is a diagram for explaining processing performed by a processing unit according to the first embodiment. FIG. 16 is another diagram for explaining processing performed by a processing unit according to the first embodiment. FIG. 17 is a diagram showing sound data (aerodynamic sound data) according to the first embodiment. FIG. 18 is a diagram showing R, which is a value indicated by a smooth function according to the first embodiment, and the amplification factor and attenuation factor of the volume of aerodynamic sound. FIG. 19 is a diagram showing divided aerodynamic sound data according to the first embodiment. FIG. 20 is a diagram showing another example of two smooth functions according to the first embodiment. FIG. 21 is a diagram showing an example in which the parameters specifying the smooth function according to the first embodiment have changed. FIG. 22 is a diagram showing another example of two smooth functions according to the first embodiment. FIG. 23 is a block diagram showing the functional configuration of an acoustic signal processing device according to a modification. FIG. 24 is a block diagram showing the functional configuration of a second processing unit according to the modification. FIG. 25 is a diagram showing aerodynamic sound data according to the modification. FIG. 26 is a conceptual diagram of processing by the second processing unit according to the modification.FIG. 27 is a block diagram showing a functional configuration of a sampling rate conversion unit according to a modified example. FIG. 28 is a state transition diagram of values ​​indicated by a smooth function according to a modified example. FIG. 29 is a block diagram showing another functional configuration of an acoustic signal processing device according to a modified example. FIG. 30 is a block diagram showing a functional configuration of an information processing device according to Embodiment 2. FIG. 31 is a diagram for explaining reading of sound data according to the conventional technology and reading of sound data according to Embodiment 2. FIG. 32 is a diagram for explaining processing performed by the information processing device according to Embodiment 2. FIG. 33 is a diagram for explaining other processing performed by the information processing device according to Embodiment 2. FIG. 34 is a functional block diagram and a diagram showing an example of steps for explaining a case where the rendering units of FIGS. 6 and 7 perform pipeline processing.

[0013] (Findings that form the basis of the present disclosure) Patent Literature 1 discloses a technology for outputting images and sounds in order to create a realistic virtual space, and also discloses a technology for changing the sound of the wind in accordance with changes in the strength of the wind in the virtual space.

[0014] A virtual space is a space in which a user (listener) exists, such as virtual reality (VR) or augmented reality (AR). The wind sound produced using the technology disclosed in Patent Document 1 is used in an application for reproducing three-dimensional sound in such a virtual space. Such controlled sound is particularly used in a virtual space in which information on the listener's 6 DoF (Degrees of Freedom) is sensed. By using the technology of Patent Document 1, natural phenomena such as wind blowing can be reproduced in the virtual space.

[0015] Incidentally, in real space, fluctuations in natural phenomena include fluctuations. Examples of natural phenomena in real space include the wind blowing, the water flowing in a river, animal behavior, etc. For example, fluctuations in natural phenomena include fluctuations in wind speed or wind direction, and fluctuations in wind speed or wind direction include fluctuations.

[0016] However, while the technology disclosed in Patent Document 1 can allow a listener to hear the sound of wind, it cannot reproduce the sound of wind that includes fluctuations in real space. Therefore, when a listener hears such wind sound, the listener feels uncomfortable, and it is difficult for the listener to obtain a sense of realism. For this reason, there is a demand for an acoustic signal processing method that can provide a sense of realism to the listener.

[0017] Therefore, the acoustic signal processing method according to the first aspect of the present disclosure includes an acquisition step of acquiring sound data indicating the waveform of a reference sound, a processing step of processing the sound data so as to change at least one of the frequency component, phase, and amplitude value of the waveform based on simulation information that simulates fluctuations in a natural phenomenon, and an output step of outputting the processed sound data.

[0018] As a result, the sound data is processed so as to change at least one of the frequency components, phase, and amplitude of the waveform based on simulation information that simulates fluctuations in natural phenomena containing fluctuations. As a result, fluctuations occur in at least one of the frequency components, phase, and amplitude in the processed sound data, and fluctuations also occur in at least one of the frequency components, phase, and amplitude in the sound represented by the processed sound data. Therefore, the listener can hear a sound in which fluctuations occur in at least one of the frequency components, phase, and amplitude, and the listener can enjoy a sense of realism without feeling any discomfort. In other words, an acoustic signal processing method that can give the listener a sense of realism is realized.

[0019] An acoustic signal processing method according to a second aspect of the present disclosure is the acoustic signal processing method according to the first aspect, wherein the reference sound is an aerodynamic sound generated by wind, and in the processing step, the sound data is processed so as to change at least one of the frequency component, the phase, and the amplitude value of the waveform based on the simulation information in which fluctuations in the wind speed of the wind are simulated.

[0020] This allows the listener to hear aerodynamic sound with fluctuations in at least one of the frequency components, phase, and amplitude, and the listener can experience a sense of realism without feeling any discomfort. In other words, an acoustic signal processing method that can give the listener a sense of realism is realized.

[0021] An acoustic signal processing method according to a third aspect of the present disclosure is the acoustic signal processing method according to the second aspect, wherein in the processing step, a smooth function that simulates fluctuations in the wind speed is determined as the simulation information, and the sound data is processed so as to change at least one of the frequency components, phase, and amplitude value of the waveform based on a value indicated by the determined smooth function.

[0022] This allows the sound data to be processed according to the values ​​indicated by the smooth function.

[0023] An acoustic signal processing method according to a fourth aspect of the present disclosure is the acoustic signal processing method according to the third aspect, in which the value indicated by the smooth function is information indicating the ratio between the wind speed of the aerodynamic sound that is the reference sound and the wind speed of the aerodynamic sound indicated by the sound data after processing in the processing step.

[0024] This allows the sound data to be processed based on the ratio between the wind speed of the aerodynamic sound, which is the reference sound, and the wind speed of the aerodynamic sound indicated by the processed sound data.

[0025] An acoustic signal processing method according to a fifth aspect of the present disclosure is the acoustic signal processing method according to the fourth aspect, wherein in the processing step, the smooth function is determined so that a parameter specifying the smooth function varies irregularly.

[0026] This allows the listener to hear aerodynamic sound with irregular fluctuations in at least one of the frequency components, phase, and amplitude, and the listener is less likely to feel uncomfortable and can enjoy a greater sense of realism.In other words, an acoustic signal processing method that can give the listener a greater sense of realism is realized.

[0027] An acoustic signal processing method according to a sixth aspect of the present disclosure is the acoustic signal processing method according to any one of the third to fifth aspects, wherein in the processing step, the sound data is processed so as to shift the frequency components of the waveform to frequencies proportional to the values ​​indicated by the determined smooth function.

[0028] This allows the listener to hear a sound with fluctuations in the frequency components, and the listener can get a sense of realism without feeling any discomfort. In other words, an acoustic signal processing method that can give the listener a sense of realism is realized.

[0029] An acoustic signal processing method according to a seventh aspect of the present disclosure is the acoustic signal processing method according to the third aspect, wherein in the processing step, the sound data is processed so as to change an amplitude value of the waveform in proportion to the α power of a value indicated by the determined smooth function.

[0030] This allows the listener to hear sounds with fluctuations in amplitude, and the listener can get a sense of realism without feeling any discomfort. In other words, an acoustic signal processing method that can give the listener a sense of realism is realized.

[0031] An acoustic signal processing method according to an eighth aspect of the present disclosure is the acoustic signal processing method according to the fourth or fifth aspect, wherein, in the processing step, the acquired sound data is divided into processing frames of a predetermined time, and the sound data is processed for each of the divided processing frames.

[0032] This realizes an acoustic signal processing method with a reduced computational processing load.

[0033] An acoustic signal processing method according to a ninth aspect of the present disclosure is the acoustic signal processing method according to the eighth aspect, wherein in the processing step, the smooth function is determined for each divided processing frame so that the value of the smooth function becomes 1.0 at the first time and the last time of the processing frame.

[0034] This prevents noise from occurring at the joint between a processing frame and the next processing frame.

[0035] An acoustic signal processing method according to a tenth aspect of the present disclosure is the acoustic signal processing method according to the ninth aspect, wherein in the processing step, parameters specifying the smooth function are determined for each of the divided processing frames.

[0036] This realizes an acoustic signal processing method with a reduced computational processing load.

[0037] An acoustic signal processing method according to an eleventh aspect of the present disclosure is the acoustic signal processing method according to the tenth aspect, in which the parameter is a time from the first time to the last time.

[0038] This allows the parameter to be the time from the first time of the processing frame to the last time of the processing frame.

[0039] An acoustic signal processing method according to a twelfth aspect of the present disclosure is the acoustic signal processing method according to the tenth aspect, in which the parameter is a value related to a maximum value of the smooth function.

[0040] This allows the parameter to be a value related to the maximum value of a smooth function.

[0041] An acoustic signal processing method according to a thirteenth aspect of the present disclosure is the acoustic signal processing method according to the tenth aspect, in which the parameter is a parameter that varies the position at which the smooth function reaches a maximum value.

[0042] This allows the parameter to vary the position where the smooth function reaches its maximum value.

[0043] An acoustic signal processing method according to a fourteenth aspect of the present disclosure is the acoustic signal processing method according to the tenth aspect, in which the parameter is a parameter that varies the steepness of variation of the smooth function.

[0044] This allows the parameter to be a parameter that changes the steepness of the change in the smooth function.

[0045] An acoustic signal processing method according to a fifteenth aspect of the present disclosure is the acoustic signal processing method according to the tenth aspect, wherein, in the processing step, a first parameter and a second parameter that specify the smooth function are determined, the acquired sound data is processed so as to change at least one of a frequency component, a phase, and an amplitude value of the waveform based on the smooth function specified by the determined first parameter, and the acquired sound data is processed so as to change at least one of a frequency component, a phase, and an amplitude value of the waveform based on the smooth function specified by the determined second parameter, and in the output step, the sound data processed based on the smooth function specified by the determined first parameter is output to a first output channel, and the sound data processed based on the smooth function specified by the determined second parameter is output to a second output channel.

[0046] This allows different sound data to be output for each output channel.

[0047] An acoustic signal processing method according to a sixteenth aspect of the present disclosure is the acoustic signal processing method according to any one of the tenth to fifteenth aspects, in which the aerodynamic sound is sound generated by the wind colliding with an object, and in the processing step, the parameters are determined by simulating the characteristics of the wind speed of the wind.

[0048] This allows the parameters to be determined by simulating fluctuations in wind speed containing fluctuations, and the sound data can be processed to change at least one of the frequency components, phase, and amplitude of the waveform based on a smooth function specified by the parameters.

[0049] An acoustic signal processing method according to a seventeenth aspect of the present disclosure is the acoustic signal processing method according to any one of the tenth to fifteenth aspects, wherein the aerodynamic sound is sound generated when the wind collides with the ear of a listener who hears the aerodynamic sound, and the processing step determines the parameters by simulating the properties of the wind direction.

[0050] This allows parameters to be determined that simulate fluctuations in wind direction, and the sound data can be processed to vary at least one of the frequency components, phase, and amplitude of the waveform based on a smooth function specified by the parameters.

[0051] An acoustic signal processing method according to an eighteenth aspect of the present disclosure is the acoustic signal processing method according to the eighth aspect, in which the maximum value of the smooth function does not exceed three.

[0052] This allows the maximum value of the smooth function to be 3 or less.

[0053] An acoustic signal processing method according to a nineteenth aspect of the present disclosure is the acoustic signal processing method according to the eighth aspect, in which the minimum value of the smooth function does not fall below zero.

[0054] This allows the minimum value of the smooth function to be 0 or greater.

[0055] An acoustic signal processing method according to a twentieth aspect of the present disclosure is the acoustic signal processing method according to the eighth aspect, which includes a receiving step of receiving instructions specifying Va, the wind speed of the wind, and Vp, the instantaneous wind speed of the wind, and in the processing step, determining the smooth function so that its maximum value is Vp / Va.

[0056] This allows for a smooth function maximum and Vp / Va.

[0057] An acoustic signal processing method according to a twenty-first aspect of the present disclosure is the acoustic signal processing method according to the eighth aspect, in which the average value for the predetermined period of time is 3 seconds.

[0058] This allows the average value of the predetermined time, which is the duration of the processing frame, to be 3 seconds.

[0059] An acoustic signal processing method according to a twenty-second aspect of the present disclosure is the acoustic signal processing method according to the sixteenth aspect, in which the object is an object having a shape resembling an ear.

[0060] This makes it possible to collect aerodynamic sounds using, for example, a dummy head microphone.

[0061] A computer program according to a twenty-third aspect of the present disclosure is a computer program for causing a computer to execute the acoustic signal processing method according to any one of the first to twenty-second aspects.

[0062] This allows the computer to execute the above-described acoustic signal processing method in accordance with the computer program.

[0063] An acoustic signal processing device according to a twenty-fourth aspect of the present disclosure includes an acquisition unit that acquires sound data indicating the waveform of a reference sound, a processing unit that processes the sound data so as to change at least one of the frequency component, phase, and amplitude value of the waveform based on simulation information that simulates fluctuations in a natural phenomenon, and an output unit that outputs the processed sound data.

[0064] As a result, the sound data is processed so as to change at least one of the frequency components, phase, and amplitude of the waveform based on simulation information that simulates fluctuations in natural phenomena containing fluctuations. As a result, fluctuations occur in at least one of the frequency components, phase, and amplitude in the processed sound data, and the sound represented by the processed sound data also has fluctuations in at least one of the frequency components, phase, and amplitude. Therefore, the listener can hear a sound in which fluctuations occur in at least one of the frequency components, phase, and amplitude, and the listener can enjoy a sense of realism without feeling any discomfort. In other words, an acoustic signal processing device that can give the listener a sense of realism is realized.

[0065] Furthermore, these comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0066] Hereinafter, the embodiments will be specifically described with reference to the drawings.

[0067] The embodiments described below are all comprehensive or specific examples, and the numerical values, shapes, materials, components, arrangement and connection of the components, steps, and order of steps shown in the following embodiments are merely examples and are not intended to limit the scope of the claims.

[0068] In the following description, ordinal numbers such as "first" and "second" may be attached to elements. These ordinal numbers are attached to elements to identify the elements and do not necessarily correspond to a meaningful order. These ordinal numbers may be rearranged, newly added, or removed as appropriate.

[0069] Furthermore, each figure is a schematic diagram and is not necessarily an exact illustration. Therefore, the scales and the like do not necessarily match in each figure. In each figure, the same reference numerals are used to denote substantially the same components, and redundant explanations will be omitted or simplified.

[0070] In this specification, terms indicating relationships between elements such as verticality, and numerical ranges, are not expressions that only express a strict meaning, but also expressions that include a substantially equivalent range, for example, a difference of about a few percent.

[0071] (Embodiment 1) [Example of a device to which the acoustic processing technique or encoding / decoding technique of the present disclosure can be applied] <Stereophonic sound reproduction system> Fig. 1 is a diagram showing a stereophonic (immersive audio) reproduction system A0000, which is an example of a system to which the acoustic processing or decoding process of the present disclosure can be applied. The stereophonic sound reproduction system A0000 includes an acoustic signal processing device A0001 and an audio presentation device A0002.

[0072] The acoustic signal processing device A0001 performs acoustic processing on an audio signal emitted by a virtual sound source to generate an acoustically processed audio signal to be presented to a listener (i.e., a listener). The audio signal is not limited to voice, and may be any audible sound. The acoustic processing is, for example, signal processing performed on an audio signal to reproduce one or more sound-related effects that a sound generated from a sound source experiences between the time the sound is emitted and the time the listener hears the sound. The acoustic signal processing device A0001 performs the acoustic processing based on information describing factors that cause the above-mentioned sound-related effects. The spatial information includes, for example, information indicating the positions of the sound source, the listener, and surrounding objects, information indicating the shape of the space, parameters related to sound propagation, etc. The acoustic signal processing device A0001 is, for example, a PC (personal computer), a smartphone, a tablet, a game console, or the like.

[0073] The signal after acoustic processing is presented to a listener (user) from the audio presentation device A0002. The audio presentation device A0002 is connected to the audio signal processing device A0001 via wireless or wired communication. The audio signal after acoustic processing generated by the audio signal processing device A0001 is transmitted to the audio presentation device A0002 via wireless or wired communication. If the audio presentation device A0002 is composed of multiple devices, such as a device for the right ear and a device for the left ear, the multiple devices present sound in synchronization by communication between the multiple devices or between each of the multiple devices and the audio signal processing device A0001. The audio presentation device A0002 is, for example, headphones, earphones, or a head-mounted display worn on the listener's head, or a surround speaker composed of multiple fixed speakers.

[0074] The stereophonic sound reproduction system A0000 may be used in combination with an image presentation device or a stereoscopic video presentation device that provides an ER (Extended Reality) experience, including visual VR or AR.

[0075] 1 shows an example system configuration in which the acoustic signal processing device A0001 and the audio presentation device A0002 are separate devices, but the stereophonic sound reproduction system A0000 to which the acoustic signal processing method or decoding method of the present disclosure can be applied is not limited to the configuration shown in FIG. 1. For example, the acoustic signal processing device A0001 may be included in the audio presentation device A0002, which may perform both acoustic processing and sound presentation. Furthermore, the acoustic signal processing device A0001 and the audio presentation device A0002 may share the acoustic processing described in the present disclosure, or a server connected to the acoustic signal processing device A0001 or the audio presentation device A0002 via a network may perform part or all of the acoustic processing described in the present disclosure.

[0076] In the above description, the acoustic signal processing device A0001 is referred to as such, but if the acoustic signal processing device A0001 performs acoustic processing by decoding a bit stream generated by encoding at least a portion of the data of an audio signal or spatial information used for acoustic processing, the acoustic signal processing device A0001 may also be referred to as a decoding device.

[0077] <Example of Encoding Device> FIG. 2 is a functional block diagram showing the configuration of an encoding device A0100, which is an example of an encoding device according to the present disclosure.

[0078] The input data A0101 is data to be coded, including spatial information and / or an audio signal, and is input to the encoder A0102. Details of the spatial information will be described later.

[0079] The encoder A0102 encodes the input data A0101 to generate encoded data A0103. The encoded data A0103 is, for example, a bit stream generated by the encoding process.

[0080] The memory A 0104 stores the encoded data A 0103. The memory A 0104 may be, for example, a hard disk or a solid-state drive (SSD), or may be other memory.

[0081] In the above description, a bitstream generated by an encoding process is given as an example of the encoded data A0103 stored in the memory A0104. However, data other than a bitstream may also be used. For example, the encoding device A0100 may convert a bitstream into a predetermined data format and store the converted data in the memory A0104. The converted data may be, for example, a file or multiplexed stream storing one or more bitstreams. Here, the file may have a file format such as ISOBMFF (ISO Base Media File Format). The encoded data A0103 may also be in the form of multiple packets generated by dividing the bitstream or file. When converting the bitstream generated by the encoder A0102 into data other than the bitstream, the encoding device A0100 may include a conversion unit (not shown), or the conversion process may be performed by a CPU (Central Processing Unit).

[0082] <Example of Decoding Device> FIG. 3 is a functional block diagram showing the configuration of a decoding device A 0110, which is an example of a decoding device according to the present disclosure.

[0083] The memory A0114 stores, for example, the same data as the coded data A0103 generated by the coding device A0100. The memory A0114 reads out the stored data and inputs it as input data A0113 to the decoder A0112. The input data A0113 is, for example, a bitstream to be decoded. The memory A0114 may be, for example, a hard disk or an SSD, or may be some other memory.

[0084] Note that the decoding device A0110 may not directly use the data stored in the memory A0114 as the input data A0113, but may convert the read data and generate converted data as the input data A0113. The data before conversion may be, for example, multiplexed data storing one or more bitstreams. Here, the multiplexed data may be, for example, a file having a file format such as ISOBMFF. The data before conversion may also be in the form of multiple packets generated by dividing the bitstream or file. When converting data different from the bitstream read from the memory A0114 into a bitstream, the decoding device A0110 may be provided with a conversion unit (not shown), or the conversion process may be performed by a CPU.

[0085] The decoder A0112 decodes the input data A0113 to generate an audio signal A0111 that is presented to the listener.

[0086] <Another Example of Encoding Device> Fig. 4 is a functional block diagram showing the configuration of an encoding device A0120, which is another example of an encoding device of the present disclosure. In Fig. 4, components having the same functions as those in Fig. 2 are assigned the same reference numerals, and descriptions of these components will be omitted.

[0087] The encoding device A0120 differs from the encoding device A0100 in that the encoding device A0100 stores the encoded data A0103 in a memory A0104, whereas the encoding device A0120 includes a transmitting unit A0121 that transmits the encoded data A0103 to the outside.

[0088] The transmitter A0121 transmits a transmission signal A0122 to another device or a server based on the encoded data A0103 or data in another data format generated by converting the encoded data A0103. The data used to generate the transmission signal A0122 is, for example, the bit stream, multiplexed data, file, or packet described in the encoding device A0100.

[0089] <Another Example of Decoding Device> Fig. 5 is a functional block diagram showing the configuration of a decoding device A0130, which is another example of a decoding device of the present disclosure. In Fig. 5, components having the same functions as those in Fig. 3 are assigned the same reference numerals, and descriptions of these components will be omitted.

[0090] The decoding device A0130 differs from the decoding device A0110 in that the decoding device A0110 reads the input data A0113 from a memory A0114, whereas the decoding device A0130 includes a receiving unit A0131 that receives the input data A0113 from outside.

[0091] The receiving unit A0131 receives the received signal A0132, acquires the received data, and outputs the input data A0113 to be input to the decoder A0112. The received data may be the same as the input data A0113 to be input to the decoder A0112, or may be data in a different data format from the input data A0113. If the received data is data in a different data format from the input data A0113, the receiving unit A0131 may convert the received data into the input data A0113, or a conversion unit or CPU (not shown) included in the decoding device A0130 may convert the received data into the input data A0113. The received data is, for example, a bit stream, multiplexed data, a file, or a packet, as described in the encoding device A0120.

[0092] <Functional Description of Decoder> FIG. 6 is a functional block diagram showing the configuration of a decoder A0200, which is an example of the decoder A0112 in FIG. 3 or FIG.

[0093] The input data A0113 is an encoded bitstream, and includes encoded audio data, which is an encoded audio signal, and metadata used for acoustic processing.

[0094] The spatial information management unit A0201 acquires metadata included in the input data A0113 and analyzes the metadata. The metadata includes information describing elements that act on sounds arranged in a sound space. The spatial information management unit A0201 manages spatial information necessary for acoustic processing obtained by analyzing the metadata and provides the spatial information to the rendering unit A0203. Note that, although the information used for acoustic processing is referred to as spatial information in this disclosure, it may be referred to by other names. The information used for acoustic processing may be referred to, for example, as sound space information or as scene information. Furthermore, when the information used for acoustic processing changes over time, the spatial information input to the rendering unit A0203 may be referred to as a space state, a sound space state, a scene state, or the like.

[0095] Furthermore, spatial information may be managed for each sound space or for each scene. For example, when different rooms are represented as virtual spaces, the spatial information may be managed as scenes of different sound spaces for each room, or the spatial information may be managed as different scenes depending on the situation being represented, even for the same space. In managing the spatial information, an identifier for identifying each piece of spatial information may be assigned. The spatial information data may be included in a bitstream, which is one form of input data, or the bitstream may include an identifier for the spatial information, and the spatial information data may be acquired from a source other than the bitstream. When the bitstream includes only the identifier for the spatial information, the identifier for the spatial information may be used during rendering to acquire the spatial information data stored in the memory of the acoustic signal processing device A0001 or an external server as input data.

[0096] Note that the information managed by the spatial information management unit A0201 is not limited to information included in the bitstream. For example, the input data A0113 may include, as data not included in the bitstream, data indicating the characteristics or structure of a space acquired from a software application or server providing VR or AR. Furthermore, for example, the input data A0113 may include, as data not included in the bitstream, data indicating the characteristics or position of a listener or object. Furthermore, the input data A0113 may include, as information indicating the position of the listener, information acquired by a sensor provided in a terminal including a decoding device, or information indicating the position of the terminal estimated based on information acquired by the sensor. In other words, the spatial information management unit A0201 may communicate with an external system or server to acquire spatial information and the position of the listener. Furthermore, the spatial information management unit A0201 may acquire clock synchronization information from an external system and execute a process of synchronizing with the clock of the rendering unit A0203. Note that the space in the above description may be a virtually formed space, i.e., a VR space, or may be a real space (i.e., a physical space) or a virtual space corresponding to the real space, i.e., an AR or MR (Mixed Reality). The virtual space may also be called a sound field or a sound space. Furthermore, the information indicating a position in the above description may be information such as coordinate values ​​indicating a position within the space, information indicating a relative position with respect to a predetermined reference position, or information indicating the movement or acceleration of a position within the space.

[0097] The audio data decoder A0202 decodes the encoded audio data included in the input data A0113 to obtain an audio signal.

[0098] The encoded audio data acquired by the stereophonic sound reproduction system A0000 is a bitstream encoded in a predetermined format such as MPEG-H 3D Audio (ISO / IEC 23008-3). Note that MPEG-H 3D Audio is merely one example of an encoding method that can be used to generate the encoded audio data included in the bitstream, and the encoded audio data may be included in a bitstream encoded in another encoding method. For example, the encoding method used may be a lossy codec such as MP3 (MPEG-1 Audio Layer-3), AAC (Advanced Audio Coding), WMA (Windows Media Audio), AC3 (Audio Codec-3), or Vorbis, or a lossless codec such as ALAC (Apple Lossless Audio Codec) or FLAC (Free Lossless Audio Codec), or any other encoding method may be used. For example, PCM (pulse code modulation) data may be a type of encoded audio data. In this case, the decoding process may be, for example, a process of converting an N-bit binary number into a number format (e.g., floating-point format) that can be processed by the rendering unit A0203, where the number of quantization bits of the PCM data is N.

[0099] The rendering unit A0203 receives an audio signal and spatial information as input, performs acoustic processing on the audio signal using the spatial information, and outputs an audio signal A0111 after acoustic processing.

[0100] Before starting rendering, the spatial information management unit A0201 reads metadata of the input signal, detects rendering items such as objects or sounds defined in the spatial information, and sends them to the rendering unit A0203. After starting rendering, the spatial information management unit A0201 grasps changes over time in the spatial information and the position of the listener, and updates and manages the spatial information. The spatial information management unit A0201 then sends the updated spatial information to the rendering unit A0203. The rendering unit A0203 generates and outputs an audio signal to which acoustic processing has been applied based on the audio signal included in the input data A0113 and the spatial information received from the spatial information management unit A0201.

[0101] The spatial information update process and the audio signal output process with added acoustic processing may be executed in the same thread, or the spatial information management unit A0201 and the rendering unit A0203 may be assigned to independent threads. When the spatial information update process and the audio signal output process with added acoustic processing are executed in different threads, the thread startup frequency may be set individually, or the processes may be executed in parallel.

[0102] By having the spatial information management unit A0201 and the rendering unit A0203 execute their processes in different, independent threads, computational resources can be preferentially allocated to the rendering unit A0203. Therefore, in the case of sound output processing in which even the slightest delay cannot be tolerated, for example, sound output processing in which a delay of even one sample (0.02 msec) would cause a popping noise, can be safely performed. In this case, the allocation of computational resources to the spatial information management unit A0201 is limited. However, compared to audio signal output processing, updating spatial information is a less frequent process (e.g., a process such as updating the listener's facial orientation). Therefore, unlike audio signal output processing, updating spatial information does not necessarily require an instantaneous response, and therefore limiting the allocation of computational resources does not significantly affect the acoustic quality experienced by the listener.

[0103] The spatial information may be updated periodically at preset times or intervals, or when preset conditions are met. The spatial information may be updated manually by a listener or a sound space manager, or may be triggered by a change in an external system. For example, if a listener operates a controller to instantly warp the position of their avatar, instantly advance or reverse the time, or if a virtual space manager suddenly changes the environment of the venue, the thread in which the spatial information management unit A0201 is located may be activated as a one-off interrupt process in addition to being activated periodically.

[0104] The role of the information update thread that executes the spatial information update process is, for example, to update the position or orientation of the listener's avatar placed in the virtual space based on the position or orientation of the VR goggles worn by the listener, and to update the position of objects moving in the virtual space. These tasks are handled within a processing thread that runs relatively infrequently, on the order of several tens of Hz. Processing to reflect the properties of direct sound may be performed in such an infrequently occurring processing thread. This is because the properties of direct sound change less frequently than the frequency of audio processing frames for audio output. Doing so can relatively reduce the computational load of the process and also avoid the risk of pulsive noise occurring when information is updated at an unnecessarily fast frequency.

[0105] FIG. 7 is a functional block diagram showing the configuration of a decoder A0210, which is another example of the decoder A0112 in FIG. 3 or 5.

[0106] The decoder A0210 shown in Fig. 7 differs from the decoder A0200 shown in Fig. 6 in that the input data A0113 includes an unencoded audio signal rather than encoded audio data. The input data A0113 includes a bitstream including metadata and an audio signal.

[0107] The spatial information management unit A0211 is the same as the spatial information management unit A0201 in FIG. 6, and therefore a description thereof will be omitted.

[0108] The rendering unit A0213 is the same as the rendering unit A0203 in FIG. 6, and therefore a description thereof will be omitted.

[0109] 7 is referred to as a decoder A0210 in the above description, it may also be referred to as an audio processing unit that performs audio processing. Furthermore, a device including an audio processing unit may also be referred to as an audio processing device rather than a decoding device. Furthermore, the audio signal processing device A0001 may also be referred to as an audio processing device.

[0110] <Physical configuration of audio signal processing device> Fig. 8 is a diagram showing an example of the physical configuration of an audio signal processing device. Note that the audio signal processing device in Fig. 8 may be a decoding device. Furthermore, part of the configuration described here may be provided in an audio presentation device A0002. Furthermore, the audio signal processing device shown in Fig. 8 is an example of the audio signal processing device A0001 described above.

[0111] The acoustic signal processing device of FIG. 8 includes a processor, a memory, a communication IF, a sensor, and a speaker.

[0112] The processor may be, for example, a CPU (Central Processing Unit), a DSP (Digital Signal Processor), or a GPU (Graphics Processing Unit), and the CPU, DSP, or GPU may execute a program stored in a memory to perform the audio processing or decoding processing of the present disclosure. Alternatively, the processor may be a dedicated circuit that performs signal processing on an audio signal, including the audio processing of the present disclosure.

[0113] The memory may be configured, for example, by RAM (Random Access Memory) or ROM (Read Only Memory). The memory may also include a magnetic storage medium such as a hard disk or a semiconductor memory such as an SSD (Solid State Drive). The term "memory" may also refer to an internal memory built into a CPU or GPU.

[0114] The communication IF (Interface) is a communication module compatible with a communication method such as Bluetooth (registered trademark) or WIGIG (registered trademark). The acoustic signal processing device shown in Fig. 8 has a function of communicating with other communication devices via the communication IF and acquires a bitstream to be decoded. The acquired bitstream is stored in a memory, for example.

[0115] The communication module is composed of, for example, a signal processing circuit and an antenna corresponding to the communication method. In the above example, Bluetooth (registered trademark) or WIGIG (registered trademark) is used as the communication method, but other communication methods such as LTE (Long Term Evolution), NR (New Radio), or Wi-Fi (registered trademark) may also be supported. Furthermore, the communication IF may be a wired communication method such as Ethernet (registered trademark), USB (Universal Serial Bus), or HDMI (registered trademark) (High-Definition Multimedia Interface) instead of the wireless communication method described above.

[0116] The sensor performs sensing to estimate the position or orientation of the listener. Specifically, the sensor estimates the position and / or orientation of the listener based on one or more detection results of the position, orientation, movement, velocity, angular velocity, or acceleration of a part or the entire body of the listener, such as the head, and generates position information indicating the position and / or orientation of the listener. Note that the position information may be information indicating the position and / or orientation of the listener in real space, or information indicating a displacement of the position and / or orientation of the listener based on the position and / or orientation of the listener at a predetermined time. Furthermore, the position information may be information indicating the position and / or orientation relative to the stereophonic sound reproduction system A0000 or an external device equipped with the sensor.

[0117] The sensor may be, for example, an imaging device such as a camera or a ranging device such as LiDAR (Light Detection and Ranging), and may capture an image of the listener's head movement and detect the movement of the listener's head by processing the captured image. Alternatively, the sensor may be a device that performs position estimation using wireless signals of any frequency band, such as millimeter waves.

[0118] The audio signal processing device shown in Fig. 8 may acquire position information from an external device equipped with a sensor via a communication IF. In this case, the audio signal processing device may not include a sensor. Here, the external device may be, for example, the audio presentation device A0002 described in Fig. 1 or a 3D video playback device worn on the listener's head. In this case, the sensor may be a combination of various sensors, such as a gyro sensor and an acceleration sensor.

[0119] The sensor may, for example, detect the angular velocity of rotation around at least one of three mutually perpendicular axes in the sound space as the axis of rotation as the speed of movement of the listener's head, or may detect the acceleration of displacement with at least one of the three axes as the direction of displacement.

[0120] For example, the sensor may detect the amount of rotation about at least one of three mutually orthogonal axes in the sound space as the rotation axis, or the amount of displacement about at least one of the three axes as the displacement direction, as the amount of movement of the listener's head. Specifically, the sensor detects 6 DoF (position (x, y, z) and angle (yaw, pitch, roll)) as the position of the listener. The sensor is configured by combining various sensors used for detecting movement, such as a gyro sensor and an acceleration sensor.

[0121] The sensor may be any device capable of detecting the position of the listener, such as a camera or a GPS (Global Positioning System) receiver. Alternatively, the sensor may use location information obtained by performing self-position estimation using LiDAR (Laser Imaging Detection and Ranging). For example, when the audio signal reproduction system is implemented by a smartphone, the sensor is built into the smartphone.

[0122] The sensor may also include a temperature sensor such as a thermocouple that detects the temperature of the acoustic signal processing device shown in Figure 8, and a sensor that detects the remaining charge of a battery provided in or connected to the acoustic signal processing device.

[0123] A speaker has, for example, a diaphragm, a drive mechanism such as a magnet or a voice coil, and an amplifier, and presents an audio signal after acoustic processing to a listener as sound. The speaker operates the drive mechanism in response to the audio signal (more specifically, a waveform signal indicating the waveform of the sound) amplified by the amplifier, and the drive mechanism vibrates the diaphragm. In this way, the diaphragm vibrating in response to the audio signal generates sound waves, which propagate through the air and reach the listener's ears, causing the listener to perceive the sound.

[0124] Note that, although the description has been given here of an example in which the acoustic signal processing device shown in FIG. 8 includes a speaker and presents an audio signal after acoustic processing via the speaker, the means for presenting the audio signal is not limited to the above configuration. For example, the audio signal after acoustic processing may be output to an external audio presentation device A0002 connected via a communication module. Communication via the communication module may be wired or wireless. As another example, the acoustic signal processing device shown in FIG. 8 may include a terminal for outputting an analog audio signal, and a cable such as earphones may be connected to the terminal to present the audio signal from the earphones. In the above case, the audio signal is reproduced by headphones, earphones, a head-mounted display, a neck speaker, a wearable speaker, a surround speaker composed of multiple fixed speakers, or the like, which are worn on the head or part of the body of the listener, which is the audio presentation device A0002.

[0125] <Physical Configuration of Encoding Device> Fig. 9 is a diagram showing an example of the physical configuration of an encoding device. The encoding device shown in Fig. 9 is an example of the encoding devices A0100 and A0120 described above.

[0126] The encoding device of FIG. 9 includes a processor, a memory, and a communication IF.

[0127] The processor may be, for example, a CPU (Central Processing Unit) or a DSP (Digital Signal Processor), and the CPU or GPU may execute a program stored in a memory to perform the encoding process of the present disclosure. Alternatively, the processor may be a dedicated circuit that performs signal processing on an audio signal, including the encoding process of the present disclosure.

[0128] The memory may be configured, for example, by RAM (Random Access Memory) or ROM (Read Only Memory). The memory may also include a magnetic storage medium such as a hard disk or a semiconductor memory such as an SSD (Solid State Drive). The term "memory" may also refer to an internal memory built into a CPU or GPU.

[0129] The communication IF (Interface) is a communication module compatible with a communication method such as Bluetooth (registered trademark) or WIGIG (registered trademark). The encoding device has a function of communicating with other communication devices via the communication IF and transmits an encoded bitstream.

[0130] The communication module is composed of, for example, a signal processing circuit and an antenna corresponding to the communication method. In the above example, Bluetooth (registered trademark) or WIGIG (registered trademark) is used as the communication method, but other communication methods such as LTE (Long Term Evolution), NR (New Radio), or Wi-Fi (registered trademark) may also be supported. Furthermore, the communication IF may be a wired communication method such as Ethernet (registered trademark), USB (Universal Serial Bus), or HDMI (registered trademark) (High-Definition Multimedia Interface) instead of the wireless communication method described above.

[0131] [Configuration] The following describes the configuration of the acoustic signal processing device 100 according to Embodiment 1. Fig. 10 is a block diagram showing the functional configuration of the acoustic signal processing device 100 according to this embodiment.

[0132] The acoustic signal processing device 100 according to this embodiment is a device for acquiring, processing, and outputting sound data indicating a waveform of a reference sound. By outputting the sound data, a listener can hear the sound indicated by the sound data. The acoustic signal processing device 100 according to this embodiment is a device that is used in various applications in virtual spaces, such as virtual reality or augmented reality (VR or AR), for example.

[0133] The reference sound may be any sound, for example, a sound related to a natural phenomenon. In this embodiment, the natural phenomenon is not particularly limited as long as it is a phenomenon that occurs in the natural world, but examples include phenomena such as the wind blowing, the water flowing in a river, and animal behavior. Examples of sounds related to natural phenomena include the sound caused by the wind blowing, the murmuring sound of water flowing in a river, and the cries of animals.

[0134] Here, when we focus on sounds caused by wind blowing, one example is aerodynamic sound, which is generated when wind collides with an object in a virtual space. This aerodynamic sound is generated when wind reaches and collides with, for example, the ear of a listener. In this way, aerodynamic sound is sound that originates from wind blowing in a virtual space.

[0135] In this embodiment, the reference sound is an aerodynamic sound generated by wind W. However, the reference sound is not limited to this, and may be the murmuring sound of flowing water in a river or the cry of an animal.

[0136] The wind in the virtual space is, for example, wind caused by an object in the virtual space.

[0137] 11 is a diagram showing an electric fan FN, which is an example of an object according to this embodiment, and a listener L. When the object is an object that can blow air, such as an electric fan FN, the aerodynamic sound is aerodynamic sound that is generated when the wind W generated by the electric fan FN reaches the listener L. More specifically, the aerodynamic sound is sound that is generated when the wind W blown out from the electric fan FN reaches the listener L, depending on the shape of the listener L's ear, for example.

[0138] Furthermore, for example, if the object is a moving body (for example, a vehicle), the aerodynamic sound is generated when wind W, which is generated by the movement of the object's position, reaches the listener L.

[0139] Furthermore, the wind W in the virtual space is, for example, a wind that occurs naturally in real space and is reproduced in the virtual space (hereinafter referred to as natural wind), and the location of its generation cannot be identified in the virtual space. When the wind W in the virtual space is natural wind, it can also be said to be wind that is not caused by an object.

[0140] Note that the object according to the present embodiment is not limited to the electric fan FN. The object in the virtual space is not particularly limited as long as it is included in the content (here, video as an example) displayed on the display unit 300 that displays the content executed in the virtual space.

[0141] The object may be, for example, a moving object that generates wind by moving its position. Moving objects include, for example, objects that represent plants and animals, man-made objects, or natural objects. Examples of objects that represent man-made objects include vehicles, bicycles, and airplanes. Examples of objects that represent man-made objects include sports equipment such as baseball bats and tennis rackets, and furniture such as desks, chairs, and grandfather clocks. Note that, for example, the object may be at least one of an object that can move within the content and an object that can be moved.

[0142] Furthermore, for example, the object may be an object that can blow air. Examples of such objects include a circulator, a hand fan, and an air conditioner, in addition to the electric fan FN.

[0143] The object may also be an object that generates a sound. The sound generated by the object is a sound indicated by sound data associated with the object (hereinafter, sometimes referred to as object sound data). For example, if the object is an electric fan FN, the sound generated by the object is a motor sound generated by a motor possessed by the electric fan FN. For example, if the object is an ambulance, the sound generated by the object is a siren sound emitted by the ambulance.

[0144] The acoustic signal processing device 100 processes sound data (aerodynamic sound data) that indicates the waveform of a reference sound, which is an aerodynamic sound in a virtual space, and outputs the processed sound to the headphones 200. Note that, hereinafter, sound data that indicates the waveform of the reference sound (aerodynamic sound) may be referred to as aerodynamic sound data.

[0145] Next, the headphones 200 will be described.

[0146] The headphones 200 are devices that play back aerodynamic sound, and are audio output devices that present the aerodynamic sound to the listener L. More specifically, the headphones 200 play back the aerodynamic sound based on the aerodynamic sound data output by the acoustic signal processing device 100. This allows the listener L to hear the aerodynamic sound. Note that instead of the headphones 200, other output channels such as speakers may be used.

[0147] As shown in FIG. 10, the headphones 200 include a head sensor unit 201 and an output unit 202 .

[0148] The head sensor unit 201 senses the position of the listener L, which is determined by the horizontal coordinates and vertical height in the virtual space, and outputs second position information indicating the position of the listener L of the aerodynamic sound in the virtual space to the acoustic signal processing device 100.

[0149] The head sensor unit 201 may sense 6 DoF information of the head of the listener L. For example, the head sensor unit 201 may be an inertial measurement unit (IMU), an accelerometer, a gyroscope, a magnetic sensor, or a combination thereof.

[0150] The output unit 202 is a device that reproduces the sound that reaches the listener L in the sound reproduction space. More specifically, the output unit 202 reproduces the aerodynamic sound based on aerodynamic sound data that indicates the aerodynamic sound output from the acoustic signal processing device 100.

[0151] Next, the display unit 300 will be described.

[0152] The display unit 300 is a display device that displays content (video) including objects in a virtual space. The process by which the display unit 300 displays content will be described later. The display unit 300 is realized by a display panel such as a liquid crystal panel or an organic EL (Electro Luminescence) panel, for example.

[0153] Next, the acoustic signal processing device 100 shown in Fig. 10 will be described. In this embodiment, the acoustic signal processing device 100 acquires sound data (aerodynamic sound data) indicating the waveform of a reference sound, which is an aerodynamic sound in a virtual space, processes the sound data, and outputs the processed sound to headphones 200.

[0154] As shown in FIG. 10, the acoustic signal processing device 100 includes an acquisition unit 110 , a processing unit 120 , an output unit 130 , a storage unit 140 , and a reception unit 150 .

[0155] The acquisition unit 110 acquires sound data indicating the waveform of a reference sound (aerodynamic sound). Fig. 12 is a diagram showing sound data according to this embodiment. As shown in Fig. 12, the sound data is data indicating a waveform in which, for example, time and amplitude are indicated, and in this case, it is aerodynamic sound data.

[0156] The sound data (aerodynamic sound data) is stored in the storage unit 140 , and the acquisition unit 110 acquires the sound data (aerodynamic sound data) stored in the storage unit 140 .

[0157] The acquisition unit 110 acquires first position information indicating the position of an object (e.g., an electric fan FN). If the object generates a sound, the acquisition unit 110 acquires object sound data indicating the sound. The acquisition unit 110 also acquires shape information indicating the shape of the object.

[0158] The acquiring unit 110 acquires second position information. As described above, the second position information is information indicating the position of the listener L in the virtual space.

[0159] The acquisition unit 110 may acquire, for example, sound data indicating the waveform of a reference sound, first position information, object sound data, shape information, and second position information from an input signal. Furthermore, the acquisition unit 110 may acquire sound data indicating the waveform of a reference sound, first position information, object sound data, shape information, and second position information from other sources. The input signal will be described below. Furthermore, hereinafter, sound data indicating the waveform of a reference sound (aerodynamic sound data) and object sound data may be collectively referred to as sound data.

[0160] The input signal may be composed of, for example, spatial information, sensor information, and sound data (audio signal). The above information and sound data may be included in a single input signal, or may be included in multiple separate signals. The input signal may include a bitstream composed of sound data and metadata (control information), in which case the metadata may include information identifying the spatial information and sound data.

[0161] The sound data indicating the waveform of the reference sound, the first position information, the object sound data, the shape information, and the second position information described above may be included in the input signal. More specifically, the first position information and the shape information may be included in the spatial information, and the second position information may be generated based on information acquired from sensor information. The sensor information may be acquired from the head sensor unit 201 or from another external device.

[0162] The spatial information is information about the sound space (three-dimensional sound field) created by the stereophonic sound reproduction system A0000, and is composed of information about objects included in the sound space and information about the listener. Objects include sound source objects that emit sound and act as sound sources, and non-sound-emitting objects that do not emit sound. Non-sound-emitting objects function as obstacle objects that reflect sounds emitted by sound source objects, but sound source objects may also function as obstacle objects that reflect sounds emitted by other sound source objects. Obstacle objects may also be called reflecting objects.

[0163] Information commonly assigned to sound source objects and non-sound-producing objects includes position information, shape information, and the rate of attenuation of the volume when the object reflects sound.

[0164] The position information is expressed as coordinate values ​​on three axes, for example, the X-axis, Y-axis, and Z-axis, in Euclidean space, but does not necessarily have to be three-dimensional information. The position information may also be two-dimensional information expressed as coordinate values ​​on two axes, for example, the X-axis and Y-axis. The position information of an object is determined by a representative position of a shape expressed by a mesh or voxels.

[0165] The shape information may include information about the surface material.

[0166] The attenuation rate may be expressed as a real number equal to or less than 1 or equal to or greater than 0, or may be expressed as a negative decibel value. Since sound volume is not amplified by reflection in real space, a negative decibel value is set for the attenuation rate. However, for example, to create an eerie feeling in an unreal space, an attenuation rate of equal to or greater than 1, i.e., a positive decibel value, may be set. Furthermore, different values ​​for the attenuation rate may be set for each frequency band constituting multiple frequency bands, or values ​​may be set independently for each frequency band. Furthermore, if an attenuation rate is set for each type of material on the object surface, a corresponding attenuation rate value may be used based on information about the surface material.

[0167] Furthermore, the information commonly assigned to the sound source object and the non-sound-producing object may include information indicating whether the object belongs to a living thing or whether the object is a moving object, etc. If the object is a moving object, the position information may move over time, and the changed position information or the amount of change is transmitted to the rendering units A0203 and A0213.

[0168] The information about the sound source object includes, in addition to the information commonly assigned to the sound source object and the non-sound generating object, object sound data and information necessary for radiating the object sound data into the sound space. The object sound data is data representing the sound perceived by the listener, including information about the frequency and intensity of the sound. The object sound data is typically a PCM signal, but may also be data compressed using an encoding method such as MP3. In this case, the signal must be decoded at least before reaching the generation unit (the generation unit 907 described later in FIG. 34 ). Therefore, the rendering units A0203 and A0213 may include a decoding unit (not shown). Alternatively, the signal may be decoded by the audio data decoder A0202.

[0169] At least one object sound data may be set for one sound source object, and multiple object sound data may be set for one sound source object. Furthermore, identification information for identifying each object sound data may be assigned, and the identification information for the object sound data may be stored as metadata as information about the sound source object.

[0170] The information necessary for radiating object sound data into the sound space may include, for example, information on the reference volume that serves as a reference when playing the object sound data, information on the position of the sound source object, information on the orientation of the sound source object, and information on the directionality of the sound emitted by the sound source object.

[0171] The reference volume information may be, for example, the effective value of the amplitude value of the object sound data at the sound source position when the object sound data is radiated into the sound space, and may be expressed as a floating-point decibel (dB) value. For example, when the reference volume is 0 dB, the reference volume information may indicate that sound is to be radiated into the sound space from the position indicated by the information regarding the position at the same volume without increasing or decreasing the volume of the signal level indicated by the object sound data. When the reference volume is −6 dB, the reference volume information may indicate that sound is to be radiated into the sound space from the position indicated by the information regarding the position with the volume of the signal level indicated by the object sound data reduced to about half. The reference volume information may be assigned to one object sound data or to multiple object sound data collectively.

[0172] The volume information included in the information necessary to radiate object sound data into a sound space may include, for example, information indicating time-series fluctuations in the volume of the sound source. For example, if the sound space is a virtual conference room and the sound source is a speaker, the volume transitions intermittently over a short period of time. Expressed more simply, this can be said to mean that sound portions and silent portions occur alternately. Furthermore, if the sound space is a concert hall and the sound source is a performer, the volume is maintained for a certain period of time. Furthermore, if the sound space is a battlefield and the sound source is an explosive, the volume of the explosion sound increases for only a moment and then remains silent. In this way, the volume information of the sound source includes not only information about the volume of the sound but also information about the transition of the volume of the sound, and such information may be used as information indicating the properties of the object sound data.

[0173] Here, the information on loudness transitions may be data indicating frequency characteristics in a time series. The information on loudness transitions may be data indicating the duration of a section in which sound is present. The information on loudness transitions may be data indicating a time series of the duration of a section in which sound is present and the duration of a section in which sound is absent. The information on loudness transitions may be data listing, in a time series, multiple sets of durations during which the amplitude of a sound signal can be considered steady (considered to be roughly constant) and data on the amplitude values ​​of the signal during those periods. The information on loudness transitions may be data indicating the durations during which the frequency characteristics of a sound signal can be considered steady. The information on loudness transitions may be data listing, in a time series, multiple sets of durations during which the frequency characteristics of a sound signal can be considered steady and data on the frequency characteristics during those periods. The data format of the information on loudness transitions may be, for example, data indicating the outline of a spectrogram. Furthermore, the volume that serves as a reference for the frequency characteristics may be the reference volume. The information on the reference volume and the information indicating the properties of the object sound data may be used to calculate the volume of the direct sound or reflected sound to be perceived by the listener, as well as in a selection process to select whether or not to perceive it.

[0174] Orientation information is typically expressed using yaw, pitch, and roll. Alternatively, the roll rotation may be omitted and the information may be expressed using azimuth (yaw) and elevation (pitch). Orientation information may change over time, and if it does change, it is transmitted to the rendering units A0203 and A0213.

[0175] The information about the listener is information about the position and orientation of the listener in sound space. The position information is expressed as positions on the X-, Y-, and Z-axes in Euclidean space, but does not necessarily have to be three-dimensional information and may be two-dimensional information. The information about orientation is typically expressed using yaw, pitch, and roll. Alternatively, the information about orientation may be expressed using azimuth (yaw) and elevation (pitch) without the roll rotation. The position information and orientation information may change over time, and if they change, they are transmitted to the rendering units A0203 and A0213.

[0176] The sensor information includes the amount of rotation or displacement detected by a sensor worn by the listener, as well as the position and orientation of the listener. The sensor information is transmitted to the rendering units A0203 and A0213, which update the position and orientation information of the listener based on the sensor information. The sensor information may be, for example, position information obtained by a mobile terminal performing self-position estimation using a GPS, a camera, or LiDAR (Laser Imaging Detection and Ranging). Information obtained from an external source other than a sensor via a communication module may also be detected as sensor information. Information indicating the temperature of the acoustic signal processing device 100 and information indicating the remaining battery capacity may be acquired from the sensor. Information indicating the computing resources (CPU capacity, memory resources, PC performance) of the acoustic signal processing device 100 or the audio presentation device A0002 may be acquired in real time as sensor information.

[0177] In the present embodiment, the acquisition unit 110 acquires sound data indicating the waveform of the reference sound, the first position information, the object sound data, and the shape information from the storage unit 140, but this is not limiting and the acquisition may be from a device other than the acoustic signal processing device 100 (for example, a server device 500 such as a cloud server). Furthermore, the acquisition unit 110 acquires the second position information from the headphones 200 (more specifically, the head sensor unit 201), but this is not limiting.

[0178] The first location information will be further described.

[0179] As described above, the object in the virtual space is included in the content (image) displayed on the display unit 300, and in this embodiment is, for example, the electric fan FN.

[0180] The first position information is information indicating the position of the electric fan FN in the virtual space at a given time. Note that in the virtual space, for example, the electric fan FN may be moved by a user picking up the electric fan FN and moving it. Therefore, the acquisition unit 110 continuously acquires the first position information. The acquisition unit 110 acquires the first position information, for example, each time the space information is updated by the space information management units A0201 and A0211.

[0181] Furthermore, sound data that indicates the waveform of a reference sound (aerodynamic sound) and sound data that includes object sound data associated with an object will be described.

[0182] The sound data including the object sound data and aerodynamic sound data described in this specification may be a sound signal such as PCM (Pulse Code Modulation) data, but is not limited to this and may be any information that indicates the properties of the sound.

[0183] As an example, if a sound signal is a noise signal with a volume of X decibels, the sound data related to the sound signal may be the PCM data itself representing the sound signal, or may be data consisting of information indicating that the component is a noise signal and information indicating that the volume is X decibels. As another example, if a sound signal is a noise signal with a frequency component having a predetermined peak / dip characteristic, the sound data related to the sound data may be the PCM data itself representing the sound signal, or may be data consisting of information indicating that the component is a noise signal and information indicating the peak / dip of the frequency component.

[0184] In this specification, a sound signal based on sound data means PCM data representing the sound data.

[0185] Furthermore, aerodynamic sound data, which is sound data that indicates the waveform of the reference sound, is stored in advance in storage unit 140, as described above. Aerodynamic sound is sound generated when wind W collides with an object, and in this case, it is sound generated when wind W collides with the ear of listener L. Aerodynamic sound data is data that captures the sound generated when wind W collides with a human ear or an object (model) that has a shape that imitates a human ear. In this embodiment, the aerodynamic sound data is data that captures the sound that occurs when wind reaches an object (model) that imitates a human ear. A dummy head microphone or the like is used as a model that imitates a human ear, and the aerodynamic sound data is collected.

[0186] Next, the shape information will be described.

[0187] Shape information is information that indicates the shape of an object in virtual space. The shape information indicates the shape of the object, and more specifically, indicates the three-dimensional shape of the object as a rigid body. The shape of the object is indicated by, for example, a sphere, a rectangular parallelepiped, a cube, a polyhedron, a cone, a pyramid, a cylinder, a prism, or a combination thereof. Note that the shape information may be expressed as, for example, mesh data, or a set of multiple faces each consisting of a voxel, a three-dimensional point cloud, or vertices with three-dimensional coordinates.

[0188] The first position information includes object identification information for identifying the object. The object sound data also includes object identification information, and the shape information also includes object identification information.

[0189] Therefore, even if the acquisition unit 110 acquires the first position information, object sound data, and shape information separately, the object indicated by each of the first position information, object sound data, and shape information can be identified by referencing the object identification information included in each of the first position information, object sound data, and shape information. For example, in this example, it is easy to identify that the object indicated by each of the first position information, object sound data, and shape information is the same fan FN. In other words, by referencing the three object identification information, it becomes clear that the first position information, object sound data, and shape information acquired by the acquisition unit 110 are information related to the fan FN. Therefore, the first position information, object sound data, and shape information are linked as information indicating the fan FN.

[0190] Next, the second location information will be described.

[0191] The listener L can move in the virtual space. The second position information is information indicating the position of the listener L in the virtual space at a given time. Since the listener L can move in the virtual space, the acquisition unit 110 continuously acquires the second position information. The acquisition unit 110 acquires the second position information, for example, each time the spatial information is updated by the spatial information management units A0201 and A0211.

[0192] The sound data indicating the waveform of the reference sound, the first position information, the object sound data, the shape information, and the second position information may be included in the metadata, control information, or header information included in the input signal. When the sound data including the object sound data and the aerodynamic sound data is a sound signal (PCM data), information identifying the sound signal may be included in the metadata, control information, or header information, or the sound signal may be included in information other than the metadata, control information, or header information. In other words, the acoustic signal processing device 100 (more specifically, the acquisition unit 110) may acquire the metadata, control information, or header information included in the input signal and perform acoustic processing based on the metadata, control information, or header information. The acoustic signal processing device 100 (more specifically, the acquisition unit 110) may acquire the sound data indicating the waveform of the reference sound, the first position information, the object sound data, the shape information, and the second position information, and the source of acquisition is not limited to the input signal. The sound data including the object sound data and the aerodynamic sound data and the metadata may be stored in a single input signal, or may be stored separately in multiple input signals.

[0193] Furthermore, sound signals other than sound data including object sound data and aerodynamic sound data may be stored in the input signal as audio content information. The audio content information may be encoded using MPEG-H 3D Audio (ISO / IEC 23008-3) (hereinafter referred to as MPEG-H 3D Audio) or other encoding technology. The encoding technology is not limited to MPEG-H 3D Audio, and other well-known technologies may also be used. Information such as sound data indicating the waveform of the reference sound, first position information, object sound data, shape information, and second position information may also be subject to encoding processing.

[0194] In other words, the acoustic signal processing device 100 acquires sound signals and metadata contained in an encoded bitstream. Audio content information is acquired and decoded in the acoustic signal processing device 100. In this embodiment, the acoustic signal processing device 100 functions as a decoder (e.g., decoders A0200 and A0210) included in a decoding device (e.g., decoding devices A0110 and A0130), and more specifically, functions as rendering units A0203 and A0213 included in the decoder. Note that the term "audio content information" in the present disclosure is to be interpreted as information including sound data indicating the waveform of the sound signal itself or a reference sound, first position information, object sound data, shape information, and second position information, in accordance with the technical content.

[0195] The acquisition unit 110 outputs the acquired sound data indicating the waveform of the reference sound, the first position information, the object sound data, the shape information, and the second position information to the processing unit 120 and the output unit 130 .

[0196] The processing unit 120 processes the sound data based on simulation information that simulates fluctuations in a natural phenomenon so as to change at least one of the frequency component, phase, and amplitude value of the waveform indicated by the sound data that indicates the waveform of the reference sound. In this embodiment, since the reference sound is aerodynamic sound generated by wind W, the natural phenomenon in the simulation information is the blowing of wind W. The fluctuations in the natural phenomenon refer to fluctuations in the wind W, and more specifically, fluctuations in the wind speed of the wind W. Note that the fluctuations in the natural phenomenon may also be fluctuations in the direction (wind direction) of the wind W, etc.

[0197] In real space, fluctuations in natural phenomena include fluctuations (for example, 1 / f fluctuations). Therefore, the simulated information is information that simulates fluctuations in natural phenomena that include fluctuations. In this embodiment, the simulated information is information that simulates fluctuations in the wind speed of the wind W, and more specifically, is information that expresses the fluctuations included in the fluctuations in the wind speed of the wind W.

[0198] More specifically, the simulation information is a smooth function that simulates fluctuations in wind speed. Here, the processing unit 120 determines, as the simulation information, a smooth function that simulates fluctuations in wind speed.

[0199] A smooth function means that it is differentiable and continuous, in other words, a smooth function is a function that does not have sharp points.

[0200] 13 is a diagram showing an example of a smooth function according to the present embodiment. As shown in FIG. 13, the smooth function is a sine curve as an example, but is not limited to this and may be a cosine curve or the like.

[0201] The processing unit 120 processes the sound data so as to change at least one of the frequency component, phase, and amplitude value of the waveform based on the value indicated by the smooth function determined by the processing unit 120. For example, the processing unit 120 processes the sound data so as to shift the frequency component of the waveform to a frequency proportional to the value indicated by the smooth function that simulates fluctuations in wind speed.

[0202] The value indicated by the smooth function is the value on the vertical axis shown in Figure 13, and is information indicating the ratio between the wind speed of the aerodynamic sound, which is the reference sound, and the wind speed of the aerodynamic sound indicated by the sound data after processing by the processing unit 120. In other words, the value indicated by the smooth function is a value indicating the ratio between the wind speed of the aerodynamic sound before processing and the wind speed of the aerodynamic sound after processing.

[0203] The processing unit 120 processes the sound data and outputs it to the output unit 130 .

[0204] The output unit 130 outputs the sound data processed by the processing unit 120. Here, the output unit 130 outputs the processed aerodynamic sound data to the headphones 200. This allows the headphones 200 to reproduce the aerodynamic sound indicated by the output aerodynamic sound data. In other words, the listener L can hear the aerodynamic sound.

[0205] The storage unit 140 is a storage device that stores computer programs executed by the acquisition unit 110, the processing unit 120, and the output unit 130, as well as aerodynamic sound data.

[0206] The reception unit 150 receives operations from a user (e.g., a creator of content executed in a virtual space) of the acoustic signal processing device 100. Specifically, the reception unit 150 is realized by a hardware button, but may also be realized by a touch panel or the like.

[0207] Here, the shape information according to the present embodiment will be explained again. The shape information is information used to generate an image of an object in a virtual space and is also information indicating the shape of the object (electric fan FN). In other words, the shape information is also information used to generate content (image) to be displayed on the display unit 300.

[0208] The acquisition unit 110 also outputs the acquired shape information to the display unit 300. The display unit 300 acquires the shape information output by the acquisition unit 110. The display unit 300 further acquires attribute information indicating attributes (such as color) of the object (electric fan FN) other than the shape in the virtual space. The display unit 300 may acquire the attribute information directly from a device other than the acoustic signal processing device 100 (the server device 500), or may acquire the attribute information from the acoustic signal processing device 100. The display unit 300 generates and displays content (video) based on the acquired shape information and attribute information.

[0209] Hereinafter, operation examples 1 and 2 of the acoustic signal processing method performed by the acoustic signal processing device 100 will be described.

[0210] [Operation Example 1] FIG. 14 is a flowchart of operation example 1 of the acoustic signal processing device 100 according to this embodiment.

[0211] 14 , first, the receiving unit 150 receives an operation to indicate that the simulation information is a smooth function that simulates fluctuations in wind speed (S10). The receiving unit 150 receives this operation from, for example, the user of the acoustic signal processing device 100.

[0212] Next, the acquisition unit 110 acquires sound data indicating the waveform of the reference sound (S20). In this operation example, the reference sound is an aerodynamic sound generated by wind, and the sound data indicating the waveform of the reference sound is aerodynamic sound data. This step S20 corresponds to the acquisition step.

[0213] The processing unit 120 determines a smooth function that simulates fluctuations in wind speed as simulation information that simulates fluctuations in natural phenomena (S30). The processing unit 120 may determine the simulation information in accordance with the operation received in step S10. In this operation example, the processing unit 120 determines the smooth function shown in FIG. 13 as the simulation information.

[0214] Furthermore, the processing unit 120 processes the sound data (aerodynamic sound data) so as to change at least one of the frequency component, phase, and amplitude value of the waveform based on the value (ratio) indicated by the smooth function determined by the processing unit 120 (S40).

[0215] Note that steps S30 and S40 correspond to processing steps.

[0216] The processing unit 120 outputs the processed sound data (aerodynamic sound data) to the output unit 130 .

[0217] The output unit 130 outputs the sound data (aerodynamic sound data) processed by the processing unit 120 to the headphones 200 (S50). Note that this step S50 corresponds to an output step.

[0218] This allows the listener L to hear the aerodynamic sound output from the headphones 200.

[0219] Here, the processing performed by the processing unit 120 in steps S30 and S40 will be described in more detail.

[0220] FIG. 15 is a diagram for explaining the processing performed by the processing unit 120 according to this embodiment.

[0221] Fig. 15(a) is a diagram showing the sound data shown in Fig. 12 (aerodynamic sound data D1 before processing) and the smooth function shown in Fig. 13. As Fig. 15(a) shows, the horizontal axis, which is the time axis, corresponds to the aerodynamic sound data D1 before processing and the smooth function.

[0222] Fig. 15(b) is a diagram for explaining the processing in the area surrounded by the dashed-dotted rectangle in Fig. 15(a). Fig. 15(b) shows an enlarged view of the aerodynamic noise data D1 before processing, the smooth function, and the aerodynamic noise data D11 after processing.

[0223] The aerodynamic sound data D1 before processing is indicated by a plurality of black dots in (b) of Fig. 15. Each of the plurality of black dots corresponds to the aerodynamic sound data D1 before processing shown in (a) of Fig. 15. It can also be said that each of the plurality of black dots is a sample point of the aerodynamic sound data D1 before processing.

[0224] The processing unit 120 first performs a first process, which will be described below.

[0225] The processing unit 120 determines an interpolation function that interpolates between one black point and another black point adjacent to the one black point. The interpolation function is, for example, a spline function, but is not limited to this and may be any known function. The processing unit 120 may also perform linear interpolation (straight-line interpolation) between one black point and another black point adjacent to the one black point, in which case the load of calculation processing is reduced. As shown in (b) of Figure 15, in the first processing, all of the space between two adjacent black points is interpolated.

[0226] As a result, a line is drawn by interpolating between one black dot and another black dot adjacent to that black dot, as shown in (b) of Fig. 15. The interval between multiple black dots before processing is defined as "1".

[0227] Next, the processing unit 120 performs a second process, which will be described below.

[0228] In the second processing, the processing unit 120 reads the value of one black dot, which is the pre-processing aerodynamic sound data D1 at time t, and determines the read value as the post-processing aerodynamic sound data D11 at that time t. The post-processing aerodynamic sound data D11 is indicated by multiple white dots (open dots) in Figure 15(c).

[0229] Next, the processing unit 120 reads the value of the smooth function for each unit time. For example, the processing unit 120 reads "0.5", "0.5", "0.4999", "0.4998", etc. as the values ​​of the smooth function.

[0230] The processing unit 120 determines the value of the smooth function read at the time t as the stride, and reads the value of the interpolation function at a position that is the stride ahead in time from one black dot in the unprocessed aerodynamic sound data D1 at the time t.

[0231] Furthermore, the processing unit 120 determines the value of the read interpolation function as the value of the processed aerodynamic sound data D11. At this time, the processing unit 120 determines the spacing of the processed aerodynamic sound data D11 (plurality of white dots) so that the spacing between the processed aerodynamic sound data D11 (plurality of black dots) is the same value as the spacing between the pre-processing aerodynamic sound data D1 (plurality of black dots), in other words, so that it is "1." In this way, the second processing is performed.

[0232] A specific example of this second process will be described, focusing on time t1.

[0233] The processing unit 120 reads the value of the black point B1, which is the aerodynamic sound data D1 before processing at time t1, and determines the read value as the value of the white point B11 in the aerodynamic sound data D11 after processing at time t1. In other words, the processing unit 120 uses the read value of the black point B1 as the value of the white point B11 as is.

[0234] Furthermore, the processing unit 120 reads the value of the smooth function at time t1, which is 0.5, and determines this as the stride. The aerodynamic sound data D1 before processing at time t1 is indicated by the black point B1, and the processing unit 120 reads the value of the interpolation function at a position 0.5 in time ahead from the black point B1, which is the aerodynamic sound data D1 before processing. This position is indicated as position P1 in Figure 15(b).

[0235] The processing unit 120 then determines the value of the interpolation function that has been read (the value indicated at position P1) as the value of the processed aerodynamic sound data D11. The processing unit 120 determines the spacing of the processed aerodynamic sound data D11 (plurality of white dots) so that the spacing of the processed aerodynamic sound data D11 (plurality of white dots) is the same value as the spacing of the pre-processing aerodynamic sound data D1 (plurality of black dots), which is "1."

[0236] As a result of the first and second processes, the processed aerodynamic sound data D11 has a shape that is elongated in the horizontal direction from the aerodynamic sound data D1 before processing. Therefore, the processed aerodynamic sound data D11 is sound data in which the frequency components are shifted to lower frequencies compared to the aerodynamic sound data D1 before processing.

[0237] FIG. 16 is another diagram for explaining the processing performed by the processing unit 120 according to the present embodiment.

[0238] FIG. 16(a) is a diagram showing the sound data (aerodynamic sound data) shown in FIG. 12 and the smooth function shown in FIG. 13, similar to FIG. 15(a).

[0239] Figures 16(b) and 16(c) are diagrams for explaining the processing in the area surrounded by the dashed-dotted rectangle in Figure 16(a). Figures 16(b) and 16(c) each show an enlarged view of the aerodynamic noise data D1 before processing, the smooth function, and the aerodynamic noise data D11 after processing.

[0240] The unprocessed aerodynamic sound data D1 shown in (b) and (c) of Figure 16 is also subjected to the same processing as that described using (b) of Figure 15. That is, first processing and second processing are performed.

[0241] 16(b), the processing unit 120 reads values ​​of the smooth function such as "1", "1", "1.0001", and "1.0002". Because the read values ​​of the smooth function are around 1, the processed aerodynamic sound data D11 has the same shape as the aerodynamic sound data D1 before processing. Therefore, the processed aerodynamic sound data D11 is sound data in which the frequency components are hardly shifted compared to the aerodynamic sound data D1 before processing.

[0242] 16(c), the processing unit 120 reads values ​​of the smooth function such as "1.5", "1.5", "1.4999", and "1.4998". Because the value of the read smooth function is approximately 1.5, the processed aerodynamic sound data D11 has a shape that is the same as the aerodynamic sound data D1 before processing but shrunk in the horizontal direction. Therefore, the processed aerodynamic sound data D11 is sound data in which the frequency components have been shifted to higher frequencies compared to the aerodynamic sound data D1 before processing.

[0243] As described above, the simulated information is information that simulates fluctuations in natural phenomena that include fluctuations, and more specifically, is information that expresses fluctuations due to changes in the wind speed of wind W, and in this operation example, is information that is represented by a smooth function.

[0244] In this operational example, sound data (aerodynamic sound data) indicating the waveform of a reference sound is processed so that the frequency components of the waveform change based on simulation information that simulates fluctuations in natural phenomena that include fluctuations. As a result, fluctuations occur in the frequency components of the processed aerodynamic sound data, and fluctuations also occur in the frequency components of the aerodynamic sound indicated by the processed aerodynamic sound data. Therefore, the listener L can hear aerodynamic sound with fluctuations in such frequency components, and can obtain a sense of realism without feeling any discomfort.

[0245] In addition, in step S40 of the first operation example, the following processing may be performed.

[0246] As described above, in step S40, the stride may be determined as follows: Here, the sampling frequency of the aerodynamic sound data before it is processed by the processing unit 120 is set to Fsc, the sampling frequency of the aerodynamic sound data output by the output unit 130 is set to Fso, and Fsc and Fso are assumed to be different values.

[0247] In this case, the stride should satisfy the following formula:

[0248] Smooth function value × (Fsc / Fso)

[0249] The effect of the stride satisfying the above formula will be explained below.

[0250] For example, if Fso is 48 kHz, it is advisable to downsample Fsc from 48 kHz to 16 kHz. This makes it possible to reduce the memory size to one-third when aerodynamic sound data of the same duration is stored in the storage unit 140. Furthermore, this also makes it possible to triple the duration of the aerodynamic sound data that is output when the same memory size is used, thereby reducing the sense of incongruity at the joins between the aerodynamic sound data.

[0251] Next, we will explain how aliasing distortion can be reduced. FIG. 17 is a diagram showing sound data according to this embodiment. More specifically, (a) and (b) of FIG. 17 are diagrams showing the frequency characteristics of aerodynamic sound data before processing (for example, the aerodynamic sound data D1 before processing shown in FIG. 15 ). In FIG. 17(a), the horizontal axis is a logarithmic axis, while in FIG. 17(b), the horizontal axis is a linear axis. Also, (c) of FIG. 17 is a diagram showing the frequency characteristics in which the frequency components of the aerodynamic sound data shown in FIG. 17(b) are shifted toward higher frequencies. Here, the frequency components in FIG. 17(c) are shifted to twice the frequency of the frequency components in FIG. 17(b). For example, the 2000 Hz frequency component in FIG. 17(b) is shifted toward higher frequencies to become the 4000 kHz frequency component in FIG. 17(c).

[0252] 17(a) and (b), the solid line shows the frequency characteristics when the sampling frequency of the aerodynamic sound data before processing is 16 kHz, and the dashed-dotted line shows the frequency characteristics when the sampling frequency of the aerodynamic sound data before processing is 48 kHz. Note that the dashed-dotted line overlaps with the solid line in the low frequency range, and is therefore not shown.

[0253] As shown in FIG. 17, aerodynamic sound data often has a characteristic structure in the low frequency range, and in the high frequency range, this component decreases monotonically.

[0254] 17(c), the solid line shows the frequency characteristics when the sampling frequency of the shifted aerodynamic sound data is 16 kHz, and the dashed-dotted line shows the frequency characteristics when the sampling frequency of the shifted aerodynamic sound data is 48 kHz. Note that the dashed-dotted line overlaps with the solid line in the low frequency range, so it is not shown in the figure.

[0255] When the sampling frequency of the aerodynamic sound data indicated by the dashed line is 48 kHz, frequency components exist in the frequency range of 12 kHz or higher in FIG. 17(b), and aliasing distortion indicated by the dashed line appears in FIG. 17(c).

[0256] When the sampling frequency of the aerodynamic sound data indicated by the solid line is 16 kHz, there are no frequency components in the frequency range of 12 kHz or higher in FIG. 17(b), and therefore aliasing distortion does not appear in FIG. 17(c).

[0257] In this way, it is possible to suppress the occurrence of aliasing distortion due to frequency shift.

[0258] Another advantage is that there is almost no increase in the computational resources required to reduce the memory size and suppress the appearance of aliasing distortion.

[0259] The above corresponds to the effect of the stride satisfying the above formula.

[0260] Furthermore, in Operation Example 1 of this embodiment, the aerodynamic sound data is stored in advance in the storage unit 140, but this is not limited to this. For example, the processing unit 120 may generate the aerodynamic sound data. For example, the processing unit 120 may generate the aerodynamic sound data by acquiring a noise signal and processing the acquired noise signal with each of a plurality of band emphasis filters.

[0261] [Operation Example 2] As described above, in Operation Example 1, the sound data (aerodynamic sound data) is processed so as to change the frequency components of the waveform, but this is not limiting. In Operation Example 2, the sound data (aerodynamic sound data) is processed so as to change the amplitude value of the waveform.

[0262] That is, in Operation Example 2, steps S10 to S30 are performed in the same way as in Operation Example 1. Then, in step S40, the processing unit 120 processes the sound data (aerodynamic sound data) so as to change the amplitude value of the waveform based on the value (ratio) indicated by the smooth function determined by the processing unit 120.

[0263] The amplitude value of a waveform indicates the level of volume of the aerodynamic sound indicated by the aerodynamic sound data indicated by that waveform. The following relationship exists between aerodynamic sound and the wind speed of the wind W that generates that aerodynamic sound: the volume of aerodynamic sound is proportional to the α power of the wind speed of the wind W. Therefore, the processing unit 120 processes the sound data so that the amplitude value of the waveform changes in proportion to the α power of the value indicated by the determined smooth function. The value of α differs depending on the type of aerodynamic sound.

[0264] For example, there is aerodynamic sound generated when a rod-shaped object cuts through the wind. This aerodynamic sound is generated when a baseball bat is swung. The volume of this type of aerodynamic sound is proportional to the sixth power of the wind speed (see Non-Patent Document 1).

[0265] Another example is aerodynamic noise, which occurs when wind enters the gap between one object and another. This type of aerodynamic noise is known as cavity noise. The volume of this type of aerodynamic noise is proportional to the fourth power of the wind speed (see Non-Patent Document 1).

[0266] Here, let R be the value indicated by the smooth function that simulates fluctuations in wind speed. In the case of any of the above types of aerodynamic noise, the volume of the aerodynamic noise is amplified or attenuated by a value that corresponds to R^α. That is, when R is greater than 1, the volume is amplified, and when R is less than 1, the volume is attenuated. It should be noted here that when the volume of aerodynamic noise is proportional to the α power of the wind speed, the volume of aerodynamic noise fluctuates very sharply. This sharp fluctuation will be explained using FIG. 18.

[0267] Fig. 18 is a diagram showing the relationship between R, which is a value indicated by a smooth function according to this embodiment, and the amplification rate and attenuation rate of the volume of aerodynamic noise. The two-dot chain line in Fig. 18 shows the relationship between R and the volume of aerodynamic noise when α is 6. Note that in the vicinity of R = 1, the two-dot chain line overlaps with the solid line.

[0268] As shown by the two-dot chain line in Figure 18, the amplification factor exceeds 30 dB when R = 2.0, and the attenuation factor falls below -30 dB when R = 0.5. To faithfully reproduce such a steep fluctuation would require expensive playback equipment with an extremely wide dynamic range, which would be excessive for the acoustic presentation of a virtual space.

[0269] To avoid the need for such expensive playback equipment, a threshold value r is used. As an example, in FIG. 18 , 1.3 is used as the threshold value r. For example, it is preferable that the amplification factor (attenuation factor) G be different between the section where (1 / r)<R<r and the section where R<(1 / r) and r<R. In FIG. 18 , the section where (1 / r)<R<r is indicated by a dashed rectangle. In FIG. 18 , the dashed and solid lines represent the case where the amplification factor (attenuation factor) G is different between the section where (1 / r)<R<r and the section where R<(1 / r) and r<R.

[0270] The dashed and solid lines indicate that in the range of (1 / r)<R<r, the amplification factor (attenuation factor) G satisfies the following formula.

[0271] G = R^α

[0272] Furthermore, the dashed dotted line indicates that in the ranges of R<(1 / r) and r<R, the amplification factor (attenuation factor) G satisfies the following formula.

[0273] G={r×(R / r)}^α = (r)^α×(R / r)^b

[0274] Furthermore, by setting b to a value smaller than α, a tendency close to the amplification rate (attenuation rate) G = R^α, i.e., the correct tendency, is achieved near R = 1.0, and monotonic amplification (monotonic attenuation) occurs outside the vicinity of R = 1.0, and sudden fluctuations can be avoided.

[0275] The dashed-dotted line in Figure 18 satisfies the conditions of r = 1.3 and b = 2.0. However, in this dashed-dotted line, the tendency of amplification and attenuation changes discontinuously at R = r and R = 1 / r. This may cause an unnatural feeling near R = r and R = 1 / r.

[0276] Therefore, instead of making b a constant, the value of b may be the same as α when R=r, and may be gradually made smaller than α as R increases. The solid line in Fig. 18 indicates that in the ranges R<(1 / r) and r<R, the amplification factor (attenuation factor) G satisfies the following formula:

[0277] G = (r)^α × (R / r)^b where b = α^(r / R)

[0278] By increasing or decreasing the volume according to the solid line in Figure 18, the volume can be changed sensitively in accordance with subtle fluctuations in wind speed (slight fluctuations around R = 1), and sudden fluctuations in the volume due to increases or decreases in R can be avoided.

[0279] The value of α can be set arbitrarily by a user of the acoustic signal processing device 100 (for example, the creator of the content executed in the virtual space). That is, the receiving unit 150 can receive an operation from the creator specifying the value of α, and the processing unit 120 can determine the value specified by the received operation as the value of α. By setting the value of α to a value that is significantly different from an academically correct value, such as 0.7, 1.0, 1.5, or 2.0, but that is a value that produces a "plausible" increase or decrease in the volume of aerodynamic sound in the virtual space, sudden fluctuations can be avoided. The values ​​of r and b can also be determined in a similar manner.

[0280] Furthermore, while the aerodynamic sound data was processed so that the frequency components changed in Operation Example 1 and the amplitude values ​​changed in Operation Example 2, this is not limiting. For example, the aerodynamic sound data may be processed so that the phase of the waveform changes. In this case, the processing unit 120 processes the sound data so that the phase of the waveform changes according to the value indicated by the determined smooth function.

[0281] Furthermore, at least one of the frequency component, phase, and amplitude number of the waveform may be changed. For example, two of the frequency component, phase, and amplitude number of the waveform may be changed, or all of the frequency component, phase, and amplitude number of the waveform may be changed.

[0282] In operation examples 1 and 2, the processing unit 120 may divide the sound data (aerodynamic sound data) indicating the waveform of the reference sound acquired by the acquisition unit 110 into processing frames F of a predetermined time, and process the sound data for each divided processing frame F.

[0283] Fig. 19 is a diagram showing divided aerodynamic sound data according to this embodiment. In Fig. 19, the aerodynamic sound data is divided into multiple processing frames F. Furthermore, the predetermined time Ts for each of the multiple processing frames F may be the same, or, as shown in Fig. 19, may be different from one another. In other words, Fig. 19 shows processing frames F1 to F6, which are an example of processing frames F, and predetermined times Ts1 to Ts6, which are an example of predetermined times Ts. Each of the predetermined times Ts1 to Ts6 is different from one another.

[0284] Furthermore, in the operation examples 1 and 2, the smooth function shown in FIG. 13 is used as the simulation information, but a different smooth function may be used.

[0285] For example, in step S30 of operation examples 1 and 2, the processing unit 120 determines a smooth function that simulates fluctuations in wind speed as simulation information that simulates fluctuations in a natural phenomenon. In this case, the processing unit 120 may determine the smooth function so that the parameters specifying the smooth function vary irregularly. Furthermore, the processing unit 120 determines the parameters specifying the smooth function for each divided processing frame F. That is, for example, the processing unit 120 determines the parameters specifying the smooth function corresponding to processing frame F1 shown in FIG. 19 . Similarly, the processing unit 120 determines the parameters specifying the smooth function corresponding to processing frame F2, the parameters specifying the smooth function corresponding to processing frame F3, the parameters specifying the smooth function corresponding to processing frame F4, the parameters specifying the smooth function corresponding to processing frame F5, and the parameters specifying the smooth function corresponding to processing frame F6.

[0286] Furthermore, the processing unit 120 determines a smooth function for each divided processing frame F so that the value of the smooth function is 1.0 at the first time and the last time of the processing frame F. For example, in the smooth function corresponding to the processing frame F2 of the predetermined time Ts2, the value indicated by the smooth function is 1.0 at times t2 and t3.

[0287] If the smooth function shown in FIG. 13 is F(t), F(t) is expressed by the following equation.

[0288] F(t)=H×{sin[2π×(t / T)^(x)]}^(y)+1.0 (0.0≦t<T)

[0289] An example of a parameter specifying a smooth function is the time from the first time of a processing frame F to the last time of the processing frame F, which is T in the above formula. For example, in the smooth function corresponding to processing frame F2 shown in Fig. 19, this parameter is the time from time t2 to time t3. In other words, if the smooth function is a sine curve, this parameter corresponds to one period.

[0290] Another example of a parameter specifying a smooth function is a value related to the maximum value of the smooth function, which is H in the above formula. When the smooth function is a sine curve as shown in this embodiment, another example of the parameter can also be said to be a value that determines the maximum value of the smooth function.

[0291] Another example of a parameter that specifies a smooth function is a parameter that varies the position at which the smooth function reaches its maximum value, which is x in the above formula.

[0292] Another example of a parameter that specifies a smooth function is a parameter that changes the steepness of the change in the smooth function, which is y in the above formula.

[0293] The processing unit 120 determines the parameters so that they vary irregularly, thereby determining a smooth function. For example, the processing unit 120 may determine the parameters based on random numbers.

[0294] For example, the processing unit 120 may include a random number sequence generator, and the processing unit 120 may change parameters depending on the output sequence. A truly random number sequence inherently has no regularity or reproducibility. However, since it is difficult to achieve this on a computer, the sequence generated by the random number sequence generator may be a pseudo-random number sequence generated through a deterministic calculation process. For example, a pseudo-random number sequence such as that generated by the rand() function in the C language may be used, or any other known algorithm for generating pseudo-random numbers may be used. Furthermore, a finite-length random number sequence, a finite-length pseudo-random number sequence, or a finite-length sequence created to create a sense of irregularity may be stored in the storage unit 140 and used repeatedly to create a long-term pseudo-random number sequence.

[0295] The receiving unit 150 may also receive an operation specifying a parameter value from a user (e.g., a creator of content executed in a virtual space) of the acoustic signal processing device 100. The processing unit 120 may determine the value specified by the operation received by the receiving unit 150 as the parameter.

[0296] 20A and 20B are diagrams showing other examples of two smooth functions according to this embodiment, in which the smooth functions shown in (a) and (b) are determined so that the parameters specifying the smooth functions vary irregularly.

[0297] In this case, the parameters may be determined by simulating the nature of the wind speed of the wind W. As described above, fluctuations in the wind speed of the wind W include fluctuations, that is, in real space, the wind speed is not constant but fluctuates. For example, the wind W may blow at the listener L at a first wind speed, and then blow at a second wind speed that is different from the first wind speed. In this way, the parameters may be determined by simulating the nature of the wind speed fluctuating.

[0298] The maximum value of the smooth function should not exceed 3, and the minimum value of the smooth function should not fall below 0. In other words, the value indicated by the smooth function should be between 0 and 3. The parameters should be determined so that the value indicated by the smooth function is as described above.

[0299] The reason why the maximum value of the smooth function should not exceed 3 is as follows: In real space, fluctuations in the wind speed of wind W include fluctuations, and wind W may blow at a momentary strong wind speed (instantaneous wind speed). The wind speed is, for example, a 10-minute average wind speed, and the instantaneous wind speed is, for example, a 3-second average wind speed. In such cases, it is known that the instantaneous wind speed is approximately 1.5 to 3 times the wind speed. The value indicated by the smooth function is the ratio between the wind speed of the aerodynamic sound, which is the reference sound, and the wind speed of the aerodynamic sound indicated by the sound data after processing. By setting the maximum value of the smooth function to 3 or less, wind W with a momentary strong wind speed (instantaneous wind speed), or more specifically, the aerodynamic sound caused by that wind W, can be reproduced in virtual space.

[0300] Furthermore, the wind speed of the wind W is Va, and the instantaneous wind speed of the wind W is Vp. In this case, the processing unit 120 determines a smooth function such that the maximum value of the smooth function is Vp / Va. More specifically, the processing unit 120 determines parameters that specify the smooth function such that the maximum value of the smooth function is Vp / Va. For example, the receiving unit 150 receives an instruction specifying Va, which is the wind speed of the wind W, and Vp, which is the instantaneous wind speed of the wind W, and the processing unit 120 determines parameters that specify the smooth function such that the maximum value of the smooth function is Vp / Va in accordance with the received instruction.

[0301] At this time, it is preferable that an image be displayed on the display unit of the acoustic signal processing device 100, in which words representing the strength of the wind W are linked to the wind speed and instantaneous wind speed of the wind W indicated by the words. In the image, for example, if the words are "slightly strong wind," the image is linked to a wind speed of "10 or more but less than 15 m / s" and an instantaneous wind speed of "20 m / s." In addition, in the image, for example, if the words are "strong wind," the image is linked to a wind speed of "15 or more but less than 20 m / s" and an instantaneous wind speed of "30 m / s."

[0302] A user of the acoustic signal processing device 100 (e.g., a creator of content executed in a virtual space) visually recognizes the image displayed on the display unit. The receiving unit 150 then receives from the user an instruction to specify words representing the strength of the wind W. The processing unit 120 determines the wind speed and instantaneous wind speed associated with the words specified by the received instruction as Va and Vp, and determines parameters specifying the smooth function so that the maximum value of the smooth function is Vp / Va.

[0303] Even in this case, it is possible to reproduce the wind W blowing momentarily at a strong wind speed (instantaneous wind speed), more specifically, the aerodynamic sound caused by the wind W, in the virtual space.

[0304] Furthermore, the processing unit 120 divides the aerodynamic sound data into processing frames F of a predetermined time, and the average value of this predetermined time is preferably 3 seconds. As described above, the instantaneous wind speed is, for example, the 3-second average wind speed. Therefore, by setting the average value over the predetermined time to 3 seconds, the predetermined time can be made to correspond to the time (i.e., 3 seconds) over which the instantaneous wind speed is measured, and the strong wind speed (instantaneous wind speed) of the wind W that blows momentarily in the virtual space can be made to resemble the wind that blows in real space.

[0305] Here, a smooth function when the above four parameters are changed will be described in more detail with reference to FIG.

[0306] Fig. 21 is a diagram showing an example in which parameters specifying a smooth function according to this embodiment are changed. Fig. 21(a) shows the same smooth function as Fig. 13. Fig. 21(b) shows a smooth function in which T in the above formula is changed. Fig. 21(c) shows a smooth function in which H in the above formula is changed. Fig. 21(d) shows a smooth function in which x in the above formula is changed. Fig. 21(e) shows a smooth function in which y in the above formula is changed.

[0307] Incidentally, in the above operational examples 1 and 2, the processed aerodynamic sound data is output to the headphones 200, which are one output channel, but this is not limited to this. For example, the processed aerodynamic sound data may be output to each of the first output channel and the second output channel. The first output channel outputs aerodynamic sound to one ear of the listener L, and the second output channel outputs aerodynamic sound to the other ear of the listener L.

[0308] In such a case, the processing unit 120 determines a first parameter and a second parameter, each of which specifies a smooth function. The processing unit 120 processes the acquired sound data (aerodynamic sound data) so as to change at least one of the frequency component, phase, and amplitude value of the waveform, based on the smooth function specified by the first parameter determined by the processing unit 120. This processed aerodynamic sound data is referred to as aerodynamic sound data A. The processing unit 120 processes the acquired sound data (aerodynamic sound data) so as to change at least one of the frequency component, phase, and amplitude value of the waveform, based on the smooth function specified by the second parameter determined by the processing unit 120. This processed aerodynamic sound data is referred to as aerodynamic sound data B.

[0309] The output unit 130 outputs the sound data (aerodynamic sound data A) processed based on the smooth function specified by the determined first parameter to a first output channel. The output unit 130 outputs the sound data (aerodynamic sound data B) processed based on the smooth function specified by the determined second parameter to a second output channel.

[0310] 22A and 22B are diagrams illustrating another example of two smooth functions according to the present embodiment. (a) shows a smooth function specified by a first parameter, and (b) shows a smooth function specified by a second parameter. Here, the first output channel is a channel output to the right ear, and the second output channel is a channel output to the left ear.

[0311] This allows different aerodynamic sound data to be output for each output channel.

[0312] In this case, the first parameter and the second parameter may be determined by simulating the nature of the direction (wind direction) of the wind W. As described above, fluctuations in the direction (wind direction) of the wind W include fluctuations, that is, in real space, the wind direction is not constant but fluctuates. For example, the wind W may blow from the right side of the listener L, and then blow from directly in front of the listener L. In this way, the first parameter and the second parameter may be determined by simulating the nature of the wind direction fluctuating and changing.

[0313] (Modification of Embodiment 1) A description will now be given of a modification of Embodiment 1. The following description will focus on differences from Embodiment 1, and descriptions of commonalities will be omitted or simplified.

[0314] [Configuration] First, the configuration of an acoustic signal processing device 100a according to this modification will be described. Fig. 23 is a block diagram showing the functional configuration of an acoustic signal processing device 100a according to this modification.

[0315] The acoustic signal processing device 100 a according to this modification has the same configuration as the acoustic signal processing device 100 according to the first embodiment, except that it includes a processing unit 120 a instead of the processing unit 120 .

[0316] The processing unit 120 a includes a first processing unit 121 and a second processing unit 122 .

[0317] The first processing unit 121 performs the processing of step S30 described with reference to Fig. 14. The second processing unit 122 performs the following processing based on the value indicated by the smooth function determined by the first processing unit 121.

[0318] 24 is a block diagram showing the functional configuration of the second processing unit 122 according to this modification. The second processing unit 122 includes a sampling rate conversion unit 1001, a rearrangement unit 1002, and a connection unit 1003.

[0319] The sampling rate conversion unit 1001 acquires sound data (aerodynamic sound data) indicating the waveform of the reference sound and the value indicated by the smooth function determined by the first processing unit 121 .

[0320] Based on the value indicated by the acquired smooth function, the sampling rate conversion unit 1001 converts the sampling rate of the aerodynamic sound data for each processing frame F. When the sampling rate of the aerodynamic sound data is Fs, the interval (sample interval) between sample points of the aerodynamic sound data before processing (for example, the aerodynamic sound data D1 before processing shown in FIG. 15) is 1 / Fs seconds.

[0321] If the value indicated by the smooth function is 0.5, the sampling rate conversion unit 1001 upsamples the aerodynamic sound data so that the sample interval is 0.5 times (1 / (2·Fs)), that is, the sampling rate is 2·Fs. On the other hand, if the value indicated by the smooth function is 2, the sampling rate conversion unit 1001 downsamples the aerodynamic sound data so that the sample interval is doubled (2 / Fs), that is, the sampling rate is Fs / 2. The sampling rate conversion unit outputs the aerodynamic sound data whose sampling rate has been converted to the rearrangement unit 1002.

[0322] The rearrangement unit 1002 performs processing to return the interval between aerodynamic sound data after sampling rate conversion to Fs. With this processing, when the value indicated by the smooth function is greater than 1, the aerodynamic sound data is played back at a faster speed. Conversely, when the value indicated by the smooth function is less than 1, the aerodynamic sound data is played back at a slower speed. This shifts the frequency components of the aerodynamic sound data toward higher frequencies or lower frequencies, making it possible to generate aerodynamic sound with a natural sense of fluctuation. Next, the rearrangement unit 1002 outputs the aerodynamic sound data, whose sample point positions have been rearranged, to the connection unit 1003.

[0323] The connection unit 1003 performs processing to prevent discontinuities from occurring between processing frames F. This processing will be explained using two processing frames F. The two processing frames F are a previous processing frame and a current processing frame. The current processing frame is the processing frame F that is the target of processing by the processing unit 120 at a given time, and the previous processing frame is the processing frame F immediately before the current processing frame.

[0324] The connection unit 1003 performs a windowed addition process on a plurality of sample points located after in time of the rearranged aerodynamic sound data generated from the aerodynamic sound data of the previous processing frame, and a plurality of sample points located before in time of the rearranged aerodynamic sound data generated from the aerodynamic sound data of the current processing frame. This process avoids discontinuities between processing frames F that occur due to fluctuations in the values ​​indicated by the smooth function.

[0325] FIG. 25 is a diagram showing aerodynamic sound data according to this modified example. FIG. 26 is a conceptual diagram of processing by the second processing unit 122 according to this modified example. The aerodynamic sound data is processed in units of processing frames F. Two adjacent processing frames F are set so that they partially overlap each other. This is to prevent discontinuities by performing windowed addition on one or more rearward sample points among the multiple rearranged sample points in the previous processing frame and one or more forward sample points among the multiple rearranged sample points in the current processing frame. For example, as shown in FIG. 25 , two adjacent processing frames Fn and Fn+1 partially overlap each other. More specifically, the two processing frames Fn and Fn+1 overlap from time t14 to time t13. Processing frame Fn corresponds to the previous processing frame, and processing frame Fn+1 corresponds to the current processing frame.

[0326] An example will be described in which the sampling rate of the aerodynamic sound data is Fs, the value indicated by the smooth function for processing frame Fn is 0.5, and the value indicated by the smooth function for processing frame Fn+1 is 0.75. In processing frame Fn, the value indicated by the smooth function is 0.5, so sampling rate conversion is performed so that the sampling rate becomes 2·Fs (sample interval is 1 / (2·Fs)). The rearrangement unit 1002 then rearranges the positions of the sample points after sampling rate conversion so that the sample interval becomes 1 / Fs, in other words, so that the original positions are restored. Therefore, the time length of the sample points after rearrangement is twice the time length of the sample points of the aerodynamic sound data converted by the sampling rate conversion unit 1001.

[0327] Then, one or more rearmost sample points among the rearranged sample points are windowed and added with one or more frontmost sample points among the rearranged sample points in the current processing frame, and the result is output. In this example, the value indicated by the smooth function in processing frame n+1 is 0.75, so the time length of the rearranged sample points is 4 / 3 times the time length of the sample points of the aerodynamic sound data converted by the sampling rate conversion unit 1001. Note that rearranged sample points in sections for which windowed addition is not performed are output as sound data as is.

[0328] Here, the sampling rate conversion unit 1001 will be described in more detail with reference to FIG.

[0329] 27 is a block diagram showing the functional configuration of a sampling rate conversion unit 1001 according to this modification. The sampling rate conversion unit 1001 includes an up-sampling unit 1021, a low-pass filter unit 1022, a down-sampling unit 1023, and an XY setting unit 1024.

[0330] The upsampling unit 1021 acquires sound data (aerodynamic sound data), and the XY setting unit 1024 acquires values ​​indicated by a smooth function. The XY setting unit 1024 sets an upsampling value X used by the upsampling unit 1021 and a downsampling value Y used by the downsampling unit 1023. Here, if the upsampling value is X, the upsampling unit 1021 upsamples the aerodynamic sound data by X times. If the downsampling value is Y, the downsampling unit 1023 downsamples the aerodynamic sound data by 1 / Y times. The settings of X and Y in the XY setting unit 1024 are determined so that X and Y are the smallest integers among the combinations of X and Y that result in Y / X being a value indicated by the smooth function. For example, when the value indicated by the smooth function is 0.5, (X, Y) = (2, 1), when the value indicated by the smooth function is 0.75, (X, Y) = (4, 3), and when the value indicated by the smooth function is 1.5, (X, Y) = (2, 3). Note that when X = 1, upsampling section 1021 does not perform upsampling processing and the aerodynamic sound data is output as is, and when Y = 1, downsampling section 1023 does not perform downsampling processing and the aerodynamic sound data is output as is.

[0331] The upsampling unit 1021 inserts X-1 zero values ​​between sample points. The downsampling unit 1023 thins out every Y sample points and outputs them. The low-pass filter unit 1022 performs the following processing to prevent aliasing distortion from occurring. Here, the sampling rate of the aerodynamic sound data is assumed to be Fs, and the sampling rate of the aerodynamic sound data after sampling rate conversion is assumed to be Fs'. At this time, the low-pass filter unit 1022 processes the aerodynamic sound data output from the upsampling unit 1021 using a low-pass filter with characteristics such that the cutoff frequency is min(Fs, Fs') / 2.

[0332] Furthermore, the temporal fluctuation patterns of the values ​​indicated by the smooth function are illustrated. Here, the values ​​indicated by the smooth function are expressed as one of five values. Here, fluctuation pattern 1 and fluctuation pattern 2 are explained.

[0333] In fluctuation pattern 1, the value indicated by the smooth function is one of 0.25, 0.5, 1, 2, and 4. In fluctuation pattern 2, the value indicated by the smooth function is one of 0.5, 0.75, 1, 1.5, and 2. The values ​​or numbers that the smooth function can take are not limited to those exemplified here.

[0334] FIG. 28 is a state transition diagram of the values ​​indicated by the smooth function according to this modification. That is, FIG. 28 shows the temporal transition of the values ​​indicated by the smooth function. Each circle represents a state, and when in the state p(0), p(0) is output as the value indicated by the smooth function. Furthermore, a(e, f) indicates the probability of transitioning from state e to state f. To represent natural sound fluctuations, it is desirable to set the state to only allow transitions to the state itself or adjacent states, as in this example. However, depending on the application, more severe fluctuations may be desired, so any transition may be specified, not limited to this example.

[0335] In this modification, processing may be performed to vary the amplitude values ​​of the aerodynamic sound data acquired by the sampling rate conversion unit 1001.

[0336] 29 is a block diagram showing another functional configuration of the acoustic signal processing device 100a according to this modification. Here, the processing unit 120a of the acoustic signal processing device 100a has a second processing unit 122b instead of the second processing unit 122. The second processing unit 122b has a sampling rate conversion unit 1001, an amplitude adjustment unit 1031, a rearrangement unit 1002, and a connection unit 1003.

[0337] In Fig. 29, an amplitude adjustment unit 1031 is arranged after the sampling rate conversion unit 1001. This amplitude adjustment unit 1031 corrects the amplitude value so that the amplitude value of the aerodynamic sound data after sampling rate conversion that is output from the sampling rate conversion unit 1001 fluctuates. As a method of correction, for example, the amplitude may be varied over time, as in the state transition diagram of values ​​indicated by the smooth function in Fig. 28. Alternatively, a configuration may be used in which the amplitude value is corrected by using one of a plurality of amplitude fluctuation patterns prepared in advance and multiplying the aerodynamic sound data by the amplitude fluctuation pattern.

[0338] Furthermore, the amplitude adjustment unit 1031 may be located after the rearrangement unit 1002 or after the connection unit 1003 .

[0339] (Embodiment 2) The following describes embodiment 2. The following mainly describes the differences from embodiment 1 and the modifications thereof, and omits or simplifies the description of commonalities.

[0340] [Configuration] First, a description will be given of the configuration of information processing device 600 according to this embodiment. Fig. 30 is a block diagram showing the functional configuration of information processing device 600 according to this embodiment.

[0341] The information processing device 600 includes a cyclic address section 610 , a frequency shift section 620 , a storage section 630 , a section designation section 640 , a cross-fade section 650 , and a read control section 660 .

[0342] When the duration of aerodynamic sound data is short, if this aerodynamic sound data is used repeatedly, problems arise such as noise occurring at the joints between the aerodynamic sound data. The information processing device 600 according to this embodiment is used to solve at least one of these problems.

[0343] 31A and 31B are diagrams illustrating how sound data is read out according to the conventional technology and how sound data is read out according to this embodiment, with (a) of Fig. 31 being a diagram illustrating how sound data is read out according to the conventional technology and (b) of Fig. 31 being a diagram illustrating how sound data is read out according to this embodiment.

[0344] The reading of sound data (aerodynamic sound data) according to the prior art will now be explained. In the prior art, a storage unit is provided in which aerodynamic sound data is stored, and a cyclic address unit cyclically cycles from the start address at which the aerodynamic sound data is stored in the storage unit to the end address at which the aerodynamic sound data is stored. The cyclic address unit reads the aerodynamic sound data from the storage unit and outputs it.

[0345] Next, the reading of sound data (aerodynamic sound data) according to this embodiment will be described.

[0346] Here, the aerodynamic sound data (for example, the unprocessed aerodynamic sound data D1 shown in FIG. 15) is made up of multiple sample points, and more specifically, it is made up of N sample points as shown in FIG. 31(b). Here, the first M sample points of the aerodynamic sound data and the last M sample points are cross-faded in advance to create M cross-faded sample points. In addition, the first M sample points and the last M sample points of the aerodynamic sound data are removed to create (N-2M) samples in the middle portion.

[0347] The storage unit 630 according to this embodiment stores aerodynamic sound data consisting of (N-M) samples, which are a combination of M cross-faded sample points and (N-2M) samples in the middle portion. A series of (N-M) addresses corresponding to the aerodynamic sound data consisting of (N-M) samples are set in this storage unit 630.

[0348] In this embodiment, the cyclic address section 610 cyclically cycles through the aerodynamic sound data made up of (N-M) samples stored in the storage section 630, from the start address to the end address, reads out the aerodynamic sound data, and outputs it to the frequency shift section 620. The frequency shift section 620 acquires the output aerodynamic sound data, shifts its frequency, and outputs it to an output channel of, for example, the headphones 200 according to the first embodiment.

[0349] In the information processing device 600 according to this embodiment, the first M sample points and the last M sample points are cross-faded, which makes it less likely that problems such as noise will occur at the joins between pieces of aerodynamic sound data.

[0350] Furthermore, the information processing device 600 according to this embodiment may perform the following processing: Fig. 32 is a diagram for explaining the processing performed by the information processing device 600 according to this embodiment.

[0351] (a) of Figure 32 is a diagram showing the configuration of the storage unit 630 according to this embodiment. Here, the storage unit 630 stores aerodynamic sound data (for example, the unprocessed aerodynamic sound data D1 shown in Figure 15), and is also provided with a first pointer Pt1 and a second pointer Pt2. The first pointer Pt1 indicates the read position from which the stored aerodynamic sound data is read. The second pointer Pt2 is a pointer that moves in conjunction with the first pointer Pt1, and indicates the read position from which aerodynamic sound data is read from the storage unit 630.

[0352] The section designation unit 640 designates a first section A1 and a second section A2. The second section A2 is a subsequent section adjacent to the first section A1. The second pointer Pt2 moves through a subsequent section A3 adjacent to the second section A2.

[0353] The first interval A1 and the second interval A2 may be arbitrarily set by the user of the information processing device 600. That is, a reception unit included in the information processing device 600 may receive an operation from the user specifying the first interval A1 and the second interval A2, and the interval designation unit 640 may determine the intervals specified by the received operation as the first interval A1 and the second interval A2.

[0354] The cross-fade unit 650 performs fade-in processing on the aerodynamic sound data read from the read position indicated by the first pointer Pt1, and outputs the faded-in aerodynamic sound data. The cross-fade unit 650 performs fade-out processing on the aerodynamic sound data read from the read position indicated by the second pointer Pt2, and outputs the faded-out aerodynamic sound data.

[0355] The read control unit 660 causes the cross-fade unit 650 to output aerodynamic sound data that has been faded in while the read position indicated by the first pointer Pt1 is included in the first section A1 and aerodynamic sound data is being read from the first section A1. The read control unit 660 outputs the aerodynamic sound data read from the second section A2 by the cyclic address unit 610 while the read position indicated by the first pointer Pt1 is not included in the first section A1 and aerodynamic sound data is not being read from the first section A1.

[0356] Then, the fade-in processed aerodynamic sound data output by the cross-fade unit 650, or the aerodynamic sound data read from the second section A2 by the cyclic address unit 610, is output to the frequency shift unit 620. The frequency shift unit 620 acquires the fade-in processed aerodynamic sound data that has been output, or the aerodynamic sound data that has been read from the second section A2, shifts the frequency of the data, and outputs it to an output channel of, for example, the headphones 200 according to the first embodiment.

[0357] Next, the processes shown in (b) and (c) of FIG. 32 will be described.

[0358] 32(b) is a diagram showing an example in which the first pointer Pt1 according to this embodiment moves cyclically between the first section A1 and the second section A2. In this example, the first pointer Pt1 moves cyclically between the first section A1 and the second section A2. While the read position indicated by the first pointer Pt1 is included in the first section A1, aerodynamic sound data is read from the read position indicated by the first pointer Pt1, and aerodynamic sound data is also read from the read position indicated by the second pointer Pt2 linked to the first pointer Pt1. The crossfade unit 650 performs crossfade processing on the two sets of aerodynamic sound data that have been read. Note that while the read position indicated by the first pointer Pt1 is included in the first section A1, the read position indicated by the second pointer Pt2 linked to the first pointer Pt1 may be included in the section A3 linked to the first section A1, and aerodynamic sound data may also be read from section A3.

[0359] (c) of Figure 32 is a diagram showing an example in which the second pointer Pt2 according to this embodiment circulates between the second section A2 and the section A3. In this example, the second pointer Pt2 circulates between the second section A2 and the section A3. While the read position indicated by the second pointer Pt2 is included in section A3, aerodynamic sound data is read from the read position indicated by the second pointer Pt2, and aerodynamic sound data is also read from the read position indicated by the first pointer Pt1. The crossfade unit 650 performs crossfade processing on the two sets of aerodynamic sound data that have been read. Note that while the read position indicated by the second pointer Pt2 is included in the second section A2, it is preferable that the read position indicated by the first pointer Pt1 be included in the first section A1 in conjunction with the second pointer Pt2, and aerodynamic sound data is also read from the first section A1.

[0360] Furthermore, the information processing device 600 according to this embodiment may perform the following process: Fig. 33 is a diagram for explaining another process performed by the information processing device 600 according to this embodiment.

[0361] In this other process, the interval designation unit 640 randomly updates the first interval A1 and the second interval A2. The interval designation unit 640 sequentially updates the position of the end point of the second interval A2 and the positions of the start and end points of the next first interval A1.

[0362] In other processing shown in Figure 33, the state in which aerodynamic sound data is read out transitions in the order of Figure 33(a), Figure 33(b), Figure 33(c), Figure 33(d), Figure 33(e), Figure 33(f), and Figure 33(g).

[0363] (a), (d), and (g) of Figure 33 show state 1 in which aerodynamic sound data is read out, (b) and (e) of Figure 33 show state 2 in which aerodynamic sound data is read out, and (c) and (f) of Figure 33 show state 3 in which aerodynamic sound data is read out.

[0364] In FIG. 33, state 1, state 2, and state 3 are repeated in this order.

[0365] In state 1 shown in Figure 33(a), aerodynamic sound data is read from the second section A2. At this time, the end point of the second section A2 has not been determined.

[0366] In state 2 shown in Figure 33(b), aerodynamic sound data is read from the second section A2. Then, the section designation unit 640 randomly designates the end point of the second section A2 and the next first section A1 at a predetermined timing. Note that section A3, which is linked to the next first section A1, is a subsequent section adjacent to the second section A2, so there is no need for the section designation unit 640 to designate it; it is determined automatically.

[0367] The predetermined timing may be arbitrarily set by the user of the information processing device 600. That is, a reception unit included in the information processing device 600 may receive an operation from the user instructing the predetermined timing, and the interval designation unit 640 may determine the timing instructed by the received operation as the predetermined timing.

[0368] 33(c), reading of the aerodynamic sound data from the second section A2 has finished. The cross-fade section 650 then performs cross-fade processing on the aerodynamic sound data read from the next first section A1, and the aerodynamic sound data read from the section A3 that is linked to the next first section A1.

[0369] In state 1 shown in (d) of Figure 33, aerodynamic sound data is read from the next second section A2. Note that this next second section A2 is the section that follows and is adjacent to the next first section A1 shown in (c) of Figure 33, so the section designation unit 640 does not need to designate the start point of this next second section A2; it is determined automatically. In other words, when the cross-fade processing described using (c) of Figure 33 is completed, aerodynamic sound data is read from this next second section A2. Note that at this time, just like state 1 shown in (a) of Figure 33, the end point of the second section A2 has not been determined.

[0370] In state 2 shown in (e) of Figure 33, aerodynamic sound data is read from second section A2 (equivalent to the next second section A2 shown in (d) of Figure 33). Then, at a predetermined timing, the section designation unit 640 randomly designates the end point of this second section A2 and the next first section A1. Note that section A3, which is linked to the next first section A1, is a subsequent section adjacent to second section A2, so there is no need for the section designation unit 640 to designate it, and it is determined automatically.

[0371] In state 3 shown in (f) of Figure 33, reading of the aerodynamic sound data from the second section A2 (corresponding to the next second section A2 shown in (e) of Figure 33) has finished. The crossfade section 650 then performs crossfade processing on the aerodynamic sound data read from the next first section A1, and the aerodynamic sound data read from the section A3 that is linked to the next first section A1.

[0372] In state 1 shown in (g) of Figure 33, aerodynamic sound data is read from the next second section A2. Note that this next second section A2 is the section that follows and is adjacent to the next first section A1 shown in (f) of Figure 33, so the section designation unit 640 does not need to designate the start point of this next second section A2; it is determined automatically. In other words, when the cross-fade processing described using (c) of Figure 33 is completed, aerodynamic sound data is read from this next second section A2. Note that at this time, just like state 1 shown in (a) of Figure 33, the end point of the second section A2 has not been determined.

[0373] 33, State 1, State 2, and State 3 are repeated in this order, and by randomly specifying the end point of the second section A2 and the next first section A1 in State 2, the listener L is prevented from repeatedly hearing the same aerodynamic sound. Therefore, the unnatural "rhythm" that occurs when the same aerodynamic sound is repeated is not generated.

[0374] Next, the pipeline processing will be described.

[0375] Some or all of the processing performed by the above-described acoustic signal processing device 100 may be performed as part of pipeline processing such as that described in Patent Document 2, for example. Fig. 34 is a functional block diagram and a diagram showing an example of steps for explaining a case where the rendering units A0203 and A0213 in Figs. 6 and 7 perform pipeline processing. In the explanation using Fig. 34, a rendering unit 900, which is an example of the rendering units A0203 and A0213 in Figs. 6 and 7, will be used.

[0376] Pipeline processing refers to dividing the process for adding sound effects into multiple processes and executing each process one by one in sequence. Each of the divided processes performs, for example, signal processing on an audio signal or generation of parameters used in signal processing.

[0377] The rendering unit 900 in this embodiment includes, as pipeline processing, processes for applying, for example, reverberation effects, early reflection processing, distance attenuation effects, and binaural processing. However, the above processes are merely examples, and other processes may be included, or some processes may not be included. For example, the rendering unit 900 may include diffraction processing or occlusion processing as pipeline processing, or may omit reverberation processing if it is not necessary. Furthermore, each process may be represented as a stage, and audio signals such as reflected sound generated as a result of each process may be represented as rendering items. The order of each stage in pipeline processing and the stages included in the pipeline processing are not limited to the example shown in FIG. 34 .

[0378] It should be noted that the rendering unit 900 does not necessarily have to include all of the stages shown in FIG. 34, and some stages may be omitted, or other stages may exist in addition to the rendering unit 900.

[0379] As an example of pipeline processing, we will explain the processes performed in each of the following: reverberation processing, early reflection processing, distance attenuation processing, selection processing, generation processing, and binaural processing. Each process analyzes metadata included in the input signal and calculates the parameters necessary to generate reflected sounds.

[0380] 34, the rendering unit 900 includes a reverberation processing unit 901, an early reflection processing unit 902, a distance attenuation processing unit 903, a selection unit 904, a calculation unit 906, a generation unit 907, and a binaural processing unit 905. Here, an example will be described in which the reverberation processing unit 901 performs a reverberation processing step, the early reflection processing unit 902 performs an early reflection processing step, the distance attenuation processing unit 903 performs a distance attenuation processing step, the selection unit 904 performs a selection processing step, and the binaural processing unit 905 performs a binaural processing step.

[0381] In the reverberation processing step, the reverberation processor 901 generates an audio signal indicating reverberant sound or parameters required to generate an audio signal. The reverberant sound is a sound that includes reverberant sound that reaches the listener as reverberation after the direct sound. As an example, the reverberant sound is a reverberant sound that reaches the listener after having been reflected more times (e.g., several tens of times) than the initial reflection sound, at a relatively later stage (e.g., about a hundred and several tens of milliseconds after the direct sound arrives) after the initial reflection sound described below reaches the listener. The reverberation processor 901 refers to the audio signal and spatial information included in the input signal and performs calculations using a predetermined function prepared in advance for generating reverberant sound.

[0382] The reverberation processor 901 may generate reverberation by applying a known reverberation generation method to the sound signal. One example of a known reverberation generation method is the Schroeder method, but the method is not limited to this. When applying the known reverberation generation process, the reverberation processor 901 uses the shape and acoustic characteristics of the sound reproduction space indicated by the spatial information. This allows the reverberation processor 901 to calculate parameters for generating an audio signal indicating reverberation.

[0383] In the early reflection processing step, the early reflection processing unit 902 calculates parameters for generating early reflection sounds based on spatial information. Early reflection sounds are reflected sounds that reach the listener after one or more reflections at a relatively early stage (e.g., approximately several tens of milliseconds after the direct sound arrives) after the direct sound arrives from the sound source object to the listener. The early reflection processing unit 902, for example, refers to the sound signal and metadata, and calculates the path (path length) of the reflected sound that travels from the sound source object to the listener after reflecting off the object, using the shape and size of the three-dimensional sound field (space), the positions of objects such as structures, and the reflectance of the objects. The early reflection processing unit 902 may also calculate the path (path length) of the direct sound. Information indicating the path may be used as a parameter for generating early reflection sounds, and may also be used as a parameter for the selection process of the reflected sound by the selection unit 904.

[0384] In the distance attenuation processing step, a distance attenuation processing unit 903 calculates the volume of sound reaching the listener based on the difference between the path length of the direct sound and the path length of the reflected sound calculated by the early reflection processing unit 902. The volume of sound reaching the listener attenuates in proportion to the distance to the listener (inversely proportional to the distance) relative to the volume of the sound source, so the volume of the direct sound can be obtained by dividing the volume of the sound source by the length of the path of the direct sound, and the volume of the reflected sound can be calculated by dividing the volume of the sound source by the length of the path of the reflected sound.

[0385] In the selection process step, the selection unit 904 selects a sound to be generated. The selection process may be performed based on parameters calculated in the previous steps.

[0386] When the selection process is performed as part of the pipeline process, sounds not selected in the selection process may not be subjected to subsequent processing in the pipeline process. By not performing subsequent processing on the unselected sounds, it is possible to reduce the computational load on the acoustic signal processing device 100 compared to when it is decided not to perform binaural processing on the unselected sounds.

[0387] Furthermore, when the selection processing described in this embodiment is executed as part of pipeline processing, if the order of the selection processing is set to be an earlier order among the orders of multiple processes in the pipeline processing, more of the processing after the selection processing can be omitted, thereby making it possible to reduce the amount of calculations even more. For example, if the selection processing is executed in an order earlier than the processing of the calculation unit 906 and the generation unit 907, it is possible to omit processing for aerodynamic sounds related to objects that have been determined not to be selected, making it possible to further reduce the amount of calculations in the acoustic signal processing device 100.

[0388] Additionally, parameters calculated during a part of the pipeline process that generates the rendering items may be used by the selection unit 904 or the calculation unit 906 .

[0389] In the binaural processing step, the binaural processing unit 905 performs signal processing on the audio signal of the direct sound so that the sound is perceived as reaching the listener from the direction of the sound source object. Furthermore, the binaural processing unit 905 performs signal processing so that the reflected sound is perceived as reaching the listener from an obstacle object involved in the reflection. Based on the coordinates and orientation of the listener in the sound space (i.e., the position and orientation of the listening point), a process of applying a head-related impulse response (HRIR) database (data base) is performed so that the sound reaches the listener from the position of the sound source object or the position of the obstacle object. Note that the position and direction of the listening point may change in accordance with, for example, the movement of the listener's head. Information indicating the position of the listener may also be acquired from a sensor.

[0390] The programs used for pipeline processing and binaural processing, spatial information required for acoustic processing, the HRIR DB, and other parameters such as threshold data are acquired from memory provided in the acoustic signal processing device 100 or from an external source. HRIR (Head-Related Impulse Responses) are response characteristics when a single impulse is generated. In other words, HRIR is a response characteristic converted from a frequency domain representation to a time domain representation by Fourier transforming a head-related transfer function (HRTF), which represents, as a transfer function, changes in sound caused by surrounding objects including the auricle, the human head, and shoulders. The HRIR DB is a database containing such information.

[0391] As an example of pipeline processing, the rendering unit 900 may include processing units (not shown), such as a diffraction processing unit or an occlusion processing unit.

[0392] The diffraction processing unit executes processing to generate an audio signal representing a sound including diffracted sound caused by an obstacle between the listener and the sound source object in a three-dimensional sound field (space). When there is an obstacle between the sound source object and the listener, the diffracted sound is sound that travels from the sound source object to the listener by going around the obstacle.

[0393] The diffraction processing unit, for example, refers to the sound signal and metadata, and uses the position of the sound source object in the three-dimensional sound field (space), the position of the listener, and the position, shape, and size of obstacles to calculate a path from the sound source object to the listener, bypassing obstacles, and generates diffracted sound based on that path.

[0394] The occlusion processing unit generates an audio signal that can be heard when a sound source object is located behind an obstacle object, based on the spatial information acquired in any of the steps and information such as the material of the obstacle object.

[0395] In the above-described first and second embodiments, the position information assigned to the sound source object is defined as a "point" in the virtual space, and the details of the invention have been described assuming that the sound source is a so-called "point sound source." On the other hand, as a method for defining a sound source in a virtual space, a spatially extended sound source, rather than a point sound source, may be defined as an object having length, size, or shape. In such cases, the distance between the listener and the sound source or the direction of sound arrival is not determined, so the reflected sound resulting from this may be limited to being "selected" by the selection unit 904 without analysis, or regardless of the analysis results. This is because it is possible to avoid deterioration in sound quality that may occur when the reflected sound is not selected. Alternatively, a representative point, such as the center of gravity of the object, may be defined, and the processing of the present disclosure may be applied assuming that the sound is generated from that representative point. In this case, the threshold may be adjusted according to information about the spatial extension of the sound source before the processing of the present disclosure is applied.

[0396] Next, an example of the structure of a bitstream will be described.

[0397] The bitstream includes, for example, an audio signal and metadata. The audio signal is sound data that represents sound, indicating information such as the frequency and intensity of the sound. The spatial information included in the metadata is information about the space in which a listener who listens to sound based on the audio signal is located. Specifically, the spatial information is information about a predetermined position (localization position) when a sound image of the sound is localized at a predetermined position in a sound space (e.g., in a three-dimensional sound field), that is, when the listener perceives the sound as arriving from a predetermined direction. The spatial information includes, for example, sound source object information and position information indicating the position of the listener.

[0398] The sound source object information is information about an object that generates a sound based on an audio signal, that is, that indicates an object that plays the audio signal, and is information about a virtual object (sound source object) that is placed in a sound space, which is a virtual space that corresponds to the real space in which the object is placed. The sound source object information includes, for example, information indicating the position of the sound source object placed in the sound space, information about the orientation of the sound source object, information about the directionality of the sound emitted by the sound source object, information indicating whether the sound source object belongs to a living thing, and information indicating whether the sound source object is a moving object. For example, the audio signal corresponds to one or more sound source objects indicated by the sound source object information.

[0399] As an example of the data structure of a bitstream, the bitstream is made up of, for example, metadata (control information) and an audio signal.

[0400] The audio signal and metadata may be stored in a single bitstream or in separate bitstreams, and similarly, the audio signal and metadata may be stored in a single file or in separate files.

[0401] A bitstream may exist for each sound source or for each playback time. When a bitstream exists for each playback time, multiple bitstreams may be processed in parallel at the same time.

[0402] The metadata may be attached to each bitstream, or may be attached collectively as information for controlling a plurality of bitstreams, or may be attached to each playback time.

[0403] When the audio signal and metadata are stored separately in multiple bitstreams or multiple files, the audio signal and metadata may be included in information indicating other bitstreams or files related to one or some of the bitstreams or files, or the audio signal and metadata may be included in information indicating other bitstreams or files related to each of all of the bitstreams or files. Here, the related bitstreams or files are, for example, bitstreams or files that may be used simultaneously during audio processing. Furthermore, the related bitstreams or files may include a bitstream or file that collectively describes information indicating other related bitstreams or files. Here, the information indicating other related bitstreams or files may be, for example, an identifier indicating the other bitstream, a file name indicating the other file, a URL (Uniform Resource Locator), or a URI (Uniform Resource Identifier). In this case, the acquisition unit 110 identifies or acquires the bitstream or file based on the information indicating the other related bitstreams or files. Furthermore, a bitstream may contain information indicating other related bitstreams, and may also contain information indicating a bitstream or file related to another bitstream or file, where the file containing information indicating the related bitstream or file may be, for example, a control file such as a manifest file used in content distribution.

[0404] Note that all or some of the metadata may be acquired from sources other than the bitstream of the audio signal. For example, either the metadata controlling the audio or the metadata controlling the video may be acquired from sources other than the bitstream, or both may be acquired from sources other than the bitstream. Furthermore, when the metadata controlling the video is included in the bitstream acquired by the audio signal reproduction system, the audio signal reproduction system may have a function of outputting the metadata that can be used to control the video to a display device that displays images or a 3D video reproduction device that reproduces 3D video.

[0405] Furthermore, an example of information contained in the metadata will be described.

[0406] The metadata may be information used to describe a scene represented in sound space. Here, a scene is a term that refers to the collection of all elements representing three-dimensional images and acoustic events in sound space that are modeled in an audio signal reproduction system using metadata. In other words, the metadata here may include not only information that controls sound processing, but also information that controls video processing. Of course, the metadata may include information that controls only one of sound processing and video processing, or information used to control both.

[0407] The audio signal reproduction system generates virtual sound effects by performing sound processing on the audio signal using metadata included in the bitstream and additionally acquired interactive listener position information. Here, a case where early reflection processing, obstacle processing, diffraction processing, blocking processing, and reverberation processing are performed among the sound effects will be described, but other sound processing may also be performed using the metadata. For example, the audio signal reproduction system may add sound effects such as distance attenuation effect, localization, and Doppler effect. Furthermore, information for switching all or some of the sound effects on or off, and priority information may also be added as metadata.

[0408] As an example, the encoded metadata includes information about a sound space including a sound source object and an obstacle object, and information about the localization position when the sound image of the sound is localized at a predetermined position in the sound space (i.e., perceived as sound arriving from a predetermined direction). Here, an obstacle object is an object that can affect the sound perceived by a listener by, for example, blocking or reflecting the sound emitted by the sound source object before it reaches the listener. Obstacle objects may include not only stationary objects but also animals such as people or moving objects such as machines. Furthermore, when multiple sound source objects exist in a sound space, other sound source objects can be obstacle objects for any sound source object. Non-sounding objects, such as building materials or inanimate objects, that do not emit sound, as well as sound source objects that emit sound, can be obstacle objects.

[0409] The metadata includes all or part of the information representing the shape of the sound space, shape and position information of obstacle objects present in the sound space, shape and position information of sound source objects present in the sound space, and the position and orientation of the listener in the sound space.

[0410] The sound space may be either a closed space or an open space. The metadata also includes information indicating the reflectance of structures that can reflect sound in the sound space, such as floors, walls, and ceilings, and the reflectance of obstacle objects that exist in the sound space. Here, the reflectance is the ratio of the energy of reflected sound to incident sound, and is set for each frequency band of sound. Of course, the reflectance may be set uniformly regardless of the frequency band of sound. When the sound space is an open space, parameters such as a uniform attenuation rate, diffracted sound, and early reflected sound may be used.

[0411] In the above description, reflectance was mentioned as a parameter related to an obstacle object or a sound source object included in the metadata, but information other than reflectance may also be included. For example, information other than reflectance may include information about the material of the object as metadata related to both the sound source object and the non-sound-producing object. Specifically, information other than reflectance may include parameters such as diffusion rate, transmittance, and sound absorption rate.

[0412] Information about a sound source object may include information such as volume, radiation characteristics (directivity), playback conditions, the number and type of sound sources emitted from a single object, and information specifying the sound source area of ​​the object. The playback conditions may, for example, determine whether the sound is a continuous sound or an event-triggered sound. The sound source area of ​​the object may be determined relative to the position of the listener and the position of the object, or may be determined based on the object. When the sound source area of ​​the object is determined relative to the position of the listener and the position of the object, the surface of the object the listener is looking at can be used as the reference, allowing the listener to perceive sound C as emanating from the right side of the object and sound E as emanating from the left side of the object as viewed from the listener. When the sound source area of ​​the object is determined based on the object, the sound emitted from which area of ​​the object can be fixed regardless of the direction the listener is looking. For example, when viewed from the front, the listener can perceive a high-pitched sound coming from the right side and a low-pitched sound coming from the left side. In this case, if the listener goes behind the object, the listener can perceive a low-pitched sound coming from the right side and a high-pitched sound coming from the left side as viewed from the back.

[0413] Spatial metadata can include time to early reflections, reverberation time, direct to diffuse ratio, etc. If the direct to diffuse ratio is zero, only direct sound will be perceived by the listener.

[0414] (Effects, etc.) The acoustic signal processing method according to the first embodiment includes an acquisition step of acquiring sound data indicating the waveform of a reference sound, a processing step of processing the sound data so as to change at least one of the frequency component, phase, and amplitude value of the waveform based on simulation information that simulates fluctuations in natural phenomena, and an output step of outputting the processed sound data.

[0415] As a result, the sound data is processed so as to change at least one of the frequency components, phase, and amplitude of the waveform based on simulation information that simulates fluctuations in natural phenomena containing fluctuations. As a result, fluctuations occur in at least one of the frequency components, phase, and amplitude in the processed sound data, and fluctuations also occur in at least one of the frequency components, phase, and amplitude in the sound represented by the processed sound data. Therefore, the listener L can hear a sound in which fluctuations occur in at least one of the frequency components, phase, and amplitude, and the listener L can experience a sense of realism without feeling any discomfort. In other words, an acoustic signal processing method that can give the listener L a sense of realism is realized.

[0416] In the above-described Operational Example 1 of the First Embodiment, an example in which wind W blows is used as the natural phenomenon. As described above, the simulated information is information that simulates fluctuations in a natural phenomenon that includes fluctuations, more specifically, information that expresses fluctuations due to fluctuations in the wind speed of the wind W, and in the Operational Example 1, the simulated information is information that is expressed by a smooth function.

[0417] In operation example 1, sound data (aerodynamic sound data) indicating the waveform of a reference sound is processed so that the frequency components of the waveform change based on simulation information that simulates fluctuations in natural phenomena that include fluctuations. As a result, fluctuations occur in the frequency components of the processed aerodynamic sound data, and fluctuations also occur in the frequency components of the aerodynamic sound indicated by the processed aerodynamic sound data. Therefore, the listener L can hear aerodynamic sound with fluctuations in such frequency components, and can obtain a sense of realism without feeling any discomfort. In other words, an acoustic signal processing method that can give the listener L a sense of realism is realized.

[0418] In the above-described operation example 1 of the first embodiment, the example of wind W blowing is used as the natural phenomenon, but the present invention is not limited to this, and natural phenomena such as flowing water in a river or animal behavior may also be used.

[0419] In the case of an example of a flowing river as a natural phenomenon, the listener L hears the murmuring sound of the flowing river. In this case, the simulated information is information that expresses fluctuations due to fluctuations in the flow speed or direction of the river water.

[0420] When the example of animal behavior is used as the natural phenomenon, the listener L will hear the animal's cry, etc. In this case, the simulated information is information that expresses fluctuations due to changes in the volume of the animal's cry, etc.

[0421] That is, even when a phenomenon such as flowing water in a river or animal behavior is used as the natural phenomenon, the simulated information is information that simulates the fluctuation of the natural phenomenon that includes fluctuations. Therefore, as shown in Operation Example 1, by using the simulated information, the listener L can hear a sound in which fluctuations occur in at least one of the frequency component, phase, and amplitude, and the listener L can obtain a sense of realism without feeling any discomfort. In other words, an acoustic signal processing method that can provide the listener L with a sense of realism is realized.

[0422] In the acoustic signal processing method according to the first embodiment, the reference sound is an aerodynamic sound generated by wind W, and in the processing step, the sound data is processed so as to change at least one of the frequency component, phase, and amplitude value of the waveform based on simulation information that simulates fluctuations in the wind speed of the wind W.

[0423] This allows the listener L to hear aerodynamic sound with fluctuations in at least one of the frequency components, phase, and amplitude, and the listener L can experience a sense of realism without feeling any discomfort. In other words, an acoustic signal processing method that can give the listener L a sense of realism is realized.

[0424] In the processing step of the acoustic signal processing method according to the first embodiment, a smooth function that simulates fluctuations in the wind speed of the wind W is determined as simulation information, and the sound data is processed so as to change at least one of the frequency components, phase, and amplitude value of the waveform based on the value indicated by the determined smooth function.

[0425] This allows the sound data to be processed according to the values ​​indicated by the smooth function.

[0426] In the acoustic signal processing method according to the first embodiment, the value indicated by the smooth function is information indicating the ratio between the wind speed of the aerodynamic sound, which is the reference sound, and the wind speed of the aerodynamic sound indicated by the sound data after processing in the processing step.

[0427] This allows the sound data to be processed based on the ratio between the wind speed of the aerodynamic sound, which is the reference sound, and the wind speed of the aerodynamic sound indicated by the processed sound data.

[0428] In the processing step of the acoustic signal processing method according to the first embodiment, a smooth function is determined such that a parameter specifying the smooth function varies irregularly.

[0429] This allows the listener L to hear aerodynamic sound with irregularly changing fluctuations in at least one of the frequency components, phase, and amplitude, and the listener L is less likely to feel uncomfortable and can experience a greater sense of realism. In other words, an acoustic signal processing method that can give the listener L a greater sense of realism is realized.

[0430] In the processing step of the acoustic signal processing method according to the first embodiment, sound data is processed so as to shift the frequency components of the waveform to frequencies proportional to the values ​​indicated by the determined smooth function.

[0431] This allows the listener L to hear a sound with fluctuations in the frequency components, and the listener L can get a sense of realism without feeling any discomfort. In other words, an acoustic signal processing method that can give the listener L a sense of realism is realized.

[0432] That is, as shown in Operation Example 1, sound data (aerodynamic sound data) indicating the waveform of the reference sound is processed so that the frequency components of the waveform change based on simulation information (smooth function) that simulates fluctuations in the wind speed of the wind W, which includes fluctuations. For this reason, fluctuations occur in the frequency components in the processed aerodynamic sound data, and fluctuations also occur in the frequency components of the aerodynamic sound indicated by the processed aerodynamic sound data. Therefore, the listener L can hear aerodynamic sound with fluctuations in such frequency components, and can obtain a sense of realism without feeling uncomfortable.

[0433] In the processing step of the acoustic signal processing method according to the first embodiment, sound data is processed so that the amplitude value of the waveform is changed in proportion to the α power of the value indicated by the determined smooth function.

[0434] This allows the listener L to hear a sound with fluctuations in amplitude, and the listener L can get a sense of realism without feeling any discomfort. In other words, an acoustic signal processing method that can give the listener L a sense of realism is realized.

[0435] That is, as shown in operation example 2, sound data (aerodynamic sound data) indicating the waveform of the reference sound is processed so that the amplitude value of the waveform changes in proportion to the α power of the value indicated by a smooth function, which is simulation information that simulates fluctuations in the wind speed of the wind W that includes fluctuations. For this reason, fluctuations occur in the amplitude values ​​in the processed aerodynamic sound data, and fluctuations also occur in the amplitude values ​​of the aerodynamic sound indicated by the processed aerodynamic sound data. Therefore, the listener L can hear aerodynamic sound with such fluctuations in the amplitude values, and can obtain a sense of realism without feeling any discomfort.

[0436] In the processing step of the acoustic signal processing method according to the first embodiment, acquired sound data is divided into processing frames F of a predetermined time, and the sound data is processed for each divided processing frame F.

[0437] This realizes an acoustic signal processing method with a reduced computational processing load.

[0438] In the processing step of the acoustic signal processing method according to the first embodiment, a smooth function is determined for each divided processing frame F so that the value of the smooth function becomes 1.0 at the first time and the last time of the processing frame F.

[0439] This prevents noise from occurring at the joint between the processed frame F and the next processed frame F.

[0440] In the acoustic signal processing method according to the first embodiment, a parameter specifying a smooth function is determined for each divided processing frame F in the processing step.

[0441] This realizes an acoustic signal processing method with a reduced computational processing load.

[0442] In the acoustic signal processing method according to the first embodiment, the parameter is the time from the first time to the last time.

[0443] This allows the parameter to be the time from the first time of the processing frame F to the last time of the processing frame F.

[0444] In the acoustic signal processing method according to the first embodiment, the parameter is a value related to the maximum value of a smooth function.

[0445] This allows the parameter to be a value related to the maximum value of a smooth function.

[0446] In the acoustic signal processing method according to the first embodiment, the parameter is a parameter that varies the position at which the smooth function reaches its maximum value.

[0447] This allows the parameter to vary the position where the smooth function reaches its maximum value.

[0448] In the acoustic signal processing method according to the first embodiment, the parameter is a parameter that varies the steepness of the variation of the smooth function.

[0449] This allows the parameter to be a parameter that changes the steepness of the change in the smooth function.

[0450] In the acoustic signal processing method according to the first embodiment, in the processing step, a first parameter and a second parameter that specify a smooth function are determined, and the acquired sound data is processed so as to change at least one of the frequency components, phase, and amplitude values ​​of the waveform based on the smooth function specified by the determined first parameter, and the acquired sound data is processed so as to change at least one of the frequency components, phase, and amplitude values ​​of the waveform based on the smooth function specified by the determined second parameter, and in the output step, the sound data processed based on the smooth function specified by the determined first parameter is output to a first output channel, and the sound data processed based on the smooth function specified by the determined second parameter is output to a second output channel.

[0451] This allows different sound data to be output for each output channel.

[0452] In the acoustic signal processing method according to the first embodiment, aerodynamic sound is sound generated when wind W collides with an object, and in the processing step, the characteristics of the wind speed of the wind W are simulated to determine parameters.

[0453] This allows the parameters to be determined by simulating fluctuations in the wind speed of the fluctuating wind W. Based on the smooth function specified by these parameters, the sound data can be processed to change at least one of the frequency components, phase, and amplitude of the waveform.

[0454] In the acoustic signal processing method according to the first embodiment, aerodynamic sound is sound generated when wind W collides with the ear of a listener L who hears the aerodynamic sound, and in the processing step, parameters are determined by simulating the characteristics of the wind direction of the wind W.

[0455] This allows the parameters to be determined by simulating fluctuations in the direction of the fluctuating wind W. Based on the smooth function specified by these parameters, the sound data can be processed to change at least one of the frequency components, phase, and amplitude of the waveform.

[0456] In the acoustic signal processing method according to the first embodiment, the maximum value of the smooth function does not exceed three.

[0457] This allows the maximum value of the smooth function to be 3 or less.

[0458] In the acoustic signal processing method according to the first embodiment, the minimum value of the smooth function does not fall below zero.

[0459] This allows the minimum value of the smooth function to be 0 or greater.

[0460] The acoustic signal processing method according to the first embodiment includes a receiving step for receiving instructions specifying Va, which is the wind speed of the wind W, and Vp, which is the instantaneous wind speed of the wind W, and a processing step for determining a smooth function such that the maximum value of the smooth function is Vp / Va.

[0461] This allows for a smooth function maximum and Vp / Va.

[0462] In the acoustic signal processing method according to the first embodiment, the average value of the predetermined time is 3 seconds.

[0463] This allows the average value of the predetermined time, which is the duration of the processing frame F, to be 3 seconds.

[0464] In the acoustic signal processing method according to the first embodiment, the object has a shape that resembles an ear.

[0465] This makes it possible to collect aerodynamic sounds using, for example, a dummy head microphone.

[0466] The computer program according to the first embodiment causes a computer to execute the above-described acoustic signal processing method.

[0467] This allows the computer to execute the above-described acoustic signal processing method in accordance with the computer program.

[0468] The acoustic signal processing device 100 according to the first embodiment includes an acquisition unit 110 that acquires sound data indicating the waveform of a reference sound, a processing unit 120 that processes the sound data so as to change at least one of the frequency component, phase, and amplitude value of the waveform based on simulation information that simulates fluctuations in natural phenomena, and an output unit 130 that outputs the processed sound data.

[0469] As a result, the sound data is processed so as to change at least one of the frequency component, phase, and amplitude of the waveform based on simulation information that simulates fluctuations in natural phenomena containing fluctuations. As a result, fluctuations occur in at least one of the frequency component, phase, and amplitude in the processed sound data, and fluctuations also occur in at least one of the frequency component, phase, and amplitude in the sound represented by the processed sound data. Therefore, the listener L can hear a sound in which fluctuations occur in at least one of the frequency component, phase, and amplitude, and the listener L can experience a sense of realism without feeling any discomfort. In other words, an acoustic signal processing device 100 that can give the listener L a sense of realism is realized.

[0470] (Other Embodiments) While the acoustic signal processing method and acoustic signal processing device according to aspects of the present disclosure have been described above based on embodiments and modifications, the present disclosure is not limited to these embodiments and modifications. For example, other embodiments realized by arbitrarily combining the components described in this specification or excluding some of the components may also be considered as embodiments of the present disclosure. Furthermore, modifications obtained by applying various modifications that would occur to a person skilled in the art to the above embodiments and modifications without departing from the spirit of the present disclosure, i.e., the meaning indicated by the wording of the claims, are also included in the present disclosure.

[0471] The following embodiments may also be included within the scope of one or more aspects of the present disclosure.

[0472] (1) Some of the components constituting the above-mentioned audio signal processing device may be a computer system comprising a microprocessor, ROM, RAM, hard disk unit, display unit, keyboard, mouse, etc. A computer program is stored in the RAM or hard disk unit. The microprocessor operates in accordance with the computer program to achieve its functions. Here, the computer program is composed of a combination of multiple instruction codes that indicate commands to a computer to achieve a predetermined function.

[0473] (2) Some of the components constituting the above-described audio signal processing device may be configured as a single system LSI (Large Scale Integration). A system LSI is an ultra-multifunctional LSI manufactured by integrating multiple components on a single chip, and specifically, is a computer system configured to include a microprocessor, ROM, RAM, etc. A computer program is stored in the RAM. The system LSI achieves its functions by the microprocessor operating in accordance with the computer program.

[0474] (3) Some of the components constituting the above-mentioned audio signal processing device may be configured as an IC card or a standalone module that can be attached to each device. The IC card or the module may be a computer system configured with a microprocessor, ROM, RAM, etc. The IC card or the module may include the above-mentioned ultra-multifunctional LSI. The IC card or the module achieves its functions when the microprocessor operates according to a computer program. The IC card or the module may be tamper-resistant.

[0475] (4) Furthermore, some of the components constituting the above-described acoustic signal processing device may be the computer program or the digital signal recorded on a computer-readable recording medium, such as a flexible disk, hard disk, CD-ROM, MO, DVD, DVD-ROM, DVD-RAM, BD (Blu-ray (registered trademark) Disc), semiconductor memory, etc. Alternatively, the components may be digital signals recorded on such recording media.

[0476] Furthermore, some of the components constituting the above-mentioned acoustic signal processing device may transmit the computer program or the digital signal via a telecommunications line, a wireless or wired communication line, a network such as the Internet, data broadcasting, etc.

[0477] (5) The present disclosure may be embodied as the methods described above, a computer program that implements these methods on a computer, or a digital signal that includes the computer program.

[0478] (6) The present disclosure may also be a computer system having a microprocessor and a memory, the memory storing the computer program, and the microprocessor operating in accordance with the computer program.

[0479] (7) The program or the digital signal may also be implemented by another independent computer system by recording it on the recording medium and transferring it, or by transferring the program or the digital signal via the network or the like.

[0480] The present disclosure is applicable to an acoustic signal processing method and an acoustic signal processing device, and is particularly applicable to an acoustic system and the like.

[0481] 100, 100a Acoustic signal processing device 110 Acquisition unit 120, 120a Processing unit 121 First processing unit 122, 122b Second processing unit 130 Output unit 140 Memory unit 150 Reception unit 200 Headphones 201 Head sensor unit 202 Output unit 300 Display unit 500 Server device 600 Information processing device 610 Cyclic address unit 620 Frequency shift unit 630 Memory unit 640 Section designation unit 650 Crossfade unit 660 Read control unit 900 Rendering unit 901 Reverberation processing unit 902 Early reflection processing unit 903 Distance attenuation processing unit 904 Selection unit 905 Binaural processing unit 906 Calculation unit 907 Generation unit 1001 Sampling rate conversion unit 1002 Rearrangement unit 1003 Connection unit 1021 Upsampling unit 1022 Low-pass filter unit 1023 Downsampling unit 1024 XY setting unit 1031 Amplitude adjustment unit A1 First section A2 Second section A3 Section A0000 Stereophonic sound reproduction system A0001 Acoustic signal processing device A0002 Audio presentation device A0100 Encoding device A0101 Input data A0102 Encoder A0103 Encoded data A0104 Memory A0110 Decoding device A0111 Audio signal A0112 Decoder A0113 Input data A0114 Memory A0120 Encoding device A0121 Transmitting unit A0122 Transmitted signal A0130 Decoding device A0131 Receiving unit A0132 Received signal A0200 Decoder A0201 Spatial information management unit A0202 Audio data decoder A0203 Rendering unit A0210 Decoder A0211 Spatial information management unit A0213 Rendering unit D1 Aerodynamic sound data before processing D11 Aerodynamic sound data after processing FN Electric fan F, F1, F2, F3, F4, F5, F6, Fn, Fn+1 Processing frame L Listener Pt1 First pointer Pt2 Second pointer

Claims

1. An acquisition step of acquiring sound data indicating a waveform of a reference sound; a processing step of processing the sound data so as to change at least one of a frequency component, a phase, and an amplitude value of the waveform based on simulation information in which a variation in a natural phenomenon is simulated; an output step of outputting the processed sound data; Including, Acoustic signal processing method.

2. The reference sound is an aerodynamic sound generated by wind, In the processing step, the sound data is processed so as to change at least one of a frequency component, a phase, and an amplitude value of the waveform based on the simulation information in which a fluctuation in the wind speed of the wind is simulated.

2. The method of claim 1.

3. In the processing step, determining a smooth function simulating fluctuations in the wind speed as the simulation information; processing the sound data so as to change at least one of a frequency component, a phase, and an amplitude value of the waveform based on a value indicated by the determined smooth function; 3. The method of claim 2.

4. The value indicated by the smooth function is information indicating the ratio between the wind speed of the aerodynamic sound that is the reference sound and the wind speed of the aerodynamic sound indicated by the sound data after processing in the processing step. The method of processing an acoustic signal according to claim 3.

5. In the processing step, the smooth function is determined such that parameters specifying the smooth function vary randomly. The method of processing an acoustic signal according to claim 3.

6. In the processing step, the sound data is processed so as to shift the frequency components of the waveform to a frequency proportional to the value indicated by the determined smooth function. The method of processing an acoustic signal according to claim 3.

7. In the processing step, the sound data is processed so as to change an amplitude value of the waveform in proportion to the α power of a value indicated by the determined smooth function. The method of processing an acoustic signal according to claim 3.

8. In the processing step, the acquired sound data is divided into processing frames of a predetermined time, and the sound data is processed for each of the divided processing frames.

5. The method of claim 4.

9. In the processing step, the smooth function is determined for each of the divided processing frames such that a value of the smooth function becomes 1.0 at a first time and a last time of the processing frame.

9. The method of claim 8.

10. An acquisition unit that acquires sound data indicating a waveform of a reference sound; a processing unit that processes the sound data so as to change at least one of a frequency component, a phase, and an amplitude value of the waveform based on simulation information that simulates a variation in a natural phenomenon; an output unit that outputs the processed sound data; Equipped with Acoustic signal processing device.

11. The parameters that specify the smooth function are determined according to information about wind speeds, including instantaneous wind speeds. The method of processing an acoustic signal according to claim 3.

12. The processing step includes determining the smooth function based on two or more randomly varying parameters; One or more of the two or more parameters are parameters indicating information about wind speed. The method of processing an acoustic signal according to claim 3.

13. In the processing step, the acquired sound data is divided into processing frames of a predetermined time, and the sound data is processed for each of the divided processing frames; One or more parameters among the two or more parameters are parameters indicating information about the processing frame length. The method of processing an acoustic signal according to claim 12.

14. In the processing step, Dividing the acquired sound data into processing frames of a predetermined time, and processing the sound data for each of the divided processing frames; determining parameters specifying the smooth function for each of the divided processing frames; 4. The method of claim 3.

15. The parameter is a time from a first time to a last time of the processing frame.

15. The method of processing an acoustic signal according to claim 14.

16. The parameter is a value relating to a maximum value of the smooth function.

15. The method of processing an acoustic signal according to claim 14.

17. A computer program for causing a computer to execute the acoustic signal processing method according to any one of claims 1 to 9 and 11 to 16.