Acoustic signal processing method, computer program, and acoustic signal processing device

By processing wind sound data in the virtual space, fluctuates its frequency components, phases and amplitude values, the problem of unrealistic reproduction of stroke sound in virtual space in the prior art is solved, and a higher sense of presence is achieved.

CN120019674APending Publication Date: 2025-05-16PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202380071660.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-06
Filing Date
2023-10-03
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art is difficult to reproduce the sound of wind in real space in virtual space, resulting in the listener being unable to feel the sense of presence.

Method used

By obtaining sound data representing the waveform of the reference sound, the sound data is processed based on simulation information that simulates changes in natural phenomena, so that at least one of the frequency component, phase and amplitude value of the waveform is changed, so that the fluctuating wind sound is reproduced in the virtual space.

Benefits of technology

The wind sound with a sense of presence is realized in the virtual space, reducing the possibility of the listener feeling of incongruity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120019674A_ABST
    Figure CN120019674A_ABST
Patent Text Reader

Abstract

The audio signal processing method includes: an acquisition step of acquiring audio data indicating a waveform of a reference tone; a processing step for processing the sound data so as to change at least one of the frequency component, the phase, and the amplitude value of the waveform, on the basis of simulation information in which fluctuations in the natural phenomenon are simulated; and an output step of outputting the processed sound data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an audio signal processing method, etc. Background Art

[0002] Furthermore, Patent Document 1 discloses a technique for outputting images and sounds to create a realistic virtual space. Patent Document 1 discloses a technique for changing the sound of the wind in accordance with changes in the strength of the wind in the virtual space.

[0003] Prior art literature

[0004] Patent Literature

[0005] Patent Document 1: Japanese Patent Application Laid-Open No. 1998-2151162

[0006] Patent Document 2: International Publication No. 2021 / 180938

[0007] Non-patent literature

[0008] Non-Patent Literature 1: Yoshinori Dobashi et al., Real-time rendering of aerodynamic sound using sound textures based on computational fluid dynamics, ACM Transactions on Graphics, Vol. 22, No. 3, p732-740 Summary of the Invention

[0009] Problems to be solved by the invention

[0010] However, the technology disclosed in Patent Document 1 may have difficulty in providing a sense of presence to the listener.

[0011] Therefore, an object of the present disclosure is to provide an audio signal processing method and the like that can provide a sense of presence to a listener.

[0012] Means used to solve problems

[0013] An acoustic signal processing method according to a technical solution of the present disclosure includes: an acquisition step of acquiring sound data representing a waveform of a reference sound; a processing step of processing the sound data based on simulation information simulating changes in natural phenomena so as to change at least one of the frequency component, phase, and amplitude value of the waveform; and an output step of outputting the processed sound data.

[0014] Furthermore, a computer program according to one technical aspect of the present disclosure causes a computer to execute the above-mentioned sound signal processing method.

[0015] In addition, an audio signal processing device according to a technical solution of the present disclosure includes: an acquisition unit that acquires sound data representing a waveform of a reference sound; a processing unit that processes the sound data based on simulation information that simulates changes in natural phenomena so as to change at least one of the frequency component, phase, and amplitude value of the waveform; and an output unit that outputs the processed sound data.

[0016] In addition, these inclusive or specific technical solutions can also be implemented by systems, devices, methods, integrated circuits, computer programs or non-temporary recording media such as computer-readable CD-ROMs, or by any combination of systems, devices, methods, integrated circuits, computer programs and recording media.

[0017] Effects of the Invention

[0018] According to an audio signal processing method according to a technical solution of the present disclosure, a sense of presence can be provided to listeners. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 This is a diagram showing an immersive audio reproduction system as an example of a system to which the audio processing or decoding processing of the present disclosure can be applied.

[0020] Figure 2 This is a functional block diagram showing the configuration of an encoding device as an example of the encoding device of the present disclosure.

[0021] Figure 3 This is a functional block diagram showing the configuration of a decoding device as an example of the decoding device of the present disclosure.

[0022] Figure 4 This is a functional block diagram showing the configuration of an encoding device as another example of the encoding device of the present disclosure.

[0023] Figure 5 This is a functional block diagram showing the configuration of a decoding device as another example of the decoding device of the present disclosure.

[0024] Figure 6 It means as Figure 3 or Figure 5 A functional block diagram of the structure of a decoder, which is an example of a decoder in FIG.

[0025] Figure 7 It means as Figure 3 or Figure 5 A functional block diagram of the structure of another example of a decoder in .

[0026] Figure 8This is a diagram showing an example of the physical configuration of an audio signal processing device.

[0027] Figure 9 This is a diagram showing an example of the physical structure of an encoding device.

[0028] Figure 10 This is a block diagram showing the functional configuration of the sound signal processing device according to the first embodiment.

[0029] Figure 11 This is a diagram showing a fan and a listener as an example of the objects according to the first embodiment.

[0030] Figure 12 This is a diagram showing audio data according to the first embodiment.

[0031] Figure 13 This is a diagram showing an example of a smoothing function according to the first embodiment.

[0032] Figure 14 This is a flowchart of Operation Example 1 of the sound signal processing device according to Embodiment 1.

[0033] Figure 15 This is a diagram for explaining the processing performed by the processing unit according to the first embodiment.

[0034] Figure 16 This is another diagram for explaining the processing performed by the processing unit according to the first embodiment.

[0035] Figure 17 This is a diagram showing sound data (aerodynamic sound data) according to the first embodiment.

[0036] Figure 18 This is a diagram showing R, which is a value represented by the smoothing function according to the first embodiment, and the amplification rate and attenuation rate of the volume of the aerodynamic sound.

[0037] Figure 19 This is a diagram showing divided aerodynamic sound data according to the first embodiment.

[0038] Figure 20 This is a diagram showing another example of two smoothing functions according to the first embodiment.

[0039] Figure 21 This is a diagram showing an example in which parameters for determining the smoothing function according to the first embodiment are changed.

[0040] Figure 22 This is a diagram showing another example of two smoothing functions according to the first embodiment.

[0041] Figure 23 This is a block diagram showing the functional configuration of an audio signal processing device according to a modified example.

[0042] Figure 24 This is a block diagram showing the functional configuration of the second processing unit according to the modification.

[0043] Figure 25 Graphs showing aerodynamic noise data related to a modified example.

[0044] Figure 26 This is a conceptual diagram of the processing performed by the second processing unit according to the modification.

[0045] Figure 27 This is a block diagram showing the functional configuration of a sampling rate conversion unit according to a modification.

[0046] Figure 28 This is a state transition diagram of the values ​​represented by the smoothing function according to the modification.

[0047] Figure 29 This is a block diagram showing another functional configuration of the audio signal processing device according to a modified example.

[0048] Figure 30 This is a block diagram showing the functional structure of an information processing device related to embodiment 2.

[0049] Figure 31 It is a diagram for explaining the reading of audio data according to the conventional technology and the reading of audio data according to the second embodiment.

[0050] Figure 32 This is a diagram for explaining processing performed by the information processing device according to the second embodiment.

[0051] Figure 33 This is a diagram used to illustrate other processing performed by the information processing device related to embodiment 2.

[0052] Figure 34 It is used to illustrate Figure 6 and Figure 7 A diagram showing an example of a functional block diagram and steps in which a rendering unit performs pipeline processing. DETAILED DESCRIPTION

[0053] (Understanding that forms the basis of this disclosure)

[0054] Patent Document 1 discloses a technique for outputting images and sounds to create a realistic virtual space. Patent Document 1 discloses a technique for changing the sound of the wind in accordance with changes in the strength of the wind in the virtual space.

[0055] A virtual space is a space in which users (listeners) of virtual reality (VR) or augmented reality (AR) exist. Wind sounds using the technology disclosed in Patent Document 1 are used by applications for reproducing three-dimensional sound in such virtual spaces. Such controlled sound is particularly useful in virtual spaces that sense the listener's 6DoF (degrees of Freedom) information. By utilizing the technology in Patent Document 1, natural phenomena such as wind can be reproduced in virtual spaces.

[0056] Furthermore, changes in natural phenomena in real space include fluctuations. Examples of natural phenomena in real space include the blowing of wind, the flow of rivers, and the movement of animals. For example, changes in natural phenomena include changes in wind speed or wind direction, and changes in wind speed or wind direction include fluctuations.

[0057] However, while the technology disclosed in Patent Document 1 allows listeners to hear the sound of wind, this wind sound cannot reproduce the sound of wind in real space, including undulations. Therefore, if listeners hear this wind sound, they will feel a sense of disharmony and will find it difficult to experience a sense of immersion. Therefore, there is a need for an acoustic signal processing method that can provide listeners with a sense of immersion.

[0058] Therefore, the sound signal processing method of the first technical solution of the present disclosure includes: an acquisition step of acquiring sound data representing a waveform of a reference sound; a processing step of processing the sound data based on simulation information that simulates changes in natural phenomena so as to change at least one of the frequency component, phase and amplitude value of the waveform; and an output step of outputting the processed sound data.

[0059] Thus, based on simulation information that simulates the fluctuations of natural phenomena involving fluctuations, sound data is processed to vary at least one of the frequency component, phase, and amplitude of the waveform. Consequently, at least one of the frequency component, phase, and amplitude fluctuates in the processed sound data, and at least one of the frequency component, phase, and amplitude fluctuates in the sound represented by the processed sound data. Consequently, the listener can hear a sound that fluctuates in at least one of the frequency component, phase, and amplitude, minimizing the sense of discomfort and allowing for a sense of immersive listening. This achieves an acoustic signal processing method that provides the listener with a sense of immersive listening.

[0060] The acoustic signal processing method of the second technical solution of the present disclosure is, in the signal processing method of the first technical solution, wherein the reference sound is aerodynamic sound generated by wind; and in the processing step, the sound data is processed based on the simulation information that simulates the change in wind speed of the wind so as to change at least one of the frequency component, phase and amplitude value of the waveform.

[0061] This allows the listener to hear aerodynamic sound that fluctuates in at least one of its frequency components, phase, and amplitude, minimizing discomfort and providing a sense of immersion. This achieves an acoustic signal processing method that provides a sense of immersion.

[0062] The sound signal processing method of the third technical solution of the present disclosure is, in the signal processing method of the second technical solution, wherein in the processing step, a smoothing function simulating the change in wind speed of the wind is determined as the simulation information; and based on the value represented by the determined smoothing function, the sound data is processed so that at least one of the frequency component, phase and amplitude value of the waveform changes.

[0063] This makes it possible to process audio data according to the value represented by the smoothing function.

[0064] The acoustic signal processing method of the fourth technical solution of the present disclosure is, in the signal processing method of the third technical solution, the value represented by the smoothing function is information representing the ratio of the wind speed of the aerodynamic sound serving as the reference sound to the wind speed of the aerodynamic sound represented by the sound data processed in the processing step.

[0065] This makes it possible to process the sound data based on the ratio of the wind speed of the aerodynamic sound serving as the reference sound to the wind speed of the aerodynamic sound represented by the processed sound data.

[0066] The acoustic signal processing method according to a fifth aspect of the present disclosure is the signal processing method according to the fourth aspect, wherein in the processing step, the smoothing function is determined so that a parameter defining the smoothing function changes irregularly.

[0067] This allows the listener to hear aerodynamic sound that fluctuates irregularly in at least one of its frequency components, phase, and amplitude, minimizing the listener's discomfort and providing a greater sense of immersion. This achieves an acoustic signal processing method that enhances the listener's sense of immersion.

[0068] The sound signal processing method of the sixth technical solution of the present disclosure is a signal processing method of any one of the third to fifth technical solutions, wherein, in the processing step, the sound data is processed so that the frequency component of the waveform is shifted to a frequency proportional to the value represented by the determined smoothing function.

[0069] This allows the listener to hear a sound with fluctuating frequency components, which reduces the listener's discomfort and allows the listener to experience a sense of immersion. In other words, an acoustic signal processing method that can provide the listener with a sense of immersion is achieved.

[0070] The sound signal processing method of the seventh technical solution of the present disclosure is, in the signal processing method of the third technical solution, characterized in that, in the processing step, the sound data is processed so that the amplitude value of the waveform changes proportionally to the αth power of the value represented by the determined smoothing function.

[0071] This allows the listener to hear a sound with fluctuating amplitude values, and the listener can experience a sense of presence without feeling any discomfort. In other words, an acoustic signal processing method that can provide the listener with a sense of presence is achieved.

[0072] The sound signal processing method of the eighth technical solution of the present disclosure is, in the signal processing method of the fourth or fifth technical solution, wherein in the processing step, the acquired sound data is divided into processing frames of a specified time, and the sound data is processed according to each of the divided processing frames.

[0073] This realizes an acoustic signal processing method that reduces the load of computational processing.

[0074] The sound signal processing method of the ninth technical solution of the present disclosure is the signal processing method of the eighth technical solution, wherein, in the processing step, the smoothing function is determined for each of the divided processing frames so that the value of the smoothing function becomes 1.0 at the first moment and the last moment of the processing frame.

[0075] This suppresses the occurrence of noise at the junction between a processing frame and the next processing frame.

[0076] According to a tenth aspect of the present disclosure, in the signal processing method according to the ninth aspect, in the processing step, a parameter of the smoothing function is determined for each of the divided processing frames.

[0077] This realizes an acoustic signal processing method that reduces the load of computational processing.

[0078] According to an eleventh technical solution of the present disclosure, in the signal processing method according to the tenth technical solution, the parameter is a time from the first time to the last time.

[0079] In this way, the parameter can be set to the time from the first time point of the processing frame to the last time point of the processing frame.

[0080] In the acoustic signal processing method according to a twelfth technical aspect of the present disclosure, in the signal processing method according to the tenth technical aspect, the parameter is a value related to a maximum value of the smoothing function.

[0081] This allows the parameter to be set to a value related to the maximum value of the smoothing function.

[0082] According to a 13th technical solution of the present disclosure, in the signal processing method according to the 10th technical solution, the parameter is a parameter that changes a position at which the smoothing function reaches a maximum value.

[0083] This allows the parameters to be set so as to change the position where the smoothing function reaches its maximum value.

[0084] In the acoustic signal processing method according to a fourteenth technical aspect of the present disclosure, in the signal processing method according to the tenth technical aspect, the parameter is a parameter that changes the sharpness of the change of the smoothing function.

[0085] This makes it possible to set a parameter that changes the sharpness of the change of the smoothing function.

[0086] The sound signal processing method according to the 15th technical solution of the present disclosure is, in the signal processing method according to the 10th technical solution, wherein, in the processing step, a first parameter and a second parameter of the smoothing function are determined; based on the smoothing function determined by the determined first parameter, the obtained sound data is processed so as to change at least one of the frequency component, phase and amplitude value of the waveform; based on the smoothing function determined by the determined second parameter, the obtained sound data is processed so as to change at least one of the frequency component, phase and amplitude value of the waveform; in the output step, the sound data processed based on the smoothing function determined by the determined first parameter is output to a first output channel; and the sound data processed based on the smoothing function determined by the determined second parameter is output to a second output channel.

[0087] This makes it possible to output different audio data for each output channel.

[0088] The acoustic signal processing method of the 16th technical solution of the present disclosure is a signal processing method of any one of the technical solutions 10 to 15, wherein the aerodynamic sound is the sound generated by the collision of the wind with the object; and in the processing step, the properties of the wind speed are simulated to determine the parameters.

[0089] By simulating the change in wind speed of wind including fluctuations, parameters can be determined. Based on the smoothing function determined by the parameters, the sound data can be processed to change at least one of the frequency component, phase, and amplitude of the waveform.

[0090] The acoustic signal processing method of the 17th technical solution of the present disclosure is, in the signal processing method of any one of the technical solutions 10 to 15, the aerodynamic sound is the sound produced by the wind colliding with the ears of the listener who listens to the aerodynamic sound; and in the processing step, the properties of the wind direction of the wind are simulated to determine the parameters.

[0091] By simulating the change in wind direction including undulations, parameters are determined. The sound data can be processed based on the smoothing function determined by the parameters so as to change at least one of the frequency component, phase, and amplitude of the waveform.

[0092] According to an eighteenth aspect of the present disclosure, in the signal processing method according to the eighth aspect, a maximum value of the smoothing function does not exceed 3.

[0093] This makes it possible to set the maximum value of the smoothing function to 3 or less.

[0094] According to a nineteenth aspect of the present disclosure, in the signal processing method according to the eighth aspect, a minimum value of the smoothing function is not less than zero.

[0095] This makes it possible to set the minimum value of the smoothing function to be equal to or greater than 0.

[0096] The sound signal processing method of the 20th technical solution of the present disclosure includes an acceptance step in the signal processing method of the 8th technical solution, in which an instruction of Va designated as the wind speed of the wind and Vp designated as the instantaneous wind speed of the wind is accepted; in the processing step, the smoothing function is determined so that the maximum value of the smoothing function becomes Vp / Va.

[0097] Thereby, the maximum value of the smoothing function can be set to Vp / Va.

[0098] According to a twenty-first aspect of the present disclosure, in the signal processing method according to the eighth aspect, an average value of the predetermined time is 3 seconds.

[0099] Thereby, the average value of the predetermined time, which is the time length of the processing frame, can be set to 3 seconds.

[0100] In the acoustic signal processing method according to the twenty-second technical solution of the present disclosure, in the signal processing method according to the sixteenth technical solution, the object is an object having a shape imitating an ear.

[0101] This makes it possible to collect aerodynamic sound using, for example, a dummy head-shaped microphone.

[0102] A computer program according to a twenty-third aspect of the present disclosure is a computer program for causing a computer to execute the sound signal processing method according to any one of the first to twenty-second aspects.

[0103] Thereby, the computer can execute the above-mentioned sound signal processing method according to the computer program.

[0104] The sound signal processing device of the 24th technical solution of the present disclosure comprises: an acquisition unit, which acquires sound data of a waveform representing a reference sound; a processing unit, which processes the sound data based on simulation information that simulates changes in natural phenomena so as to change at least one of the frequency component, phase and amplitude value of the waveform; and an output unit, which outputs the processed sound data.

[0105] Thus, based on simulation information that simulates the fluctuations of natural phenomena involving fluctuations, sound data is processed to change at least one of the frequency component, phase, and amplitude of the waveform. Consequently, at least one of the frequency component, phase, and amplitude fluctuates in the processed sound data, and at least one of the frequency component, phase, and amplitude fluctuates in the sound represented by the processed sound data. Consequently, the listener can hear a sound that fluctuates in at least one of the frequency component, phase, and amplitude, minimizing the sense of discomfort and allowing for a sense of immersive listening. This results in an audio signal processing device that provides the listener with a sense of immersive listening.

[0106] Furthermore, these inclusive or specific technical solutions may also be implemented by systems, devices, methods, integrated circuits, computer programs, or non-temporary recording media such as computer-readable CD-ROMs, or by any combination of systems, devices, methods, integrated circuits, computer programs, and recording media.

[0107] Hereinafter, embodiments will be described in detail with reference to the drawings.

[0108] The numerical values, shapes, materials, components, configuration positions and connection forms of components, steps, and the order of steps shown in the following embodiments are examples and are not intended to limit the present disclosure.

[0109] In the following description, elements may be assigned ordinal numbers such as 1 and 2. These ordinal numbers are assigned to elements for identification purposes and do not necessarily correspond to a meaningful order. These ordinal numbers may be replaced, newly assigned, or removed as appropriate.

[0110] In addition, each figure is a schematic diagram and does not necessarily illustrate the exact figure. Therefore, the scales, etc. in each figure are not necessarily the same. In each figure, the same reference numerals are given to substantially the same structure, and repeated descriptions are omitted or simplified.

[0111] In this specification, terms such as vertical and numerical ranges indicating relationships between elements do not represent relationships in a strict sense, but rather represent substantially equivalent ranges, including differences of several percentage points, for example.

[0112] (Implementation 1)

[0113] [Examples of Devices to Which the Audio Processing Technology or Encoding / Decoding Technology of the Present Disclosure Can Be Applied]

[0114] <Stereo Sound Reproduction System>

[0115] Figure 1 FIG1 is a diagram showing an immersive audio reproduction system A0000 as an example of a system to which the audio processing or decoding processing of the present disclosure can be applied. The immersive audio reproduction system A0000 includes an audio signal processing device A0001 and an audio presentation device A0002.

[0116] The sound signal processing device A0001 performs sound processing on the sound signal emitted by the virtual sound source to generate a sound signal after sound processing to prompt the listener (i.e., the listener). The sound signal is not limited to the human voice, as long as it is audible sound. For example, sound processing is a signal processing performed on the sound signal in order to reproduce one or more sound-related effects received from the time the sound is emitted to the time the listener hears it. The sound signal processing device A0001 performs sound processing based on information that describes the cause of the above-mentioned sound-related effects. Spatial information includes, for example, information indicating the position of the sound source, the listener, and surrounding objects, information indicating the shape of the space, parameters related to the propagation of sound, etc. The sound signal processing device A0001 is, for example, a PC (Personal Computer), a smartphone, a tablet computer, or a game console.

[0117] The signal after the sound processing is prompted to the listener (user) from the sound prompting device A0002. The sound prompting device A0002 is connected to the sound signal processing device A0001 via wireless or wired communication. The sound signal after the sound processing generated by the sound signal processing device A0001 is transmitted to the sound prompting device A0002 via wireless or wired communication. In the case where the sound prompting device A0002 is composed of multiple devices such as a device for the right ear and a device for the left ear, the multiple devices communicate with the sound signal processing device A0001 or each of the multiple devices communicates with the sound signal processing device A0001, and the multiple devices prompt the sound synchronously. The sound prompting device A0002 is, for example, headphones worn on the listener's head, earplugs, a head-mounted display, or a surround speaker composed of multiple fixed speakers.

[0118] Furthermore, the stereo sound reproduction system A0000 may be used in combination with an image presentation device or a stereoscopic image presentation device that visually provides an ER (Extended Reality) experience including VR or AR.

[0119] in addition, Figure 1 The system configuration example in which the sound signal processing device A0001 and the sound prompting device A0002 are different devices is shown. However, the stereo sound reproduction system A0000 to which the sound signal processing method or decoding method disclosed herein can be applied is not limited to the system configuration example in which the sound signal processing device A0001 and the sound prompting device A0002 are different devices. Figure 1 For example, the sound signal processing device A0001 may be included in the sound prompting device A0002, and the sound prompting device A0002 may perform both sound processing and sound prompting. Furthermore, the sound signal processing device A0001 and the sound prompting device A0002 may share the sound processing described in this disclosure, or a server connected to the sound signal processing device A0001 or the sound prompting device A0002 via a network may perform part or all of the sound processing described in this disclosure.

[0120] In the above description, the sound signal processing device A0001 is referred to as the sound signal processing device A0001. However, when the sound signal processing device A0001 performs sound processing by decoding a bit stream generated by encoding data of at least a portion of spatial information used in a sound signal or sound processing, the sound signal processing device A0001 may also be referred to as a decoding device.

[0121] <Example of Encoding Device>

[0122] Figure 2 This is a functional block diagram showing the configuration of an encoding device A0100 as an example of an encoding device according to the present disclosure.

[0123] Input data A0101 is data to be encoded, including spatial information and / or a sound signal, input to encoder A0102. Details of spatial information will be described later.

[0124] The encoder A0102 encodes the input data A0101 to generate encoded data A0103. The encoded data A0103 is, for example, a bit stream generated by the encoding process.

[0125] The memory A0104 stores the encoded data A0103. The memory A0104 may be, for example, a hard disk or an SSD (Solid-State Drive), or other memory.

[0126] In addition, in the above description, a bit stream generated by encoding processing is cited as an example of the encoded data A0103 stored in the memory A0104, but it can also be data other than a bit stream. For example, the encoding device A0100 can also store the converted data generated by converting the bit stream into a specified data format in the memory A0104. The converted data can also be, for example, a file or multiplexed stream that stores one or more bit streams. Here, the file is a file having a file format such as ISOBMFF (ISO Base Media File Format). In addition, the encoded data A0103 can also be in the form of multiple packets generated by dividing the above-mentioned bit stream or file. When the bit stream generated by the encoder A0102 is converted into data different from the bit stream, the encoding device A0100 can also have a conversion unit not shown in the figure, or the conversion processing can be performed by the CPU (Central Processing Unit).

[0127] <Example of Decoding Device>

[0128] Figure 3 This is a functional block diagram showing the configuration of a decoding device A0110 as an example of a decoding device according to the present disclosure.

[0129] Memory A0114 stores, for example, the same data as encoded data A0103 generated by encoding device A0100. Memory A0114 reads the stored data and inputs it as input data A0113 to decoder A0112. Input data A0113 is, for example, a bitstream to be decoded. Memory A0114 can be, for example, a hard disk, an SSD, or other memory.

[0130] Alternatively, the decoding device A0110 may use the data stored in the memory A0114 as input data A0113, rather than the data stored therein as is, but may use transformed data generated by transforming the data read therefrom as input data A0113. The data before transformation may be, for example, multiplexed data storing one or more bit streams. Here, the multiplexed data may be a file having a file format such as ISOBMFF. Furthermore, the data before transformation may be in the form of multiple packets generated by segmenting the aforementioned bit stream or file. When data different from the bit stream read from the memory A0114 is transformed into a bit stream, the decoding device A0110 may include a transformation unit (not shown), or the transformation process may be performed by the CPU.

[0131] The decoder A0112 decodes the input data A0113 to generate a sound signal A0111 to be presented to the listener.

[0132] <Another Example of Encoding Device>

[0133] Figure 4 This is a functional block diagram showing the configuration of an encoding device A0120 as another example of the encoding device disclosed herein. Figure 4 For Figure 2 The same function as the Figure 2 The same reference numerals are used for the components, and descriptions of these components are omitted.

[0134] The coding device A0100 stores the coded data A0103 in the memory A0104 , but the coding device A0120 is different from the coding device A0100 in that it includes a transmission unit A0121 that transmits the coded data A0103 to the outside.

[0135] Transmitter A0121 transmits a transmission signal A0122 to another device or server based on coded data A0103 or data in another data format generated by transforming coded data A0103. The data used to generate transmission signal A0122 is, for example, the bitstream, multiplexed data, file, or packet described in connection with coding device A0100.

[0136] <Another Example of Decoding Device>

[0137] Figure 5 This is a functional block diagram showing the configuration of a decoding device A0130 as another example of the decoding device disclosed herein. Figure 5 For Figure 3 The same function as the Figure 3 The same reference numerals are used for the components, and descriptions of these components are omitted.

[0138] The decoding device A0110 reads the input data A0113 from the memory A0114 . However, the decoding device A0130 differs from the decoding device A0110 in that it includes a receiving unit A0131 that receives the input data A0113 from the outside.

[0139] The receiving unit A0131 receives the received signal A0132, obtains received data, and outputs input data A0113 to the decoder A0112. The received data may be the same as the input data A0113 input to the decoder A0112, or may be in a different data format from the input data A0113. If the received data is in a different data format from the input data A0113, the receiving unit A0131 may convert the received data into the input data A0113, or a conversion unit (not shown) or CPU included in the decoding device A0130 may convert the received data into the input data A0113. The received data may be, for example, the bitstream, multiplexed data, file, or packet described above for the encoding device A0120.

[0140] <Decoder Functional Description>

[0141] Figure 6 It means as Figure 3 or Figure 5 A functional block diagram showing the structure of a decoder A0200 as an example of the decoder A0112 in FIG.

[0142] Input data A0113 is a coded bit stream, and includes coded audio data, which is a coded audio signal, and metadata used in audio processing.

[0143] The spatial information management unit A0201 obtains the metadata contained in the input data A0113 and parses the metadata. The metadata contains information describing the elements that act on the sound and are configured in the sound space. The spatial information management unit A0201 manages the spatial information required for the sound processing obtained by parsing the metadata, and provides the spatial information to the rendering unit A0203. In addition, in the present disclosure, the information used in the sound processing is referred to as spatial information, but it can also be called something else. The information used in the sound processing can be called, for example, sound space information or scene information. In addition, in the case where the information used in the sound processing changes over time, the spatial information input to the rendering unit A0203 can also be called a spatial state, a sound space state, a scene state, etc.

[0144] Furthermore, spatial information can be managed by sound space or by scene. For example, when different rooms are represented as virtual spaces, the spatial information of each room can be managed as a scene of a different sound space, or the spatial information can be managed as a scene that is different depending on the occasion of representation even though it is the same space. In the management of spatial information, an identifier for identifying each piece of spatial information can also be assigned. The data of the spatial information can be included in a bit stream as a form of input data, or the bit stream can include an identifier of the spatial information and the data of the spatial information can be obtained from outside the bit stream. In the case where only the identifier of the spatial information is included in the bit stream, the identifier of the spatial information can be used during rendering to obtain the data of the spatial information stored in the memory of the sound signal processing device A0001 or in an external server as input data.

[0145] Furthermore, the information managed by the spatial information management unit A0201 is not limited to information contained in the bitstream. For example, the input data A0113 may include data representing spatial characteristics or structure obtained from a software application or server providing VR or AR, as data not included in the bitstream. Furthermore, for example, the input data A0113 may include data representing the characteristics or position of a listener or object, as data not included in the bitstream. Furthermore, the input data A0113 may include information representing the listener's position obtained by sensors included in a terminal including a decoding device, or information representing the terminal's position inferred based on information obtained by the sensors. In other words, the spatial information management unit A0201 may communicate with an external system or server to obtain spatial information and the listener's position. Furthermore, the spatial information management unit A0201 may obtain clock synchronization information from an external system and perform processing synchronized with the clock of the rendering unit A0203. In addition, the space in the above description may be a virtually formed space, i.e., a VR space, or a real space (i.e., a real space) or a virtual space corresponding to a real space, i.e., AR or MR (Mixed Reality). In addition, the virtual space may also be referred to as a sound field or a sound space. In addition, the information indicating the position in the above description may be information such as coordinate values ​​indicating a position in the space, information indicating a relative position relative to a predetermined reference position, or information indicating movement or acceleration of a position in the space.

[0146] The sound data decoder A0202 decodes the encoded sound data included in the input data A0113 to obtain a sound signal.

[0147] The coded audio data obtained by the stereo sound reproduction system A0000 is, for example, a bitstream encoded in a format specified by MPEG-H 3D Audio (ISO / IEC 23008-3). MPEG-H 3D Audio is merely one example of a coding method that can be used to generate the coded audio data contained in the bitstream, and may also include bitstreams and coded audio data encoded using other coding methods. For example, the coding method used may be a non-reversible codec such as MP3 (MPEG-1 Audio Layer-3), AAC (Advanced Audio Coding), WMA (Windows Media Audio), AC3 (Audio Codec-3), or Vorbis; a reversible codec such as ALAC (Apple Lossless Audio Codec) or FLAC (Free Lossless Audio Codec); or any other coding method may be used. For example, PCM (pulse code modulation) data may be used as one type of coded audio data. In this case, the decoding process may be a process of converting an N-bit binary number into a number format (eg, floating point format) that can be processed by the rendering unit A0203, for example, when the number of quantization bits of the PCM data is N.

[0148] The rendering unit A0203 takes the sound signal and spatial information as input, performs acoustic processing on the sound signal using the spatial information, and outputs the sound signal A0111 after the acoustic processing.

[0149] Before rendering begins, the spatial information management unit A0201 reads metadata from the input signal, detects rendering items such as objects and sounds specified by the spatial information, and sends them to the rendering unit A0203. After rendering begins, the spatial information management unit A0201 monitors changes in the spatial information and the listener's position over time, updates the spatial information, and manages it. The spatial information management unit A0201 then sends the updated spatial information to the rendering unit A0203. Based on the sound signal included in the input data A0113 and the spatial information received from the spatial information management unit A0201, the rendering unit A0203 generates and outputs an acoustically processed sound signal.

[0150] The spatial information update process and the audio signal output process with added acoustic processing can be executed by the same thread, or the spatial information management unit A0201 and the rendering unit A0203 can be assigned to separate threads. If the spatial information update process and the audio signal output process with added acoustic processing are handled by different threads, the thread activation frequency can be set independently, or the processes can be executed in parallel.

[0151] By having the spatial information management unit A0201 and the rendering unit A0203 execute processing in separate, independent threads, computing resources can be preferentially allocated to the rendering unit A0203. This allows for safe implementation of sound processing where even minimal delay is unacceptable, such as sound processing that generates small noise even with a delay of 1 sample (0.02 msec). In this case, the allocation of computing resources to the spatial information management unit A0201 is restricted. However, updating spatial information is a lower-frequency process than outputting sound signals (for example, updating the orientation of the listener's face). Therefore, unlike outputting sound signals, which requires instantaneous response, limiting the allocation of computing resources does not significantly impact the sound quality for the listener.

[0152] Spatial information updates can be performed periodically at pre-set times or intervals, or when pre-set conditions are met. Furthermore, spatial information updates can be performed manually by the listener or the administrator of the sound space, or triggered by changes in external systems. For example, if a listener manipulates a controller to momentarily tilt their avatar's standing position, or if the virtual space administrator suddenly changes the scene environment, the thread configured with the spatial information management unit A0201 can be activated as a one-shot interrupt process in addition to the regular activation.

[0153] The information update thread that performs the updating process of spatial information is responsible for, for example, updating the position or orientation of the listener's avatar arranged in the virtual space based on the position or orientation of the VR goggles worn by the listener, and updating the position of objects moving in the virtual space, etc., and is provided in a processing thread that is started at a relatively low frequency of about tens of Hz. It is also possible to perform processing that reflects the properties of direct sound with such a low-occurrence processing thread. This is because the frequency of changes in the properties of direct sound is lower than the frequency of occurrence of audio processing frames used for audio output. On the contrary, doing so can relatively reduce the computational load of the processing, and can also avoid the risk of generating impulse noise if the information is updated at an unnecessarily high frequency.

[0154] Figure 7 It means as Figure 3 or Figure 5 A functional block diagram showing the structure of a decoder A0210 which is another example of the decoder A0112 in FIG.

[0155] Figure 7 The decoder A0210 shown is different from the decoder A0210 in that the input data A0113 does not include coded audio data but includes an uncoded audio signal. Figure 6 The decoder A0200 shown is different. Input data A0113 includes a bit stream including metadata and an audio signal.

[0156] Space Information Management Department A0211 and Figure 6 The spatial information management unit A0201 is the same as that of FIG. 1 , so the description is omitted.

[0157] Rendering Department A0213 and Figure 6 The rendering part A0203 is the same, so the description is omitted.

[0158] In addition, in the above description Figure 7 The structure of the decoder A0210 can be called a decoder, but it can also be called an audio processing unit that performs audio processing. In addition, the device including the audio processing unit can also be called an audio processing device instead of a decoder. In addition, the audio signal processing device A0001 can also be called an audio processing device.

[0159] <Physical Structure of the Acoustic Signal Processing Device>

[0160] Figure 8 : is a diagram showing an example of the physical structure of the sound signal processing device. Figure 8 The audio signal processing device may also be a decoding device. In addition, a part of the structure described here may also be equipped in the sound prompting device A0002. In addition, Figure 8 The illustrated sound signal processing device is an example of the above-mentioned sound signal processing device A0001.

[0161] Figure 8 The audio signal processing device includes a processor, a memory, a communication IF, a sensor, and a speaker.

[0162] The processor may be, for example, a CPU (Central Processing Unit), a DSP (Digital Signal Processor), or a GPU (Graphics Processing Unit). The CPU, DSP, or GPU may execute a program stored in memory to implement the audio processing or decoding processing disclosed herein. Alternatively, the processor may be a dedicated circuit that performs signal processing on sound signals, including the audio processing disclosed herein.

[0163] Memory is composed of, for example, RAM (Random Access Memory) or ROM (Read Only Memory). It can also include magnetic storage media such as hard disks or semiconductor memory such as SSDs (Solid State Drives). Furthermore, the term "memory" also includes internal memory embedded in the CPU or GPU.

[0164] The communication IF (Interface) is a communication module corresponding to a communication method such as Bluetooth (registered trademark) or WIGIG (registered trademark). Figure 8 The acoustic signal processing device shown has a function of communicating with another communication device via a communication IF, and acquires a bit stream to be decoded, and stores the acquired bit stream in, for example, a memory.

[0165] The communication module is composed of, for example, a signal processing circuit and an antenna corresponding to the communication method. In the above example, Bluetooth (registered trademark) or WIGIG (registered trademark) is cited as an example of a communication method, but it can also correspond to communication methods such as LTE (Long Term Evolution), NR (New Radio) or Wi-Fi (registered trademark). In addition, the communication IF may not be a wireless communication method as described above, but a wired communication method such as Ethernet (registered trademark), USB (Universal Serial Bus), HDMI (registered trademark) (High-Definition Multimedia Interface), etc.

[0166] The sensor performs sensing to estimate the position or orientation of the listener. Specifically, the sensor estimates the listener's position and / or orientation based on one or more detection results of the position, orientation, movement, speed, angular velocity, or acceleration of a part or the entire body, such as the listener's head, and generates position information representing the listener's position and / or orientation. Furthermore, the position information may be information representing the listener's position and / or orientation in real space, or information representing a displacement of the listener's position and / or orientation relative to the listener's position and / or orientation at a specified point in time. Furthermore, the position information may be information representing the relative position and / or orientation to the stereophonic sound reproduction system A0000 or an external device equipped with a sensor.

[0167] The sensor may be, for example, an imaging device such as a camera or a distance measuring device such as LiDAR (Light Detection and Ranging), which may capture the movement of the listener's head and detect the movement of the listener's head by processing the captured image. Alternatively, a device that estimates position using wireless signals in any frequency band, such as millimeter waves, may be used as a sensor.

[0168] in addition, Figure 8 The acoustic signal processing device shown in FIG. 1 may also obtain position information from an external device having a sensor via a communication IF. In this case, the acoustic signal processing device may not include a sensor. Here, the external device is, for example, Figure 1 The sound presentation device A0002 described in , or the stereoscopic image reproduction device worn on the head of the listener, etc. In this case, the sensor is composed of a combination of various sensors such as a gyro sensor and an acceleration sensor.

[0169] For example, as the speed of the listener's head movement, the sensor can detect the angular velocity of rotation with at least one of the three axes orthogonal to each other in the sound space as the rotation axis, and can also detect the acceleration of displacement with at least one of the above three axes as the displacement direction.

[0170] For example, as a measure of the listener's head movement, the sensor can detect both rotation about at least one of three mutually orthogonal axes within the sound space, and displacement about at least one of these three axes. Specifically, the sensor detects 6DoF (position (x, y, z) and angles (yaw, pitch, roll)) as the listener's position. The sensor is constructed by combining various sensors for motion detection, such as gyroscopes and accelerometers.

[0171] The sensor only needs to detect the listener's position and can be implemented by a camera or a GPS (Global Positioning System) receiver. Position information obtained by self-estimation using LiDAR (Laser Imaging Detection and Ranging) or the like can also be used. For example, when implementing the audio signal playback system in a smartphone, the sensor is built into the smartphone.

[0172] In addition, the sensor can also include detection Figure 8 A temperature sensor such as a thermocouple for detecting the temperature of the audio signal processing device shown, and a sensor for detecting the remaining amount of a battery included in or connected to the audio signal processing device, etc.

[0173] A speaker comprises a driving mechanism, such as a diaphragm, magnet, or voice coil, and an amplifier, and produces acoustically processed sound signals for the listener as sound. The speaker activates the driving mechanism based on the sound signal (more specifically, a waveform signal representing the sound waveform) amplified by the amplifier, which in turn causes the diaphragm to vibrate. The diaphragm vibrates in response to the sound signal, generating sound waves. These sound waves propagate through the air and reach the listener's ears, where they are perceived as sound.

[0174] In addition, here are Figure 8 The example of the case where the audio signal processing device shown in the figure has a speaker and prompts the audio signal after audio processing via the speaker is described, but the prompting mechanism of the audio signal is not limited to the above-mentioned structure. For example, the audio signal after audio processing can also be output to the external audio prompting device A0002 connected by the communication module. The communication performed by the communication module can be either wired or wireless. In addition, as another example, it can also be Figure 8 The audio signal processing device shown has a terminal for outputting an analog audio signal, and a cable of an earphone or the like is connected to the terminal to produce an audio signal from the earphone or the like. In the above-described case, the audio presentation device A0002, such as headphones, earphones, a head-mounted display, a neck speaker, a wearable speaker, or a surround speaker system composed of a plurality of fixed speakers, which is worn on the listener's head or a part of the body, reproduces the audio signal.

[0175] <Physical Structure of Encoding Device>

[0176] Figure 9 : is a diagram showing an example of the physical structure of the encoding device. Figure 9 The illustrated encoding device is an example of the encoding devices A0100 and A0120 described above.

[0177] Figure 9 The encoding device includes a processor, a memory and a communication IF.

[0178] The processor is, for example, a CPU (Central Processing Unit) or a DSP (Digital Signal Processor). The encoding process of the present disclosure may be implemented by executing a program stored in a memory by the CPU or GPU. Alternatively, the processor may be a dedicated circuit that performs signal processing on audio signals, including the encoding process of the present disclosure.

[0179] Memory is composed of, for example, RAM (Random Access Memory) or ROM (Read Only Memory). It can also include magnetic storage media such as hard disks or semiconductor memory such as SSDs (Solid State Drives). Furthermore, it can also include internal memory embedded in the CPU or GPU.

[0180] The communication IF (Interface) is a communication module compatible with a communication method such as Bluetooth (registered trademark) or WIGIG (registered trademark). The encoding device has a function of communicating with other communication devices via the communication IF and transmits an encoded bit stream.

[0181] The communication module is composed of, for example, a signal processing circuit and an antenna corresponding to the communication method. In the above examples, Bluetooth (registered trademark) or Wi-Fi (registered trademark) are used as examples of communication methods, but it can also correspond to communication methods such as LTE (Long Term Evolution), NR (New Radio), or Wi-Fi (registered trademark). In addition, the communication IF may not be a wireless communication method as described above, but a wired communication method such as Ethernet (registered trademark), USB (Universal Serial Bus), HDMI (registered trademark) (High-Definition Multimedia Interface), etc.

[0182] [constitute]

[0183] Next, the configuration of the sound signal processing device 100 according to the first embodiment will be described. Figure 10 1 is a block diagram showing the functional configuration of the sound signal processing device 100 according to the present embodiment.

[0184] The sound signal processing device 100 according to this embodiment is a device for acquiring, processing, and outputting sound data representing the waveform of a reference sound. By outputting this sound data, a listener can hear the sound represented by this sound data. For example, the sound signal processing device 100 according to this embodiment is used in various applications within virtual spaces, such as virtual reality or augmented reality (VR or AR).

[0185] The reference sound can be any sound, for example, a sound related to a natural phenomenon. In this embodiment, the natural phenomenon is not particularly limited as long as it occurs in nature, and examples include the sound of wind, the flow of a river, the movement of animals, and the like. Examples of sounds related to natural phenomena include the sound of wind, the gurgling sound of a river, and the calls of animals.

[0186] Here, focusing on the sound produced by wind, we can cite aerodynamic sound, which is produced by wind colliding with objects in a virtual space. This aerodynamic sound is produced when wind, for example, reaches and strikes a listener's ears. Thus, aerodynamic sound originates from the wind blowing in a virtual space.

[0187] In the present embodiment, the reference sound is the aerodynamic sound generated by the wind W. However, the present invention is not limited thereto, and the reference sound may be the gurgling sound of a flowing river or the cry of an animal.

[0188] As an example, the wind in the virtual space is the wind caused by objects in the virtual space.

[0189] Figure 11 This figure shows a fan FN and a listener L, as an example of an object related to this embodiment. When the object is an object capable of blowing air, such as a fan FN, aerodynamic sound is produced when the wind W generated by the fan FN reaches the listener L. More specifically, aerodynamic sound is produced when the wind W blown out from the fan FN reaches the listener L and corresponds to the shape of the listener L's ear, for example.

[0190] Furthermore, for example, when the object is a moving object (eg, a vehicle), aerodynamic sound is aerodynamic sound generated when wind W generated by the movement of the object reaches the listener L.

[0191] Furthermore, as an example, the wind W in the virtual space is wind that is generated naturally in the real space and is reproduced in the virtual space (hereinafter referred to as natural wind), and is wind whose generating location in the virtual space cannot be determined. If the wind W in the virtual space is natural wind, it can also be said that it is not wind caused by an object.

[0192] The object in this embodiment is not limited to the fan FN. The object in the virtual space is not particularly limited as long as it is an object included in the content (here, as an example, a video) displayed on the display unit 300 that displays the content executed in the virtual space.

[0193] For example, an object can be a moving object that generates wind by its positional movement. Examples of moving objects include objects representing plants, animals, man-made objects, and natural objects. Examples of objects representing man-made objects include cars, bicycles, and airplanes. Furthermore, examples of objects representing man-made objects include sporting goods such as baseball bats and tennis rackets, and furniture such as tables, chairs, and clocks. Furthermore, as an example, an object can be at least one of an object that can move within the content and an object that is moved.

[0194] Furthermore, for example, the object may be an object capable of blowing air. Examples of such an object include a circulator, a fan, an air conditioner, and the like in addition to the fan FN described above.

[0195] Alternatively, an object may be one that generates sound. The sound generated by an object is represented by sound data associated with the object (hereinafter sometimes referred to as object sound data). For example, if the object is a fan FN, the sound generated by the object is the motor sound generated by the fan FN's motor. Alternatively, if the object is an ambulance, the sound generated by the object is the sound of the ambulance's siren.

[0196] The acoustic signal processing device 100 processes sound data (aerodynamic sound data) representing the waveform of a reference sound of aerodynamic sound in a virtual space, and outputs the processed sound data to the earphone 200. In the following, the sound data representing the waveform of the reference sound (aerodynamic sound) may be referred to as aerodynamic sound data.

[0197] Next, the earphones 200 will be described.

[0198] Headphones 200 are a device that reproduces aerodynamic sounds and serves as a sound output device that presents aerodynamic sounds to a listener L. More specifically, headphones 200 reproduce aerodynamic sounds based on aerodynamic sound data output by the acoustic signal processing device 100. This allows listener L to hear the aerodynamic sounds. Furthermore, other output channels, such as speakers, may be used in addition to headphones 200.

[0199] like Figure 10 As shown, the headset 200 includes a head sensor unit 201 and an output unit 202 .

[0200] The head sensor unit 201 senses the position of the listener L in the virtual space, which is determined by horizontal coordinates and vertical height, and outputs second position information indicating the position of the listener L of aerodynamic sound in the virtual space to the audio signal processing device 100 .

[0201] The head sensor unit 201 may sense 6DoF information of the head of the listener L. For example, the head sensor unit 201 may be an inertial measurement unit (IMU), an accelerometer, a gyroscope, a magnetic sensor, or a combination thereof.

[0202] The output unit 202 is a device that reproduces, in a sound reproduction space, the sound reaching the listener L. More specifically, the output unit 202 reproduces the aerodynamic sound based on the aerodynamic sound data representing the aerodynamic sound output from the acoustic signal processing device 100 .

[0203] Next, the display unit 300 will be described.

[0204] The display unit 300 is a display device that displays content (images) including objects in a virtual space. The processing used by the display unit 300 to display content will be described later. The display unit 300 is implemented by a display panel such as a liquid crystal panel or an organic EL (ElectroLuminescence) panel.

[0205] Furthermore, Figure 10 In this embodiment, the acoustic signal processing device 100 acquires and processes sound data (aerodynamic sound data) representing a waveform of a reference sound of aerodynamic sound in a virtual space, and outputs the processed sound data to the earphone 200 .

[0206] like Figure 10 As shown, the sound signal processing device 100 includes an acquisition unit 110 , a processing unit 120 , an output unit 130 , a storage unit 140 , and a reception unit 150 .

[0207] The acquisition unit 110 acquires sound data representing the waveform of the reference sound (pneumatic sound). Figure 12 : is a diagram showing the sound data related to this embodiment. Figure 12 As shown, the sound data is data showing a waveform showing time and amplitude, for example, and is aerodynamic sound data in this case.

[0208] The sound data (pneumatic sound data) is stored in the storage unit 140 , and the acquisition unit 110 acquires the sound data (pneumatic sound data) stored in the storage unit 140 .

[0209] The acquisition unit 110 acquires first position information indicating the position of an object (eg, fan FN). If the object generates sound, the acquisition unit 110 acquires object sound data indicating the sound. The acquisition unit 110 also acquires shape information indicating the shape of the object.

[0210] The acquisition unit 110 acquires the second position information. As described above, the second position information is information indicating the position of the listener L in the virtual space.

[0211] The acquisition unit 110 may, for example, acquire sound data representing the waveform of the reference sound, the first position information, the target sound data, the shape information, and the second position information from the input signal. Furthermore, the acquisition unit 110 may acquire sound data representing the waveform of the reference sound, the first position information, the target sound data, the shape information, and the second position information from other sources. The input signal will be described below. In the following, the sound data representing the waveform of the reference sound (aerodynamic sound data) and the target sound data may be collectively referred to as sound data.

[0212] The input signal may be composed of, for example, spatial information, sensor information, and sound data (sound signal). Furthermore, the above information and sound data may be contained in a single input signal or in multiple different signals. The input signal may also include a bitstream composed of sound data and metadata (control information). In this case, the metadata may also include spatial information and information for identifying the sound data.

[0213] The audio data representing the waveform of the reference sound, the first position information, the target sound data, the shape information, and the second position information described above may also be included in the input signal. More specifically, the first position information and shape information may be included in the spatial information, and the second position information may be generated based on information obtained from sensor information. The sensor information may be obtained from the head sensor unit 201 or from another external device.

[0214] Spatial information is information related to the sound space (three-dimensional sound field) created by the stereo sound reproduction system A0000. It consists of information about objects contained in the sound space and information about the listener. Objects include sound source objects, which emit sound and serve as sound sources, and non-sound-emitting objects, which do not emit sound. Non-sound-emitting objects function as obstruction objects that reflect the sound emitted by sound source objects, but they can also function as sound source objects that reflect the sound emitted by other sound source objects. Obstruction objects are also referred to as reflection objects.

[0215] Information commonly assigned to a sound source object and a non-sound emitting object includes position information, shape information, and a volume attenuation rate when the object reflects sound.

[0216] Position information is represented by coordinate values ​​along three axes in Euclidean space, such as the X, Y, and Z axes, but it does not necessarily need to be three-dimensional information. Position information can also be two-dimensional information represented by coordinate values ​​along two axes, such as the X and Y axes. Object position information is determined by the representative position of a shape represented by a grid or voxel.

[0217] The shape information may also include information about the material of the surface.

[0218] The attenuation rate can be expressed as a real number below 1 or above 0, or as a negative decibel value. Since the volume is not amplified by reflection in real space, the attenuation rate is set to a negative decibel value. However, to create a sense of horror in an unreal space, for example, an attenuation rate of 1 or above, i.e., a positive decibel value, can be deliberately set. Furthermore, the attenuation rate can be set to a different value for each of the multiple frequency bands, or a value can be set independently for each frequency band. Furthermore, when setting the attenuation rate for each type of surface material, the corresponding attenuation rate value can be used based on information related to the surface material.

[0219] Furthermore, the information assigned to both the sound source object and the non-sound emitting object may include information indicating whether the object is a living being or whether the object is a moving object. If the object is a moving object, its position information may also change over time, and the changed position information or the amount of change is transmitted to the rendering units A0203 and A0213.

[0220] The information related to the sound source object includes the object sound data and the information required to radiate the object sound data into the sound space, in addition to the information given to the sound source object and the non-sounding object mentioned above. The object sound data is data that represents the sound perceived by the listener, such as information related to the frequency and strength of the sound. The object sound data is typically a PCM signal, but it can also be data compressed using an encoding method such as MP3. In this case, it is necessary to at least Figure 34 Since the signal is decoded before the generation unit 907 described later in the figure, a decoding unit (not shown) may be included in the rendering unit A0203 and A0213. Alternatively, the signal may be decoded by the sound data decoder A0202.

[0221] At least one target sound data item may be set for each sound source object, or multiple target sound data items may be set. In addition, identification information for identifying each target sound data item may be assigned, and the identification information of the target sound data item may be stored as metadata as information related to the sound source object.

[0222] The information required to radiate the object sound data into the sound space may include, for example, information on a reference volume serving as a reference when reproducing the object sound data, information related to the position of the sound source object, information related to the orientation of the sound source object, and information related to the directionality of the sound emitted by the sound source object.

[0223] The reference volume information may be, for example, the effective value of the amplitude of the target sound data at the sound source position when the target sound data is emitted into the sound space. It may also be expressed as a decibel (dB) value with a floating decimal point. For example, if the reference volume is 0dB, the reference volume information may indicate that the sound should be emitted into the sound space at the original volume from the position indicated by the position information, without increasing or decreasing the volume of the signal level indicated by the target sound data. If the reference volume is -6dB, the reference volume information may indicate that the sound should be emitted into the sound space from the position indicated by the position information, with the volume of the signal level indicated by the target sound data reduced to approximately half. The reference volume information may be assigned to a single target sound data item or to multiple target sound data items simultaneously.

[0224] The volume information included in the information required to radiate the target sound data into the sound space may include, for example, information indicating the temporal changes in the volume of the sound source. For example, when the sound space is a virtual conference room and the sound source is a speaker, the volume changes intermittently over a short period of time. To put it more simply, it can be said that the sound and silent parts are produced alternately. Furthermore, when the sound space is a concert hall and the sound source is a performer, the volume is maintained for a certain length of time. Furthermore, when the sound space is a battlefield and the sound source is an explosive, the volume of the explosion increases only for a moment, then becomes silent and continues. Thus, the information on the volume of the sound source includes not only information on the volume of the sound, but also information on the changes in the volume of the sound. Such information can also be used as information indicating the nature of the target sound data.

[0225] Here, the information on the changes in sound volume can also be data representing frequency characteristics in a time series. The information on the changes in sound volume can also be data representing the duration of sound intervals. The information on the changes in sound volume can also be data representing the duration of sound intervals and the duration of silent intervals. The information on the changes in sound volume can also be data that lists, in a time series, multiple sets of data representing durations during which the amplitude of a sound signal can be considered constant (considered approximately constant) and the amplitude values ​​of the signal during those durations. The information on the changes in sound volume can also be data that lists, in a time series, multiple sets of data representing durations during which the frequency characteristics of a sound signal can be considered constant and the frequency characteristics during those durations. The information on the changes in sound volume can also be data representing, for example, the general shape of a spectrogram. Furthermore, the reference volume can be a volume that serves as a reference for the frequency characteristics. The information indicating the reference volume and the information indicating the properties of the target sound data are used not only to calculate the volume of the direct sound or reflected sound perceived by the listener, but also to select whether to perceive the sound to the listener.

[0226] Information related to orientation is typically expressed as yaw, pitch, and roll. Alternatively, roll rotation can be omitted and represented as azimuth (yaw) and elevation (pitch). Orientation information can also change over time and, if so, is transmitted to rendering units A0203 and A0213.

[0227] Information related to the listener is information related to the position information and orientation of the listener in the sound space. The position information is represented by the position of the X-axis, Y-axis, and Z-axis in the Euclidean space, but it does not necessarily have to be three-dimensional information and can also be two-dimensional information. Information related to orientation is typically represented by yaw, pitch, and roll. Alternatively, the rotation of roll can be omitted and represented by azimuth (yaw) and elevation (pitch). The position information and orientation information can also change over time and are transmitted to the rendering units A0203 and A0213 when changes occur.

[0228] Sensor information is information including the amount of rotation or displacement detected by the sensor worn by the listener and the position and orientation of the listener. The sensor information is transmitted to the rendering units A0203 and A0213, and the rendering units A0203 and A0213 update the information on the position and orientation of the listener based on the sensor information. Sensor information can also use, for example, location information obtained by a portable terminal using GPS, a camera, or LiDAR (Laser Imaging Detection and Ranging) to estimate its own position. In addition, information obtained from the outside via a communication module other than the sensor can also be detected as sensor information. Information indicating the temperature of the sound signal processing device 100 and information indicating the remaining battery level can also be obtained from the sensor as sensor information. Information indicating the computing resources (CPU capacity, memory resources, PC performance) of the sound signal processing device 100 or the sound prompt device A0002 can also be obtained in real time as sensor information.

[0229] In this embodiment, the acquisition unit 110 acquires audio data representing the waveform of the reference sound, the first position information, the target sound data, and the shape information from the storage unit 140. However, this is not limiting. The acquisition unit 110 may also acquire the audio data from a device other than the audio signal processing device 100 (e.g., a server device 500 such as a cloud server). Furthermore, the acquisition unit 110 acquires the second position information from the earphones 200 (more specifically, the head sensor unit 201), but this is not limiting.

[0230] Next, the first position information will be described.

[0231] As described above, the object in the virtual space is included in the content (image) displayed by the display unit 300 , and in the present embodiment, is, for example, the fan FN.

[0232] The first position information indicates the location of the fan FN in the virtual space at a certain point in time. Furthermore, the fan FN may be moved in the virtual space, for example, by a user moving the fan FN in their hand. Therefore, the acquisition unit 110 continuously acquires the first position information. For example, the acquisition unit 110 acquires the first position information each time the space information is updated by the space information management units A0201 and A0211.

[0233] Furthermore, the following describes sound data including sound data representing a waveform of a reference sound (aerodynamic sound) and target sound data associated with an object.

[0234] The sound data including the target sound data and the aerodynamic sound data described in this specification may be a sound signal such as PCM (Pulse Code Modulation) data, but is not limited thereto and may be any information indicating the properties of the sound.

[0235] For example, if a sound signal is a noise signal with a volume of X decibels, the sound data associated with the sound signal may be the PCM data representing the sound signal itself, or data consisting of information indicating that the sound signal component is a noise signal and information indicating that the volume is X decibels. For another example, if a sound signal is a noise signal with a frequency component having a predetermined Peak / Dip ratio, the sound data associated with the sound signal may be the PCM data representing the sound signal itself, or data consisting of information indicating that the sound signal component is a noise signal and information indicating the Peak / Dip ratio of the frequency component.

[0236] In this specification, a sound signal based on sound data refers to PCM data representing the sound data.

[0237] Furthermore, as described above, aerodynamic sound data, which is sound data representing the waveform of the reference sound, is pre-stored in the storage unit 140. Aerodynamic sound is sound produced by the impact of wind W on an object, in this case, the impact of wind W on the ear of the listener L. Aerodynamic sound data is data collected by collecting the sound produced by the impact of wind W on a human ear or an object (model) having a shape that mimics a human ear. In this embodiment, aerodynamic sound data is data collected by collecting the sound produced by wind reaching an object (model) that mimics a human ear. The aerodynamic sound data is collected using a dummy head-shaped microphone or the like as a model that mimics a human ear.

[0238] Next, shape information will be described.

[0239] Shape information represents the shape of an object in virtual space. Shape information represents the object's shape, more specifically, the three-dimensional shape of the object as a rigid body. Examples of object shapes include spheres, cuboids, cubes, polyhedrons, cones, pyramids, cylinders, prisms, and combinations thereof. Shape information can also be represented as mesh data, voxels, a three-dimensional point group, or a collection of multiple faces consisting of vertices with three-dimensional coordinates.

[0240] Furthermore, the first position information includes object identification information for identifying the object. Furthermore, the object sound data also includes object identification information, and the shape information also includes object identification information.

[0241] Therefore, even if the acquisition unit 110 separately acquires the first position information, the target sound data, and the shape information, it can identify the objects represented by the first position information, the target sound data, and the shape information by referencing the object identification information contained in each of the first position information, the target sound data, and the shape information. For example, it can be easily identified that the objects represented by the first position information, the target sound data, and the shape information are the same fan FN. In other words, by referencing the three pieces of object identification information, the first position information, the target sound data, and the shape information acquired by the acquisition unit 110 clearly indicate that the first position information, the target sound data, and the shape information are information related to the fan FN. Thus, the first position information, the target sound data, and the shape information are associated as information representing the fan FN.

[0242] Next, the second position information will be described.

[0243] Listener L can move within the virtual space. The second positional information indicates the location of listener L within the virtual space at a given point in time. Furthermore, because listener L can move within the virtual space, acquisition unit 110 continuously acquires the second positional information. For example, acquisition unit 110 acquires the second positional information each time spatial information is updated by spatial information management units A0201 and A0211.

[0244] Furthermore, the aforementioned audio data representing the waveform of the reference sound, the first position information, the target sound data, the shape information, and the second position information may also be included in metadata, control information, or header information included in the input signal. If the audio data including the target sound data and aerodynamic sound data is an audio signal (PCM data), information identifying the audio signal may also be included in the metadata, control information, or header information, or the audio signal may be included in a source other than the metadata, control information, or header information. In other words, the audio signal processing device 100 (more specifically, the acquisition unit 110) may also acquire metadata, control information, or header information included in the input signal and perform audio processing based on the metadata, control information, or header information. Furthermore, the audio signal processing device 100 (more specifically, the acquisition unit 110) only needs to acquire the aforementioned audio data representing the waveform of the reference sound, the first position information, the target sound data, the shape information, and the second position information; the acquisition source is not limited to the input signal. The audio data and metadata including the target sound data and aerodynamic sound data may be stored in a single input signal or in multiple input signals.

[0245] In addition, a sound signal other than the sound data including the target sound data and the aerodynamic sound data may be stored in the input signal as audio content information. The audio content information may also be subjected to encoding processing such as MPEG-H 3D Audio (ISO / IEC23008-3) (hereinafter referred to as MPEG-H 3D Audio). In addition, the technology used in the encoding processing is not limited to MPEG-H 3D Audio, and other well-known technologies may also be used. In addition, the above-mentioned sound data representing the waveform of the reference sound, the first position information, the target sound data, the shape information, and the second position information may also be used as the encoding processing object.

[0246] Specifically, the audio signal processing device 100 obtains the audio signal and metadata included in the encoded bitstream. The audio signal processing device 100 obtains and decodes the audio content information. In this embodiment, the audio signal processing device 100 functions as a decoder (e.g., decoders A0200 and A0210) included in a decoding device (e.g., decoding devices A0110 and A0130), and more specifically, as a rendering unit A0203 and A0213 included in the decoder. The term "audio content information" in this disclosure is to be interpreted as meaning information that can be replaced with the audio signal itself, audio data representing the waveform of the reference sound, first position information, target sound data, shape information, and second position information, depending on the technical content.

[0247] The acquisition unit 110 outputs the acquired audio data indicating the waveform of the reference sound, the first position information, the target sound data, the shape information, and the second position information to the processing unit 120 and the output unit 130 .

[0248] Based on simulation information simulating the fluctuations of a natural phenomenon, the processing unit 120 processes the sound data so as to change at least one of the frequency components, phase, and amplitude of the waveform represented by the sound data representing the waveform of the reference sound. In this embodiment, since the reference sound is aerodynamic sound generated by wind W, the natural phenomenon in the simulation information is the blowing of wind W. The fluctuation of the natural phenomenon is the fluctuation of wind W, more specifically, the fluctuation of the wind speed of wind W. Alternatively, the fluctuation of the natural phenomenon may be a change in the direction (wind direction) of wind W.

[0249] In real space, the fluctuations of natural phenomena include fluctuations (e.g., 1 / f fluctuations). Therefore, simulation information simulates the fluctuations of natural phenomena that include fluctuations. In this embodiment, the simulation information simulates the fluctuations of the wind speed of wind W, more specifically, information that represents the fluctuations included in the fluctuations of the wind speed of wind W.

[0250] More specifically, the simulation information is a smooth function that simulates the fluctuation of the wind speed. Here, the processing unit 120 determines the smooth function that simulates the fluctuation of the wind speed as the simulation information.

[0251] A smooth function is a function that is differentiable and continuous. In other words, a smooth function is a function that does not have sharp points.

[0252] Figure 13 is a diagram showing an example of a smoothing function related to this embodiment. Figure 13 As shown, the smoothing function is a sine curve as an example, but is not limited to this, and may be a cosine curve or the like.

[0253] Processing unit 120 processes the sound data based on the value represented by the smoothing function determined by processing unit 120 to change at least one of the frequency component, phase, and amplitude value of the waveform. For example, processing unit 120 processes the sound data to shift the frequency component of the waveform to a frequency proportional to the value represented by the smoothing function that simulates the fluctuation of wind speed.

[0254] The value represented by the smooth function is Figure 13 The value shown on the vertical axis is information indicating the ratio of the wind speed of the aerodynamic sound serving as the reference sound to the wind speed of the aerodynamic sound represented by the sound data processed by processing unit 120. In other words, the value represented by the smoothing function is the ratio of the wind speed of the aerodynamic sound before processing to the wind speed of the aerodynamic sound after processing.

[0255] The processing unit 120 processes the audio data and outputs the processed audio data to the output unit 130 .

[0256] The output unit 130 outputs the sound data processed by the processing unit 120. Here, the output unit 130 outputs the processed aerodynamic sound data to the earphones 200. This allows the earphones 200 to reproduce the aerodynamic sound represented by the output aerodynamic sound data. In other words, the listener L can hear the aerodynamic sound.

[0257] The storage unit 140 is a storage device that stores computer programs and the like executed by the acquisition unit 110 , the processing unit 120 , and the output unit 130 , as well as pneumatic sound data.

[0258] The receiving unit 150 receives an operation from a user (for example, a creator of content executed in a virtual space) of the sound signal processing device 100. Specifically, the receiving unit 150 is implemented by hardware buttons, but may also be implemented by a touch panel or the like.

[0259] Here, we will further explain the shape information related to this embodiment. Shape information is used to generate the image of an object in virtual space and represents the shape of the object (fan FN). In other words, shape information is also used to generate the content (image) displayed on display unit 300.

[0260] The acquisition unit 110 also outputs the acquired shape information to the display unit 300. The display unit 300 acquires the shape information output by the acquisition unit 110. The display unit 300 also acquires attribute information representing attributes (such as color) other than the shape of the object (fan FN) in virtual space. The display unit 300 may directly acquire the attribute information from a device other than the sound signal processing device 100 (the server device 500), or may acquire it from the sound signal processing device 100. The display unit 300 generates content (video) based on the acquired shape information and attribute information, and displays it.

[0261] Hereinafter, operation examples 1 and 2 of the sound signal processing method performed by the sound signal processing device 100 will be described.

[0262] [Action Example 1]

[0263] Figure 14 This is a flowchart of an operation example 1 of the sound signal processing device 100 according to the present embodiment.

[0264] like Figure 14 As shown, first, the accepting unit 150 accepts an operation indicating that the simulation information is a smooth function that simulates the fluctuation of wind speed (S10). The accepting unit 150 accepts this operation from, for example, a user of the sound signal processing device 100.

[0265] Next, the acquisition unit 110 acquires sound data representing the waveform of the reference sound (S20). In this operation example, the reference sound is aerodynamic sound generated by wind, and the sound data representing the waveform of the reference sound is aerodynamic sound data. This step S20 corresponds to the acquisition step.

[0266] The processing unit 120 determines a smoothing function that simulates the change of wind speed as simulation information that simulates the change of natural phenomena (S30). The processing unit 120 can determine the simulation information according to the operation received in step S10. In this operation example, the processing unit 120 determines Figure 13 The smoothing function shown serves as the simulation information.

[0267] Furthermore, the processing unit 120 processes the sound data (aerodynamic sound data) based on the value (ratio) represented by the smoothing function determined by the processing unit 120 so as to change at least one of the frequency component, phase, and amplitude value of the waveform ( S40 ).

[0268] Note that steps S30 and S40 correspond to processing steps.

[0269] The processing unit 120 outputs the processed sound data (aerodynamic sound data) to the output unit 130 .

[0270] The output unit 130 outputs the sound data (aerodynamic sound data) processed by the processing unit 120 to the earphone 200 (S50). Note that this step S50 corresponds to an output step.

[0271] This allows the listener L to listen to the aerodynamic sound output from the earphones 200 .

[0272] Here, the processing in steps S30 and S40 performed by the processing unit 120 will be described in more detail.

[0273] Figure 15 This is a diagram for explaining the processing performed by the processing unit 120 according to this embodiment.

[0274] Figure 15 (a) means Figure 12 The sound data shown in (the aerodynamic sound data D1 before processing) and Figure 13 The graph of the smooth function shown in . Figure 15 As shown in (a) of FIG. 1 , the aerodynamic sound data D1 before processing and the smoothing function correspond to each other along the horizontal axis, which is the time axis.

[0275] Figure 15 (b) is used to illustrate Figure 15 (a) is a diagram of the processing in the area surrounded by the single-dot dashed rectangle. Figure 15 In (b), the aerodynamic sound data D1 before processing, the smoothing function, and the aerodynamic sound data D11 after processing are magnified and shown.

[0276] The aerodynamic sound data D1 before processing is Figure 15 The multiple black dots in (b) represent the Figure 15 The aerodynamic sound data D1 before processing is shown in (a). In addition, each of the plurality of black dots can be regarded as a sample point of the aerodynamic sound data D1 before processing.

[0277] The processing unit 120 first performs a first process. The first process will be described below.

[0278] The processing unit 120 determines an interpolation function for interpolating between one black dot and another black dot adjacent to the one black dot. The interpolation function is, for example, a spline function, but is not limited thereto and may be a well-known function. In addition, the processing unit 120 may also perform linear interpolation (straight line interpolation) between one black dot and another black dot adjacent to the one black dot. In this case, the load of the computational processing is reduced. Figure 15 As shown in (b), in the first process, all the values ​​between two adjacent black dots are interpolated.

[0279] Therefore, if Figure 15As shown in (b), one black dot and another black dot adjacent to it are interpolated and displayed as a line. In addition, the interval between the plurality of black dots before the processing is defined as "1".

[0280] Next, the processing unit 120 performs the second process. Hereinafter, the second process will be described.

[0281] In the second process, the processing unit 120 reads the value of one black dot of the pneumatic sound data D1 before processing at time t, and defines the read value as the pneumatic sound data D11 after processing at the time t. Figure 15 (c) is represented by multiple white dots (hollow dots).

[0282] Next, the processing unit 120 reads the value of the smoothing function for each unit time. For example, the processing unit 120 reads "0.5," "0.5," "0.4999," and "0.4998," etc. as the value of the smoothing function.

[0283] The processing unit 120 determines the value of the smoothing function read at time t as a stride, and reads the value of the interpolation function at a position that is a distance corresponding to the stride from a black dot of the pre-processed aerodynamic sound data D1 at time t.

[0284] The processing unit 120 then determines the value of the read interpolation function as the value of the processed aerodynamic sound data D11. At this time, the processing unit 120 determines the spacing of the processed aerodynamic sound data D11 (multiple white dots) so that the spacing is the same as the spacing of the pre-processed aerodynamic sound data D1 (multiple black dots), that is, "1." In this way, the second process is performed.

[0285] A specific example of the second process will be described focusing on time t1.

[0286] The processing unit 120 reads the value of black point B1 in the pre-processed aerodynamic sound data D1 at time t1 and determines the read value as the value of white point B11 in the post-processed aerodynamic sound data D11 at time t1. In other words, the processing unit 120 uses the read value of black point B1 as is as the value of white point B11.

[0287] Furthermore, the processing unit 120 reads the value of the smoothing function at time t1, which is 0.5, and determines it as the step size. The aerodynamic sound data D1 before processing at time t1 is represented by the black dot B1. The processing unit 120 reads the value of the interpolation function at a position that is a distance corresponding to 0.5 from the black dot B1 of the aerodynamic sound data D1 before processing. This position is Figure 15 In (b) it is represented by position P1.

[0288] Next, the processing unit 120 determines the read interpolation function value (the value indicated by position P1) as the value of the processed aerodynamic sound data D11. The processing unit 120 determines the intervals of the processed aerodynamic sound data D11 (multiple white dots) so that the intervals of the processed aerodynamic sound data D11 (multiple white dots) are equal to "1," which is the same value as the intervals of the pre-processed aerodynamic sound data D1 (multiple black dots).

[0289] Through the first and second processes, the processed aerodynamic sound data D11 becomes a laterally extended shape of the pre-processed aerodynamic sound data D1. Therefore, the processed aerodynamic sound data D11 becomes sound data whose frequency components are shifted to a lower frequency range compared to the pre-processed aerodynamic sound data D1.

[0290] Figure 16 This is another diagram for explaining the processing performed by the processing unit 120 according to this embodiment.

[0291] Figure 16 (a) and Figure 15 (a) is similar, indicating that Figure 12 The sound data (aerodynamic sound data) shown in Figure 13 Plot of the smoothing function shown.

[0292] Figure 16 (b) and (c) are used to illustrate Figure 16 (a) is a diagram of the processing in the area surrounded by the single-dot dashed rectangle. Figure 15 In each of (b) and (c), the aerodynamic sound data D1 before processing, the smoothing function, and the aerodynamic sound data D11 after processing are enlarged and shown.

[0293] exist Figure 16 The aerodynamic sound data D1 before processing shown in (b) and (c) is also processed using Figure 15 The same processing as that described in (b) is performed. That is, the first processing and the second processing are performed.

[0294] exist Figure 16 In (b), the processing unit 120 reads "1," "1," "1.0001," and "1.0002," etc., as the smoothing function value. Since the read smoothing function value is approximately 1, the processed aerodynamic sound data D11 has the same shape as the pre-processed aerodynamic sound data D1. Therefore, the processed aerodynamic sound data D11 is sound data with substantially no shift in frequency components compared to the pre-processed aerodynamic sound data D1.

[0295] exist Figure 16In (c), the processing unit 120 reads "1.5," "1.5," "1.4999," and "1.4998," among other values, as the smoothing function values. Since the smoothing function values ​​read are approximately 1.5, the processed aerodynamic sound data D11 becomes a laterally contracted version of the pre-processed aerodynamic sound data D1. Consequently, the processed aerodynamic sound data D11 becomes sound data with frequency components shifted toward higher frequencies compared to the pre-processed aerodynamic sound data D1.

[0296] As described above, the simulation information is information simulating the fluctuation of natural phenomena including fluctuations, more specifically, information expressing fluctuations due to changes in the wind speed of the wind W, and is information represented by a smoothing function in this operation example.

[0297] In this example operation, sound data (aerodynamic sound data) representing the waveform of a reference sound is processed based on simulation information that simulates the fluctuations of natural phenomena involving fluctuations, thereby varying the frequency components of the waveform. Consequently, the frequency components of the processed aerodynamic sound data fluctuate, and the frequency components of the aerodynamic sound represented by the processed aerodynamic sound data also fluctuate. Consequently, the listener L can hear the aerodynamic sound with such fluctuating frequency components, thereby minimizing any sense of discomfort and providing a sense of presence.

[0298] Furthermore, in step S40 of operation example 1, it is preferable to perform the following processing.

[0299] As described above, in step S40, the step size is preferably determined as follows: Here, the sampling frequency of the aerodynamic sound data before being processed by the processing unit 120 is denoted as Fsc, and the sampling frequency of the aerodynamic sound data output by the output unit 130 is denoted as Fso, and Fsc and Fso are assumed to be different values.

[0300] In this case, the step size preferably satisfies the following equation.

[0301] Smoothing function value × (Fsc / Fso)

[0302] The following describes the effect of the step size satisfying the above formula.

[0303] For example, if Fso is 48 kHz, it is preferable to downsample Fsc from 48 kHz to 16 kHz. This allows, for example, the storage unit 140 to store aerodynamic sound data of the same length, reducing the memory size to one-third. Furthermore, for example, when the same memory size is used, the time length of the output aerodynamic sound data is tripled, thus reducing the awkwardness of the transitions between aerodynamic sound data.

[0304] Furthermore, a case where the aliasing distortion can be reduced will be described. Figure 17 is a diagram showing the audio data related to this embodiment. More specifically, Figure 17 (a) and (b) are the aerodynamic sound data before processing (e.g. Figure 15 The frequency characteristics of the aerodynamic sound data D1 before processing are shown in FIG. Figure 17 In (a), the horizontal axis is a logarithmic axis. Figure 17 In (b), the horizontal axis is a linear axis. In addition, Figure 17 (c) means Figure 17 (b) shows the frequency characteristics of the aerodynamic sound data after the frequency component is shifted to the high frequency side. Figure 17 The frequency component in (c) is Figure 17 The result of shifting the frequency component in (b) to 2 times the frequency. For example, Figure 17 The frequency component of 2000Hz in (b) is shifted to the high frequency side to become Figure 17 (c) The frequency component of 4000kHz.

[0305] exist Figure 17 In (a) and (b), the solid line represents the frequency characteristics when the sampling frequency of the aerodynamic sound data before processing is 16 kHz, and the dashed-dotted line represents the frequency characteristics when the sampling frequency of the aerodynamic sound data before processing is 48 kHz. The dashed-dotted line overlaps with the solid line in the low-frequency domain and is therefore not shown.

[0306] like Figure 17 As shown, aerodynamic sound data often exhibits a characteristic structure in the low-frequency domain, and its components decrease monotonically in the high-frequency domain.

[0307] exist Figure 17 In (c), the solid line represents the frequency characteristics when the sampling frequency of the shifted aerodynamic sound data is 16 kHz, and the dashed-dotted line represents the frequency characteristics when the sampling frequency of the shifted aerodynamic sound data is 48 kHz. The dashed-dotted line overlaps with the solid line in the low-frequency domain and is therefore not shown.

[0308] When the sampling frequency of the aerodynamic sound data indicated by the single-dot chain line is 48 kHz, Figure 17 In (b), there are frequency components in the frequency domain above 12kHz. Figure 17 In (c), folding distortion appears as indicated by the dotted line.

[0309] When the sampling frequency of the aerodynamic sound data indicated by the solid line is 16 kHz, Figure 17 In (b), there is no frequency component in the frequency domain above 12kHz, so Figure 17 There is no aliasing distortion in (c).

[0310] In this way, the occurrence of aliasing distortion caused by frequency shift can be suppressed.

[0311] Furthermore, there is an effect that the aforementioned reduction in memory size and the increase in computational resources required to suppress the occurrence of aliasing distortion are hardly achieved.

[0312] The above is equivalent to the effect brought about by the step size satisfying the above formula.

[0313] In Operation Example 1 of this embodiment, the aerodynamic sound data is pre-stored in the storage unit 140, but the present invention is not limited thereto. For example, the processing unit 120 may generate the aerodynamic sound data. For example, the processing unit 120 may generate the aerodynamic sound data by acquiring a noise signal and processing the acquired noise signal using a plurality of frequency band enhancement filters.

[0314] [Action Example 2]

[0315] As described above, in Operation Example 1, the sound data (pneumatic sound data) is processed to change the frequency component of the waveform, but the present invention is not limited thereto. In Operation Example 2, the sound data (pneumatic sound data) is processed to change the amplitude value of the waveform.

[0316] That is, in Action Example 2, steps S10 to S30 are performed similarly to Action Example 1. Next, in step S40, the processing unit 120 processes the sound data (aerodynamic sound data) based on the value (ratio) represented by the smoothing function determined by the processing unit 120 to change the amplitude value of the waveform.

[0317] The amplitude value of a waveform refers to the volume of the aerodynamic sound represented by the aerodynamic sound data represented by the waveform. Aerodynamic sound has the following relationship with the wind speed W that generates the aerodynamic sound. The volume of aerodynamic sound is proportional to the wind speed W raised to the power of α. Therefore, the processing unit 120 processes the sound data so that the amplitude value of the waveform changes proportionally to the value represented by the determined smoothing function raised to the power of α. The value of α varies depending on the type of aerodynamic sound.

[0318] For example, there is aerodynamic sound produced by a stick-shaped object cutting through the wind. This aerodynamic sound is produced by the swinging of a baseball bat, etc. The volume of this type of aerodynamic sound is proportional to the sixth power of the wind speed (see Non-Patent Document 1).

[0319] Another example is aerodynamic sound, which is produced when wind blows through the gap between an object and another object. This aerodynamic sound is called cavity sound. The volume of this type of aerodynamic sound is proportional to the fourth power of the wind speed (see Non-Patent Document 1).

[0320] Here, let the value represented by the smoothing function that simulates the change in wind speed be R. In the case of any of the above-mentioned types of aerodynamic sound, the volume of the aerodynamic sound is amplified or attenuated by a value corresponding to R^α. That is, it is amplified when R is greater than 1 and attenuated when R is less than 1. It should be noted here that when the volume of the aerodynamic sound is proportional to the α-th power of the wind speed, the volume of the aerodynamic sound changes very rapidly. Use Figure 18 to explain this rapid change.

[0321] Figure 18 is a graph showing R, which represents the value as a smoothing function in this embodiment, and the amplification factor and attenuation factor of the volume of the aerodynamic sound. Figure 18 The double-dashed line in

[0322] shows the relationship between R and the volume of the aerodynamic sound when α is 6. In addition, near R = 1, the double-dashed line overlaps with the solid line. Figure 18 As shown by the double-dashed line in

[0323] when R = 2.0, the amplification factor exceeds 30 dB, and when R = 0.5, the attenuation factor is lower than -30 dB. In order to faithfully reproduce such a rapid change, an expensive reproduction device with a very wide dynamic range is required, but such a reproduction device is excessive for the audio presentation in the virtual space. Figure 18 Figure 18 Figure 18

[0324]

[0325]

[0326]

[0327] G = R^α

[0328]

[0329] Figure 18 In addition, for the single-dashed line in the range of R < (1 / r) and r < R, the amplification factor (attenuation factor) G satisfies the following formula.

[0330] Figure 18

[0331]

[0332] Figure 18

[0333]

[0334]

[0335]

[0336]

[0337] Figure 19 Figure 19 Figure 19 Figure 19

[0338] Figure 13

[0339] Figure 19

[0340]

[0341] Figure 13

[0342]

[0343] Figure 19

[0344]

[0345]

[0346]

[0347]

[0348]

[0349]

[0350] Figure 20 Figure 20

[0351]

[0352]

[0353]

[0354]

[0355]

[0356]

[0357]

[0358]

[0359] Figure 21

[0360] Figure 21 Figure 21 Figure 13 Figure 21 Figure 21 Figure 21 Figure 21

[0361]

[0362]

[0363]

[0364] Figure 22 Figure 22 Figure 22

[0365]

[0366]

[0367]

[0368]

[0369]

[0370] Figure 23

[0371]

[0372]

[0373] Figure 14

[0374] Figure 24

[0375]

[0376] Figure 15

[0377]

[0378]

[0379]

[0380]

[0381] Figure 25 Figure 26 Figure 25

[0382]

[0383]

[0384] Figure 27

[0385] Figure 27

[0386]

[0387]

[0388]

[0389]

[0390] Figure 28 Figure 28

[0391] [[ID=;]] [[ID=;]]

[0392] [[ID=;]] Figure 29 [[ID=;]] [[ID=;]]

[0393] [[IDIn addition, by setting b to a value smaller than α, near R = 1.0, a trend close to the magnification (attenuation rate) G = R^α, that is, the correct trend, is achieved, and it is monotonically magnified (monotonically attenuated) outside the vicinity of R = 1.0, and sharp changes can be avoided.

[0329] Figure 18 The single-dot chain line in satisfies the conditions of r = 1.3 and b = 2.0. However, in this single-dot chain line, the trend of magnification and attenuation changes discontinuously at R = r and R = 1 / r. Therefore, a sense of incongruity may sometimes occur near R = r and near R = 1 / r.

[0330] Therefore, instead of setting b as a constant, the value of b can be made the same as the value of α at the position of R = r and gradually become smaller than the value of α as R increases. Figure 18 In Figure 18 , for the solid line in the interval where R < (1 / r) and r < R, the magnification (attenuation rate) G satisfies the following formula.

[0331] G = (r)^α × (R / r)^b, where b = α^(r / R)

[0332] By increasing and decreasing the volume along with the Figure 18 solid line in Figure 18 , the volume can change sensitively according to the subtle changes in the wind speed (the minute changes near R = 1), and sharp changes in the volume due to the increase and decrease of R can be avoided.

[0333] In addition, the value of α can be arbitrarily set by the user of the audio signal processing device 100 (for example, the producer of the content executed in the virtual space). That is, the reception unit 150 can receive an operation indicating the value of α from the producer, and the processing unit 120 can determine the value indicated by the received operation as the value of α. By setting the value of α to a value that is significantly different from the academically correct value, such as 0.7, 1.0, 1.5, or 2.0, but is used to present the increase and decrease of the volume of "realistic" pneumatic sound in the virtual space, sharp changes can be avoided. In addition, the values of r and b can be determined in the same way.

[0334] In addition, in operation example 1, the pneumatic sound data was processed to change the frequency components, and in operation example 2, the pneumatic sound data was processed to change the amplitude values, but it is not limited to this. For example, the pneumatic sound data can also be processed to change the phase of the waveform. In this case, the processing unit 120 processes the pneumatic sound data so that the phase of the waveform changes according to the value represented by the determined smoothing function.

[0335] In addition, at least one of the frequency components, phase, and amplitude number of the waveform only needs to change. For example, two of the frequency components, phase, and amplitude number of the waveform can change, or all of the frequency components, phase, and amplitude number of the waveform can change.

[0336] In Operation Examples 1 and 2, the processing unit 120 may divide the sound data (aerodynamic sound data) representing the waveform of the reference sound acquired by the acquisition unit 110 into processing frames F of a predetermined time and process the sound data for each of the divided processing frames F.

[0337] Figure 19 : is a diagram showing the divided aerodynamic sound data according to this embodiment. Figure 19 In the embodiment, the pneumatic sound data is divided into a plurality of processing frames F. In addition, the prescribed time Ts of each of the plurality of processing frames F may be the same or may be the same as Figure 19 are different from each other. Figure 19 , processing frames F1 to F6 are shown as examples of the processing frames F, and predetermined times Ts1 to Ts6 are shown as examples of the predetermined time Ts. The predetermined times Ts1 to Ts6 are different from each other.

[0338] In addition, in the operation examples 1 and 2, the simulation information is used Figure 13 The smoothing function shown is shown, but different smoothing functions may be used.

[0339] For example, in step S30 of operation examples 1 and 2, the processing unit 120 determines a smoothing function that simulates the change in wind speed as simulation information that simulates the change in natural phenomena. In this case, the processing unit 120 may determine the smoothing function so that the parameters that determine the smoothing function change irregularly. Furthermore, the processing unit 120 determines the parameters that determine the smoothing function for each processed frame F after division. That is, for example, the processing unit 120 determines to determine the parameters that determine the smoothing function. Figure 19 Parameters of the smoothing function corresponding to processing frame F1 are shown. Similarly, processing unit 120 determines parameters of the smoothing function corresponding to processing frame F2, parameters of the smoothing function corresponding to processing frame F3, parameters of the smoothing function corresponding to processing frame F4, parameters of the smoothing function corresponding to processing frame F5, and parameters of the smoothing function corresponding to processing frame F6.

[0340] Furthermore, the processing unit 120 determines a smoothing function for each of the divided processing frames F so that the value of the smoothing function is 1.0 at the first and last times of the processing frame F. For example, in the smoothing function corresponding to the processing frame F2 at the predetermined time Ts2, the value represented by the smoothing function is 1.0 at time t2 and time t3.

[0341] If you set Figure 13 The smoothing function shown is F(t), and F(t) is expressed by the following formula.

[0342] F(t)=H×{sin[2π×(t / T)^(x)]}^(y)+1.0(0.0≤t <T)

[0343] An example of a parameter for determining the smoothing function is the time from the first moment of processing frame F to the last moment of processing frame F, which is T in the above equation. Figure 19 In the smoothing function corresponding to the processing frame F2 shown, this parameter is the time from time t2 to time t3. In other words, when the smoothing function is a sine curve, this parameter is equivalent to one cycle.

[0344] Another example of a parameter that determines a smoothing function is a value related to the maximum value of the smoothing function, which is H in the above equation. As in this embodiment, when the smoothing function is a sine curve, another example of the parameter may be a value that determines the maximum value of the smoothing function.

[0345] Another example of a parameter that determines a smoothing function is a parameter that changes the position at which the smoothing function reaches its maximum value, which is x in the above equation.

[0346] Another example of a parameter that determines a smoothing function is a parameter that changes the sharpness of a change in the smoothing function, which is y in the above equation.

[0347] The processing unit 120 determines the parameters so that the parameters vary irregularly, thereby determining the smoothing function. For example, the processing unit 120 may determine the parameters based on random numbers.

[0348] For example, the processing unit 120 may also be provided with a random number sequence generator, which changes parameters according to its output sequence. Here, a true random number sequence originally has no regularity and no reproducibility. However, owing to being difficult to realize this on a computer, the sequence generated by the random number sequence generator may be a pseudorandom number sequence generated by a deterministic computational process. For example, the pseudorandom number sequence generated by the rand() function in the C language may also be used, or any known algorithm for generating pseudorandom numbers may be used in addition. In addition, a random number sequence of finite length, a pseudorandom number sequence of finite length, or a sequence of finite length made in order to present an irregular sense may be stored in the storage unit 140 and used repeatedly as a pseudorandom number sequence for a long time.

[0349] Alternatively, the accepting unit 150 may accept an operation indicating a parameter value from a user of the audio signal processing device 100 (e.g., a creator of content executed in a virtual space). The processing unit 120 may determine the value indicated by the operation accepted by the accepting unit 150 as the parameter.

[0350] Figure 20 This is a diagram showing another example of two smoothing functions according to the present embodiment. Figure 20 The smoothing functions represented by (a) and (b) are determined so that the parameters defining the smoothing functions vary irregularly.

[0351] Furthermore, parameters can be determined by simulating the properties of the wind speed of wind W. As mentioned above, the changes in the wind speed of wind W include fluctuations. That is, in real space, wind speed is not constant but fluctuates. For example, wind W may strike listener L at a first speed and then, afterward, at a second speed different from the first. This fluctuating property of wind speed can be simulated to determine parameters.

[0352] The maximum value of the smoothing function may not exceed 3, and the minimum value of the smoothing function may not be less than 0. That is, the value represented by the smoothing function may be greater than or equal to 0 and less than or equal to 3. Parameters may be determined so that the value represented by the smoothing function becomes as described above.

[0353] The reason why the maximum value of the smoothing function can be no more than 3 is as follows. In the real space, the change in the wind speed of the wind W includes fluctuations, and sometimes a wind W with a relatively strong wind speed (instantaneous wind speed) blows instantly. The wind speed is, for example, the average wind speed for 10 minutes, and the instantaneous wind speed is, for example, the average wind speed for 3 seconds. In such a case, it is known that the instantaneous wind speed is about 1.5 to 3 times the wind speed. The value represented by the smoothing function is the ratio of the wind speed of the aerodynamic sound as the reference sound to the wind speed of the aerodynamic sound represented by the processed sound data. By setting the maximum value of the smoothing function to be less than 3, the wind W with a relatively strong wind speed (instantaneous wind speed) that blows instantly, more specifically, the aerodynamic sound caused by the wind W, can be reproduced in the virtual space.

[0354] Furthermore, let the wind speed of wind W be Va, and let the instantaneous wind speed of wind W be Vp. In this case, processing unit 120 determines a smoothing function so that the maximum value of the smoothing function is Vp / Va. More specifically, processing unit 120 determines parameters of the smoothing function so that the maximum value of the smoothing function is Vp / Va. For example, receiving unit 150 receives an instruction specifying Va as the wind speed of wind W and Vp as the instantaneous wind speed of wind W. Processing unit 120, in accordance with the received instruction, determines parameters of the smoothing function so that the maximum value of the smoothing function is Vp / Va.

[0355] In this case, the display unit of the acoustic signal processing device 100 may display an image that associates a word indicating the strength of the wind W with the wind speed and instantaneous wind speed of the wind W indicated by the word. For example, in this image, if the word is "slightly strong wind," the image is associated with a wind speed of "10 or more and less than 15 (m / s)" and an instantaneous wind speed of "20 (m / s)." Furthermore, in this image, if the word is "strong wind," the image is associated with a wind speed of "15 or more and less than 20 (m / s)" and an instantaneous wind speed of "30 (m / s)."

[0356] The user of the sound signal processing device 100 (e.g., the creator of content executed in the virtual space) recognizes the image displayed on the display unit. Next, the receiving unit 150 receives an instruction from the user specifying a word representing the strength of wind W. The processing unit 120 determines the wind speed and instantaneous wind speed associated with the word specified in the received instruction as Va and Vp, and determines the parameters of the smoothing function so that the maximum value of the smoothing function is Vp / Va.

[0357] In this case, the wind W of a strong wind speed (instantaneous wind speed) blowing instantly, and more specifically, the aerodynamic sound caused by the wind W, can be reproduced in the virtual space.

[0358] Furthermore, the processing unit 120 divides the aerodynamic sound data into processing frames F of a predetermined time. The average value of the predetermined time may be 3 seconds. As described above, the instantaneous wind speed is, for example, the 3-second average wind speed. Therefore, by setting the average value of the predetermined time to 3 seconds, the predetermined time can be aligned with the time for measuring the instantaneous wind speed (i.e., 3 seconds), and the wind W of a strong wind speed (instantaneous wind speed) that blows instantly in the virtual space can be made closer to the wind blowing in the real space.

[0359] Here, use Figure 21 The smoothing function when the above four parameters are changed will be described in more detail.

[0360] Figure 21 : is a diagram showing an example in which the parameters for determining the smoothing function according to this embodiment are changed. Figure 21 In (a), it is shown that Figure 13 Same smoothing function. Figure 21 (b) is a graph showing a smooth function in which T in the above formula is changed. Figure 21 (c) is a graph showing a smooth function in which H of the above formula is changed. Figure 21 (d) is a graph showing a smooth function when x in the above equation is changed. Figure 21 (e) is a graph showing a smooth function in which y in the above equation is changed.

[0361] Furthermore, in the above-described operation examples 1 and 2, the processed aerodynamic sound data is output to the earphone 200 as one output channel. However, the present invention is not limited thereto. For example, the processed aerodynamic sound data may be output to a first output channel and a second output channel, respectively. The first output channel outputs the aerodynamic sound to one ear of the listener L, and the second output channel outputs the aerodynamic sound to the other ear of the listener L.

[0362] In this case, the processing unit 120 determines a first parameter and a second parameter that define a smoothing function. Based on the smoothing function determined by the first parameter determined by the processing unit 120, the processing unit 120 processes the acquired sound data (aerodynamic sound data) to change at least one of the frequency component, phase, and amplitude of the waveform. This processed aerodynamic sound data is referred to as aerodynamic sound data A. Based on the smoothing function determined by the second parameter determined by the processing unit 120, the processing unit 120 processes the acquired sound data (aerodynamic sound data) to change at least one of the frequency component, phase, and amplitude of the waveform. This processed aerodynamic sound data is referred to as aerodynamic sound data B.

[0363] The output unit 130 outputs the sound data (aerodynamic sound data A) processed based on the smoothing function determined by the determined first parameter to the first output channel. The output unit 130 outputs the sound data (aerodynamic sound data B) processed based on the smoothing function determined by the determined second parameter to the second output channel.

[0364] Figure 22 This is a diagram showing another example of two smoothing functions according to the present embodiment. Figure 22 (a) represents a smooth function determined by the first parameter, Figure 22 (b) represents a smoothing function determined by the second parameter. Here, the first output channel is a channel output to the right ear, and the second output channel is a channel output to the left ear.

[0365] This makes it possible to output different aerodynamic sound data for each output channel.

[0366] Furthermore, in this case, the first and second parameters can be determined by simulating the direction (wind direction) of wind W. As described above, changes in the direction (wind direction) of wind W include fluctuations. That is, in real space, wind direction is not constant but fluctuates. For example, wind W may blow from the right side of listener L and then from the front of listener L. In this way, the first and second parameters can be determined by simulating the fluctuating nature of wind direction.

[0367] (Variation of Embodiment 1)

[0368] Hereinafter, a modification of Embodiment 1 will be described. Hereinafter, the description will focus on the differences from Embodiment 1, and the description of the common points will be omitted or simplified.

[0369] [constitute]

[0370] First, the configuration of an acoustic signal processing device 100 a according to this modification will be described. Figure 23 : is a block diagram showing the functional configuration of an audio signal processing device 100 a according to this modification.

[0371] An acoustic signal processing device 100 a according to this modification has the same configuration as the acoustic signal processing device 100 according to the first embodiment, except that a processing unit 120 a is provided instead of the processing unit 120 .

[0372] The processing unit 120 a includes a first processing unit 121 and a second processing unit 122 .

[0373] The first processing unit 121 performs Figure 14 The second processing unit 122 performs the following processing based on the value represented by the smoothing function determined by the first processing unit 121.

[0374] Figure 24 100 is a block diagram showing the functional configuration of the second processing unit 122 according to this modification. The second processing unit 122 includes a sampling rate converter 1001 , a rearrangement unit 1002 , and a connection unit 1003 .

[0375] The sampling rate conversion unit 1001 acquires sound data (aerodynamic sound data) representing the waveform of the reference sound and a value represented by the smoothing function determined by the first processing unit 121 .

[0376] The sampling rate conversion unit 1001 converts the sampling rate of the aerodynamic sound data for each processing frame F based on the value represented by the obtained smoothing function. When the sampling rate of the aerodynamic sound data is Fs, the aerodynamic sound data before processing (for example Figure 15 The interval between sample points (sample interval) of the aerodynamic sound data D1 before processing is 1 / Fs seconds.

[0377] When the value indicated by the smoothing function is 0.5, the sampling rate conversion unit 1001 upsamples the pneumatic sound data to a sampling rate of 2·Fs, using a sampling interval of 0.5 times (1 / (2·Fs)). Furthermore, when the value indicated by the smoothing function is 2, the sampling rate conversion unit 1001 downsamples the pneumatic sound data to a sampling interval of 2 times (2 / Fs), using a sampling rate of Fs / 2. The sampling rate conversion unit outputs the pneumatic sound data after the sampling rate conversion to the reconfiguration unit 1002.

[0378] The reconfiguration unit 1002 restores the interval between the sample rate-converted aerodynamic sound data to Fs. This process causes the aerodynamic sound data to be reproduced quickly when the value represented by the smoothing function is greater than 1. On the other hand, when the value represented by the smoothing function is less than 1, the aerodynamic sound data is reproduced slowly. This shifts the frequency components of the aerodynamic sound data toward higher or lower frequencies, generating aerodynamic sound with a natural, fluctuating feel. The reconfiguration unit 1002 then outputs the aerodynamic sound data, after reconfiguring the sample point positions, to the connection unit 1003.

[0379] The connection unit 1003 performs processing to suppress discontinuities between processing frames F. This processing is described here using two processing frames F. The two processing frames F are the previous processing frame and the current processing frame. The current processing frame is the processing frame F being processed by the processing unit 120 at that moment, and the previous processing frame is the processing frame F immediately preceding the current processing frame.

[0380] The connecting unit 1003 performs windowed addition processing on multiple temporally subsequent sample points of the reconfigured aerodynamic sound data generated from the aerodynamic sound data of the previous processing frame and multiple temporally preceding sample points of the reconfigured aerodynamic sound data generated from the aerodynamic sound data of the current processing frame. This processing avoids discontinuities between processing frames F caused by fluctuations in the value represented by the smoothing function.

[0381] Figure 25 : is a graph showing aerodynamic sound data related to this modification. Figure 26 This is a conceptual diagram of the processing performed by the second processing unit 122 of this modification. The aerodynamic sound data is processed in units of processing frames F. In addition, two adjacent processing frames F are set so that they partially overlap each other. This is to perform windowing and addition on one or more sample points located at the rear of the multiple sample points after the previous processing frame is reconfigured, and one or more sample points located at the front of the multiple sample points after the current processing frame is reconfigured, so as to avoid discontinuity. For example, Figure 25 As shown, two adjacent processing frames Fn and Fn+1 partially overlap. More specifically, from time t14 to time t13, the two processing frames Fn and Fn+1 overlap. Furthermore, processing frame Fn corresponds to the previous processing frame, and processing frame Fn+1 corresponds to the current processing frame.

[0382] An example will be described in which the sampling rate of aerodynamic sound data is set to Fs, the value represented by the smoothing function for processing frame Fn is set to 0.5, and the value represented by the smoothing function for processing frame Fn+1 is set to 0.75. In processing frame Fn, the value represented by the smoothing function is 0.5, so the sampling rate is converted to 2·Fs (the sampling interval is 1 / (2·Fs)). The relocation unit 1002 then relocates the positions of the sample points after the sampling rate conversion so that the sampling interval becomes 1 / Fs, restoring the original sampling interval. Therefore, the time length of the relocated sample points is twice the time length of the aerodynamic sound data sample points converted by the sampling rate conversion unit 1001.

[0383] Next, windowing and summing are performed on one or more sample points located at the rear of the rearranged sample points and one or more sample points located at the front of the rearranged sample points in the current processing frame, and the resulting data is output. In this example, since the value represented by the smoothing function in processing frame n+1 is 0.75, the time length of the rearranged sample points is 4 / 3 the time length of the sample points of the aerodynamic sound data converted by the sampling rate conversion unit 1001. Furthermore, the rearranged sample points in the intervals not subjected to windowing and summing are output as sound data.

[0384] Here, use Figure 27 The sampling rate conversion unit 1001 will be described in more detail.

[0385] Figure 27 1021 , a low-pass filter 1022 , a down-sampling unit 1023 , and an XY setting unit 1024 .

[0386] The upsampling unit 1021 obtains sound data (aerodynamic sound data), and the XY setting unit 1024 obtains the value represented by the smoothing function. The XY setting unit 1024 sets the upsampling value X used by the upsampling unit 1021 and the downsampling value Y used by the downsampling unit 1023. Here, when the upsampling value is X, the upsampling unit 1021 upsamples the aerodynamic sound data by a factor of X. When the downsampling value is Y, the downsampling unit 1023 downsamples the aerodynamic sound data by a factor of 1 / Y. The settings of X and Y in the XY setting unit 1024 are determined to be the smallest integers for the combination of X and Y that results in Y / X being the value represented by the smoothing function. For example, when the value represented by the smoothing function is 0.5, (X, Y) = (2, 1); when the value represented by the smoothing function is 0.75, (X, Y) = (4, 3); and when the value represented by the smoothing function is 1.5, (X, Y) = (2, 3). When X=1, the upsampling unit 1021 does not perform upsampling but outputs the pneumatic sound data as is, and when Y=1, the downsampling unit 1023 does not perform downsampling but outputs the pneumatic sound data as is.

[0387] The upsampling unit 1021 inserts X-1 zero values ​​between sample points. The downsampling unit 1023 thins out every Y sample points and outputs them. The low-pass filter unit 1022 performs the following processing to prevent aliasing. Here, the sampling rate of the aerodynamic sound data is assumed to be Fs, and the sampling rate of the aerodynamic sound data after sampling rate conversion is assumed to be Fs'. At this point, the low-pass filter unit 1022 processes the aerodynamic sound data output from the upsampling unit 1021 using a low-pass filter with a cutoff frequency of min(Fs, Fs') / 2.

[0388] Furthermore, the temporal variation patterns of the values ​​represented by the smoothing function are illustrated. Here, the values ​​represented by the smoothing function are represented by one of five values. Here, variation patterns 1 and 2 are described.

[0389] In variation pattern 1, the value represented by the smoothing function is one of 0.25, 0.5, 1, 2, and 4. In variation pattern 2, the value represented by the smoothing function is one of 0.5, 0.75, 1, 1.5, and 2. The values ​​or number of values ​​that the smoothing function can represent are not limited to those exemplified here.

[0390] also, Figure 28 This is a state transition diagram of the value represented by the smoothing function of this modification. That is, Figure 28This indicates the temporal transition of the value represented by the smoothing function. Each circle represents a state, and in the state p(0), p(0) is output as the value represented by the smoothing function. Furthermore, a(e, f) represents the probability of transitioning from state e to state f. To represent natural sound fluctuations, it is desirable to only recognize transitions to the state itself or to adjacent states, as in this example. However, depending on the application, there may be cases where more dramatic fluctuations are desired, so this is not a limitation; any transition can be specified.

[0391] In addition, in this modification, a process of adding a change to the amplitude value of the aerodynamic sound data acquired by the sampling rate conversion unit 1001 may be performed.

[0392] Figure 29 This is a block diagram showing other functional components of the acoustic signal processing device 100a according to this variation. Here, the processing unit 120a of the acoustic signal processing device 100a includes a second processing unit 122b in place of the second processing unit 122. The second processing unit 122b includes a sampling rate converter 1001, an amplitude adjuster 1031, a reconfiguration unit 1002, and a connector 1003.

[0393] exist Figure 29 In the example, an amplitude adjustment unit 1031 is provided at the rear of the sampling rate conversion unit 1001. The amplitude adjustment unit 1031 corrects the amplitude value so that the amplitude value of the aerodynamic sound data after sampling rate conversion output from the sampling rate conversion unit 1001 fluctuates. As a correction method, for example, Figure 28 Alternatively, the amplitude value may be modified by multiplying the aerodynamic sound data by one of a plurality of amplitude variation patterns prepared in advance.

[0394] In addition, the amplitude adjustment unit 1031 may be located after the rearrangement unit 1002 or after the connection unit 1003 .

[0395] (Implementation Method 2)

[0396] Hereinafter, a description will be given of Embodiment 2. The following description will focus on the differences from Embodiment 1 and its modifications, and descriptions of the common points will be omitted or simplified.

[0397] [constitute]

[0398] First, the configuration of the information processing device 600 according to this embodiment will be described. Figure 30 This is a block diagram showing the functional configuration of the information processing device 600 according to this embodiment.

[0399] The information processing device 600 includes a loop address unit 610 , a frequency shift unit 620 , a storage unit 630 , a section specifying unit 640 , a crossfade unit 650 , and a readout control unit 660 .

[0400] If the duration of aerodynamic sound data is short, repeated use of the aerodynamic sound data may generate noise at the junction of the aerodynamic sound data. The information processing device 600 according to this embodiment solves at least one of these problems.

[0401] Figure 31 It is a diagram for explaining the reading of audio data according to the conventional technique and the reading of audio data according to the present embodiment. Figure 31 (a) is a diagram for explaining the reading of audio data in the prior art. Figure 31 (b) is a diagram for explaining the reading of audio data according to this embodiment.

[0402] The following describes the reading of sound data (pneumatic sound data) in the prior art. In the prior art, a storage unit is provided to store the pneumatic sound data, and a loop address unit loops from the address of the starting point of the pneumatic sound data storage to the address of the ending point of the pneumatic sound data storage in the storage unit. The loop address unit reads the pneumatic sound data from the storage unit and outputs it.

[0403] Next, the reading of sound data (aerodynamic sound data) according to this embodiment will be described.

[0404] Here, aerodynamic sound data (e.g. Figure 15 The aerodynamic sound data D1 before processing is composed of a plurality of sample points. More specifically, Figure 31 As shown in (b), the data consists of N sample points. Here, the first M sample points and the last M sample points of the aerodynamic sound data are cross-faded to create the M cross-faded sample points. Furthermore, (N - 2M) samples are created from the middle portion of the aerodynamic sound data after removing the first M sample points and the last M sample points.

[0405] In the storage unit 630 of this embodiment, aerodynamic sound data consisting of (N-M) samples, which is a combination of M cross-faded sample points and (N-2M) intermediate samples, is stored. A series of (N-M) addresses corresponding to the aerodynamic sound data consisting of (N-M) samples is set in the storage unit 630.

[0406] In this embodiment, the loop address unit 610 loops through the aerodynamic sound data consisting of (N-M) samples stored in the storage unit 630 from the start address to the end address, reads the aerodynamic sound data, and outputs it to the frequency shift unit 620. The frequency shift unit 620 takes the output aerodynamic sound data, shifts its frequency, and outputs it to, for example, the output channel of the earphone 200 according to Embodiment 1.

[0407] In the information processing device 600 according to this embodiment, since the first M sample points and the last M sample points are cross-faded, problems such as noise generation at the junctions between aerodynamic sound data are less likely to occur.

[0408] Furthermore, the information processing device 600 according to this embodiment preferably performs the following processing. Figure 32 This is a diagram for explaining processing performed by the information processing device 600 according to this embodiment.

[0409] Figure 32 (a) is a diagram showing the structure of the storage unit 630 of this embodiment. Here, the storage unit 630 stores aerodynamic sound data (for example, Figure 15 1 ). Furthermore, a first pointer Pt1 and a second pointer Pt2 are provided. The first pointer Pt1 indicates the position at which the stored aerodynamic sound data is read. The second pointer Pt2 moves in conjunction with the first pointer Pt1 and indicates the position at which the aerodynamic sound data is read from the storage unit 630.

[0410] The section designation unit 640 designates a first section A1 and a second section A2. The second section A2 is a subsequent section adjacent to the first section A1. The second pointer Pt2 moves to a subsequent section A3 adjacent to the second section A2.

[0411] Furthermore, the first section A1 and the second section A2 may be arbitrarily set by a user of the information processing device 600. Specifically, the receiving unit of the information processing device 600 may receive an operation from the user indicating the first section A1 and the second section A2, and may determine the sections indicated by the received operation to be the first section A1 and the second section A2 by the section specifying unit 640.

[0412] The crossfade unit 650 performs a fade-in process on the aerodynamic sound data read from the readout position indicated by the first pointer Pt1 and outputs the aerodynamic sound data after the fade-in process. The crossfade unit 650 performs a fade-out process on the aerodynamic sound data read from the readout position indicated by the second pointer Pt2 and outputs the aerodynamic sound data after the fade-out process.

[0413] While the reading position indicated by the first pointer Pt1 is included in the first interval A1 and aerodynamic sound data is being read from the first interval A1, the readout control unit 660 causes the crossfade unit 650 to output aerodynamic sound data after fading in. While the reading position indicated by the first pointer Pt1 is not included in the first interval A1 and aerodynamic sound data is not being read from the first interval A1, the readout control unit 660 outputs aerodynamic sound data read from the second interval A2 by the loop address unit 610.

[0414] Next, the aerodynamic sound data after the fade-in processing output by the crossfade unit 650 or the aerodynamic sound data read from the second interval A2 by the loop address unit 610 is output to the frequency shift unit 620. The frequency shift unit 620 receives the aerodynamic sound data after the fade-in processing output or the aerodynamic sound data read from the second interval A2, shifts the frequency thereof, and outputs the data to, for example, the output channel of the earphone 200 according to the first embodiment.

[0415] Then, Figure 32 The processing shown in (b) and (c) is explained.

[0416] Figure 32 (b) is a diagram showing an example of the first pointer Pt1 of the present embodiment cycling through the first interval A1 and the second interval A2. In this example, the first pointer Pt1 cycles through the first interval A1 and the second interval A2. While the readout position indicated by the first pointer Pt1 is included in the first interval A1, aerodynamic sound data is read out from the readout position indicated by the first pointer Pt1, and aerodynamic sound data is also read out from the readout position indicated by the second pointer Pt2 linked to the first pointer Pt1. The crossfade unit 650 performs crossfading processing on the two readout aerodynamic sound data. In addition, while the readout position indicated by the first pointer Pt1 is included in the first interval A1, the readout position indicated by the second pointer Pt2 can be linked to the first pointer Pt1 so that the aerodynamic sound data can also be read out from the interval A3.

[0417] Figure 32(c) is a diagram showing an example of the second pointer Pt2 of the present embodiment circulating in the second interval A2 and the interval A3. In this example, the second pointer Pt2 circulates in the second interval A2 and the interval A3. During the period when the readout position represented by the second pointer Pt2 is included in the interval A3, the aerodynamic sound data is read out from the readout position represented by the second pointer Pt2, and the aerodynamic sound data is also read out from the readout position represented by the first pointer Pt1. The cross-fade unit 650 performs cross-fading processing on the two readout aerodynamic sound data. In addition, during the period when the readout position represented by the second pointer Pt2 is included in the second interval A2, the readout position represented by the first pointer Pt1 can be included in the first interval A1 in conjunction with the second pointer Pt2, and the aerodynamic sound data can also be read out from the first interval A1.

[0418] Furthermore, the information processing device 600 according to this embodiment can perform the following processing. Figure 33 It is a diagram for explaining other processing performed by the information processing device 600 according to this embodiment.

[0419] In this other process, the section specifying unit 640 randomly updates the first section A1 and the second section A2. The section specifying unit 640 sequentially updates the end point of the second section A2 and the start and end points of the next first section A1.

[0420] exist Figure 33 In the other processes shown, the pneumatic sound data is read as Figure 33 (a) Figure 33 (b) Figure 33 (c) Figure 33 (d) Figure 33 (e), Figure 33 (f), Figure 33 (g) Sequential transfer.

[0421] Figure 33 (a), (d) and (g) respectively represent the state 1 where the aerodynamic sound data is read out, Figure 33 (b) and (e) respectively represent the state 2 where the aerodynamic sound data is read out, Figure 33 (c) and (f) respectively represent state 3 where the aerodynamic sound data is read.

[0422] exist Figure 33 In the above example, state 1, state 2, and state 3 are repeated in this order.

[0423] exist Figure 33 In the state 1 shown in (a), the aerodynamic sound data is read from the second interval A2. In addition, at this time, the end point of the second interval A2 has not been determined.

[0424] exist Figure 33In state 2 shown in (b), aerodynamic sound data is read from the second interval A2. Next, the interval designation unit 640 arbitrarily designates the end point of the second interval A2 and the next first interval A1 at a predetermined timing. Furthermore, since interval A3, which is linked to the next first interval A1, is a subsequent interval adjacent to the second interval A2, it is automatically determined without requiring designation by the interval designation unit 640.

[0425] The predetermined timing may be arbitrarily set by the user of information processing device 600. That is, the receiving unit included in information processing device 600 may receive an operation from the user to instruct the predetermined timing, and section designation unit 640 may determine the timing indicated by the received operation as the predetermined timing.

[0426] exist Figure 33 In state 3 shown in (c), the reading of the aerodynamic sound data from the second interval A2 is completed. Then, the crossfade unit 650 performs a crossfade process on the aerodynamic sound data read from the next first interval A1 and the aerodynamic sound data read from the interval A3 linked to the next first interval A1.

[0427] exist Figure 33 In the state 1 represented by (d), the aerodynamic sound data is read from the next second interval A2. In addition, since the next second interval A2 is Figure 33 Therefore, the interval designation unit 640 does not need to designate the starting point of the next second interval A2, which is automatically determined. Figure 33 When the crossfade process described in (c) is completed, the aerodynamic sound data is read from the next second interval A2. Figure 33 Similarly to the state 1 shown in (a), the end point of the second section A2 has not been determined.

[0428] exist Figure 33 In the state 2 shown in (e), from the second interval A2 (equivalent to Figure 33 The aerodynamic sound data is read for the next second interval A2 shown in (d). Next, the interval designation unit 640 arbitrarily designates the end point of the second interval A2 and the next first interval A1 at a predetermined timing. Since interval A3, which is linked to the next first interval A1, is a subsequent interval adjacent to the second interval A2, it is automatically determined without requiring designation by the interval designation unit 640.

[0429] exist Figure 33 In the state 3 represented by (f), from the second interval A2 (equivalent to Figure 33The reading of the aerodynamic sound data of the next second interval A2 shown in (e) is completed. Then, the crossfade unit 650 performs a crossfade process on the aerodynamic sound data read from the next first interval A1 and the aerodynamic sound data read from the interval A3 linked to the next first interval A1.

[0430] exist Figure 33 In the state 1 shown in (g), the aerodynamic sound data is read from the next second interval A2. In addition, the next second interval A2 is the same as Figure 33 Therefore, the interval designation unit 640 does not need to designate the starting point of the next second interval A2, which is automatically determined. Figure 33 When the crossfade process described in (c) is completed, the aerodynamic sound data is read from the next second interval A2. Figure 33 Similarly to the state 1 shown in (a), the end point of the second section A2 has not been determined.

[0431] like Figure 33 As shown, states 1, 2, and 3 are repeated in this order. By arbitrarily specifying the end point of the second interval A2 and the next first interval A1 in state 2, the listener L is prevented from repeatedly hearing the same aerodynamic sound. Thus, the unnatural "rhythm" caused by the repetition of the same aerodynamic sound is not generated.

[0432] Next, pipeline processing will be described.

[0433] Part or all of the processing performed by the above-described sound signal processing device 100 may be performed as part of pipeline processing as described in Patent Document 2, for example. Figure 34 It is used to illustrate Figure 6 and Figure 7 This is a functional block diagram and an example of the steps for the rendering units A0203 and A0213 to perform pipeline processing. Figure 34 In the description, use Figure 6 and Figure 7 The rendering unit 900 as an example of the rendering units A0203 and A0213 will be described.

[0434] Pipeline processing is a process of dividing the process of applying an acoustic effect into multiple processes and executing each process in sequence. Each of the divided processes may perform, for example, signal processing of an audio signal or generation of parameters used in the signal processing.

[0435] The rendering unit 900 of this embodiment includes, for example, implementing reverberation effects, initial reflection processing, distance attenuation effects, binaural processing and other processing as pipeline processing. However, the above-mentioned processing is an example, and it may also include processing other than it, or may not include some of the processing. For example, the rendering unit 900 may include diffraction processing or occlusion processing as pipeline processing, for example, it may be omitted when reverberation processing is not required. In addition, each processing may be represented as a stage, and the sound signal such as the reflected sound generated by the result of each processing may be represented as a rendering item. The order of each stage in the pipeline processing and the stages included in the pipeline processing are not limited to Figure 34 Example shown.

[0436] In addition, the rendering unit 900 may not include Figure 34 Of all the stages shown, some stages may be omitted, or other stages may exist in addition to the rendering unit 900 .

[0437] As an example of pipeline processing, the following describes the processes performed in reverberation, early reflections, distance attenuation, selection, generation, and binaural processing. Each process analyzes metadata contained in the input signal and calculates the parameters required to generate reflected sound.

[0438] In addition, Figure 34 In the embodiment, the rendering unit 900 includes a reverberation processing unit 901, an initial reflection processing unit 902, a distance decay processing unit 903, a selection unit 904, a calculation unit 906, a generation unit 907, and a binaural processing unit 905. Here, an example is described in which the reverberation processing unit 901 performs a reverberation processing step, the initial reflection processing unit 902 performs an initial reflection processing step, the distance decay processing unit 903 performs a distance decay processing step, the selection unit 904 performs a selection processing step, and the binaural processing unit 905 performs a binaural processing step.

[0439] In the reverberation processing step, the reverberation processing unit 901 generates a sound signal representing reverberation sound or parameters required for generating the sound signal. Reverberation sound is a sound that arrives at the listener as reverberation after the direct sound, including the reverberation sound. As an example, reverberation sound arrives at the listener relatively late (for example, about a hundred or so milliseconds from the arrival of the direct sound) after the initial reflected sound, described later, arrives at the listener after a greater number of reflections (for example, several dozen times) than the initial reflected sound. The reverberation processing unit 901 calculates the reverberation sound using a pre-prepared function for generating reverberation sound, referring to the sound signal and spatial information included in the input signal.

[0440] The reverberation processing unit 901 may also apply a known reverberation generation method to the sound signal to generate reverberation. As an example, the known reverberation generation method is the Schroeder method, but the method is not limited thereto. Furthermore, when applying the known reverberation generation process, the reverberation processing unit 901 utilizes the shape and acoustic characteristics of the sound reproduction space represented by the spatial information. This allows the reverberation processing unit 901 to calculate parameters for generating a sound signal representing reverberation.

[0441] In the initial reflection processing step, the initial reflection processing unit 902 calculates the parameters for generating the initial reflected sound based on the spatial information. The initial reflected sound is the reflected sound that reaches the listener after one or more reflections in the relatively early period (for example, about tens of ms from the time when the direct sound arrives) after the direct sound arrives from the sound source object. The initial reflection processing unit 902, for example, refers to the sound signal and metadata, and uses the shape, size, position of objects such as structures and the reflectivity of the objects in the three-dimensional sound field (space) to calculate the path (the length of the path) of the reflected sound that is reflected from the sound source object and reaches the listener. In addition, the initial reflection processing unit 902 can also calculate the path (the length of the path) of the direct sound. The information representing the path can also be used as a parameter for generating the initial reflected sound, and can also be used as a parameter for the selection processing of the reflected sound in the selection unit 904.

[0442] In the distance attenuation processing step, the distance attenuation processing unit 903 calculates the volume reaching the listener based on the difference between the length of the direct sound path and the length of the reflected sound path calculated by the initial reflection processing unit 902. The volume reaching the listener attenuates in direct proportion to the distance from the sound source (and inversely proportional to the distance). Therefore, the volume of the direct sound can be calculated by dividing the volume of the sound source by the length of the direct sound path, and the volume of the reflected sound can be calculated by dividing the volume of the sound source by the length of the reflected sound path.

[0443] In the selection process step, the selection unit 904 selects a sound to be generated. The selection process may be performed based on the parameters calculated in the previous step.

[0444] When selection processing is performed as part of the pipeline processing, sounds not selected in the selection processing may not be subject to processing subsequent to the selection processing in the pipeline processing. By not performing processing subsequent to the selection processing on unselected sounds, the computational load of the audio signal processing device 100 can be reduced compared to a case where only binaural processing is not performed on unselected sounds.

[0445] Furthermore, when the selection process described in this embodiment is executed as part of a pipeline process, if the selection process is ordered so that it is executed earlier in the order of the multiple processes in the pipeline process, more processing subsequent to the selection process can be omitted, thereby further reducing the amount of computation. For example, if the calculation unit 906 and the generation unit 907 execute the selection process earlier than the earlier processing, processing of aerodynamic sounds associated with objects determined not to be selected can be omitted, further reducing the amount of computation in the acoustic signal processing device 100.

[0446] Furthermore, parameters calculated in a part of the pipeline process for generating a rendering item may be used in the selection unit 904 or the calculation unit 906 .

[0447] In the binaural processing step, the binaural processing unit 905 performs signal processing on the sound signal of the direct sound so that it is perceived as sound reaching the listener from the direction of the sound source object. Furthermore, the binaural processing unit 905 performs signal processing so that the reflected sound is perceived as sound reaching the listener from the obstacle object related to the reflection. Based on the coordinates and orientation of the listener in the sound space (that is, the position and orientation of the listening point), processing is performed by applying the HRIR (Head-Related Impulse Responses) DB (Data base) so that the sound reaches the listener from the position of the sound source object or the position of the obstacle object. In addition, the position and direction of the listening point can change, for example, to match the movement of the listener's head. In addition, information indicating the position of the listener can also be obtained from the sensor.

[0448] Programs used in pipeline processing and binaural processing, spatial information required for audio processing, the HRIR database, threshold data, and other parameters are acquired from the audio signal processing device 100's memory or from an external source. HRIR (Head-Related Impulse Responses) are the response characteristics when a single impulse is generated. Specifically, HRIR is the response characteristic obtained by Fourier transforming the head-related transfer function, which represents the changes in sound generated by surrounding objects such as the concha, head, and shoulders, from the frequency domain to the time domain. The HRIR database is a database containing this information.

[0449] Furthermore, as an example of pipeline processing, the rendering unit 900 may include a processing unit (not shown), such as a diffraction processing unit or an occlusion processing unit.

[0450] The diffraction processing unit generates a sound signal representing a sound including diffracted sound caused by an obstacle between a listener and a sound source object in a three-dimensional sound field (space). Diffracted sound is sound that, when there is an obstacle between the sound source object and the listener, travels from the sound source object to the listener, bypassing the obstacle.

[0451] The diffraction processing unit, for example, refers to the sound signal and metadata, uses the position of the sound source object in the three-dimensional sound field (space), the position of the listener, and the position, shape and size of the obstacle, etc., to calculate the path from the sound source object around the obstacle to reach the listener, and generates diffracted sound based on the path.

[0452] The occlusion processing unit generates a sound signal that can be vaguely heard when a sound source object is located on the opposite side of the obstacle object, based on the spatial information and information such as the material of the obstacle object acquired in a certain step.

[0453] Furthermore, in the above-described embodiments 1 and 2, the position information assigned to the sound source object defines a "point" within the virtual space, assuming a so-called "point sound source" for the detailed description of the invention. Alternatively, as a method of defining a sound source in a virtual space, there are also cases where a sound source that is spatially extended, rather than a point source, is defined as an object with length, size, or shape. In such cases, since the distance between the listener and the sound source or the direction of arrival of the sound is uncertain, the resulting reflected sound does not need to be analyzed, or, regardless of the analysis results, can be limited to the "selection" process in the aforementioned selection unit 904. This avoids the potential degradation of sound quality caused by not selecting reflected sound. Alternatively, a representative point, such as the center of gravity of the object, can be set and the processing of the present disclosure can be applied as if the sound originates from this representative point. In this case, the processing of the present disclosure can also be applied after adjusting the threshold based on the spatial extension information of the sound source.

[0454] Next, an example of the structure of a bit stream is described.

[0455] The bitstream includes, for example, sound signals and metadata. The sound signal is sound data that represents the sound, indicating information such as the frequency and strength of the sound. The spatial information included in the metadata is information about the space in which the listener is located when listening to the sound based on the sound signal. Specifically, the spatial information is information about the predetermined position (localization position) when the sound image of the sound is localized at a predetermined position in the sound space (for example, within a three-dimensional sound field), that is, when the listener perceives the sound as arriving from a predetermined direction. The spatial information includes, for example, sound source object information and position information indicating the listener's position.

[0456] Sound source object information is information about an object that produces sound based on a sound signal, that is, an object that reproduces the sound signal. It is information related to a virtual object (sound source object) located in the sound space, a virtual space corresponding to the real space in which the object is located. This sound source object information includes, for example, information indicating the position of the sound source object in the sound space, information regarding the orientation of the sound source object, information regarding the directionality of the sound emitted by the sound source object, information indicating whether the sound source object is a living being, and information indicating whether the sound source object is a moving object. For example, a sound signal corresponds to one or more sound source objects indicated by the sound source object information.

[0457] As an example of the data structure of a bit stream, the bit stream is composed of metadata (control information) and an audio signal, for example.

[0458] The audio signal and metadata can be stored in one bitstream or in multiple bitstreams. Similarly, the audio signal and metadata can be stored in one file or in multiple files.

[0459] A bitstream may exist for each sound source or for each playback time. In the case where a bitstream exists for each playback time, multiple bitstreams may be processed in parallel.

[0460] Metadata can be added for each bitstream or as information for controlling multiple bitstreams. Metadata can also be added for each playback time.

[0461] When audio signals and metadata are stored separately in multiple bitstreams or files, the audio signals and metadata may be included in information indicating other bitstreams or files associated with one or some bitstreams or files. Alternatively, the audio signals and metadata may be included in information indicating other bitstreams or files associated with all bitstreams or files. Here, associated bitstreams or files may be, for example, bitstreams or files that may be used simultaneously during audio processing. Furthermore, associated bitstreams or files may include bitstreams or files that also describe information indicating other associated bitstreams or files. Information indicating other associated bitstreams or files may include, for example, identifiers indicating the other bitstreams, file names indicating the other files, URLs (Uniform Resource Locators) or URIs (Uniform Resource Identifiers). In this case, the acquisition unit 110 identifies or acquires the bitstreams or files based on the information indicating the other associated bitstreams or files. Furthermore, a bitstream may include information indicating other associated bitstreams, and information indicating bitstreams or files associated with other bitstreams or files may also be included within the bitstream. Here, the file including information indicating the associated bitstream or file may be, for example, a control file such as a manifest file used for content distribution.

[0462] Furthermore, all or part of the metadata may be obtained from sources other than the audio signal bitstream. For example, either metadata for controlling the audio or metadata for controlling the video may be obtained from sources other than the bitstream, or both metadata may be obtained from sources other than the bitstream. Furthermore, if metadata for controlling the video is included in the bitstream obtained by the audio signal reproduction system, the audio signal reproduction system may also include a function for outputting metadata that can be used to control the video to a display device that displays images or a stereoscopic video reproduction device that reproduces stereoscopic images.

[0463] Next, examples of information included in metadata will be described.

[0464] Metadata can also be used to describe the scene represented by the sound space. Here, "scene" refers to the collection of all elements of the three-dimensional image and acoustic events in the sound space modeled by the sound signal reproduction system using metadata. Specifically, the metadata mentioned here includes not only information for controlling audio processing, but also information for controlling video processing. Of course, metadata can include information for controlling only one of the audio and video processing, or information used for controlling both.

[0465] The sound signal reproduction system uses metadata included in the bitstream and additional interactive listener position information to perform audio processing on the sound signal, generating virtual sound effects. While this description describes the sound effects as including initial reflection processing, obstacle processing, diffraction processing, occlusion processing, and reverberation processing, other audio processing techniques can also be performed using metadata. For example, the sound signal reproduction system can incorporate sound effects such as distance attenuation, localization, and the Doppler effect. Furthermore, metadata can be added to the system to toggle all or some of the sound effects on and off, as well as priority information.

[0466] In addition, as an example, the encoded metadata includes: information related to the sound space containing the sound source object and the obstacle object, and information related to the positioning position when the sound image of the sound is positioned at a specified position in the sound space (that is, so that it is perceived as a sound arriving from a specified direction). Here, the obstacle object is an object that may affect the sound perceived by the listener, such as by blocking or reflecting the sound, during the period from the sound emitted by the sound source object to the sound reaching the listener. In addition to stationary objects, obstacle objects may also include animals such as humans or moving objects such as machinery. In addition, when there are multiple sound source objects in the sound space, for any sound source object, other sound source objects may become obstacle objects. Non-sound-emitting objects that do not emit sound, such as building materials or inanimate objects, and sound source objects that emit sound may also become obstacle objects.

[0467] The metadata includes all or part of the information representing the shape of the sound space, the shape information and position information of obstacle objects existing in the sound space, the shape information and position information of sound source objects existing in the sound space, and the position and orientation of the listener in the sound space.

[0468] The sound space can be either closed or open. In addition, the metadata includes information indicating the reflectivity of structures that can reflect sound in the sound space, such as floors, walls, or ceilings, as well as the reflectivity of obstacles present in the sound space. Here, the reflectivity is the energy ratio of reflected sound to incident sound, and is set for each frequency band of sound. Of course, the reflectivity can also be set uniformly regardless of the frequency band of the sound. In the case of an open space, for example, uniformly set parameters such as attenuation rate, diffracted sound, or initial reflected sound can also be used.

[0469] While reflectivity is listed as a parameter included in metadata related to obstacle objects or sound source objects, information other than reflectivity can also be included. For example, information other than reflectivity can also include information related to the object's material as metadata related to both sound source objects and non-sound-emitting objects. Specifically, information other than reflectivity can also include parameters such as diffusivity, transmittance, and sound absorption.

[0470] Information related to the sound source object may also include volume, radiation characteristics (directivity), reproduction conditions, the number and type of sound sources emitted from an object, and information specifying the sound source area in the object. The reproduction conditions may also include, for example, whether the sound is continuously flowing or the sound is triggered by an event. The sound source area in the object can be set by the relative relationship between the listener's position and the object's position, or it can be set based on the object. When the sound source area in the object is set by the relative relationship between the listener's position and the object's position, the listener can perceive that sound C is emitted from the right side of the object and sound E is emitted from the left side when the listener is looking at it, based on the surface of the object viewed by the listener. When the sound source area in the object is set based on the object, it is possible to determine which area of ​​the object emits which sound regardless of the direction the listener is looking. For example, the listener can perceive that when viewing the object from the front, high notes are emitted from the right side and low notes are emitted from the left side. In this case, if the listener goes around to the back of the object, the listener can perceive that when viewing from the back, low notes are emitted from the right side and high notes are emitted from the left side.

[0471] Metadata related to space may include time until initial reflection, reverberation time, ratio of direct sound to diffuse sound, etc. When the ratio of direct sound to diffuse sound is zero, the listener can perceive only direct sound.

[0472] (Effects, etc.)

[0473] The sound signal processing method according to embodiment 1 includes: an acquisition step of acquiring sound data representing a waveform of a reference sound; a processing step of processing the sound data based on simulation information simulating changes in natural phenomena so as to change at least one of the frequency component, phase, and amplitude value of the waveform; and an output step of outputting the processed sound data.

[0474] Thus, based on simulation information simulating the fluctuations of natural phenomena involving fluctuations, sound data is processed to vary at least one of the frequency component, phase, and amplitude of the waveform. Consequently, at least one of the frequency component, phase, and amplitude fluctuates in the processed sound data, and at least one of the frequency component, phase, and amplitude fluctuates in the sound represented by the processed sound data. Consequently, listener L can hear a sound in which at least one of the frequency component, phase, and amplitude fluctuates, and listener L can experience a sense of presence without feeling any discomfort. In other words, an acoustic signal processing method capable of providing listener L with a sense of presence is achieved.

[0475] In Action Example 1 of the above-described embodiment 1, the natural phenomenon used is the blowing of wind W. As described above, the simulation information simulates the fluctuations of a natural phenomenon including fluctuations, and more specifically, represents the fluctuations caused by changes in the wind speed of wind W. In Action Example 1, this information is represented by a smoothing function.

[0476] In Example 1, sound data (aerodynamic sound data) representing the waveform of a reference sound is processed based on simulation information that simulates the fluctuations of natural phenomena involving fluctuations, thereby varying the frequency components of the waveform. Consequently, the frequency components of the processed aerodynamic sound data fluctuate, and the frequency components of the aerodynamic sound represented by the processed aerodynamic sound data also fluctuate. Consequently, the listener L can hear the aerodynamic sound with fluctuating frequency components, thereby minimizing the sense of discomfort and providing a sense of immersiveness. In other words, an acoustic signal processing method capable of providing the listener L with a sense of immersiveness is achieved.

[0477] Furthermore, in Operation Example 1 of the above-described Embodiment 1, the blowing of wind W is used as an example of a natural phenomenon. However, the present invention is not limited thereto, and natural phenomena such as the flow of river water and the movement of animals may also be used.

[0478] When the example of the flow of a river is used as a natural phenomenon, the listener L hears the gurgling sound of the river. In this case, the analog information is information expressing fluctuations caused by changes in the flow rate of the river or changes in the direction of the flow of the river.

[0479] When the behavior of an animal is used as an example of a natural phenomenon, the listener L hears the animal's cry, etc. In this case, the pseudo information is information expressing fluctuations caused by changes in the volume of the animal's cry, etc.

[0480] Specifically, when using natural phenomena such as the flow of a river or the movement of animals, the simulated information also simulates the fluctuations of natural phenomena involving fluctuations. Therefore, as shown in Example 1, by using the simulated information, listener L can hear a sound that fluctuates in at least one of its frequency components, phase, and amplitude, while also minimizing the sense of discomfort and allowing listener L to experience a sense of presence. This achieves an acoustic signal processing method that provides listener L with a sense of presence.

[0481] In the acoustic signal processing method according to embodiment 1, the reference sound is the aerodynamic sound generated by the wind W; in the processing step, the sound data is processed based on simulation information that simulates the change in wind speed of the wind W so as to change at least one of the frequency component, phase and amplitude value of the waveform.

[0482] As a result, the listener L can hear aerodynamic sound that fluctuates in at least one of its frequency component, phase, and amplitude, and the listener L can experience a sense of presence without feeling any discomfort. In other words, an acoustic signal processing method that can provide the listener L with a sense of presence is achieved.

[0483] In the sound signal processing method related to embodiment 1, in the processing step, a smoothing function that simulates the change in wind speed of wind W is determined as simulation information; based on the value represented by the determined smoothing function, the sound data is processed so that at least one of the frequency component, phase and amplitude value of the waveform changes.

[0484] This makes it possible to process audio data according to the value represented by the smoothing function.

[0485] In the acoustic signal processing method according to the first embodiment, the value represented by the smoothing function is information representing the ratio of the wind speed of the aerodynamic sound serving as the reference sound to the wind speed of the aerodynamic sound represented by the sound data processed in the processing step.

[0486] This makes it possible to process the sound data based on the ratio of the wind speed of the aerodynamic sound serving as the reference sound to the wind speed of the aerodynamic sound represented by the processed sound data.

[0487] In the acoustic signal processing method according to the first embodiment, in the processing step, a smoothing function is determined so that a parameter defining the smoothing function changes irregularly.

[0488] As a result, the listener L can hear aerodynamic sound that fluctuates irregularly in at least one of its frequency components, phase, and amplitude. This reduces the sense of discomfort felt by the listener L and allows for a better sense of presence. In other words, an acoustic signal processing method that provides the listener L with a better sense of presence is achieved.

[0489] In the acoustic signal processing method according to the first embodiment, in the processing step, the audio data is processed so that the frequency component of the waveform is shifted to a frequency proportional to a value represented by the determined smoothing function.

[0490] As a result, the listener L can hear the sound with fluctuating frequency components, and can obtain a sense of immersion without causing discomfort to the listener L. In other words, an acoustic signal processing method capable of providing the listener L with a sense of immersion is realized.

[0491] Specifically, as shown in Example 1, the sound data (aerodynamic sound data) representing the waveform of the reference sound is processed based on simulation information (smoothing function) that simulates the fluctuations in wind speed of the fluctuating wind W, thereby varying the frequency components of the waveform. Consequently, the frequency components of the processed aerodynamic sound data fluctuate, and the frequency components of the aerodynamic sound represented by the processed aerodynamic sound data also fluctuate. Consequently, the listener L can hear the aerodynamic sound with such fluctuating frequency components, thereby minimizing the sense of discomfort and providing a sense of presence.

[0492] In the acoustic signal processing method according to the first embodiment, in the processing step, the audio data is processed so that the amplitude value of the waveform changes in proportion to the α-th power of the value represented by the determined smoothing function.

[0493] As a result, the listener L can hear the sound with fluctuating amplitude values, and can obtain a sense of immersion without causing discomfort to the listener L. In other words, an acoustic signal processing method capable of providing the listener L with a sense of immersion is realized.

[0494] Specifically, as shown in Example 2, the sound data (aerodynamic sound data) representing the waveform of the reference sound is processed so that the amplitude of the waveform varies proportionally to the αth power of the value represented by the smoothing function, which is simulated information simulating the fluctuations in wind speed of the fluctuating wind W. Consequently, the amplitude of the processed aerodynamic sound data fluctuates, and the amplitude of the aerodynamic sound represented by the processed aerodynamic sound data also fluctuates. Consequently, the listener L can hear the aerodynamic sound with such fluctuating amplitude, thus minimizing the sense of discomfort and providing a sense of presence.

[0495] In the sound signal processing method according to the first embodiment, in the processing step, the acquired sound data is divided into processing frames F of a predetermined time, and the sound data is processed for each of the divided processing frames F.

[0496] This realizes an acoustic signal processing method that reduces the load of computational processing.

[0497] In the acoustic signal processing method according to the first embodiment, in the processing step, a smoothing function is determined for each divided processing frame F so that the value of the smoothing function becomes 1.0 at the first and last times of the processing frame F.

[0498] This suppresses the generation of noise at the junction between the processing frame F and the next processing frame F.

[0499] In the acoustic signal processing method according to the first embodiment, in the processing step, parameters of a smoothing function are determined for each of the divided processing frames F.

[0500] This realizes an acoustic signal processing method that reduces the load of computational processing.

[0501] In the acoustic signal processing method according to the first embodiment, the parameter is the time from the first time to the last time.

[0502] In this way, the time from the first time of the processing frame F to the last time of the processing frame F can be set as a parameter.

[0503] In the acoustic signal processing method according to the first embodiment, the parameter is a value related to the maximum value of the smoothing function.

[0504] This allows the parameter to be set to a value related to the maximum value of the smoothing function.

[0505] In the acoustic signal processing method according to the first embodiment, the parameter is a parameter that changes the position at which the smoothing function reaches a maximum value.

[0506] This allows the parameters to be set so as to change the position where the smoothing function reaches its maximum value.

[0507] In the acoustic signal processing method according to the first embodiment, the parameter is a parameter that changes the sharpness of the change of the smoothing function.

[0508] This makes it possible to set a parameter that changes the sharpness of the change of the smoothing function.

[0509] In the acoustic signal processing method according to embodiment 1, in a processing step, a first parameter and a second parameter are determined to define a smoothing function; based on the smoothing function determined by the determined first parameter, acquired sound data is processed so as to change at least one of the frequency component, phase, and amplitude value of the waveform; based on the smoothing function determined by the determined second parameter, acquired sound data is processed so as to change at least one of the frequency component, phase, and amplitude value of the waveform; in an output step, the sound data processed based on the smoothing function determined by the determined first parameter is output to a first output channel; and the sound data processed based on the smoothing function determined by the determined second parameter is output to a second output channel.

[0510] This makes it possible to output different audio data for each output channel.

[0511] In the acoustic signal processing method according to the first embodiment, aerodynamic sound is the sound generated by the collision of wind W with an object; in the processing step, the properties of the wind speed of wind W are simulated to determine the parameters.

[0512] Parameters are determined by simulating the change in wind speed of the fluctuating wind W. The sound data can be processed based on the smoothing function determined by the parameters to change at least one of the frequency component, phase, and amplitude of the waveform.

[0513] In the acoustic signal processing method according to the first embodiment, aerodynamic sound is sound produced by the wind W hitting the ears of a listener who listens to the aerodynamic sound. In the processing step, the properties of the wind direction of the wind W are simulated to determine the parameters.

[0514] Parameters are determined by simulating the change in the direction of the wind W including the undulation. The sound data can be processed based on the smoothing function determined by the parameters so as to change at least one of the frequency component, phase, and amplitude of the waveform.

[0515] In the acoustic signal processing method according to the first embodiment, the maximum value of the smoothing function does not exceed 3.

[0516] This makes it possible to set the maximum value of the smoothing function to 3 or less.

[0517] In the acoustic signal processing method according to the first embodiment, the minimum value of the smoothing function is not less than zero.

[0518] This makes it possible to set the minimum value of the smoothing function to be equal to or greater than 0.

[0519] The sound signal processing method according to embodiment 1 includes an acceptance step in which an instruction for specifying Va as the wind speed of wind W and Vp as the instantaneous wind speed of wind W is accepted; in the processing step, a smoothing function is determined so that the maximum value of the smoothing function becomes Vp / Va.

[0520] Thereby, the maximum value of the smoothing function can be set to Vp / Va.

[0521] In the acoustic signal processing method according to the first embodiment, the average value of the predetermined time is 3 seconds.

[0522] Thereby, the average value of the predetermined time, which is the time length of the processing frame, can be set to 3 seconds.

[0523] In the acoustic signal processing method according to the first embodiment, the object has a shape imitating an ear.

[0524] This makes it possible to collect aerodynamic sound using, for example, a dummy head-shaped microphone or the like.

[0525] The computer program according to the first embodiment causes a computer to execute the above-described sound signal processing method.

[0526] Thereby, the computer can execute the above-mentioned sound signal processing method according to the computer program.

[0527] The sound signal processing device 100 according to embodiment 1 includes: an acquisition unit 110 for acquiring sound data representing a waveform of a reference sound; a processing unit 120 for processing the sound data based on simulation information that simulates changes in natural phenomena so as to change at least one of the frequency component, phase, and amplitude value of the waveform; and an output unit 130 for outputting the processed sound data.

[0528] Thus, based on simulation information that simulates the fluctuations of natural phenomena involving fluctuations, the sound data is processed to change at least one of the frequency component, phase, and amplitude of the waveform. Consequently, at least one of the frequency component, phase, and amplitude fluctuates in the processed sound data, and at least one of the frequency component, phase, and amplitude fluctuates in the sound represented by the processed sound data. Consequently, the listener L can hear the sound in which at least one of the frequency component, phase, and amplitude fluctuates, and the listener L is less likely to experience a sense of discomfort and can experience a sense of immersiveness. In other words, an audio signal processing device 100 capable of providing the listener L with a sense of immersiveness is achieved.

[0529] (Other embodiments)

[0530] While the acoustic signal processing method and acoustic signal processing device of the technical solution of the present disclosure have been described above based on the embodiments, the present disclosure is not limited to these embodiments and their variations. For example, other embodiments achieved by arbitrarily combining the components described in this specification or excluding certain components may also serve as embodiments of the present disclosure. Furthermore, variations of the above-described embodiments that are conceivable by those skilled in the art without departing from the spirit of the present disclosure, that is, without departing from the meaning of the claims, are also encompassed by the present disclosure.

[0531] In addition, the following aspects are also included in the scope of one or more technical solutions of the present disclosure.

[0532] (1) A portion of the components constituting the above-mentioned sound signal processing device may be a computer system composed of a microprocessor, ROM, RAM, a hard disk unit, a display unit, a keyboard, a mouse, etc. A computer program is stored in the RAM or the hard disk unit. The microprocessor operates according to the computer program to achieve its function. Here, the computer program is composed of a combination of multiple command codes representing instructions to the computer in order to achieve a predetermined function.

[0533] (2) A portion of the components that make up the above-mentioned audio signal processing device may be formed by a single system LSI (Large Scale Integration). A system LSI is a highly multifunctional LSI manufactured by integrating multiple components onto a single chip. Specifically, it is a computer system composed of a microprocessor, ROM, and RAM. The RAM stores a computer program. The microprocessor operates according to the computer program, and the system LSI achieves its functions.

[0534] (3) A portion of the components constituting the above-mentioned audio signal processing device may be constituted by an IC card or a single module that is detachable from each device. The above-mentioned IC card or module is a computer system composed of a microprocessor, ROM, RAM, etc. The above-mentioned IC card or module may also include the above-mentioned ultra-multifunctional LSI. The above-mentioned IC card or module achieves its function by the microprocessor operating according to the computer program. The IC card or module may also be tamper-resistant.

[0535] (4) Furthermore, a portion of the components constituting the above-mentioned audio signal processing device may be a product in which the above-mentioned computer program or the above-mentioned digital signal is recorded on a computer-readable recording medium, such as a floppy disk, hard disk, CD-ROM, MO, DVD, DVD-ROM, DVD-RAM, BD (Blu-ray (registered trademark) Disc), semiconductor memory, or the like. Furthermore, the digital signal recorded on such a recording medium may also be used.

[0536] Furthermore, a portion of the components constituting the sound signal processing device may be transmitted by transmitting the computer program or the digital signal via a telecommunication line, a wireless or wired communication line, a network such as the Internet, data broadcasting, or the like.

[0537] (5) The present disclosure may also be the methods shown above. In addition, the present disclosure may also be a computer program that implements these methods on a computer, or a digital signal composed of the computer program.

[0538] (6) Furthermore, the present disclosure may be a computer system including a microprocessor and a memory, wherein the memory stores the computer program and the microprocessor operates according to the computer program.

[0539] (7) In addition, the program or the digital signal may be recorded on the recording medium and transferred, or the program or the digital signal may be transferred via the network, etc., so that the program or the digital signal can be implemented by another independent computer system.

[0540] Industrial Applicability

[0541] The present disclosure can be utilized in an acoustic signal processing method and an acoustic signal processing device, and can be particularly applied to an acoustic system and the like.

[0542] Description of labels

[0543] 100, 100a Audio signal processing device

[0544] 110 Acquisition Department

[0545] 120, 120a processing unit

[0546] 121 Processing Department 1

[0547] 122, 122b Second processing unit

[0548] 130 Output unit

[0549] 140 Storage

[0550] 150 Reception Department

[0551] 200 headphones

[0552] 201 Head sensor unit

[0553] 202 Output

[0554] 300 Display Unit

[0555] 500 Server Installation

[0556] 600 Information Processing Device

[0557] 610 loop address unit

[0558] 620 Frequency Shift Unit

[0559] 630 Storage Department

[0560] 640 Section Designation Department

[0561] 650 Crossfade Department

[0562] 660 Readout control unit

[0563] 900 Rendering Department

[0564] 901 Reverberation Processing Unit

[0565] 902 Initial reflection processing unit

[0566] 903 Distance Attenuation Processing Unit

[0567] 904 Selection Department

[0568] 905 Binaural Processing Unit

[0569] 906 Computing Department

[0570] 907 Generation Department

[0571] 1001 Sampling rate conversion unit

[0572] 1002 Relocation Department

[0573] 1003 Connection

[0574] 1021 Upsampling Unit

[0575] 1022 Low-pass filter unit

[0576] 1023 downsampling unit

[0577] 1024 XY setting unit

[0578] 1031 Amplitude Adjustment Unit

[0579] A1 Section 1

[0580] A2 Section 2

[0581] A3 range

[0582] A0000 Stereo Sound Reproduction System

[0583] A0001 Audio signal processing device

[0584] A0002 Sound prompt device

[0585] A0100 Encoding Device

[0586] A0101 Input Data

[0587] A0102 encoder

[0588] A0103 Encoded Data

[0589] A0104 Memory

[0590] A0110 decoding device

[0591] A0111 Sound signal

[0592] A0112 decoder

[0593] A0113 Input Data

[0594] A0114 Memory

[0595] A0120 Encoding Device

[0596] A0121 Sending Department

[0597] A0122 Send signal

[0598] A0130 Decoding Device

[0599] A0131 Receiving Department

[0600] A0132 Receive signal

[0601] A0200 decoder

[0602] A0201 Spatial Information Management Department

[0603] A0202 Sound Data Decoder

[0604] A0203 Rendering Department

[0605] A0210 decoder

[0606] A0211 Spatial Information Management Department

[0607] A0213 Rendering Department

[0608] D1 Aerodynamic sound data before processing

[0609] D11 processed aerodynamic sound data

[0610] FN Fan

[0611] F, F1, F2, F3, F4, F5, F6, Fn, Fn+1 processing frame

[0612] L Listener

[0613] Pt1 first pointer

[0614] Pt2 second pointer

Claims

1. A method for processing an acoustic signal, wherein: include: An acquisition step of acquiring sound data representing a waveform of a reference sound; A processing step of processing the sound data based on simulation information simulating changes in natural phenomena so as to change at least one of the frequency component, phase and amplitude value of the waveform; as well as The output step is to output the processed sound data.

2. The sound signal processing method according to claim 1, wherein: The reference sound is a pneumatic sound generated by wind. In the processing step, the sound data is processed based on the simulation information simulating the change in wind speed of the wind so as to change the at least one of the frequency component, the phase, and the amplitude value of the waveform.

3. The sound signal processing method according to claim 2, wherein: In the processing step, As the simulation information, a smoothing function simulating the change in wind speed of the wind is determined, Based on the determined value represented by the smoothing function, the audio data is processed so that at least one of the frequency component, the phase, and the amplitude value of the waveform is changed.

4. The sound signal processing method according to claim 3, wherein: The value represented by the smoothing function is information representing a ratio of the wind speed of the aerodynamic sound as the reference sound to the wind speed of the aerodynamic sound represented by the sound data processed in the processing step.

5. The sound signal processing method according to claim 4, wherein: In the processing step, the smoothing function is determined so that a parameter defining the smoothing function changes irregularly.

6. The sound signal processing method according to claim 3, wherein: In the processing step, the audio data is processed so that the frequency component of the waveform is shifted to a frequency proportional to the value represented by the determined smoothing function.

7. The sound signal processing method according to claim 3, wherein: In the processing step, the audio data is processed so that the amplitude value of the waveform changes in proportion to the α-th power of the value represented by the determined smoothing function.

8. The sound signal processing method according to claim 4, wherein: In the processing step, the acquired audio data is divided into processing frames of a predetermined time, and the audio data is processed for each of the divided processing frames.

9. The sound signal processing method according to claim 8, wherein: In the processing step, the smoothing function is determined for each of the divided processing frames so that the value of the smoothing function becomes 1.0 at the first time and the last time of the processing frame.

10. The sound signal processing method according to claim 9, wherein: In the processing step, the parameters of the smoothing function are determined for each of the segmented processing frames.

11. The sound signal processing method according to claim 10, wherein: The parameter is the time from the initial moment to the final moment.

12. The sound signal processing method according to claim 10, wherein: The parameter is a value related to the maximum value of the smoothing function.

13. The sound signal processing method according to claim 10, wherein: The parameter is a parameter of the position change that causes the smoothing function to reach a maximum value.

14. The sound signal processing method according to claim 10, wherein: The parameter is a parameter that changes the sharpness of the change of the smooth function.

15. The sound signal processing method according to claim 10, wherein: In the processing step, Determining the first parameter and the second parameter of the smoothing function, Based on the smoothing function determined by the determined first parameter, the acquired sound data is processed so as to change at least one of the frequency component, phase and amplitude value of the waveform, Based on the smoothing function determined by the determined second parameter, the acquired sound data is processed so as to change at least one of the frequency component, phase and amplitude value of the waveform, In the output step, outputting the sound data processed based on the smoothing function determined by the determined first parameter to a first output channel, The audio data processed based on the smoothing function determined by the determined second parameter is output to a second output channel.

16. The sound signal processing method according to claim 10, wherein: The aerodynamic sound is the sound produced by the collision of the wind with an object. In the processing step, the properties of the wind speed are simulated to determine the parameters.

17. The sound signal processing method according to claim 10, wherein: The aerodynamic sound is a sound generated by the wind colliding with the ears of a listener who listens to the aerodynamic sound. In the processing step, the properties of the wind direction are simulated to determine the parameters.

18. The sound signal processing method according to claim 8, wherein: The maximum value of the smoothing function does not exceed 3.

19. The sound signal processing method according to claim 8, wherein: The minimum value of the smoothing function is not less than 0.

20. The sound signal processing method according to claim 8, wherein: The method comprises an accepting step of accepting an instruction to designate Va as the wind speed of the wind and Vp as the instantaneous wind speed of the wind, In the processing step, the smoothing function is determined so that the maximum value of the smoothing function becomes Vp / Va.

21. The sound signal processing method according to claim 8, wherein: The average value of the prescribed time is 3 seconds.

22. The sound signal processing method according to claim 16, wherein: The object is an object having a shape imitating an ear.

23. A computer program for causing a computer to execute the sound signal processing method according to any one of claims 1 to 22.

24. An audio signal processing device, wherein: have: an acquisition unit that acquires audio data representing a waveform of a reference sound; a processing unit that processes the sound data based on simulation information that simulates changes in natural phenomena so as to change at least one of a frequency component, a phase, and an amplitude value of the waveform; as well as The output unit outputs the processed sound data.

Citation Information

Patent Citations

  • Apparatus and method for rendering a sound scene using pipeline stages

    WO2021180938A1