Acoustic processing method, acoustic processing device and acoustic processing program

DE112023005115T5Pending Publication Date: 2025-10-16SONY GROUP CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
DE112023005115
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

Conventional methods for simulating realistic acoustic spaces in virtual environments face challenges in balancing realism with computational load, as more realistic simulations increase processing demands, and existing techniques either require high computational resources or compromise on physical phenomena to achieve real-time performance.

Method used

A sound processing method that uses adaptive rectangular decomposition to focus wave acoustic simulations only on the area surrounding sound reflection points, reducing the computational load by limiting calculations to specific regions and employing impulse response generation and convolution to enhance tone quality and accuracy of sound propagation.

Benefits of technology

This approach allows for the creation of realistic acoustic spaces with reduced computational requirements, accurately reproducing scattering phenomena and improving tone quality while maintaining efficient processing, thereby enhancing the auditory experience in virtual environments without excessive resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present acoustic processing method includes a search step for searching for a path of a sound from a sound source to a sound receiving point in a virtual space, an identification step for identifying a reflection position of the sound on the path, and a processing step for performing processing on the sound at the sound receiving point based on a result of a wave motion acoustic simulation regarding the reflection of the sound with a surrounding area of ​​the reflection position as a calculation area.
Need to check novelty before this filing date? Find Prior Art

Description

Acoustic processing method, acoustic processing device, and acoustic processing program

[0001] The present disclosure relates to a sound processing method, a sound processing device, and a sound processing program.

[0002] Technologies for simulating sound in virtual spaces are being actively developed, including wave acoustic simulation and geometric acoustic simulation.

[0003] JP 2000-267675 A JP 2019-165845 A

[0004] “Efficient and Accurate Sound Propagation Using Adaptive Rectangular Decomposition” N.Raghuvanshi et al., IEEE(2009) “Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations” M. Raissi, P. Perdikaris, GE Karniadakis, Journal of Computational Physics 378(2019)

[0005] According to conventional techniques, it is possible to realize a realistic acoustic space in a virtual space. However, the more realistic the acoustic space that is realized, the greater the computational load becomes.

[0006] Therefore, the present disclosure proposes a sound processing method, a sound processing device, and a sound processing program that can realize a realistic sound space with a small calculation load.

[0007] It should be noted that the above problem or object is merely one of multiple problems or objects that can be solved or achieved by multiple embodiments disclosed in this specification.

[0008] In order to solve the above problems, one form of acoustic processing method according to the present disclosure includes a search step of searching for a sound path from a sound source to a sound receiving point in a virtual space, an identification step of identifying a reflection point of the sound on the path, and a processing step of performing processing related to the sound at the sound receiving point based on the results of a wave acoustic simulation related to the reflection of the sound, with the surrounding area of ​​the reflection point as the calculation area.

[0009] 1 is a diagram illustrating an overview of information processing of this embodiment. FIG. 1 is a diagram for explaining an overview of acoustic processing of this embodiment. FIG. 1 is a diagram for explaining an overview of acoustic processing of this embodiment. FIG. 2 is a diagram illustrating an example of the configuration of a sound processing device according to this embodiment. A flowchart showing audio signal generation processing of this embodiment. FIG. 2 is a diagram for explaining acoustic simulation in a virtual space. FIG. 3 is a diagram for explaining acoustic simulation in a virtual space. FIG. 4 is a diagram schematically showing the state of reflection in acoustic simulation. FIG. 4 is a diagram showing an example of a synthesized and output audio signal. FIG. 5 is a flowchart showing an example of processing when there are multiple initial reflected sounds, diffracted sounds, and transmitted sounds. A flowchart showing reflected sound generation processing of this embodiment. A flowchart showing impulse response generation processing of this embodiment. FIG. 5 is a diagram illustrating an example of setting a wave calculation region. FIG. 6 is a diagram illustrating an example of a two-dimensional wave calculation region. FIG. 6 is a diagram illustrating wave calculation of this embodiment. FIG. 7 is a diagram illustrating an example of a wave front calculation. FIG. 7 is a diagram illustrating an example of an impulse response calculation. FIG. 8 is a diagram illustrating an example of a response in free space. FIG. 9 is a diagram illustrating an input waveform. FIG. 10 is a diagram illustrating an example of removal of unnecessary components. FIG. 11 is a diagram illustrating the effect of this embodiment. FIG. 12 is a diagram illustrating a state in which a virtual sound source and a virtual sound receiving point are arranged inside a surrounding region. FIG. 2 is a hardware configuration diagram illustrating an example of a computer that realizes the functions of the sound processing device.

[0010] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the following embodiments, the same components are designated by the same reference numerals, and redundant description will be omitted.

[0011] One or more embodiments (including examples and modified examples) described below can be implemented independently. However, at least a portion of the embodiments described below may be implemented in appropriate combination with at least a portion of another embodiment. These embodiments may include novel features that are different from one another. Therefore, these embodiments may contribute to solving different purposes or problems and may produce different effects.

[0012] The present disclosure will be described in the following order: 1. Overview 2. Configuration of sound processing device 3. Operation of sound processing device 3-1. Audio signal generation processing 3-2. Reflected sound generation processing 3-3. Impulse response generation processing 4. Effects 5. Modifications 6. Hardware configuration example 7. Conclusion

[0013] <<1. Overview>> Content using virtual spaces, such as the Metaverse, games, movies, and animation, is being actively produced. As virtual space technology develops, there is a growing demand for sounds reproduced in virtual spaces to be realistic and in line with reality. In research into auditory recognition of sound-producing bodies in three-dimensional virtual spaces, there are many methods for simulating the sounds they produce. Among these, spatial sound field reproduction technology using physical simulation is particularly well-designed based on physical phenomena in real space, and therefore has a high level of auditory reproducibility when reproducing sounds in virtual spaces. Physical simulation also demonstrates excellent performance when reproducing sounds that do not actually exist.

[0014] Physical simulation techniques in the field of acoustics include, for example, wave acoustic simulation, which models and calculates the wave properties of sound, and geometric acoustic simulation, which geometrically models and calculates the energy propagation of sound.

[0015] Wave acoustic simulation (also called wave propagation simulation) is known for its ability to accurately represent microscopic wave phenomena, and is particularly advantageous in representing scattering phenomena that occur when sound waves collide with objects with minute surface shapes. However, this method discretizes space for calculations, and the fineness of the discretization affects the frequencies that can be calculated. In other words, when using wave acoustic simulation, the amount of calculation increases as the sound band becomes higher, which requires finer discretization. This makes wave acoustic simulation difficult to use in applications that require real-time performance.

[0016] Known geometric acoustic simulation methods include the ray tracing method, which tracks the trajectory of sound rays emitted from a sound source, and the image method, which assumes sound propagation due to geometric reflections at boundaries. Geometric acoustic simulation does not require spatial division, so there is little correlation between bandwidth and computational effort. Therefore, geometric acoustic simulation has a relatively light computational load, and can be expected to achieve high-speed processing in the real-time processing required for games and other applications. However, geometric acoustic simulation omits the representation of some physical phenomena during modeling. Therefore, when using such methods, the reproduced sound may sound unnatural to the ear.

[0017] That is, with conventional technology, it is difficult to realize a realistic acoustic space with a small calculation load.

[0018] Therefore, in this embodiment, the above problem is solved by the following method.

[0019] Fig. 1 is a diagram showing an overview of information processing according to this embodiment. The information processing according to this embodiment is executed by a sound processing device 100 shown in Fig. 1. The sound processing device 100 is an example of the sound processing device according to this embodiment. The sound processing device 100 is an information processing device such as a personal computer (PC), a server device, a tablet terminal, or a game console. The sound processing device 100 is used by a user U who uses / creates content related to a virtual space (for example, a game or a metaverse).

[0020] The sound processing device 100 has output units such as a display and a speaker, and outputs various information to the user U. For example, the sound processing device 100 displays a virtual three-dimensional space (hereinafter simply referred to as a "virtual space") of a game or the like on the display, and outputs an audio signal generated by the information processing of this embodiment from the speaker. Alternatively, the sound processing device 100 displays a user interface of software related to sound production on the display, and outputs an audio signal generated in accordance with an operation of the user U from the speaker.

[0021] In this embodiment, the sound processing device 100 calculates how a sound output from a sound source object, which is a sound emission point, will be reproduced at a listening point (sound receiving point) in a virtual space, and reproduces the calculated sound. For example, the sound processing device 100 performs an acoustic simulation in the virtual space, and performs processing to make sounds emitted in the virtual space closer to those in the real world or to reproduce the reverberation desired by the sound producer.

[0022] FIG. 1 shows a virtual space V1 in a game. For example, the virtual space V1 is displayed on a display provided in the sound processing device 100. In the virtual space V1, a listening point (sound receiving point) is set along with the position (coordinates) of a sound source object (sound generating point). The listening point is the position where a user U virtually hears sound in the virtual space V1. In real space, various physical phenomena cause a difference between the sound observed near the sound source and the sound observed at the listening point. Therefore, the sound processing device 100 virtually reproduces (simulates) real physical phenomena in the virtual space V1 and generates an audio signal appropriate for the space to enhance the realism of the sound experienced by the user U in the virtual space V1.

[0023] Here, the audio signal generated by the sound processing device 100 will be described. Graph G1 schematically shows the sound intensity when a pulsed sound emitted from a sound source object is observed at a listening point. At the listening point, the direct sound is observed first, followed by diffracted sounds of the direct sound, etc. Then, at the listening point, the first-order reflected sound reflected at the boundary of the virtual space and the transmitted sound that has passed through the object are observed. Reflected sounds are observed every time the sound reflects at a boundary, and for example, first- to third-order reflected sounds, which are considered to be early reflected sounds, are observed. Then, at the listening point, higher-order reflected sounds, which are considered to be late reverberation sounds, are observed. Because the sound emitted from the sound source decays over time, graph G1 depicts an envelope (decay curve) that asymptotically approaches 0, with the direct sound at its peak.

[0024] In this embodiment, we focus on reflected sound that occurs when a sound wave collides with an object such as a wall. The sound processing device 100 calculates microscopic wave phenomena that occur during reflection (for example, phenomena such as minute scattering and interference that occur near the point where reflection occurs) using wave sound simulation. In this case, the sound processing device 100 performs the wave sound simulation only in the area surrounding the reflection point. This makes it possible to realize a realistic acoustic space with a small calculation load.

[0025] This processing will be explained below using the drawings. Figures 2 and 3 are diagrams for explaining an overview of the sound processing of this embodiment. Figures 2 and 3 show a virtual space V1 in which a sound source and a listening point are set. In the following explanation, the sound source placed in the virtual space is referred to as a sound-emitting point S. The listening point placed in the virtual space is referred to as a sound-receiving point T. The sound processing device 100 searches for a sound path from the sound-emitting point S to the sound-receiving point T in the virtual space V1. Then, the sound processing device 100 simulates the propagation of sound along each path.

[0026] First, a simulation of a path in which one reflection occurs will be described. Fig. 2 shows a path in which one reflection occurs among the sound paths from the sound-emitting point S to the sound-receiving point T. The sound processing device 100 identifies a reflection point P1 of the sound on the path. The sound processing device 100 then extracts a surrounding area A1 of the reflection point P1 from the virtual space V1. The surrounding area A1 is a hemispherical area on the sound source side (the side of the sound-emitting point S) of two areas obtained by dividing a sphere centered on the reflection point P1 by a sound-reflecting surface (a wall surface in the example of Fig. 2).

[0027] The sound processing device 100 then performs processing related to sound reflection (e.g., wave sound simulation) using the surrounding area A1 as a calculation area. Specifically, the sound processing device 100 places a virtual sound emission point S1 at a point on the spherical surface of the surrounding area A1 that is closest to the sound emission point S. The sound processing device 100 also places a virtual sound receiving point T1 at a point on the spherical surface of the surrounding area A1 that is closest to the sound receiving point T. The sound processing device 100 then performs wave sound simulation limited to the surrounding area A1. For example, the sound processing device 100 performs a propagation simulation of sound that reaches the sound receiving point T1 from the sound emission point S1. As a result, the sound processing device 100 acquires the waveform of the sound that reaches the sound receiving point T1 from the sound emission point S1. The sound processing device 100 calculates the virtual output sound from the sound emission point S1 used in the simulation by performing a predetermined process on the output sound from the sound emission point S. For example, the sound processing device 100 may calculate a virtual output sound from the sound generation point S1 by performing an adjustment process (e.g., gain adjustment and delay adjustment) on the output sound from the sound generation point S based on the attenuation / time delay caused by the propagation distance between the sound generation point S and the sound generation point S1 (or the time delay from the sound generation point S to the reflection point P1, and the attenuation from the sound generation point S to the sound generation point S1).

[0028] Then, the sound processing device 100 calculates the sound propagating to the sound receiving point T based on the calculation result of the sound propagating to the sound receiving point T1. For example, the sound processing device 100 calculates the waveform of the sound propagating to the sound receiving point T by performing gain adjustment and delay adjustment on the output sound from the sound receiving point T1 based on the propagation distance between the sound receiving point T1 and the sound receiving point T. Alternatively, the sound processing device 100 may calculate the waveform of the sound propagating to the sound receiving point T by performing delay adjustment based on the time delay from the reflection point P1 to the sound receiving point T, and gain adjustment based on attenuation from the sound receiving point T1 to the sound receiving point T.

[0029] Next, a simulation of a path on which multiple reflections occur will be described. Fig. 3 shows a sound path from a sound emission point S to a sound receiving point T on which two reflections occur. The sound processing device 100 identifies sound reflection points P2 and P3 on the path. The sound processing device 100 then extracts a surrounding area A2 of the reflection point P2 from the virtual space V1. The sound processing device 100 also extracts a surrounding area A3 of the reflection point P3 from the virtual space V1. The surrounding areas A2 and A3 are hemispherical areas on the sound source side (the side of the sound emission point S) of two areas obtained by dividing a sphere centered on the reflection point P2 or P3 by a sound reflection surface (a wall surface in the example of Fig. 2 ).

[0030] The sound processing device 100 then calculates sound reflection using the surrounding area A2 as a calculation area. Specifically, the sound processing device 100 places a tentative sound emission point S2 at a point on the spherical surface of the surrounding area A2 that is closest to the sound emission point S. The sound processing device 100 also places a tentative sound receiving point T2 at a point on the spherical surface of the surrounding area A2 that is closest to the next reflection point P3. The sound processing device 100 then performs a wave sound simulation (a simulation of sound propagating from the sound emission point S2 to the sound receiving point T2) limited to the surrounding area A2. At this time, the sound processing device 100 calculates a tentative output sound from the sound emission point S2 to be used in the simulation by performing predetermined processing on the output sound from the sound emission point S. For example, the sound processing device 100 may calculate the tentative output sound from the sound emission point S2 by performing adjustment processing (e.g., gain adjustment and delay adjustment) on the output sound from the sound emission point S based on the attenuation / time delay caused by the propagation distance between the sound emission points S and S2. Alternatively, the sound processing device 100 may calculate a virtual output sound from the sound emission point S2 by performing delay adjustment based on the time delay between the sound emission point S and the reflection point P2, and gain adjustment based on the attenuation between the sound emission point S and the sound emission point S2.

[0031] Furthermore, the sound processing device 100 calculates sound reflection using the surrounding area A3 as a calculation area. Specifically, the sound processing device 100 places a tentative sound emission point S3 at a point on the spherical surface of the surrounding area A3 that is closest to the immediately preceding reflection point P2. The sound processing device 100 also places a tentative sound receiving point T3 at a point on the spherical surface of the surrounding area A3 that is closest to the sound receiving point T. The sound processing device 100 then performs a wave sound simulation (a simulation of sound propagating from the sound emission point S3 to the sound receiving point T3) limited to the surrounding area A3. At this time, the sound processing device 100 calculates a tentative output sound from the sound emission point S3 used in the simulation by performing predetermined processing on the output sound from the sound receiving point T2. For example, the sound processing device 100 may calculate the tentative output sound from the sound emission point S3 by performing adjustment processing (e.g., gain adjustment and delay adjustment) on the output sound from the sound emission point S based on the attenuation / time delay caused by the propagation distance between the sound receiving point T2 and the sound emission point S3. Alternatively, the sound processing device 100 may calculate a virtual output sound from the sound emission point S3 by performing delay adjustment based on the time delay between the reflection point P2 (or the sound receiving point T2) and the reflection point P3, and gain adjustment based on the attenuation from the sound receiving point T2 to the sound emission point S3.

[0032] Then, the sound processing device 100 calculates the sound propagating to the sound receiving point T based on the calculation result of the sound propagating to the sound receiving point T3. For example, the sound processing device 100 calculates the waveform of the sound propagating to the sound receiving point T by performing gain adjustment and delay adjustment on the output sound from the sound receiving point T3 based on the propagation distance between the sound receiving point T3 and the sound receiving point T. Alternatively, the sound processing device 100 may calculate the waveform of the sound propagating to the sound receiving point T by performing delay adjustment based on the time delay between the reflection point P3 (or the sound receiving point T3) and the sound receiving point T, and gain adjustment based on the attenuation from the sound receiving point T3 to the sound receiving point T.

[0033] The sound processing device 100 calculates the sound heard at the sound receiving point T by synthesizing waveforms of multiple sounds that propagate to the sound receiving point T via multiple paths.

[0034] According to this embodiment, the sound processing device 100 performs a simulation of sound propagation limited to the area around the reflection point, and calculates the sound propagating to the listening point (sound receiving point) based on the simulation results. Therefore, the sound processing device 100 can realize a realistic acoustic space with a small calculation load.

[0035] The outline of this embodiment has been described above, and the sound processing device 100 according to this embodiment will now be described in detail.

[0036] In addition, the term "acoustic processing" that appears in the following description can be replaced with "acoustic signal processing," "audio processing," "audio signal processing," "sound processing," "sound signal processing," "speech / voice processing," or "speech / voice signal processing."

[0037] In the following description, the term "audio signal" may be replaced with "acoustic signal," "sound signal," or "speech / voice signal." Similarly, the term "audio data" may be replaced with "acoustic data," "sound data," or "speech / voice data."

[0038] <<2. Configuration of Sound Processing Device>> First, the configuration of the sound processing device 100 will be described in detail with reference to the drawings.

[0039] The sound processing device 100 is an information processing device (computer) that performs processing related to content relating to virtual space (for example, games, the Metaverse, etc.) For example, the sound processing device 100 is a user terminal used for games, the Metaverse, etc.

[0040] The sound processing device 100 is typically a game console (e.g., a dedicated game console), but is not limited to a game console. Any type of computer can be used for the sound processing device 100. For example, the sound processing device 100 may be a mobile terminal such as a mobile phone, a smart device (smartphone or tablet), a PDA (Personal Digital Assistant), or a notebook PC. The sound processing device 100 may also be an imaging device or a car navigation device. The sound processing device 100 may also be an M2M (Machine to Machine) device or an IoT (Internet of Things) device. The sound processing device 100 may also be a wearable device such as a smart watch. As long as a user can play a game on the device, these devices can also be considered game consoles (general-purpose game consoles).

[0041] The sound processing device 100 may also be an xR device such as an AR (Augmented Reality) device, a VR (Virtual Reality) device, or an MR (Mixed Reality) device. In this case, the xR device may be a glasses-type device such as AR glasses or MR glasses, or a head-mounted device such as a VR head-mounted display. When the sound processing device 100 is an xR device, the sound processing device 100 may be a standalone device consisting only of a user-worn part (e.g., glasses). The sound processing device 100 may also be a terminal-linked device consisting of a user-worn part (e.g., glasses) and a terminal part (e.g., a smart device) linked to the user-worn part. In this case, the sound processing device 100 may also be an information processing device (e.g., a game console) connected to a user-worn part (e.g., a head-mounted display) via a wired or wireless connection. If a user can play a game on the device, these devices can also be considered game consoles (general-purpose game consoles).

[0042] The sound processing device 100 may also be an information processing device connected to a user terminal and transmitting the results of processing related to the virtual space to the user terminal. For example, the sound processing device 100 may be a server device connected to a user terminal (e.g., a game console) via a network. In this case, the sound processing device 100 may be an application server or a web server. The sound processing device 100 may also be a PC server, a mid-range server, or a mainframe server. The sound processing device 100 may also be an information processing device that performs data processing (edge ​​processing) near the user terminal. For example, the sound processing device 100 may be an information processing device attached to or built into a base station. Of course, the sound processing device 100 may also be an information processing device that performs cloud computing.

[0043] Fig. 4 is a diagram showing an example of the configuration of a sound processing device 100 according to this embodiment. As shown in Fig. 4, the sound processing device 100 includes a communication unit 110, a storage unit 120, a control unit 130, an input unit 140, and an output unit 150. Note that the configuration shown in Fig. 4 is a functional configuration, and the hardware configuration may be different from this. Furthermore, the functions of the sound processing device 100 may be statically or dynamically distributed and implemented in multiple physically separated configurations. For example, the sound processing device 100 may be configured by multiple information processing devices (computers) connected via communication.

[0044] The communication unit 110 is a communication interface for communicating with other devices. The communication unit 110 may be a network interface or a device connection interface. For example, the communication unit 110 may be a LAN (Local Area Network) interface such as a NIC (Network Interface Card), or a USB (Universal Serial Bus) interface configured by a USB host controller, a USB port, etc. The communication unit 110 may be a wired interface or a wireless interface. The communication unit 110 exchanges information with other information devices, etc. via the network N.

[0045] Here, the network N is, for example, a communication network such as a local area network (LAN), a wide area network (WAN), a cellular network, a fixed telephone network, a regional Internet Protocol (IP) network, or the Internet. The network N may include a wired network or a wireless network. The network N may also include a core network. The core network is, for example, an Evolved Packet Core (EPC) or a 5G Core network (5GC). Of course, the network N may also be a data network connected to the core network. The data network may be a service network of a telecommunications carrier, for example, an IP Multimedia Subsystem (IMS) network. The data network may also be a private network such as an in-house network or a home network.

[0046] The storage unit 120 is a storage device that stores various types of information. For example, the storage unit 120 is a data readable / writable storage device such as a dynamic random access memory (DRAM), a static random access memory (SRAM), a semiconductor memory (e.g., a flash memory), or a hard disk. The storage unit 120 may be an optical drive such as a Blu-ray (registered trademark) drive, a DVD drive, or a CD drive. The storage unit 120 stores information for processing related to virtual space (hereinafter referred to as virtual space information). For example, the storage unit 120 stores sound source information, object information, sound receiving point information, and setting information as virtual space information.

[0047] Here, the sound source information is, for example, position information of the sound emission point S, data of the sound output from the sound emission point S (e.g., waveform data), and information on the directivity of the sound output from the sound emission point S. The sound source information may also include information on the type of sound. If the sound is dialogue, the sound source information may include data on the character emitting the sound, as well as metadata such as angry voices and laughter. The object information is, for example, information on the position, shape, material, sound absorption coefficient, and acoustic impedance of the object. The sound receiving point information is, for example, information such as the position of the sound receiving point T, directional characteristics, and HRTF (Head Related Transfer Function) which represents the characteristics from the sound source S to both ears of the listener. The positions of the sound emission point S and the sound receiving point T may be expressed using a Cartesian coordinate system, a cylindrical coordinate system, or a spherical coordinate system. The setting information is information on a playback device that plays content related to a virtual space, information on a platform for creating the content, and specific scene information within the content. The virtual space information may also include, for example, CAD data of a virtual space V1 arbitrarily created by a creator, as shown in Figure 1, or CAD data of a virtual space V1 in a scene of a game in which a user U is operating a character in the game, etc. In addition to CAD data, the virtual space information may also include voxel data, mesh data, point cloud data, etc.

[0048] Note that the information stored in the storage unit 120 is not limited to these pieces of information. The virtual space information stored in the storage unit 120 may also include setting information related to sound settings (for example, setting information related to diffracted sound, reflected sound, transmitted sound, and late reverberation sound). The virtual space information may be input to the storage unit 120 via the communication unit 110 or the input unit 140.

[0049] The control unit 130 is a controller that controls each unit of the sound processing device 100. The control unit 130 is realized by a processor such as a CPU (Central Processing Unit), MPU (Micro Processing Unit), GPU (Graphics Processing Unit), APU (Accelerated Processing Unit), or DSP (Digital Signal Processor). For example, the control unit 130 is realized by executing various programs (e.g., sound processing programs according to the present disclosure) stored in a storage device inside the sound processing device 100 using RAM (Random Access Memory) or the like as a working area. The CPU, MPU, GPU, APU, ASIC, and FPGA can all be considered as controllers.

[0050] As shown in FIG. 4 , the control unit 130 includes an acquisition unit 131, a search unit 132, an identification unit 133, a signal processing unit 134, and an output control unit 135. Each block (acquisition unit 131 to output control unit 135) constituting the control unit 130 is a functional block that indicates a function of the control unit 130. These functional blocks may be software blocks or hardware blocks. For example, each of the above-described functional blocks may be a software module realized by software (including a microprogram), or may be a circuit block on a semiconductor chip (die). Of course, each functional block may be a processor or an integrated circuit. The control unit 130 may be configured as a functional unit different from the above-described functional blocks. The method of configuring the functional blocks is arbitrary. The operation of each block constituting the control unit 130 will be described later.

[0051] The input unit 140 is an input device that accepts various inputs from the outside. For example, the input unit 140 is a data input interface for inputting various data related to a network. In this case, the data input interface may be a device connection interface such as a USB interface. For example, the input unit 140 may be an operation device for a user to perform various operations, such as a keyboard, a pointing device (e.g., a mouse), or operation keys. The input unit 140 may be a microphone for a user to input sound. The input unit 140 may also be a camera for a user to input a line of sight or a gesture. If a touch panel is adopted as an input / output device for the sound processing device 100, the touch panel is also included in the input unit 140. In this case, the user performs various operations by touching the screen with a finger or a stylus.

[0052] The output unit 150 is a device that outputs various types of information to the outside, such as sound, light, vibration, and images. For example, the output unit 150 may be an acoustic device such as a speaker, or a display device such as a display. The display device may be, for example, a liquid crystal display or an organic electroluminescence display (OLED). The output unit 150 may be a touch panel display device. In this case, the input unit 140 and the output unit 150 may be considered to be an integrated configuration. The output unit 150 may also be an output unit of an xR device.

[0053] The output unit 150 performs various outputs to the user under the control of the control unit 130. For example, the output unit 150 outputs synthesized sound from headphones, earphones, speakers, a head mounted display (HMD), or the like under the control of the output control unit 135. Alternatively, the output unit 150 displays a virtual space or a user interface on a display under the control of the output control unit 135.

[0054] <<3. Operation of Sound Processing Device>> Having described the configuration of the sound processing device 100 above, the operation of the sound processing device 100 will now be described. Hereinafter, the information processing executed by the sound processing device 100 will be described in detail with reference to the drawings. First, a general process for generating an audio signal in a virtual space will be described with reference to FIGS. Then, the sound processing according to this embodiment will be described in detail with reference to FIG. 10 and subsequent figures.

[0055] <3-1. Audio Signal Generation Processing> First, a general process for generating an audio signal in a virtual space will be described.

[0056] 5 is a flowchart showing the audio signal generation process of this embodiment. The audio signal generation process is initiated by an event in the virtual three-dimensional space (for example, an event that occurs as a result of an operation by the user U / the progress of the game).

[0057] In the following description, the sound processing device 100 is assumed to execute the audio signal generation process, but the audio signal generation process may be executed by a single computer or by multiple computers working together. For example, one computer (e.g., a server device) among multiple computers (e.g., a user terminal and a server device) connected via a network may execute some steps of the audio signal generation process, and another computer (e.g., a user terminal) may execute the remaining steps.

[0058] In the following description, a scene in a game is assumed as an example. Specifically, a virtual space V2 shown in FIGS. 6A and 6B is assumed. FIGS. 6A and 6B are diagrams for explaining acoustic simulation in a virtual space. The virtual space V2 is a virtual three-dimensional space for a game or the like. FIG. 6A is a perspective view of the virtual space V2, and FIG. 6B is a plan view of the virtual space V2. In the following description, an XYZ coordinate system may be used for ease of understanding. In the drawings, the X-axis and Y-axis directions are horizontal, and the Z-axis direction is vertical.

[0059] In the virtual space V2, a sound emission point S, which is a position that serves as a sound source, and a sound receiving point T, which is a position that serves as a listening point, are set. The sound emission point S is the position of an object that emits any sound toward the sound receiving point T (i.e., an object that can serve as a sound source). For example, the sound emission point S is the coordinates where the sound source object is located. Furthermore, the sound receiving point T is the position where the sound output from the sound emission point S is observed. For example, the sound receiving point T is the position where the game character operated by the user U is located (more precisely, the coordinates of the position corresponding to the head of the game character).

[0060] Furthermore, a plurality of three-dimensional objects are placed in the virtual space V2. For example, an object J2, which is a furniture-type object, and an object J3, which is a human-type object, are placed in the virtual space V2. Also placed in the virtual space V2 are walls J1 and the like, which serve as boundaries that define the virtual space V2. Note that acoustic characteristics, as will be described later, are also set for boundaries such as the wall J1, and therefore in this embodiment, boundaries such as the wall J1 are also treated as one of the virtual objects placed in the virtual space V2.

[0061] 6A and 6B , sound is generated at the position of a sound emission point S. The sound processing device 100 generates an audio signal at a sound receiving point T while taking into consideration the propagation characteristics of sound in the virtual space V2. The audio signal generation process of this embodiment will be described below with reference to the flowchart of FIG.

[0062] First, the sound processing device 100 acquires audio data of a sound emitted from a sound-emitting point S (step S101). When a vocalization event at the sound-emitting point is initiated, the sound processing device 100 acquires pre-recorded audio data of the sound to be emitted. The audio data is, for example, audio data recorded in an in-game library. Considering that the audio data will be subjected to subsequent signal processing, data that does not include reverberation (dry source) is preferable. Alternatively, the audio data may be a simplified sound presentation that omits some of the subsequent signal processing, or a wet source that adds reverberation to express a distinctive tone. Furthermore, the sound processing device 100 may not only reference audio data from a library, but also acquire audio data generated on demand, such as by a synthesizer. Alternatively, the sound processing device 100 may acquire audio data generated on demand by simulating audio production from a structural physical simulation using FEM (Finite Element Method) or the like.

[0063] Next, the sound processing device 100 acquires spatial data of the space in which the sound emission point S and the sound receiving point T exist (step S102). For example, the sound processing device 100 acquires three-dimensional data of the virtual space V2 and metadata (e.g., material information) attached to the three-dimensional data. The three-dimensional data includes, for example, coordinates indicating the boundary shapes of the walls of the virtual space V2 and the coordinates of objects placed within the virtual space V2. If the sound emission point S and the sound receiving point T exist in a closed space, the sound processing device 100 acquires three-dimensional data of the entire closed space, or three-dimensional data of the space on the path from the sound emission point S to the sound receiving point T and the entire space connected to that path. Note that, since the amount of data increases when the three-dimensional data includes detailed textures for video display, the sound processing device 100 may acquire data in which color information, fine surface shapes, etc. are simplified in a form that maintains the acoustic parameters required for subsequent physical simulation. It is preferable that the three-dimensional data used in the path search described below be polygon information that does not include fine surface irregularities (e.g., polygon information without texture).

[0064] Next, the sound processing device 100 calculates the sound path from the sound emission point S to the sound receiving point T (step S103). For example, the sound processing device 100 calculates the sound propagation path from the sound emission point S to the sound receiving point T using the acquired three-dimensional data. At this time, the path of the direct sound is a line-of-sight path from the sound emission point S to the sound receiving point T. Note that if a transmission phenomenon such as a wall is included, the path including the obstacle is calculated even if there is no line-of-sight path. In the examples of Figures 6A and 6B, the propagation path of the direct sound from the sound emission point S to the sound receiving point T is shown by a solid line.

[0065] Furthermore, the sound processing device 100 calculates a path that includes reflection boundaries and the like along which the early reflected sound reaches the sound receiving point T. For example, the sound processing device 100 determines the propagation path of the early reflected sound using a ray tracing method, which is a geometric physical simulation method. Note that in this embodiment, the early reflected sound includes not only sound that reaches the sound receiving point T after reflecting once on a boundary surface, but also sound that reaches the sound receiving point T after reflecting a small number of times, such as two or three times.

[0066] FIG. 7 shows an example of propagation paths calculated by the sound processing device 100. FIG. 7 is a diagram schematically illustrating reflections in a sound simulation. In the example shown in FIG. 7, path R1 is the propagation path of a direct sound from a sound emission point S to a sound receiving point T. Path R2 is the propagation path of a primary reflected sound that reflects once at a boundary B1 and reaches the sound receiving point T. Path R3 is the propagation path of a secondary reflected sound that reflects twice at boundaries B2 and B3 and reaches the sound receiving point T. In FIG. 7, paths R1 to R3 are shown as examples of propagation paths. However, actual calculations also include reflections from floors, ceilings, other wall surfaces, object surfaces, and the like.

[0067] Diffracted sound reaching the sound receiving point T from the sound emission point S can be obtained not only by the shortest path from the sound emission point S to the sound receiving point T, but also by a path that detours around the end of an obstacle located between the sound emission point S and the sound receiving point T. In the examples of Figures 6A and 6B, object J3 is an example of an obstacle located between the sound emission point S and the sound receiving point T. In this case, the sound processing device 100 can, for example, obtain a detour route that detours around the obstacle in the shape of an arc, or can obtain the detour route by connecting the sound emission point S and the end of the object J3, and the end of the object J3 and the sound receiving point T, with a straight line.

[0068] When directivity is set for the sound source, the sound processing device 100 may omit sound rays in directions where there is no sound radiation according to the directivity, or may change the density of sound rays according to the strength of radiation. By performing this processing, it is possible to reduce the impact on auditory sensation while reducing subsequent processing performed for each path.

[0069] Based on the obtained information, the sound processing device 100 generates an audio signal to be heard at the sound receiving point T. For example, the sound processing device 100 performs generation processes such as direct sound generation, early reflected sound generation, diffracted sound generation, and late reverberation sound generation. These generation processes will be described below.

[0070] First, the sound processing device 100 generates a sound signal based on a direct sound from among the sound signals heard at the sound receiving point T (step S104). Specifically, the sound processing device 100 receives the acquired sound data as an input, and generates a sound signal using parameters such as the spatial size of the sound source, directivity, attenuation according to the distance from the sound source to the listening point, and a deviation in observation time due to propagation time. For example, if the sound source, that is, the sound source, is a non-directional point sound source, and the distance between the sound source S and the sound receiving point T is x, the attenuation is 20log 10 x. The propagation time can be calculated by dividing the distance x by the speed of sound c. If a content creator desires to emphasize the sense of distance between the sound-emitting point S and the sound-receiving point T, the sound processing device 100 can also change the amount of attenuation without being bound by physical phenomena in the real world.

[0071] Next, the sound processing device 100 generates a sound signal based on the early reflected sound from the sound signals heard at the sound receiving point T (step S105). Specifically, the sound processing device 100 receives the acquired sound data as input, and calculates the attenuation and time lag according to the propagation path length in the medium based on the propagation path calculated in advance, as in the case of direct sound. Furthermore, the sound processing device 100 performs signal processing that simulates reflection at a boundary surface based on information about the boundary surface that reflects the input signal. For example, the sound processing device 100 may provide a predetermined attenuation for each frequency depending on the sound absorption coefficient of the boundary surface that reflects the input signal.

[0072] For example, for the reflected sound corresponding to path R2 shown in FIG. 7 , the section from the sound generation point S to the boundary B1 and the section from the boundary B1 to the sound receiving point T can be considered as the propagation path of the medium. The sound processing device 100 calculates the attenuation and time lag from the distance of each of these propagation paths. Note that the sound processing device 100 may calculate attenuation by storing sound absorption coefficient data for each frequency for the boundary B1 in a library in advance and reading out that data as needed. The sound processing device 100 may generate an audio signal of the reflected sound corresponding to path R2 by applying attenuation according to the sound absorption coefficient to the input signal in addition to the attenuation on the previous propagation path. In this case, if the effects of plate vibration and resonance at the boundary B1 and nonlinear phenomena are not taken into consideration, the sound processing device 100 may calculate attenuation, etc. by combining the section from the sound generation point S to the boundary B1 and the propagation path from the boundary B1 to the sound receiving point T.

[0073] Similarly, the sound processing device 100 performs calculations for reflected sound corresponding to path R3 shown in FIG. 7 . That is, the sound processing device 100 performs calculations according to the propagation paths from the sound emission point S to boundary B2, from boundary B2 to boundary B3, and from boundary B3 to sound receiving point T. When calculating sound reflected at multiple boundary surfaces such as boundary B2 and boundary B3, the sound processing device 100 may perform calculations of attenuation according to the sound absorption coefficient all at once if nonlinear behavior according to sound pressure at the boundary surfaces is not taken into consideration. Digital filters such as IIR (Infinite Impulse Response) and FIR (Finite Impulse Response) may be used to calculate attenuation for each frequency according to the sound absorption coefficient. Note that in the process of generating early reflected sounds, the sound processing device 100 may appropriately use a technique such as ray tracing to simulate the reflection from each wall surface.

[0074] Next, the sound processing device 100 generates an audio signal related to the diffracted sound (step S106). As with the early reflection sound, the sound processing device 100 generates an audio signal related to the diffracted sound using a geometric method that calculates attenuation and time lag based on the propagation path. It is known that the attenuation of diffracted sound increases as the frequency increases in the sound diffraction phenomenon. Therefore, the sound processing device 100 may generate the diffracted sound by applying a filter with different attenuation amounts depending on the frequency. To adopt a more accurate method based on the physical laws of wave phenomena, the sound processing device 100 can use not only the geometric simulation method described above, but also a wave sound simulation method from the sound generation point S to the sound receiving point T. In this case, the amount of calculation may be greater than in the case of geometric simulation. Therefore, the user U may use a device with sufficient computing power as the sound processing device 100, or may prepare a library in which the characteristics of representative propagation path shapes are calculated in advance.

[0075] Next, the sound processing device 100 generates late reverberation sounds (step S107). Late reverberation sounds refer to the sound emitted from the sound source (sound source) at the sound receiving point T after repeated reflections and diffractions in space, excluding the early reflection sounds. This reflection and diffraction continues until the sound is completely attenuated (for example, until the quantization error on the computer can be considered to be zero). Here, the completely attenuated sound may be determined as a discrimination limit in the listener's own perception by the listener operating the user interface of the sound processing device 100 for the reproduced sound including the late reverberation sounds generated by the sound processing device 100.

[0076] In a simulation, it is also possible to perform calculations until the sound is completely attenuated. However, since the amount of calculation increases with each reflection, the sound processing device 100 may calculate the early reflection sound by considering up to a predetermined number of reflections as early reflections. The sound processing device 100 may then calculate subsequent sounds using a different method and combine them with the early reflection sound, etc. Here, the predetermined number of reflections included in the early reflection sound may be set by the creator, or may be dynamically changed or set within the sound processing device 100 depending on the type of content or virtual space to be reproduced, the size of the space, the sound absorption coefficient of boundary surfaces within the space, object surfaces, etc. Note that methods for calculating late reverberation sound, such as a method of calculating reverberation time using a statistical method based on the size of the space, the sound absorption coefficient of boundary surfaces within the space, object surfaces, etc., are also known in the field of architectural acoustics.

[0077] To generate late reverberation sounds, the sound processing device 100 can use a method using a comb filter, which uses acquired audio data as input, multiplies the input by a predetermined gain, and feeds back the input signal so that the input signal attenuates and repeats at a constant cycle, or a method using impulse response data stored in a library or the like convolved with the input signal. The sound processing device 100 may also generate late reverberation sounds using other known techniques. As described above, late reverberation sounds have a significant impact on a user's auditory perception of space and are therefore important in content creation. Many existing modular reverberation generators allow adjustment of late reverberation settings (parameters), such as late reverberation level, late reverberation delay time, decay time, frequency-specific decay ratio, echo density, and modal density.

[0078] The sound processing device 100 generates the direct sound, the early reflection sound, the diffracted sound, and the late reverberation sound, and then synthesizes these signals. The sound processing device 100 then outputs the synthesized audio signal to a speaker or the like as the sound observed at the sound receiving point T (step S108). Fig. 8 is a diagram showing an example of the synthesized and output audio signal. Once the output is complete, the sound processing device 100 ends the audio signal generation process.

[0079] The generation processes from step S104 to step S106 above may be switched around because there is no dependency on the order of the steps. Also, when considering the case where diffraction occurs after reflection in the propagation path, the signal generated by first calculating the reflection attenuation at the boundary from the sound generation point may be used as input for calculating the diffracted sound.

[0080] 5, the sound processing device 100 may add processing using a porting simulation technique that simulates sound transmission through a boundary surface or the characteristics of a small sound passing through a space such as a window or door of a building, thereby enabling the sound processing device 100 to realize sound representation that is closer to phenomena occurring in real space.

[0081] Furthermore, when there are multiple reflected sounds, diffracted sounds, and transmitted sounds on the path, the sound processing device 100 may perform these calculations for each location where each phenomenon occurs, and after all calculations have been completed, may synthesize and output the audio signal. Fig. 9 is a flowchart showing an example of processing when there are multiple early reflected sounds, diffracted sounds, and transmitted sounds. The audio signal generation processing shown in Fig. 9 may be executed by a single computer, or may be executed by multiple computers working together.

[0082] The flowchart shown in FIG. 9 includes new steps S109 and S110, as compared to the flowchart shown in FIG. 5 . That is, the sound processing device 100 processes transmitted sound in addition to processing direct sound, early reflected sound, and diffracted sound. Specifically, the sound processing device 100 generates an audio signal based on diffracted sound, and then generates an audio signal based on transmitted sound (step S109). The sound processing device 100 then determines whether all of the early reflected sound, diffracted sound, and transmitted sound have been calculated (step S110). If all of the early reflected sound, diffracted sound, and transmitted sound have not been calculated (step S110: No), the sound processing device 100 returns to step S105. If all of the early reflected sound, diffracted sound, and transmitted sound have been calculated (step S110: Yes), the sound processing device 100 proceeds to step S107. The generation processes of steps S104, S105, S106, and S109 can be interchanged.

[0083] If there are multiple sound-emitting points S, the sound processing device 100 may perform the entire process described above in parallel for the sounds output from the multiple sound-emitting points S. Alternatively, the sound processing device 100 may perform sequential processing by delaying the time until output until all processing is completed, and may further combine and output each combined signal in which the time series of the sounds arriving at the sound-receiving point T are aligned.

[0084] 5 to 9, an example of a general-purpose audio signal generation process has been shown. Next, the audio signal generation process according to this embodiment will be described in detail. In this embodiment, the description will focus on the process of generating early reflection sounds (step S105 shown in FIGS. 5 and 9) out of the above-mentioned audio signal generation process.

[0085] Conventional systems, as explained using Fig. 7, calculate paths using geometric acoustics, calculate attenuation / time delay due to propagation distance, and simulate the behavior of reflections at boundaries. To this end, conventional systems replace the sound absorption coefficient (or attenuation / reflection coefficient) included in the material information set at boundaries with a filter such as IIR (Infinite Impulse Response), and then use that filter to perform signal processing to represent reflected sounds.

[0086] This method has many similarities to the ray tracing method, which is widely used in the imaging field, and is relatively easy to implement on a computer. However, with the limited number of rays due to constraints on computational resources, it is difficult for a computer to use this method to simulate phenomena caused by the fine shape of a wall surface (e.g., phenomena such as scattering and interference). Therefore, conventional systems generally treat such shapes as smooth walls and replace them with a simple filter (e.g., simple IIR). In this case, a computer cannot represent these phenomena (e.g., scattering and interference) that can occur in real physics.

[0087] Furthermore, conventional systems perform signal processing that assumes boundaries expressed only by simple values ​​(e.g., sound absorption coefficients using only the real part) and that does not consider changes in sound absorption coefficients due to angle. If a computer were to use this type of signal processing to represent reflected sound from second-order and subsequent reflections, highly correlated signals (e.g., simple time-delayed signals) would be synthesized. This would result in the formation of a comb filter, causing unexpected peaks and dips in frequencies that would not occur in real-world physical phenomena.

[0088] In other words, using conventional methods, it is difficult to create a realistic acoustic space (e.g., an acoustic space that reproduces phenomena such as scattering and interference, and / or an acoustic space with few unexpected peaks and dips) with a small computational load.

[0089] Therefore, in this embodiment, the sound processing device 100 generates a short-term impulse response by using a wave acoustic simulation of wave behavior near a wall surface (or a machine learning model trained based on the results of the wave acoustic simulation).The sound processing device 100 then expresses reflected sound by convolving this short-term impulse response with input sound.This allows the sound processing device 100 to create a realistic acoustic space with a small calculation load.

[0090] The reflected sound generation process of this embodiment will be described in detail below. FIG. 10 is a flowchart showing the reflected sound generation process of this embodiment. The process shown in FIG. 10 is executed, for example, in step S105 of the audio signal generation process shown in FIG. 5 or FIG. 9. The reflected sound generation process may be executed by a single computer, or may be executed by multiple computers working together. In the following description, it is assumed that the control unit 130 (acquisition unit 131 to output control unit 135) of the sound processing device 100 executes the reflected sound generation process. The reflected sound generation process will be described below with reference to the flowchart of FIG. 10.

[0091] First, the acquisition unit 131 of the sound processing device 100 acquires data to be used in the reflected sound generation process (step S201). For example, the acquisition unit 131 acquires the audio data acquired in the above-mentioned step S101 and the spatial data acquired in the above-mentioned step S102. More specifically, the acquisition unit 131 acquires, for example, three-dimensional shape data required for path search, texture data of objects where reflection may occur, and material / sound absorption coefficient data assigned to the objects.

[0092] The search unit 132 of the sound processing device 100 searches for a sound path from a sound source (sound emission point S) to a listening point (sound receiving point T) in the virtual space V2 (step S202). This process may be the same as the process shown in step S103 described above. For example, the search unit 132 may search for a sound path using a geometric acoustic simulation method (e.g., a ray tracing method or a virtual image method). Here, the search unit 132 may select in advance from the searched paths a path that includes reflections for subsequent processing.

[0093] The identifying unit 133 of the sound processing device 100 identifies a point on the path where sound reflection occurs (hereinafter simply referred to as a reflection point) (step S203). For example, the identifying unit 133 identifies a path that includes reflection from among the paths searched in step S202. Then, the identifying unit 133 identifies the sound reflection point on that path.

[0094] Next, the control unit 130 (for example, the signal processing unit 134) of the sound processing device 100 performs signal processing, which is one of the features of this embodiment, in steps S204 and S205.

[0095] First, the control unit 130 of the sound processing device 100 performs processing related to sound reflection. For example, the control unit 130 generates an impulse response (IR) for each reflection point by performing calculations limited to the area surrounding the reflection point (step S204). For example, the control unit 130 generates a short-term impulse response based on characteristic information related to the sound at the reflection point (for example, at least one of the following (1) to (5)).

[0096] (1) Shape of the reflection point (2) Material of the reflection point (3) Sound absorption coefficient data of the reflection point (4) Angle from the reflection surface at the reflection point to the sound source (or the immediately preceding reflection / diffraction point) (5) Information on the distance from the reflection point to the sound source (or the immediately preceding reflection / diffraction point)

[0097] The characteristic information about the sound at the reflection point may include texture data set at the reflection point (for example, texture data of an object surface (for example, a wall surface)). In this case, the control unit 130 may generate characteristic information about the reflection point (for example, shape, material, or sound absorption coefficient) based on the texture data. Then, the control unit 130 may generate an impulse response based on the characteristic information (for example, shape, material, or sound absorption coefficient).

[0098] The impulse response generation process in step S204 will be described in detail later.

[0099] Next, the signal processing unit 134 of the sound processing device 100 performs processing related to the sound at the sound receiving point T. Specifically, the signal processing unit 134 performs processing to convolve the impulse response generated in step S204 with the audio data acquired in step S201 (step S205). At this time, the signal processing unit 134 may perform the convolution processing in the time domain or in the frequency domain. For example, the signal processing unit 134 may perform the convolution processing in the frequency domain using an FFT (Fast Fourier Transform) and then convert the processed signal back into a time domain signal using an IFFT (Inverse Fast Fourier Transform).

[0100] Next, the output control unit 135 of the sound processing device 100 outputs the sound (reflected sound) processed in step S205 to, for example, the storage unit 120 (step S206). When the output is completed, the control unit 130 of the sound processing device 100 returns the process to the sound signal generation process (for example, the sound signal generation process shown in FIG. 5 or 9).

[0101] <3-3. Impulse Response Generation Processing> An example of the reflected sound generation processing has been shown above using Fig. 10. Next, the impulse response generation processing will be described in detail.

[0102] Fig. 11 is a flowchart showing the impulse response generation process of this embodiment. The process shown in Fig. 11 is executed, for example, in step S204 of the reflected sound generation process shown in Fig. 10. The impulse response generation process may be executed by a single computer, or may be executed by multiple computers working together. In the following description, it is assumed that the control unit 130 (acquisition unit 131 to output control unit 135) of the sound processing device 100 executes the impulse response generation process. The impulse response generation process will be described below with reference to the flowchart in Fig. 11.

[0103] First, the acquisition unit 131 of the sound processing device 100 acquires data to be used in the impulse response generation process (step S301). For example, the acquisition unit 131 acquires the spatial data acquired in the above-mentioned step S102. More specifically, the acquisition unit 131 acquires, from the spatial data, for example, shape data near the reflecting wall surface, information on the angle between the wall surface and the sound source, etc., and information on the distance from the wall surface to the sound source, etc.

[0104] Next, the signal processing unit 134 of the sound processing device 100 sets a region that will be the calculation region for the wave calculation (step S302). At this time, the signal processing unit 134 may extract a region around the reflection point in the virtual space V2 according to a predetermined criterion, and set the extracted region as the wave calculation region. At this time, the region that will be the wave calculation region (the region around the reflection point) may be a hemispherical region on the sound source side (sound receiving point side) of two regions obtained by dividing a sphere centered on the reflection point by a sound reflection surface.

[0105] Fig. 12 is a diagram showing an example of setting a wave calculation region. In the example of Fig. 12, a hemispherical region (space) centered on a reflection point P4 and extending toward the sound source side (sound receiving point T side) is set as the calculation region for wave calculation. More specifically, the signal processing unit 134 sets as the wave calculation region a hemispherical region centered on the reflection point P4 on a path including reflections from the sound generation point S to the sound receiving point T, and a hemispherical region (surrounding region A4 shown in Fig. 12) including the shape of a wall surface (object J1 shown in Fig. 12) that serves as a reflecting surface.

[0106] If the object J1 (wall surface) serving as a reflective surface has a curved surface or ridge lines except for minute irregularities, the signal processing unit 134 may perform a shape change process to bend the curved surface or corners in the opposite direction so that the object J1 (wall surface) becomes approximately flat, and then set a hemispherical space as the wave calculation domain. At this point, information such as the material (sound speed, density) and sound absorption coefficient may be set for the object J1 (wall surface).

[0107] The hemispherical region serving as the wave calculation region is desirably large enough to accommodate the required calculation volume, since it also affects the calculation volume during wave calculation. When a virtual sound source (virtual sound source) and a virtual sound receiving point are set in the hemispherical region, the calculation range can be considered to be from the first impulse of the reflected and scattered waves observed at the virtual sound receiving point until the appearance of scattered waves with a certain peak value. In this case, the calculation volume during wave calculation is strongly correlated with the roughness of the wall surface shape, the magnitude of the wall surface irregularities, and the sound absorption coefficient of the wall surface. Therefore, the signal processing unit 134 may set the wave calculation region based on the results of calculations performed on multiple wall surfaces within a sufficiently large hemisphere.

[0108] Furthermore, when setting boundary conditions for wave calculation, the signal processing unit 134 may set the outside of the hemisphere to a non-reflecting boundary, a PML (Perfectly Matched Layer), or the like. In this case, signal processing using a time waveform, which will be described later, is also possible. Therefore, if setting these boundary conditions causes inconvenience, the signal processing unit 134 may extend the hemisphere outside the hemisphere where the virtual sound source and virtual sound receiving point are located, and set the boundary of the hemisphere to any value.

[0109] Note that wave calculations require a large amount of calculation even in a small, limited space such as a sphere. Because three-dimensional calculations have a greater computational load than two-dimensional calculations, the signal processing unit 134 may use a two-dimensional plane as the wave calculation region rather than a three-dimensional space. For example, the signal processing unit 134 may use a semicircular region as the wave calculation region rather than a hemispherical region. In this case, the signal processing unit 134 may define a plane in the three-dimensional space that includes three points: the reflection point, the sound source (sound-emitting point), and the sound-receiving point, and may use a cross section of the three-dimensional space cut by this plane as the two-dimensional plane used in the wave calculation.

[0110] 13A and 13B are diagrams showing an example of a two-dimensional wave calculation region. In the example of Fig. 13A and 13B, the signal processing unit 134 sets, as the wave calculation region, a semicircular region (surrounding region A5) around the reflection point P5 on a plane including the reflection point P5, the sound-emitting point S, and the sound-receiving point T.

[0111] As shown in FIG. 13B , if the reflecting surface (object J1) has a shape in which regular irregularities appear in one direction (the Z-axis direction in the example of FIG. 13B ), the signal processing unit 134 may use a cross section cut along a plane perpendicular to the cross section (the XY plane in the example of FIG. 13B ) as the two-dimensional plane used in the wave calculation. In this case, when the sound source and the sound receiving point are mapped onto a two-dimensional plane, differences in the incidence angle are likely to result in differences in the influence of scattered waves caused by the irregularities. This is preferable because it reduces the likelihood of an unpleasant auditory sensation. Instead of three-dimensional data, two-dimensional data simulating the cross-sectional shape of the wall surface may be prepared in advance for the wave calculation. Alternatively, the signal processing unit 134 may estimate the irregularities / cross section of the reflecting surface based on texture data using machine learning or the like. The signal processing unit 134 may then use the estimated results in the wave calculation. This enables a harmonious representation of appearance and sound.

[0112] Returning to the flowchart of FIG. 11 , the description continues. The signal processing unit 134 determines the sound source shape, which serves as an initial condition for wave calculation, depending on the distance from the wall surface to the sound source, etc. Specifically, the signal processing unit 134 determines whether the distance from the sound source to the wall surface is equal to or greater than a threshold (step S303). If the distance is equal to or greater than the threshold (step S303: Yes), the signal processing unit 134 places a line sound source or a surface sound source as a temporary sound source in the surrounding area (step S304). If not (step S303: No), the signal processing unit 134 places a point sound source as a temporary sound source in the surrounding area (step S305). At this time, the signal processing unit 134 may place the temporary sound source (and temporary sound receiving point) on the boundary of the surrounding area (wave calculation area) (for example, on the spherical surface of the hemispherical surrounding area or on the circumference of the semicircular surrounding area).

[0113] When a point sound source exists in space, the wavefront spreads as a spherical wave. That is, the wavefront spreads in a spherical shape with a radius equal to the distance from the sound source. Therefore, when the sound source is at a certain distance, the radius of the sphere becomes very large. Often, in situations where the radius of the sphere becomes large like this, a computer performing wave calculations may regard the wavefront from a certain distance as a plane wave. Therefore, even when a point sound source is set as the sound source, if the distance between the wall surface and the sound source is greater than a certain distance, the signal processing unit 134 sets a surface sound source or a line sound source that behaves like a plane wave in the wave calculation domain as a temporary sound source during wave calculation. On the other hand, if the sound source is close to the wall surface, the signal processing unit 134 sets a point sound source as a temporary sound source in the wave calculation domain. Note that if the sound source size and sound source type (plane wave, spherical wave, cylindrical wave, directional, etc.) are given to the sound source in advance, the signal processing unit 134 may prioritize them.

[0114] Next, the signal processing unit 134 calculates the response from the virtual sound source to the virtual sound receiving point based on the conditions set in steps S301 to S305 (step S306). At this time, the signal processing unit 134 may calculate a time response waveform until the peak value of the scattered wave falls below a certain level. FIG. 14 is a diagram for explaining the wave calculation of this embodiment. In the example of FIG. 14, the wave calculation region is a hemispherical region, but it may also be a semicircular region. For the sake of simplicity of calculation and convenience of calculation amount, the wave calculation region is generally a hemispherical or semicircular region, but it is not limited to this, and a wave calculation region of a rectangular shape, a triangle, or the like may also be set. Furthermore, the wave calculation region may be changed or set to any shape by the creator. In the example of FIG. 14, the sound generation point S4, which is a virtual sound source, and the sound receiving point T4, which is a virtual sound receiving point, are set equidistant from the reflection point P4.

[0115] The signal processing unit 134 inputs a short-duration wideband signal (e.g., a rectangular pulse, a Gaussian pulse, or a sinusoidal pulse) to the sound-emitting point S4. The signal processing unit 134 then calculates how a wavefront transiently propagates from the sound-emitting point S4, collides with a wall surface, and is reflected and scattered, thereby obtaining a sound pressure time waveform at the sound-receiving point T4. This transient response calculation can be performed using acoustic simulation methods such as finite-difference time-domain (FDTD), constrained interpolation profile (CIP), and finite element method (FEM). This allows the signal processing unit 134 to obtain a response from the sound-emitting point S4 to the sound-receiving point T4 in the wave calculation region (surrounding region A4).

[0116] The signal processing unit 134 can also obtain a time-domain signal by performing an inverse Fourier transform on the results of frequency analysis at multiple frequencies. When using this method, the signal processing unit 134 determines the spatial discretization width (mesh size) and the time discretization width (sampling rate) according to the expected band of reflected and scattered sound. Generally, a coarser resolution reduces the upper limit frequency that can be accurately calculated, thereby reducing the computational load. Even with the same discretization width, the amount of calculation increases or decreases depending on the size of the hemisphere or semicircle. Therefore, the signal processing unit 134 may change the discretization width depending on the size of the wave calculation domain. When changing the time discretization width, the signal processing unit 134 must match the sampling rate with the sampling rate of the sound source in subsequent convolution processing, etc. Therefore, the signal processing unit 134 upsamples or downsamples the output signal and uses it for subsequent processing. The signal processing unit 134 may also switch the order of this processing with subsequent processing. Note that this content is not limited to calculations in the frequency domain. This content can also be applied to calculations in the time domain (for example, calculations in the transient response).

[0117] 15A and 15B are diagrams illustrating examples of wavefront calculations. Each of FIGS. 15A and 15B shows a wavefront extracted from a certain time period of a transient response. The examples of FIGS. 15A and 15B show space calculations using FEM. FIG. 15A is a calculation example for an uneven wall surface (object J4) where scattering is likely to occur, while FIG. 15B is a calculation example for a smooth wall surface (object J5) where fine scattering is unlikely to occur. In both of the examples of FIGS. 15A and 15B, the sound-emitting point S6 and the sound-receiving point T6 are positioned equidistant from the reflection point P6. Furthermore, in the examples of FIGS. 15A and 15B, the wave calculation region (surrounding region A6) extends outside the hemispherical surface (the dotted line portion shown in FIGS. 15A and 15B) where the sound-emitting point S6 and the sound-receiving point T6 are located. When a computer represents reflected sound using conventional geometric techniques, scattering cannot be taken into account, and the computer can only simulate a wall surface such as that shown in FIG. 15B.

[0118] 16A and 16B are diagrams showing examples of calculations of impulse responses. Fig. 16A is a time signal of the response of the wall surface (object J4) shown in Fig. 15A, and Fig. 16B is a time signal of the response of the wall surface (object J4) shown in Fig. 15B. The time response of the reflected and scattered sound shown in Fig. 16A is more complex than the time response shown in Fig. 16B. This is a feature of the reflected sound representation of this embodiment, which takes into account the unevenness of the wall surface.

[0119] Returning to the flowchart of FIG. 11 , the description will be continued. The signal processing unit 134 generates a response in free space from the virtual sound source to the virtual sound receiving point (step S307). At this time, the signal processing unit 134 may generate the response in free space by performing calculations while replacing the boundary conditions of the wall surfaces used in the processing of step S306 with sound absorbing walls / PML, etc. Alternatively, the signal processing unit 134 may generate the response in free space by defining another space in which only the distance between the virtual sound source and the virtual sound receiving point is the same and calculating a transient response. The same data as that used in step S306 is used as input to the virtual sound source.

[0120] FIG. 17 is a diagram showing an example of a response in free space. FIGS. 18A and 18B are diagrams showing input waveforms. FIG. 18A shows the time response, and FIG. 18B shows the spectrum. The signal processing unit 134 inputs a Gaussian pulse as shown in FIGS. 18A and 18B to a virtual sound source. The response (FIG. 17) of the simulation result shows unwanted dullness in the falling edge and residual DC components. These can be attributed to several factors, including numerical errors in the simulation, numerical dispersion, and the difference between the ideal boundary and the implemented boundary. This particular phenomenon also appears in the calculation results in step S306. Therefore, the signal processing unit 134 performs subsequent processing to remove unwanted components from the reflected waveform.

[0121] Returning to the flowchart of FIG. 11 , the description continues. The signal processing unit 134 removes the direct sound (also referred to as the indirect sound) from the response generated in step S306 (step S308). The waveform of the direct sound included in the response generated in step S306 and the waveform of the direct sound included in the response generated in step S307 are very similar. Therefore, the signal processing unit 134 can remove the direct sound by taking the difference between the signal generated in step S306 and the signal generated in step S307. Note that, for example, if the calculation in step S307 confirms that the direct sound does not contain unnecessary components, or if a simulation method that is less likely to produce unnecessary components is used, the signal processing unit 134 can also perform the process of removing the direct sound by processing using a signal generated by delaying the input waveform used in step S306 by the propagation time of the direct wave.

[0122] The signal processing unit 134 generates an inverse filter that converts the short-duration wideband signal (e.g., a rectangular pulse, a Gaussian pulse, or a sinusoidal pulse) introduced in step S306 back into an impulse, and performs filtering based on the inverse filter (step S309). At this time, the filter generated by the signal processing unit 134 may be a signal that converts back into an impulse by convolving with the input signal. Alternatively, the filter generated by the signal processing unit 134 may be a signal that converts back into an impulse by convolving with the free-space response generated in step S307. The signal processing unit 134 processes the signal in step S308 using the generated filter to remove unnecessary components in the input signal and the simulation.

[0123] FIG. 19 is a diagram showing an example of removing unnecessary components. FIG. 19 shows a waveform (solid line) obtained by signal processing the simulation results of reflected and scattered sound including unnecessary components (short dashed line) using an inverse filter generated based on the free space propagation simulation results of step S307 (long dashed line). Here, the simulation results of reflected and scattered sound including unnecessary components are the simulation results of step S306, such as the waveform shown in FIG. 16A. The free space propagation simulation results are the simulation results of step S307, such as the waveform shown in FIG. 17. It can be seen that these processes (processing of steps S308 and S309) generate a waveform in which only reflected and scattered sound is extracted.

[0124] Returning to the flowchart of FIG. 11 , we continue the explanation. Up until the processing of step S105, the signal processing unit 134 has calculated the path to the reflection point through geometric simulation. For this reason, it is desirable to convolute only the impulse response of the reflection point in step S205. Meanwhile, in the waveform generated by the processing of steps S306 to S309, a propagation time equivalent to the diameter of a hemisphere, which is set for the sake of waveform calculation, is added to the beginning of the reflection and scattering responses. Therefore, the signal processing unit 134 removes the waveform equivalent to the propagation time calculated by dividing the set diameter of the hemisphere by the speed of sound from the waveform generated in step S309 (step S310). However, due to the uneven shape of the wall surface, it is possible that the reflected and scattered components may arrive several samples earlier than the simply calculated propagation time. Therefore, if a convex shape exists inside the portion set as the reflection point, the signal processing unit 134 subtracts several samples from the propagation time from the waveform generated in step S309.

[0125] Next, the signal processing unit 134 cuts out the impulse response waveform obtained up to step S310 into a signal of a length that is easy to handle in step S205 and thereafter (step S311). As described in step S306, the signal processing unit 134 calculates, for example, the time response waveform until the peak value of the scattered wave falls below a certain level. However, when the signal processing unit 134 performs the convolution process in step S205, it often performs calculations in the frequency domain using FFT. Therefore, it is desirable for the signal processing unit 134 to set the number of samples of the response waveform to a power of two. On the other hand, when considering use in games, etc., interactive tone changes are also expected. Therefore, to avoid cumbersome processing and / or unnatural tone changes in the convolution process, it is considered to set the signal length of the impulse response to one frame or less.

[0126] Therefore, the signal processing unit 134 cuts out from the signal (response waveform) the portion from the beginning of the response waveform up to the first power-of-two sample that appears after the peak value of the scattered wave falls below the threshold. Alternatively, the signal processing unit 134 cuts out from the signal (response waveform) the portion from the beginning of the response waveform up to the largest power-of-two sample that fits within a frame. As an example of the former, consider a case where the audio sample rate is 48 kHz. In this case, if the peak value falls below the threshold 5 ms from the beginning of the signal, the signal will fall below the threshold at 240 samples. In this case, the next power-of-two number that appears is 256, so the signal processing unit 134 cuts out from the signal up to 256 samples from the beginning of the waveform. As an example of the latter, consider a case where the game frame rate is 60 fps. In this case, one frame is approximately 16.7 ms. Since the largest power-of-two sample that is less than one frame in length is 512 samples, the signal processing unit 134 cuts out from the signal up to 512 samples from the beginning of the waveform.

[0127] Next, the signal processing unit 134 performs fade-out processing on a predetermined number of samples at the end of the waveform extracted in step S311 (step S312). The signal processing unit 134 performs simple cutting when extracting the waveform in step S311. Therefore, the signal has characteristics similar to those obtained by applying a rectangular window. Therefore, the signal processing unit 134 processes the signal so that the absolute value of the signal gradually decreases until the value of the last sample becomes zero.

[0128] Next, the output control unit 135 of the sound processing device 100 outputs the impulse response generated in the processes up to step S312 to, for example, the storage unit 120 (step S313). When the output is completed, the control unit 130 of the sound processing device 100 returns the process to the reflected sound generation process.

[0129] <<4. Effects>> The sound processing device 100 of this embodiment simulates sound propagation using the area surrounding the reflection point as the calculation area, and calculates the sound propagating to the listening point (sound receiving point) based on the simulation results.

[0130] For example, the sound processing device 100 generates an impulse response by calculation limited to the area surrounding the reflection point, taking into account the unevenness of the wall surface expressed by texture, etc., and performs a convolution operation. Therefore, the sound processing device 100 can use the same path search method as the acoustic simulation using a flat wall surface that has been used to express conventional virtual spaces, and can highly reproduce the tone of sound when it hits an object and is reflected and scattered, with a small calculation load.

[0131] Furthermore, the sound processing device 100 performs a simulation of sound propagation taking into account wave behavior limited to the area surrounding the reflection point, and therefore can more accurately reproduce physical phenomena (e.g., scattering phenomena that change depending on the angle of incidence and angle of reflection) that accompany changes in the orientation of a sound source virtually placed in the area surrounding the reflection point relative to the reflection surface, with less computational load.

[0132] FIG. 20 is a diagram for explaining the effect of this embodiment. FIG. 20 is a diagram showing changes in the wavefront (and waveform at the sound receiving point) due to changes in angle or surface shape. The center of FIG. 20 shows the wavefront (upper side of the diagram) and the waveform at the sound receiving point (lower side of the diagram) under reference conditions. The left side of FIG. 20 shows the wavefront and the response waveform at the sound receiving point when the angle of incidence of sound on the reflecting surface is changed. The right side of FIG. 20 shows the wavefront and the response waveform at the sound receiving point when the shape of the reflecting surface is changed. The waveform at the bottom of FIG. 20 is an example of processing in step S306. As can be seen from FIG. 20, the sound processing device 100 can accurately reproduce scattering phenomena caused by changes in angle or surface shape.

[0133] Furthermore, in this embodiment, in an environment where reflections occur at multiple locations, a different filter is used for each reflecting object (or for each relative position between the reflecting object and the sound source). Therefore, by using the method of this embodiment, the correlation between sounds arriving via multiple paths is weakened. As a result, the sound processing device 100 of this embodiment can reduce peaks, dips, and echoes that are likely to occur in the past when reproducing such scenes.

[0134] <<5. Modifications>> The above-described embodiment is merely an example, and various modifications and applications are possible. Modifications of the above-described embodiment will be described below.

[0135] <5-1. Pre-generation of impulse responses> In the above-described embodiment (step S204 and steps S301 to S313), the sound processing device 100 generates impulse responses when playing / executing content (for example, when an event occurs in a virtual three-dimensional space). However, the sound processing device 100 may generate impulse responses in advance. For example, the sound processing device 100 may perform the process of step S204 (steps S301 to S313) in advance and create tables for each predetermined condition (for example, for each wavefront shape, wall surface, angle, and texture phase).

[0136] More specifically, the sound processing device 100 generates impulse responses in advance for each surface data of an object for which reflected sound may be generated for each map / scene. For example, the sound processing device 100 performs the process of step S204 (steps S301 to S313) in advance for each surface data of the object. This enables the sound processing device 100 to respond to content (e.g., games and the Metaverse) in which the situation changes in real time and responsiveness is required.

[0137] The sound processing device 100 may generate impulse responses in advance for each sound source wavefront shape, for each angle at which reflection can occur, and / or for each positional relationship on a hemispherical flat portion of the uneven shape of the object surface, and generate a table based on that information. During the processing of step S105 in which reflected sound is actually generated, the sound processing device 100 may refer to the table based on the wall surface ID, angle ID, and positional relationship ID during the processing of step S203, and read out the pre-generated impulse responses.

[0138] The angles in the table may be discrete. In this case, when the sound processing device 100 needs data (impulse response) of an angle not included in the table, it may read data of an angle close to the required angle. Alternatively, the sound processing device 100 may use multiple data of angles close to the required angle to supplement the data of the required angle. For example, assume that data for every 10-degree angle is stored in the table. When data of a 25-degree angle is required, the sound processing device 100 may read data of 20 degrees and data of 30 degrees and average the data in the time waveform to generate data (impulse response) of the required angle. Alternatively, the sound processing device 100 may generate an impulse response of a required angle by applying an arbitrary gain to the impulse response of a certain angle or by applying a frequency change using a filter such as an IIR.

[0139] Similarly, when an impulse response of a wavefront shape that is not in the table is required, the sound processing device 100 may generate an impulse response of the required wavefront shape using data of another wavefront shape.

[0140] The computer that generates the impulse response in advance does not necessarily have to be a computer that plays / executes content related to the virtual space. That is, the sound processing device 100 that generates the impulse response in advance may be a computer that plays / executes content, or may be a computer other than a computer that executes content.

[0141] 5-2. Location of Execution of Impulse Response Generation Process In the above-described embodiment, the impulse response generation process (step S204 and steps S301 to S313) is performed by a computer that executes the content (e.g., a user terminal). However, the impulse response generation process may be performed by a computer other than the computer that executes the content (e.g., a server device connected to the user terminal via a network).

[0142] In step S204 (steps S301 to S313), the sound processing device 100 performs wave calculations in a limited area. This significantly reduces the calculation load compared to conventional methods that simulate the entire space. However, in recent years, it has become possible to handle content that handles three-dimensional data (e.g., games and videos) on mobile devices such as smartphones. Such mobile devices require even greater reductions in calculation load and power consumption.

[0143] Therefore, the impulse response generation process (step S204 and steps S301 to S313) may be performed by a computer (e.g., a server device on the cloud) connected via a network to the computer that plays / executes the content. For example, a mobile terminal such as a smartphone transmits data required for the impulse response generation process (e.g., shape data of a reflective surface that generates reflected and scattered sound, and the required sampling rate of the impulse response) to the server device on the cloud. The server device, having received the data, performs the impulse response generation process of this embodiment. The server device then transmits the generated impulse response to the mobile terminal.

[0144] 5-3. Creating a Database of Impulse Responses The sound processing device 100 may create a database of impulse responses for content / textures created by the user.

[0145] Recently, crafting games, in which users can freely create buildings and objects within the game, have become popular. In these games, it is difficult for content creators who provide the platform to adjust the sound quality in advance, so it is desirable for the sound quality to be automatically adjusted within the game. In addition, there are many cases where data created by users can be used by other users on the platform.

[0146] Therefore, the sound processing device 100 generates an impulse response using an object (e.g., a building) created by the user or its surface data (texture, surface shape) as input. In this case, the sound processing device 100 that generates the impulse response may be a server device on the cloud or a user terminal in the user's local environment. The sound processing device 100 then handles the generated impulse response as data associated with the object (or its surface data). This reduces the processing load on the computer used by other users who touch the object.

[0147] <5-4. Distribution of Pre-Generated Impulse Responses> The sound processing device 100 may process pre-generated impulse responses so that they can be distributed. For example, the sound processing device 100 may store impulse response data associated with an object (or its surface data) in a container together with format information. Examples of the format information include an encoding format, a sampling rate, a quantization bit rate, and an input wavefront shape. Furthermore, the sound processing device 100 may store the container in a database so that it can be used in other content or so that it can be used by other users.

[0148] 5-5. Normalization of Impulse Response In the above-described embodiment (steps S204 to S205), the sound processing device 100 generates an impulse response and performs a process of convolving the generated impulse response. At this time, the sound processing device 100 may normalize the waveform of the impulse response based on the maximum sound pressure. Then, when performing a process of convolving the normalized impulse response, the sound processing device 100 may adjust the gain of the impulse response.

[0149] That is, the sound processing device 100 normalizes the output waveform of step S204 so that the maximum sound pressure is 1 (or the maximum value of the data type for sound pressure retention). This allows the sound processing device 100 to handle the impulse response data separately as a normalization parameter and a waveform representing timbre. The normalization parameter is a normalization parameter that becomes a gain in the processing of the subsequent step S205. This modification is applicable both to the case where step S204 is processed sequentially and to the case where the impulse responses generated in step S204 are compiled into a table.

[0150] This allows the creator to adjust only the gain section when he / she wants to exaggerate a particular reflected sound, thereby improving the creator's workability.

[0151] <5-6. Suppression of Specific Reflected Sound Calculations> When the sound processing device 100 handles sound that reaches the sound receiving point after multiple reflections, when it handles sound that has multiple paths from the sound source to the sound receiving point, or when it handles many sound sources in a single scene, the calculation load of the reflected sound on the sound processing device 100 becomes large. In these cases, this may have a negative impact on the progress of the game or the progress of content production, or may make it impossible to process audio signals appropriately. Therefore, the sound processing device 100 may reduce the calculation load of the reflected sound calculations by suppressing calculations of specific reflected sounds.

[0152] For example, the sound processing device 100 may exclude from processing the sound-related process (e.g., the process of steps S204 to S205) any of the multiple paths from the sound source to the sound receiving point that have more than a predetermined number of reflections.

[0153] Furthermore, the sound processing device 100 may exclude from the processing target of sound processing (for example, the processing of steps S204 to S205) a path in which the sum of the amount of sound attenuation due to reflection satisfies a predetermined criterion, among multiple paths from a sound source to a sound receiving point. For example, after performing a path search, when reading out an impulse response from the table, the sound processing device 100 may exclude from the processing target sound propagating through a path including a response whose gain information is equal to or less than a predetermined value, or a path in which the total gain of the paths is equal to or less than a predetermined value.

[0154] This allows the sound processing device 100 to reduce the calculation load of the reflected sound calculation. Note that the thresholds for the number of reflections and the sum of the attenuation amounts (for example, (1) the value of the variable N when sounds that have reflected N times or more are excluded from the processing target, and (2) the value of the variable M when paths for which the sum of the attenuation amounts of sounds due to reflections is below M are excluded from the processing target) may be set arbitrarily by the creator, or may be dynamically changed or set within the sound processing device 100 depending on the type of content or virtual space to be reproduced, the size of the space, the sound absorption coefficients of boundary surfaces in the space, object surfaces, etc.

[0155] <5-7. Sound Source Shape> In the above-described embodiment (steps S304 and S305), the sound processing device 100 arranged a line sound source or a point sound source in the wave calculation region (surrounding region). However, the sound source shape is not limited to a line sound source or a point sound source. For example, the sound processing device 100 may arrange the shape of the sound source arranged in the wave calculation region to a shape other than a line sound source or a point sound source in accordance with the directivity of the original sound source. For example, when a sound is emitted from an avatar having a specific shape in the Metaverse space, the sound processing device 100 may adjust the shape of the sound source to the shape of the avatar.

[0156] 5-8. Use of Machine Learning In the above embodiment (step S306), the sound processing device 100 generates an impulse response based on the processing result of the wave sound simulation. However, the sound processing device 100 may generate in advance a learning model that learns the relationship between the input and output of the wave sound simulation, and generate an impulse response using the generated learning model.

[0157] For example, the sound processing device 100 generates a learning model using data used in a wave acoustic simulation and the simulation results as training data. Here, the training data may include at least one of an output from a wave acoustic simulator that performs wave calculations from a virtual sound source to a virtual sound receiving point, actual measurement data, and a physical theoretical formula.

[0158] The learning model is, for example, a machine learning model such as a neural network model. The neural network model is composed of layers called an input layer, an intermediate layer (or hidden layer), and an output layer, each of which includes a plurality of nodes, and each node is connected via an edge. Each layer has a function called an activation function, and each edge is weighted. The learning model has one or more intermediate layers (or hidden layers). When the learning model is a neural network model, learning the learning model means setting, for example, the number of intermediate layers (or hidden layers), the number of nodes in each layer, or the weight of each edge. The sound processing device 100 may learn the learning model using a method such as backpropagation.

[0159] Here, the neural network model may be a model based on deep learning. In this case, the neural network model may be a model in a form called a deep neural network (DNN). Furthermore, the neural network model may be a model in a form called a convolution neural network (CNN), a recurrent neural network (RNN), or a long short-term memory (LSTM). Of course, the neural network model is not limited to these types of models.

[0160] Furthermore, the form of the learning model is not limited to a neural network. For example, the learning model may be a model based on reinforcement learning. In reinforcement learning, behaviors (settings) that maximize value are learned through trial and error. Alternatively, the learning model may be a logistic regression model.

[0161] The learning model may be composed of a plurality of models. For example, the learning model may be composed of a plurality of neural network models. More specifically, the learning model may be composed of a plurality of neural network models selected from, for example, CNN, RNN, and LSTM. When the learning model is composed of a plurality of neural network models, these plurality of neural network models may be in a subordinate relationship or in a parallel relationship.

[0162] Note that various learning algorithms can be used for training the learning model. For example, the sound processing device 100 may perform training of the learning model using a learning algorithm such as a neural network, a support vector machine, clustering, reinforcement learning, a random forest, or a decision tree.

[0163] Then, the sound processing device 100 may generate an impulse response using the generated learning model, thereby reducing the calculation load of the process in step S306.

[0164] Note that multiple learning models may be prepared. For example, multiple learning models with different forms may be prepared. If the learning model is a neural network such as a DNN, multiple learning models with different network structures may be prepared. The sound processing device 100 may select a learning model to use for generating an impulse response based on the acoustic characteristics (e.g., directivity) of the sound source or metadata attached to the sound source. The metadata may include, for example, the spatial coordinates where the sound is generated, the directivity, the type of sound, the character emitting the sound, or the attributes of the sound (e.g., angry voice, laughter, etc.). By changing the learning model used for generating the impulse response based on the acoustic characteristics (e.g., directivity) of the sound source or the metadata attached to the sound source, the sound processing device 100 can generate highly accurate impulse responses that match the characteristics of the sound source with a low computational load.

[0165] The computer that generates the learning model does not necessarily have to be a computer that plays / executes content related to a virtual space. That is, the sound processing device 100 that generates the learning model may be a computer that plays / executes content, or may be a computer other than a computer that executes content.

[0166] <5-9. Arrangement of Sound Source and / or Sound Receiving Point> In the above-described embodiment (step S305), the sound processing device 100 arranges the boundary of the surrounding area (wave calculation area). For example, the sound processing device 100 arranges a virtual sound source and a virtual sound receiving point on the spherical surface of a hemispherical surrounding area or on the circumference of a semicircular surrounding area. However, the sound processing device 100 may arrange a virtual sound source and a virtual sound receiving point (or the sound source and sound receiving point themselves) at any location within the surrounding area. In this case, the distance from the reflection point to the virtual sound source (or sound source) and the distance from the reflection point to the virtual sound receiving point (or sound receiving point) do not need to be equal. This enables more accurate reproduction of reflected and scattered sound at the sound receiving point than when performing wave calculations with a uniformly determined distance between the reflection point and the virtual sound source (or virtual sound receiving point).

[0167] 21 is a diagram showing a state in which a virtual sound source and a virtual sound receiving point are arranged within the surrounding area A7. In the example of FIG. 21 , the sound processing device 100 arranges a virtual sound source, namely, a sound emission point S7, at a distance L1 from the reflection point P7, and arranges a virtual sound receiving point, namely, a sound receiving point T7, at a distance L2 from the reflection point P7. In the example of FIG. 21 , L1 and L2 are different lengths, but they may be the same length. In other words, the sound processing device 100 may arrange the sound emission point S7 and the sound receiving point T7 at the same distance from the reflection point P7. Furthermore, the sound processing device 100 may arrange one of the sound emission point S7 and the sound receiving point T7 on the boundary of the surrounding area A7.

[0168] Furthermore, the sound source and sound receiving point located inside the surrounding area may not be a virtual sound source and virtual sound receiving point, but may be the sound source and sound receiving point themselves (i.e., the sound emission point S and sound receiving point T). In this case, one of the sound emission point S and sound receiving point T may be located on the boundary of the surrounding area. When the sound emission point S, the sound receiving point T, or both are close to the reflection point, the sound processing device 100 can directly calculate the wave phenomenon from the sound emission point S to the sound receiving point T.

[0169] The radius of the surrounding area can be set to the same value as when the sound-emitting point S and the sound-receiving point T are far from the reflection point. However, when both the sound-emitting point S and the sound-receiving point T are inside the radius, the sound processing device 100 may narrow this radius. This allows the sound processing device 100 to further reduce the calculation load.

[0170] <5-10. Input of Data for Wave Calculation> In the above embodiment (step S301), the sound processing device 100 acquired shape data of the vicinity of the wall surface on which sound is reflected as information on the reflecting surface for use in wave calculation of reflection and scattering. However, the information on the reflecting surface is not limited to this. For example, the sound processing device 100 may acquire information on the reflecting surface by cutting out a portion that will become a reflecting surface from three-dimensional data of the virtual space. Alternatively, the sound processing device 100 may acquire information on the reflecting surface by cutting out a portion that will become a reflecting surface from texture data. Alternatively, the sound processing device 100 may acquire information on the reflecting surface by converting texture data into shape data.

[0171] Furthermore, three-dimensional or two-dimensional shape data for wave calculation (for generating impulse response waveforms) may be prepared in advance. In this case, the sound processing device 100 may be configured so that a user (or an external computer) can input this shape data for wave calculation. For example, the sound processing device 100 may be configured so that, when a creator adjusts reflected sound in a content production environment, the creator can acquire and input this shape data for wave calculation (e.g., shape data of an object such as a wall surface) via, for example, a network. This facilitates content production.

[0172] 5-11. Acquisition of Object Data from Library In the above-described embodiment (step S201 / step S301), the sound processing device 100 acquires data of objects for which reflection may occur (for example, shape data, texture data, and material / sound absorption coefficient data assigned to the object) as data for calculating reflected sound. For example, the sound processing device 100 acquires shape data, texture data, and material / sound absorption coefficient data assigned to the object as data for calculating reflected sound. When data of objects that will become reflective surfaces (for example, shape data, texture, etc. of wall surfaces) is not set / designated in the virtual space to be processed, the sound processing device 100 may be configured to allow a creator to read and use any object data stored in a library. The sound processing device 100 may automatically select object data from the library.

[0173] Here, the object data stored in the library may be linked not only with shape data and material information but also with parameters related to reflection (for example, frequency characteristics of reflected sound, waveforms such as behavior of scattered waves, and spectrum). Furthermore, the library may store not only object data but also the waveform data of the impulse response itself, which is acquired in advance by the computer by performing the processing of step S204.

[0174] <5-12. Reflected Sound Adjustment Function> The sound processing device 100 may be configured so that a creator can edit data (e.g., object data or waveform data) stored in the above-mentioned library. For example, the sound processing device 100 may be provided with a UI (User Interface) that allows a creator to arbitrarily change object data (e.g., material parameters of an object). This allows a creator to change the frequency characteristics of reflected sound to desired characteristics. Furthermore, the sound processing device 100 may be provided with a UI that allows a creator to parametrically change the size of the object's unevenness, the scattering repetition frequency, etc., for adjusting scattering behavior.

[0175] <5-13. Searching for a path of reflected sound> In the above-described embodiment (step S103 / step S202), the sound processing device 100 searches for a path of a sound using a geometric physical simulation method (for example, a ray tracing method). However, the method of path search is not limited to the above-described method.

[0176] For example, the sound processing device 100 may search for sound paths by regarding, among object surfaces in the virtual space, object surfaces whose characteristics meet predetermined criteria as flat and uniform. For example, the sound processing device 100 may search for sound paths by regarding, among object surfaces (e.g., wall surfaces) in the virtual space, object surfaces whose irregularities are below a predetermined level as flat and uniform. This simplifies the reflection paths, allowing the sound processing device 100 to reduce the number of paths required for calculating reflected sound. As a result, the sound processing device 100 can reduce the calculation load for processing reflected sound.

[0177] Furthermore, the sound processing device 100 may search for a sound path based on information about the directivity of the sound at the sound source. For example, the sound processing device 100 may exclude paths in directions where sound is not emitted (or directions in which the strength of sound waves is equal to or less than a predetermined threshold) from the search processing targets. This allows the sound processing device 100 to reduce the number of paths involved in the calculation of reflected sound, thereby reducing the calculation load of the reflected sound processing.

[0178] Similarly, the sound processing device 100 may search for a sound path based on at least one of information on the directivity of sound at the sound receiving point, information on the position of the sound receiving point in the virtual space, and information on the distance between the sound receiving point and an object in the virtual space. For example, the sound processing device 100 may exclude from the search processing target a path in a direction where no sound is received (or a direction in which the sound receiving sensitivity does not satisfy a predetermined standard). Furthermore, for example, if the position of the sound receiving point in the virtual space is surrounded by walls on all sides except one, the sound processing device 100 may exclude from the search processing target a path in a direction other than the one side. Furthermore, for example, if the distance between the sound receiving point and an object in the virtual space is closer than a predetermined threshold, the sound processing device 100 may exclude from the search processing target a path in a direction where the object is located. This allows the sound processing device 100 to reduce the number of paths required for calculating reflected sound, thereby reducing the calculation load of reflected sound processing.

[0179] The sound processing device 100 may also search for a sound path based on a ray tracing technique, which enables the sound processing device 100 to perform a highly accurate path search.

[0180] 5-14. Other Modifications The control device that controls the sound processing device 100 of this embodiment may be realized by a dedicated computer system or a general-purpose computer system.

[0181] For example, an acoustic processing program for executing the above-described operations is stored in a computer-readable recording medium such as an optical disk, a semiconductor memory, a magnetic tape, or a flexible disk and distributed. Then, for example, the program is installed in a computer and the above-described processing is executed to configure a control device. In this case, the control device may be a device external to the sound processing device 100 (e.g., a personal computer). Alternatively, the control device may be a device internal to the sound processing device 100 (e.g., the control unit 130).

[0182] The communication program may also be stored in a disk device provided in a server on a network such as the Internet, and may be downloaded to a computer. The above-described functions may also be realized by cooperation between an operating system (OS) and application software. In this case, the components other than the OS may be stored on a medium and distributed, or may be stored on a server and downloaded to a computer.

[0183] Furthermore, among the processes described in the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using a known method. In addition, the information including the processing procedures, specific names, various data, and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified. For example, the various information shown in each drawing is not limited to the information shown in the drawings.

[0184] Furthermore, the components of each device shown in the figure are conceptual functional units and do not necessarily have to be physically configured as shown. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of the devices can be functionally or physically distributed or integrated in any unit depending on various loads, usage conditions, etc. Note that this distribution or integration configuration may also be performed dynamically.

[0185] The above-described embodiments can be combined as appropriate within the scope of the present invention without causing any inconsistency in the processing content. The order of the steps shown in the flowcharts of the above-described embodiments can be changed as appropriate.

[0186] Furthermore, for example, the present embodiment can also be implemented as any configuration that constitutes an apparatus or system, such as a processor as a system LSI (Large Scale Integration), a module using multiple processors, a unit using multiple modules, a set in which other functions are added to a unit, or the like (i.e., a configuration of a part of an apparatus).

[0187] In this embodiment, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device in which multiple modules are housed in a single housing, are both systems.

[0188] Furthermore, for example, this embodiment can have a cloud computing configuration in which one function is shared and processed jointly by a plurality of devices via a network.

[0189] <<6. Hardware Configuration Example>> An information device such as the sound processing device 100 according to the above-described embodiment is realized by, for example, a computer 1000 configured as shown in Fig. 22. Fig. 22 is a hardware configuration diagram showing an example of the computer 1000 that realizes the functions of the sound processing device 100. The computer 1000 has a CPU 1100, a RAM 1200, a ROM (Read Only Memory) 1300, an SSD (Solid State Drive) 1400, a communication interface 1500, and an input / output interface 1600. The components of the computer 1000 are connected by a bus 1050.

[0190] The CPU 1100 operates and controls each component based on programs stored in the ROM 1300 or the SSD 1400. For example, the CPU 1100 loads the programs stored in the ROM 1300 or the SSD 1400 into the RAM 1200 and executes processing corresponding to the various programs.

[0191] The ROM 1300 stores boot programs such as a Basic Input Output System (BIOS) that is executed by the CPU 1100 when the computer 1000 is started, and programs that depend on the hardware of the computer 1000 .

[0192] The SSD 1400 is a computer-readable recording medium that non-temporarily records programs executed by the CPU 1100 and data used by such programs. Specifically, the SSD 1400 is a recording medium that records an information processing program according to the present disclosure, which is an example of the program data 1450. The SSD 1400 may be another non-temporary recording medium, such as a hard disk drive (HDD).

[0193] The communication interface 1500 is an interface for connecting the computer 1000 to an external network 1550 (e.g., the Internet). For example, the CPU 1100 receives data from other devices and transmits data generated by the CPU 1100 to other devices via the communication interface 1500.

[0194] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the CPU 1100 receives data from input devices such as a touch panel, keyboard, mouse, microphone, and camera via the input / output interface 1600. The CPU 1100 also transmits data to output devices such as a display, speaker, and printer via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs and the like recorded on a predetermined recording medium. Examples of media include optical recording media such as DVDs (Digital Versatile Discs) and PDs (Phase Change Rewritable Disks), magneto-optical recording media such as MOs (Magneto-Optical Disks), tape media, magnetic recording media, and semiconductor memories.

[0195] For example, when the computer 1000 functions as the sound processing device 100 according to this embodiment, the CPU 1100 of the computer 1000 executes an information processing program loaded onto the RAM 1200 to realize the functions of the control unit 130, etc. Also, the information processing program according to the present disclosure and data in the storage unit 120 are stored in the SSD 1400. Note that the CPU 1100 reads and executes the program data 1450 from the SSD 1400, but as another example, the CPU 1100 may obtain these programs from another device via an external network 1550.

[0196] <<7. Conclusion>> As described above, the sound processing device 100 of this embodiment searches for a sound path from a sound source to a sound receiving point in a virtual space and identifies the sound reflection points on the path. Then, the sound processing device 100 performs processing related to the sound at the sound receiving point based on the results of processing related to sound reflection, with the surrounding area of ​​the reflection point as the calculation area. For example, the sound processing device 100 performs processing to generate an impulse response of the reflection point using the surrounding area as the calculation area, and performs processing to convolve the generated impulse response with input audio data. Because the sound processing device 100 performs processing related to sound reflection by limiting it to the surrounding area, it can realize a realistic acoustic space with a small calculation load.

[0197] Although the embodiments of the present disclosure have been described above, the technical scope of the present disclosure is not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present disclosure. Furthermore, components of different embodiments and modifications may be combined as appropriate.

[0198] Furthermore, the effects of each embodiment described in this specification are merely examples and are not intended to be limiting, and other effects may also be obtained.

[0199] The present technology can also be configured as follows. (1) An acoustic processing method comprising: a search step of searching for a sound path from a sound source to a sound receiving point in a virtual space; an identification step of identifying a reflection point of the sound on the path; and a processing step of performing processing related to the sound at the sound receiving point based on a result of a wave acoustic simulation related to the reflection of the sound, with an area surrounding the reflection point as a calculation area. (2) The acoustic processing method according to (1), wherein the processing step performs processing to generate an impulse response at the reflection point as the wave acoustic simulation related to the reflection of the sound, and processing to convolve the impulse response as the processing related to the sound at the sound receiving point. (3) The acoustic processing method according to (2), wherein the processing step generates a signal with a time of one frame or less as the impulse response. (4) The acoustic processing method according to (2) or (3), wherein the waveform of the impulse response is a waveform normalized based on a maximum sound pressure, and the processing step performs gain adjustment of the impulse response when performing processing to convolve the impulse response. (5) The acoustic processing method according to any one of (2) to (4), wherein the processing step generates the impulse response based on characteristic information related to the sound at the reflection point. (6) The acoustic processing method according to (5), wherein the characteristic information of the reflection point includes at least one of the shape, material, and sound absorption coefficient of the reflection point, the angle from the reflection surface at the reflection point to the sound source, and distance information from the reflection point to the sound source. (7) The acoustic processing method according to (5) or (6), wherein the processing step generates shape data of an object surface based on texture data of the object surface in the virtual space, and generates the impulse response based on the shape data.(8) The acoustic processing method according to any one of (2) to (7), wherein the surrounding area is a hemispherical area closer to the sound source than two areas obtained by dividing a sphere centered on the reflection point by the sound reflection surface, and the processing step generates the impulse response based on the result of the wave acoustic simulation from a point on the spherical surface of the hemispherical area that is closest to the sound source or the previous reflection point to a point on the spherical surface of the hemispherical area that is closest to the sound receiving point or the next reflection point. (9) The acoustic processing method according to any one of (2) to (7), wherein the surrounding area is a semicircular planar area, and the processing step generates the impulse response based on the result of a two-dimensional wave acoustic simulation in the planar area. (10) The acoustic processing method according to (9), wherein the semicircular planar area is included in a plane that includes three points: the sound source, the reflection point, and the sound receiving point. (11) The acoustic processing method according to any one of (2) to (10), wherein the processing step generates the impulse response using a learning model that has learned the relationship between input and output of the wave acoustic simulation. (12) The acoustic processing method according to (11), wherein a plurality of learning models are prepared, and wherein the processing step selects a learning model to be used for generating the impulse response according to acoustic characteristics of the sound source or metadata assigned to the sound source. (13) The acoustic processing method according to any one of (1) to (12), wherein the processing step makes it possible to exclude from the sound-related processing any path that has a number of reflections exceeding a predetermined number of times from among multiple paths from the sound source to the sound-receiving point. (14) The acoustic processing method according to any one of (1) to (12), wherein the processing step makes it possible to exclude from the sound-related processing any path that has a sum of sound attenuation due to reflections that satisfies a predetermined criterion from among multiple paths from the sound source to the sound-receiving point. (15) The acoustic processing method according to any one of (1) to (14), wherein the searching step searches for the path of the sound by assuming that, among object surfaces in the virtual space, object surfaces whose characteristics satisfy predetermined criteria are flat and uniform.(16) The acoustic processing method according to any one of (1) to (14), wherein the search step changes the search process or the search result based on information on directivity of the sound at the sound source. (17) The acoustic processing method according to any one of (1) to (16), wherein the search step changes the search process or the search result when at least one of the directivity of the sound at a sound receiving point, the position of the sound receiving point in the virtual space, and the distance between the sound receiving point and an object in the virtual space satisfies a predetermined criterion. (18) The acoustic processing method according to any one of (1) to (17), wherein the search step searches for the path of the sound by ray tracing. (19) An acoustic processing device comprising: a search unit that searches for the path of the sound from the sound source to the sound receiving point in the virtual space; an identification unit that identifies a reflection point of the sound on the path; and a processing unit that performs processing related to the sound at the sound receiving point based on a result of processing related to the reflection of the sound, using an area surrounding the reflection point as a calculation area. (20) An acoustic processing program for causing a computer to function as: a search unit that searches for a sound path from a sound source to a sound receiving point in a virtual space; an identification unit that identifies a reflection point of the sound on the path; and a processing unit that performs processing related to the sound at the sound receiving point based on the results of processing related to the reflection of the sound, with the surrounding area of ​​the reflection point as a calculation area.

[0200] V1, V2 Virtual space 100 Sound processing device 110 Communication unit 120 Storage unit 130 Control unit 131 Acquisition unit 132 Search unit 133 Identification unit 134 Signal processing unit 135 Output control unit 140 Input unit 150 Output unit A1 to A7 Surrounding area P1, P2, P3 Reflection point S, S1 to S7 Sound emission point T, T1 to T7 Sound reception point U User J1 to J5 Object

Claims

1. An acoustic processing method comprising: a search step of searching for a sound path from a sound source to a sound receiving point in a virtual space; an identification step of identifying a reflection point of the sound on the path; and a processing step of performing processing related to the sound at the sound receiving point based on the results of a wave acoustic simulation related to the reflection of the sound, with the surrounding area of ​​the reflection point as the calculation area.

2. The acoustic processing method according to claim 1, wherein the processing step comprises: generating an impulse response at the reflection point as a wave acoustic simulation related to the reflection of the sound; and convolving the impulse response as processing related to the sound at the sound receiving point.

3. The acoustic processing method according to claim 2, wherein the processing step generates a signal having a time period of one frame or less as the impulse response.

4. The acoustic processing method according to claim 2, wherein the waveform of the impulse response is a waveform normalized based on a maximum sound pressure, and the processing step adjusts a gain of the impulse response when performing a convolution process on the impulse response.

5. The acoustic processing method according to claim 2, wherein the processing step generates the impulse response based on characteristic information about the sound at the reflection point.

6. The acoustic processing method according to claim 5, wherein the characteristic information of the reflection point includes at least one of the shape, material, and sound absorption coefficient of the reflection point, the angle from the reflection surface at the reflection point to the sound source, and distance information from the reflection point to the sound source.

7. The acoustic processing method according to claim 5, wherein the processing step generates shape data of the object surface based on texture data of the object surface in the virtual space, and the processing step generates the impulse response based on the shape data.

8. The acoustic processing method of claim 2, wherein the surrounding area is one of two areas obtained by dividing a sphere centered on the reflection point by the sound reflection surface, and the processing step generates the impulse response based on the result of the wave acoustic simulation from a point on the spherical surface of the hemispherical area that is closest to the sound source or the immediately preceding reflection point to a point on the spherical surface of the hemispherical area that is closest to the sound receiving point or the next reflection point.

9. The acoustic processing method according to claim 2, wherein the surrounding area is a semicircular planar area, and the processing step generates the impulse response based on the results of the two-dimensional wave acoustic simulation in the planar area.

10. The acoustic processing method according to claim 9, wherein the semicircular planar area is included in a plane that includes three points: the sound source, the reflection point, and the sound receiving point.

11. The acoustic processing method according to claim 2, wherein the processing step generates the impulse response using a learning model that has learned the relationship between the input and output of the wave acoustic simulation.

12. The acoustic processing method according to claim 11, wherein a plurality of learning models are prepared, and the processing step selects a learning model to be used for generating the impulse response according to the acoustic characteristics of the sound source or metadata assigned to the sound source.

13. The acoustic processing method according to claim 1, wherein the processing step makes it possible to exclude from the processing target for the sound processing, among the multiple paths from the sound source to the sound receiving point, paths in which the number of reflections exceeds a predetermined number.

14. The acoustic processing method according to claim 1, wherein the processing step makes it possible to exclude from the processing target for the sound processing, among a plurality of paths from the sound source to the sound receiving point, paths for which the sum of the amount of sound attenuation due to reflection satisfies a predetermined criterion.

15. The acoustic processing method according to claim 1, wherein the search step searches for the path of the sound by assuming that, among the surfaces of objects in the virtual space, the surfaces of objects whose characteristics satisfy predetermined criteria are flat and uniform.

16. The acoustic processing method according to claim 1, wherein the search step changes the search process or the search result based on information about the directionality of the sound at the sound source.

17. The acoustic processing method according to claim 1, wherein the search step changes the search process or the search result when at least one of the directivity of the sound at the sound receiving point, the position of the sound receiving point in the virtual space, and the distance between the sound receiving point and an object in the virtual space satisfies a predetermined criterion.

18. The acoustic processing method according to claim 1, wherein the searching step searches for the path of the sound by ray tracing.

19. An acoustic processing device comprising: a search unit that searches for a sound path from a sound source to a sound receiving point in a virtual space; an identification unit that identifies a reflection point of the sound on the path; and a processing unit that performs processing related to the sound at the sound receiving point based on the results of processing related to the reflection of the sound, with the surrounding area of ​​the reflection point being used as a calculation area.

20. An acoustic processing program that causes a computer to function as: a search unit that searches for a sound path from a sound source to a sound receiving point in a virtual space; an identification unit that identifies the point on the path where the sound is reflected; and a processing unit that performs processing related to the sound at the sound receiving point based on the results of processing related to the reflection of the sound, with the area surrounding the reflection point as the calculation area.

Citation Information

Patent Citations

  • 2019-165845

  • 2000-267675