Acoustic processing method, acoustic processing apparatus, and acoustic processing program
By limiting the time and space of acoustic simulation, combining fluctuations and geometric acoustic methods to generate diffraction sound, the problems of large calculation volume and low simulation accuracy are solved, and the realism of sound in virtual space is improved.
Patent Information
- Application Number
- CN202380082631.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-05
- Filing Date
- 2023-11-16
- Publication Date
- 2025-07-11
AI Technical Summary
When the existing acoustic processing methods simulate the sound field, wave acoustic simulation calculations are large, making it difficult to efficiently process high-frequency sounds; while geometric acoustic methods are difficult to accurately simulate low-frequency diffraction and scattering phenomena, resulting in unreal sound experience in the virtual space.
By limiting the time or space of acoustic simulations, combining wave acoustic and geometric acoustic methods, searching for diffraction paths and generating diffraction sounds, using finite difference time domain method (FDTD) or physical information neural network (PINN) to reduce the computational volume while fidelity in the sound experience.
While reducing the computational load, the realism and nature of sound in virtual space are improved, especially the acoustic effect simulation accuracy in the low and high frequency ranges.
Smart Images

Figure CN120303956A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an acoustic processing method, an acoustic processing device, and an acoustic processing program. Background Art
[0002] With the development of virtual space technologies such as the metaverse and games, the sounds reproduced in virtual spaces also need to be based on actual realism. There are many studies on audibly identifying sound sources in three-dimensional space and technologies for simulating sound sources. Among them, a spatial sound field reproduction technology that is designed based on physical phenomena in the real space and uses physical simulations, etc., has high reproducibility when reproduced audibly in the virtual space in the real space, and is also excellent in terms of spatial identifiability even when reproducing sounds in a space that does not actually exist.
[0003] Prior Art Documents
[0004] Patent Documents
[0005] Patent Document 1: JP 2000-267675 A
[0006] Patent Document 2: JP 2019-165845 A
[0007] Non-Patent Documents
[0008] Non-Patent Document 1: "Efficient and Accurate Sound Propagation Using AdaptiveRectangular Decomposition" N. Raghuvanshi et al., IEEE (2009)
[0009] Non-Patent Document 2: "Physics-informed neural networks: A deep learningframework for solving forward and inverse problems involving nonlinearpartial differential equations" M. Raissi, P. Perdikaris, G. E. Karniadakis, Journalof Computational Physics 378 (2019) Summary of the Invention
[0010] Technical Problem
[0011] There are various physical simulation methods in the sound field, such as wave acoustic simulation that models and calculates the wave characteristics of sound, and geometric acoustic simulation that geometrically models and calculates the energy propagation of sound.
[0012] It is known that the former well exhibits obvious microwave phenomena and is particularly advantageous in expressing scattering phenomena and the like when sound waves collide with an object having a fine surface shape. Although the space is discretized and calculated, the fineness of the discretization affects the computable frequency, and thus the amount of calculation increases as higher frequencies require finer discretization for calculation.
[0013] As the latter, there are known the ray tracing method that tracks the trajectories of sound rays emitted from a sound source, the virtual image method that assumes sound propagates through geometric reflection at a boundary, and the like. Since spatial partitioning is not required, there is no correlation between the frequency band and the amount of calculation as in wave acoustic simulation. However, since the wave nature of sound is ignored and the energy propagation is geometrically modeled, there is a drawback that it is difficult to simulate phenomena such as diffraction and scattering that are particularly apparent at low frequencies.
[0014] Therefore, the present disclosure proposes an acoustic processing method, an acoustic processing device, and an acoustic processing program that can provide a realistic sound experience with a small amount of calculation.
[0015] Solution to the problem
[0016] According to the present disclosure, an acoustic processing method executed by a computer includes: searching for a diffraction path from a sound source to a sound reception point; and generating diffracted sound obtained at the sound reception point based on wave acoustic simulation in which the simulation time is limited according to the distance of the diffraction path. According to the present disclosure, there is provided an acoustic processing device that performs the information processing defined in the acoustic processing method and an acoustic processing program that implements the information processing defined in the acoustic processing method. Brief description of the drawings
[0017] Figure 1 is a diagram showing an overview of the information processing according to the first embodiment.
[0018] Figure 2 is a flowchart showing the audio signal generation process.
[0019] Figure 3A is a diagram for explaining acoustic simulation in a virtual space.
[0020] Figure 3B is a schematic diagram for explaining acoustic simulation in a virtual space.
[0021] Figure 4 is a diagram schematically showing the reflection state in acoustic simulation.
[0022] Figure 5It is a diagram showing an embodiment of synthesizing and outputting an audio signal.
[0023] Figure 6 It is a flowchart showing an embodiment of processing in the presence of multiple early reflections, diffraction sounds, and transmission sounds.
[0024] Figure 7 It is a diagram for explaining the diffraction phenomenon.
[0025] Figure 8 It is a diagram showing an embodiment of a plug-in representing diffraction by an LPF.
[0026] Figure 9 It is a diagram showing an insertion screen representing the diffraction phenomenon in Wwise.
[0027] Figure 10 It is a diagram showing a configuration example of a virtual space.
[0028] Figure 11 It is a diagram that uses FDTD or the like for simulation, extracts only the diffraction sound from the obtained impulse response, and displays the frequency characteristics.
[0029] Figure 12 It is a diagram showing an embodiment of the configuration of an acoustic processing device.
[0030] Figure 13 It is a diagram showing an embodiment of a processing flow.
[0031] Figure 14 It is a diagram showing an embodiment of a processing flow using an AI model.
[0032] Figure 15 It is a diagram showing an embodiment of a processing flow using the cloud.
[0033] Figure 16 It is a diagram for explaining the characteristics of the audio signal generation process of the second embodiment.
[0034] Figure 17 It is a diagram showing an embodiment of the configuration of a subspace.
[0035] Figure 18 It is a diagram showing an embodiment of a processing flow.
[0036] Figure 19 It is a diagram for explaining an embodiment of processing in the case where a real sound source is not included in the subspace.
[0037] Figure 20 It is a block diagram showing the flow of the diffraction sound generation process.
[0038] Figure 21 It is a diagram showing an embodiment of a processing flow.
[0039] Figure 22 It is a diagram showing an embodiment in which the subspace is configured as a spherical or circular space centered on the diffraction site.
[0040] Figure 23 It is a diagram showing an embodiment in which the real sound source and the real sound receiving point are outside the subspace.
[0041] Figure 24 It is a diagram for explaining an embodiment of a method for setting a subspace for a complex object.
[0042] Figure 25 It is a diagram for explaining an embodiment of a method for setting a subspace for a complex object.
[0043] Figure 26 It is a diagram for explaining the features of the audio signal generation process of the third embodiment.
[0044] Figure 27 It is a diagram showing an embodiment of the directivity of the sound source.
[0045] Figure 28 It is a diagram for explaining the relationship between the size of the pronunciation site and the diffracted sound.
[0046] Figure 29 It is a diagram for explaining a method for setting a subspace in the case where there is a gap between two objects.
[0047] Figure 30 It is a diagram for explaining a method for setting a subspace in the case where there is a gap between two objects.
[0048] Figure 31 It is a diagram for explaining an operation example in the case where the sound source and the sound receiving point are on the same coordinate plane.
[0049] Figure 32 It is a diagram showing an embodiment of distance attenuation in two-dimensional and three-dimensional spaces.
[0050] Figure 33 It is a diagram for explaining the gain adjustment of each diffraction path.
[0051] Figure 34 It is a diagram showing an embodiment in which diffraction paths exist in three directions.
[0052] Figure 35 It is a diagram for explaining an operation example in the case where the sound source and the sound receiving point are not on the same coordinate plane.
[0053] Figure 36It is a schematic diagram for explaining the generation process of diffracted sound through the edge of the side of an object.
[0054] Figure 37 It is a schematic diagram for explaining the generation process of diffracted sound through the edge of the side of an object.
[0055] Figure 38 It is a diagram for explaining the generation process of diffracted sound through the edge of the upper part of an object.
[0056] Figure 39 It is a diagram for explaining the generation process of diffracted sound through the edge of the upper part of an object.
[0057] Figure 40 It is a diagram for explaining the integration process of multiple IRs (impulse responses) with different diffraction paths.
[0058] Figure 41 It is a schematic diagram for explaining another calculation example when the sound source and the sound receiving point are not in the same coordinate plane.
[0059] Figure 42 It is a diagram for explaining an embodiment of IR generation based on encoded object information.
[0060] Figure 43 It is a hardware configuration diagram showing an embodiment of a computer that implements the functions of an acoustic processing device. Detailed Description of the Embodiment
[0061] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In the following embodiments, the same parts are denoted by the same reference numerals, and redundant descriptions will be omitted.
[0062] Note that the description will be given in the following order.
[0063] [1. First Embodiment: Limitation of Simulation Time]
[0064] [1-1. Overview of Information Processing]
[0065] [1-2. Audio Signal Generation Processing by Acoustic Simulation]
[0066] [1-3. Diffraction]
[0067] [1-4. Features of the Audio Signal Generation Processing in the First Embodiment]
[0068] [1-5. Configuration Embodiment of the Acoustic Processing Device]
[0069] [1-6. Embodiment of the Processing Flow 1]
[0070] [1-7. Embodiment of the Processing Flow 2]
[0071] [1 - 8. Embodiment of the processing flow 3]
[0072] [2. Second Embodiment: Restriction on the simulation space]
[0073] [2 - 1. Characteristics of the audio signal generation process in the second embodiment]
[0074] [2 - 2. Embodiment of the subspace configuration]
[0075] [2 - 3. Relationship between the real sound source, real sound receiving point and the subspace]
[0076] [2 - 4. Delay / gain adjustment]
[0077] [2 - 5. Embodiment of the processing flow]
[0078] [2 - 6. Circular / spherical subspace]
[0079] [2 - 7. Embodiment of setting the size of the subspace]
[0080] [3. Third Embodiment: Generation of high - order diffracted sound based on the combination of multiple subspaces]
[0081] [4. Variant 1: Restriction on the audio signal generation process based on the directivity of the sound source][5. Variant 2: Restriction on the audio signal generation process based on the size of the sound - producing part][6. Variant 3: Mesh size of the subspace]
[0082] [7. Modification 4: Example of setting the subspace according to the size of the gap between objects]
[0083] [8. Variant 5: Two - dimensional subspace]
[0084] [8 - 1. Case where the sound source and the sound receiving point are on the same coordinate plane]
[0085] [8 - 2. Case where the sound source and the sound receiving point are not on the same coordinate plane]
[0086] [9. Variant 6: Generation of IR based on the encoded object information]
[0087] [10. Embodiment of the hardware configuration]
[0088] [11. Effects]
[0089] [1. First Embodiment: Restriction on the simulation time]
[0090] [1 - 1. Overview of information processing]
[0091] Figure 1It is a diagram showing an overview of the information processing according to the first embodiment.
[0092] The information processing of this embodiment is executed by Figure 1 the acoustic processing device 100 shown in. The acoustic processing device 100 is an embodiment of an acoustic processing device according to the present disclosure. The acoustic processing device 100 is, for example, an information processing device such as a personal computer (PC), a server device, a game console, or a tablet terminal. The acoustic processing device 100 is used by a user U who uses / creates content related to a virtual space (e.g., a game or the metaverse).
[0093] Hereinafter, the processing of audio information in a virtual three-dimensional space in game content, which is interactive content of user operations, will be mainly described. However, the acoustic processing device 100 of the present disclosure can be used to pre-combine video and audio, generate audio for moving image content reproduced and viewed by the user U, generate content including audio information reproducing a sound field in a virtual space, and present the audio information.
[0094] The acoustic processing device 100 includes an output unit such as a display or a speaker, and outputs various types of information to the user U. For example, the acoustic processing device 100 displays a virtual three-dimensional space (hereinafter, simply referred to as "virtual space") of a game or the like on the display, and outputs an audio signal generated by the information processing of this embodiment from the speaker. Alternatively, the acoustic processing device 100 displays a user interface of software related to acoustic generation on the display, and outputs an audio signal generated according to the operation of the user U from the speaker.
[0095] In this embodiment, the sound processing device 100 calculates how the sound output from a sound source object, which is the sound generation point of the sound, will be reproduced at a listening point (sound reception point) in the virtual space, and reproduces the calculated sound. For example, the acoustic processing device 100 performs an acoustic simulation in the virtual space, and performs processing to make the sound emitted in the virtual space approximate to the real world or reproduce the sound desired by the sound designer.
[0096] Figure 1 The virtual space V in the game is shown. For example, the virtual space V is displayed on a display included in the acoustic processing device 100. In the virtual space V, the position (coordinates) of the listening point (sound reception point) and the sound source object (sound emission point) are set together. The listening point is the position where the user U virtually listens to the sound in the virtual space V1. In the actual space, due to various physical phenomena, a difference occurs between the sound observed near the sound source and the sound observed at the listening point. Therefore, the acoustic processing device 100 virtually reproduces (simulates) the actual physical phenomena in the virtual space V, and generates an audio signal suitable for the space to enhance the realistic feeling of the sound experienced by the user U in the virtual space V.
[0097] Here, the audio signal generated by the acoustic processing device 100 will be described. The graph G schematically shows the intensity of the sound when observing the pulsed sound emitted from the sound source object at the listening point. At the listening point, the direct sound is first observed, and then the diffracted sound of the direct sound, etc. are observed. Then, at the monitoring point, the first reflection sound reflected by the boundary of the virtual space and the transmission sound passing through the object, etc. are observed. Every time the reflection sound is reflected at the boundary, the reflection sound is observed. For example, the first-order to third-order reflection sounds regarded as early reflection sounds are observed. Thereafter, at the listening point, the high-order reflection sounds regarded as late reverberation sounds are observed. Since the sound emitted from the sound source decays over time, the graph G depicts an envelope (attenuation curve) that peaks at the direct sound and asymptotically approaches 0.
[0098] [1-2. Audio signal generation processing by acoustic simulation]
[0099] Hereinafter, the information processing performed by the acoustic processing device 100 will be described in detail with reference to the drawings. First, the general generation processing of the audio signal in the virtual space will be described with reference to Figures 2 to 6 After that, the acoustic processing according to the present embodiment will be described in detail with reference to Figure 7 and the subsequent drawings.
[0100] First, the general generation processing of the audio signal in the virtual space will be described. Figure 2 is a flowchart showing the audio signal generation processing according to the present embodiment. The audio signal generation processing starts with an event in the virtual three-dimensional space (for example, an event that occurs due to the operation of the user U / game progress).
[0101] In the following description, it is assumed that the acoustic processing device 100 performs the audio signal generation processing, but the audio signal generation processing can be executed by one computer or can be executed by multiple computers in cooperation. For example, one computer (for example, the server device) among multiple computers (for example, the user terminal and the server device) connected via a network can execute some steps of the audio signal generation processing, and another computer (for example, the user terminal) can execute the remaining steps.
[0102] In addition, in the following description, a scenario in the game is taken as an example for illustration. Specifically, it is assumed that in Figure 3A and Figure 3B the virtual space V shown. Figure 3A and Figure 3B are diagrams for explaining the acoustic simulation in the virtual space. The virtual space V is a virtual three-dimensional space such as a game. Figure 3A is a perspective view of the virtual space V, Figure 3BThis is a plan view of the virtual space V. Note that in the following description, the XYZ coordinate system can be used for description for ease of understanding. In the figure, the X-axis direction and the Y-axis direction are horizontal directions, and the Z-axis direction is the vertical direction.
[0103] In the virtual space V2, a sound generation point set as the position of the sound source S and a sound reception point T as the listening point are set. The sound source S is an object that emits arbitrary sounds relative to the sound reception point T (i.e., an object that can be a sound source). For example, the sound generation point is the coordinate where the sound source object is configured. In addition, the sound reception point T is the position where the sound output from the sound source S is observed. For example, the sound reception point T is the position where the game character operated by the user U is located (more specifically, the coordinate corresponding to the position of the head of the game character).
[0104] In addition, a plurality of three-dimensional objects OB are arranged in the virtual space V. For example, in the virtual space V, an object J2 as a furniture-type object OB and an object J3 as a human-type object OB are arranged. In addition, in the virtual space V, a wall J1 or the like that constitutes the boundary of the virtual space V is also arranged. In addition, since the sound characteristics described later are also set on the boundary of the wall J1 or the like, in this embodiment, the boundary of the wall J1 or the like is also regarded as one of the objects OB arranged in the virtual space V.
[0105] In Figure 3A and Figure 3B In the embodiment of, a sound is generated at the position of the sound source S. The acoustic processing device 100 generates an audio signal at the sound reception point T while considering the propagation characteristics of the sound in the virtual space V. Hereinafter, the audio signal generation process of this embodiment will be described with reference to the Figure 2 flowchart of.
[0106] First, the acoustic processing device 100 acquires the audio data of the sound emitted from the sound source S (step S101). The acoustic processing device 100 acquires the audio data that pre-records the sound to be emitted at the start of the sound emission event of the sound source S. The audio data is, for example, the audio data recorded in the library in the game. Note that, considering being subjected to signal processing in a subsequent stage, the audio data is preferably data that does not include reverberation, etc. (dry source). On the other hand, a simple sound rendering that omits a part of the signal processing in a subsequent stage or a wet source that adds reverberation, etc. to express a characteristic tone can be used. In addition, the acoustic processing device 100 can not only refer to the audio data from the library, but also acquire the audio data generated as needed (such as a synthesizer). Alternatively, the acoustic processing device 100 can acquire the audio data generated as needed by simulating the sound from the structural physical simulation using the finite element method (FEM), etc.
[0107] Next, the acoustic processing device 100 acquires spatial data of the space where the sound source S and the sound reception point T are located (step S102). For example, the acoustic processing device 100 acquires three-dimensional data of the virtual space V and metadata (e.g., material information) added to the three-dimensional data. The three-dimensional data includes, for example, coordinates representing the boundary shape of the wall surface, etc., of the virtual space V and coordinates of objects arranged in the virtual space V. When the sound source S and the sound reception point T are in a closed space, the acoustic processing device 100 acquires three-dimensional data of the entire closed space or three-dimensional data of the entire space on the path from the sound source S to the sound reception point T and the entire space continuous with this space. It should be noted that since the data volume increases when the three-dimensional data includes detailed textures for video display, the acoustic processing device 100 can acquire data in a form that simplifies color information, the fine shape of the surface, etc., so as to maintain the acoustic parameters required for physical simulation in the subsequent stage. The three-dimensional data for path search described later is preferably polygon information that does not include fine surface irregularities (e.g., polygon information to which no texture is applied).
[0108] Next, the acoustic processing device 100 calculates the sound path from the sound source S to the sound reception point T (step S103). For example, the acoustic processing device 100 calculates the propagation path of the sound from the sound source S to the sound reception point T using the acquired three-dimensional data. At this time, the direct sound path is the line-of-sight path from the sound source S to the sound reception point T. It should be noted that in the case including a transmission phenomenon such as a wall, even when there is no line-of-sight path, a path including obstacles is calculated. In Figure 3A and Figure 3B the embodiment of, the propagation path of the direct sound from the sound source S to the sound reception point T is represented by a solid line.
[0109] In addition, the acoustic processing device 100 calculates a path including a reflection boundary, etc., until the early reflected sound reaches the sound reception point T. For example, the acoustic processing device 100 obtains the propagation path of the early reflected sound using a ray method or the like as a geometric physical simulation method. Further, in the present embodiment, the early reflected sound includes not only the sound that reaches the sound reception point T after being reflected once on the boundary surface, but also the sound that reaches the sound reception point T after being reflected a small number of times such as 2 or 3 times.
[0110] Figure 4 An embodiment of the propagation path calculated by the acoustic processing device 100 is shown. Figure 4 is a diagram schematically showing the reflection state in the acoustic simulation. In Figure 4In the illustrated embodiment, the path R1 is the direct sound propagation path from the sound source S to the sound receiving point T. Additionally, the path R2 is the propagation path of the first-reflected sound that is reflected once at the boundary B1 and reaches the sound receiving point T. Further, the path R3 is the propagation path of the second-reflected sound that is reflected twice at the boundaries B2 and B3 and reaches the sound receiving point T. In Figure 4 , the paths R1 to R3 are shown as propagation paths. However, the actual calculation includes reflections on the ground, ceiling, other wall surfaces, object surfaces, etc.
[0111] In addition to the shortest path from the sound source S to the sound receiving point T, diffracted sound reaching the sound receiving point T from the sound source S is obtained by bypassing the end of the obstacle located between the sound source S and the sound receiving point T. In Figure 3A and Figure 3B 's example, the object J3 is an example of an obstacle located between the sound source S and the sound receiving point T. In this case, for example, the acoustic processing device 100 can obtain a detour path that bypasses the obstacle in an arc, and can also obtain a detour path by connecting the sound source S and the end of the object J3 and the end of the object J3 and the sound receiving point T with a straight line.
[0112] When setting directivity for the sound source, the acoustic processing device 100 can omit sound rays in the direction where there is no sound radiation according to the directivity, or can change the density of the sound rays according to the radiation intensity. By performing this process, the subsequent processing performed on each path can be reduced while reducing the impact on audibility.
[0113] Based on the obtained information, the acoustic processing device 100 generates an audio signal to be heard at the sound receiving point T. For example, the acoustic processing device 100 performs generation processes such as direct sound generation, early reflected sound generation, diffracted sound generation, and late reverberant sound generation. Hereinafter, these generation processes will be described.
[0114] First, the acoustic processing device 100 generates an audio signal based on the direct sound in the audio signal heard at the sound receiving point T (step S104). Specifically, the acoustic processing device 100 uses the acquired audio data as input and uses the spatial size and directivity of the sound source, the attenuation according to the distance from the sound source to the listening point, and the observation time offset due to the propagation time as parameters to generate an audio signal. For example, when the sound source S is an omnidirectional point source and the distance between the sound source S and the sound receiving point T is X, the attenuation amount is obtained by 20log10x. Further, the propagation time is obtained by dividing the distance x by the speed of sound c. It should be noted that when the content creator expects to emphasize the expression of the distance feeling between the sound source S and the sound receiving point T, the acoustic processing device 100 can change the attenuation amount without being restricted by the physical phenomenon in the real space.
[0115] Subsequently, the acoustic processing device 100 generates an audio signal based on the early reflected sound in the audio signal heard at the sound reception point T (step S105). Specifically, similar to the case of the direct sound, the acoustic processing device 100 uses the acquired audio data as input and calculates the attenuation and time shift corresponding to the propagation path length in the medium based on the pre-calculated propagation path. In addition, the acoustic processing device 100 performs signal processing that simulates reflection at the boundary surface based on the information of the boundary surface that reflects the input signal. For example, the acoustic processing device 100 can apply a predetermined attenuation to each frequency according to the sound absorption coefficient of the boundary surface of the reflected input signal.
[0116] For example, for the reflected sound corresponding to the Figure 4 path R2 shown in, the part from the sound source S to the boundary B1 and the part from the boundary B1 to the sound reception point T can be regarded as the propagation paths of the medium. The acoustic processing device 100 obtains the attenuation and time shift according to each distance of the propagation path. In addition, the sound processing device 100 can also calculate the attenuation by appropriately reading out data by previously storing data such as the sound absorption coefficient data of each frequency in a library for the boundary B1. The acoustic processing device 100 can generate an audio signal of the reflected sound corresponding to the path R2 by giving the attenuation of the input signal according to the sound absorption coefficient together with the attenuation on the previous propagation path. At this time, without considering the influence of the plate vibration or resonance at the boundary B1 and the nonlinear phenomenon, the acoustic processing device 100 can calculate the attenuation, etc. by adding the part from the sound source S to the boundary B1 and the propagation path from the boundary B1 to the sound reception point T.
[0117] Similarly, the acoustic processing device 100 performs the calculation of the reflected sound corresponding to the Figure 4 path R3 shown in. That is, the acoustic processing device 100 performs calculations based on the propagation paths from the sound source S to the boundary B2, from the boundary B2 to the boundary B3, and from the boundary B3 to the sound reception point T. When calculating the sound reflected at a plurality of boundary surfaces such as the boundary B2 and the boundary B3, the acoustic processing device 100 can jointly calculate the attenuation according to the sound absorption coefficient without considering the nonlinear behavior according to the sound pressure, etc. at the boundary surface. For example, digital filters such as infinite impulse response (IIR) or finite impulse response (FIR) can be used to calculate the attenuation of each frequency according to the sound absorption coefficient. Note that in the process of generating the early reflected sound, the acoustic processing device 100 can appropriately simulate the reflection state on each wall surface using a method such as ray tracing.
[0118] Subsequently, the acoustic processing device 100 generates an audio signal related to the diffracted sound (step S106). Similar to the early reflected sound, the acoustic processing device 100 generates an audio signal related to the diffracted sound by using a geometric method that calculates attenuation or time shift based on the propagation path. Note that in the sound diffraction phenomenon, it is known that the attenuation amount of the diffracted sound increases as the frequency increases. Therefore, the acoustic processing device 100 can generate the diffracted sound by applying a filter with different attenuation amounts according to the frequency. Note that if a more accurate method based on the physical laws of wave phenomena is adopted, the acoustic processing device 100 can use not only the above geometric simulation method but also the wave acoustic simulation method from the sound source S to the sound reception point T. In this case, since there is a possibility that the amount of calculation increases compared to the geometric simulation, the user U can use a device with sufficient computing power as the acoustic processing device 100 or prepare a library in which the characteristics of the representative propagation path shape are pre-calculated.
[0119] Subsequently, the acoustic processing device 100 generates the late reverberation (step S107). The late reverberation refers to the part of the sound emitted from the sound source S that reaches the sound reception point T through repeated reflections and diffractions in space, excluding the early reflected sound. This reflection or diffraction continues until the sound is completely attenuated (for example, the state where the quantization error on the computer can be regarded as 0). Here, the state where the sound is attenuated can be the state of the user interface of the acoustic processing device 100 when the listener operates it with respect to the reproduced sound including the late reverberation generated on the acoustic processing device 100, and this state is determined as the discrimination limit in the listener's perception.
[0120] In addition, in the simulation, it is also possible to calculate until the sound is completely attenuated. However, since the amount of calculation increases with each number of reflections, the acoustic processing device 100 can calculate the early reflected sound by regarding up to a predetermined number of reflections as the early reflection. Then, the acoustic processing device 100 can calculate the subsequent sound by another method and synthesize the calculated sound with the early reflected sound, etc. Here, the predetermined number of reflections included in the early reflected sound can be set on the creator side, or can be dynamically changed and set in the acoustic processing device 100 according to the content to be reproduced, the type of virtual space, the size of the space, the boundary surface in the space, the sound absorption coefficient on the object surface, etc. It should be noted that as a method for calculating the late reverberation sound, a method of calculating the reverberation time statistically according to the size of the space, the boundary surface in the space, the sound absorption coefficient on the object surface, etc. is also known in the field of architectural acoustics.
[0121] Regarding the generation of late reverberation sound, the acoustic processing device 100 can use methods such as the comb filter method, the method of convolving the input signal with the data of the impulse response stored in a library, etc. Among them, the acquired audio data is input into the comb filter, feedback is performed on the input by multiplying by a predetermined gain, and the input signal is repeated at a constant period while being attenuated. It should be noted that the acoustic processing device 100 can use other known techniques to generate late reverberation sound. As described above, the late reverberation sound has a great impact on the audible recognition space of the user and is therefore of great significance in content generation. In many existing modular reverberation generators, the late reverberation level, late reverberation delay time, decay time, ratio of the amount of attenuation for each frequency, echo density, modal density, etc. can be adjusted as the late reverberation sound settings (parameters).
[0122] The acoustic processing device 100 generates direct sound, early reflection sound, diffraction sound, and late reverberation sound respectively, and then synthesizes these signals. Then, the acoustic processing device 100 outputs the synthesized audio signal as the sound observed at the sound reception point T to a speaker or the like (step S108). Figure 5 It is a diagram showing an example of synthesizing and outputting an audio signal. When the output is completed, the acoustic processing device 100 ends the audio signal generation process.
[0123] It should be noted that the generation process according to the above steps S104 to S106 does not depend on the order between the respective steps and can therefore be replaced. In addition, when considering the case where diffraction occurs after reflection in the propagation path, the signal generated by first calculating the reflection attenuation from the sound source to the boundary can be used as the input for calculating the diffraction sound.
[0124] It should be noted that although not considered in the embodiment of FIG. 3, the acoustic processing device 100 can add processing using the projection simulation of the sound on the boundary surface or the port simulation technology for simulating the characteristics when a small sound passes through a space (for example, a window or a door of a building). As a result, the acoustic processing device 100 can achieve a sound expression closer to the real space phenomenon.
[0125] In addition, in the case where there are multiple reflection sounds, multiple diffraction sounds, and multiple projection sounds on the path, the acoustic processing device 100 can perform these calculations for each location where each phenomenon occurs, and can synthesize and output the audio signal after performing all the calculations. Figure 6 It is a flowchart showing an embodiment of the processing in the case where there are multiple early reflection sounds, diffraction sounds, and transmission sounds. Figure 6 The audio signal generation process shown in can be executed by one computer or can be executed by multiple computers in cooperation.
[0126] Compared with Figure 2 the flowchart shown, steps S109 and S110 are re-added toFigure 6 The flowchart shown. That is, the acoustic processing device 100 processes transmitted sound in addition to direct sound, early reflected sound, and diffracted sound. Specifically, the acoustic processing device 100 generates an audio signal based on the diffracted sound, and then generates an audio signal based on the transmitted sound (step S109). Then, the acoustic processing device 100 determines whether all of the early reflected sound, diffracted sound, and transmitted sound have been calculated (step S110). If all values have not been calculated (step S110: No), the acoustic processing device 100 returns the processing to step S105. If all values have been calculated (step S110: Yes), the acoustic processing device 100 advances the processing to step S107. Note that the generation processes in steps S104, S105, S106, and S109 described above can be interchanged.
[0127] Note that in the case where there are multiple sound sources S, the acoustic processing device 100 can perform the entire above-described processing in parallel for the sounds output from the multiple sound sources S. Alternatively, the acoustic processing device 100 can delay the time until output until the entire processing is completed, perform the processing sequentially, and further synthesize and output each synthesized signal that is time-sequentially aligned with the arriving sound at the sound reception point T.
[0128] [1-3. Diffraction]
[0129] The overall flow of signal generation and reproduction has been described above. Hereinafter, the diffracted phenomenon, which is the gist of the present disclosure, will be described in detail.
[0130] In the present disclosure, attention is focused on realistically reproducing a phenomenon such as diffraction, in which when a sound wave impinges on an object OB such as an obstacle, the sound wave bends around to the back side of the obstacle, or fine scattering and interference generated near where diffraction occurs using a wave acoustics simulation method in the above-described audio signal generation process. As described above, wave acoustics simulation has a problem of high computational load, but in the present disclosure, in order to improve such a problem, time or space restrictions are imposed on the wave acoustics simulation. This can improve the huge amount of calculation in wave acoustics simulation and provide the user U with a diffracted phenomenon with a more natural timbre.
[0131] First, the diffracted phenomenon will be described. Figure 7 is a diagram for explaining the diffracted phenomenon. The diffracted phenomenon means as Figure 7As shown, in the presence of an object OB as an obstacle, there is a phenomenon where sound bends around to the back of the object OB. The edge of the object OB is a diffraction part DF. The diffraction part DF refers to a continuous region where the diffraction phenomenon occurs. The diffraction phenomenon is related to the wavelength of sound and has the characteristic that the attenuation amount decreases as the frequency decreases and increases as the frequency increases. In a virtual space V such as a game, there are many scenarios where there is an obstacle between a sound source S and a sound reception point T. It is important to reproduce and generate diffracted sound without discomfort.
[0132] The sound experience of a user U in a virtual space V such as a game is usually generated by geometric-acoustics-based simulations such as the virtual image method or the ray-tracing method. Although geometric simulations based on acoustics have the advantage of a smaller computational amount compared to wave simulations based on acoustics, there are the following disadvantages: it is difficult to simulate phenomena such as diffraction and scattering that are particularly prominent at low frequencies because the simulation is a method of geometrically modeling energy propagation while ignoring the wave characteristics of sound.
[0133] As a method for expressing diffraction based on geometric acoustics, a geometric-acoustics model using the uniform theory of diffraction (UTD) was disclosed as "Modeling Acoustics in Virtual Environments Using the Uniform Theory of Diffraction, N. Nicolas, T. Funkhouser, A. Ngan, I. Carlbom" in 2001. However, UTD is a theoretical model that assumes that the edge of the obstacle where diffraction occurs has an infinite length, and since there is actually no obstacle with an infinite-length edge, this theoretical model causes a larger error as the edge becomes shorter.
[0134] As described above, models using UTD have been studied, but in audio middleware used in actual games, etc., the diffraction phenomenon is simply implemented by a simple low-pass filter (LPF) of the infinite impulse response (IIR) type. As described above, since the diffraction phenomenon has a small sound attenuation amount at low frequencies and a large sound attenuation amount at high frequencies, the diffraction phenomenon can be easily implemented by an LPF.
[0135] Figure 8 An embodiment of a plug-in representing diffraction by an LPF is shown. In this LPF, frequencies below 1 kHz are set to directly have volume, and frequencies above 1 kHz are set to decrease by -6 dB each time the octave increases. In addition, Figure 9 A plug-in screen representing the diffraction phenomenon in the audio middleware Wave Works Interactive Sound Engine (Wwise) used in games, etc., is shown. Although it is the same as Figure 8Although it may seem different at first glance, it can be seen that the vertical axis represents the amount of LPF applied, the horizontal axis represents the size of the obstacle, and diffraction is represented by the LPF.
[0136] As described above, the diffraction phenomenon is usually reproduced by a simple LPF, and the diffracted sound generated by the LPF is significantly different from the actual diffracted sound and causes discomfort, which reduces the immersion of the user U. In a game, when a character moves behind an obstacle, due to the LPF, the sound suddenly changes, so there are many scenes where users feel uncomfortable.
[0137] [1-4. Features of the audio signal generation process of the first embodiment]
[0138] As described above, it is difficult to well represent the diffracted sound by using the geometric acoustics-based simulation using UTD. Therefore, in the present invention, the diffracted sound is generated by a wave equation-based acoustic wave-based simulation. In a three-dimensional orthogonal coordinate system, the wave equation is represented by the following equation (1).
[0139]
[0140] Here, p is the sound pressure, c is the speed of sound, t is the time, and x, y, and z are variables in the three-dimensional orthogonal coordinate system. As a simulation method based on this wave equation, there is the finite-difference time-domain method (FDTD). FDTD can obtain the solution by discretizing the space and time to be simulated, expanding each differential term of the wave equation into a differential equation, and performing sequential calculations. In FDTD, calculations are performed using a staggered grid that gives the sound pressure and particle velocity in a semi-integer coordinate system. Here, FDTD has been described as a simulation method based on the wave equation in the time domain, but the finite element method (FEM) or the boundary element method (BEM) can also be used.
[0141] Since the simulation based on wave acoustics can express the diffraction phenomenon, a more natural tone can be provided to the user than the diffracted sound based on the conventional geometric acoustics basis or created by IIR. For example, Figure 11 is shown in Figure 10 the frequency characteristics from the sound source S to the sound reception point T in the virtual space V. Figure 10 is a diagram showing an example of the structure of the virtual space V. Figure 11 is a diagram showing the simulation performed using FDTD or the like, extracting only the diffracted sound from the obtained impulse response (IR) and displaying the frequency characteristics.
[0142] As can be seen from Figure 10 and Figure 11 , even such a simple diffraction phenomenon has complex frequency characteristics. The reason for having such complex frequency characteristics is that fine scattering or reflection occurs at the edge part where diffraction occurs. It can be seen that the frequency characteristics are related to Figure 8There are significant differences in the frequency characteristics representing the diffraction phenomenon through the LPF. As described above, by simulating the diffraction phenomenon based on wave acoustics, a more realistic natural diffraction sound can be presented to the user U.
[0143] Generally, in FDTD etc., the entire space is discretized, and before a certain time point including late reverberation, the acoustic pressure and particle velocity of the entire space are calculated step by step for each time step. Therefore, the calculation amount increases, and it is difficult to generate the IR in real time and convolve the IR with the sound source in games etc. Here, real time means that the processing can be executed within 16 ms obtained according to, for example, the frame rate of 60 fps often used in games.
[0144] Since the present disclosure focuses on the diffraction sound, it is sufficient to perform the simulation until the time when the diffraction sound can be observed. The time when the diffraction sound can be observed can be calculated based on the distance (path length) x of the diffraction path DR from the sound source S to the sound reception point T. The diffraction path DR represents the propagation path of the sound output from the sound source S after diffraction at the diffraction part DF to the sound reception point T. For example, the time t0 of the first observed diffraction sound can be obtained by dividing the distance x by the sound speed c (t0 = x / c). Therefore, if the simulation is performed until about t = 2 × t0 seconds, a diffraction phenomenon including fine scattering and reflection components at the edge can be obtained.
[0145] The system developer can arbitrarily set the length of the simulation time. For example, the time until the diffraction sound first reaches the sound reception point T can be calculated as the arrival time (t0) based on the distance x of the diffraction path DR, and a time equal to or greater than 1 times and equal to or less than 5 times (preferably, equal to or less than 3 times) the arrival time can be set as the simulation time. Optionally, since processing such as convolution of the sound source and the IR is often calculated in the frequency domain using FFT, the simulation time can be set to a sampling number with a power of 2. For example, the simulation time can be set to a sampling number of 128 samples or more and 131072 samples or less. In the case of a sampling rate of 48 kHz, this value is about 2.7 ms or more and 2.7 seconds or less. In addition, the simulation time can be set by a method other than the above method.
[0146] Therefore, the IR in a specific space is simulated before the time step (about 2 × t0) when the diffraction phenomenon can be observed, and the IR of about 2 × t0 seconds is held as a table. Therefore, only the process of convolving the time signal of the sound source S with the IR is performed, and a natural diffraction sound can be presented to the user U in real time. Hereinafter, the configuration and the acoustic processing method of the acoustic processing device 100 of the present disclosure will be specifically described.
[0147] [1-5. Configuration Embodiment of Acoustic Processing Device]
[0148] Figure 12 FIG. is a diagram showing an example of the configuration of the acoustic processing apparatus 100.
[0149] The acoustic processing apparatus 100 includes a communication unit 110, a storage unit 120, a control unit 130, and an output unit 140. Note that the acoustic processing apparatus 100 may include an input device (e.g., a pointing device such as a touch panel, keyboard, or mouse, a voice input microphone, or an image input camera (gaze or gesture input)) that performs various operation inputs from a user U who operates the acoustic processing apparatus 100, etc.
[0150] The communication unit 110 is implemented by, for example, a network interface card (NIC), etc. The communication unit 110 is connected to a network N (Internet, near field communication (NFC), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.) in a wired or wireless manner, and transmits information to other information devices, etc. via the network N and receives information from other information devices, etc.
[0151] The storage unit 120 is implemented by, for example, a semiconductor storage element (such as a random access memory (RAM) or a flash memory) or a storage device (such as a hard disk or an optical disk). The storage unit 120 stores various types of data, such as audio data output from a sound source S, shape data of an object, preset reverberation settings, and sound absorption coefficient settings.
[0152] The control unit 130 is implemented by, for example, a central processing unit (CPU), a microprocessing unit (MPU), a digital signal processor (DSP), etc. that execute a program (e.g., an acoustic processing program according to the present disclosure) stored inside the acoustic processing apparatus 100 using a RAM, etc. as a working area. In addition, the control unit 130 is a controller and may be implemented by, for example, an integrated circuit (such as an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA)).
[0153] The control unit 130 generates an audio signal according to the acoustic processing program and outputs the audio signal to the output unit 140. The output unit 140 includes a headphone 150 and a speaker 160. The output unit 140 outputs audio data to the user U via the headphone 150 and the speaker 160.
[0154] The control unit 130 includes a path search unit 131, an identification unit 132, a signal processing unit 139, and an audio data reproduction unit 135, and implements or executes the functions and effects of the information processing described below. Note that the internal configuration of the control unit 130 is not limited to the configuration shown in Figure 12 and may be another configuration as long as it executes the information processing described later.
[0155] <1. Path Search Unit>
[0156] The path search unit 131 searches for the shortest path from the sound source S to the sound reception point T. For example, Figure 10 the shortest path in represents the path along which the sound output from the sound source S diffracts at the diffraction site DF and reaches the sound reception point T. In Figure 10 , the shortest path is the same as the diffraction path DR. For example, the path search unit 131 searches for the shortest path based on the space information, object information, sound source information, and sound reception point information obtained from the storage unit 120.
[0157] For example, the space information may include CAD data of the virtual space V arbitrarily created by the creator as shown in Figure 1 . The space information may include CAD data of the virtual space V in a specific scene of a particular game in which the user U is operating a character such as in a game. In addition to the CAD data, the space information may include voxel data, mesh data, point cloud data, etc.
[0158] The object information includes, for example, information about the position, texture, material, size, density, internal sound velocity, sound absorption coefficient, acoustic impedance, etc. of the object OB. The texture information includes information about the surface shape of the object OB. Here, the position can be represented by a Cartesian coordinate system, a cylindrical coordinate system, or a spherical coordinate system.
[0159] For example, the sound source information includes the position of the sound source S, the directivity characteristics, the time waveform data of the sound source S, the direction pointed by the sound source S, etc. The sound source information may include information about the type of sound. In the case where the sound is a line, data of the character who utters the line or metadata such as an angry sound or a laughing sound may be included in the sound source information. Here, the method of representing the position of the sound source S can be a Cartesian coordinate system, a cylindrical coordinate system, or a spherical coordinate system.
[0160] For example, the sound reception point information includes the position of the sound reception point T, the directivity characteristics, the direction in which the sound reception point T (the listener) is facing, and the head-related transfer function (HRTF) expressing the characteristics of both ears from the sound source S to the listener. The position of the sound reception point T can be represented by a Cartesian coordinate system, a cylindrical coordinate system, or a spherical coordinate system.
[0161] As a method for path search, for example, a geometry-based method such as ray tracing is used. The path search unit 131 searches for the shortest path from the position of the sound source S to the position of the sound reception point T and outputs the searched path as path information.
[0162] <2. Identification unit>
[0163] The recognition unit 132 identifies the position where diffraction occurs (diffraction part DF) based on the path information output from the path search unit 131. The recognition unit 132 identifies whether an object exists on the path, that is, whether there is a diffraction part DF, according to the input spatial information, object information, and path information. When there is no diffraction part DF, no diffraction phenomenon occurs, and thus, no diffraction sound generation process is performed. When there is a diffraction part DF, the shortest path including the diffraction part DF is the diffraction path DR, and the signal processing unit 139 performs a diffraction sound generation process.
[0164] <3. Signal Processing Unit>
[0165] The signal processing unit 139 generates the diffraction sound obtained at the sound reception point T based on a wave acoustics simulation. Although the wave acoustics simulation can reproduce a real sound field, it is difficult to apply the wave acoustics simulation to games and the like that require real-time characteristics due to high computational load. In the present disclosure, such a problem is improved by restricting the time or space or both the time and space in which the simulation is executed. For example, the signal processing unit 139 restricts the simulation time according to the distance of the diffraction path DR and generates the diffraction sound obtained at the sound reception point T based on the wave acoustics simulation.
[0166] The signal processing unit 139 includes an IR generation unit 133 and a convolution unit 134. The IR generation unit 133 obtains an impulse response (IR) representing the propagation information of the sound from the sound source S to the sound reception point T. The convolution unit 134 convolves the IR obtained from the IR generation unit 133 with the time waveform data of the sound source S. As a result, the diffraction sound obtained at the sound reception point T is produced.
[0167] For example, the storage unit 120 stores a table that defines the IR for each sound reception point T. The IR is pre-calculated based on the sound source information, object information, and sound reception point information. For the calculation, methods such as FDTD, FEM, or BEM are used. The IR generation unit 133 detects the positional relationship among the sound source S, the sound reception point T, and the diffraction part DF. The IR generation unit 133 reads the IR data corresponding to the detected positional relationship from the table. The convolution unit 134 convolves the read IR with the time waveform data of the sound source S. The convolution unit 134 may further convolve a time signal representing HRTF in order to present a more realistic sound.
[0168] <4. Audio Data Reproduction Unit>
[0169] The audio data reproduction unit 135 reproduces, from a reproduction device (e.g., the headphones 150 or the speaker 160 used by the user U), the diffraction sound generated by the convolution unit 134 and the Figure 6The synthesized output sound obtained from the early reflected sound, transmitted sound, and late reverberant sound generated in each step. As a result, the sound is presented to the user U.
[0170] [1-6. Processing Flow Example 1]
[0171] Figure 13 It is a diagram showing an embodiment of the processing flow.
[0172] The acoustic processing device 100 acquires spatial information, object information, sound source information, and sound reception point information from an input device (not shown) (step S201). The path search unit 131 outputs path information of the shortest path from the sound source S to the sound reception point T using a geometric shape-based method such as ray tracing (step S202).
[0173] The identification unit 132 identifies the position where diffraction occurs on the path (diffraction part DF) based on the path information output from the path search unit 131 (step S203). For example, the identification unit 132 identifies whether there is an object OB on the path, that is, whether there is a diffraction part DF, according to the input spatial information, object information, and path information (step S204). In the case where there is no diffraction part DF (step S204: No), the diffraction phenomenon does not occur, so this process ends (step S208).
[0174] In the case where there is a diffraction part DF (step S204: Yes), the path including the diffraction part DF becomes a diffraction path DR, and the IR generation unit 133 reads the corresponding IR from a table that holds the IRs pre-calculated using FDTD, FEM, or BEM (step S205). The convolution unit 134 convolves the read IR with the time waveform data of the sound source S to generate a diffracted sound (step S206). The convolution unit 134 can further convolve the time signal representing the HRTF to present a realistic sound to the user U.
[0175] The audio data reproduction unit 135 reproduces the synthesized output sound including the diffracted sound generated by the convolution unit 134 using a reproduction device (e.g., headphones 150 or speakers 160) used by the user U (step S207). As a result, the sound is presented to the user U. It should be noted that the reproduction device is not limited to the above examples and can be headphones, a head-mounted display (HMD), etc.
[0176] [1-7. Processing Flow Example 2]
[0177] In the above processing flow example 1, the table for storing IR is used for IR generation processing. In this method, since the table needs to hold IR data representing the diffraction phenomena of various objects OB, the data volume becomes huge. Therefore, a method of using a pre-trained AI model to generate IR instead of the table can also be considered.
[0178] Figure 14 FIG. is a diagram showing an embodiment of a processing flow using an AI model. The difference from the processing flow example 1 is that the IR generation unit 133 uses the AI model to generate IR (steps S209 and S210).
[0179] The AI model is a trained neural network that learns the relationship between the simulation conditions of the wave acoustics simulation and the simulation results obtained under the simulation conditions as the relationship between the input data and the correct answer data.
[0180] For example, the AI model can be a model that outputs IR when input with sound source position information, object information, sound receiving point information, the angle information from the sound source S to the edge (diffraction field DF) of the object OB, the angle information from the edge to the sound receiving point T, and the like. For example, based on the path information output from the path search unit 131 and the information of the diffraction part DF output from the recognition unit 132, the angle information (incident angle information) from the sound source S to the edge and the angle information (exit angle information) from the edge to the sound receiving point T can be calculated.
[0181] The AI model can be a physics-informed neural network (PINN) model that incorporates physical information such as the wave equation. For PINN, refer to "Physics-Informed Neural Networks: A Deep Learning Framework for Solving Forward and Inverse Problems Involving Nonlinear Partial Differential Equations, M. Raissi, P. Perdikaris, G. E. Karniadakis, 2019".
[0182] The IR generation unit 133 prepares the sound source information, object information, sound receiving point information, etc. as input information (step S209). The IR generation unit 133 calculates the incident angle information and the exit angle information based on the path information and the information of the diffraction part DF. The IR generation unit 133 inputs the input information, the incident angle information, and the exit angle information into the AI model, and outputs the IR corresponding to the diffraction part DF (step S210). The steps other than steps S209 and S210 are similar to those in the processing flow example 1.
[0183] It should be noted that in Process Example 1 and Process Example 2, a pre-computed IR is used to generate an IR corresponding to the scenario. This method is effective in cases where it is difficult to perform the huge calculations required for simulation sequentially. However, in cases where various operations can be processed in real time due to the development of computing resources, it is not necessarily required to maintain the tables of the IR and the AI model. In such a case, the IR can be generated by inputting the spatial information, object information, sound source information, and sound reception point information of the virtual space V for each scenario and directly performing wave-based simulation.
[0184] To save computing resources, it is preferable to limit the simulation time according to the distance of the diffraction path DR. For example, the IR generation unit 133 calculates the time until the diffracted sound first reaches the sound reception point T based on the distance of the diffraction path DR as the arrival time t0. The IR generation unit 133 sets the simulation time to, for example, a time equal to or greater than 1 times the arrival time t0 and equal to or less than 5 times the arrival time t0. Alternatively, since processing such as convolution of the sound source and the IR is often calculated in the frequency domain using FFT, the simulation time can be set to the number of samples having a power of 2. For example, the simulation time can be set to the number of samples above 128 samples and below 131072 samples. In the case where the sampling rate is 48 kHz, this value is about 2.7 milliseconds or more and 2.7 seconds or less.
[0185] [1-8. Process Example 3]
[0186] The three modes of "maintaining a table of IR", "AI model for outputting IR", and "directly using wave acoustics simulation" have been described above. When considering games, etc., it is necessary to ensure the allocation of memory for storing the table of IR, the AI model, and the algorithm for wave acoustics simulation such as FDTD on the CPU or GPU. In the case of reducing memory occupancy, it is conceivable to process some or all of the necessary operations in the cloud.
[0187] Figure 15 It is a diagram showing an example of a process flow using the cloud. The difference from Process Example 1 and Process Example 2 is that the cloud is used to generate the IR (step S211).
[0188] In this method, the table of the IR, the AI model, or the table of an algorithm such as FDTD for wave acoustics simulation is stored in a server on the cloud. The server acquires the sound source information, object information, sound reception point information, and path information from a terminal (client terminal) such as a game console serving as a client, and calculates the IR corresponding to the diffracted part (step S211). The client convolves the time waveform data of the IR acquired from the server with the sound source S and presents the sound to the user U. The steps other than step S211 are similar to those in Process Example 1.
[0189] In Figure 15 the embodiment, only the calculation of IR is performed by the server, while the convolution operation is performed by the client terminal. However, both the IR calculation and the convolution operation can be performed by the server. In this case, the client transmits the time waveform data of the sound source S to the server together with the sound source information, object information, sound reception point information, and path information. The server performs a convolution operation on the time waveform data of the corresponding IR and the sound source S to generate diffracted sound, and transmits the diffracted sound to the client.
[0190] In this method, the client terminal only needs to perform the process of playing sound to the user U. Therefore, the computational load of the client terminal can be reduced. The above method can be implemented when the communication between the client terminal and the server is executed at high speed at a level where discomfort is not generated during the game.
[0191] [2. Second Embodiment: Limitation of the Simulation Space]
[0192] Generally, in FDTD or the like, the entire space is discretized, and the sound pressure and particle velocity of the entire space are calculated step by step in time until a certain time point including late reverberation. Therefore, the computational load increases. In the first embodiment, by focusing on the fact that the simulation is only performed until the time when the diffracted sound can be observed, the simulation time is limited according to the distance of the diffraction path DR. In the second embodiment, the simulation is only performed in the subspace where the diffraction phenomenon occurs, and the entire space is not discretized.
[0193] [2-1. Features of the Audio Signal Generation Process of the Second Embodiment]
[0194] Figure 16 is a diagram for explaining the features of the audio signal generation process of the second embodiment.
[0195] This embodiment is different from the first embodiment in that the space for performing the wave acoustics simulation is limited to the periphery of the diffraction part DF. The IR generation unit 133 acquires the diffraction part DF on the object OB where the diffraction phenomenon occurs. The IR generation unit 133 generates IR in the subspace SB by selectively performing the wave acoustics simulation on the subspace SB including the diffraction part DF.
[0196] In a general wave acoustics simulation, since the entire region of the virtual space V (the entire space TS) is simulated, the computational load is high (see Figure 16(left figure in). In this embodiment, a small space around the diffraction part DF in the entire space TS is set as the subspace SB to be simulated. For example, the storage unit 120 stores the size of the virtual space V that can be simulated in one frame period. The IR generation unit 133 obtains, from the storage unit 120, the size that can be simulated in one frame period as the maximum size. The IR generation unit 133 sets the subspace SB as a space having a size equal to or less than the maximum size.
[0197] For example, the subspace SB is set as a space including at least one of the sound source S or the sound receiving point T. As an example, the subspace SB is set as a space including at least one of the sound source S or the sound receiving point T on the boundary with the external space (the virtual space V outside the subspace SB). The sound source S can be an actual sound source (real sound source RS) serving as a sound generation source, or can be a virtual sound source VS (see Figure 23 ) provided on the diffraction path DR between the real sound source RS and the diffraction part DF. The sound receiving point T can be an actual sound receiving point (real sound receiving point RL) serving as a sound listening position (e.g., the position of the user U's ear), or can be a virtual sound receiving point VL (see Figure 23 ) provided on the diffraction path DR between the actual sound receiving point RL and the diffraction part DF. In Figure 16 's embodiment, the subspace SB is set as a small space including the real sound source RS, the diffraction part DF, and the real sound receiving point RL.
[0198] In the case where the real sound source RS, the diffraction part DF, and the real sound receiving point RL are arranged at positions close enough to enable real-time simulation, needless to say, the subspace SB can be configured as a space internally including the real sound source RS, the diffraction part DF, and the real sound receiving point RL (a configuration where neither the real sound source RS nor the real sound receiving point RL is on the boundary of the subspace SB).
[0199] [2-2. Configuration Embodiments of Subspace]
[0200] Figure 17 is a diagram showing configuration embodiments of the subspace SB.
[0201] The identification unit 132 specifies the diffraction part DF based on the path information obtained from the path search unit 131, and determines the subspace SB to be input to the IR generation unit 133. The subspace SB can be a predefined space, or can be set by a creator such as a sound designer who actually produces sounds such as games.
[0202] In Figure 17In the embodiment, the subspace SB is set to include part or all of the object OB containing the diffraction part DF, the sound source S, and the sound reception point T. For example, in the case of considering a two-dimensional space, the subspace SB can be set to a square, rectangle, circle, or semi-circle. In the case of considering a three-dimensional space, the subspace SB can be set to a cube, cuboid, sphere, or hemisphere. The above shapes are examples, and the subspace SB can be set to various shapes other than the above examples.
[0203] Figure 18 is a diagram showing an embodiment of the processing flow. The difference from the first embodiment lies in the setting process of the subspace SB (step S212), the setting process of the simulation time (step S213), and the generation process of the IR (step S214).
[0204] The identification unit 132 sets the subspace SB including the sound source S, the diffraction part DF, and the sound reception point T (step S212). The subspace SB can be predefined according to the diffraction part DF or can be arbitrarily defined by the creator. The identification unit 132 determines the simulation time based on the distance of the diffraction path DR (step S213). For example, using the distance x of the diffraction path DR and the speed of sound c, the simulation time t is set to t = 2×x / c.
[0205] The IR generation unit 133 receives information on the subspace SB for performing the simulation, information on the time t for simulating the diffraction phenomenon, sound source information, sound reception point information, and object information. The IR generation unit 133 performs a wave acoustics simulation based on the received information and generates an IR (step S214). As the simulation method, methods such as FDTD, FEM, and BEM are used.
[0206] Note that an AI model that can use the input object information, sound source information, and sound reception point information and output an IR can be used to perform the IR generation process. The AI model can be the above PINN model. The IR generation process can be executed on the cloud.
[0207] [2-3. Relationship between the real sound source, the real sound reception point, and the subspace]
[0208] In Figure 16 the embodiment, the subspace SB is set to a small space including the real sound source RS and the real sound reception point RL. However, the subspace SB can be set to a small space that does not include one or both of the real sound source RS and the real sound reception point RL.
[0209] Figure 19 is a simple diagram for explaining a processing example in the case where the real sound source RS is not included in the subspace SB. Figure 20 is a block diagram showing the flow of the generation process of the diffracted sound. Figure 20The difference from the first embodiment is that the IR generation unit 133 performs delay / gain adjustment of the IR.
[0210] [2-4. Delay / Gain Adjustment]
[0211] In the case where the subspace SB does not include the true sound source RS, the identification unit 132 sets a point on the boundary of the subspace SB that intersects the diffraction path DR as the virtual sound source VS. The identification unit 132 calculates the distance between the true sound source RS and the virtual sound source VS as the attenuation distance x b . The identification unit 132 uses the speed of sound c according to the attenuation distance x b to calculate the delay time τ (τ = x b / c). The identification unit 132 calculates the amplitude attenuation rate according to the attenuation distance x b . For the calculation of the amplitude attenuation rate, an inverse square law such as "obtaining a -6dB attenuation when the distance is doubled" or an attenuation curve uniquely set by the creator can be used.
[0212] The IR generation unit 133 sets the virtual sound source VS and the true sound reception point RL as the sound source S and the sound reception point T of the wave acoustics simulation. The IR generation unit 133 performs a wave acoustics simulation on the subspace SB and generates the IR between the sound source S and the sound reception point T. The IR generation process has a time limit similar to that of the first embodiment. The simulation time is limited according to the distance of the diffraction path DR from the sound source S (virtual sound source VS) to the sound reception point T (true sound reception point RL). For example, using the distance (x - x b ) obtained by subtracting the attenuation distance x from the distance x of the diffraction path DR b ) and the speed of sound c, the simulation time t is calculated as t = 2×(x - x b ) / c. The IR generation unit 133 adds the amplitude attenuation rate and the delay time τ according to the attenuation distance x b to the IR.
[0213] Note that it is assumed that the IR obtained by the wave acoustics simulation is normalized so that the absolute value of the maximum amplitude is 1 or the maximum value of the data type in which the sound pressure is held. The normalization coefficient used here can be saved as data in the memory. Then, the IR obtained by adding the delay time τ to the normalized IR and further multiplying the IR by the amplitude attenuation rate is the IR initially desired to be obtained.
[0214] When the real sound reception point RL is set outside the subspace SB, the above idea can be similarly applied. For example, in the case where the subspace SB does not include the real sound reception point RL, the recognition unit 132 sets a point on the boundary of the subspace SB that intersects the diffraction path DR as the virtual sound reception point VL. The IR generation unit 133 adds the amplitude attenuation rate and the delay time τ to the IR according to the attenuation distance x between the real sound reception point RL and the virtual sound reception point VL b Add the amplitude attenuation rate and the delay time τ to the IR.
[0215] In the above method, the subspace SB for which simulation is to be performed is further reduced. Therefore, the computational amount is reduced, and real-time speech processing becomes easier to perform.
[0216] [2-5. Embodiment of the processing flow]
[0217] Figure 21 It is a diagram showing an embodiment of the processing flow. The difference from the first embodiment lies in the presetting process (steps S221 to S222) and the IR generation / delay addition / gain adjustment process (steps S231 to S232).
[0218] In this flow, the setting of the subspace SB and the amplitude attenuation rate is performed as a preset. The storage unit 120 stores the setting information about the subspace SB and the amplitude attenuation rate. The recognition unit 132 performs the setting of the shape and size of the subspace SB based on the setting information acquired from the storage unit 120 (step S221). The IR generation unit 133 performs the setting related to the method of calculating the amplitude attenuation rate based on the setting information acquired from the storage unit 120 (step S222).
[0219] After that, the diffraction part DF is detected in the same method as the first embodiment (steps S201 to S204). The recognition unit 132 extracts a small space having a shape and size according to the setting information from the entire space TS with the diffraction part DF as the center. The recognition unit 132 sets the extracted small space as the subspace SB that is the target of the wave acoustics simulation. The recognition unit 132 determines whether the real sound source RS and the real sound reception point RL exist in the subspace SB (step S223).
[0220] In the case where the real sound source RS and the real sound reception point RL exist in the subspace SB (step S223: Yes), the recognition unit 132 sets the real sound source RS and the real sound reception point RL as the sound source S and the sound reception point T in the wave acoustics simulation.
[0221] The identification unit 132 determines the simulation time t (step S224) using the distance x of the diffraction path DR and the speed of sound c (t = 2×x / c). The IR generation unit 133 performs a wave acoustics simulation such as FTDT, FEM, or BEM on the subspace SB to generate an IR (step S225). The convolution unit 134 convolves the time waveform data of the real sound source RS and the IR to generate a diffracted sound (step S206). The audio data reproduction unit 135 presents the synthesized output sound including the generated diffracted sound to the user U via the headphones 150 and the speaker 160 (step S207).
[0222] In the case where the real sound source RS and the real sound reception point RL do not exist in the subspace SB (step S223: No), the identification unit 132 sets a virtual sound source (virtual sound source VS) or a virtual sound reception point (virtual sound reception point VL) on the boundary of the subspace SB. For example, in the case where the real sound source RS does not exist in the subspace SB, the identification unit 132 sets the virtual sound source VS on the diffraction path DR between the real sound reception point RL and the diffraction part DF. The identification unit 132 sets the virtual sound source VS as the sound source S in the wave acoustics simulation. In the case where the real sound reception point RL does not exist in the subspace SB, the identification unit 132 sets the virtual sound reception point VL on the diffraction path DR between the real sound reception point RL and the diffraction part DF. The identification unit 132 sets the virtual sound reception point VL as the sound reception point T in the wave acoustics simulation.
[0223] The identification unit 132 calculates the distance from the real sound source RS or the real sound reception point RL that does not exist in the subspace SB to the boundary of the subspace SB as the attenuation distance x b (step S228). For example, in the case where the real sound source RS does not exist in the subspace SB, the identification unit 132 calculates the distance from the real sound source RS to the boundary of the subspace SB (virtual sound source VS) as the attenuation distance x b s . In the case where the real sound reception point RL does not exist in the subspace SB, the identification unit 132 calculates the distance from the real sound reception point RL to the boundary of the subspace SB (virtual sound reception point VL) as the attenuation distance x b m .
[0224] The identification unit 132 calculates the delay time τ based on the attenuation distance x b (step S229). For example, in the case where the real sound source RS does not exist in the subspace SB, the identification unit 132 calculates the value obtained by dividing the attenuation distance x b s by the speed of sound c as the delay time τ s。When the true sound reception point RL does not exist in the subspace SB, the recognition unit 132 calculates the value obtained by dividing the attenuation distance x b m by the speed of sound c as the delay time τ m 。When both the true sound source RS and the true sound reception point RL do not exist in the subspace SB, the delay time is τ s +τ m ,and the distance considering the amplitude attenuation is x b s +x b m 。
[0225] The IR generation unit 133 performs a wave acoustics simulation such as FTDT, FEM, or BEM on the subspace SB to generate an IR (step S230). The IR generation unit 133 adds the delay time τ to the generated IR (step S231). In addition, the IR generation unit 133 adds the amplitude attenuation rate (gain) corresponding to the attenuation distance x b to the generated IR (step S232). The convolution unit 134 convolves the time waveform data of the true sound source RS and the IR to which the delay time τ and the amplitude attenuation rate are applied, and generates a diffracted sound (step S206). The audio data reproduction unit 135 presents the synthesized output sound including the generated diffracted sound to the user U via the headphones 150 and the speaker 160 (step S207).
[0226] Note that the order of addition of the delay time τ and the addition of the amplitude attenuation rate (gain adjustment) can be changed. In the above process, the gain adjustment is performed using a preset amplitude attenuation rate. However, when the creator wants to amplify the sound from a certain sound source, the gain parameter can also be adjusted and the adjustment result can be reflected in the sound.
[0227] [2-6. Circular / Spherical Subspace]
[0228] Figure 22 is a schematic diagram showing an embodiment in which the subspace SB is configured as a spherical or circular space centered on the diffraction site DF. In Figure 22 ,since the subspace SB is shown as a two-dimensional plane, the subspace SB is a circle represented by the polar coordinate system (r, θ), but in the case of a three-dimensional space, the subspace SB is a sphere represented by the spherical coordinate system (r, θ, φ).
[0229] The recognition unit 132 designates the edge of the object OB as the diffraction site DF and sets a circle with a radius r from the center O of the edge. The IR generation unit 133 simulates the IR from the sound source S on the circumference to the sound reception point T by FDTD or the like, and can obtain the IR by adjusting the delay and gain using the above method.
[0230] As an advantage of the circular or spherical subspace SB, since the distance (=r) from the sound source S or the sound reception point T on the subspace SB to the center O of the edge is constant, the delay time τ can be easily obtained. In addition, by making the subspace SB circular, the incident angle θ from the sound source S to the edge can also be calculated s and the angle θ from the edge to the sound reception point T m of the angle information. When considering a three-dimensional space (θ s , φ s ), there is a pair of horizontal angle θ and elevation angle φ of (θ m , φ m ). The angle information obtained here can also be used as input information for performing numerical simulations such as FDTD, can also be used as input data in the case of adopting an AI model, and can also be used as table information when creating a table for storing IR.
[0231] Figure 23 is a schematic diagram showing an embodiment in which the real sound source RS and the real sound reception point RL are outside the subspace SB.
[0232] Calculate the distance r from the real sound source RS to the center O of the edge based on the sound source information and the position information of the diffraction part DF sO . Calculate the distance r from the real sound reception point RL to the center O of the edge based on the sound reception point information and the position information of the diffraction part DF mO . The radius of the subspace SB is set to r. Therefore, the distance from the real sound source RS to the virtual sound source VS is r sO -r. The distance from the real sound reception point RL to the virtual sound reception point VL is r mO -r. By dividing these distances by the speed of sound c, the delay time τ can be easily obtained. Then, if the calculated delay time τ is added to the IR, the correct IR delay time can be obtained.
[0233] In Figure 23 the embodiment, the case where the real sound source RS and the real sound reception point RL exist outside the subspace SB has been considered. However, a similar concept can be applied to the case where the real sound source RS and the real sound reception point RL exist inside the subspace SB. In this case, since the radius r is large, the delay time to be adjusted is the value obtained by adding (r - r sO ) / c and (r - r mO ) / c. Then, by cutting the head signal of the IR (such as FDTD) obtained by simulation by the calculated delay time, an IR with the correct delay time can be obtained.
[0234] [2-7. Embodiment of setting the size of the subspace]
[0235] The size (radius r) of the subspace SB can be set as follows. For example, the frame rate often used in games is 60 fps. The game console performs some processing every 16 ms. Therefore, it can be considered to determine the radius r based on the 16-ms frame period.
[0236] Suppose the virtual sound source VS and the virtual sound reception point VL are arranged on the circumference of a circle with radius r. In this case, the diffraction path DR from the virtual sound source VS to the virtual sound reception point VL is 2×r. Using the speed of sound c, the arrival time t0 is t0 = 2×r / c. As described above, it is considered that the simulation time t needs to be about t = 2×t0 seconds. When using the value of 16 ms, 4×r / c = 16 ms is obtained, and when c = 340 m / s, r = 1.36 m is obtained. In this way, even if a large virtual space V is provided, it is only necessary to simulate the space of a circle or a sphere with a radius of 1.36 m around each diffraction part DF calculated by the recognition unit 132.
[0237] The recognition unit 132 can set the size of the subspace SB based on the size of the object OB where the diffraction phenomenon occurs. For example, the recognition unit 132 can set the subspace SB as a space including a preset ratio of the surface area or volume of the object OB. For example, if it can cover more than 50% of the surface area or volume of each object OB, the outline of the object OB can be known. Therefore, it is also conceivable to use a value of more than 50% as an index and use a circle or a sphere with a radius depending on the size of each object OB to set the subspace SB. In this case, since the subspace SB corresponding to the large object OB becomes larger, there is a disadvantage of an increase in the calculation load.
[0238] A complex object OB having a plurality of diffraction parts DF can be processed as follows. Figure 24 and Figure 25 are diagrams for explaining an example of a method for setting the subspace SB for the complex object OB.
[0239] In Figure 24 's embodiment, a plurality of diffraction parts DF exist in the same object OB. The recognition unit 132 sets the subspace SB to include all the plurality of diffraction parts DF. However, in the case of a large object OB, the subspace SB also increases, and the calculation load increases. Therefore, it is also conceivable to perform the wave acoustics simulation only in the subspace SB where the radius r is preset to a fixed value of a specific value. Figure 25 The state at this time is shown.
[0240] The recognition unit 132 extracts one or more diffraction parts DF that can be included in the subspace SB from the plurality of diffraction parts DF included in the same object OB. The IR generation unit 133 selectively performs the wave acoustics simulation on the extracted one or more diffraction parts DF. For example, in setting asFigure 25 In the case of the subspace SB shown, only the diffraction phenomenon occurring at the central diffraction site DF among the multiple diffraction sites DF is simulated. Alternatively, only the diffraction phenomenon occurring at the diffraction site DF corresponding to the peak among the multiple diffraction sites DF is simulated. Figure 25 The peak in [] refers to the diffraction site DF at the center of the uppermost side when three diffraction sites DF are connected by a straight line. The diffraction phenomenon of another diffraction site DF present in the original diffraction path DR is not simulated. However, this way of thinking is also possible when there is no discomfort with respect to the audibility of the generated diffraction sound.
[0241] [3. Third Embodiment: Generation of Higher-Order Diffraction Sound Based on a Combination of Multiple Subspaces]
[0242] Figure 26 is a diagram for explaining the features of the audio signal generation process of the third embodiment.
[0243] In the first and second embodiments, an example in which only one object OB exists in the diffraction path DR from the sound source S to the sound reception point T, that is, an example in which diffraction occurs only once, has been described. In the third embodiment, a case in which multiple diffraction sites DF exist in the diffraction path DR from the sound source S to the sound reception point T will be described. The multiple diffraction sites DF are detected by the identification unit 132.
[0244] Hereinafter, in the case of distinguishing multiple diffraction sites DF, an identification number is added after the reference numeral. The same applies to the case of distinguishing between subspaces SB, between virtual sound sources VS, and between virtual sound reception points VL. Diffraction sites DF, virtual sound sources VS, and virtual sound reception points VL belonging to the same subspace SB are assigned the same identification number as the identification number of the subspace SB.
[0245] The identification unit 132 detects multiple diffraction sites DF that cause the sound output from the sound source S to diffract toward the sound reception point T. The identification unit 132 classifies the detected diffraction sites DF into multiple groups based on the above maximum size. The identification unit 132 sets a subspace SB for each classified group. The IR generation unit 133 adds an amplitude attenuation rate and a delay time corresponding to the distance between the subspaces SB along the diffraction path DR to the IR.
[0246] Figure 26 An example in which two diffraction sites DF are present is shown. The identification unit 132 sets a subspace SB for each diffraction site DF. In the Figure 26 embodiment, the subspace SB is ensured to be rectangular, but the subspace SB can be a closed space other than a rectangle, such as a circle.
[0247] When ensuring multiple subspaces SB by the recognition unit 132, the IR generation unit 133 assumes that the virtual sound source VS-1 and the virtual sound reception point VL-1 are located on the boundary of the subspace SB-1, inputs a Gaussian pulse from the virtual sound source VS-1, and obtains the IR representing the diffraction phenomenon by a method such as FDTD. Here, other signals such as time-stretched pulse (TSP) can be used as long as the IR other than the Gaussian pulse can be simulated.
[0248] If the method for obtaining the IR is a trained AI model, the model makes the input be the angle information from the virtual sound source VS to the edge (diffraction field DF), object information, and the angle information from the edge to the virtual sound reception point VL, and the output be the IR. Alternatively, an input similar to the learned AI model can be given, and the IR can be obtained from a table maintained by pre-simulating the IR.
[0249] The IR generation unit 133 inputs the IR obtained in the subspace SB-1 to the virtual sound source VS-2, and obtains the IR up to the virtual sound reception point VL-2 using FDTD, AI model or table. The IR generation unit 133 obtains the distances on the diffraction path DR from the real sound source RS to the subspace SB-1, the distances on the diffraction path DR from the subspace SB-1 to the subspace SB-2, and the distances on the diffraction path DR from the subspace SB-2 to the real sound reception point RL as attenuation distances, and adds the delay time τ and the amplitude attenuation rate corresponding to the attenuation distances to the IR. As a result, the IR representing the diffracted sound in the case where there are two diffraction sites DF can be obtained by low computation. In addition, similarly, for the virtual sound source VS-2, a Gaussian pulse is input, the IR up to the virtual sound reception point VL-2 is obtained using FDTD, AI model or table, and the method of performing convolution operation with the IR obtained in the subspace SB-1 is also possible. In this case, the IR calculations for each of the subspaces SB-1 and SB-2 can be processed in parallel. The method of adding the delay time τ and the amplitude attenuation rate corresponding to the attenuation distance is similar. Even when the number of diffraction sites DF is three or more, these concepts are similar.
[0250] On the other hand, as the number of diffraction sites DF increases, the amplitude of the IR representing the diffracted sound decreases. Therefore, a threshold value for the amplitude of the IR can also be set, and the diffracted sound is not calculated when the amplitude is equal to or less than the threshold value. Since the magnitude of the amplitude depends on the distance (path length) of the diffraction path DR from the true sound source RS to the true sound reception point RL, the necessity for calculation can be determined by setting the threshold value to the path length. For example, the signal processing unit 139 can limit the generation of diffracted sound through a diffraction path DR having a length exceeding a preset allowable length. The limitation includes, for example, partial or complete prohibition. The allowable length can be set arbitrarily. When setting a threshold value for the path length, it is determined whether calculation is necessary without calculating the IR. Therefore, the calculation burden is reduced.
[0251] Alternatively, a threshold value for the number of diffraction sites DF can also be set to determine whether calculation is necessary. For example, the signal processing unit 139 can limit the generation of diffracted sound through a number of subspaces SB exceeding a preset allowable number. The allowable number can be set arbitrarily. For example, in the case where there are three or more diffraction sites DF, a method such as not performing calculation can be considered. Moreover, in this method, since the necessity for calculation can be easily determined, the calculation burden is reduced.
[0252] [4. Modification Example 1: Limitation on Audio Signal Generation Processing Based on the Directivity of the Sound Source]
[0253] Although the directivity of the sound source S has not been mentioned so far, in the direction where the sound source S has directivity and the sound is not strongly emitted, the generation process of the diffracted sound is unnecessary. Therefore, the signal processing unit 139 can obtain an allowable propagation direction allowing sound propagation based on the directivity of the sound source S, and can limit the generation of diffracted sound propagating in a direction other than the allowable propagation direction. Information about the directivity of the sound source S is included in the sound source information.
[0254] Figure 27 is a diagram showing an example of the directivity of the sound source S. In Figure 27 the example, a directivity characteristic is shown in which sound is strongly emitted in the 0-degree direction and no sound is emitted in the 180-degree direction. As described above, since the sound is not strongly emitted in the 180-degree direction with respect to the sound source S, the diffracted sound generated in this direction is small. Therefore, processing for ignoring the propagation of sound in the 180-degree direction can be performed.
[0255] [5. Modification Example 2: Limitation on Audio Signal Generation Processing Based on the Size of the Sound Generation Site]
[0256] The necessity for the generation process of the diffracted sound can be determined based on the size of the sound generation site SE. The sound generation site SE refers to a position on an object that outputs sound. Figure 28It is a diagram for explaining the relationship between the size of the sound generation part SE and the diffracted sound.
[0257] When the size of the sound generation part SE is larger than the size of the object OB where the diffraction phenomenon occurs, it is considered that the direct sound directly reaching the sound reception point T dominates over the diffracted sound. Therefore, in this case, the process of calculating the diffracted sound can be omitted. For example, the signal processing unit 139 calculates the ratio of the direct sound reaching the sound reception point T from the sound generation part SE via the diffraction part DF. The signal processing unit 139 can limit the ratio of the direct sound to be greater than the generation of the diffracted sound of the sound generation part SE that allows the standard. The allowable standard can be set arbitrarily.
[0258] [6. Modification Example 3: Mesh Size of the Subspace]
[0259] The IR generation unit 133 can change the mesh size of the subspace SB used in the wave acoustics simulation according to the size of the object OB where the diffraction phenomenon occurs. When the size of the object is large, only low-frequency sounds are diffracted. Therefore, for example, when the obstacle is 5 m, if it is assumed that only diffraction up to 340 Hz (the frequency with a wavelength of 1 m), which is 1 / 5 of the size frequency, occurs, the mesh size can be set to 10 cm or the like. Generally, in FDTD, it is considered preferable to set the mesh size to the mesh size of the wavelength from 1 / 20 to 1 / 10 of the upper limit frequency of interest. When the size of the obstacle becomes small, diffraction occurs up to a higher frequency, so it is preferable to set the mesh size more finely.
[0260] [7. Modification Example 4: Embodiment of Setting the Subspace According to the Size of the Gap between Objects]
[0261] Figure 29 and Figure 30 It is a simple diagram for explaining the method of setting the subspace SB when there is a gap between two objects OB.
[0262] When there is a gap between two objects OB and it is desired to obtain the IR of the diffracted sound generated in the gap, the recognition unit 132 sets the subspace SB to include all the edges (diffraction parts DF-1 and DF-2) of each object OB. This makes it possible to accurately obtain the IR of the diffracted sound due to the gap.
[0263] On the other hand, when the distance between the edges is long enough, as Figure 30As shown, the identification unit 132 sets the subspace SB centered on the edge (diffraction part DF-1) that is the shortest path between the sound source S and the sound reception point T. The identification unit 132 calculates the IR between the virtual sound source VS-1 and the virtual sound reception point VL-1, and performs delay / gain adjustment on the calculated IR to obtain the IR from the real sound source RS to the real sound reception point RL. It is possible to judge whether the edge is sufficiently separated based on a threshold value. As the threshold value, for example, a value such as 3 m is used.
[0264] [8. Modification Example 5: Two-Dimensional Subspace]
[0265] The subspace SB in the three-dimensional virtual space V can be set as a three-dimensional space such as a sphere. However, in order to reduce the memory occupancy and calculation amount of the subspace SB, it is also conceivable to convert the three-dimensional subspace SB into a two-dimensional subspace SB and approximately perform the calculation of the IR in the three-dimensional space in the two-dimensional space. Specific embodiments will be described below.
[0266] [8-1. Case Where the Sound Source and the Sound Reception Point Are on the Same Coordinate Plane]
[0267] Figure 31 is a diagram for explaining an operation embodiment in the case where the sound source S and the sound reception point T are on the same coordinate plane.
[0268] The three-dimensional coordinates of the sound source S and the sound reception point T and the configuration of the width, thickness, etc. of the object OB are as follows. For example, the object OB is a wall that can diffract the sound from the sound source S at the upper and side edges.
[0269] · Sound source: (2.0 m, 1.5 m, 1.2 m)
[0270] · Sound reception point: (3.0 m, 1.5 m, 1.2 m)
[0271] · Width of the object in the x direction: 0.3 m
[0272] · Width of the object in the y direction: 2.0 m
[0273] · Width of the object in the z direction: 1.5 m
[0274] The sound source S and the sound reception point T exist on the same xy plane (z = 1.2 m). Therefore, the diffraction path DR of the diffracted sound that first reaches the sound reception point T through the side edge of the object OB also exists on the same xy plane as the sound source S and the sound reception point T. Figure 31In the example, the coordinates of the diffraction part DF-1 (the edge on the object OB side) are (2.65 m, 2.0 m). The recognition unit 132 sets a circular subspace SB-1 centered on the diffraction part DF-1 as the center O1 on the xy plane where z = 1.2 m. The IR generation unit 133 performs a wave acoustics simulation on the subspace SB-1 and generates an IR (IR_xy) in the xy two-dimensional space.
[0275] Similarly, the sound source S and the sound reception point T exist on the same zx plane (y = 1.5 m). Therefore, the diffraction path DR of the diffracted sound that first reaches the sound reception point T through the upper edge of the object OB also exists on the same zx plane as the sound source S and the sound reception point T. In Figure 31 In the example, the coordinates of the diffraction part DF-2 (the upper edge of the object OB) are (2.65 m, 1.5 m). The recognition unit 132 sets a circular subspace SB-2 centered on the diffraction part DF-2 as the center O2 on the zx plane where y = 1.5 m. The IR generation unit 133 performs a wave acoustics simulation on the subspace SB-2 and generates an IR (IR_zx) in the zx two-dimensional space.
[0276] Between the propagation in three-dimensional space and the propagation in two-dimensional space, the way of sound attenuation is different. The IR generation unit 133 adjusts the gain of the IR obtained in the two-dimensional space based on the difference in distance attenuation between the two-dimensional and three-dimensional spaces. For example, the distance attenuation in three-dimensional space is theoretically 1 / 2 times the amplitude when the distance is doubled. When the distance is doubled, the distance attenuation in two-dimensional space theoretically becomes 1 / √2 times the amplitude. In Figure 31 In the embodiment of, since the simulation is performed in two-dimensional space, correction by distance attenuation is necessary.
[0277] Figure 32 is a diagram showing an example of distance attenuation in two-dimensional space and three-dimensional space.
[0278] The distance attenuation curve is theoretically a curve as Figure 32 shown. Even in the case of propagation at the same distance, the attenuation is greater in the case of three-dimensional extended propagation than in the case of two-dimensional extended propagation. The magnitude of the attenuation can be calculated based on physical laws, but there may be cases where the sound designer wants to emphasize the sound of the game production venue. In this case, the distance attenuation curve can be set separately by the sound designer.
[0279] In Figure 31 In the embodiment of, the distance from the sound source S via the edge on the object OB side (diffraction part DF-1) to the sound reception point T is 1.43 m. The distance is obtained based on the path information output from the path search unit 131. Refer to Figure 32, at a distance of 1.43 m, the distance attenuation in three-dimensional space is -28.5 dB (0.0376 on a linear scale). The distance from the sound source S via the upper edge of the object OB (diffraction part DF-2) to the sound reception point T is 1.18 m. Refer to Figure 32 , at a distance of 1.18 m, the distance attenuation in three-dimensional space is -26.8 dB (linear scale of 0.0457).
[0280] Figure 33 is a diagram for explaining the gain adjustment of each diffraction path DR.
[0281] The IR generation unit 133 performs gain adjustment on IR_xy obtained in the subspace SB-1 on the xy plane and IR_zx obtained in the subspace SB-2 on the zx plane based on Figure 31 the distance attenuation information therein. Figure 33 The left side of shows the status of the gain adjustment of IR_xy. Figure 33 The middle part of shows the status of the gain adjustment of IR_zx. Figure 33 The right side of shows the status of the integration process of IR_xy and IR_zx after the gain adjustment.
[0282] The IR generation unit 133 can calculate the IR of the diffracted sound on the side and the diffracted sound on the top of the observation object OB by adding IR_xy and IR_zx after the gain adjustment on the time axis. The IR generation unit 133 can calculate the IR with the maximum value normalized to 1 as needed (see Figure 33 the upper right part of ). In the Figure 33 embodiment, by dividing the data value by 0.0457, the maximum value becomes 1. This process is equivalent to +26.8 dB. The sound designer can also perform the final gain adjustment by considering the sound balance. The normalization parameter when the maximum value is 1 and the value of +26.8 dB (linear scale of 0.0457) may be of reference value for the sound designer to adjust the gain subsequently, so they can be stored separately in the storage unit 120.
[0283] Figure 34 is a diagram showing an example where the diffraction path DR exists in three directions.
[0284] In Figure 34In the embodiment, there are diffraction paths DR passing through the left edge (diffraction part DF-1) of the object OB, diffraction paths DR passing through the right edge (diffraction part DF-2) of the object OB, and diffraction paths DR passing through the upper edge (diffraction part DF-3) of the object OB. The recognition unit 132 sets sub-spaces SB-1, SB-2, and SB-3 corresponding to the diffraction part DF-1, the diffraction part DF-2, and the diffraction part DF-3. The IR generation unit 133 performs a wave acoustics simulation for each sub-space SB. The IR generation unit 133 performs gain adjustment on the IRs obtained for each sub-space SB, and integrates the three IRs after the gain adjustment on the time axis.
[0285] In Figure 31 and 34 In the embodiment shown, the diffracted sound generated in the three-dimensional space is processed in the two-dimensional plane. The recognition unit 132 sets a two-dimensional space including the sound source S and the sound reception point T on the same coordinate plane as the sub-space SB. The IR generation unit 133 calculates the IR in the sub-space SB configured as the two-dimensional space. Therefore, the diffracted sound in the three-dimensional space can be approximately calculated while reducing the amount of calculation.
[0286] [8-2. Case where the sound source and the sound reception point are not on the same coordinate plane]
[0287] Figure 35 is a diagram for explaining an operation example in the case where the sound source S and the sound reception point T are not on the same coordinate plane.
[0288] The three-dimensional coordinates of the sound source S and the sound reception point T and the configuration of the width, thickness, etc. of the object OB are as follows. For example, the object OB is a wall that can diffract the sound from the sound source S at the upper edge and the side edge.
[0289] · Sound source: (2.0 m, 1.5 m, 1.2 m)
[0290] · Sound reception point: (3.0 m, 1.0 m, 1.0 m)
[0291] · Width of the object in the x direction: 0.3 m
[0292] · Width of the object in the y direction: 2.0 m
[0293] · Width of the object in the z direction: 1.5 m
[0294] Since the sound source S and the sound reception point T have different z coordinates, they do not exist in the same xy plane. Therefore, assuming that the sound source S and the sound reception point T are in the same xy plane, the IR generation unit 133 first passes through the same as in Figure 31A method similar to the method in generates the IR. Thereafter, the IR generation unit 133 corrects the IR based on the difference between the z coordinate of the sound source S and the sound reception point T. As a result, the IR of the diffracted sound passing through the side edge (diffraction part DF-1) of the object OB is generated. Similarly, since the y coordinates are different, the sound source S and the sound reception point T do not exist in the same zx plane. Therefore, the IR generation unit 133 uses a similar method to generate the IR of the diffracted sound passing through the upper edge (diffraction part DF-2: see Figure 38 ) of the object OB. Specific description will be given below with reference to Figures 36 to 40 .
[0295] Figure 36 And Figure 37 are diagrams for explaining the generation process of the diffracted sound passing through the side edge (diffraction part DF-1) of the object OB.
[0296] The identification unit 132 projects the sound reception point T that is not in the same coordinate plane as the sound source S onto the same coordinate plane as the sound source S. In the example of Figure 36 , the coordinate plane to be projected is the xy plane where z = 1.2 m. In Figure 36 , the reference numeral PJ1 represents the projection point on the xy plane obtained by projecting the sound reception point T onto the same xy plane as the sound source S. The identification unit 132 sets the two-dimensional space on the xy plane including the sound source S and the projection point PJ1 as the subspace SB-1.
[0297] The IR generation unit 133 performs a wave acoustics simulation on the set two-dimensional subspace SB-1 and calculates the IR_xy in the subspace SB-1 (see the lower right of Figure 36 ). The identification unit 132 calculates the difference between the distance on the diffraction path DR between the sound source S and the sound reception point T and the distance on the diffraction path between the sound source S and the projection point PJ1. The identification unit 132 calculates the delay time τ1 by dividing this difference by the speed of sound. The identification unit 132 obtains the amplitude attenuation rate corresponding to the difference value using the distance attenuation curve in Figure 32 . The IR generation unit 133 adds the amplitude attenuation rate corresponding to the difference value and the delay time τ1 to the IR_xy. Thus, the IR_xy1 with the delay time τ1 and gain adjusted is obtained (see the right side of Figure 37 ). In this embodiment, according to Figure 32 , the value of the amplitude attenuation rate corresponding to the distance in the three-dimensional space is 0.0279, so the maximum amplitude is adjusted to 0.0279.
[0298] For example, in Figure 36In the embodiment, the coordinates of the projection point PJ1 are (3.0 m, 1.0 m, 1.2 m). The coordinates of the sound reception point T are (3.0 m, 1.0 m, 1.0 m). The coordinates of the diffraction part DF-1 on the xy plane (z = 1.2 m) are (2.65 m, 2.0 m, 1.2 m). The distance on the diffraction path between the sound source S and the projection point PJ1 is 1.8795 m. The distance on the diffraction path DR between the sound source S and the sound reception point T is 1.8983 m. Therefore, the difference is 0.0187 m. Assuming the speed of sound is 340 m / s, the delay time τ1 is 0.055 milliseconds. The time waveform given for the delay time τ1 is the time waveform indicated by the dotted line on the Figure 37 right side.
[0299] Figure 38 and Figure 39 are diagrams for explaining the generation process of the diffracted sound passing through the upper edge (diffraction part DF-2) of the object OB.
[0300] The identification unit 132 projects the sound reception point T not on the same coordinate plane as the sound source S onto the same coordinate plane as the sound source S. In Figure 38 the embodiment, the coordinate plane to be projected is the zx plane of y = 1.5 m. In Figure 38 , the reference numeral PJ2 represents the projection point on the zx plane obtained by projecting the sound reception point T onto the same zx plane as the sound source S. The identification unit 132 sets the two-dimensional space on the zx plane including the sound source S and the projection point PJ2 as the subspace SB-2.
[0301] The IR generation unit 133 performs a wave acoustics simulation on the set two-dimensional subspace SB-2 and calculates the IR_zx in the subspace SB-2 (see the lower right of Figure 38 ). The identification unit 132 calculates the difference between the distance on the diffraction path DR between the sound source S and the sound reception point T and the distance on the diffraction path between the sound source S and the projection point PJ2. The identification unit 132 calculates the delay time τ2 by dividing this difference by the speed of sound. The identification unit 132 obtains the amplitude attenuation rate corresponding to the difference value using the distance attenuation curve in Figure 32 . The IR generation unit 133 adds the amplitude attenuation rate corresponding to the difference value and the delay time τ2 to the IR_zx. Thus, the IR_zx1 with the delay time τ2 and the adjusted gain is obtained (see the right side of Figure 39 ). In this embodiment, according to Figure 32 , since the value of the amplitude attenuation rate corresponding to the distance in the three-dimensional space is 0.0316, the maximum amplitude is adjusted to 0.0316.
[0302] For example, in Figure 38In the embodiment, the coordinates of the projection point PJ2 are (3.0 m, 1.5 m, 1.0 m). The coordinates of the sound reception point T are (3.0 m, 1.0 m, 1.0 m). The coordinates of the diffraction part DF-2 on the zx plane (y = 1.5 m) are (2.65 m, 1.5 m, 1.5 m). The distance on the diffraction path between the sound source S and the projection point PJ2 is 1.4304 m. The distance on the diffraction path DR between the sound source S and the sound reception point T is 1.6090 m. Therefore, the difference is 0.1787 m. Assuming the speed of sound is 340 m / s, the delay time τ2 is 0.525 milliseconds. The time waveform given for the delay time τ2 is the time waveform shown by the dotted line on the Figure 39 right side.
[0303] Figure 40 is a diagram for explaining the integration process of multiple IRs having different diffraction paths DR.
[0304] The IR generation unit 133 adds the delayed and gain-adjusted IR_xy and IR_zx on the time axis. As a result, the IR of the composite diffraction sound obtained by combining the diffraction sound from the side of the object OB and the diffraction sound from the upper part of the object OB can be calculated.
[0305] Figure 41 is a diagram for explaining another calculation example in the case where the sound source S and the sound reception point T are not on the same coordinate plane.
[0306] The three-dimensional coordinates of the sound source S and the sound reception point T are as follows. Configurations such as the width and thickness of the object OB are similar to those in Figure 31 the configuration.
[0307] · Sound source: (2.0 m, 1.5 m, 1.2 m)
[0308] · Sound reception point: (3.0 m, 1.0 m, 1.0 m)
[0309] In this method, a two-dimensional subspace SB is set in the plane PL including the diffraction path DR connecting the sound source S and the sound reception point T at the shortest distance. In the Figure 41 embodiment, the plane obtained by connecting the sound source S and the sound reception point T with a straight line, observing the straight line on the zx plane, and translating the straight line in the y direction is the plane PL for setting the subspace SB. For example, the identification unit 132 sets the two-dimensional space including the sound source S and the sound reception point T on the same plane PL as the subspace SB. The IR generation unit 133 calculates the IR in the subspace SB configured as a two-dimensional space. As a result, the IR of the diffraction sound passing through the side edge (diffraction part DF) of the object OB is generated.
[0310] A similar method is used in the case of generating an IR of diffracted sound passing through the upper edge of the object OB. In this case, the plane PL of the subspace SB is a plane obtained by connecting the sound source S and the sound receiving point T with a straight line, observing the straight line in the xy plane, and translating the straight line in the z direction.
[0311] [9. Modification Example 6: Generating IR Based on Encoded Object Information]
[0312] Figure 42 is a diagram for explaining an embodiment of generating an IR based on encoded object information.
[0313] The signal processing unit 139 generates an IR using an AI model including an encoding unit 137 and an IR prediction unit 138. The encoding unit 137 compresses data of various object shapes and converts the data into a specific feature quantity. The IR prediction unit 138 outputs an IR based on the feature quantity obtained from the encoding unit 137.
[0314] The encoding unit 137 encodes the object information (diffraction site information) related to the diffraction site DF and outputs a feature quantity. The input is the diffraction site information, and the output is a dimensionally compressed feature quantity. The diffraction site information includes information such as the texture, material, size, density, internal sound velocity, sound absorption coefficient (reflectivity), acoustic impedance, etc. of the diffraction site DF. The IR prediction unit 138 uses the feature quantity output from the encoder, the angle from the sound source S to the diffraction site DF, and the angle from the diffraction site DF to the sound receiving point T as inputs, and outputs an IR representing only the diffracted sound corresponding to the inputs.
[0315] [10. Embodiment of Hardware Configuration]
[0316] For example, an information device such as the acoustic processing device 100 according to the above embodiment is implemented by a computer 1000 having a configuration as shown in Figure 43 . Figure 43 is a hardware configuration diagram showing an embodiment of a computer 1000 that implements the functions of the acoustic processing device 100. The computer 1000 includes a CPU 1100, a RAM 1200, a read-only memory (ROM) 1300, a solid-state drive (SSD) 1400, a communication interface 1500, and an input / output interface 1600. Each unit of the computer 1000 is connected by a bus 1050.
[0317] The CPU 1100 operates based on a program stored in the ROM 1300 or the SSD 1400 and controls each unit. For example, the CPU 1100 expands the program stored in the ROM 1300 or the SSD 1400 into the RAM 1200 and executes processing corresponding to various programs.
[0318] The ROM 1300 stores a boot program, such as a basic input / output system (BIOS) executed by the CPU 1100 when the computer 1000 is activated, a program depending on the hardware of the computer 1000, and the like.
[0319] The SSD 1400 is a computer-readable recording medium that non-transiently records a program executed by the CPU 1100, data used by the program, and the like. Specifically, the SSD 1400 is a recording medium that records an information processing program according to the present disclosure as an example of the program data 1450. The SSD 1400 may be another non-transitory recording medium such as a hard disk drive (HDD).
[0320] The communication interface 1500 is an interface for the computer 1000 to connect to an external network 1550 (e.g., the Internet). For example, the CPU 1100 receives data from other devices or sends data generated by the CPU 1100 to other devices via the communication interface 1500.
[0321] The input / output interface 1600 is an interface for connecting an input / output device 1650 and the computer 1000. For example, the CPU 1100 receives data from an input device such as a touch panel, keyboard, mouse, microphone, or camera via the input / output interface 1600. In addition, the CPU 1100 sends data to an output device such as a display, speaker, or printer via the input / output interface 1600. In addition, the input / output interface 1600 can be used as a medium interface for reading a program or the like recorded in a predetermined recording medium (medium). For example, the medium is an optical recording medium such as a digital versatile disc (DVD) or a phase change rewritable optical disc (PD), a magneto-optical recording medium such as a magneto-optical disc (MO), a tape medium, a magnetic recording medium, a semiconductor memory, or the like.
[0322] For example, when the computer 1000 functions as the acoustic processing device 100 according to the present embodiment, the CPU 1100 of the computer 1000 implements the functions of the control unit 130 and the like by executing an information processing program loaded onto the RAM 1200. In addition, the SSD 1400 stores the information processing program and data according to the present disclosure in the storage unit 120. Note that the CPU 1100 reads the program data 1450 from the SSD 1400 and executes the program data, but as another embodiment, these programs can be obtained from another device via the external network 1550.
[0323] [11. Effects]
[0324] The acoustic processing method of the present disclosure includes a diffraction path search process and a diffracted sound generation process. In the diffraction path search process, a diffraction path DR from a sound source S to a sound reception point T is searched. In the diffracted sound generation process, the diffracted sound obtained at the sound reception point T is generated based on a wave simulation, where the simulation time is constrained according to the distance of the diffraction path DR. The acoustic processing device of the present disclosure executes the information processing defined in the acoustic processing method. The acoustic processing program of the present disclosure implements the information processing defined in the acoustic processing method.
[0325] According to this configuration, an appropriate time constraint according to the distance of the diffraction path DR is applied to the wave acoustic simulation. Therefore, a realistic sound experience with a small amount of calculation is provided.
[0326] In the diffracted sound generation process, based on the distance of the diffraction path DR, the time until the diffracted sound first reaches the sound reception point T is calculated as the arrival time t0. In the diffracted sound generation process, a time equal to or greater than 1 times the arrival time t0 and equal to or less than 5 times the arrival time t0 is set as the simulation time.
[0327] According to this configuration, the simulation can be executed until the time when the diffracted sound can be sufficiently observed, while minimizing the amount of calculation.
[0328] The diffracted sound generation process includes an IR generation process. In the IR generation process, a diffraction part DF on an object OB where a diffraction phenomenon occurs is obtained. In the IR generation process, a wave acoustic simulation is selectively executed for a subspace SB including the diffraction part DF, thereby generating an IR in the subspace SB.
[0329] According to this configuration, an appropriate spatial restriction dedicated to the vicinity of the diffraction part DF is applied to the simulation. Therefore, a realistic sound experience with a small amount of calculation is provided.
[0330] In the IR generation process, an AI model is used to generate an IR. The AI model is a trained neural network that learns the relationship between the simulation conditions of the wave acoustic simulation and the simulation results obtained under the simulation conditions as the relationship between input data and correct answer data.
[0331] According to this configuration, it is not necessary to store a huge amount of pre-prepared data. Therefore, processing can be performed gently without putting pressure on the memory or the like.
[0332] The AI model encodes the object information related to the diffraction part DF and outputs a feature amount. The AI model outputs an IR based on the output feature amount.
[0333] According to this configuration, an appropriate IR that captures the characteristics of the diffraction part DF is produced.
[0334] The subspace SB is a space including at least one of a sound source S and a sound reception point T.
[0335] According to this configuration, the subspace SB to be simulated can be made as small as possible.
[0336] The sound source S is a virtual sound source VS provided on a diffraction path DR between a real sound source RS and a diffraction part DF. In the IR generation process, an amplitude attenuation rate and a delay time corresponding to an attenuation distance between the real sound source RS and the virtual sound source VS are added to the IR.
[0337] According to this configuration, the subspace SB to be simulated can be further limited to the vicinity of the diffraction part DF.
[0338] The sound reception point T is a virtual sound reception point provided on the diffraction path DR between a real sound reception point RL and the diffraction part DF. In the IR generation process, an amplitude attenuation rate and a delay time corresponding to an attenuation distance between the real sound reception point RL and the virtual sound reception point are added to the IR.
[0339] According to this configuration, the subspace SB to be simulated can be further limited to the vicinity of the diffraction part DF.
[0340] The acoustic processing method of the present disclosure includes a process of setting the subspace SB. In the process of setting the subspace SB, a size that can be simulated in one frame period is obtained as a maximum size. In the process of setting the subspace SB, the subspace SB is set as a space having a size equal to or smaller than the maximum size.
[0341] According to this structure, appropriate simulation is performed within a range where real-time processing is possible. Therefore, natural diffracted sound can be presented in real time.
[0342] In the process of setting the subspace SB, the subspace SB is set as a space including a surface area or a volume of a preset ratio of an object OB.
[0343] According to this configuration, since the surface area or the volume of the object OB is covered to a certain extent, the contour of the object OB is well understood. Therefore, appropriate diffracted sound is generated.
[0344] In the process of setting the subspace SB, one or more diffraction parts DF that can be included in the subspace SB are extracted from a plurality of diffraction parts DF included in the same object OB. In the IR generation process, wave acoustics simulation is selectively performed on the extracted one or more diffraction parts DF.
[0345] According to this configuration, the sound quality and the processing speed can be balanced.
[0346] The subspace SB is a spherical or circular space centered on the diffraction part DF.
[0347] According to this configuration, calculation is facilitated.
[0348] The setting process of the subspace SB detects a plurality of diffraction parts DF that diffract the sound output from the sound source S toward the sound reception point T. In the setting process of the subspace SB, the plurality of detected diffraction parts DF are classified into a plurality of groups based on the maximum size. In the setting process of the subspace SB, the subspace SB is set for each of the classified groups. In the IR generation process, an amplitude attenuation rate and a delay time corresponding to the distance between the subspace SB along the diffraction path DR are added to the IR.
[0349] According to this configuration, realistic high-order diffracted sound is generated with a small amount of calculation.
[0350] The generation process of the diffracted sound enables the generation of diffracted sound to be restricted by the number of subspaces SB exceeding a preset allowable number.
[0351] According to this structure, the generation of diffracted sound with a large number of diffractions and an extremely small amplitude is restricted. As a result, the amount of calculation can be suppressed without degrading the sound quality.
[0352] The generation process of the diffracted sound enables the generation of diffracted sound to be restricted by the diffraction path DR having a length exceeding a preset allowable length.
[0353] According to this configuration, the generation of diffracted sound with an excessively long propagation distance and an extremely reduced amplitude is restricted. As a result, the amount of calculation can be suppressed without degrading the sound quality.
[0354] In the IR generation process, the grid size of the subspace SB used in the wave acoustics simulation is changed according to the size of the object OB.
[0355] According to this structure, simulation can be appropriately performed while suppressing the amount of calculation.
[0356] In the setting process of the subspace SB, the sound reception point T not on the same coordinate plane as the sound source S is projected onto the coordinate plane. In the setting process of the subspace SB, the two-dimensional space on the coordinate plane including the sound source S and the projected point on the coordinate plane is set as the subspace SB. The IR generation process calculates the IR in the subspace SB. In the IR generation process, an amplitude attenuation rate and a delay time corresponding to the difference between the distance on the diffraction path DR between the sound source S and the sound reception point T and the distance on the diffraction path DR between the sound source S and the projected point are added to the IR.
[0357] According to this configuration, IR calculation in three-dimensional space can be performed in a pseudo two-dimensional space. Therefore, the amount of calculation can be reduced.
[0358] In the setup process of the subspace SB, a two-dimensional space including the sound source S and the sound reception point T on the same plane PL is set as the subspace SB. The IR generation process calculates the IR in the subspace SB.
[0359] According to this configuration, IR calculation in three-dimensional space can be performed in a pseudo two-dimensional space. Therefore, the amount of calculation can be reduced.
[0360] In the diffracted sound generation process, an allowable propagation direction in which sound is allowed to propagate is obtained based on the directivity of the sound source S. The diffracted sound generation process enables the generation of diffracted sound propagating in directions other than the allowable propagation direction to be restricted.
[0361] According to this configuration, the generation of diffracted sound in the direction where the sound emission is greatly reduced according to the directivity is restricted. As a result, the amount of calculation can be suppressed without degrading the sound quality.
[0362] In the diffracted sound generation process, the ratio of the direct sound reaching the sound reception point T from the sound generation site SE via the diffraction site DF is calculated. The diffracted sound generation process enables the generation of diffracted sound from the sound generation site SE where the ratio of the direct sound becomes greater than the allowable standard to be restricted.
[0363] According to this structure, the generation of diffracted sound is restricted for the sound generation site where the direct sound dominates over the diffracted sound. As a result, the amount of calculation can be suppressed without degrading the sound quality.
[0364] It should be noted that the effects described in this specification are merely examples and are not restrictive, and other effects may be provided.
[0365] [Supplementary Explanation]
[0366] It should be noted that the present technology can also adopt the following configurations. (1)
[0368] An acoustic processing method executed by a computer, the acoustic processing method including:
[0369] Searching for a diffraction path from the sound source to the sound reception point; and
[0370] Generating diffracted sound obtained at the sound reception point based on a wave acoustics simulation in which the simulation time is restricted based on the distance of the diffraction path. (2)
[0372] According to the acoustic processing method of (1), wherein,
[0373] In the diffracted sound generation process, the time until the diffracted sound first reaches the sound reception point is calculated as the arrival time based on the distance of the diffraction path, and a time equal to or greater than 1 times and equal to or less than 5 times the arrival time is set as the simulation time. (3)
[0375] The acoustic processing method according to (1) or (2), wherein
[0376] The generation process of the diffracted sound includes: obtaining the diffracted part on the object where the diffraction phenomenon occurs; and selectively performing wave acoustic simulation on the subspace including the diffracted part to generate the impulse response in the subspace. (4)
[0378] The acoustic processing method according to (3), wherein
[0379] In the generation process of the impulse response, a trained AI model is used to generate the impulse response. In the AI model, the relationship between the simulation conditions of the wave acoustic simulation and the simulation results obtained under the simulation conditions is learned as the relationship between the input data and the correct answer data. (5)
[0381] The acoustic processing method according to (4), wherein
[0382] The AI model encodes the object information related to the diffracted part to output a feature quantity, and outputs the impulse response based on the feature quantity. (6)
[0384] The acoustic processing method according to any one of (3) to (5), wherein
[0385] The subspace is a space including at least one of the sound source and the sound receiving point. (7)
[0387] The acoustic processing method according to (6), wherein
[0388] The sound source is a virtual sound source arranged on the diffraction path between the real sound source and the diffracted part, and
[0389] The generation process of the impulse response adds the amplitude attenuation rate and delay time corresponding to the attenuation distance between the real sound source and the virtual sound source to the impulse response. (8)
[0391] The acoustic processing method according to (6) or (7), wherein
[0392] The sound receiving point is a virtual sound receiving point arranged on the diffraction path between the real sound receiving point and the diffracted part, and
[0393] The generation process of the impulse response adds the amplitude attenuation rate and delay time corresponding to the attenuation distance between the real sound receiving point and the virtual sound receiving point to the impulse response. (9)
[0395] The acoustic processing method according to any one of (3) to (8) further includes:
[0396] Obtaining the size that can be simulated in one frame period as the maximum size, and setting the subspace as a space with a size equal to or smaller than the maximum size. (10)
[0398] The acoustic processing method according to (9), wherein
[0399] In the setting process of the subspace, the subspace is set as a space including the surface area or volume of the object at a preset ratio. (11)
[0401] The acoustic processing method according to (9), wherein
[0402] In the setting process of the subspace, one or more diffraction parts that can be included in the subspace are extracted from multiple diffraction parts included in the same object, and
[0403] In the generation process of the impulse response, wave acoustics simulation is selectively performed on the extracted one or more diffraction parts. (12)
[0405] The acoustic processing method according to any one of (9) to (11), wherein
[0406] The subspace is a spherical or circular space centered on the diffraction part. (13)
[0408] The acoustic processing method according to any one of (9) to (12), wherein
[0409] The setting process of the subspace includes:
[0410] Detecting multiple diffraction parts that diffract the sound output from the sound source towards the sound receiving point;
[0411] Based on the maximum size, classifying the detected multiple diffraction parts into multiple groups; and
[0412] Setting a subspace for each group in the classified groups, and
[0413] The generation process of the impulse response adds an amplitude attenuation rate and a delay time corresponding to the distance between the subspaces along the diffraction path to the impulse response. (14)
[0415] The acoustic processing method according to (13), wherein
[0416] The generation process of diffracted sound enables the generation of diffracted sound passing through a subspace with a number exceeding a preset allowable number to be restricted. (15)
[0418] An acoustic processing method according to any one of (3) to (14), wherein
[0419] The generation process of diffracted sound enables the generation of diffracted sound passing through a diffraction path with a length exceeding a preset allowable length to be restricted. (16)
[0421] An acoustic processing method according to any one of (3) to (15), wherein
[0422] In the generation process of the impulse response, the grid size of the subspace used in the wave acoustic simulation is changed according to the size of the object. (17)
[0424] An acoustic processing method according to any one of (3) to (16), wherein
[0425] The setting process of the subspace includes:
[0426] Projecting a sound receiving point that is not in the same coordinate plane as the sound source onto the coordinate plane; and
[0427] Setting the two-dimensional space on the coordinate plane including the sound source and the projection point on the coordinate plane as the subspace, and
[0428] The generation process of the impulse response includes:
[0429] Calculating the impulse response in the subspace; and
[0430] Adding an amplitude attenuation rate and a delay time corresponding to the difference between the distance on the diffraction path between the sound source and the sound receiving point and the distance on the diffraction path between the sound source and the projection point to the impulse response. (18)
[0432] An acoustic processing method according to any one of (3) to (16), wherein
[0433] In the setting process of the subspace, a two-dimensional space including the sound source and the sound receiving point on the same plane is set as the subspace, and
[0434] The generation process of the impulse response calculates the impulse response in the subspace. (19)
[0436] An acoustic processing method according to any one of (1) to (18), wherein
[0437] The generation process of diffracted sound enables obtaining allowed propagation directions for sound propagation based on the directivity of the sound source, and restricts the generation of diffracted sound propagating in directions other than the allowed propagation directions. (20)
[0439] The acoustic processing method according to any one of (1) to (19), wherein
[0440] The generation process of diffracted sound enables calculating the ratio of direct sound from the sound generation site passing through the diffraction site to the sound reception point, and restricts the generation of diffracted sound for the sound generation site where the ratio of direct sound becomes greater than the allowed standard. (21)
[0442] An acoustic processing device, comprising:
[0443] A path search unit that searches for a diffraction path from the sound source to the sound reception point; and
[0444] A signal processing unit that generates diffracted sound obtained at the sound reception point based on a fluctuating acoustic simulation in which the simulation time is restricted according to the distance of the diffraction path. (22)
[0446] An acoustic processing program implemented by a computer, the acoustic processing program comprising:
[0447] Searching for a diffraction path from the sound source to the sound reception point; and
[0448] Generating diffracted sound obtained at the sound reception point based on a fluctuating acoustic simulation in which the simulation time is restricted according to the distance of the diffraction path.
[0449] List of reference numerals
[0450] 100 Acoustic processing device
[0451] 131 Path search unit
[0452] 139 Signal processing unit
[0453] DF Diffraction site
[0454] DR Diffraction path
[0455] OB Object
[0456] PJ1, PJ2 Projection point
[0457] PL Plane
[0458] RL Real sound reception point
[0459] RS Real sound source
[0460] S Sound source
[0461] SB subspace
[0462] SE sound generation site
[0463] T sound reception point
[0464] t0 arrival time
[0465] VL virtual sound reception point
[0466] VS virtual sound source.
Claims
1. An acoustic processing method executed by a computer, the acoustic processing method comprising: Searching for a diffraction path from a sound source to a sound reception point; And Generating diffracted sound acquired at the sound reception point based on a fluctuating acoustic simulation in which the simulation time is limited according to the distance of the diffraction path.
2. The acoustic processing method according to claim 1, wherein, In the generation process of the diffracted sound, the time until the diffracted sound first reaches the sound reception point is calculated based on the distance of the diffraction path as the arrival time, and a time equal to or greater than 1 times and equal to or less than 5 times the arrival time is set as the simulation time.
3. The acoustic processing method according to claim 1, wherein, The generation process of the diffracted sound includes: acquiring a diffraction part on an object where a diffraction phenomenon occurs; and selectively performing the fluctuating acoustic simulation on a subspace including the diffraction part to generate an impulse response in the subspace.
4. The acoustic processing method according to claim 3, wherein, In the generation process of the impulse response, an AI model that has been learned is used to generate the impulse response. In the AI model, the relationship between the simulation conditions of the fluctuating acoustic simulation and the simulation results obtained under the simulation conditions is learned as the relationship between input data and correct answer data.
5. The acoustic processing method according to claim 4, wherein, The AI model encodes object information related to the diffraction part to output a feature quantity, and outputs the impulse response based on the feature quantity.
6. The acoustic processing method according to claim 3, wherein, The subspace is a space including at least one of the sound source and the sound reception point.
7. The acoustic processing method according to claim 6, wherein, The sound source is a virtual sound source provided on the diffraction path between the real sound source and the diffraction part, and The generation process of the impulse response adds an amplitude attenuation rate and a delay time corresponding to the attenuation distance between the real sound source and the virtual sound source to the impulse response.
8. The acoustic processing method according to claim 6, wherein, The sound reception point is a virtual sound reception point provided on the diffraction path between the real sound reception point and the diffraction part, and The generation process of the impulse response adds an amplitude attenuation rate and a delay time corresponding to the attenuation distance between the real sound reception point and the virtual sound reception point to the impulse response.
9. The acoustic processing method according to claim 3, further comprising: Acquiring a size that can be simulated in one frame period as the maximum size, and setting the subspace as a space having a size equal to or smaller than the maximum size.
10. The acoustic processing method according to claim 9, wherein, In the setting process of the subspace, the subspace is set as a space including a preset proportion of the surface area or volume of the object.
11. The acoustic processing method according to claim 9, wherein, In the setup process of the subspace, one or more diffractive parts that can be included in the subspace are extracted from multiple diffractive parts included in the same object, and in the generation process of the impulse response, the wave acoustics simulation is selectively performed on the one or more extracted diffractive parts.
12. The acoustic processing method according to claim 9, wherein, The subspace is a spherical or circular space centered on the diffractive part.
13. The acoustic processing method according to claim 9, wherein, The setup process of the subspace includes: Detecting multiple diffractive parts that diffract the sound output from the sound source towards the sound receiving point; Classifying the detected multiple diffractive parts into multiple groups based on the maximum size; and Setting the subspace for each group of the classified groups, and The generation process of the impulse response adds an amplitude attenuation rate and a delay time corresponding to the distance between the subspaces along the diffractive path to the impulse response.
14. The acoustic processing method according to claim 13, wherein, The generation process of the diffracted sound enables restricting the generation of the diffracted sound passing through the subspaces with a number exceeding a preset allowable number.
15. The acoustic processing method according to claim 3, wherein, The generation process of the diffracted sound enables restricting the generation of the diffracted sound passing through the diffractive path with a length exceeding a preset allowable length.
16. The acoustic processing method according to claim 3, wherein, In the generation process of the impulse response, the grid size of the subspace used in the wave acoustics simulation is changed according to the size of the object.
17. The acoustic processing method according to claim 3, wherein, The setup process of the subspace includes: Projecting the sound receiving point not in the same coordinate plane as the sound source onto the coordinate plane; and Setting the two-dimensional space on the coordinate plane including the sound source and the projection point on the coordinate plane as the subspace, and The generation process of the impulse response includes: Calculating the impulse response in the subspace; and Adding an amplitude attenuation rate and a delay time corresponding to the difference between the distance on the diffractive path between the sound source and the sound receiving point and the distance on the diffractive path between the sound source and the projection point to the impulse response.
18. The acoustic processing method according to claim 3, wherein, In the setup process of the subspace, the two-dimensional space including the sound source and the sound receiving point on the same plane is set as the subspace, and The generation process of the impulse response calculates the impulse response in the subspace.
19. The acoustic processing method according to claim 1, wherein, The generation process of the diffracted sound enables obtaining an allowable propagation direction for allowing sound propagation based on the directivity of the sound source, and restricting the generation of the diffracted sound propagating in directions other than the allowable propagation direction.
20. The acoustic processing method according to claim 1, wherein, The generation process of the diffracted sound enables the calculation of the ratio of the direct sound that reaches the sound receiving point from the sound generation site through the diffraction site, and restricts the generation of the diffracted sound for the sound generation site where the ratio of the direct sound becomes greater than the allowable standard.
21. An acoustic processing device, comprising: a path search unit that searches for a diffraction path from a sound source to a sound receiving point; and a signal processing unit that generates diffracted sound obtained at the sound receiving point based on a fluctuating acoustic simulation in which the simulation time is restricted according to the distance of the diffraction path.
22. An acoustic processing program implemented by a computer, the acoustic processing program comprising: searching for a diffraction path from a sound source to a sound receiving point; and generating diffracted sound obtained at the sound receiving point based on a fluctuating acoustic simulation in which the simulation time is restricted according to the distance of the diffraction path.
Citation Information
Patent Citations
Acoustical signal processor
JP2000267675A