Audible simulation system and audible simulation method
Patent Information
- Application Number
- JP2025027523
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2026-09-04
AI Technical Summary
【0007】 本開示の可聴化シミュレーションシステム、および可聴化シミュレーション方法によれば、仮想空間ごとの音質がユーザに認知されやすい。
Smart Images

Figure 2026141131000001_ABST
Abstract
Description
[Technical Field]
[0001] This disclosure relates to an audible simulation system for reproducing sound in a virtual space, and to an audible simulation method. [Background technology]
[0002] One example of a virtual space provisioning device involves overlaying three-dimensional objects onto parts of the virtual space that have a different shape from the real world. The virtual space provisioning device performs acoustic simulation of the virtual space based on ambient sounds collected in the real world. This allows the virtual space provisioning device to help users viewing the virtual space perceive more accurate ambient sounds (see, for example, Patent Documents 1 and 2). [Prior art documents] [Patent Documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2017-146762 [Patent Document 2] Japanese Patent Publication No. 2024-72072 [Overview of the project] [Problems that the invention aims to solve]
[0004] On the other hand, human cognitive characteristics mean that differences between two virtual spaces can be perceived either auditorily or visually. In the latter case, the sound quality of each virtual space perceived auditorily is unlikely to be retained in the user's memory. Therefore, there is still a strong need for virtual space providers to allow users to perceive differences in sound quality between the virtual space in question and other virtual spaces when providing sound experiences within that virtual space. [Means for solving the problem]
[0005] The audible simulation system for solving the above problems is an audible simulation system that reproduces sound in an indoor space in a virtual space. This audible simulation system has a setting unit that changes the acoustic parameters of the indoor space, including the arrangement of the sound source and the sound receiving point in the indoor space, and the arrangement of objects installed in the indoor space, with each of the sound source and the sound receiving point in the indoor space as the target; a space generation unit that changes the virtual space to be displayed on a display device using the arrangement of the target dynamically changed by the setting unit; a sound generation unit that estimates the audible sound reaching the sound receiving point from the original sound of the sound source using the acoustic parameters dynamically changed by the setting unit; and a sound quality presentation unit that presents the properties of the audible sound that can be perceived from the audible sound estimated by the sound generation unit.
[0006] The audible simulation method for solving the above problems is an audible simulation method that reproduces sound in an indoor space in a virtual space. This audible simulation method is based on the sound source and the sound receiving point in the indoor space, and includes dynamically changing the acoustic parameters of the indoor space, including the arrangement of the objects and the arrangement of objects installed in the indoor space; changing the virtual space displayed on a display device using the dynamically changed arrangement of the objects; estimating the audible sound reaching the sound receiving point from the original sound of the sound source using the dynamically changed acoustic parameters; and presenting the properties of the audible sound that can be perceived from the estimated audible sound. [Effects of the Invention]
[0007] According to the audible simulation system and audible simulation method disclosed herein, the sound quality of each virtual space is easily perceived by the user. [Brief explanation of the drawing]
[0008] [Figure 1] Figure 1 is a functional block diagram showing the audible simulation system. [Figure 2] Figure 2 is a schematic diagram illustrating the virtual space. [Figure 3]FIG. 3 is a schematic diagram illustrating cells for explaining response data. [Figure 4] FIG. 4 is a configuration diagram showing the configuration of object data. [Figure 5] FIG. 5 is a configuration diagram showing the configuration of convolution data. [Figure 6] FIG. 6 is a graph showing cardioid directivity. [Figure 7] FIG. 7 is a flowchart showing the flow of an audibilization simulation method. [Figure 8] FIG. 8 is a sequence chart showing the flow of sound generation processing. [Figure 9] FIG. 9 is a schematic diagram showing an output result of an audibilization simulation system. [Figure 10] FIG. 10 is a sequence chart showing another example of the flow of sound generation processing. Mode for Carrying Out the Invention
[0009] As an embodiment of an audibilization simulation system and an audibilization simulation method, the audibilization simulation system and the audibilization simulation method for generating an indoor space 32R used for property viewing will be described below.
[0010] Overview of System As shown in FIG. 1, the audibilization simulation system 10 includes a control unit 20 and a storage unit 30. The control unit 20 includes a setting unit 21, a space generation unit 22, a sound generation unit 23, a sound quality presentation unit 24, an original sound presentation unit 25, and a history management unit 26. The storage unit 30 stores original sound data 31, space data 32, object data 33, response data 34, convolution data 35, and conversion data 36.
[0011] The audibilization simulation system 10 communicates with one or more terminal devices 50 via a network. The terminal device 50 includes a projection device 51, a sound output device 52, and an input device 53.
[0012] The audible simulation system 10 causes each terminal device 50 to reproduce the field of view 51A of the indoor space 32R (see Figure 2) in the virtual space, and the audible sound that reaches the sound receiving point P2 (see Figure 3) in the indoor space 32R. The audible simulation system 10 also causes each terminal device 50 to present the properties of the audible sound that can be perceived from the audible sound.
[0013] The audible simulation system 10, in accordance with instructions for dynamic repositioning of the sound source P1 and the sound receiving point P2, causes the terminal device 50 to reproduce a visual field image 51A and audible sound corresponding to the repositioning, and also causes the terminal device 50 to present the properties of the audible sound that can be perceived from the audible sound.
[0014] The audible simulation system 10, in accordance with the instruction to modify the original sound emitted by the sound source P1, causes the terminal device 50 to reproduce a visual image 51A and an audible sound corresponding to the modification, and also causes the terminal device 50 to present the properties of the audible sound that can be perceived from the audible sound.
[0015] The audible simulation system 10, in accordance with instructions to change the arrangement of objects or the acoustic characteristics of objects, causes the terminal device 50 to reproduce a visual field image 51A and audible sound corresponding to the changes, and also causes the terminal device 50 to present the properties of the audible sound that can be perceived from the audible sound.
[0016] [Configuration of the control unit 20] The control unit 20 comprises a processor and memory. The storage unit 30 stores processing data and instructions. The control unit 20 executes instructions using the processing data. The control unit 20 includes a network controller for communicating with the terminal device 50 via a network.
[0017] The control unit 20 functions as a setting unit 21, a spatial generation unit 22, a sound generation unit 23, a sound quality presentation unit 24, a source sound presentation unit 25, and a history management unit 26 by executing commands. The control unit 20, by executing commands, obtains data corresponding to the sound receiving point P2 from spatial data 32 for generating the indoor space 32R. The control unit 20 obtains data corresponding to the sound source P1 and the sound receiving point P2 from original sound data 31, object data 33, response data 34, convolution data 35, and transformation data 36 for reproducing audible sound. The control unit 20 generates a field image 51A and audible sound corresponding to the sound source P1 and the sound receiving point P2 using various data corresponding to the sound source P1 and the sound receiving point P2. The control unit 20 calculates the properties of the audible sound that can be perceived from the generated audible sound.
[0018] The control unit 20 communicates with the terminal device 50 to cause the projection device 51 to reproduce a field of view image 51A of the indoor space 32R, and to cause the sound output device 52 to reproduce audible sound. The control unit 20 also communicates with the terminal device 50 to cause the projection device 51 to display an index 24A (see Figure 9) indicating the properties of audible sound.
[0019] [Configuration of memory unit 30] The original sound data 31 is a dry source acquired in advance in an acoustic environment. The original sound data 31 may be speech or animal sounds, or environmental sounds such as city sounds or room sounds. The original sound data 31 may be artificial sounds such as musical sounds or signal sounds, physical sounds such as collision sounds or plosive sounds, or white noise with sounds of various frequency bands mixed evenly. The original sound data 31 may consist of multiple dry sources that are different from each other, so associating the dry source with the original sound identifier.
[0020] Spatial data 32 is a three-dimensional model for constructing a virtual space. Spatial data 32 includes a model for constructing an indoor space 32R. Spatial data 32 includes dimensions and shape of the indoor space 32R. Spatial data 32 includes data representing the side wall surfaces 32W, floor surface, and ceiling surface that demarcate the indoor space 32R. Spatial data 32 may be represented in polygon format, voxel format, or point cloud format.
[0021] As shown in Figure 2, the indoor space 32R is divided by the side walls 32W, the floor, and the ceiling. The view 51A of the indoor space 32R may include partitions 32P. Partitions 32P create passages in the indoor space 32R or divide spaces connected by passages. The view 51A of the indoor space 32R may also include independent spaces 32A. Independent spaces 32A include meeting rooms separated by transparent partitions or studies separated by walls. The view 51A of the indoor space 32R may also include sound-absorbing members 32B. Sound-absorbing members 32B are installed on the side walls 32W, the floor, the ceiling, etc. The indoor space 32R may include an avatar HA that acts as a surrogate for the user viewing the indoor space 32R, or an avatar HA that acts as a surrogate for the user guiding the indoor space 32R.
[0022] As shown in Figure 3, the indoor space 32R is divided into multiple cells 32C by multiple grids 32G. Each cell 32C may be arranged in a rectangular grid along the horizontal plane at the average height of a human ear. The sound source P1 and the sound receiving point P2 belong to one of the cells 32C.
[0023] As shown in Figure 4, object data 33 associates the physical property data of an object with an object identifier. The physical property data of an object may be the sound absorption coefficient of the object for each of different frequencies, the mass of the object, or the density of the object. The physical property data of an object may also be the transmission loss coefficient, scattering coefficient, or acoustic impedance of the object for each of different frequencies. For example, object identifier M1 may be associated with a sound absorption coefficient of "0.1" in frequency band A and a sound absorption coefficient of "0.15" in frequency band B. Object identifier M2 may be associated with a sound absorption coefficient of "0.1" in frequency band A and a sound absorption coefficient of "0.2" in frequency band B. The physical property data of an object may include the amount of light in the object, the optical spectrum of the object, or the microstructure and layer structure on the surface of the object.
[0024] The object data 33 associates the object identifier with object data representing the object. The object data representing the object includes the object's position, dimensions, shape, and surface texture. The surface texture is image data that represents the visual pattern and color characteristics of the object's surface. The object may be a partition 32P, an independent space 32A separated by partition 32P, or a sound-absorbing member 32B. The object may be an avatar HA acting as a surrogate for a viewing user, or an avatar HA acting as a surrogate for a guidance user. The object may be a structure fixed in the indoor space 32R, or a mobile body moving within the indoor space 32R. The position of the object representing the object belongs to at least one cell 32C.
[0025] The response data 34 includes a spatial impulse response. The spatial impulse response represents the propagation characteristics from one cell 32C that divides the spatial data 32 to another cell 32C other than the said cell 32C. The response data 34 includes spatial impulse responses between each of the four directions at the center of one cell 32C and each of the four directions at the center of the other cell 32C. The four directions at the center of cell 32C are the directions from that cell 32C to the other cells 32C located in front of, behind, to the left and right of that cell 32C. The spatial impulse response is defined in 16 pairs (= 4 directions × 4 directions) between one cell 32C and another cell 32C. The response data 34 associates one spatial impulse response with a combination of one direction of one cell 32C and one direction of the other cell 32C other than the said cell 32C.
[0026] The response data 34 includes spatial impulse responses for all combinations of each cell 32C that divides the spatial data 32 and other cells 32C. Furthermore, the response data 34 includes spatial impulse responses for all combinations of cells 32C where an object may be placed and the object identifier of that object. In other words, the response data 34 includes spatial impulse responses for all combinations of the dimensions and shape of the indoor space 32R, the sound source P1, the receiving point P2, the position of the object, and the object identifier of that object. The impulse responses are obtained in advance by geometric acoustic analysis such as the ray method or beam tracing method, or by wave acoustic analysis such as the finite element method or boundary element method.
[0027] As shown in Figure 5, the convolution data 35 includes convolution components. The convolution components are obtained by convolving the waveform data of an energy cardioid characteristic sound obtained from a dry source into a spatial impulse response. The convolution data 35 includes multiple component identifiers for each combination of source identifier and cell 32C. The convolution data 35 associates the convolution components with the position data of another cell 32C and each component identifier.
[0028] For example, component identifier R1 is associated with the position data "0,0" and the convolution component "TD001". "TD001" consists of 16 sets of convolution components. The convolution components of "TD001" are obtained from 16 sets of spatial impulse responses associated with combinations of one cell 32C and cell 32C represented by "0,0". Component identifier R2 is associated with the position data "0,1" and the convolution component "TD002". "TD002" also consists of 16 sets of convolution components. The convolution components of "TD002" are obtained from 16 sets of spatial impulse responses associated with combinations of one cell 32C and cell 32C represented by "0,1".
[0029] The convolution data 35 includes convolution components for all combinations of each cell 32C and other cells 32Cs other than the cell 32C in question. Furthermore, the convolution data 35 includes convolution components for all combinations of cells 32C in which an object may be placed and the object identifier of that object. In other words, the convolution data 35 includes convolution components for all combinations of the dimensions and shape of the indoor space 32R, the sound source P1, the receiving point P2, the position of the object, the object identifier of that object, the original sound identifier, etc.
[0030] The converted data 36 includes a head impulse response. The head impulse response represents the propagation characteristics of sound in the direction of arrival, from the vicinity of the human head to both ears. [Configuration of the control unit 20] Returning to Figure 1, the setting unit 21 sets various parameters for generating the field image 51A of the indoor space 32R and the audible sound of the indoor space 32R. The various parameters include the dimensions and shape of the indoor space 32R. The various parameters include the sound source P1 and the sound receiving point P2 in the indoor space 32R. The various parameters include the position of an object in the indoor space 32R and the object identifier of that object. The various parameters include the user's position, viewpoint, direction of gaze, and the direction of the head facing forward in the indoor space 32R. The various parameters may include the position the user is fixated on. The various parameters may include the original sound identifier. The setting unit 21 receives a change request from the terminal device 50 and updates the various parameters according to the request.
[0031] The dimensions and shape of the indoor space 32R, the sound source P1, the user's position (receiving point P2), the direction of the head's front, the position of an object, the object's identifier, and the original sound identifier constitute the acoustic parameters.
[0032] The spatial generation unit 22 identifies the field of view in the indoor space 32R and the objects placed within the field of view based on various parameters dynamically changed by the setting unit 21. The spatial generation unit 22 identifies the field of view in the indoor space 32R that is compatible with the direction of the user's gaze, centered on the user's position in the indoor space 32R. The spatial generation unit 22 generates a field of view image 51A, which is an image containing objects within the identified field of view, from the spatial data 32 and object data 33. The spatial generation unit 22 may generate the field of view image 51A in such a way that it simulates the propagation of light in the indoor space 32R and the optical effects such as reflection and scattering on the surface of objects in the indoor space 32R.
[0033] [Configuration of the sound generation unit 23] The sound generation unit 23 comprises a source processing unit 23A, a propagation processing unit 23B, and a sound receiving processing unit 23C (see Figure 8). The sound generation unit 23 performs static sound generation and dynamic sound generation (see Figure 8). Static sound generation generates audible sound in an indoor space 32R with fixed acoustic parameters. Static sound generation includes the reproduction of the directivity of the sound source system by the source processing unit 23A, the reproduction of the directivity of the propagation system by the propagation processing unit 23B, and the reproduction of the directivity of the sound receiving system by the sound receiving processing unit 23C.
[0034] The directivity reproduction of the sound source system involves referencing the dry source associated with the original sound identifier of the acoustic parameters and converting the dry source into energy cardioid characteristic sound in the four directions (front, back, left, and right) (step S41). This conversion is necessary because the sound source is given an energy cardioid characteristic in the impulse response analysis. As part of the directivity reproduction of the sound source system, the waveform data of the energy cardioid characteristic sound is calculated by multiplying the directivity of the dry source in the four directions by a conversion matrix. The matrix used to convert the dry source into cardioid characteristic sound is the first conversion matrix.
[0035] As shown in FIG. 6, the one-directional cardioid curve 36A has strong directivity at 0° in the front direction and weak directivity at 180° in the rear direction. The cardioid curve 36A is represented by the following Equation 1 using the component r and the angle θ. The components r in four directions on the cardioid curve 36A are 1.0 in the front direction, 0.5 in each of the left and right directions, and 0 in the rear direction.
[0036]
Math
[0037] Here, the sound pressures of cardioid sound sources in four directions (front F t , right R t , back B t , left L t ), the combined sound pressures in four directions (front F S , right R S , back B S , left L S ) are represented by the following Equation 2. For the convenience of performing ray tracing simulation, when an energy-based relational expression is converted into a sound pressure-based relational expression, the relational expression of Equation 2 is represented by Equation 3 using the square root of the conversion matrix. Note that Equation 4 represents the contribution component A of the cardioid sound source in Equation 3. Equation 5 represents the inverse matrix A -1 of the contribution component A. Then, when this inverse matrix A -1 is applied to dry sources (front F, right R, back B, left L) and the result (front F', right R', back B', left L') is further substituted into the sound pressure of the cardioid sound source, the combined four-direction sound pressures match the dry sources.
[0038] In this way, the directivity of a dry source is reproduced by the four-direction sound pressures converted from the dry source. Based on the above, the first conversion matrix for converting a dry source into cardioid characteristic sound is the inverse matrix A -1 , which is represented by Equation 5. The conversion of a dry source to a cardioid sound source is represented by Equation 6.
[0039]
Math
[0040]
number
[0041]
number
[0042]
number
[0043]
number
[0044] The propagation system's directional reconstruction refers to response data 34 to identify 16 sets of impulse responses between one cell 32C and another cell 32C. The propagation system's directional reconstruction convolves the waveform data of an energy cardioid characteristic sound onto the identified 16 sets of spatial impulse responses. The propagation system's directional reconstruction sums up all the convolutional components received by each directional component in the receiving system to generate a four-directional distributed sound (front F'', right R'', rear B'', left L'').
[0045] Furthermore, the propagation processing unit 23B performs directivity reproduction of the sound source system and directivity reproduction of the propagation system for each object position and dry source. As a result, the propagation processing unit 23B generates convolution data 35 having convolution components associated with acoustic parameters.
[0046] The directional reproduction of the receiving system converts the four-directional distributed sound into virtual sound sources positioned in the four directions of the receiving system. First, the directional reproduction of the receiving system converts the four-directional distributed sound into a coordinate system in which the direction in front of the head is defined as 0°. That is, the directional reproduction of the receiving system applies a rotation matrix to the four-directional distributed sound, using a rotation matrix that contributes to rotating the center of the receiving system in the four-directional cardioid characteristics counterclockwise around the axis of rotation.
[0047] For example, the distribution of sound from a 4-directional cardioid sound source (previous r F , right r R , after r B , left r L The result is given by Equation 7. The four-way distributed sound represented by Equation 7 is a sound source rotated counterclockwise by an angle θ' (previous r F ', right r R ', after r B ', left r L ') is represented by Equation 8. The distribution tone (pre-r) represented by Equation 7 F ~Left r L ) represented by Equation 8 (previous r F '~Left r L The rotational transformation that converts to ') is represented by equation 9. And, as shown in equation 10, the distribution tone (pre-r) of equation 9 is represented by equation 9. F ~Left r L By substituting the four-directional distributed sound (front F'' to left L'') into the given formula, the sound source after rotational transformation in the receiving system (front F'', right R''', rear B''', left L''') is obtained.
[0048]
number
[0049]
number
[0050]
number
[0051]
number
[0052] The directivity of the receiving system is then reproduced by reproducing the directivity of the rotation-transformed four-directional distributed sound from the four-directional plane wave virtual sound. The contribution to making the directivity of the receiving point P2 a sound pressure cardioid is expressed by Equation 11. In Equation 11, the sound pressure (pre-F) t ~Left L t ) represents the sound received by the cardioid characteristics in each direction, and also the sound pressure (preceding F). S ~Left L S ) represents the sound from four plane wave sound sources. Note that Equation 12 represents the contributing component B of the cardioid sound source in Equation 11.
[0053] Here, similar to the sound source system, the sound source (front F'''~left L''') obtained by convolving the impulse response into the cardioid-recorded distributed sound source is used to increase the sound pressure (front F t ~Left L t Substituting this into the formula, the sound pressure of a plane wave source (front F) S ~Left L S ) does not match the sound source (front F'''~left L'''). Therefore, similar to the sound source system, the inverse matrix B of the contributing component B is used for the sound source (front F'''~left L''''). -1 By applying this beforehand, the sound pressure of the plane wave sound source (pre-F) S ~Left L S The correspondence between the sound source (front F'''~left L''') and the sound source is achieved. However, since the contributing component B satisfies det(B)=0, there is no inverse matrix for contributing component B. For this reason, as shown in equation 13, the Moore-Penrose pseudo-inverse matrix B is used as a substitute for the inverse matrix. + The following is used. That is, the conversion formula that converts a cardioid-recorded sound source (front F'''~left L''') into a four-directional plane wave sound source (front F''''~left L'''') as a virtual sound source is expressed by Equation 14.
[0054] From the above, the directivity reproduction of the receiving system consists of a rotational transformation process of the cardioid characteristics at the receiving point P2 and a process of transforming the cardioid sound source into a virtual plane wave, and their contributions are expressed by Equation 15. Here, the rotational transformation matrix is the pseudo-inverse matrix B + The result obtained by multiplying by is the second transformation matrix.
[0055] Then, the directional characteristics of the receiving system are reproduced by referring to the conversion data 36, and the pseudo-inverse matrix B + The virtual sounds from four directions, obtained by multiplying these signals, are convolved with head impulse responses corresponding to each direction. This allows the receiving system's directional reproduction to generate the sounds heard by the left and right ears when the virtual sounds from four directions are radiated. The receiving system's directional reproduction then generates the audible sound heard by the left and right ears by adding up the four convolved sounds.
[0056]
number
[0057]
number
[0058]
number
[0059]
number
[0060]
number
[0061] The dynamic sound generation identifies the cell 32C to which the receiving point P2 belongs by referring to the acoustic parameters set in the setting unit 21. The dynamic sound generation also identifies three cells 32C adjacent to the cell 32C to which the receiving point P2 belongs, in order of proximity to the receiving point P2. The dynamic sound generation also refers to the convolution data 35 and reads out the convolution components associated with the cell 32C to which the sound source P1 belongs and each of the four identified cells 32C. Then, the dynamic sound generation extracts data of the length required to reproduce the audible sound from the four read-out convolution components.
[0062] Here, the receiving point P2 is likely to be different from the center of cell 32C to which the receiving point P2 belongs. To interpolate this positional difference, dynamic sound generation generates an interpolated sound at the receiving point P2 using the extracted convolutional components. Dynamic sound generation performs linear interpolation using waveform data extracted from the four convolutional components, such that the contribution increases as it gets closer to the receiving point P2. In this way, dynamic sound generation generates an interpolated sound at the receiving point P2 that changes sensitively in response to changes in the receiving point P2. Then, by performing directional reproduction of the receiving system, dynamic sound generation generates an audible sound with dynamically changed acoustic parameters.
[0063] [Configuration of the sound quality presentation unit 24] The sound quality presentation unit 24 uses the audible sound generated by the sound generation unit 23 to index the properties of the audible sound. The sound quality presentation unit 24 transmits the indexed properties of the audible sound, along with the visual field image 51A and the audible sound, to the terminal device 50. The sound quality presentation unit 24 may display the index 24A indicating the properties of the audible sound on the projection device 51, or output it as sound to the sound output device 52.
[0064] The properties of audible sound can be either the clarity of the audible sound or an evaluation index value for spatial impression. The clarity of audible sound is the Speech Transmission Index (STI). The clarity of audible sound is indexed by a numerical value between 0 and 1. The more extraneous sounds such as reflected sound and noise are included in the audible sound at the receiving point P2, the closer the clarity of the audible sound is to 0. The closer the audible sound at the receiving point P2 is to the dry source, the closer the clarity of the audible sound is to 1.
[0065] The sound quality presentation unit 24 may be configured to allow selection of one of the following frequency characteristics when determining the STI value: general human characteristics, characteristics of a user guiding the indoor space 32R, or a user's selection when listening to the indoor space 32R. The sound quality presentation unit 24 may select one of these frequency characteristics based on a selection request from the input device 53 of the terminal device 50.
[0066] When determining the STI value, the sound quality display unit 24 may set the position of the user emitting the sound as the sound source P1, such as the position of the user guiding the indoor space 32R. When determining the STI value, the sound quality display unit 24 may set a pre-set position in the indoor space 32R as the sound source P1.
[0067] The evaluation index values for spatial impression may also be temporal index values such as the reverberation time of audible sound or the initial delay time of audible sound. The evaluation index values for spatial impression may also be spatial index values such as the magnitude of reflected sound from the side relative to the direct sound from in front of the receiving point P2, or the magnitude of reflected sound from behind and to the sides of the receiving point P2. The sound quality presentation unit 24 may, for example, overlay the STI value onto the visual field image 51A.
[0068] The properties of audible sound may be the direction of arrival of the audible sound, the propagation path of the audible sound, or the volume of the audible sound in the left and right ears. The properties of audible sound may also be the volume absorbed by an object or the volume blocked by an object. The sound quality display unit 24 calculates, for example, the time-averaged equivalent noise level from the audible sound in the left and right ears over a predetermined time width. The sound quality display unit 24 may then display the time-averaged values for the left and right ears using level bars superimposed on the visual field image 51A. The sound quality display unit 24 also calculates, for example, the maximum value of the equivalent noise level from the audible sound in the left and right ears. The sound quality display unit 24 may then display the maximum value on the level bars showing the time-averaged values for the left and right ears. The sound quality display unit 24 may, for example, superimpose the trajectory of a sound ray obtained by sound ray analysis onto the visual field image 51A. The sound quality display unit 24 may, for example, superimpose the amount of energy absorbed by sound ray analysis onto the sound-absorbing member 32B, or display it as a contour.
[0069] The original sound presentation unit 25 may receive a selection request from the terminal device 50 and transmit a plurality of dry sources constituting the original sound data 31 to the terminal device 50, causing the terminal device 50 to reproduce each dry source. The original sound presentation unit 25 may receive a selection request from the terminal device 50 and cause the terminal device 50 to select a dry source to be used for generating audible sound. The original sound presentation unit 25 may also set the selected dry source as the dry source to be used for generating audible sound in the setting unit 21. Furthermore, the original sound presentation unit 25 may receive a transmission request from the terminal device 50 and transmit the dry source used for generating audible sound to the terminal device 50. The original sound presentation unit 25 may also cause the terminal device 50 to reproduce the transmitted dry source.
[0070] The history management unit 26 stores the history of sound sources P1, sound receiving points P2, the positions of each object, object identifiers, original sound identifiers, and audible sounds generated based on the settings of the setting unit 21, over time. The history management unit 26 receives a transmission request from the terminal device 50 and transmits the history to the terminal device 50. The history management unit 26 displays the transmitted history on the terminal device 50.
[0071] The history management unit 26 may store comments 24B from users who have viewed the indoor space 32R in association with the history. The history management unit 26 may receive a registration request from the terminal device 50 and receive comments 24B related to audible sounds. The history management unit 26 may receive a transmission request from the terminal device 50 and send the comments 24B to the terminal device 50 along with the history. The history management unit 26 may display the transmitted comments 24B along with the history on the terminal device 50.
[0072] The accuracy of reproduction related to the directivity of audible sound at the receiving point P2 can be improved by subdividing the directivity of the sound at the sound source P1 and the directivity of the audible sound at the receiving point P2. On the other hand, subdividing the directivity of the sound at the sound source P1 and the directivity of the audible sound at the receiving point P2 leads to an enormous amount of computation required to reproduce the audible sound, reducing real-time performance. Furthermore, dynamic changes to acoustic parameters such as the sound source P1 and the receiving point P2 make it difficult to derive the spatial impulse response in real time. In this regard, associating the spatial impulse response with a combination of one cell 32C and another cell 32C makes it easier to derive the spatial impulse response in accordance with the subdivided directivity of the sound at the sound source P1 and the subdivided directivity of the audible sound at the receiving point P2. Furthermore, associating the spatial impulse response with a combination of one cell 32C and another cell 32C reduces the enormous computational cost required to derive the spatial impulse response.
[0073] [Terminal device 50] Returning to Figure 1, the terminal device 50 comprises a projection device 51, a sound output device 52, and an input device 53. The terminal device 50 further comprises a processing circuit and storage. The processing circuit comprises a processor and memory. The storage stores processing data and instructions. The processing circuit executes instructions using the processing data. The processing circuit comprises a network controller for connecting to the audible simulation system 10 via a network. By communicating with the audible simulation system 10, the processing circuit acquires various data for reproducing the field image 51A of the indoor space 32R and the audible sound corresponding to the sound receiving point P2. The processing circuit comprises a controller for causing the projection device 51, the sound output device 52, and the input device 53 to perform processing.
[0074] The terminal device 50 may be a head-mounted display (HMD) device worn on the user's head. The HMD device provides the user wearing the HMD device with a visual field image 51A and audible sound. The terminal device 50 may consist of a display device, a speaker, and a computer that controls the output of the display device and the speaker. The display device may be various types of displays such as flat or curved displays, or it may be a projection display that projects the visual field image 51A onto a panoramic screen or a dome screen.
[0075] The projection device 51 is an example of a display device. The projection device 51 receives data from the audible simulation system 10 for reproducing the field of view image 51A of the indoor space 32R through processing by a processing circuit. The projection device 51 reproduces the field of view image 51A of the indoor space 32R based on the data for reproducing the field of view image 51A of the indoor space 32R.
[0076] The sound output device 52 receives data for reproducing audible sound in the indoor space 32R from the audible simulation system 10 through processing by a processing circuit. Based on the data for reproducing audible sound in the indoor space 32R, the sound output device 52 reproduces audible sound in the indoor space 32R. The sound output device 52 transmits audible sound as sound vibrations to the left and right ears of a person. The sound output device 52 may be various earphone devices that use gas conduction, where vibration drivers are placed near the left and right ears, or solid conduction such as bone conduction or cartilage conduction. The sound output device 52 may be a channel-controlled type that simulates the direction of arrival by vibration drivers installed in the listening space of the indoor space 32R, or a spatial-controlled type that controls the sound field of the listening space. The sound output device 52 is an example of a sound output unit.
[0077] The input device 53 transmits data for reproducing audible sound to the audible sound simulation system 10 through processing by the processing circuit. The input device 53 dynamically changes the data for reproducing audible sound. The data for reproducing audible sound includes acoustic parameters. The data for reproducing audible sound includes the user's position in the indoor space 32R, the user's viewpoint, and the direction of the user's gaze. The data for reproducing audible sound may also include the position the user is fixated on. The position the user is fixated on may be calculated by the processing circuit of the terminal device 50 from the user's viewpoint and the direction of the user's gaze. The data for reproducing audible sound may also include the position, dimensions, shape, and surface texture of objects. The input device 53 may be a touch panel, a keyboard, a mouse, or a controller operated by the user's hand.
[0078] If the terminal device 50 is an HMD device, the HMD device includes a display that displays a field of view image 51A. The field of view image 51A displayed by the display has parallax between the field of view image 51A for the left eye and the field of view image 51A for the right eye in order to provide a three-dimensional field of view image 51A with a sense of depth. The display constitutes a projection device 51. The HMD device includes a speaker that outputs audible sound. The speaker constitutes a sound output device 52. The HMD device includes a sensor that detects the movement of the user's head. The movement of the user's head includes the position of the head and the orientation of the head. The sensor constitutes an input device 53. The HMD device may also include a tracking device that detects the movement of the user's hands and the movement of the user's eyes. The tracking device constitutes an input device 53.
[0079] [Method for simulating the creation of audible sounds] The following describes a method of audible simulation performed by the audible simulation system 10, in which the audible simulation system 10 performs dynamic sound generation in response to changes in the receiving point P2.
[0080] As shown in Figure 7, the audible simulation method includes spatial generation (step S11), operation decision (step S12), receiving point processing (step S13), sound generation (step S14), output processing (step S15), and termination decision (step S16).
[0081] As shown in Figure 8, the control unit 20 generates convolution data 35 as a preprocessing step 40A prior to dynamic sound generation. Specifically, the source processing unit 23A converts all dry sources included in the original sound data 31 into energy cardioid characteristic sounds (step S41). Next, the propagation processing unit 23B refers to the spatial data 32, object data 33, and response data 34 and calculates the convolution component associated with each acoustic parameter, which is a combination of object position and dry source. As a result, the audible simulation system 10 generates convolution data 35 (step S42).
[0082] Returning to Figure 7, in the spatial generation (step S11), the control unit 20 identifies the field of view range in the indoor space 32R and the objects to be placed within the field of view range based on various parameters set in the setting unit 21. The control unit 20 generates audible sound at the sound receiving point P2 using convolutional components associated with acoustic parameters. The audible simulation system 10 then causes the terminal device 50 to reproduce the visual image and audible sound.
[0083] In the operation decision (step S12), the control unit 20 determines whether or not the terminal device 50 requests a change in the sound receiving point P2. Until the control unit 20 accepts the request to change the sound receiving point P2, the audible simulation system 10 causes the terminal device 50 to maintain the visual image and audible sound (NO in step S12).
[0084] When the control unit 20 receives a request to change the sound receiving point P2, in the sound receiving point processing (step S13), the control unit 20 identifies the field of view in the indoor space 32R and the objects placed within the field of view based on the changed sound receiving point P2. Then, in sound generation (step S14), the control unit 20 generates an audible sound at the sound receiving point P2 using convolutional components associated with acoustic parameters including the changed sound receiving point P2.
[0085] That is, returning to Figure 8, the propagation processing unit 23B, as a dynamic sound generation 40B, refers to the acoustic parameters including the modified receiving point P2 and the convolution data 35, and reads out the convolution component associated with the acoustic parameters. The propagation processing unit 23B uses the read-out convolution component to generate an interpolated sound corresponding to the modified receiving point P2 (step S43). The receiving processing unit 23C converts the interpolated sound, which is a four-directional distributed sound, into virtual sounds arranged in the four directions of the receiving system (step S44). Then, the receiving processing unit 23C refers to the conversion data 36 and convolves the head impulse response corresponding to each direction into the four-directional virtual sounds (step S45). As a result, the control unit 20 generates the sound that would be heard by the left and right ears if the four-directional virtual sounds were radiated, and by adding up the four-directional convolution sounds, generates the audible sound that would be heard by the left and right ears of the modified receiving point P2.
[0086] Returning to Figure 7, in the output processing (step S15), the control unit 20 uses an audible sound corresponding to the modified receiving point P2 to index the properties of the audible sound. The control unit 20 causes the terminal device 50 to reproduce the visual image and audible sound corresponding to the modified receiving point P2, and also causes the terminal device 50 to output an index 24A indicating the properties of the audible sound corresponding to the modified receiving point P2. The audible simulation system 10 then causes the terminal device 50 to reproduce the visual image and audible sound corresponding to the modified receiving point P2 until the control unit 20 receives a request to stop the reproduction for the indoor space 32R (NO in step S16).
[0087] As shown in Figure 9, the visual image of the indoor space 32R reproduced on the terminal device 50 by the output processing (step S15) displays an index 24A indicating the properties of audible sound superimposed on the objects of the indoor space 32R. At this time, the control unit 20 may receive a registration request from the terminal device 50 and receive a comment 24B related to audible sound. The control unit 20 may also display the comment 24B sent from the terminal device 50 on the terminal device 50 along with the source identifier in the history.
[0088] [effect] According to the above embodiment, the following effects can be obtained. (1) The terminal device 50 outputs a visual image of the indoor space 32R, an audible sound, and an index 24A corresponding to the change in the sound receiving point P2. Therefore, the sound quality for each acoustic parameter is easily perceived by the user.
[0089] (2) The audible simulation system 10 includes convolution data 35, and the convolution components associated with acoustic parameters including the modified receiving point P2 are read out to the sound generation unit 23. This reduces the enormous amount of computation required for deriving the spatial impulse response and for convolution using the spatial impulse response.
[0090] (3) Multiple mutually different dry sources are configured to be selectable by the terminal device 50. This enables changing the dry source for each sound receiving point P2, and changing the dry source for each object's position. Furthermore, in constructing a real space similar to the indoor space 32R, the acoustic characteristics of the real space can be understood in more detail.
[0091] (4) The terminal device 50 outputs comments 24B regarding the audible sound at the modified receiving point P2. The terminal device 50 also outputs the source identification of the dry source used to generate the audible sound at the modified receiving point P2. This makes it easier to share simulation results among different users. It also makes it easier to compare simulation results among different acoustic parameters.
[0092] [Example of changes] The above embodiment can be implemented with the following modifications. Furthermore, these modifications can be combined to the extent that they do not contradict each other technically.
[0093] [Sound generation section 23] As shown in Figure 10, the sound generation unit 23 may separate the cardioid characteristic sound into a direct sound component 31A and a reflected sound component 31B, and generate an audible sound by using the convolution of the spatial impulse response only on the reflected sound component 31B.
[0094] For example, the response data 34 includes a spatial impulse response of the reflected sound for each combination of one cell 32C and another cell 32C different from that cell 32C. The propagation processing unit 23B performs preprocessing 40A, separating the components other than the direct sound component 31A in the cardioid characteristic sound as the reflected sound component 31B (step S41). The propagation processing unit 23B calculates the convolution component of the reflected sound by convolving the spatial impulse response into the reflected sound component 31B, and generates convolution data 35 composed of the convolution component of the reflected sound (step S42).
[0095] The propagation processing unit 23B, in dynamic sound generation, refers to the acoustic parameters set in the setting unit 21 to identify the cell 32C to which the receiving point P2 belongs. Then, the propagation processing unit 23B and the receiving processing unit 23C refer to the convolution data 35 to read out the convolution component of the reflected sound, and use the read out convolution component to generate an interpolated sound with the reflected sound component 31B (step S43). The receiving processing unit 23C uses the direct sound component 31A in the cardioid characteristic sound to rotate the direct sound component 31A in a coordinate system defined as 0° in the direction of the front of the head. Then, the receiving processing unit 23C adds the rotated direct sound component 31A and the interpolated sound with the reflected sound component 31B to convert them into virtual sounds in the four directions at the receiving point P2 (step S44).
[0096] According to this example of modification, the computational cost required to generate the convolution data 35 in preprocessing 40A is reduced. The sound generation unit 23 may read out the impulse response associated with the acoustic parameters dynamically changed by the setting unit 21 and use it for sound convolution to generate audible sound. That is, the audible simulation system 10 may omit the convolution data 35 from the data stored in the memory unit 30 and perform sound convolution to generate audible sound each time the setting unit 21 dynamically changes the acoustic parameters.
[0097] The sound generation unit 23 may cause the sound output device 52 to output a reference tone for adjusting the volume of the sound output device 52. The listener, simultaneously with the reference tone of the sound output device 52, adjusts the output of the sound output device 52 to match the volume of the reference tone played by another sound output device 52 that serves as a volume reference, for example, a specific type of smartphone or a headphone-type playback device with calibrated volume. In this case, the sound generation unit 23 may receive the adjustment amount from the terminal device 50 and output an audible sound again at the volume adjusted based on the adjustment amount.
[0098] [Virtual Space] The virtual space where the audible sound at the receiving point P2 is estimated may be a shared space within a building, such as an entrance or lobby; an equipment space, such as a machine room or electrical room; or a service space, such as a restroom or warehouse. The virtual space where the audible sound at the receiving point P2 is estimated may also be an open space, such as a terrace or courtyard.
[0099] The objects included in the acoustic parameters may be side walls 32W in the indoor space 32R, floors, or ceilings. The objects included in the acoustic parameters may be fixed structures installed in a shared space, equipment installed in a facility space, or objects installed in an open space.
[0100] [others] The audible simulation system 10 may display an icon that emits sound near the sound source P1 in the field of view 51A of the indoor space 32R, or it may make the icon blink. The audible simulation system 10 may also display a color indicating that the object set as the sound source P1 is the sound source P1 in the field of view 51A of the indoor space 32R.
[0101] The audible simulation system 10 may consist of a single server device equipped with multiple virtual servers, or it may consist of a set of multiple physical servers. The hardware elements constituting the audible simulation system 10 may be implemented by various circuit elements. Each function of the audible simulation system 10 may be implemented by a circuit including one or more processing circuits.
[0102] The audible simulation system 10 and the terminal device 50 include a processor. The processor may include a processor that executes readable instructions. The processor is a processing circuit comprising transistors and non-transistor circuits. The processor may also be a CPU configured to execute a program stored in memory. The processor may also be a general-purpose processor configured to perform the functions it possesses, or a general-purpose processor programmed to perform the functions it possesses. The processor may also be a special-purpose processor, an integrated circuit, or an ASIC. The processor may be implemented using a circuit that includes a combination of these various processors.
[0103] At least one block constituting the audible simulation system 10 may implement two or more functions according to the flowchart showing the audible simulation method, or it may implement two or more functions simultaneously, or it may implement two or more functions in reverse order.
[0104] • At least one block in the flowchart showing the audible simulation method may execute two or more steps simultaneously, or two or more steps in reverse order. At least one process constituting the audible simulation method may be implemented in the audible simulation system 10 as a module of code that makes the process executable, or as a segment.
[0105] At least one process for implementing the audio simulation method is stored in a computer-readable medium. The computer-readable medium may be ROM, FLASH memory, a hard disk, or an optical disc. The computer-readable medium may also be network storage or cloud storage.
[0106] The network may be a public network such as the internet, a private network such as a local area network or wide area network, or a combination of both. The network may be a wired network or a wireless network.
[0107] [Note] The technical ideas derived from the above embodiments and their modifications are described below. [Note 1] An audible simulation system that reproduces sound in an indoor space in a virtual space, Each of the sound source and sound receiving point in the indoor space is a target, and the setting unit changes the acoustic parameters of the indoor space, including the arrangement of the targets and the arrangement of objects installed in the indoor space. A space generation unit that modifies the virtual space to be displayed on the display device using the arrangement of the target dynamically changed by the setting unit, A sound generation unit estimates the audible sound reaching the sound receiving point from the original sound of the sound source using the acoustic parameters dynamically changed by the setting unit, The system includes a sound quality presentation unit that presents the properties of the audible sound that can be perceived from the audible sound estimated by the sound generation unit, A sound-making simulation system characterized by the following features.
[0108] [Note 2] The aforementioned properties of the audible sound are at least one selected from the group consisting of clarity and spatial impression evaluation index values. The audible simulation system described in Appendix 1.
[0109] [Note 3] The properties of the audible sound are at least one selected from the direction of arrival, the propagation path, the volume in the left and right ears, the volume absorbed by the object, and the volume blocked by the object. The audible simulation system described in Appendix 1.
[0110] [Note 4] The presentation of the properties of the audible sound is performed using at least one of the display by the display device and the sound output by the sound generation unit. An audible simulation system as described in any one of the appendices 1 through 3.
[0111] [Note 5] The sound generation unit, The system divides the virtual space into multiple cells, and when the sound source and the sound receiving point are placed in separate cells, the system stores the impulse response between the sound source and the sound receiving point in relation to the acoustic parameters. From this storage unit, the system reads the impulse response associated with the dynamically changed arrangement of the target, and uses it for sound convolution to generate the audible sound. An audible simulation system as described in any one of the appendices 1 through 4.
[0112] [Note 6] The acoustic parameters include the arrangement of the multiple objects installed in the indoor space, The sound generation unit reads the impulse response associated with the arrangement of the object dynamically changed by the setting unit from the storage unit, which stores the impulse response between the sound source and the sound receiving point in relation to the acoustic parameters when the sound source, the sound receiving point, and each object are arranged in separate cells, and uses it for sound convolution to generate the audible sound. The audible simulation system described in Appendix 5.
[0113] [Note 7] The storage unit stores the impulse response between the four directions at the sound source and the four directions at the sound receiving point. The sound generation unit uses a first transformation matrix that converts the components in each direction of the sound source into components having cardioid directivity in that direction so as to maintain the directivity of the sound source in each direction, and as a convolution of the sound, it convolves the components in each direction obtained by applying the first transformation matrix to the original sound with the impulse response in that direction. The audible simulation system described in Appendix 6.
[0114] [Note 8] The sound generation unit, A second transformation matrix is used to transform the components in each direction at the receiving point into components having cardioid directivity in that direction, so as to maintain the directivity at the receiving point in each direction. The second transformation matrix is then applied to the components in each direction into which the impulse response is convolved, and the head-of-direction transfer function for that direction is further convolved into the components in each direction. The audible simulation system described in Appendix 7.
[0115] [Note 9] The virtual space is divided into multiple cells, The sound generation unit, A storage unit is used to store convolutional components for each acoustic parameter when the sound source, the receiving point, and each object are arranged in separate cells, wherein the convolutional components are the result obtained by convolving the impulse response between the sound source and the receiving point in each arrangement with the sound of the sound source, and the setting unit reads out the convolutional components associated with the acoustic parameters that have been dynamically changed, and estimates the audible sound reaching the receiving point using the convolutional components. The audible simulation system described in Appendix 1.
[0116] [Note 10] The system further includes a sound presentation unit that allows for the selection and presentation of multiple mutually different original sounds, The sound generation unit estimates the audible sound using the original sound selected by the original sound presentation unit. An audible simulation system as described in any one of the items from Appendix 1 to Appendix 9.
[0117] [Note 11] The sound generation unit, The sound output unit outputs the audible sound and a reference sound, receives an adjustment amount for the audible sound relative to the volume of the reference sound, and outputs the audible sound again at the volume adjusted based on the adjustment amount. The audible simulation system described in Appendix 1 to Appendix 10.
[0118] [Note 12] The system further includes a history management unit that stores over time the virtual space displayed by the space generation unit and the audible sound estimated by the sound generation unit in the virtual space. An audible simulation system as described in any one of the appendices 1 through 11.
[0119] [Note 13] The history management unit receives comments regarding the virtual space displayed by the space generation unit and the audible sound estimated by the sound generation unit in the virtual space, and stores the comments in association with the virtual space and the audible sound. The audible simulation system described in Appendix 12.
[0120] [Note 14] A server connected to multiple terminal devices, The system comprises the setting unit, the space generation unit, the sound generation unit, and the sound quality presentation unit, The virtual space is displayed on the display device of each terminal device. The sound output section of each terminal device is made to output the audible sound. An audible simulation system as described in any one of the appendices 1 through 13.
[0121] [Note 15] An audible simulation system that reproduces sound in an indoor space in a virtual space, Each of the sound source and sound receiving point in the indoor space is a target, and the setting unit changes the acoustic parameters of the indoor space, including the arrangement of the targets and the arrangement of objects installed in the indoor space. A space generation unit that modifies the virtual space to be displayed on the display device using the arrangement of the target dynamically changed by the setting unit, The system comprises: a storage unit that divides the virtual space into multiple cells and stores the acoustic parameters associated with the acoustic parameters when the sound source, the sound receiving point, and each object are placed in separate cells, and a setting unit reads the impulse response associated with the acoustic parameters dynamically changed by the setting unit from the storage unit, and uses a first transformation matrix that converts the components in each direction of the sound source into components having cardioid directivity in that direction so as to maintain the directivity of the sound source, and then convolves the read-out impulse response in that direction into the components in each direction obtained by applying the first transformation matrix to the original sound, thereby estimating the audible sound reaching the sound receiving point from the original sound of the sound source. A sound-making simulation system characterized by the following features.
[0122] [Note 16] An audible simulation system that reproduces sound in an indoor space in a virtual space, Each of the sound source and sound receiving point in the indoor space is a target, and the setting unit changes the acoustic parameters of the indoor space, including the arrangement of the targets and the arrangement of objects installed in the indoor space. A space generation unit that modifies the virtual space to be displayed on the display device using the arrangement of the target dynamically changed by the setting unit, The system comprises a storage unit that stores convolutional components for each acoustic parameter when the virtual space is divided into multiple cells and the sound source, the sound receiving point, and each object are arranged in separate cells, wherein the convolutional components are the result obtained by convolving the impulse response between the sound source and the sound receiving point in each arrangement with the sound of the sound source, and a sound generation unit that reads out the convolutional components associated with the acoustic parameters dynamically changed by the setting unit and estimates the audible sound reaching the sound receiving point from the original sound of the sound source. A sound-making simulation system characterized by the following features.
[0123] [Note 17] A method for reproducing sound in an indoor space within a virtual environment, Each of the sound source and the sound receiving point in the aforementioned indoor space is a target, and the acoustic parameters of the indoor space, including the arrangement of the targets and the arrangement of objects installed in the indoor space, are to be dynamically changed. The virtual space is divided into a plurality of cells, and the sound source, the receiving point, and each of the objects are arranged in separate cells. The impulse response between the sound source and the receiving point is stored in association with the acoustic parameters of the acoustic parameters. The impulse response associated with the dynamically changed acoustic parameters is read from a storage unit that stores the impulse response associated with the acoustic parameters of the acoustic parameters. A first transformation matrix is used to convert the components in each direction of the sound source into components having cardioid directivity in that direction so as to maintain the directivity of the sound source. The read-out impulse response in that direction is convolved into the components in each direction obtained by applying the first transformation matrix to the original sound, thereby estimating the audible sound reaching the receiving point from the original sound of the sound source. A method for simulating audibility, characterized by the following features.
[0124] [Note 18] An audible simulation system that reproduces sound in an indoor space in a virtual space, Each of the sound source and sound receiving point in the indoor space is a target, and the setting unit changes the acoustic parameters of the indoor space, including the arrangement of the targets and the arrangement of objects installed in the indoor space. A space generation unit that modifies the virtual space to be displayed on the display device using the arrangement of the target dynamically changed by the setting unit, The virtual space is divided into a plurality of cells, and a storage unit is used to store a convolution component for each acoustic parameter when the sound source, the receiving point, and each object are placed in separate cells, wherein the convolution component is the result obtained by convolving the impulse response between the sound source and the receiving point at the acoustic parameter associated with the convolution component with the sound of the sound source, and the setting unit reads out the convolution component associated with the acoustic parameter that has been dynamically changed, and estimates the audible sound reaching the receiving point from the original sound of the sound source. A method for simulating audibility, characterized by the following features.
[0125] The audible simulation systems described in appendices 15 and 16 above, and the audible simulation methods described in appendices 17 and 18, do not require the presentation of the properties of audible sounds that can be perceived from audible sounds. The configurations described in appendices 15 to 18 reduce the enormous amount of computation required for deriving the impulse response and for convolution using the impulse response. [Explanation of Symbols]
[0126] P1...Sound source P2... Receiving point 10…Audible Simulation System 20... Control Unit 21...Settings section 22…Space generation section 23...Sound generation section 24…Sound quality presentation section 25…Original sound presentation section 26…History Management Department 30...Storage section 31... Original audio data 32…Spatial data 32C...Cell 32R…Indoor space 33...Object data 34…Response data 35…Convolutional data 36...Converted data 50…Terminal device 51...projection device 52…Sound output device
Claims
1. An audible simulation system that reproduces sound in an indoor space in a virtual space, Each of the sound source and sound receiving point in the indoor space is a target, and the setting unit changes the acoustic parameters of the indoor space, including the arrangement of the targets and the arrangement of objects installed in the indoor space. A space generation unit that modifies the virtual space to be displayed on the display device using the arrangement of the target dynamically changed by the setting unit, A sound generation unit estimates the audible sound reaching the sound receiving point from the original sound of the sound source using the acoustic parameters dynamically changed by the setting unit, The system includes a sound quality presentation unit that presents the properties of the audible sound that can be perceived from the audible sound estimated by the sound generation unit, A sound-making simulation system characterized by the following features.
2. The properties of the audible sound are at least one selected from the group consisting of clarity and spatial impression evaluation index values. The audible simulation system according to claim 1.
3. The properties of the audible sound are at least one selected from the direction of arrival, the propagation path, the volume in the left and right ears, the volume absorbed by the object, and the volume blocked by the object. The audible simulation system according to claim 1.
4. The presentation of the properties of the audible sound is performed using at least one of the display by the display device and the sound output by the sound generation unit. The audible simulation system according to claim 1.
5. The sound generation unit, The setting unit reads the impulse response associated with the acoustic parameters when the virtual space is divided into multiple cells and the sound source and the sound receiving point are placed in separate cells, and uses this to perform sound convolution for generating the audible sound. The audible simulation system according to claim 1.
6. The acoustic parameters include the arrangement of the multiple objects installed in the indoor space, The sound generation unit reads the impulse response associated with the acoustic parameters dynamically changed by the setting unit from the storage unit, which stores the impulse response between the sound source and the sound receiving point in association with the acoustic parameters when the sound source, the sound receiving point, and each of the objects are arranged in separate cells, and uses it for sound convolution to generate the audible sound. The audible simulation system according to claim 5.
7. The storage unit stores the impulse response between the four directions at the sound source and the four directions at the sound receiving point. The sound generation unit uses a first transformation matrix that converts the components in each direction of the sound source into components having cardioid directivity in that direction so as to maintain the directivity of the sound source in each direction, and as a convolution of the sound, it convolves the components in each direction obtained by applying the first transformation matrix to the original sound with the impulse response in that direction. The audible simulation system according to claim 6.
8. The sound generation unit, A second transformation matrix is used to transform the components in each direction at the receiving point into components having cardioid directivity in that direction, so as to maintain the directivity at the receiving point in each direction. The second transformation matrix is then applied to the components in each direction into which the impulse response is convolved, and the head-of-direction transfer function for that direction is further convolved into the components in each direction. The audible simulation system according to claim 7.
9. The virtual space is divided into multiple cells, The sound generation unit, A storage unit is used to store a convolution component for each acoustic parameter when the sound source, the receiving point, and the object are arranged in separate cells, wherein the convolution component is the result obtained by convolving the impulse response between the sound source and the receiving point in the acoustic parameter with the sound of the sound source, and the setting unit reads out the convolution component associated with the acoustic parameter that has been dynamically changed, and estimates the audible sound reaching the receiving point using the convolution component. The audible simulation system according to claim 1.
10. The system further includes a sound presentation unit that allows for the selection and presentation of multiple mutually different original sounds, The sound generation unit estimates the audible sound using the original sound selected by the original sound presentation unit. The audible simulation system according to any one of claims 1 to 9.
11. The sound generation unit, The sound output unit outputs the audible sound and a reference sound, receives an adjustment amount for the audible sound relative to the volume of the reference sound, and outputs the audible sound again at the volume adjusted based on the adjustment amount. The audible simulation system according to claim 10.
12. The system further includes a history management unit that stores over time the virtual space displayed by the space generation unit and the audible sound estimated by the sound generation unit in the virtual space. The audible simulation system according to any one of claims 1 to 9.
13. The history management unit receives comments regarding the virtual space displayed by the space generation unit and the audible sound estimated by the sound generation unit in the virtual space, and stores the comments in association with the virtual space and the audible sound. The audible simulation system according to claim 12.
14. A server connected to multiple terminal devices, The system comprises the setting unit, the space generation unit, the sound generation unit, and the sound quality presentation unit, The virtual space is displayed on the display device of each terminal device. The sound output section of each terminal device is made to output the audible sound. The audible simulation system according to any one of claims 1 to 9.
15. A method for reproducing sound in an indoor space within a virtual environment, Each of the sound source and the sound receiving point in the aforementioned indoor space is a target, and the acoustic parameters of the indoor space, including the arrangement of the targets and the arrangement of objects installed in the indoor space, are to be dynamically changed. This includes modifying the virtual space displayed on the display device using the dynamically changed arrangement of the object, estimating the audible sound reaching the receiving point from the original sound of the sound source using the dynamically changed acoustic parameters, and presenting the properties of the audible sound that can be perceived from the estimated audible sound. A method for simulating audibility, characterized by the following features.
Citation Information
Patent Citations
Image display type simulation service providing system and image display type simulation service providing method
JP2017146762A
Information processing device
JP2024072072A