Information processing device, information processing program, and information processing system
The integration of a machine learning model to process both physical and perceptual quantities in reverberation processing addresses the challenge of unrealistic sound in virtual environments, enabling more immersive and accurate acoustic experiences.
Patent Information
- Application Number
- PCT/JP2025/017206
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-22
- Filing Date
- 2025-05-12
- Publication Date
- 2025-11-27
AI Technical Summary
Conventional reverberation processing methods struggle to incorporate human cognition and perception into acoustic space reproduction, whether through manual adjustments or physics-based information processing, leading to unrealistic sound experiences in virtual environments.
An information processing device that integrates a machine learning model to generate impulse responses by combining physical and perceptual quantities, allowing operators to input desired voice parameters and adjust them intuitively using a user interface.
Enables the creation of a more realistic acoustic space by aligning sound effects with human perception, improving immersion and accuracy in virtual environments.
Smart Images

Figure JP2025017206_27112025_PF_FP_ABST
Abstract
Description
Information processing device, information processing program, and information processing system
[0001] The present disclosure relates to an information processing device, an information processing program, and an information processing system.
[0002] As virtual space technologies such as the Metaverse and games continue to develop, there is a growing demand for realistic experiences of the sounds reproduced in virtual spaces. Reproducing sound in such virtual spaces requires the generation and processing of reverberation. To realize next-generation virtual space audio experiences, it is particularly desirable to incorporate human cognition and perception into reverberation processing. The temporal and frequency information that represents reverberation characteristics is called the impulse response.
[0003] Conventional reverberation processing methods involve manual adjustments by sound engineers relying on their own intuition, or methods that recreate reverberation through physics-based information processing. One proposed technology for acoustic design in virtual spaces is to use machine learning models to generate acoustic signals at the sound receiving point using structural data such as the structure of the virtual space, the coordinates of objects placed in the virtual space, and acoustic impedance.
[0004] International Publication No. 2023 / 182024
[0005] However, it is difficult to automate reverberation processing based on methods that rely on the operator's intuition. Furthermore, methods that reproduce reverberation using information processing that is faithful to physics have difficulty incorporating human cognition and perception. Therefore, it is difficult to reproduce a realistic acoustic space with either method.
[0006] Furthermore, technology that uses a machine learning model to generate acoustic signals using inputs such as the structure of a virtual space, object coordinates, and structural data does not take into account human cognition and perception separately from physical quantities, making it difficult to appropriately incorporate human cognition and perception into reverberation processing. In other words, even with this technology, it is difficult to reproduce a realistic acoustic space.
[0007] The present disclosure has been made in consideration of the above-described circumstances, and provides an information processing device, an information processing program, and an information processing system that can reproduce a realistic acoustic space.
[0008] The information processing device disclosed herein includes the following units: the receiving unit receives input of information related to a user's desired voice parameter and a physical quantity; the processing unit inputs the user's desired voice parameter corresponding to the information related to the user's desired voice parameter and the physical quantity to a trained machine learning model, and outputs a voice signal in which one or more voice parameters correspond to the input user's desired voice parameter and the voice parameters other than the input user's desired voice parameter are adjusted.
[0009] FIG. 1 is a block diagram of a reverberation processing device according to a first embodiment. FIG. 2 is a diagram illustrating the relationship between the direction of arrival of a sound image, the interaural level difference, and the interaural time difference in the case of a binaural signal. FIG. 3 is a diagram illustrating an example of a user interface. FIG. 4 is a flowchart of impulse response generation processing by the reverberation processing device according to the first embodiment. FIG. 5 is a diagram illustrating processing related to input of perception quantity related information. FIG. 6 is a diagram illustrating an example of adjusting the size of a sound image in a game engine. FIG. 7 is a block diagram of another example of a reverberation processing device. FIG. 8 is a conceptual diagram of learning processing according to a second embodiment. FIG. 9 is a flowchart of reverberation processing by the reverberation processing device according to the second embodiment. FIG. 10 is a conceptual diagram of learning processing according to a third embodiment. FIG. 11 is a flowchart of reverberation processing by the reverberation processing device according to the third embodiment. FIG. 12 is a conceptual diagram of learning processing according to a fourth embodiment. FIG. 13 is a flowchart of reverberation processing by the reverberation processing device according to the fourth embodiment. FIG. 14 is a hardware configuration diagram illustrating an example of a computer that realizes an arithmetic device of the reverberation processing device, which is an information processing device according to each of the embodiments or modified examples.
[0010] Preferred embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. In this specification and the drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted. The description will be given in the following order.
[0011] 1. First Embodiment 1.1. Overview of the Reverberation Processor According to the First Embodiment 1.2. Details of the Reverberation Processor 1.3. Impulse Response Generation Processing 1.4. Effects 1.5. Modifications of the First Embodiment 2. Second Embodiment 2.1. Reverberation Processor According to the Second Embodiment 2.2. Reverberation Processing According to the Second Embodiment 2.3. Effects 3. Third Embodiment 3.1. Reverberation Processor According to the Third Embodiment 3.2. Reverberation Processing According to the Third Embodiment 3.3. Effects 4. Fourth Embodiment 4.1. Reverberation Processor According to the Fourth Embodiment 4.2. Reverberation Processing According to the Fourth Embodiment 4.3. Effects 5. Other Modifications 5.1. First Modification 5.2. Second Modification 5.3. Third Modification 6. Hardware Configuration
[0012] <1. First embodiment> <1.1. Overview of reverberation processing device according to the first embodiment> Conventionally, reverberation processing methods have mainly been based on heuristic creation by sound engineers or on reverberation reproduction using physically accurate information processing. Techniques for reverberation reproduction using physically accurate information processing include classical reverberation generation methods such as the ray tracing method, the finite domain difference method, the finite element method, and methods based on room acoustics theory. In these classical reverberation generation methods, impulse responses are often calculated by acoustic simulation using a physical model of the desired space.
[0013] However, heuristic reverberation processing relies on the sound engineer's individual judgment to adjust reverberation and other aspects. Impulse responses contain various parameters. Therefore, adjusting only one parameter often results in an unnatural sound. Changing even one reverberation characteristic requires adjusting multiple parameters. Furthermore, the expertise required for such adjustments is unique to each acoustic engineer. Therefore, it is difficult to automate heuristic reverberation processing as is. Furthermore, the relationship between human spatial perception of sound and the physical quantities used in simulation is difficult to clearly define. Therefore, reverberation reproduction methods based on information processing that faithfully follows physics have difficulty incorporating human cognition and perception into reverberation processing.
[0014] 1 is a block diagram of a reverberation processing device according to a first embodiment. The reverberation processing device 1 according to this embodiment receives input of physical quantities including a 3D model, sound source positions, and sound receiving point positions specified by an operator such as a sound engineer. The reverberation processing device 1 also receives input of perceptual quantities that directly represent information related to human cognition and recognition, separate from the physical quantities specified by the operator. The definitions of physical quantities and perceptual quantities will be explained below. The reverberation processing device 1 receives input of the acquired physical quantities and perceptual quantities, and obtains an impulse response output, which is an inference result, from a trained machine learning model 16.
[0015] An operator can obtain the impulse response output from the reverberation processing device 1, and if the impulse response can produce the desired sound effect, the operator can use the impulse response to reproduce the reverberation. For example, in the production of a game, the operator can assign the reverberation reproduced from the obtained impulse response to a specified scene in the game. Here, the physical quantities and perceptual quantities used in the following explanation will be explained.
[0016] Physical quantities are physical models and parameters required to generate impulse responses that represent reverberation when using a reverberation reproduction method that relies on information processing that is faithful to physics. Specifically, physical quantities include the following information:
[0017] For example, in the case of a sound reproduction method using sound simulation, the following information corresponds to physical quantities: Here, sound simulation includes wave sound simulation methods that are methods based on wave equations, such as the finite difference method or the finite element method, and geometric sound simulation methods that are methods based on spatial geometric shapes, such as the ray method, the back tracing method, or the virtual image method.
[0018] In this case, examples of physical quantities include a spatial geometric model including a 3D (dimensional) model of the target space, a 3D mesh, 3D voxels, and dimensions, volume, and surface area of the space and each object in the space. Boundary conditions, including the sound absorption coefficient, acoustic reflectance, acoustic impedance, acoustic admittance, and density of the wall material, are also physical quantities. Propagation materials, including sound speed, temperature, humidity, specific heat ratio, density, and bulk modulus, are also physical quantities. Here, the sound propagating material is usually air. A sound source model, including sound source coordinates, sound source directivity, sound source frequency characteristics, and the number of sound rays in the case of the sound ray method, is also a physical quantity. A sound receiving point model, including sound receiving point coordinates, sound receiving point directivity, and the size of the sound receiving sphere in the case of the sound ray method, is also a physical quantity.
[0019] In the case of a probabilistic method based on room acoustics theory, the physical quantities include the volume of the space, the surface area, the complexity of the space, the speed of sound, and the density of reflected sound.
[0020] In other cases, impulse responses or reverb filters generated using acoustic simulation techniques or stochastic techniques based on room acoustics theory may be used. In this case, the impulse responses or reverb filters generated by conventional reverberation processing techniques, such as acoustic simulation techniques, are input to the model. In this case, the input impulse responses, reverb filters, or filter coefficients of the reverb filters correspond to physical quantities.
[0021] <1.1.2. Perception quantity> Perception quantity is a parameter that can be calculated from the output impulse response based on the spatial acoustic perception characteristics of humans. This perception quantity is an example of an "audio parameter." Information related to spatial acoustic perception contained in the impulse response is divided into information related to the basic sound image and information related to reverberation.
[0022] The perceptual quantities related to the basic sound image include the interaural level difference (ILD), the interaural time difference (ITD), and the direction of arrival (DoA) of the sound image. The interaural level difference and the interaural time difference are perceptual quantities related to binaural signals reproduced by headphones, earphones, etc. The direction of arrival of the sound image is a perceptual quantity related to the sound of binaural signals and multi-channel signals.
[0023] The interaural level difference (ILD) is also called the interaural intensity difference (IID). The interaural level difference is a perceptual quantity related to the sound image direction localization, and the sound image localization in the high frequency band is a clue. The interaural level difference is expressed, for example, by the following equation (1). Here, X L (t) and X R (t) are the impulse response signals of the left and right channels of the binaural signal, respectively.
[0024]
[0025] The interaural level difference for the time interval T is expressed by the following equation (2).
[0026]
[0027] The interaural level difference depending on the frequency f is expressed by the following equation (3).
[0028]
[0029] The interaural time difference (ITD) is also called the interaural phase difference (IPD). The interaural time difference is a perceptual quantity related to the directional localization of a sound image, and the sound image localization in the low frequency band is a clue. The interaural time difference is expressed, for example, by the following equation (4). Here, τ R and τ L are the times at which the first pulse of the impulse response arrives in the left and right channels, respectively, i.e., the time of the first silent section of the signal.
[0030]
[0031] The direction of arrival of a sound image (DoA) is the direction of a localized sound image, in other words, the angle of the localized sound image. In the case of a multi-channel signal, the direction of arrival of a sound image can be approximately calculated from the signal strength and phase of the multi-channel signal.
[0032] 2 is a diagram showing the relationship between the sound image arrival direction, the interaural level difference, and the interaural time difference in the case of a binaural signal. The angle θ in FIG. 2 is the sound image arrival direction of the sound from the sound source 101. Also, the distance d L is the propagation distance from the sound source 101 to the left ear of the person P. Also, the distance d R is the propagation distance from the sound source 101 to the right ear of the person P. When the sound source 101 is located in front of the person P, the distance d L = distance d R Distance d L and distance d R The difference between the two and the diffraction effect of the person P's head causes an interaural level difference and an interaural time difference.
[0033] On the other hand, the perceptual quantities related to reverberation include reverberation time (RT), early decay time (EDT), clarity, and initial time delay gap (ITDG).
[0034] Reverberation time (RT) is the perception of how reverberation sounds, or in other words, the size of a space. 60 , R.T. 30 and R.T. 20 There are others. 60 , R.T. 30 and R.T. 20 are the times from the start of reverberation until the reverberation energy drops by 60 dB, 30 dB, and 20 dB, respectively. 60 In the calculation of RT 60 =RT 30 x2 or RT 60 =RT 20 An approximation such as x3 may be used.
[0035] Early decay time (EDT) is a perceptual quantity that is close to reverberation time. 10 × 6, where RT 10 is the time from the onset of reverberation until the reverberation energy drops by 10 dB.
[0036] Clarity is the perceived ease of listening to speech. 50と The intelligibility index is generally known. 50 is the ratio of the direct sound and early reflected sound energy up to 50 ms to the late reverberation energy after 50 ms. 50 is expressed by the following equation (5).
[0037]
[0038] Also, C 80 There is also an index written as: 80 is the ratio of the direct sound and early reflected sound energy up to 80 ms to the late reverberation energy after 80 ms. 80 is expressed by the following equation (6).
[0039]
[0040] There is also another index denoted as D, which is the ratio of the energy of the direct sound and early reflected sound up to 50 ms to the energy of the entire impulse response. D is expressed by the following equation (7).
[0041]
[0042] The time delay of first reflections (ITDG) is a measure of the perceived spaciousness of a space. The time delay of first reflections is the time difference between the direct sound and the first reflection.
[0043] There are also perceptual quantities related to both the sound image and the reverberation, including the direct-to-reverberation ratio (DRR) and the interaural cross-correlation coefficient (IACC). The interaural cross-correlation coefficient is a perceptual quantity related to binaural signals.
[0044] The direct-to-reverberant ratio (DRR) is a measure of the perceived spaciousness of a space and the distance of a sound image. The direct-to-reverberant ratio is the energy ratio between direct sound and reverberant sound.
[0045] The interaural cross-correlation coefficient (IACC) is a perceptual quantity related to the apparent source width (ASW) and listener envelope (LEV). The interaural cross-correlation between direct sound and early reflected sound is proportional to the width of the apparent sound image. The interaural cross-correlation coefficient is related to the perception of the size of the auditory sound image. Furthermore, the interaural cross-correlation coefficient of late reverberant sound is related to the listener envelope. The interaural cross-correlation coefficient is expressed by the following equation (8). Here, IACCt 1 t 2 is the time t 1 From time t 2 is the interaural cross-correlation coefficient in the time interval from 1 t 2 (τ) is the x L (t) and x R (t) is the cross-correlation function with (t).
[0046]
[0047] There are the following differences between physical quantities and perceptual quantities. First, perceptual quantities are information that is directly linked to human hearing. On the other hand, physical quantities have a complex relationship with human spatial cognition. For example, the greater the time delay of the first reflection, which is a perceptual quantity, the larger the perceived space. In contrast, the dimensions of a room, which are physical quantities, are affected by factors such as the sound absorption coefficient and the position of the sound source, and are therefore not directly related to the perceived size of the space.
[0048] Second, although a perceptual quantity is uniquely determined when an impulse response is determined, the impulse response is not uniquely determined. When generating an impulse response for a certain input set, the impulse response output for the same physical quantity value is identical. On the other hand, even if the impulse response output for the same perceptual quantity value is different, it is possible that the impulse response will be different.
[0049] In other words, while it is difficult to adjust the reverberation to match human perception and cognition when adjusting the physical quantity, it is relatively easy to adjust the reverberation to match human perception and cognition by adjusting the perceived quantity. Furthermore, since different outputs can be obtained even when the same perceived quantity is used, the operator can select the reverberation they want from among these different outputs, making it possible to obtain a more appropriate reverberation.
[0050] In the following description, reverberation processing methods that do not properly consider human cognition and perception are referred to as "non-perceptual reverberation generation methods." Non-perceptual reverberation generation methods include heuristic reverberation processing, reverberation reproduction methods that use information processing that is faithful to physics, and machine learning-based reverberation generation methods that do not consider information related to human cognition and perception separately from physical quantities. In the following description, sound engineers, creators, and other people who adjust the sound using the reverberation processing device 1 are collectively referred to as "workers."
[0051] <1.2. Details of the reverberation processing device> Next, the reverberation processing device 1 will be described in detail with reference to Fig. 1. As shown in Fig. 1, the reverberation processing device 1 includes a user interface providing unit 11, an output unit 14, a processing unit 15, a machine learning model 16, and a receiving unit 20.
[0052] The user interface providing unit 11 has in advance a user interface format used for inputting perceptual quantities. The user interface providing unit 11 receives input of physical quantities from the physical quantity acquiring unit 12. Then, the user interface providing unit 11 generates a user interface for inputting each perceptual quantity using a predetermined format based on the acquired physical quantities, etc. Then, the user interface providing unit 11 transmits the generated user interface to the audio adjustment terminal 2 to display it on the screen, thereby providing the user interface to the operator. In this way, the user interface providing unit 11 allows the operator to input perceptual quantities using the user interface.
[0053] Fig. 3 is a diagram showing an example of a user interface. The user interface 110 shown in Fig. 3 is used for inputting the interaural level difference. For example, the user interface providing unit 11 has the format shown in Fig. 3 in which an adjustment slider 112 is arranged on a slide bar 111 as the format of the user interface for inputting the interaural level difference. The slider 112 can be moved to any position on the slide bar 111. Furthermore, a setting value window 114 displays the absolute value of the interaural level difference according to the position of the slider 112 on the slide bar 111.
[0054] The user interface providing unit 11 calculates the interaural level difference before adjustment as a reference value using a non-perceptual reverberation generation method for the acquired physical quantity. Next, the user interface providing unit 11 generates a user interface 110 for inputting the interaural level difference by placing reference value information 113 indicating the reference value on a slide bar 111 in the format shown in Fig. 3. The user interface providing unit 11 then transmits the generated user interface 110 for inputting the interaural level difference to the acoustic adjustment terminal 2.
[0055] The operator inputs various perceptual quantities according to his / her own recognition and perception using the user interface displayed on the screen of the audio adjustment terminal 2. The perceptual quantities according to the operator's recognition and perception input by the operator are examples of "user-desired audio parameters."
[0056] For example, when inputting the interaural level difference, the operator uses a mouse or the like to move a slider 112 on the user interface 110 shown in Fig. 3 to specify the desired interaural level difference. The operator can confirm the set interaural level difference on the display of a setting value window 114. The operator can also use reference value information 113 as a reference for adjusting the interaural level difference.
[0057] In this way, the user interface providing unit 11 generates and provides a user interface for inputting a user's desired voice parameter value through a predetermined operation. More specifically, the user interface providing unit 11 generates and provides a user interface 110 that has a slide bar 111 indicating continuous values of the predetermined voice parameter and a slider 112 indicating a specific value on the slide bar 111, and for inputting the user's desired voice parameter value by moving the slider 112 to a position on the slide bar 111 that indicates the user's desired voice parameter value. The user interface providing unit 11 also displays the value of the predetermined voice parameter based on an impulse response obtained from a physical quantity on the slide bar 111 as reference value information 113.
[0058] Although the above description has been given with reference to an example in which the user interface providing unit 11 calculates the reference value using a non-perceptual reverberation generation method for the physical quantity, the present invention is not limited to this. For example, the user interface providing unit 11 may input the acquired physical quantity into the machine learning model 16 to acquire the reference value.
[0059] Returning to FIG. 1 , the explanation will be continued. For example, among the perceived quantities, there is information whose adjustable range can be calculated from a physical quantity or an impulse response. For such perceived quantities, the user interface providing unit 11 may display an adjustment range in the user interface. Therefore, the user interface providing unit 11 may display an adjustable range of the perceived quantity in the user interface. Furthermore, the user interface providing unit 11 may limit the setting range of the perceived quantity to the displayed adjustable range.
[0060] For example, the user interface providing unit 11 calculates an adjustable range of the interaural level difference from the physical quantity. Then, the user interface providing unit 11 can display an adjustable range 122 indicating the calculated adjustable range on a slide bar 121, as in the user interface 120 shown in FIG. 3 . In this case, the user interface providing unit 11 may configure the user interface 120 to restrict the slider 123 from being moved beyond the adjustable range 122. This allows the operator to adjust the interaural level difference by moving the slider 123 within the adjustable range 122, thereby avoiding adjustment of a perceptual quantity that contradicts the physical quantity. In this way, the user interface providing unit 11 displays the adjustable range 122 of the value of a predetermined audio parameter on the slide bar 121.
[0061] Furthermore, the user interface providing unit 11 may generate a user interface that allows a user to input a relative value, such as a ratio of increase or decrease from a reference value, instead of an absolute value of the perceived amount. For example, the user interface providing unit 11 may generate a user interface 130 that has a slide bar 131 that allows a user to set a relative value within a range of ±20% with the calculated reference value as the median, like the user interface 130 shown in FIG.
[0062] Continuing the explanation, returning to Fig. 1, the receiving unit 20 receives the physical quantities and perceptual quantities input to the reverberation processing device 1. The receiving unit 20 includes a physical quantity acquiring unit 12 and a perceptual quantity acquiring unit 13.
[0063] The physical quantity acquisition unit 12 acquires physical quantities from setting information 3 of the virtual space in the game engine, etc. For example, the physical quantity acquisition unit 12 can acquire physical quantities such as the volume and surface area of a 3D model and space from 3D asset data in the game engine. In addition, the physical quantity acquisition unit 12 can acquire physical quantities such as sound absorption coefficient and acoustic impedance from information on the material and type of walls of the space for which an impulse response is to be generated.
[0064] 1 is an example of a source of physical quantities, and the physical quantity acquisition unit 12 may acquire the physical quantities from other sources. For example, the physical quantity acquisition unit 12 may acquire the physical quantities through input by an operator from the sound adjustment terminal 2. The physical quantity acquisition unit 12 then outputs the acquired physical quantities to the user interface providing unit 11 and the processing unit 15.
[0065] The perceptual quantity acquisition unit 13 can acquire the perceptual quantity 22 from an input from the sound adjustment terminal 2 by an operator using a user interface. For example, the perceptual quantity acquisition unit 13 acquires from the sound adjustment terminal 2 a plurality of types of perceptual quantities input using user interfaces for various perceptual quantities, including the user interface 110 for inputting the interaural level difference shown in FIG. 3 . The perceptual quantity acquisition unit 13 then outputs the acquired perceptual quantity 22 to the processing unit 15.
[0066] In this way, the receiving unit 20 receives input of the values of the audio parameters desired by the user input using the user interface. The receiving unit 20 also receives input of any one or a combination of the following audio parameters: interaural level difference, interaural time difference, sound image arrival direction, reverberation time, early decay time, intelligibility, time delay of the first reflection, direct-to-direct ratio, and interaural cross-correlation coefficient.
[0067] In this embodiment, the perception quantity acquisition unit 13 acquires the perception quantity input by the operator using the user interface, but the method of acquiring the perception quantity is not limited to this. For example, when acquiring the impulse response of a game scene, the operator may specify the localization position of a sound image in a game image generated by rendering the three-dimensional space set in the game into a two-dimensional space, and the perception quantity acquisition unit 13 may acquire the perception quantity based on the specification.
[0068] Alternatively, if the localization position of a sound image is determined within the game engine, the perception quantity acquisition unit 13 may acquire that position and use it as perceptual information. For example, if an object exists within a closed space within the game, it is conceivable that sound reverberates within the space, resulting in a location different from the object being the physically correct localization position of the sound image. In this case, although the location is physically correct, a discrepancy occurs between the localization position of the sound image and the position of the object that is the source of the sound, as perceived by humans. Therefore, the perception quantity acquisition unit 13 acquires the position of the object from the game engine and uses that information as the perceived quantity, thereby obtaining an impulse response that matches the display of the game.
[0069] Furthermore, the perception quantity acquisition unit 13 may use image recognition technology to identify an object to which a sound image is to be localized from an image to which an impulse response is assigned, and may acquire the position of the identified object to use as the perception quantity. For example, in the case of an image of ventriloquism, the perception quantity acquisition unit 13 may use image recognition technology to recognize a ventriloquist's doll as an object to which a sound image is to be localized, and may use the position of the ventriloquist's doll as the perception quantity for the voice of the ventriloquist's doll.
[0070] The machine learning model 16 receives input of physical quantities and perceptual quantities, infers and outputs an impulse response corresponding to the input physical quantities and perceptual quantities. The machine learning model 16 can use, for example, a dataset including pre-prepared physical quantities and perceptual quantities and training data of the corresponding impulse responses for learning. For example, the training data of the impulse responses for predetermined physical quantities and perceptual quantities can specify an impulse response desired by an operator for the predetermined physical quantities and perceptual quantities. The machine learning model 16 then performs learning by adjusting parameters so as to reduce the error between the impulse response output when the physical quantities and perceptual quantities included in the dataset are input and the training data of the impulse responses.
[0071] The trained machine learning model 16 receives input of physical quantities and perceptual quantities as inference target data from the processing unit 15. Then, the machine learning model 16 performs inference by forward propagating the input physical quantities and perceptual quantities through the network, and obtains impulse responses corresponding to the input physical quantities and perceptual quantities. Then, the machine learning model 16 outputs the impulse responses as the inference results.
[0072] The processing unit 15 receives an input of a physical quantity from the physical quantity acquisition unit 12. The processing unit 15 also receives an input of a perceptual quantity from the perceptual quantity acquisition unit 13. Next, the processing unit 15 inputs the acquired physical quantity and perceptual quantity to the trained machine learning model 16. Then, the processing unit 15 acquires an impulse response output as an inference result from the machine learning model 16. Thereafter, the processing unit 15 outputs the impulse response corresponding to the physical quantity and the perceptual quantity to the output unit 14.
[0073] Here, when impulse responses are acquired at several points for the same location and sound image, prioritizing the perceived amount may result in a discrepancy between the desired position, the position determined by the impulse response, and the position of the obtained impulse response. Therefore, when impulse responses are acquired at multiple points, if the localization position of the sound image based on the already acquired impulse response is close to the desired position, the processing unit 15 may reflect the perceived amount when the impulse response was acquired, and allow the operator to make corrections later.
[0074] Furthermore, although the perceived position and the position of the sound image change from moment to moment in an analog manner, the impulse responses for each physical quantity and perceived quantity are instantaneous data, and the impulse responses may be digital data. Therefore, when digital data is acquired as the impulse response, the processing unit 15 may generate a continuous impulse response by interpolating between the individual impulse responses.
[0075] The output unit 14 acquires impulse responses corresponding to the input physical quantities and perceptual quantities from the processing unit 15. Then, the output unit 14 transmits the impulse responses corresponding to the input physical quantities and perceptual quantities to the sound adjustment terminal 2. The operator can check the reverberation corresponding to the input physical quantities and perceptual quantities by having the sound adjustment terminal 2 reproduce the reverberation corresponding to the impulse response transmitted from the reverberation processing device 1.
[0076] 4 is a flowchart of the impulse response generation process performed by the reverberation processing device 1 according to the first embodiment. Next, the flow of the impulse response generation process performed by the reverberation processing device 1 according to the present embodiment will be described with reference to FIG.
[0077] The physical quantity acquisition unit 12 acquires physical quantities from, for example, setting information 3 of a virtual space in a game engine (step S1).
[0078] The operator inputs a desired perceived amount using the user interface provided by the user interface providing unit 11 at the sound adjustment terminal 2 (step S2).
[0079] The perception amount acquiring unit 13 acquires the perception amount input by the operator from the sound adjustment terminal 2 (step S3).
[0080] The processing unit 15 receives input of physical quantities from the physical quantity acquisition unit 12. The processing unit 15 also receives input of perceptual quantities from the perceptual quantity acquisition unit 13. Next, the processing unit 15 inputs the acquired physical quantities and perceptual quantities to the trained machine learning model 16. The machine learning model 16 performs inference by forward propagating the physical quantities and perceptual quantities through the network (step S4). Thereafter, the processing unit 15 outputs an impulse response, which is the inference result output from the machine learning model 16, to the output unit 14.
[0081] The output unit 14 acquires the impulse response from the processing unit 15. Then, the output unit 14 transmits the impulse response corresponding to the input physical quantity and perceptual quantity to the sound adjustment terminal 2 (step S5).
[0082] <1.4. Effects> As described above, the reverberation processing device 1 according to this embodiment inputs, in addition to physical quantities, perceptual quantities, which are information based on human spatial acoustic perception characteristics and are parameters that can be calculated from the output impulse response, to the machine learning model 16. The reverberation processing device 1 then acquires and outputs an impulse response inferred by the machine learning model 16 based on the input perceptual quantities and physical quantities.
[0083] Conventional reverberation generation and editing processes use the physical quantities of space as input, which can deviate from human intuition and require a great deal of processing experience to achieve the desired sound. In contrast, the reverberation processing device 1 of this embodiment uses perceptual quantities that directly represent the reverberation felt by humans, allowing input of information based on human perception, enabling more intuitive reverberation generation and editing. In other words, operators can generate and edit reverberation more intuitively.
[0084] Furthermore, conventional highly immersive reverberation processing often involves acoustic simulation using a 3D spatial model. However, when the video playback device is a device that outputs two-dimensional images, such as a flat display, a mismatch between the visual and / or auditory senses may occur. This is because the video is a two-dimensional image, while the sound is three-dimensional spatial audio. In contrast, the reverberation processing device 1 according to this embodiment can calculate a perceptual amount from the rendered two-dimensional video signal and adjust the impulse response using this perceptual amount, thereby eliminating the mismatch between the visual and auditory senses and delivering a more immersive reverberation sound to the user. This allows for an impulse response that more closely matches the perception of the user, recreating a more realistic acoustic space.
[0085] Furthermore, conventional non-perceptual reverberation generation methods may not be able to express certain physical phenomena. For example, in probabilistic methods based on room acoustics theory, the sound source model is a point sound source with no volume, making it impossible to change the size of the sound image. Furthermore, in machine learning-based methods, if a bias exists in the dataset, it is not possible to generate an appropriate impulse response for inputs outside the parameter range included in the dataset. In contrast, the reverberation processing device 1 of this embodiment performs learning so that the perceptual amount becomes a target value. Therefore, it is possible to generate reverberation corresponding to perceptual amounts that are not included in the correct impulse response generated by conventional reverberation sound processing. Therefore, the reverberation processing device 1 of this embodiment can improve the reverberation expression ability compared to conventional reverberation sound processing.
[0086] The reverberation processing device 1 according to this embodiment also provides a user interface for inputting perceived quantities. By using the user interface, the operator can intuitively and easily input the desired perceived quantities, thereby obtaining impulse responses that better suit the operator's own recognition and perception. This makes it possible to reproduce a more realistic acoustic space.
[0087] <1.5. Modification of First Embodiment> Next, a modification of the first embodiment will be described. The user interface providing unit 11 according to this modification uses, as an input perceptual quantity, perceptual quantity-related information relating to human perception and cognition that can be mapped to the perceptual quantity described above.
[0088] The perception quantity acquisition unit 13 may acquire the size of the sound image, the distance of the sound image, the height of the sound image, etc. as the perception quantity related information about the sound image. The size of the sound image is information proportional to the interaural cross-correlation coefficient among the perception quantities. The distance of the sound image is information related to the direct-indirect ratio among the perception quantities. The height of the sound image is information partially included in the direction from which the sound image arrives among the perception quantities.
[0089] Furthermore, the perception quantity acquisition unit 13 may acquire, as the perception quantity related information regarding reverberation, the size of the acoustic space, the reverberation length, the intensity of the direct sound, the ease of speech intelligibility, etc. The size of the acoustic space is information related to the early decay time, the time delay of the first reflection, and the direct-to-indirect ratio, which are among the perception quantities. The reverberation length is information proportional to the reverberation time and early decay time, which are among the perception quantities. The intensity of the direct sound is information proportional to the direct-to-indirect ratio, which are among the perception quantities. The ease of speech intelligibility is information proportional to the clarity, which is among the perception quantities.
[0090] The perceptual quantity acquisition unit 13 maps the acquired perceptual quantity related information to the perceptual quantity, and acquires the perceptual quantity corresponding to the acquired perceptual quantity related information. Thereafter, the perceptual quantity acquisition unit 13 outputs the acquired perceptual quantity to the processing unit 15.
[0091] In this case, the user interface providing unit 11 generates a user interface for inputting the perceptual quantity related information. The user interface providing unit 11 then transmits the generated user interface to the sound adjustment terminal 2 and provides it to the operator. The operator inputs the perceptual quantity related information corresponding to the desired reverberation using the user interface for inputting the perceptual quantity related information provided by the sound adjustment terminal 2.
[0092] 5 is a diagram showing a process related to input of perception amount related information. For example, the user interface providing unit 11 generates a user interface 140 used to input the magnitude of a sound image, which is part of the perception amount related information. In this case, the user interface providing unit 11 generates a user interface 140 that allows the user to specify whether to decrease or increase the magnitude of the sound image. The operator moves a slider on the user interface 140 to input the desired sound image magnitude.
[0093] The perception quantity acquisition unit 13 acquires information on the size of a sound image specified using the user interface 140. Next, as shown in process 141, the perception quantity acquisition unit 13 calculates the width of an apparent sound image from the size of the acquired sound image, and calculates an interaural cross-correlation coefficient (IACC) corresponding to the size of the input sound image by mapping the calculated width of the apparent sound image to an interaural cross-correlation coefficient (IACC). In this case, the perception quantity acquisition unit 13 converts the size of the sound image, represented by a magnitude value indicated by a slide bar 142, into an interaural cross-correlation coefficient ranging from 0 to 1, indicated by a slide bar 143.
[0094] In addition, the user interface providing unit 11 may generate a user interface used to input the size of a sound image, with a different method of specifying values. FIG. 6 is a diagram showing an example of adjusting the size of a sound image in a game engine. For example, as shown in FIG. 6, the user interface providing unit 11 may generate a user interface 150 that allows a user to input a value for the size of a sound image by specifying the size of an image corresponding to the image of the size of the sound image. The user interface providing unit 11 generates a user interface 150 that displays a scene from game production and places spheres 151 to 153, which are images indicating the size of a sound image, on objects that serve as sound sources. For example, as shown by arrow 154, an operator uses a mouse or the like to change the size of sphere 151 to match the desired size of the sound image.
[0095] In this way, the user interface providing unit 11 displays the size of the sound image as the size of the image, i.e., the size of the spheres 151 to 153, and displays the image size in an adjustable manner, thereby generating and providing a user interface 150 for inputting the user's desired size of the sound image by changing the image size to the user's desired size of the sound image.
[0096] The perception quantity acquisition unit 13 receives an input of the size of the sound image corresponding to the size of the sphere 151. In this case, the perception quantity acquisition unit 13 receives an input of a magnitude value as the sound image size, as indicated by a slide bar 156 in process 155 in Fig. 6. Then, as shown in process 155, the perception quantity acquisition unit 13 converts the sound image size indicated by the slide bar 156 into an interaural cross-correlation coefficient indicated by a slide bar 157.
[0097] Furthermore, the input of the size of the sound image may be automatic. For example, the perception amount acquisition unit 13 may acquire the size of the object that is the source of the sound in the game as the size of the sound image from the setting information 3 of the virtual space in the game engine.
[0098] In this way, the receiving unit 20 receives input of information and physical quantities related to the user's desired voice parameters. Here, the perceived quantity and perceived quantity-related information that the receiving unit 20 can acquire can be said to be information that affects how the human ear hears. That is, the receiving unit 20 receives input of, as information related to the voice parameters, perceived quantity-related information that affects how the human ear hears, and that is any of the voice parameters or any of the sound image size, sound image distance, sound image height, acoustic space width, reverberation length, direct sound intensity, or voice audibility corresponding to the voice parameters. Furthermore, the receiving unit 20 calculates the user's desired voice parameters corresponding to the perceived quantity-related information based on the perceived quantity-related information input by the user.
[0099] In addition, the processing unit 15 inputs the user's desired voice parameters and physical quantities corresponding to information regarding the user's desired voice parameters into the trained machine learning model 16, and outputs a voice signal in which one or more voice parameters correspond to the input user's desired voice parameters and voice parameters other than the input user's desired voice parameters are adjusted.
[0100] In some cases, the information related to perceived volume may be easier for the operator to recognize than the perceived volume. In such cases, inputting the information related to perceived volume makes it easier for the operator to input information to obtain the desired reverberation. Therefore, the reverberation desired by the operator can be properly reproduced, and a more realistic acoustic space can be reproduced.
[0101] <2. Second embodiment> <2.1. Reverberation processing device according to the second embodiment> Next, a reverberation processing device 1 according to the second embodiment will be described. FIG. 7 is a block diagram showing another example of a reverberation processing device. The reverberation processing device 1 according to this embodiment executes learning of a machine learning model 16. The reverberation processing device 1 further includes a dataset storage unit 17 and a learning execution unit 18 in addition to the units of the first embodiment. In the following explanation, description of the operation of the units similar to those of the first embodiment will be omitted.
[0102] The dataset storage unit 17 stores a dataset 170 used for training the machine learning model 16. The dataset 170 includes, for example, a plurality of pairs of physical quantities obtained from acoustic simulation or actual measurement and impulse responses generated using a non-perceptual reverberation generation method based on the predetermined physical quantities. Hereinafter, the predetermined physical quantities and the impulse responses based on the predetermined physical quantities will be referred to as "paired data."
[0103] 8 is a conceptual diagram of the learning process according to the second embodiment. The learning execution unit 18 acquires data pairs from the dataset storage unit 17. The learning execution unit 18 also acquires perceptual quantities. Here, the learning execution unit 18 may randomly generate perceptual quantities, or may acquire perceptual quantities input by an operator or the like using the sound adjustment terminal 2. Hereinafter, the perceptual quantities acquired by the learning execution unit 18 as learning data will be referred to as "target perceptual quantities."
[0104] Next, the learning execution unit 18 inputs the physical quantity and the target perceptual quantity included in the data pair to the machine learning model 16, and acquires an impulse response output from the machine learning model 16 (step S11). Here, the impulse response output from the machine learning model 16 is referred to as an "output impulse response."
[0105] Next, the learning execution unit 18 compares the teacher data of the impulse response, which is the impulse response included in the data pair, with the output impulse response, and calculates the physical error L physic (Step S12). The learning execution unit 18 also calculates the perceptual quantity from the output impulse response (Step S13). The learning execution unit 18 then compares the target perceptual quantity with the calculated perceptual quantity to obtain a perceptual error L perceptual is calculated (step S14).
[0106] Thereafter, the learning execution unit 18 calculates a loss function, where L=λ physi L physic +λperceptualL perceptual It is. physic is a loss term for evaluating the physical error between the training data of the impulse response that reflects a predetermined physical quantity and the output impulse response. perceptual is a loss term that evaluates the perceptual error between the target perceptual quantity and the perceptual quantity calculated from the output impulse response. physi and λperceptual are respectively L physic and L perceptual By training the machine learning model 16 so that the loss function L is minimized, an impulse response that satisfies the target values of both the physical quantity and the perceptual quantity is output from the machine learning model 16 during inference. perceptual is set as a differentiable function of the input of the learning model. That is, the learning execution unit 18 performs error backpropagation in the machine learning model 16 and updates the weights in the machine learning model 16 (step S15).
[0107] In this way, the machine learning model 16 in this embodiment is trained to minimize the error between the output impulse response, which has the values of a specified physical quantity and a specified voice parameter as input, and the teacher data of the impulse response obtained from the specified physical quantity, and the error between the value of the voice parameter calculated from the output impulse response and the value of the specified voice parameter.
[0108] 9 is a flowchart of the reverberation processing by the reverberation processing device 1 according to the second embodiment. Next, the flow of the reverberation processing by the reverberation processing device 1 according to the second embodiment will be described again with reference to FIG.
[0109] The learning execution unit 18 acquires data pairs from the dataset 170 stored in the dataset storage unit 17 (step S101).
[0110] Furthermore, the learning execution unit 18 randomly generates a target perception amount (step S102).
[0111] Next, the learning execution unit 18 inputs the physical quantity and the perceptual quantity included in the data pair to the machine learning model 16 (step S103).
[0112] The machine learning model 16 performs inference by forward propagating the input physical quantities and target perceptual quantities through the network (step S104).
[0113] The learning execution unit 18 acquires the output impulse response output from the machine learning model 16 (step S105).
[0114] Next, the learning execution unit 18 compares the teacher data of the impulse response included in the data pair with the output impulse response to obtain the physical error L physic is calculated (step S106).
[0115] Furthermore, the learning execution unit 18 calculates the perceived amount from the output impulse response (step S107).
[0116] Next, the learning execution unit 18 compares the target perceptual quantity with the calculated perceptual quantity to obtain a perceptual error L perceptual is calculated (step S108).
[0117] Next, the learning execution unit 18 calculates the loss function L=λ physi L physic +λperceptualL perceptual is calculated (step S109).
[0118] Next, the learning execution unit 18 causes the machine learning model 16 to perform error backpropagation (step S110).
[0119] Then, the learning execution unit 18 updates the weights in the machine learning model 16 (step S111).
[0120] Thereafter, the learning execution unit 18 determines whether learning has ended based on whether a predetermined ending condition has been met (step S112). The ending condition is given, for example, by an upper limit on the number of epochs or a loss threshold. If learning has not ended (step S112: No), the learning execution unit 18 returns to step S101. On the other hand, if it is determined that learning has ended (step S112: Yes), the learning execution unit 18 ends learning of the machine learning model 16.
[0121] 2.3. Effects As described above, the reverberation processing device 1 according to this embodiment causes the machine learning model 16 to learn using physical quantities, impulse responses and perceptual quantities obtained based on the physical quantities by a non-perceptual reverberation generation technique or the like. This makes it possible to generate a trained machine learning model 16 that has undergone appropriate learning, and to obtain accurate inference results from the trained machine learning model 16. Therefore, it is possible to appropriately reproduce the reverberation desired by the operator, and to reproduce a more realistic acoustic space.
[0122] 3. Third Embodiment 3.1. Reverberation Processor According to Third Embodiment Next, a reverberation processor 1 according to a third embodiment will be described. The reverberation processor 1 according to this embodiment is also represented by the block diagram in Fig. 7. The reverberation processor 1 according to this embodiment uses, for learning, data pairs in which the perceptual amount calculated from training data of impulse responses matches the perceptual amount to be input to the machine learning model 16.
[0123] 10 is a conceptual diagram of the learning process according to the third embodiment. The learning execution unit 18 acquires data pairs from the data set storage unit 17. The learning execution unit 18 also acquires a target perception amount.
[0124] Next, the learning unit 18 acquires training data of impulse responses included in the data pairs in the machine learning model 16. Next, the learning unit 18 calculates a perceptual quantity from the acquired training data of impulse responses (step S21). Next, the learning unit 18 compares the target perceptual quantity with the calculated perceptual quantity (step S22). Then, the learning unit 18 determines whether the target perceptual quantity matches the calculated perceptual quantity (step S23).
[0125] If the target perceptual quantity does not match the calculated perceptual quantity (step S23: No), the learning execution unit 18 acquires a data pair different from the current data pair from the data set storage unit 17 and selects another impulse response (step S24). Thereafter, the learning execution unit 18 repeats steps S21 to S23.
[0126] On the other hand, if the target perceptual quantity matches the calculated perceptual quantity (step S23: Yes), the learning execution unit 18 inputs the physical quantity included in the selected data pair and the perceptual quantity other than the target perceptual quantity to the machine learning model 16. Then, the learning execution unit 18 acquires an output impulse response from the machine learning model 16 (step S25).
[0127] Next, the learning execution unit 18 compares the teacher data of the impulse response with the output impulse response to obtain the physical error L physic (Step S26). Then, the learning execution unit 18 calculates the loss function. Here, the physical quantity θ physic and the training data of the impulse response corresponding to the target perceptual quantity θperceptual is D(θ physic , θperceptual). In this embodiment, the learning execution unit 18 uses θ physic Using the loss function L = λ physi L physic The learning execution unit 18 calculates the physical quantity θ physic and a perceptual quantity θperceptual different from the target perceptual quantity θperceptual * For D(θ physic ,θperceptual * ) is used as the correct answer and the machine learning model 16 is made to learn.physic ,θperceptual * ) does not exist in the dataset 170, the machine learning model 16 relies on the latent space interpolation capability of machine learning to find the input physical quantity θ physic and perceptual quantity θ perceptual * The machine learning model 16 outputs an impulse response interpolated from corresponding outputs of nearby values of the combination of (a) and (b). This is a general characteristic of the machine learning model 16. Next, the learning execution unit 18 performs error backpropagation in the machine learning model 16 and updates the weights in the machine learning model 16 (step S27).
[0128] In this way, the learning execution unit 18 in this embodiment is trained to minimize the error between the output impulse response, which has as input the values of a specified physical quantity and a voice parameter other than the specified physical quantity, and the impulse response teacher data, when the value of the voice parameter calculated based on the impulse response teacher data obtained from the specified physical quantity matches the value of the specified voice parameter.
[0129] 11 is a flowchart of the reverberation processing by the reverberation processing device 1 according to the third embodiment. Next, the flow of the reverberation processing by the reverberation processing device 1 according to the third embodiment will be described again with reference to FIG.
[0130] The learning execution unit 18 randomly generates a target perception amount (step S201).
[0131] Furthermore, the learning execution unit 18 acquires data pairs from the dataset 170 stored in the dataset storage unit 17 (step S202).
[0132] Next, the learning execution unit 18 calculates the amount of perception from the teacher data of the impulse response included in the data pair (step S203).
[0133] Next, the learning execution unit 18 determines whether the target perceptual quantity and the calculated perceptual quantity match (step S204). If the target perceptual quantity and the calculated perceptual quantity do not match (step S204: No), the learning execution unit 18 returns to step S202.
[0134] On the other hand, if the target perceptual quantity matches the calculated perceptual quantity (step S204: Yes), the learning execution unit 18 inputs the physical quantity and perceptual quantities other than the target perceptual quantity included in the data pair into the machine learning model 16 (step S205).
[0135] The machine learning model 16 performs inference by forward propagating the input physical quantities and perceptual quantities through the network (step S206).
[0136] The learning execution unit 18 acquires the output impulse response output from the machine learning model 16 (step S207).
[0137] Next, the learning execution unit 18 compares the teacher data of the impulse response included in the data pair with the output impulse response to obtain the physical error L physic is calculated (step S208).
[0138] Next, the learning execution unit 18 calculates the loss function L=λ physi L physic is calculated (step S209).
[0139] Next, the learning execution unit 18 causes the machine learning model 16 to perform error backpropagation (step S210).
[0140] Then, the learning execution unit 18 updates the weights in the machine learning model 16 (step S211).
[0141] Thereafter, the learning execution unit 18 determines whether the learning has ended based on whether a predetermined end condition has been satisfied (step S212). If the learning has not ended (step S212: No), the learning execution unit 18 returns to step S201. On the other hand, if the learning has ended (step S212: Yes), the learning execution unit 18 ends the learning of the machine learning model 16.
[0142] 3.3. Effects As described above, the reverberation processing device 1 according to this embodiment causes the machine learning model 16 to learn using training data of impulse responses in which the calculated perceptual quantity matches the target perceptual quantity. This method also makes it possible to generate a trained machine learning model 16 that has undergone appropriate learning, and to obtain accurate inference results from the trained machine learning model 16. Therefore, it is possible to appropriately reproduce the reverberation desired by the operator, and to reproduce a more realistic acoustic space.
[0143] <4. Fourth embodiment> <4.1. Reverberation processor according to the fourth embodiment> Next, a reverberation processor 1 according to the fourth embodiment will be described. The reverberation processor 1 according to this embodiment is also represented by the block diagram in Fig. 7. The reverberation processor 1 according to this embodiment edits training data of impulse responses in accordance with the amount of perception, and compares the impulse responses to calculate a perceptual error.
[0144] 12 is a conceptual diagram of the learning process according to the fourth embodiment. The learning execution unit 18 acquires data pairs from the data set storage unit 17. The learning execution unit 18 also acquires a target perception amount.
[0145] Next, the learning execution unit 18 inputs the physical quantity and the target perceptual quantity included in the data pair to the machine learning model 16, and obtains an output impulse response from the machine learning model 16 (step S31). Next, the learning execution unit 18 compares the output impulse response with the teacher data of the impulse response, which is the impulse response included in the data pair, to obtain a physical error L physic is calculated (step S32).
[0146] The learning execution unit 18 also edits the teacher data of the impulse response so that the calculated perceptual quantity becomes the target perceptual quantity (step S33). physic and the target perceptual quantity θ perceptual, and the physical quantity θ physic and the output impulse response to the target perceptual quantity θperceptual is defined as D(θ physic , θperceptual). For example, the learning execution unit 18 physic and perceptual quantity θ perceptual *is input to the machine learning model 16. Then, the learning execution unit 18 calculates D(θ physic , θperceptual) is subjected to processing H to generate pseudo-teaching data D that satisfies the input perceptual amount. ~ (θperceptual * ) = H(D(θ physic , θperceptual), θperceptual * This process H is a non-perceptual reverberation generation method, and it is not possible to consider both physical quantities and perceptual quantities at the same time. Therefore, in process H, physical quantities are not considered. Therefore, the pseudo-teaching data D ~ (θperceptual * ) the input physical quantity θ physic Next, the learning execution unit 18 compares the output impulse response with the edited impulse response to determine the perceptual error L modified is calculated (step S34).
[0147] Then, the learning execution unit 18 calculates the loss function. In this case, the learning execution unit 18 calculates the loss function as L=λ physic L physic +λ modified L modified Here, L modified is the impulse response D of the pseudo-teaching data that does not reflect physical quantities but reflects perceptual quantities. ~ (θ physic ) and a loss term evaluating the error of the output impulse response from the machine learning model 16, and λ modified is its weight. physic is the training data D(θ) of the original impulse response that does not reflect the perceptual quantity but reflects the physical quantity. physic , θperceptual). modified and L physic In order to minimize both D(θ physic , θperceptual) and D ~ (θ physic ) and outputs the intermediate value between them. Then, the learning execution unit 18 causes the machine learning model 16 to perform error backpropagation, and updates the weights in the machine learning model 16 (step S35).
[0148] In this way, the machine learning model 16 according to this embodiment is trained to minimize the error between the output impulse response, which has the values of predetermined physical quantities and predetermined voice parameters as input, and the teacher data of the impulse response obtained from the predetermined physical quantities, and the error between the teacher data of the impulse response edited based on the values of the predetermined voice parameters and the output impulse response.
[0149] 13 is a flowchart of the reverberation processing by the reverberation processing device 1 according to the fourth embodiment. Next, the flow of the reverberation processing by the reverberation processing device 1 according to the fourth embodiment will be described again with reference to FIG.
[0150] The learning execution unit 18 acquires data pairs from the dataset 170 stored in the dataset storage unit 17 (step S301).
[0151] The learning execution unit 18 also randomly generates a target perception amount (step S302).
[0152] Next, the learning execution unit 18 inputs the physical quantity and the perceptual quantity included in the data pair to the machine learning model 16 (step S303).
[0153] The machine learning model 16 performs inference by forward propagating the input physical quantities and target perceptual quantities through the network (step S304).
[0154] The learning execution unit 18 acquires the output impulse response output from the machine learning model 16 (step S305).
[0155] Next, the learning execution unit 18 compares the teacher data of the impulse response included in the data pair with the output impulse response to obtain the physical error L physic is calculated (step S306).
[0156] Furthermore, the learning execution unit 18 edits the output impulse response based on the target perception amount (step S307).
[0157] Next, the learning execution unit 18 compares the edited impulse response with the output impulse response to obtain the perceptual error L modifiedis calculated (step S308).
[0158] Next, the learning execution unit 18 calculates the loss function L=λ physic L physic +λ modified L modified is calculated (step S309).
[0159] Next, the learning execution unit 18 causes the machine learning model 16 to perform error backpropagation (step S310).
[0160] Then, the learning execution unit 18 updates the weights in the machine learning model 16 (step S311).
[0161] Thereafter, the learning execution unit 18 determines whether learning has ended based on whether a predetermined ending condition has been met (step S312). The ending condition is given, for example, by an upper limit on the number of epochs or a loss threshold. If learning has not ended (step S312: No), the learning execution unit 18 returns to step S301. On the other hand, if it is determined that learning has ended (step S312: Yes), the learning execution unit 18 ends learning of the machine learning model 16.
[0162] 4.3. Effects As described above, the reverberation processing device 1 according to this embodiment edits the output impulse response in accordance with the target perceptual quantity, calculates the perceptual error using the edited impulse response and the original output impulse response, and causes the machine learning model 16 to perform training. This method also makes it possible to generate a trained machine learning model 16 that has been appropriately trained, and to obtain accurate inference results from the trained machine learning model 16. Therefore, it is possible to appropriately reproduce the reverberation desired by the operator, and to reproduce a more realistic acoustic space.
[0163] The reverberation processing device 1 may also use some of the learning processes described in the second to fourth embodiments. If there is a contradiction between the target physical quantity and the target perceptual quantity and no impulse response exists that satisfies both, the machine learning model 16 simultaneously minimizes both errors, resulting in errors in the physical quantity and perceptual quantity of the output impulse response.
[0164] 5. Other Modifications 5.1. First Modification In the above embodiments, the reverberation processor 1 outputs an impulse response for reproducing reverberation, but other sounds may be added to the reverberation. For example, a reverberant sound source can be obtained by convolving an impulse response with a dry sound source. Therefore, the reverberation processor 1 may add the dry sound source to its input in addition to the physical quantity and perceptual quantity.
[0165] For example, the physical quantity acquisition unit 12 acquires a physical quantity and a dry sound source. In this case, the machine learning model 16 receives the physical quantity, the perceptual quantity, and the dry sound source as input data and receives the reverberant sound source as output data. The processing unit 15 inputs the physical quantity, the perceptual quantity, and the dry sound source to the machine learning model 16 and acquires the reverberant sound source from the machine learning model 16 as a corresponding output. The output unit 14 transmits the reverberant sound source to the audio adjustment terminal 2 and provides it to an operator.
[0166] 5.2. Second Modification Although in the above embodiments the operator inputs the perceived amount by specifying information about the perceived amount, this is not limiting. For example, the reverberation processing device 1 may input information in the virtual space other than the perceived amount by mapping it to the perceived amount.
[0167] For example, the user interface providing unit 11 generates a user interface in which information in the virtual space other than the perceived amount, such as the size of the space, can be adjusted with a slider, and provides the user with the generated user interface. In this case, the user interface providing unit 11 uses sensory expressions such as "small to large" or "narrow to wide" as inputs to the sliders in the user interface.
[0168] The perceptual quantity acquisition unit 13 acquires information in the virtual space input using the user interface. Here, the information in the virtual space acquired by the perceptual quantity acquisition unit 13 is expressed as x∈(x min , x max ) x min is the minimum value of the virtual space information that can be input using the user interface. max is the maximum value of the information in the virtual space that can be input using the user interface. In this case, the perception amount acquisition unit 13 converts the input information into a perception amount y∈(y min , ymax ) can be mapped to y min is the minimum value of the perceptual quantity to be mapped. max is the maximum value of the perceptual quantity to be mapped.
[0169] The perception amount acquisition unit 13 can calculate the perception amount from the acquired virtual space information by linear mapping using the following equation (9).
[0170]
[0171] Furthermore, the perception amount acquisition unit 13 can calculate the perception amount from the acquired virtual space information by nonlinear mapping using the following equation (10).
[0172]
[0173] Furthermore, the perception amount acquisition unit 13 can calculate the perception amount from the acquired virtual space information by selective mapping using the following equation (11).
[0174]
[0175] Alternatively, the perception amount acquiring unit 13 may calculate the perception amount using non-independent mapping in which the mapped perception amount depends on parameters other than the acquired information in the virtual space. In this case, information other than the information in the virtual space on which the mapped perception amount depends may also be input using a user interface provided by the user interface providing unit 11.
[0176] After mapping the information in the virtual space to a perceptual quantity as described above, the perceptual quantity acquisition unit 13 outputs the mapped perceptual quantity to the processing unit 15, causing the machine learning model 16 to perform inference and output an impulse response.
[0177] 5.3. Third Modification The above description has been given on the assumption that the processing of the inference phase for generating reverberation using the trained machine learning model 16 is performed by a single information processing device such as a client-end computer. However, in addition to this configuration, the information processing device may perform the reverberation generation processing by having a client-end computer, such as a mobile device, which has difficulty in performing the inference processing due to its performance, and a server-end computer, which performs the inference processing of the trained machine learning model 16, work together.
[0178] For example, the client-end computer may be an audio adjustment terminal 2 having a user interface providing unit 11, a physical quantity acquiring unit 12, and a perceptual quantity acquiring unit 13. In this case, the server-end computer has a processing unit 15, a machine learning model 16, and an output unit 14.
[0179] The client-end computer executes the following processes. The user interface providing unit 11 generates a user interface and displays the user interface on a display screen (not shown). The physical quantity acquiring unit 12 acquires a physical quantity 21 and transmits it to the server-end computer. The perceptual quantity acquiring unit 13 acquires a perceptual quantity input using the user interface and an input device (not shown) of the client-end computer and transmits it to the server-end computer. The client-end computer then receives an impulse response corresponding to the transmitted physical quantity and perceptual quantity from the server-end computer. The client-end computer then uses the received impulse response to reflect the results in an application such as a game. This client-end computer is an example of a "first information processing device."
[0180] The server-end computer executes the following process. The processing unit 15 receives physical quantities and perceptual quantities from the client-end computer. The processing unit 15 then inputs the acquired physical quantities and perceptual quantities into the machine learning model 16 to obtain an impulse response output from the machine learning model 16. The output unit 14 transmits the impulse response to the client-end computer. This server-end computer is an example of a "second information processing device."
[0181] 6. Hardware Configuration FIG. 14 is a hardware configuration diagram showing an example of a computer that realizes a calculation device of a reverberation processing device that is an information processing device according to each of the embodiments and modifications.
[0182] The computer 1000 includes a CPU 1100, a RAM 1200, a ROM (Read Only Memory) 1300, a HDD (Hard Disk Drive) 1400, a communication interface 1500, and an input / output interface 1600. The components of the computer 1000 are connected to each other via a bus 1050.
[0183] The CPU 1100 operates and controls each component based on programs stored in the ROM 1300 or the HDD 1400. For example, the CPU 1100 loads the programs stored in the ROM 1300 or the HDD 1400 into the RAM 1200 and executes processing corresponding to the various programs.
[0184] The ROM 1300 stores boot programs such as a Basic Input Output System (BIOS) that is executed by the CPU 1100 when the computer 1000 is started, and programs that depend on the hardware of the computer 1000 .
[0185] HDD 1400 is a computer-readable recording medium that non-temporarily records programs executed by CPU 1100 and data used by such programs. Specifically, HDD 1400 is a recording medium that records an application program according to the present disclosure, which is an example of program data 1450.
[0186] The communication interface 1500 is an interface for connecting the computer 1000 to an external network 1550 (e.g., the Internet). For example, the CPU 1100 receives data from other devices and transmits data generated by the CPU 1100 to other devices via the communication interface 1500.
[0187] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the CPU 1100 receives data from an input device such as a keyboard or a mouse via the input / output interface 1600. The CPU 1100 also transmits data to an output device such as a display or a printer via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs and the like recorded on a predetermined recording medium. Examples of media include optical recording media such as DVDs (Digital Versatile Discs) and PDs (Phase Change Rewritable Disks), magneto-optical recording media such as MOs (Magneto-Optical Disks), tape media, magnetic recording media, and semiconductor memories.
[0188] Although the CPU 1100 reads and executes the program data 1450 from the HDD 1400, as another example, the CPU 1100 may obtain these programs from other devices via an external network 1550.
[0189] Although the preferred embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings, the technical scope of the present disclosure is not limited to such examples. It is clear that a person skilled in the art of the present disclosure can conceive of various modified or altered examples within the scope of the technical idea described in the claims, and it is understood that these also naturally fall within the technical scope of the present disclosure.
[0190] Furthermore, the effects described herein are merely descriptive or exemplary and are not limiting. In other words, the technology according to the present disclosure may achieve other effects that will be apparent to those skilled in the art from the description of this specification, in addition to or in place of the above-described effects.
[0191] The present technology can also be configured as follows.
[0192] (1) An information processing device comprising: a receiving unit that receives input of information related to a user's desired voice parameter and a physical quantity; and a processing unit that inputs the user's desired voice parameter corresponding to the information related to the user's desired voice parameter and the physical quantity into a trained machine learning model, and outputs a voice signal in which one or more voice parameters, which correspond to the input user's desired voice parameter and the voice parameter other than the input user's desired voice parameter, are adjusted. (2) The information processing device according to (1), wherein the receiving unit receives input of perception quantity related information, as information related to the voice parameter, which is information that affects how a person hears things, and is any one of the voice parameters or any one of the size of a sound image, distance of a sound image, height of a sound image, width of an acoustic space, length of reverberation, intensity of a direct sound, or ease of hearing a voice corresponding to the voice parameter. (3) The information processing device according to (2), wherein the receiving unit receives as input, as the speech parameters, one or a combination of an interaural level difference, an interaural time difference, a direction of arrival of a sound image, a reverberation time, an early decay time, an intelligibility, a time delay of a first reflection, a direct-to-direction ratio, or an interaural cross-correlation coefficient. (4) The information processing device according to (2), wherein the receiving unit calculates, based on the perception quantity related information input by a user, speech parameters desired by the user corresponding to the perception quantity related information. (5) The information processing device according to any one of (1) to (4), wherein the machine learning model is trained to minimize an error between an output impulse response having input values of a predetermined physical quantity and a predetermined speech parameter and training data of an impulse response obtained from the predetermined physical quantity, and an error between a speech parameter value calculated from the output impulse response and a value of the predetermined speech parameter. (6) The information processing device according to any one of (1) to (4), wherein the machine learning model is trained to minimize an error between an output impulse response that uses values of voice parameters other than the predetermined physical quantity and the predetermined voice parameter as input and the teacher data of the impulse response when the value of the voice parameter calculated based on teacher data of the impulse response obtained from the predetermined physical quantity matches the value of the predetermined voice parameter.(7) The information processing device according to any one of (1) to (4), wherein the machine learning model is trained to minimize an error between an output impulse response having inputs of a predetermined physical quantity and a predetermined voice parameter value and teacher data of an impulse response obtained from the predetermined physical quantity, and an error between the output impulse response and teacher data of the impulse response edited based on the predetermined voice parameter value. (8) The information processing device according to any one of (1) to (7), further comprising a user interface providing unit that generates and provides a user interface for inputting a value of a voice parameter desired by the user through a predetermined operation, wherein the receiving unit receives an input of the value of the voice parameter desired by the user input using the user interface. (9) The information processing device according to (8), wherein the user interface providing unit has a slide bar that indicates consecutive values of the predetermined voice parameter and a slider that indicates a specific value on the slide bar, and generates and provides the user interface for inputting the value of the voice parameter desired by the user by moving the slider to a position on the slide bar that indicates the value of the voice parameter desired by the user. (10) The information processing device according to (9), wherein the user interface providing unit displays an adjustable range of the value of the predetermined voice parameter on the slide bar. (11) The information processing device according to (9) or (10), wherein the user interface providing unit displays the value of the predetermined voice parameter based on an impulse response obtained from the physical quantity on the slide bar as reference value information. (12) The information processing device according to (8), wherein the user interface providing unit generates and provides a user interface that indicates the size of a sound image by the size of an image, displays the size of the image in an adjustable manner, and changes the size of the image to the size of the sound image desired by the user to input the size of the sound image desired by the user.(13) An information processing system having a first information processing device and a second information processing device, wherein the first information processing device receives input of information and physical quantities related to a user's desired voice parameters and transmits the input voice parameters and the physical quantities to the second information processing device, and the second information processing device receives the voice parameters and the physical quantities from the first information processing device, and inputs the user's desired voice parameters and the physical quantities corresponding to information related to the user's desired voice parameters into a trained machine learning model, and outputs a voice signal in which the voice parameters other than the input user's desired voice parameters, which correspond to the input user's desired voice parameters, among one or more voice parameters, are adjusted. (14) An information processing program that causes a computer to execute a process of receiving input of information and physical quantities related to a user's desired voice parameters and inputs the user's desired voice parameters and the physical quantities corresponding to information related to the user's desired voice parameters into a trained machine learning model, and outputting a voice signal in which the voice parameters other than the input user's desired voice parameters, which correspond to the input user's desired voice parameters, among one or more voice parameters, are adjusted.
[0193] REFERENCE SIGNS LIST 1 Reverberation processing device 2 Acoustic adjustment terminal 3 Virtual space setting information 11 User interface providing unit 12 Physical quantity acquisition unit 13 Perceptual quantity acquisition unit 14 Output unit 15 Processing unit 16 Machine learning model 17 Data set storage unit 20 Receiving unit 170 Data set
Claims
1. An information processing device having: a receiving unit that receives input of information related to a user's desired voice parameters and physical quantities; and a processing unit that inputs the user's desired voice parameters corresponding to the information related to the user's desired voice parameters and the physical quantities into a trained machine learning model, and outputs a voice signal in which one or more voice parameters correspond to the input user's desired voice parameters and the voice parameters other than the input user's desired voice parameters are adjusted.
2. The information processing device of claim 1, wherein the receiving unit receives, as information about the voice parameters, information that affects how a sound is heard by humans, and which receives input of perception-related information, which is any of the voice parameters or any of the size of the sound image, distance of the sound image, height of the sound image, width of the acoustic space, length of reverberation, intensity of direct sound, or ease of hearing the voice corresponding to the voice parameters.
3. The information processing device according to claim 2, wherein the receiving unit receives input of one or a combination of the interaural level difference, interaural time difference, sound image arrival direction, reverberation time, early decay time, clarity, time delay of first reflection, direct-to-direct ratio, or interaural cross-correlation coefficient as the audio parameters.
4. The information processing device according to claim 2, wherein the receiving unit calculates the user's desired voice parameters corresponding to the perception amount related information based on the perception amount related information input by the user.
5. The information processing device described in claim 1, wherein the machine learning model is trained to minimize the error between an output impulse response having input values of a predetermined physical quantity and a predetermined voice parameter and training data of an impulse response obtained from the predetermined physical quantity, and the error between the value of a voice parameter calculated from the output impulse response and the value of the predetermined voice parameter.
6. The information processing device of claim 1, wherein the machine learning model is trained to minimize the error between an output impulse response that uses values of voice parameters other than the specified physical quantity and the specified voice parameter as input and the teacher data of the impulse response when the value of the voice parameter calculated based on teacher data of the impulse response obtained from the specified physical quantity matches the value of the specified voice parameter.
7. The information processing device described in claim 1, wherein the machine learning model is trained to minimize the error between an output impulse response that uses a predetermined physical quantity and a predetermined voice parameter value as input and training data of an impulse response obtained from the predetermined physical quantity, and the error between the training data of the impulse response edited based on the value of the predetermined voice parameter and the output impulse response.
8. An information processing device as described in claim 1, further comprising a user interface providing unit that generates and provides a user interface for inputting the value of the voice parameter desired by the user through a predetermined operation, and wherein the receiving unit receives input of the value of the voice parameter desired by the user input using the user interface.
9. The information processing device of claim 8, wherein the user interface providing unit has a slide bar indicating successive values of a predetermined voice parameter and a slider indicating a specific value on the slide bar, and generates and provides the user interface for inputting the value of the voice parameter desired by the user by moving the slider to a position on the slide bar indicating the value of the voice parameter desired by the user.
10. The information processing device according to claim 9, wherein the user interface providing unit displays an adjustable range of the value of the predetermined voice parameter on the slide bar.
11. The information processing device according to claim 9, wherein the user interface providing unit displays the value of the predetermined voice parameter based on the impulse response obtained from the physical quantity on the slide bar as reference value information.
12. An information processing device as described in claim 8, wherein the user interface providing unit generates and provides a user interface for inputting the user's desired size of the sound image by indicating the size of the sound image with the size of an image, displaying the image in an adjustable size, and changing the size of the image to the user's desired size of the sound image.
13. An information processing system having a first information processing device and a second information processing device, wherein the first information processing device receives input of information and physical quantities related to a user's desired voice parameters, and transmits the input voice parameters and physical quantities to the second information processing device, and the second information processing device receives the voice parameters and the physical quantities from the first information processing device, inputs the user's desired voice parameters and the physical quantities corresponding to the information related to the user's desired voice parameters into a trained machine learning model, and outputs a voice signal in which one or more voice parameters correspond to the input user's desired voice parameters and the voice parameters other than the input user's desired voice parameters are adjusted.
14. An information processing program that causes a computer to execute a process of receiving input of information and physical quantities related to a user's desired voice parameters, inputting the user's desired voice parameters corresponding to the information related to the user's desired voice parameters and the physical quantities into a trained machine learning model, and outputting a voice signal in which one or more voice parameters correspond to the input user's desired voice parameters and in which the voice parameters other than the input user's desired voice parameters have been adjusted.
Citation Information
Patent Citations
Voice broadcast control method and device, and air conditioner
CN111260864A
Audio playback device, audio feedback system and method
JP2006512819A
Information processing device, sound reproduction device, information processing system, information processing method, and virtual sound source generation device
JP2023132236A
Rendering scene-aware audio using neural network-based acoustic analysis
US20210136510A1
Signal processing device, signal processing method, and program
WO2018079850A1