Information processing device, information processing method, information processing system, and program

The system addresses sound quality issues in reverberation removal by using machine learning to adjust direct sound components, enabling high-quality recordings in non-studio environments and reducing processing time and costs.

JP7736061B2Active Publication Date: 2025-09-09SONY GROUP CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023517147
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-04-30
Filing Date
2022-03-22
Publication Date
2025-09-09
Estimated Expiration
2042-03-22

AI Technical Summary

Technical Problem

Existing dereverberation technologies do not adequately consider sound quality, requiring professional equipment and extensive processing time to remove reverberation from recordings made in environments lacking sufficient recording facilities, such as homes, and often result in suboptimal sound quality.

Method used

A system that uses machine learning to generate training data by convolving measurement signals from non-professional recording devices with reference signals to adjust the average level and delay of direct sound components, training a learning model to remove reverberation while maintaining sound quality, and applying these adjustments to recordings made in reverberant environments.

Benefits of technology

Enables high-quality recordings without professional equipment by removing reverberation effectively, allowing for optimal addition of reverberation characteristics, thus improving sound quality and reducing processing time and costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007736061000001
    Figure 0007736061000001
  • Figure 0007736061000002
    Figure 0007736061000002
  • Figure 0007736061000003
    Figure 0007736061000003
Patent Text Reader

Abstract

The present invention makes it possible to perform reverberation processing that takes sound quality into account, for example. This information processing device includes a learning processing unit which: generates training data by convoluting a measurement signal representing an acoustic characteristic, collected by means of a sound collecting unit identical to one used to collect an observed signal, with a reference signal having a sound quality and a reverberation characteristic that are different from the measurement signal; generates teacher data by adapting an average level and a delay value of a convoluted signal, generated by convoluting a direct sound component of the measurement signal with the reference signal, to the reference signal; and trains a learning model that performs reverberation processing of the observed signal collected by the sound collecting unit, using the training data and the teacher data as input data.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an information processing device, an information processing method, an information processing system, and a program. [Background technology]

[0002] Techniques for removing reverberation from speech are known. For example, Patent Documents 1 and 2 listed below disclose dereverberation techniques that use machine learning. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2018-146683 [Patent Document 2] WO2019-026973 publication Summary of the Invention [Problem to be solved by the invention]

[0004] However, the techniques disclosed in the above-mentioned Patent Documents 1 and 2 merely remove reverberation from the speech of a conference or the like to make it easier to hear, and do not take sound quality into consideration.

[0005] An object of the present disclosure is to propose an information processing device, an information processing method, an information processing system, and a program that are capable of performing reverberation processing taking sound quality into consideration. [Means for solving the problem]

[0006] The present disclosure provides, for example, The information processing device has a learning processing unit that generates training data by convolving a measurement signal representing acoustic characteristics collected by the same sound collection unit as that used to collect an observation signal with a reference signal having sound quality and reverberation characteristics different from those of the observation signal, generates teacher data by adapting the average level and delay value of the convolved signal generated by convolving a direct sound component of the measurement signal with the reference signal, and trains a learning model that performs reverberation processing of the observation signal collected by the sound collection unit using the training data and the teacher data as input data.

[0007] The present disclosure provides, for example, This information processing method performs a learning process in which a measurement signal representing acoustic characteristics picked up by the same sound pickup unit as that used to pick up an observation signal is convolved with a reference signal having sound quality and reverberation characteristics different from those of the observation signal to generate training data, a direct sound component of the measurement signal is convolved with the reference signal to generate convolved signal, and the average level and delay value of the convolved signal are adapted to the reference signal to generate training data, and a learning model that performs reverberation processing of the observation signal picked up by the sound pickup unit is trained using the training data and the training data as input data.

[0008] The present disclosure provides, for example, This program causes a computer to execute a learning process that generates training data by convolving a measurement signal, which represents acoustic characteristics and is collected by the same sound collection unit as that used to collect an observation signal, with a reference signal that has sound quality and reverberation characteristics different from those of the observation signal, generates teacher data by adapting the average level and delay value of the convolved signal, which is generated by convolving the direct sound component of the measurement signal with the reference signal, to the reference signal, and learns a learning model that performs reverberation processing of the observation signal collected by the sound collection unit using the training data and the teacher data as input data.

[0009] The present disclosure provides, for example, a dereverberation unit that receives the second sound data collected by the microphone and outputs third sound data obtained by removing the reverberation components from the second sound data, using a learning model that is machine-learned using first data generated by convolving the first sound data collected by the microphone with dry data that does not contain reverberation, and second data generated by convolving the direct sound component of the first sound data with the dry data and adapting the average level and delay value of the convolution signal to the dry data; a reverberation component feature extraction unit that extracts feature data representing a reverberation component of the second acoustic data by using the second acoustic data and the third acoustic data; a reverberation adding unit that inputs sound source data, adds reverberation characteristics based on the feature data, and outputs the result; The information processing device has the following.

[0010] The present disclosure provides, for example, a learning model machine-learned using first data generated by convolving first sound data collected by a microphone with dry data that does not contain reverberation, and second data generated by convolving a direct sound component of the first sound data with the dry data and adapting the average level and delay value of the convolved signal to the dry data, to input the second sound data collected by the microphone, and output third sound data in which the reverberation component has been removed from the second sound data; extracting feature data representing a reverberation component of the second acoustic data using the second acoustic data and the third acoustic data; Input the sound source data, add reverberation characteristics based on the feature data, and output the result This is an information processing method in which processing is carried out by a computer.

[0011] The present disclosure provides, for example, a learning model machine-learned using first data generated by convolving first sound data collected by a microphone with dry data that does not contain reverberation, and second data generated by convolving a direct sound component of the first sound data with the dry data and adapting the average level and delay value of the convolved signal to the dry data, to input the second sound data collected by the microphone, and output third sound data in which the reverberation component has been removed from the second sound data; extracting feature data representing a reverberation component of the second acoustic data using the second acoustic data and the third acoustic data; Input the sound source data, add reverberation characteristics based on the feature data, and output the result It is a program that causes a computer to execute a process.

[0012] The present disclosure provides, for example, A microphone and a learning model that is machine-learned using first data generated by convolving first sound data collected by the microphone with dry data that does not include reverberation, and second data generated by convolving a direct sound component of the first sound data with the dry data and adapting the average level and delay value of a convolved signal to the dry data; a sound source output unit that outputs sound source data; a memory that stores second acoustic data collected by the microphone and the sound source data; a dereverberation unit that uses the learning model to input the second acoustic data and output third acoustic data obtained by removing reverberation components from the second acoustic data; a reverberation component feature extraction unit that extracts feature data representing a reverberation component of the second acoustic data by using the second acoustic data and the third acoustic data; a reverberation adding unit that inputs the sound source data, adds reverberation characteristics based on the feature data, and outputs the result; It is an information processing system having the above. [Brief explanation of the drawings]

[0013] [Figure 1] FIG. 1 is a diagram illustrating an example of the configuration of an information processing system according to an embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of the configuration of an information processing device according to an embodiment. [Figure 3] FIG. 3 is a diagram illustrating an example of the configuration of the teacher data generating unit. [Figure 4] FIG. 4 is a diagram showing an example of the waveform of a general impulse response signal. [Figure 5] FIG. 5 is a diagram showing an example of the waveform of an impulse response signal when sound is collected in an anechoic chamber using a smartphone. [Figure 6] FIG. 6 is a diagram for explaining an example of extraction of a direct sound component of an impulse response signal collected in an anechoic room. [Figure 7] FIG. 7 is a diagram for explaining measurement of the delay value. [Figure 8] FIG. 8 is a diagram for explaining an example of extraction of a direct sound component of an impulse response signal collected for learning. [Figure 9] FIG. 9 is a diagram for explaining the addition of characteristics. [Figure 10] FIG. 10 is a flowchart showing the flow of the learning process performed by the information processing device. [Figure 11] FIG. 11 is a diagram illustrating an example of a hardware configuration of an information processing device. [Figure 12] FIG. 12 is a diagram showing an example of the configuration of an information processing system during recording according to a modified example. [Figure 13] FIG. 13 is a diagram showing an example of the configuration of an information processing device during editing according to a modified example. [Figure 14] FIG. 14 is a diagram showing an example of a waveform of an input signal. [Figure 15] FIG. 15 is a diagram showing an example waveform of a dereverberated signal. [Figure 16] FIG. 16 is a diagram showing an example of the waveform of a reverberation component. [Figure 17] FIG. 17 is a diagram showing another example of the configuration of an information processing device during editing according to a modified example. DETAILED DESCRIPTION OF THE INVENTION

[0014] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. Note that the embodiments described below are preferred specific examples of the present disclosure, and the contents of the present disclosure are not limited to these embodiments. The description will be given in the following order. <1. Background> <2. One embodiment> <3. Modifications>

[0015] <1. Background> First, the background of the present disclosure will be described. Content production is undergoing a transformation. Traditionally, content that requires recording of sound, such as music (e.g., commercial content), has been produced in a location with a sufficient recording environment (e.g., a studio). However, in recent years, there has been an increase in remote work and independent internet distribution due to infection control measures. Therefore, there is a demand for an environment that allows easy content production that prioritizes sound quality, even in a location that does not have a sufficient recording environment (e.g., at home).

[0016] For example, in music genres such as pop, orchestras can often be recorded at the homes of individual musicians, but vocals are generally recorded in studios. However, in the future, there is a demand for easy recording of vocals at home.

[0017] The problem here is how to eliminate reverberation when recording in a location that lacks a sufficient recording environment, such as a home. For example, in music production, it is common to first record without reverberation and then add reverberation later to complete the song. However, as mentioned above, conventional dereverberation technology only made the sound easier to hear, and there were no dereverberation functions that were clearly at a production-quality level. As a result, studio engineers had to spend time listening to the recording and then eliminating reverberation. Furthermore, even if reverberation could be eliminated, recording using non-professional microphones, such as those on smartphones, can still result in sound quality issues compared to recording using professional equipment in a studio.

[0018] Therefore, currently, audio sources recorded at home or elsewhere must be brought to a studio, where an engineer must use professional music production equipment to check the sound quality and remove reverberation while listening to the sound, requiring a huge amount of processing time. Furthermore, in some cases, the quality required for commercial content may not be achieved, and the recording may have to be repeated in the studio. For these reasons, there is a demand for reverberation processing technology that takes sound quality into consideration.

[0019] <2. One embodiment> [2-1. System configuration example] 1 is a diagram showing an example of the configuration of an information processing system (information processing system 1) according to an embodiment. The information processing system 1 is a system for music production. The information processing system 1 includes a user device 2, an information processing device 3, and a studio device 4.

[0020] The user device 2 is a device used by a user. The user device 2 is configured as a terminal device such as a smartphone, tablet, or personal computer. The user device 2 has a communication function that enables communication with the information processing device 3. The user device 2 has a sound collection unit 20 such as a microphone, and has a sound collection function that can collect external sounds. For example, if the user device 2 is a smartphone, the microphone of the smartphone can be used as the sound collection unit 20.

[0021] Specifically, the user device 2 transmits to the information processing device 3 an observation signal (in this example, a vocal signal obtained by collecting vocal sounds by a vocalist) collected by the sound collection unit 20 in a location that does not have a sufficient recording environment (for example, the user's home). The user device 2 has the sound collection unit 20, which allows the user device 2 to easily transmit the vocal signal to the information processing device 3. The sound collection unit 20 may be a separate unit that is connected to the user device 2, or may be included in a device other than the user device 2. The point is that as long as the vocal signal that is transmitted to the information processing device 3 is collected by the sound collection unit 20, it does not matter whether the user device 2 is connected to the sound collection unit 20.

[0022] The information processing device 3 is a device that performs reverberation processing of an input signal, and is provided, for example, as a cloud server on a network such as the Internet. Specifically, the information processing device 3 has a function of removing reverberation and recording system characteristics using machine learning. The recording system characteristics here refer to characteristics that are unintentionally added during recording, such as noise characteristics and unnecessary characteristics resulting from equipment such as microphones and speakers used for recording. The information processing device 3 has a learning processing unit 30 (details of which will be described later) that trains a learning model that removes the reverberation and recording system characteristics.

[0023] The information processing device 3 also has a communication function that enables communication with the user device 2 and the studio device 4. Specifically, the information processing device 3 receives a vocal signal transmitted from the user device 2, removes reverberation and recording system characteristics from the received vocal signal to generate a processed signal, and transmits the generated processed signal to the studio device 4. Note that this series of processes may be real-time processes.

[0024] The studio equipment 4 is a device having commercial content production functions such as mixing and mastering. The studio equipment 4 is configured, for example, with a personal computer, and is installed in a location with a sufficient recording environment, such as a studio. The studio equipment 4 has a communication function that enables communication with the information processing device 3. The studio equipment 4 receives processed signals transmitted from the information processing device 3 and performs processing such as mixing and mastering.

[0025] As shown in the figure, the processed signal processed by the information processing device 3 may be transmitted from the information processing device 3 to the user device 2. This allows for, for example, personal use in music production (when studio production is not required). The processed signal may be transmitted to both the user device 2 and the studio device 4. Alternatively, there may be multiple user devices 2, with different user devices 2 transmitting the vocal signal and receiving the processed signal. For example, a vocal signal may be transmitted from a smartphone, and the processed signal may be received by a personal computer at home.

[0026] As described above, the information processing device 3 does not have to be provided on a network, but may have the functions of the user device 2 and the studio device 4 and be usable (function) as at least one of the user device 2 and the studio device 4. In this way, the information processing device 3 may be provided on the operating side operated by the user or on the studio side. The connections between the devices including the sound collection unit 20 described above (for example, connection modes such as wired or wireless, communication methods, etc.) are not limited to specific ones.

[0027] [2-2. Example of configuration of information processing device] 2 is a diagram showing an example of the configuration of an information processing device 3 according to an embodiment. As described above, the information processing device 3 has a learning processing unit 30, which has a learning data generation unit 31, a teacher data generation unit 32, and a DNN learning unit 33. The following description will be given taking as an example a case where reverberation is removed from a vocal signal recorded by a musician at home using a smartphone.

[0028] The learning data generation unit 31 generates learning data (learning sound source) for learning by convolving the recorded IR data for learning with the dry data (dry sound source). The learning data generated by the learning data generation unit 31 is input to the DNN learning unit 33. The dry data and the recorded IR data for learning are prepared in advance so that they can be processed by the information processing device 3. These data are stored and prepared in advance, for example, in a storage unit (not shown) of the information processing device 3.

[0029] Dry data is a reference signal used during training. Dry data is high-quality, reverberant-free vocal audio (e.g., audio recorded with professional equipment in a recording studio) recorded separately from the training IR data. High-quality audio here specifically refers to audio with various characteristics (e.g., frequency characteristics, noise characteristics, etc.) sufficient for commercial content production, such as a sampling frequency of 48 kHz or higher (e.g., 96 kHz, 192 kHz, etc.). A large amount of dry data (e.g., approximately 500 songs) of various vocal voices (e.g., male, female, child, adult, Japanese, English, etc.) is prepared to support dereverberation of a wide variety of vocal signals.

[0030] The recorded IR data for learning is a measurement signal representing acoustic characteristics collected and measured using a smartphone at a musician's home. This smartphone is used to collect the vocal signal (observation signal) described above. The measurement signal is, for example, an impulse response signal, and can be recorded using a TSP (Time Stretched Pulse) signal. Note that the measurement signal used to collect the impulse response signal is not limited to a TSP signal, and other signals such as an M-sequence (Maximum Length Sequence) signal may also be used. The recorded IR data for learning includes a large amount (e.g., approximately 300 songs) of impulse responses collected in various locations (e.g., various locations in a workroom or living room) so that it can be used to dereverberate vocal signals collected in various environments.

[0031] The teacher data generation unit 32 generates teacher data (teacher sound source) for learning using the dry data. The teacher data generated by the teacher data generation unit 32 is input to the DNN learning unit 33. The DNN learning unit 33 uses the learning data and the teacher data as input data to learn and construct a learning model using a DNN (Deep Neural Network) that dereverberates vocal signals.

[0032] For example, by using dry data as training data, it is possible to cancel out the characteristics of smartphone recordings and achieve high-quality recordings. However, simply using dry data as training data and training the system will not learn dereverberation, and will only be able to learn to adjust the level. This is because the delay and level of the dry sound source (target sound source) inherent in the various training data do not match the training data. Since DNN training progresses in the direction of reducing the average error, it will focus on matching these and will not be able to learn the essential dereverberation.

[0033] Therefore, the teacher data generation unit 32 generates teacher data by adjusting the delay and level of the dry data using the IR data recorded for learning and the IR data recorded in an anechoic chamber. The IR data recorded in an anechoic chamber is a measurement signal (e.g., an impulse response signal) collected and measured in an anechoic chamber using a smartphone used to collect a vocal signal (observation signal). The IR data recorded in an anechoic chamber is, for example, stored and prepared in advance in a storage unit of the information processing device 3.

[0034] This anechoic chamber-recorded IR data and the above-mentioned learning-recorded IR data may be recorded using the same recording system (e.g., recording equipment such as a microphone) as that used to record the vocal signal (observation signal). It is preferable to use the same equipment as that used to actually record the vocal signal, but it is sufficient if the recording characteristics are the same, including within the tolerance range. For example, it is not limited to the same model, and a different model may be used. The recording system for the anechoic chamber-recorded IR data and the recording system for the learning-recorded IR data (e.g., equipment such as a microphone and speaker) should have the same characteristics.

[0035] 3 is a diagram showing an example of the configuration of the teacher data generation unit 32. The teacher data generation unit 32 includes a direct sound extraction unit 321, a delay value measurement unit 322, a direct sound convolution unit 323, an average level ratio calculation unit 324, and a characteristic addition unit 325.

[0036] The direct sound extraction unit 321 extracts direct sound components from the anechoic chamber-recorded IR data. The extraction of this direct sound component will now be described in detail. Figure 4 shows an example of the waveform of a typical impulse response signal, and Figure 5 shows an example of the waveform of an impulse response signal when sound is collected in an anechoic chamber using a smartphone. In the diagram, the horizontal axis represents time, and the vertical axis represents level.

[0037] As shown in Figure 4, a typical impulse response signal has a waveform that appears in the following order: direct sound (impulse), early reflections, and late reverberation. Early reflections are the direct sound reflected off the floor, walls, ceiling, etc., while late reverberation is the subsequent reverberation that decays as it reflects repeatedly. Because a typical impulse response signal has this kind of waveform, the characteristics of the direct sound component contained within it can be obtained by convolving the delay of the direct sound component's impulse (specifically, the delay from output to detection) and level with the dry data, and this can be used as training data for learning.

[0038] However, the impulse response signal picked up by a smartphone has a waveform like that shown in Figure 5, making it impossible to clearly separate the direct sound. The reason for this waveform is that to obtain a waveform like that shown in Figure 4, the equipment used for recording must also have flat characteristics across the processing band. When picking up sound with a smartphone, a smartphone microphone is essential, and speakers that can be carried around and used for recording at home have a narrower band and are not as flat as speakers installed in professional studios. As a result, the characteristics of the smartphone microphone and the speaker used during impulse measurement are added, resulting in a waveform like that shown in Figure 5. This problem is not limited to sound picking up with a smartphone.

[0039] Therefore, the direct sound extraction unit 321 estimates the direct sound component from the anechoic chamber-recorded IR data and extracts the estimated direct sound component from the anechoic chamber-recorded IR data. Specifically, the direct sound extraction unit 321 cuts the anechoic chamber-recorded IR data to the number of samples (N) actually required.

[0040] FIG. 6 is a diagram illustrating an example of extraction of a direct sound component from an impulse response signal collected in an anechoic chamber. The direct sound extraction unit 321, for example, cuts the length of the impulse response signal within a range where the frequency characteristics do not change (which may include a tolerance). In other words, the range where the frequency characteristics do not change is estimated to be the direct sound component. For example, the length L of the extracted portion P is adjusted so that the frequency characteristics of the convolution of the entire signal W and the frequency characteristics of the convolution of the extracted portion P are the same. Then, the portion other than the extracted portion P is cut, and the required number of samples (N) is extracted from the anechoic chamber-recorded IR data. These N samples of anechoic chamber-recorded IR data are sent to the delay value measurement unit 322 shown in FIG. Note that the estimation and extraction of the direct sound component by the direct sound extraction unit 321 may be performed in advance. Alternatively, the required number of samples (N) may be manually estimated and set (stored) based on the direct sound component.

[0041] The delay value measurement unit 322 convolves these N samples of anechoic chamber-recorded IR data with the dry data to obtain a delay value (D) of the learning-recorded IR data. FIG. 7 is a diagram for explaining the measurement of the delay value (D). The delay value measurement unit 322 compares the dry data (dry sound source) with the convolved signal (convoluted sound source) to obtain the delay value (D). The delay value (D) obtained in this manner is supplied to the characteristic addition unit 325.

[0042] Meanwhile, the direct sound convolution unit 323 extracts direct sound components from the learning-use recorded IR data and convolves the extracted N samples of learning-use recorded IR data with the dry data to generate a convolution signal (convolution sound source). FIG. 8 is a diagram for explaining an example of extracting direct sound components from an impulse response signal collected for learning. As shown in the figure, the direct sound convolution unit 323 extracts the beginning portion of the learning-use recorded IR data (IR for convolution). Specifically, the direct sound convolution unit 323 extracts the above-mentioned N samples. Then, the direct sound convolution unit 323 convolves the extracted N samples of learning-use recorded IR data with the dry data to generate a convolution signal. This convolution signal is sent to the average level ratio calculation unit 324 shown in FIG. 3.

[0043] 3, the average level ratio calculation unit 324 calculates the average level ratio (A) using the average level of the dry data S1 and the average level of the convolution signal S2 sent from the direct sound convolution unit 323. The average level ratio (A) is calculated by the following formula. A = (S2 average level) / (S1 average level) The calculated average level ratio (A) is supplied to the characteristic adding section 325 .

[0044] The characteristic adding unit 325 adjusts the delay and level of the dry data to generate training data. Specifically, the characteristic adding unit 325 adds characteristics of a delay value (D) and an average level ratio (A) to the dry data.

[0045] 9 is a diagram for explaining the addition of characteristics. As shown in the figure, a high-quality sound source without reverberation (dry data S1) is represented by X(n), and its average level (average level of original sound) is represented by Tref. Also, the average level of a sound source (convoluted signal S2) obtained by convoluting the direct sound component of the recorded IR data for learning estimated and extracted as described above is represented by Tconv, and the delay value (D) is represented by Δ. In this case, the teacher sound source Y(n) used for learning can be calculated by the following formula: Y(n)=(Tconv / Tref)×X(n-Δ) It should be noted that (Tconv / Tref) is the average level ratio (A).

[0046] The teacher sound source (teacher data) generated in this way is sent to the DNN learning unit 33 shown in Figure 2, as described above, and used for learning. All combinations of recorded IR data for learning and dry data are learned using the learning sound source and the teacher sound source. This makes it possible to output a processed signal in which reverberation and recording characteristics have been removed from the vocal signal when a vocal signal is input.

[0047] The learning in the DNN learning unit 33 may use time-axis data or frequency-axis data. When frequency-axis data is used, learning may include phase, or only amplitude values. Amplitude values ​​may be learned using linear amplitude values ​​or normalized dB values. In the latter case, learning can be performed well even for reverberation characteristics in which amplitude values ​​decrease exponentially over time. Furthermore, linear amplitude weighting may be used to efficiently process signals with higher levels that are more easily heard. This maintains processing accuracy without degrading sound quality, enabling good evaluations not only numerically but also subjectively.

[0048] [2-3. Processing flow example] 10 is a flowchart showing the flow of the learning process by the information processing device 3. When the learning process starts, the information processing device 3 acquires necessary data such as dry data and recorded IR data for learning (step S1). Specifically, this data is acquired by reading it from a storage unit, for example.

[0049] Then, the information processing device 3 generates learning data using the above-described learning data generating unit 31 (step S2). Furthermore, the information processing device 3 generates teacher data using the teacher data generating unit 32. That is, the following processing is performed.

[0050] The information processing device 3 extracts direct sound components from the anechoic chamber recorded IR data using the direct sound extraction unit 321 (step S3). Next, the information processing device 3 measures the delay value (D) using the delay value measurement unit 322 (step S4).

[0051] Furthermore, the information processing device 3 extracts direct sound components (specifically, N samples) from the recorded IR data for learning using the direct sound convolution unit 323, and convolves the extracted direct sound components with the dry data to generate a convolution signal (step S5). Next, the information processing device 3 calculates the average level ratio (A) using the average level ratio calculation unit 324 (step S6). Then, the information processing device 3 adds the characteristics of the delay value (D) and the average level ratio (A) to the dry data using the characteristic addition unit 325 to generate training data (step S7).

[0052] Finally, the information processing device 3 causes the DNN learning unit 33 to learn a learning model using DNN with the learning data and teacher data as input data, and ends the process, thereby constructing the learning model.

[0053] [2-4. Hardware configuration example] 11 shows an example of the hardware configuration of an information processing device (computer) 3. The information processing device 3 has a control unit 101, a storage unit 102, an input unit 103, a communication unit 104, and an output unit 105, which are interconnected by a bus.

[0054] The control unit 101 is composed of, for example, a CPU (Central Processing Unit), RAM (Random Access Memory), and ROM (Read Only Memory). The ROM stores programs that are read and run by the CPU. The RAM is used as a work memory for the CPU. The CPU controls the entire information processing device 3 by executing various processes and issuing commands in accordance with the programs stored in the ROM. The control unit 101 executes, for example, the reverberation processing and learning processing described above.

[0055] The storage unit 102 is a storage medium configured, for example, by an HDD (Hard Disc Drive), SSD (Solid State Drive), semiconductor memory, etc., and stores content data such as image data, video data, audio data, text data, etc., as well as data such as programs (for example, applications). The storage unit 102 stores, for example, programs that cause the information processing device 3 to execute the various processes described above, as well as data such as dry data, anechoic chamber-recorded IR data, and learning-recorded IR data.

[0056] It should be noted that these programs and data do not have to be stored in the storage unit 102. For example, the information processing device 3 may read and use the programs and data stored in a readable storage medium. Examples of this storage medium include an optical disk, a magnetic disk, a semiconductor memory, and a HDD that are detachable from the information processing device 3. Alternatively, the programs and data may be stored in a device connected to a network such as the Internet, and the information processing device 3 may read and use the programs and data from there. The programs may be, for example, a plug-in program that adds some or all of the above-described processing to an existing program.

[0057] The input unit 103 is a device for inputting various types of information to the information processing device 3. When information is input by the input unit 103, the control unit 101 performs various processes corresponding to the input information. The input unit 103 may be a mouse and a keyboard, a touch panel, a touch screen integrated with a monitor, physical buttons, etc. The input unit 103 may also be a sensor such as a microphone. Note that the input of various types of information to the information processing device 3 may be configured to be performed via the communication unit 104, which will be described later.

[0058] The communication unit 104 is a communication module that communicates with other devices and the Internet using a predetermined communication standard. Communication methods include wireless LANs (Local Area Networks) such as Wi-Fi (Wireless Fidelity), 4G (fourth generation mobile communication system), broadband, Bluetooth (registered trademark), and the like.

[0059] The output unit 105 is a device for outputting various types of information from the information processing device 3. The output unit 105 is a display for displaying images and videos, a speaker for outputting sound, etc. Note that the output of various types of information from the information processing device 3 may be configured to be performed via the communication unit 104.

[0060] [2-5. Summary] In this embodiment, training data is generated by convolving a measurement signal (e.g., an impulse response signal) collected by a smartphone microphone used to collect a vocal signal (observation signal) with dry data (reference signal) that has higher sound quality and less reverberation than the vocal signal. Furthermore, training data is generated by adapting the average level and delay value of the convolved signal, which is generated by convolving the direct sound component of the measurement signal with the dry data, to the dry data. Then, a learning model that dereverberates vocal signals is trained using the generated training data and training data as input data.

[0061] In this way, by using training data in which the delay and level of the dry data have been adjusted to match the training data, it is possible to remove reverberation while taking into consideration sound quality, even when recording in a reverberant location such as a home living room. By removing the reverberation components, studio engineers and others can freely add reverberation that is optimal for the content.

[0062] Because the measurement signal is recorded using the same recording system (specifically, a smartphone) as used to record vocals, the characteristics of the recording system can be canceled out. This allows high-quality recordings to be made without using expensive professional microphones. For example, when recording with a smartphone, recording can be done using only the smartphone, without the need for expensive professional recording equipment.

[0063] In addition, by performing a large amount of training based on dry data of various vocal voices, dereverberation processing can be performed regardless of the vocalist. Also, by recording IR data for training in various locations where vocal voices are expected to be recorded, dereverberation processing can be performed regardless of the location. Therefore, dereverberation that takes sound quality into consideration can be performed regardless of the recording environment.

[0064] <3. Modifications> Although the embodiments of the present disclosure have been specifically described above, the present disclosure is not limited to the above-described embodiments, and various modifications based on the technical concepts of the present disclosure are possible. For example, various modifications such as those described below are possible. Furthermore, one or more of the following modified aspects can be arbitrarily selected and combined as appropriate. Furthermore, the configurations, methods, processes, shapes, materials, and numerical values ​​of the above-described embodiments can be combined or substituted with each other without departing from the spirit of the present disclosure. Furthermore, one thing can be divided into two or more things, and two or more things can be combined into one thing. Furthermore, some parts can be omitted.

[0065] While the above-described embodiment illustrates the removal of reverberation and recording characteristics from vocals in music production, the sound (observed signal) from which reverberation and recording characteristics are removed is not limited to this, and any sound from which reverberation and recording characteristics are to be removed may be used. For example, other sounds used in music production, such as the sounds of individual instruments in an orchestra, may be used. Furthermore, the present invention is not limited to music content, and may also be applied to content such as movies, television programs, games, and announcements. In this case, the present invention can be applied to the removal of reverberation and recording characteristics from sounds such as narration, audio dubbing (post-recording), insert songs, and announcements. It can also be applied to sounds from voice recorders, online conference systems, translation devices, car navigation systems, and the like. In these cases, the dry data can be a sound source that matches the sound from which reverberation and recording characteristics are to be removed.

[0066] Furthermore, in the above-described embodiment, high-quality dry data that does not include reverberation is used as the reference signal. However, the reference signal is not limited to this, and any sound that can be used as a learning target may be used. For example, if the reference signal is recorded with a higher quality sound than the vocal signal (observed signal), the sound quality of the processed signal can be improved. Furthermore, for example, a reference signal that intentionally includes reverberation may be used to add reverberation to the observed signal. In other words, this may be applied to reverberation processing other than dereverberation.

[0067] In the above-described embodiment, a home is assumed as the location where the vocal signal is collected, and a smartphone is assumed as the recording device. However, the recording environment (location, device, etc.) is not limited to this and may be any environment. For example, the recording location may be a family member, relative, or acquaintance's home, a personal studio, a music room, a car, a rental space, an event hall, a karaoke room, etc. In this case, the learning IR data may be collected in a location where sound collection is expected. Furthermore, for example, the microphone for sound collection may be of any type, such as a microphone attached to a headset, tablet, laptop, smart home appliance, or video camera, a desktop microphone, or a handheld microphone. In this case, the user device 2 may be configured with any of these devices.

[0068] Furthermore, in the above-described embodiment, it is assumed that one type of microphone (specifically, a smartphone microphone) is used to collect sound. However, multiple types of microphones (e.g., smartphones from multiple manufacturers) may be used to collect sound. In this case, it is sufficient to collect a measurement signal for each microphone. This allows dereverberation to be performed appropriately depending on the microphone used.

[0069] Furthermore, the information processing device 3 in the above-described embodiment has been exemplified as having a function for removing reverberation and recording system characteristics (hereinafter referred to as dereverberation technology), but this dereverberation technology can also be used in new music production technology, for example. Details thereof will be described below.

[0070] In music production, especially in pop music, the revenue model is shifting from packaged media to live performances. When recording live music, each sound source is generally recorded as a separate track. For example, acoustic sound sources (such as vocals and drums) are recorded using a microphone, while the output of electronic instruments (such as keyboards) is recorded as is (directly from the instrument without using a microphone). In addition to this, it is common to use multiple audience microphones to record the reverberations in the venue, the cheers of the audience, etc.

[0071] For example, in the case of pop music, acoustic sound sources (especially vocals) are often recorded very dry, with little reverberation, due to the sound source being located very close to the microphone and the fact that dynamic microphones are the norm. Therefore, in the past, in order to match this sound source (e.g., microphone signal) with the atmosphere of the venue, it was common to mix the signal from the audience microphone mentioned above to create a reverberant component. In this way, when producing live sound sources, it is common to use multiple audience microphones to capture the reverberation of the venue, the cheers of the audience, etc., and then mix these signals with each sound source as needed to add the atmosphere of the venue.

[0072] However, in recent years, pre-mix audio sources (especially vocal sources) have frequently undergone correction processes such as pitch correction, which can result in problems such as double sounds being heard when the processed audio source is mixed with the audience microphone signal. In such cases, the audience microphone signal cannot be used, and engineers must separately create (artificially generate, for example) a reverberation that matches the acoustics of the venue and use this to create the acoustics. This makes it difficult to accurately reproduce the acoustics of the venue, which has also led to problems with work efficiency.

[0073] The dereverberation technology described above can be used to add reverberation in such cases. In other words, by using the dereverberation technology described above to extract and estimate the characteristics of the reverberation in the venue and adding this effect to the sound source such as vocals, it is possible to add the reverberation of the actual venue without mixing in the signal from the audience microphone as described above.

[0074] Specifically, first, the dereverberation technology described above is used to derive reverberation from the signals of the audience microphones, and the reverberation components are extracted from the signals. Next, the reverberation characteristics of the venue are generated inversely from the extracted reverberation components. The generated reverberation characteristics are then used, for example, as coefficients for adding reverberation to the sound source (e.g., the modified sound source of vocals, guitar, etc.). In this way, the reverberation of the actual venue can be added to the sound source. This will be described in detail below with reference to the drawings.

[0075] [Recording configuration example] FIG. 12 shows an example of the configuration of an information processing system (information processing system 1A) according to a modified example. Note that FIG. 12 shows an example of the configuration during recording. The information processing system 1A is, for example, a system for music production. The information processing system 1A has an information processing device 3A, multiple (1st to Nth) microphones 5, and multiple (1st to Mth) sound source output units 6. Note that in the case of general live recording, multiple microphones 5 and multiple sound source output units 6 are used as shown in the figure, but these may also be singular.

[0076] The information processing device 3A records and edits various sound sources, sounds, etc. The information processing device 3A has, for example, the same hardware configuration (see FIG. 11) as the information processing device 3 of the above-described embodiment. For example, the information processing device 3A realizes music production using a DAW (Digital Audio Workstation). The information processing device 3A is connected to each microphone 5 and each sound source output unit 6 via, for example, an audio interface (not shown) such as a mixer.

[0077] The microphones 5 pick up sounds from the venue, such as the reverberation of the venue and the voices of the audience (sounds from locations where reverberation to be added occurs), and output them as microphone signals. The microphones 5 are, for example, directional audience microphones, and are installed to pick up sounds from various locations in the venue. For example, a pair of microphones 5 for the L channel and the R channel are installed in multiple locations with different sound pickup targets so as to pick up sounds from audience seats near the stage, audience seats far from the stage, above the audience seats, etc. The output signals of each microphone 5 (microphone outputs 1 to N) are input to the information processing device 3A.

[0078] The sound source output unit 6 outputs a sound source signal (sound source data). The sound source output unit 6 is composed of, for example, a microphone that collects sounds from sound sources such as vocals, drums, and piano and outputs them as a microphone signal, or a musical instrument (for example, a keyboard) that can directly output (line output) a sound source signal. The output signals (1st to Mth sound source outputs) of each sound source output unit 6 are input to the information processing device 3A.

[0079] The information processing device 3A has a recording unit 7 and a memory 8. The recording unit 7 records each signal input from each microphone 5 and each sound source output unit 6 in the memory 8. For example, the recording unit 7 is configured to include an ADC (Analog-to-Digital Converter) and the like, and converts each signal input from each microphone 5 and each sound source output unit 6 into recorded data and stores it in the memory 8.

[0080] The memory 8 is configured with a storage device such as an HDD or SSD, and stores the recorded data converted by the recording unit 7. The memory 8 may be any device capable of recording the recorded data, and may be, for example, an optical disc such as a DVD (Digital Versatile Disc), a portable semiconductor memory, or another removable recording medium.

[0081] [Editing configuration example] 13 shows an example of the configuration during editing. In this modification, the same information processing device 3A is used for recording and editing, but these may be different. If different devices are used for recording and editing, it is sufficient if the data recorded by the recording device can be used by the editing device, for example, by data communication or by attaching and detaching a recording medium.

[0082] The information processing device 3A has a function of adding reverberation (including sounds from the venue such as audience voices and applause) to a recorded sound source. The information processing device 3A has a dereverberation unit 9, a reverberation component feature extraction unit 10, a user interface unit 11, a reverberation coefficient derivation unit 12, and a reverberation addition unit 13.

[0083] The dereverberation unit 9 performs dereverberation using the above-mentioned dereverberation technology. The information processing device 3A has, for example, the trained learning model described in the embodiment (see FIGS. 2 and 3, etc.), and the dereverberation unit 9 can use the trained model.

[0084] Specifically, this learning model was previously machine-trained using first data (learning data) generated by convolving first acoustic data (measurement signal) collected by a specified microphone 5 (see Figure 12) with dry data (reference signal) that does not contain reverberation, and second data (teaching data) generated by convolving the direct sound component of the first acoustic data (measurement signal) with the dry data (reference signal) and adapting the average level and delay value of the convolved signal to the dry data (reference signal).

[0085] The dereverberation unit 9 uses this learning model to input second acoustic data (observation signal) picked up by the above-mentioned predetermined microphone 5, and outputs third acoustic data in which the reverberation component has been removed from the second acoustic data (observation signal). To put it simply, the dereverberation unit 9 inputs the microphone signal of the above-mentioned predetermined microphone 5 stored in the memory 8, and dereverberates the input microphone signal (e.g., an input signal having the characteristics shown in FIG. 14) using the above-mentioned learning model, and outputs a signal from which the reverberation has been removed (e.g., a dereverberated signal having the characteristics shown in FIG. 15). The third acoustic data (dereverberated signal) output by the dereverberation unit 9 is input to the reverberation component feature extraction unit 10.

[0086] The reverberation component feature extraction unit 10 uses the second acoustic data and the third acoustic data to extract feature data representing the reverberation components of the second acoustic data. This feature data may be, for example, a feature data of an impulse response. The reverberation component feature extraction unit 10 receives the microphone signal (the original signal before dereverberation by the dereverberation unit 9) of the predetermined microphone 5 stored in the memory 8, extracts the difference between the microphone signal and the dereverberation signal received from the dereverberation unit 9, calculates the feature data of the reverberation component (e.g., the reverberation component having the characteristics shown in FIG. 16), and generates feature data representing the calculated feature data. The feature data extracted (generated) by the reverberation component feature extraction unit 10 is input to the reverberation coefficient derivation unit 12.

[0087] The user interface unit 11 generates control data in response to user operations. The user interface unit 11 is composed of input devices such as a keyboard, a mouse, and physical buttons, and is used by an engineer to actually listen to the sound. The user interface unit 11 is used to allow the user to control the settings for adding reverberation in the reverberation unit 13. Examples of settings for adding reverberation include adjusting the wet and dry levels. Specifically, the user interface unit 11 generates control data (control data for adding settings) for controlling settings such as reverberation time (e.g., 1 second) and reverberation amount (e.g., -40 dB). The control data generated by the user interface unit 11 is input to the reverberation unit 13.

[0088] The reverberation coefficient derivation unit 12 (characteristics determination unit) determines the characteristics of reverberation to be added to the sound source data in the reverberation adding unit 13. In this configuration example, these characteristics are determined to be the characteristics of the reverberation components extracted by the reverberation component feature extraction unit 10. Specifically, the reverberation coefficient derivation unit 12 converts the feature data input from the reverberation component feature extraction unit 10 into coefficient data that can be used in the subsequent reverberation adding unit 13. The coefficient data converted by the reverberation coefficient derivation unit 12 is input to the reverberation adding unit 13.

[0089] The reverberation unit 13 adds reverberation to sound source data and outputs the resulting sound. The reverberation unit 13 inputs sound source data (for example, sound source data after correction processing such as the above-mentioned pitch correction) from the memory 8, adds reverberation to the input sound source data based on the feature data extracted (generated) by the reverberation component feature extraction unit 10, and outputs the resulting sound source data. Specifically, the reverberation unit 13 synthesizes reverberation components having the characteristics of the above-mentioned feature data into the input sound source data using coefficient data input from the reverberation coefficient derivation unit 12. At this time, the reverberation unit 13 changes the settings for adding reverberation in accordance with control data input from the user interface unit 11.

[0090] The sound source data to which reverberation is added is not limited to that recorded during the recording process as described above, but may be separately prepared. Sound sources heard at live venues are designed to sound consistent depending on the location, and may not be optimal for music production. In such cases, music production requires separate sound sources, but this system can also accommodate such situations. The reverberated sound source data output from the reverberation unit 13 is stored in the memory 8, for example, and used for music production.

[0091] As described above, in this modified example, first, the dereverberation unit 9 inputs the second acoustic data collected by the specified microphone 5 and outputs third acoustic data in which the reverberation components have been removed from the second acoustic data, using a learning model that has been machine-trained using first data generated by convolving first acoustic data collected by a specified microphone 5 with dry data that does not contain reverberation, and second data generated by convolving the direct sound component of the first acoustic data with the dry data and adapting the average level and delay value of the convolved signal to the dry data.

[0092] Next, the reverberation component feature extraction unit 10 extracts feature data representing the reverberation components of the second acoustic data using the second acoustic data and the third acoustic data. After that, the reverberation adding unit 13 inputs the sound source data, adds reverberation with characteristics based on the feature data (specifically, characteristics represented by the feature data) to the input sound source data, and outputs the resulting data.

[0093] This allows the reverberation of an actual venue to be added without mixing the signal from the audience microphone with the recorded sound source, as described above. Therefore, even if reverberation is added, the problem of double sounds as described above does not occur. This eliminates the need to consider and add similar reverberation, significantly reducing the engineer's workload. Furthermore, by using the dereverberation technology of one embodiment described above in the dereverberation unit 9, sound quality is improved and more accurate dereverberation is performed, enabling more accurate processing.

[0094] [Other configuration examples when editing] 17 shows another example of the configuration during editing. In the description of this example, the same reference numerals are used to designate components that are the same as or of the same quality as those in the above description, and redundant description will be omitted where appropriate.

[0095] The information processing device 3A in this configuration example includes a memory 8, multiple (1st to Nth) dereverberation units 9, multiple (1st to Nth) reverberation component feature extraction units 10, a user interface unit 11, a reverberation coefficient derivation unit 12A, multiple (1st to Mth) reverberation units 13, a mixer unit 14, and a total reverberation unit 15. The number of dereverberation units 9, reverberation component feature extraction units 10, and reverberation units 13 may be changed as appropriate.

[0096] Each dereverberation unit 9 receives second acoustic data (first to Nth microphone signals in the figure) collected by a corresponding microphone 5 (see FIG. 12) and outputs third acoustic data in which reverberation components have been removed from the second acoustic data. For example, the first dereverberation unit 9 receives from the memory 8 the first microphone signal collected and recorded by the first microphone 5, and outputs a dereverberated signal in which reverberation has been removed from the input first microphone signal. The first dereverberation unit 9 uses a learning model that has been machine-learned using the first acoustic data (measurement signal) collected by the first microphone 5. The same applies to the second to Nth dereverberation units 9.

[0097] The third acoustic data output from each dereverberation unit 9 is input to the corresponding reverberation component feature extraction unit 10. For example, the third acoustic data output from the first dereverberation unit 9 is input to the first reverberation component feature extraction unit 10.

[0098] Each reverberation component feature extraction unit 10 extracts feature data representing the reverberation component of the second acoustic data using the second acoustic data and the third acoustic data. For example, the first reverberation component feature extraction unit 10 receives the first microphone signal from the memory 8 and extracts feature data representing the reverberation component of the first microphone signal using the received first microphone signal and the dereverberated signal output from the first dereverberation unit 9. The feature data extracted by each reverberation component feature extraction unit 10 is input to the reverberation coefficient derivation unit 12A.

[0099] As described above, the user interface unit 11 generates control data in response to user operations and is used to allow the user to control the reverberation settings of the multiple reverberation units 13 and the total reverberation unit 15. Specifically, the user interface unit 11 generates control data (control data for reverberation settings) for controlling the reverberation settings. This control data is input to the corresponding multiple reverberation units 13 and total reverberation unit 15.

[0100] Furthermore, the user interface unit 11 in this configuration example is used to allow the user to control the selection of characteristics in the reverberation coefficient derivation unit 12A, which will be described later. Specifically, the user interface unit 11 generates control data (control data for characteristic selection) that controls the selection of this characteristic. This control data is input to the reverberation coefficient derivation unit 12A.

[0101] The reverberation coefficient derivation unit 12A (characteristics determination unit) determines the characteristics of reverberation to be added to sound source data in each of the multiple reverberation adding units 13 and the total reverberation adding unit 15. The reverberation coefficient derivation unit 12A selectively determines these characteristics, for example, using feature data extracted by each of the multiple reverberation component feature extracting units 10. Specifically, the reverberation coefficient derivation unit 12A selects the characteristics of reverberation to be added to the sound source data in accordance with control data (control data for characteristic selection) input from the user interface unit 11.

[0102] The reverberation coefficient derivation unit 12A determines, for example, the characteristics of feature data selected from the feature data extracted by the multiple reverberation component feature extraction units 10 as the reverberation characteristics to be added to the sound source data. In other words, the reverberation coefficient derivation unit 12A selects, from the feature data extracted by the multiple reverberation component feature extraction units 10, characteristics to be added to the sound source data by each of the multiple reverberation units 13 and the total reverberation unit 15. The reverberation coefficient derivation unit 12A then outputs coefficient data of the selected feature data to the corresponding multiple reverberation units 13 and the total reverberation unit 15. To achieve this, for example, each feature data may be expanded as coefficient data as described above so that it can be used in each of the multiple reverberation units 13 and the total reverberation unit 15 at the subsequent stages.

[0103] The reverberation coefficient derivation unit 12A may determine the average characteristics of the features represented by the feature data extracted by the multiple reverberation component feature extraction units 10 as the reverberation characteristics to be added to the sound source data, and output the coefficient data for adding the determined reverberation characteristics to the corresponding multiple reverberation units 13 and the total reverberation unit 15. The average characteristics may be determined by using a representative value of the average characteristics represented by each feature data. Alternatively, instead of using a single representative value, it may be possible to perform processing such as partial averaging in several patterns, such as weak, medium, and strong. In other words, the partial average of the features may be used to select the reverberation characteristics to be added from among multiple candidates. This selection may be performed, for example, in accordance with the control data (control data for characteristic selection) described above.

[0104] Each coefficient data generated and output by the reverberation coefficient derivation unit 12A is input to the corresponding multiple reverberation units 13 and total reverberation unit 15. Each reverberation unit (sound source-specific reverberation unit) 13 inputs each sound source data (e.g., sound source data after the above-mentioned correction processing) from the memory 8, adds reverberation to the input sound source data, and outputs reverberated sound source data. The reverberated sound source data output by each reverberation unit 13 is input to a mixer unit 14. The mixer unit 14 inputs and synthesizes the reverberated sound source data output from each of the multiple reverberation units 13. The mixed sound source data synthesized by the mixer unit 14 is output to the total reverberation unit 15.

[0105] The total reverberation unit 15 adds reverberation to mixed sound source data and outputs the resulting data. The total reverberation unit 15 receives the mixed sound source data synthesized by the mixer unit 14, adds reverberation to the input mixed sound source data based on the feature data extracted (generated) by each of the multiple reverberation component feature extraction units 10, and outputs the resulting mixed sound source data. Specifically, like the reverberation unit 13, the total reverberation unit 15 synthesizes reverberation components into the mixed sound source data using coefficient data input from the reverberation coefficient derivation unit 12A. The reverberated mixed sound source data output from the total reverberation unit 15 is stored in the memory 8, for example.

[0106] In addition to the effects shown in FIG. 13, this configuration example has the following advantages. By processing the second acoustic signals picked up by the multiple microphones 5, multiple feature data can be used. The characteristics of the recorded reverberation vary depending on the type, installation location, direction, etc. of the microphones 5. Therefore, multiple feature data representing reverberation components with slightly different characteristics are input to the reverberation coefficient derivation unit 12A. This increases the number of reverberation characteristics to be added to the sound source data, allowing the engineer to select the quality of the reverberation from among them. In other words, it is possible to use sounds from various venues (realistic reverberations), broadening the range of expression available to the engineer.

[0107] Furthermore, by including multiple reverberation units 13, it is possible to add reverberation to each sound source, and furthermore, by including the mixer unit 14 and the total reverberation unit 15, it is possible to add reverberation to mixed sound source data that is a mixture of sound source data with reverberation. Therefore, it is possible to add detailed reverberation that is in line with reality, for example, by adding strong and short reverberation to individual sound sources and weak and long reverberation to mixed sound sources.

[0108] The present disclosure can also be configured as follows. (1) An information processing device having a learning processing unit that generates training data by convolving a measurement signal representing acoustic characteristics collected by the same sound collection unit as that used to collect an observation signal with a reference signal having sound quality and reverberation characteristics different from those of the observation signal, generates teacher data by adapting an average level and a delay value of the convolved signal generated by convolving a direct sound component of the measurement signal with the reference signal, and trains a learning model that performs reverberation processing of the observation signal collected by the sound collection unit using the training data and the teacher data as input data. (2) The direct sound component is estimated from the measurement signal recorded in an anechoic chamber. The information processing device described in (1). (3) The delay value is a delay value of a convoluted signal generated by extracting the estimated direct sound component and convolving it with the reference signal. (2) An information processing device according to the present invention. (4) The dereverberation process is a dereverberation process that uses a signal with characteristics that do not include reverberation as the reference signal. An information processing device according to any one of (1) to (3). (5) The reference signal is recorded as a sound with higher quality than the observed signal. An information processing device according to any one of (1) to (4). (6) The observed signal and the reference signal are vocal signals. An information processing device according to any one of (1) to (5). (7) The observation signal and the measurement signal are picked up by a microphone of a terminal device. An information processing device according to any one of (1) to (6). (8) An information processing method that performs a learning process to generate training data by convolving a measurement signal representing acoustic characteristics picked up by the same sound pickup unit as that used to pick up an observation signal with a reference signal having sound quality and reverberation characteristics different from those of the observation signal, generate teacher data by adapting the average level and delay value of the convolved signal, which is generated by convolving a direct sound component of the measurement signal with the reference signal, to the reference signal, and learn a learning model that performs reverberation processing of the observation signal picked up by the sound pickup unit using the training data and the teacher data as input data. (9) A program that causes a computer to execute a learning process that generates training data by convolving a measurement signal, which represents acoustic characteristics and is picked up by the same sound pickup unit as that used to pick up an observation signal, with a reference signal whose sound quality and reverberation characteristics are different from those of the observation signal, generates teacher data by adapting the average level and delay value of the convolved signal, which is generated by convolving the direct sound component of the measurement signal with the reference signal, to the reference signal, and learns a learning model that performs reverberation processing of the observation signal picked up by the sound pickup unit using the training data and the teacher data as input data. (10) a dereverberation unit that receives the second sound data collected by the microphone and outputs third sound data obtained by removing the reverberation components from the second sound data, using a learning model that is machine-learned using first data generated by convolving the first sound data collected by the microphone with dry data that does not contain reverberation, and second data generated by convolving the direct sound component of the first sound data with the dry data and adapting the average level and delay value of the convolution signal to the dry data; a reverberation component feature extraction unit that extracts feature data representing a reverberation component of the second acoustic data by using the second acoustic data and the third acoustic data; a reverberation adding unit that inputs sound source data, adds reverberation characteristics based on the feature data, and outputs the result; An information processing device having the above. (11) a plurality of the dereverberation units and the reverberation component feature extraction units; a characteristic determining unit that selectively determines the characteristics of the reverberation to be added to the sound source data using the feature data extracted by each of the plurality of reverberation component feature extracting units; (10) An information processing device according to (10). (12) a user interface unit that generates control data in response to a user operation; The characteristic determination unit selects the characteristics of the reverberation to be added to the sound source data in accordance with the control data. (11) An information processing device according to (11). (13) The characteristic determination unit determines the characteristic of feature data selected from the feature data extracted by the plurality of reverberation component feature extraction units as the reverberation characteristic to be added to the sound source data. The information processing device according to (11) or (12). (14) a plurality of the dereverberation units and the reverberation component feature extraction units; a characteristic determining unit that determines an average characteristic of the feature data extracted by each of the plurality of reverberation component feature extracting units as a reverberation characteristic to be added to the sound source data; An information processing device according to any one of (10) to (12). (15) a plurality of the reverberation units; a mixer unit that mixes the reverberated sound source data output from each of the plurality of reverberation units; a total reverberation adding unit that adds reverberation having characteristics based on the feature data to the reverberated sound source data synthesized by the mixer unit and outputs the resultant data; have An information processing device according to any one of (10) to (14). (16) The sound source data is data after correction processing. An information processing device according to any one of (10) to (15). (17) a user interface unit that generates control data in response to a user operation; The reverberation adding unit changes the setting for adding reverberation in accordance with the control data. An information processing device according to any one of (10) to (16). (18) a learning model machine-learned using first data generated by convolving first sound data collected by a microphone with dry data that does not contain reverberation, and second data generated by convolving a direct sound component of the first sound data with the dry data and adapting the average level and delay value of the convolved signal to the dry data, to input the second sound data collected by the microphone, and output third sound data in which the reverberation component has been removed from the second sound data; extracting feature data representing a reverberation component of the second acoustic data using the second acoustic data and the third acoustic data; Input the sound source data, add reverberation characteristics based on the feature data, and output the result An information processing method in which processing is performed by a computer. (19) a learning model machine-learned using first data generated by convolving first sound data collected by a microphone with dry data that does not contain reverberation, and second data generated by convolving a direct sound component of the first sound data with the dry data and adapting the average level and delay value of the convolved signal to the dry data, to input the second sound data collected by the microphone, and output third sound data in which the reverberation component has been removed from the second sound data; extracting feature data representing a reverberation component of the second acoustic data using the second acoustic data and the third acoustic data; Input the sound source data, add reverberation characteristics based on the feature data, and output the result A program that causes a computer to perform a process. (20) A microphone and a learning model that is machine-learned using first data generated by convolving first sound data collected by the microphone with dry data that does not include reverberation, and second data generated by convolving a direct sound component of the first sound data with the dry data and adapting the average level and delay value of a convolved signal to the dry data; a sound source output unit that outputs sound source data; a memory that stores second acoustic data collected by the microphone and the sound source data; a dereverberation unit that uses the learning model to input the second acoustic data and output third acoustic data obtained by removing reverberation components from the second acoustic data; a reverberation component feature extraction unit that extracts feature data representing a reverberation component of the second acoustic data by using the second acoustic data and the third acoustic data; a reverberation adding unit that inputs the sound source data, adds reverberation characteristics based on the feature data, and outputs the result; An information processing system having the above. [Explanation of symbols]

[0109] 1,1A...information processing system, 3,3A...information processing device, 5...microphone, 6...sound source output unit, 8...memory, 9...dereverberation unit, 10...reverberation component feature extraction unit, 11...user interface unit, 12,12A...reverberation coefficient derivation unit, 13...reverberation addition unit, 14...mixer unit, 15...total reverberation addition unit, 30...learning processing unit, 31...learning data generation unit, 20...sound collection unit, 32...teaching data generation unit, 33...DNN learning unit, 321...direct sound extraction unit, 322...delay value measurement unit, 323...direct sound convolution unit, 324...average level ratio calculation unit, 325...characteristic addition unit

Claims

1. An information processing device having a learning processing unit that generates training data by convolving a measurement signal representing acoustic characteristics collected by the same sound collection unit as that used to collect an observation signal with a reference signal having sound quality and reverberation characteristics different from those of the observation signal, generates teacher data by adapting an average level and a delay value of the convolved signal generated by convolving a direct sound component of the measurement signal with the reference signal, and trains a learning model that performs reverberation processing of the observation signal collected by the sound collection unit using the training data and the teacher data as input data.

2. The direct sound component is estimated from the measurement signal recorded in an anechoic chamber. The information processing device according to claim 1 .

3. The delay value is a delay value of a convoluted signal generated by extracting the estimated direct sound component and convolving it with the reference signal. The information processing device according to claim 2 .

4. The dereverberation process is a dereverberation process that uses a signal with characteristics that do not include reverberation as the reference signal. The information processing device according to claim 1 .

5. The reference signal is recorded as a sound with higher quality than the observed signal. The information processing device according to claim 1 .

6. The observed signal and the reference signal are vocal signals. The information processing device according to claim 1 .

7. The observation signal and the measurement signal are picked up by a microphone of a terminal device. The information processing device according to claim 1 .

8. An information processing method that performs a learning process to generate training data by convolving a measurement signal representing acoustic characteristics picked up by the same sound pickup unit as that used to pick up an observation signal with a reference signal having sound quality and reverberation characteristics different from those of the observation signal, generate teacher data by adapting the average level and delay value of the convolved signal, which is generated by convolving a direct sound component of the measurement signal with the reference signal, to the reference signal, and learn a learning model that performs reverberation processing of the observation signal picked up by the sound pickup unit using the training data and the teacher data as input data.

9. A program that causes a computer to execute a learning process that generates training data by convolving a measurement signal, which represents acoustic characteristics and is picked up by the same sound pickup unit as that used to pick up an observation signal, with a reference signal whose sound quality and reverberation characteristics are different from those of the observation signal, generates teacher data by adapting the average level and delay value of the convolved signal, which is generated by convolving the direct sound component of the measurement signal with the reference signal, to the reference signal, and learns a learning model that performs reverberation processing of the observation signal picked up by the sound pickup unit using the training data and the teacher data as input data.

10. a dereverberation unit that receives the second sound data collected by the microphone and outputs third sound data obtained by removing the reverberation components from the second sound data, using a learning model that has been machine-learned using first data generated by convolving the first sound data collected by the microphone with dry data that does not contain reverberation, and second data generated by convolving the direct sound component of the first sound data with the dry data and adapting the average level and delay value of the convolved signal to the dry data; a reverberation component feature extraction unit that extracts feature data representing a reverberation component of the second acoustic data using the second acoustic data and the third acoustic data; a reverberation adding unit that inputs sound source data, adds reverberation characteristics based on the feature data, and outputs the result; An information processing device having the above.

11. a plurality of the dereverberation units and the reverberation component feature extraction units; a characteristic determining unit that selectively determines the characteristics of the reverberation to be added to the sound source data using the feature data extracted by each of the plurality of reverberation component feature extracting units; The information processing device according to claim 10.

12. a user interface unit that generates control data in response to a user operation; The characteristic determination unit selects the characteristics of the reverberation to be added to the sound source data in accordance with the control data. The information processing device according to claim 11.

13. The characteristic determination unit determines the characteristic of feature data selected from the feature data extracted by the plurality of reverberation component feature extraction units as the reverberation characteristic to be added to the sound source data. The information processing device according to claim 11.

14. a plurality of the dereverberation units and the reverberation component feature extraction units; a characteristic determining unit that determines an average characteristic of the feature data extracted by each of the plurality of reverberation component feature extracting units as a reverberation characteristic to be added to the sound source data; The information processing device according to claim 10.

15. a plurality of the reverberation units; a mixer unit that mixes the reverberated sound source data output from each of the plurality of reverberation units; a total reverberation adding unit that adds reverberation having characteristics based on the feature data to the reverberated sound source data synthesized by the mixer unit and outputs the resultant data; have The information processing device according to claim 10.

16. The sound source data is data after correction processing. The information processing device according to claim 10.

17. a user interface unit that generates control data in response to a user operation; The reverberation adding unit changes the setting for adding reverberation in accordance with the control data. The information processing device according to claim 10.

18. a learning model machine-learned using first data generated by convolving first sound data collected by a microphone with dry data that does not contain reverberation, and second data generated by convolving a direct sound component of the first sound data with the dry data and adapting the average level and delay value of a convolved signal generated by convolving the direct sound component of the first sound data with the dry data, to the dry data; and extracting feature data representing a reverberation component of the second acoustic data using the second acoustic data and the third acoustic data; Input the sound source data, add reverberation characteristics based on the feature data, and output the result An information processing method in which processing is performed by a computer.

19. a learning model machine-learned using first data generated by convolving first sound data collected by a microphone with dry data that does not contain reverberation, and second data generated by convolving a direct sound component of the first sound data with the dry data and adapting the average level and delay value of a convolved signal generated by convolving the direct sound component of the first sound data with the dry data, to the dry data; and extracting feature data representing a reverberation component of the second acoustic data using the second acoustic data and the third acoustic data; Input the sound source data, add reverberation characteristics based on the feature data, and output the result A program that causes a computer to perform a process.

20. A microphone and a learning model that is machine-learned using first data generated by convolving first sound data collected by the microphone with dry data that does not include reverberation, and second data generated by convolving a direct sound component of the first sound data with the dry data and adapting the average level and delay value of a convolved signal to the dry data; a sound source output unit that outputs sound source data; a memory that stores second acoustic data collected by the microphone and the sound source data; a dereverberation unit that receives the second acoustic data and removes reverberation components from the second acoustic data using the learning model, and outputs third acoustic data; a reverberation component feature extraction unit that extracts feature data representing a reverberation component of the second acoustic data using the second acoustic data and the third acoustic data; a reverberation adding unit that inputs the sound source data, adds reverberation characteristics based on the feature data, and outputs the result; An information processing system having the above.

Citation Information

Patent Citations

  • Processor, voice recognition device, voice recognition system and voice recognition method

    JP2009276365A

  • Acoustic processor

    JP2014077916A

  • Signal processor, signal processing method and signal processing program

    JP2018146683A

  • Signal processing device using neural network, signal processing method using neural network, and signal processing program

    WO2019026973A1