Signal processing device, signal processing method, and signal processing program

The signal processing device enhances sound source separation by converting mixed acoustic signals into feature quantities and using ambient sound features to estimate masks, eliminating preparatory processing and ensuring efficient separation of multiple acoustic signals.

JP7880349B2Active Publication Date: 2026-06-25PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
Filing Date
2022-10-11
Publication Date
2026-06-25

AI Technical Summary

Technical Problem

Conventional sound source separation techniques require complex preparatory processing to create auxiliary information and may suffer performance degradation when ambient noise is not used in training, leading to inefficient separation of multiple acoustic signals.

Method used

A signal processing device that converts mixed acoustic signals into feature quantities, estimates masks using ambient sound features, and separates acoustic signals without prior auxiliary information, utilizing multiple acoustic models to enhance real-time separation.

Benefits of technology

Eliminates the need for complex preparatory processing and prevents performance degradation by enabling real-time separation of ambient and target acoustic signals, improving the efficiency of sound source separation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007880349000001
    Figure 0007880349000001
  • Figure 0007880349000002
    Figure 0007880349000002
  • Figure 0007880349000003
    Figure 0007880349000003
Patent Text Reader

Abstract

A signal processing device (1) comprises: a mixed acoustic signal acquisition unit (11) that acquires a mixed acoustic signal including a plurality of acoustic signals; a mixed feature amount conversion unit (12) that converts the mixed acoustic signal into a mixed feature amount; a mask estimation unit (15) that estimates a plurality of masks on the basis of the mixed feature amount; an acoustic signal conversion unit (16) that converts a plurality separated feature amounts calculated by using the masks into a plurality of separated acoustic signals; an environment sound section estimation unit (18) that estimates an environment sound section including only an environment sound on the basis of the separated acoustic signals; an environment acoustic signal extraction unit (19) that extracts, from the mixed acoustic signal, a mixed acoustic signal in the environment sound section as an environment acoustic signal; and an environment sound feature amount conversion unit (14) that converts the environment acoustic signal into an environment sound feature amount. The mask estimation unit (15) estimates the masks on the basis of the mixed feature amount weighted by using the environment sound feature amount.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a technique for separating a plurality of acoustic signals from a mixed acoustic signal.

Background Art

[0002] For example, in Patent Document 1, there are a conversion unit that converts an input mixed acoustic signal into a plurality of first internal states, and when auxiliary information regarding the acoustic signal of a target sound source is input, a second internal state that is a weighted sum of the plurality of first internal states is generated based on the auxiliary information, and when the auxiliary information is not input, a weighting unit that generates the second internal state by selecting any one of the plurality of first internal states, and a mask estimation unit that estimates a mask based on the second internal state. A signal processing device is disclosed.

[0003] However, in the above conventional technology, there is a possibility that complicated preparation processing for creating auxiliary information regarding the acoustic signal of a target sound source may be required, and there is also a possibility that the performance of separating a plurality of acoustic signals from a mixed acoustic signal may deteriorate, and further improvement has been required.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

[0005] The present disclosure has been made to solve the above problems, and an object thereof is to provide a technique that eliminates complicated preparation processing for creating auxiliary information regarding the acoustic signal of a target sound source and can prevent a deterioration in the performance of separating a plurality of acoustic signals from a mixed acoustic signal.

[0006] The signal processing device according to this disclosure includes: a mixed acoustic signal acquisition unit that acquires a mixed acoustic signal including a plurality of acoustic signals; a mixed feature conversion unit that converts the mixed acoustic signal into a mixed feature quantity that shows the characteristics of the mixed acoustic signal; a mask estimation unit that estimates a plurality of masks corresponding to each of the plurality of acoustic signals based on the mixed feature quantity; an acoustic signal conversion unit that calculates a plurality of separate feature quantities corresponding to each of the plurality of acoustic signals from the mixed feature quantity using the plurality of masks and converts the calculated plurality of separate feature quantities into a plurality of separate acoustic signals; an ambient sound interval estimation unit that estimates an ambient sound interval that includes only acoustic signals showing ambient sound in the entire input interval of the mixed acoustic signal based on the plurality of separate acoustic signals; an ambient sound signal extraction unit that extracts the mixed acoustic signal of the estimated ambient sound interval from the mixed acoustic signal as an ambient sound signal; and an ambient sound feature conversion unit that converts the ambient sound signal into an ambient sound feature quantity that shows the characteristics of the ambient sound signal. The mask estimation unit weights the mixed feature quantity using the ambient sound feature quantity and estimates the plurality of masks based on the weighted mixed feature quantity.

[0007] According to this disclosure, the complicated preparation process required to create auxiliary information regarding the acoustic signal of the target sound source in advance is eliminated, and a decrease in the performance of separating multiple acoustic signals from a mixed acoustic signal can be prevented. [Brief explanation of the drawing]

[0008] [Figure 1] This is a block diagram showing the configuration of a signal processing device in an embodiment of the present disclosure. [Figure 2] This is a block diagram showing the configuration of the learning device in an embodiment of the present disclosure. [Figure 3] This is a flowchart illustrating the sound source separation process of the signal processing device in this embodiment. [Figure 4] This is a flowchart illustrating the learning process of the learning device in this embodiment. [Modes for carrying out the invention]

[0009] (Knowledge that forms the basis of this disclosure) In the conventional technology described above, when sound source separation is performed using auxiliary information of the target sound source, it is necessary to record the audio of the target sound source in advance and generate auxiliary information from the recorded audio of the target sound source. This may require complicated preparatory processing to create auxiliary information regarding the acoustic signal of the target sound source in advance.

[0010] Furthermore, in the conventional techniques described above, when blind source separation is performed, if noise (ambient sound) that was not used to train the neural network model is included in the mixed acoustic signal, the performance of separating multiple acoustic signals from the mixed acoustic signal may decrease.

[0011] To address the above challenges, the following technologies are disclosed.

[0012] (1) A signal processing device according to one aspect of the present disclosure includes: a mixed acoustic signal acquisition unit that acquires a mixed acoustic signal including a plurality of acoustic signals; a mixed feature conversion unit that converts the mixed acoustic signal into a mixed feature quantity that shows the characteristics of the mixed acoustic signal; a mask estimation unit that estimates a plurality of masks corresponding to each of the plurality of acoustic signals based on the mixed feature quantity; an acoustic signal conversion unit that calculates a plurality of separate feature quantities corresponding to each of the plurality of acoustic signals from the mixed feature quantity using the plurality of masks and converts the calculated plurality of separate feature quantities into a plurality of separate acoustic signals; an ambient sound interval estimation unit that estimates an ambient sound interval that includes only acoustic signals showing ambient sound in the entire input interval of the mixed acoustic signal based on the plurality of separate acoustic signals; an ambient sound signal extraction unit that extracts the mixed acoustic signal of the estimated ambient sound interval from the mixed acoustic signal as an ambient sound signal; and an ambient sound feature conversion unit that converts the ambient sound signal into an ambient sound feature quantity that shows the characteristics of the ambient sound signal, wherein the mask estimation unit weights the mixed feature quantity using the ambient sound feature quantity and estimates the plurality of masks based on the weighted mixed feature quantity.

[0013] In this configuration, a mixed acoustic signal containing only ambient sound signals is extracted from the mixed acoustic signal as an ambient sound signal. The mixed features are weighted using ambient sound features that represent the characteristics of the ambient sound signal, and multiple masks are estimated based on the weighted mixed features. Therefore, multiple masks are estimated in real time using the ambient sound signals extracted from the mixed acoustic signal, and the mixed acoustic signal is separated into multiple separate acoustic signals using the estimated masks. This eliminates the need for the cumbersome preparatory processing required to create auxiliary information about the acoustic signals of the target sound source in advance, as in conventional techniques, and prevents a decrease in the performance of separating multiple acoustic signals from the mixed acoustic signal.

[0014] (2) In the signal processing apparatus described in (1) above, the mixed feature conversion unit includes a first acoustic model that outputs the mixed feature when the mixed acoustic signal is input, the mask estimation unit includes a second acoustic model that outputs the plurality of masks when the mixed feature is input, the acoustic signal conversion unit includes a third acoustic model that outputs the plurality of separated acoustic signals when the calculated plurality of separated feature quantities are input, and the ambient sound feature conversion unit may include a fourth acoustic model that outputs the ambient sound feature when the ambient acoustic signal is input.

[0015] In this configuration, a mixed acoustic signal is input to the first acoustic model, and mixed features are output from the first acoustic model. The mixed features are then input to the second acoustic model, and multiple masks are output from the second acoustic model. Furthermore, the calculated multiple separation features are input to the third acoustic model, and multiple separation acoustic signals are output from the third acoustic model. Finally, an ambient acoustic signal is input to the fourth acoustic model, and ambient sound features are output from the fourth acoustic model.

[0016] Therefore, mixed features can be easily estimated by the first acoustic model, multiple masks can be easily estimated by the second acoustic model, multiple separated acoustic signals can be easily estimated by the third acoustic model, and ambient sound features can be easily estimated by the fourth acoustic model.

[0017] (3) The signal processing apparatus described in (2) above further comprises: a learning acoustic signal acquisition unit that acquires a learning mixed acoustic signal and a plurality of correct acoustic signals corresponding to the correct answers of a plurality of acoustic signals included in the learning mixed acoustic signal; and a parameter update unit that updates the parameters of the first acoustic model, the second acoustic model, the third acoustic model, and the fourth acoustic model, wherein the mixed feature conversion unit inputs the learning mixed acoustic signal to the first acoustic model and acquires the mixed features output from the first acoustic model; the ambient sound feature conversion unit inputs a correct ambient acoustic signal indicating an ambient sound corresponding to the correct answer among the plurality of correct acoustic signals to the fourth acoustic model and acquires the ambient sound features output from the fourth acoustic model; and the mask estimation unit uses the ambient sound features output from the fourth acoustic model to perform the The acoustic signal conversion unit may weight the mixed features output from the first acoustic model, input the weighted mixed features into the second acoustic model, obtain the plurality of masks output from the second acoustic model, calculate a plurality of separation features corresponding to each of the plurality of ground truth acoustic signals from the mixed features using the plurality of masks output from the second acoustic model, input the calculated plurality of separation features into the third acoustic model, obtain the plurality of separation acoustic signals output from the third acoustic model, and the parameter update unit may calculate the error between each of the plurality of acoustic signals output from the third acoustic model and each of the plurality of ground truth acoustic signals, and update the parameters of the first acoustic model, second acoustic model, third acoustic model and fourth acoustic model based on the calculated plurality of errors.

[0018] According to this configuration, a learning mixed acoustic signal and a plurality of correct acoustic signals corresponding to the correct answers of the plurality of acoustic signals included in the learning mixed acoustic signal are obtained. The learning mixed acoustic signal is input into the first acoustic model, and a mixed feature amount is output from the first acoustic model. A correct environmental acoustic signal indicating environmental sound corresponding to the correct answer among the plurality of correct acoustic signals is input into the fourth acoustic model, and an environmental sound feature amount is output from the fourth acoustic model. The mixed feature amount output from the first acoustic model is weighted using the environmental sound feature amount output from the fourth acoustic model. The weighted mixed feature amount is input into the second acoustic model, and a plurality of masks are output from the second acoustic model. Using the plurality of masks output from the second acoustic model, a plurality of separated feature amounts corresponding to the plurality of correct acoustic signals are calculated from the mixed feature amount. The calculated plurality of separated feature amounts are input into the third acoustic model, and a plurality of separated acoustic signals are output from the third acoustic model. An error between each of the plurality of acoustic signals output from the third acoustic model and each of the plurality of correct acoustic signals is calculated. Based on the calculated plurality of errors, the parameters of the first acoustic model, the second acoustic model, the third acoustic model, and the fourth acoustic model are updated.

[0019] Therefore, the first acoustic model, the second acoustic model, the third acoustic model, and the fourth acoustic model can be learned using the learning mixed acoustic signal and the plurality of correct acoustic signals corresponding to the correct answers of the plurality of acoustic signals included in the learning mixed acoustic signal, and the estimation accuracy of the first acoustic model, the second acoustic model, the third acoustic model, and the fourth acoustic model can be improved.

[0020] (4) In the signal processing device according to any one of (1) to (3) above, the plurality of acoustic signals may include an acoustic signal indicating the environmental sound and an acoustic signal indicating a voice other than the environmental sound.

[0021] According to this configuration, an acoustic signal indicating environmental sound and an acoustic signal indicating a voice other than the environmental sound can be separated from the mixed acoustic signal.

[0022] (5) In the signal processing device according to (4) above, the voice other than the ambient sound may be a voice spoken by a person.

[0023] According to this configuration, an acoustic signal indicating ambient sound and an acoustic signal indicating a voice spoken by a person can be separated from the mixed acoustic signal.

[0024] (6) In the signal processing device according to (4) above, the voice other than the ambient sound may be a sound emitted by a specific object.

[0025] According to this configuration, an acoustic signal indicating ambient sound and an acoustic signal indicating a sound emitted by a specific object can be separated from the mixed acoustic signal.

[0026] (7) In the signal processing device according to any one of (1) to (6) above, the ambient acoustic signal extraction unit may store the extracted ambient acoustic signal in a memory, and the ambient sound feature quantity conversion unit may read the ambient acoustic signal from the memory and convert the read ambient acoustic signal into an ambient sound feature quantity.

[0027] According to this configuration, every time a mixed acoustic signal is acquired, the extracted ambient acoustic signal is stored in the memory, and an ambient sound feature quantity is generated using the ambient acoustic signal stored in the memory. Therefore, every time a mixed acoustic signal is acquired, a plurality of masks can be estimated in real time using the ambient sound feature quantity, and a plurality of separated acoustic signals can be accurately separated from the mixed acoustic signal using the plurality of masks.

[0028] (8) The signal processing device according to any one of (1) to (7) above may further include an acoustic signal output unit that outputs the plurality of separated acoustic signals converted by the acoustic signal conversion unit.

[0029] According to this configuration, since the plurality of separated acoustic signals that have been converted are output, signal processing such as speech recognition processing can be performed using the plurality of output separated acoustic signals.

[0030] Furthermore, this disclosure can be implemented not only as a signal processing device having the characteristic configuration described above, but also as a signal processing method that performs characteristic processing corresponding to the characteristic configuration of the signal processing device. It can also be implemented as a computer program that causes a computer to execute the characteristic processing included in such a signal processing method. Therefore, the same effects as the above-described signal processing device can be achieved in the following other embodiments.

[0031] (9) A signal processing method according to another aspect of the present disclosure involves a computer acquiring a mixed acoustic signal including a plurality of acoustic signals, converting the mixed acoustic signal into a mixed feature quantity that shows the characteristics of the mixed acoustic signal, estimating a plurality of masks corresponding to each of the plurality of acoustic signals based on the mixed feature quantity, calculating a plurality of separation feature quantities corresponding to each of the plurality of acoustic signals from the mixed feature quantity using the plurality of masks, converting the calculated plurality of separation feature quantities into a plurality of separation acoustic signals, estimating an ambient sound section that includes only acoustic signals showing ambient sounds in the entire input section of the mixed acoustic signal based on the plurality of separation acoustic signals, extracting the mixed acoustic signal of the estimated ambient sound section from the mixed acoustic signal as an ambient acoustic signal, converting the ambient acoustic signal into an ambient sound feature quantity that shows the characteristics of the ambient acoustic signal, weighting the mixed feature quantity using the ambient sound feature quantity in the estimation of the plurality of masks, and estimating the plurality of masks based on the weighted mixed feature quantity.

[0032] (10) A signal processing program according to another aspect of the present disclosure includes a mixed acoustic signal acquisition unit that acquires a mixed acoustic signal including a plurality of acoustic signals; a mixed feature conversion unit that converts the mixed acoustic signal into a mixed feature quantity that shows the characteristics of the mixed acoustic signal; a mask estimation unit that estimates a plurality of masks corresponding to each of the plurality of acoustic signals based on the mixed feature quantity; an acoustic signal conversion unit that calculates a plurality of separate feature quantities corresponding to each of the plurality of acoustic signals from the mixed feature quantity using the plurality of masks and converts the calculated plurality of separate feature quantities into a plurality of separate acoustic signals; an ambient sound interval estimation unit that estimates an ambient sound interval that includes only acoustic signals showing ambient sounds in the entire input interval of the mixed acoustic signal based on the plurality of separate acoustic signals; an ambient sound signal extraction unit that extracts the mixed acoustic signal of the estimated ambient sound interval from the mixed acoustic signal as an ambient sound signal; and an ambient sound feature conversion unit that converts the ambient acoustic signal into an ambient sound feature quantity that shows the characteristics of the ambient sound signal, wherein the mask estimation unit weights the mixed feature quantity using the ambient sound feature quantity and estimates the plurality of masks based on the weighted mixed feature quantity.

[0033] (11) A non-temporary computer-readable recording medium that records a signal processing program according to another aspect of the present disclosure includes: a mixed acoustic signal acquisition unit that acquires a mixed acoustic signal including a plurality of acoustic signals; a mixed feature conversion unit that converts the mixed acoustic signal into a mixed feature quantity that shows the characteristics of the mixed acoustic signal; a mask estimation unit that estimates a plurality of masks corresponding to each of the plurality of acoustic signals based on the mixed feature quantity; and an acoustic signal that calculates a plurality of separation feature quantities corresponding to each of the plurality of acoustic signals from the mixed feature quantity using the plurality of masks, and converts the calculated plurality of separation feature quantities into a plurality of separate acoustic signals. The computer functions as a number conversion unit, an ambient sound interval estimation unit that estimates an ambient sound interval that includes only the ambient sound signals representing ambient sound in the entire input interval of the mixed acoustic signal based on the plurality of separated acoustic signals, an ambient sound signal extraction unit that extracts the estimated mixed acoustic signal of the ambient sound interval from the mixed acoustic signal as an ambient sound signal, and an ambient sound feature quantity conversion unit that converts the ambient sound signal into ambient sound feature quantities that represent the characteristics of the ambient sound signal. The mask estimation unit weights the mixed feature quantities using the ambient sound feature quantities and estimates the plurality of masks based on the weighted mixed feature quantities.

[0034] Embodiments of this disclosure will be described below with reference to the attached drawings. Note that the following embodiments are merely examples of the disclosure and do not limit the technical scope of this disclosure.

[0035] (Embodiment) Figure 1 is a block diagram showing the configuration of the signal processing device 1 in an embodiment of the present disclosure.

[0036] The signal processing device 1 separates multiple acoustic signals from a mixed acoustic signal. The mixed acoustic signal contains multiple acoustic signals. These multiple acoustic signals include, for example, an acoustic signal representing ambient sound and an acoustic signal representing speech other than ambient sound. Speech other than ambient sound is, for example, a person's voice.

[0037] The signal processing device 1 shown in Figure 1 comprises a mixed acoustic signal acquisition unit 11, a mixed feature conversion unit 12, an ambient acoustic signal storage unit 13, an ambient sound feature conversion unit 14, a mask estimation unit 15, an acoustic signal conversion unit 16, an acoustic signal output unit 17, an ambient sound interval estimation unit 18, and an ambient acoustic signal extraction unit 19.

[0038] The mixed acoustic signal acquisition unit 11, the mixed feature conversion unit 12, the ambient sound feature conversion unit 14, the mask estimation unit 15, the acoustic signal conversion unit 16, the acoustic signal output unit 17, the ambient sound interval estimation unit 18, and the ambient acoustic signal extraction unit 19 are implemented by a processor. The processor consists of, for example, a CPU (Central Processing Unit).

[0039] The ambient acoustic signal storage unit 13 is implemented by memory. The memory consists of, for example, ROM (Read Only Memory) or EEPROM (Electrically Erasable Programmable Read Only Memory).

[0040] The signal processing device 1 may be, for example, a computer, smartphone, tablet computer, or server. Furthermore, the signal processing device 1 may be incorporated into other devices such as a car navigation system or home appliances.

[0041] The mixed acoustic signal acquisition unit 11 acquires a mixed acoustic signal that includes multiple acoustic signals. For example, the mixed acoustic signal includes a first acoustic signal that represents ambient sounds around a person and a second acoustic signal that represents a person's voice. The mixed acoustic signal acquisition unit 11 may be connected to a microphone (not shown). The microphone picks up sounds from multiple sound sources, converts them into acoustic signals, and outputs the converted acoustic signals as a mixed acoustic signal to the signal processing device 1. For example, the microphone picks up a person's voice and ambient sounds around the person. The mixed acoustic signal acquisition unit 11 acquires the mixed acoustic signal from the microphone.

[0042] Furthermore, the mixed acoustic signal acquisition unit 11 acquires the mixed acoustic signal for a predetermined period at predetermined intervals. For example, the mixed acoustic signal acquisition unit 11 may acquire a 10-second mixed acoustic signal every 10 seconds.

[0043] In this embodiment, the mixed acoustic signal acquisition unit 11 directly acquires the mixed acoustic signal picked up by the microphone from the microphone, but this disclosure is not particularly limited thereto. For example, the mixed acoustic signal picked up by the microphone or the like may be recorded on a computer-readable recording medium. The mixed acoustic signal acquisition unit 11 may acquire the mixed acoustic signal from a computer-readable recording medium. Examples of computer-readable recording media include semiconductor memory, hard disk drives, optical discs, or USB (Universal Serial Bus) memory. Furthermore, the mixed acoustic signal acquisition unit 11 may acquire the mixed acoustic signal from other devices via a network such as the Internet.

[0044] The mixed feature conversion unit 12 converts the mixed acoustic signal acquired by the mixed acoustic signal acquisition unit 11 into mixed features that represent the characteristics of the mixed acoustic signal. Mixed features are features that represent the mixed acoustic signal as a vector or matrix, for example, an embedding vector. The mixed feature conversion unit 12 includes a first acoustic model that outputs mixed features when a mixed acoustic signal is input. The first acoustic model is, for example, a convolutional neural network, a recurrent neural network, a long short-term memory network, or a deep neural network. The first acoustic model converts the input mixed acoustic signal into mixed features and outputs them. The first acoustic model is machine-learned by the learning device 2, which will be described later.

[0045] The mixed feature conversion unit 12 inputs the mixed acoustic signal to the first acoustic model and obtains the mixed features output from the first acoustic model. The mixed feature conversion unit 12 outputs the mixed features converted from the mixed acoustic signal to the mask estimation unit 15 and the acoustic signal conversion unit 16.

[0046] The ambient sound signal storage unit 13 stores the mixed sound signal of the ambient sound section, which contains only the sound signals indicating ambient sounds in the entire input section of the mixed sound signal, as the ambient sound signal. The ambient sound signal storage unit 13 temporarily stores the ambient sound signal. The ambient sound signals stored in the ambient sound signal storage unit 13 are updated at predetermined intervals.

[0047] The ambient sound feature conversion unit 14 converts the ambient sound signal into ambient sound features that represent the characteristics of the ambient sound signal. The ambient sound feature conversion unit 14 reads the ambient sound signal from the ambient sound signal storage unit 13 and converts the read ambient sound signal into ambient sound features. Ambient sound features are features that represent the ambient sound signal as a vector or matrix, for example, an embedding vector. The ambient sound feature conversion unit 14 includes a fourth acoustic model that outputs ambient sound features when an ambient sound signal is input. The fourth acoustic model is, for example, a convolutional neural network, a recurrent neural network, a long short-term memory network, or a deep neural network. The fourth acoustic model is machine-learned by the learning device 2, which will be described later.

[0048] The ambient sound feature conversion unit 14 inputs the ambient acoustic signal to the fourth acoustic model and obtains ambient sound features output from the fourth acoustic model. These ambient sound features correspond to auxiliary information. The ambient sound feature conversion unit 14 outputs the ambient sound features converted from the ambient acoustic signal to the mask estimation unit 15.

[0049] The mask estimation unit 15 estimates multiple masks corresponding to each of the multiple acoustic signals based on the mixed features converted by the mixed feature conversion unit 12. The mask estimation unit 15 includes a second acoustic model that outputs multiple masks when mixed features are input. The second acoustic model is, for example, a convolutional neural network, a recurrent neural network, a long short-term memory network, or a deep neural network. The second acoustic model is machine-learned by the learning device 2, which will be described later. The mask estimation unit 15 also weights the mixed features using the ambient sound features converted by the ambient sound feature conversion unit 14, and estimates multiple masks based on the weighted mixed features. The multiple masks are, for example, time-frequency masks.

[0050] The mask estimation unit 15 inputs a weighted mixed feature using ambient sound features into the second acoustic model and obtains multiple masks corresponding to each of the multiple acoustic signals output from the second acoustic model. The mask estimation unit 15 outputs the multiple masks estimated from the mixed feature to the acoustic signal conversion unit 16.

[0051] By weighting the mixed features with the ambient sound features, it is possible to accurately estimate masks for extracting acoustic signals representing ambient sounds and masks for extracting acoustic signals representing non-ambient sounds.

[0052] For example, if the mixed acoustic signal includes a first acoustic signal representing ambient sounds around a person and a second acoustic signal representing a person's voice, the mask estimation unit 15 estimates a first mask for extracting the first acoustic signal representing ambient sounds, and also estimates a second mask for extracting the second acoustic signal representing a person's voice, based on the mixed features converted by the mixed feature conversion unit 12.

[0053] The acoustic signal conversion unit 16 uses the mask estimation unit 15 to calculate multiple separate features corresponding to each of the multiple acoustic signals from the mixed features converted by the mixed feature conversion unit 12. Separate features are features that represent the acoustic signals included in the mixed acoustic signal as vectors or matrices, for example, embedding vectors.

[0054] The acoustic signal conversion unit 16 masks the mixed features using multiple masks estimated by the mask estimation unit 15 and calculates multiple separated features corresponding to each of the multiple acoustic signals.

[0055] Furthermore, the acoustic signal conversion unit 16 converts the calculated separation features into multiple separation acoustic signals. The acoustic signal conversion unit 16 includes a third acoustic model that outputs multiple separation acoustic signals when the calculated separation features are input. The third acoustic model is, for example, a convolutional neural network, a recurrent neural network, a long short-term memory network, or a deep neural network. The third acoustic model is machine-learned by the learning device 2, which will be described later.

[0056] The acoustic signal conversion unit 16 inputs the calculated separation features into the third acoustic model and obtains multiple separation acoustic signals output from the third acoustic model. The acoustic signal conversion unit 16 outputs the multiple separation acoustic signals converted from the multiple separation features to the acoustic signal output unit 17 and the ambient sound interval estimation unit 18.

[0057] For example, the acoustic signal conversion unit 16 uses the first mask estimated by the mask estimation unit 15 to calculate a first separated feature corresponding to the first acoustic signal from the mixed feature, and uses the second mask estimated by the mask estimation unit 15 to calculate a second separated feature corresponding to the second acoustic signal from the mixed feature. The acoustic signal conversion unit 16 calculates a first separated feature corresponding to the first acoustic signal by multiplying the mixed feature and the first mask in each time-frequency component, and calculates a second separated feature corresponding to the second acoustic signal by multiplying the mixed feature and the second mask in each time-frequency component. Furthermore, the acoustic signal conversion unit 16 converts the calculated first separated feature into a first separated acoustic signal, and converts the calculated second separated feature into a second separated acoustic signal.

[0058] The acoustic signal output unit 17 outputs multiple separated acoustic signals converted by the acoustic signal conversion unit 16. The acoustic signal output unit 17 outputs multiple separated acoustic signals separated from the mixed acoustic signal. The acoustic signal output unit 17 may output all of the multiple separated acoustic signals, or it may output only some of the multiple separated acoustic signals.

[0059] For example, the acoustic signal output unit 17 outputs a first separated acoustic signal representing ambient sound and a second separated acoustic signal representing human voice, which have been converted by the acoustic signal conversion unit 16. By separating ambient sound and human voice, ambient sound such as factory noise, in-car noise, or outside-car noise can be removed from the input mixed acoustic signal, and only human voice can be extracted. The second separated acoustic signal representing human voice can be used, for example, for speech recognition. The first separated acoustic signal representing ambient sound can be used, for example, to detect events occurring around a person. The acoustic signal output unit 17 may output both the first separated acoustic signal and the second separated acoustic signal, or it may output either the first separated acoustic signal or the second separated acoustic signal.

[0060] The ambient sound interval estimation unit 18 estimates an ambient sound interval that contains only the acoustic signals representing ambient sounds in the entire input interval of the mixed acoustic signal, based on a plurality of separated acoustic signals converted by the acoustic signal conversion unit 16. For example, the ambient sound interval estimation unit 18 estimates an ambient sound interval that contains only the acoustic signals representing ambient sounds in the entire input interval of the mixed acoustic signal by subtracting the interval of a second separated acoustic signal representing a human voice from the interval of a first separated acoustic signal representing ambient sounds.

[0061] Furthermore, the ambient sound interval estimation unit 18 may, through voice activity detection (VAD) processing, identify voice segments containing human voices and non-voice segments containing sounds other than human voices from the entire input section of each of the multiple acoustic signals, and estimate the intervals consisting only of non-voice segments that do not overlap with voice segments as ambient sound intervals. For example, the ambient sound interval estimation unit 18 may, through VAD processing, identify voice segments and non-voice segments from the entire input section of a first separated acoustic signal representing ambient sound, and also identify voice segments and non-voice segments from the entire input section of a second separated acoustic signal representing human voices. Then, the ambient sound interval estimation unit 18 may estimate the intervals consisting only of non-voice segments that do not overlap with voice segments from the entire input section of the mixed acoustic signal as ambient sound intervals.

[0062] The ambient sound signal extraction unit 19 extracts the mixed sound signal of the ambient sound interval estimated by the ambient sound interval estimation unit 18 from the mixed sound signal as the ambient sound signal. The ambient sound signal extraction unit 19 stores the extracted ambient sound signal in the ambient sound signal storage unit 13. The ambient sound signal extraction unit 19 stores the ambient sound signal in the ambient sound signal storage unit 13 at predetermined intervals and updates the ambient sound signal in the ambient sound signal storage unit 13. The predetermined interval is the interval at which the mixed sound signal is acquired.

[0063] In this way, ambient sound signals are stored in the ambient sound signal storage unit 13 at predetermined intervals, and the ambient sound signals stored in the ambient sound signal storage unit 13 are converted into ambient sound feature quantities that represent the characteristics of the ambient sound signals. These converted ambient sound feature quantities are then used to estimate multiple masks. Therefore, multiple acoustic signals can be separated from a mixed acoustic signal using ambient sound that changes in real time.

[0064] Next, the configuration of the learning device 2 in the embodiment of this disclosure will be described.

[0065] Figure 2 is a block diagram showing the configuration of the learning device 2 in an embodiment of the present disclosure.

[0066] The learning device 2 learns the parameters of each acoustic model (e.g., a neural network) of the mixed feature conversion unit 12, the ambient sound feature conversion unit 14, the mask estimation unit 15, and the acoustic signal conversion unit 16.

[0067] The learning device 2 shown in Figure 2 comprises a learning acoustic signal acquisition unit 21, a mixed feature conversion unit 12, an ambient sound feature conversion unit 14, a mask estimation unit 15, an acoustic signal conversion unit 16, and a parameter update unit 22. In the learning device 2, components identical to those in the signal processing device 1 are denoted by the same reference numerals, and their descriptions are omitted.

[0068] The learning acoustic signal acquisition unit 21, the mixed feature conversion unit 12, the ambient sound feature conversion unit 14, the mask estimation unit 15, the acoustic signal conversion unit 16, and the parameter update unit 22 are implemented by a processor. The processor consists of, for example, a CPU.

[0069] The learning device 2 may be, for example, a computer or a server. In this embodiment, the signal processing device 1 and the learning device 2 are different devices, but the signal processing device 1 may include the learning acoustic signal acquisition unit 21 and parameter update unit 22 of the learning device 2. In other words, the signal processing device 1 may have the functions of the learning device 2.

[0070] The learning acoustic signal acquisition unit 21 acquires a learning mixed acoustic signal and multiple correct acoustic signals that correspond to the correct answers of multiple acoustic signals included in the learning mixed acoustic signal. The learning acoustic signal acquisition unit 21 outputs the multiple correct acoustic signals to the parameter update unit 22, outputs the learning mixed acoustic signal to the mixed feature conversion unit 12, and outputs the correct ambient acoustic signal, which represents the correct ambient sound among the multiple correct acoustic signals, to the ambient sound feature conversion unit 14.

[0071] The learning acoustic signal acquisition unit 21 may be connected to a microphone (not shown). The microphone individually captures sounds from multiple sound sources, converts each into an acoustic signal, and outputs each converted acoustic signal to the signal processing device 1 as a correct acoustic signal. For example, the microphone individually captures a person's voice and ambient sounds. The microphone also captures a sound that is a mixture of multiple correct acoustic signals and multiple identical sounds, converts it into an acoustic signal, and outputs the converted acoustic signal to the signal processing device 1 as a learning mixed acoustic signal. The learning acoustic signal acquisition unit 21 acquires the learning mixed acoustic signal and multiple correct acoustic signals from the microphone. The learning acoustic signal acquisition unit 21 also uses the learning mixed acoustic signal and multiple correct acoustic signals as a single training data set and acquires multiple training data sets.

[0072] In this embodiment, the learning acoustic signal acquisition unit 21 directly acquires the learning mixed acoustic signal and multiple correct acoustic signals picked up by the microphone from the microphone, but this disclosure is not particularly limited thereto. For example, the learning mixed acoustic signal and multiple correct acoustic signals picked up by the microphone or the like may be recorded on a computer-readable recording medium. The learning acoustic signal acquisition unit 21 may acquire the learning mixed acoustic signal and multiple correct acoustic signals from a computer-readable recording medium. Furthermore, the learning acoustic signal acquisition unit 21 may acquire the learning mixed acoustic signal and multiple correct acoustic signals from other devices via a network such as the Internet.

[0073] The parameter update unit 22 updates the parameters of the first acoustic model, the second acoustic model, the third acoustic model, and the fourth acoustic model.

[0074] The mixed feature conversion unit 12 converts the training mixed acoustic signal acquired by the training acoustic signal acquisition unit 21 into mixed features that represent the characteristics of the training mixed acoustic signal. The mixed feature conversion unit 12 inputs the training mixed acoustic signal acquired by the training acoustic signal acquisition unit 21 into the first acoustic model and acquires the mixed features output from the first acoustic model.

[0075] The ambient sound feature conversion unit 14 converts the ground truth ambient sound signal, which represents the correct ambient sound among the multiple ground truth acoustic signals acquired by the learning acoustic signal acquisition unit 21, into ambient sound features that represent the characteristics of the ground truth ambient sound signal. The ambient sound feature conversion unit 14 inputs the ground truth ambient sound signal, which represents the correct ambient sound among the multiple ground truth acoustic signals acquired by the learning acoustic signal acquisition unit 21, into the fourth acoustic model and acquires the ambient sound features output from the fourth acoustic model.

[0076] The mask estimation unit 15 weights the mixed features using the ambient sound features converted by the ambient sound feature conversion unit 14, and estimates multiple masks corresponding to each of the multiple ground truth acoustic signals based on the weighted mixed features. The mask estimation unit 15 weights the mixed features output from the first acoustic model using the ambient sound features output from the fourth acoustic model, inputs the weighted mixed features into the second acoustic model, and obtains multiple masks output from the second acoustic model.

[0077] The acoustic signal conversion unit 16 uses multiple masks output from the second acoustic model to calculate multiple separate features corresponding to each of the multiple ground truth acoustic signals from the mixed features. The acoustic signal conversion unit 16 masks the mixed features using multiple masks estimated by the mask estimation unit 15 to calculate multiple separate features corresponding to each of the multiple ground truth acoustic signals. The acoustic signal conversion unit 16 also converts the calculated multiple separate features into multiple separate acoustic signals. The acoustic signal conversion unit 16 inputs the calculated multiple separate features into the third acoustic model and obtains multiple separate acoustic signals output from the third acoustic model.

[0078] The parameter update unit 22 calculates the error between each of the multiple separated acoustic signals output from the third acoustic model and each of the multiple ground truth acoustic signals acquired by the learning acoustic signal acquisition unit 21, and updates the parameters of the first acoustic model of the mixed feature conversion unit 12, the second acoustic model of the mask estimation unit 15, the third acoustic model of the acoustic signal conversion unit 16, and the fourth acoustic model of the ambient sound feature conversion unit 14 based on the calculated multiple errors. The parameter update unit 22 updates the parameters of the first acoustic model, second acoustic model, third acoustic model, and fourth acoustic model using backpropagation. More specifically, the parameter update unit 22 calculates the average error between each of the multiple separated acoustic signals output from the third acoustic model and each of the multiple ground truth acoustic signals, and updates the parameters of the first acoustic model, second acoustic model, third acoustic model, and fourth acoustic model so that the average of the calculated multiple errors is minimized.

[0079] Each part of the learning device 2 processes multiple training data, thereby repeatedly updating the parameters of the first acoustic model, second acoustic model, third acoustic model, and fourth acoustic model, and the first acoustic model, second acoustic model, third acoustic model, and fourth acoustic model are learned.

[0080] The mixed feature conversion unit 12, which includes a pre-trained first acoustic model; the mask estimation unit 15, which includes a pre-trained second acoustic model; the acoustic signal conversion unit 16, which includes a pre-trained third acoustic model; and the ambient sound feature conversion unit 14, which includes a pre-trained fourth acoustic model, are mounted on the signal processing device 1.

[0081] Next, the sound source separation process of the signal processing device 1 in this embodiment will be described.

[0082] Figure 3 is a flowchart illustrating the sound source separation process of the signal processing device 1 in this embodiment.

[0083] First, in step S1, the mixed acoustic signal acquisition unit 11 acquires a mixed acoustic signal that includes multiple acoustic signals. For example, the mixed acoustic signal includes a first acoustic signal that represents ambient sounds around a person and a second acoustic signal that represents a person's voice. Note that the second acoustic signal may represent not only the voice of one person but also the voices of multiple people.

[0084] Next, in step S2, the mixed feature conversion unit 12 converts the mixed acoustic signal acquired by the mixed acoustic signal acquisition unit 11 into mixed features that represent the characteristics of the mixed acoustic signal. At this time, the mixed feature conversion unit 12 inputs the mixed acoustic signal into the trained first acoustic model and acquires the mixed features output from the first acoustic model.

[0085] Next, in step S3, the ambient sound feature conversion unit 14 reads out the ambient sound signal storage unit 13, which represents only ambient sounds.

[0086] Next, in step S4, the ambient sound feature conversion unit 14 converts the ambient sound signal read from the ambient sound signal storage unit 13 into ambient sound features that represent the characteristics of the ambient sound signal. At this time, the ambient sound feature conversion unit 14 inputs the ambient sound signal to the trained fourth acoustic model and obtains the ambient sound features output from the fourth acoustic model.

[0087] Next, in step S5, the mask estimation unit 15 weights the mixed features using the ambient sound features converted by the ambient sound feature conversion unit 14.

[0088] Next, in step S6, the mask estimation unit 15 estimates multiple masks corresponding to each of the multiple acoustic signals based on a mixed feature weighted using ambient sound features. At this time, the mask estimation unit 15 inputs the mixed feature weighted using ambient sound features into a trained second acoustic model and obtains multiple masks corresponding to each of the multiple acoustic signals output from the second acoustic model. For example, the mask estimation unit 15 inputs the mixed feature weighted using ambient sound features into a trained second acoustic model and obtains a first mask corresponding to the first acoustic signal and a second mask corresponding to the second acoustic signal output from the second acoustic model.

[0089] In the initial sound source separation process, the ambient sound signal is not stored in the ambient sound signal storage unit 13, and the mask estimation unit 15 cannot weight the mixed features using the ambient sound features. Therefore, in the initial sound source separation process, the mask estimation unit 15 may estimate multiple masks corresponding to each of the multiple sound signals based on the mixed features converted by the mixed feature conversion unit 12, without weighting them using the ambient sound features. Then, in subsequent sound source separation processes, the mask estimation unit 15 may estimate multiple masks corresponding to each of the multiple sound signals based on the mixed features weighted using the ambient sound features.

[0090] Next, in step S7, the acoustic signal conversion unit 16 uses the mask estimation unit 15 to calculate multiple separation features corresponding to each of the multiple acoustic signals from the mixed features converted by the mixed feature conversion unit 12. At this time, the acoustic signal conversion unit 16 calculates multiple separation features corresponding to each of the multiple acoustic signals by multiplying the mixed features converted by the mixed feature conversion unit 12 and each of the multiple masks estimated by the mask estimation unit 15 in each time-frequency component. For example, the acoustic signal conversion unit 16 calculates a first separation feature corresponding to the first acoustic signal by multiplying the mixed features converted by the mixed feature conversion unit 12 and the first mask estimated by the mask estimation unit 15 in each time-frequency component, and calculates a second separation feature corresponding to the second acoustic signal by multiplying the mixed features converted by the mixed feature conversion unit 12 and the second mask estimated by the mask estimation unit 15 in each time-frequency component.

[0091] Next, in step S8, the acoustic signal conversion unit 16 converts the calculated separation features into multiple separated acoustic signals. At this time, the acoustic signal conversion unit 16 inputs the calculated separation features into a trained third acoustic model and obtains the multiple separated acoustic signals output from the third acoustic model. For example, the acoustic signal conversion unit 16 inputs the calculated first separation feature into a trained third acoustic model and obtains the first separated acoustic signal output from the third acoustic model, and also inputs the calculated second separation feature into a trained third acoustic model and obtains the second separated acoustic signal output from the third acoustic model.

[0092] Next, in step S9, the acoustic signal output unit 17 outputs a plurality of separated acoustic signals converted by the acoustic signal conversion unit 16. For example, the acoustic signal output unit 17 outputs a first separated acoustic signal and a second separated acoustic signal converted by the acoustic signal conversion unit 16.

[0093] Next, in step S10, the ambient sound interval estimation unit 18 estimates an ambient sound interval that contains only the acoustic signals indicating ambient sounds in the entire input section of the mixed acoustic signal, based on the multiple separated acoustic signals converted by the acoustic signal conversion unit 16. For example, the ambient sound interval estimation unit 18 estimates an ambient sound interval that contains only the acoustic signals indicating ambient sounds in the entire input section of the mixed acoustic signal, based on the first separated acoustic signal and the second separated acoustic signal converted by the acoustic signal conversion unit 16.

[0094] Next, in step S11, the ambient sound signal extraction unit 19 extracts the mixed sound signal of the ambient sound interval estimated by the ambient sound interval estimation unit 18 from the mixed sound signal acquired by the mixed sound signal acquisition unit 11 as the ambient sound signal.

[0095] Next, in step S12, the ambient sound signal extraction unit 19 stores the extracted ambient sound signal in the ambient sound signal storage unit 13. When the processing in step S12 is completed, the process returns to step S1.

[0096] In this way, from the mixed acoustic signal, the mixed acoustic signal of the ambient sound section containing only the acoustic signals representing ambient sound is extracted as the ambient acoustic signal, the mixed features are weighted using ambient sound features that represent the characteristics of the ambient acoustic signal, and multiple masks are estimated based on the weighted mixed features. Therefore, since multiple masks are estimated in real time using the ambient acoustic signal extracted from the mixed acoustic signal, and the mixed acoustic signal is separated into multiple separated acoustic signals using the estimated multiple masks, the complicated preparatory processing required to create auxiliary information about the acoustic signal of the target sound source in advance, as in conventional techniques, is eliminated, and a decrease in the performance of separating multiple acoustic signals from the mixed acoustic signal can be prevented.

[0097] Furthermore, by estimating ambient noise and using ambient noise features as auxiliary information to represent the characteristics of that ambient noise, it is possible to accurately separate sound sources while adapting each acoustic model to the usage environment in real time.

[0098] Next, the learning process of the learning device 2 in this embodiment will be described.

[0099] Figure 4 is a flowchart illustrating the learning process of the learning device 2 in this embodiment.

[0100] First, in step S21, the learning acoustic signal acquisition unit 21 acquires a learning mixed acoustic signal and a plurality of correct acoustic signals. For example, the plurality of correct acoustic signals include a first correct acoustic signal that represents ambient sounds around a person and a second correct acoustic signal that represents a human voice.

[0101] Next, in step S22, the mixed feature conversion unit 12 converts the training mixed acoustic signal acquired by the training acoustic signal acquisition unit 21 into mixed features that represent the characteristics of the training mixed acoustic signal. At this time, the mixed feature conversion unit 12 inputs the training mixed acoustic signal acquired by the training acoustic signal acquisition unit 21 into the untrained first acoustic model and acquires the mixed features output from the first acoustic model.

[0102] Next, in step S23, the ambient sound feature conversion unit 14 converts the correct ambient sound signal, which represents the correct ambient sound among the multiple correct acoustic signals acquired by the learning acoustic signal acquisition unit 21, into ambient sound features that represent the characteristics of the correct ambient sound signal. At this time, the ambient sound feature conversion unit 14 inputs the correct ambient sound signal, which represents the correct ambient sound signal among the multiple correct acoustic signals acquired by the learning acoustic signal acquisition unit 21, into an untrained fourth acoustic model and acquires ambient sound features output from the fourth acoustic model.

[0103] Next, in step S24, the mask estimation unit 15 weights the mixed features using the ambient sound features converted by the ambient sound feature conversion unit 14.

[0104] Next, in step S25, the mask estimation unit 15 estimates multiple masks corresponding to each of the multiple ground truth acoustic signals based on a mixed feature weighted using ambient sound features. At this time, the mask estimation unit 15 inputs the mixed feature weighted using ambient sound features into an untrained second acoustic model and obtains multiple masks corresponding to each of the multiple ground truth acoustic signals output from the second acoustic model. For example, the mask estimation unit 15 inputs the mixed feature weighted using ambient sound features into an untrained second acoustic model and obtains a first mask corresponding to the first ground truth acoustic signal and a second mask corresponding to the second ground truth acoustic signal output from the second acoustic model.

[0105] Next, in step S26, the acoustic signal conversion unit 16 uses the mask estimation unit 15 to calculate multiple separation features corresponding to each of the multiple ground truth acoustic signals from the mixed features converted by the mixed feature conversion unit 12. At this time, the acoustic signal conversion unit 16 calculates multiple separation features corresponding to each of the multiple ground truth acoustic signals by multiplying the mixed features converted by the mixed feature conversion unit 12 and each of the multiple masks estimated by the mask estimation unit 15 in each time-frequency component. For example, the acoustic signal conversion unit 16 calculates a first separation feature corresponding to the first ground truth acoustic signal by multiplying the mixed features converted by the mixed feature conversion unit 12 and the first mask estimated by the mask estimation unit 15 in each time-frequency component, and calculates a second separation feature corresponding to the second ground truth acoustic signal by multiplying the mixed features converted by the mixed feature conversion unit 12 and the second mask estimated by the mask estimation unit 15 in each time-frequency component.

[0106] Next, in step S27, the acoustic signal conversion unit 16 converts the calculated separation features into multiple separated acoustic signals. At this time, the acoustic signal conversion unit 16 inputs the calculated separation features into an untrained third acoustic model and obtains multiple separated acoustic signals output from the third acoustic model. For example, the acoustic signal conversion unit 16 inputs the calculated first separation feature into an untrained third acoustic model and obtains the first separated acoustic signal output from the third acoustic model, and also inputs the calculated second separation feature into an untrained third acoustic model and obtains the second separated acoustic signal output from the third acoustic model.

[0107] Next, in step S28, the parameter update unit 22 calculates the error between each of the multiple separated acoustic signals output from the third acoustic model and each of the multiple correct acoustic signals acquired by the learning acoustic signal acquisition unit 21. For example, the parameter update unit 22 calculates the error between the first separated acoustic signal output from the third acoustic model and the first correct acoustic signal, and also calculates the error between the second separated acoustic signal output from the third acoustic model and the second correct acoustic signal.

[0108] Next, in step S29, the parameter update unit 22 calculates the average of the multiple errors that have been calculated. For example, the parameter update unit 22 calculates the average of the error between the first separated sound signal and the first correct sound signal, and the error between the second separated sound signal and the second correct sound signal.

[0109] Next, in step S30, the parameter update unit 22 updates the parameters of the first acoustic model of the mixed feature conversion unit 12, the second acoustic model of the mask estimation unit 15, the third acoustic model of the acoustic signal conversion unit 16, and the fourth acoustic model of the ambient sound feature conversion unit 14 so that the average of the multiple calculated errors is minimized.

[0110] Each training data set includes a mixed acoustic signal for learning and multiple correct acoustic signals. The learning acoustic signal acquisition unit 21 acquires one of the training data sets. Then, steps S21 to S30 are performed for all of the training data sets, and the first acoustic model, second acoustic model, third acoustic model, and fourth acoustic model are trained.

[0111] In this way, a training mixed acoustic signal and multiple ground truth acoustic signals corresponding to the correct answers of the multiple acoustic signals contained in the training mixed acoustic signal are obtained. The training mixed acoustic signal is input to the first acoustic model, and a mixed feature is output from the first acoustic model. A ground truth ambient acoustic signal, which represents the correct ambient sound among the multiple ground truth acoustic signals, is input to the fourth acoustic model, and an ambient sound feature is output from the fourth acoustic model. The mixed feature output from the first acoustic model is weighted using the ambient sound feature output from the fourth acoustic model. The weighted mixed feature is input to the second acoustic model, and multiple masks are output from the second acoustic model. Separation features corresponding to each of the multiple acoustic signals are calculated from the mixed feature using the multiple masks output from the second acoustic model. The calculated separation features are input to the third acoustic model, and multiple separated acoustic signals are output from the third acoustic model. The error between each of the multiple acoustic signals output from the third acoustic model and each of the multiple ground truth acoustic signals is calculated. Based on the multiple errors calculated, the parameters of the first, second, third, and fourth acoustic models are updated.

[0112] Therefore, the first, second, third, and fourth acoustic models can be trained using a training mixed acoustic signal and multiple ground truth acoustic signals corresponding to the correct answers of multiple acoustic signals included in the training mixed acoustic signal, thereby improving the estimation accuracy of the first, second, third, and fourth acoustic models.

[0113] In this embodiment, sounds other than ambient sounds may be sounds emitted by specific objects. Sounds emitted by specific objects may be, for example, the sound of a siren from a police vehicle, fire truck, or ambulance. The learning device 2 learns the first to fourth acoustic models using a mixed learning acoustic signal obtained by mixing an acoustic signal indicating the sound of a siren with an acoustic signal indicating ambient sounds other than the sound of a siren, so that the signal processing device 1 can separate and output the sound of a siren from ambient sounds other than the sound of a siren.

[0114] Although the embodiments described above explain the case where the multiple masks are time-frequency masks, this disclosure is not limited to this. For example, the multiple masks may be vectors that show the contribution of each element of the mixed feature to each acoustic signal.

[0115] In each of the above embodiments, each component may be implemented by dedicated hardware or by executing a software program suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory. Furthermore, the program may be executed by another independent computer system by recording and transferring the program to a recording medium, or by transferring the program via a network.

[0116] Some or all of the functions of the apparatus according to the embodiments of this disclosure are typically implemented as an integrated circuit, or LSI (Large Scale Integration). These may be individually integrated onto a single chip, or some or all of them may be integrated onto a single chip. Furthermore, the integration is not limited to LSIs, but may also be implemented using dedicated circuits or general-purpose processors. An FPGA (Field Programmable Gate Array) that can be programmed after LSI manufacturing, or a reconfigurable processor that can reconfigure the connections and settings of circuit cells inside the LSI may also be used.

[0117] Furthermore, some or all of the functions of the apparatus according to the embodiments of this disclosure may be realized by a processor such as a CPU executing a program.

[0118] Furthermore, all figures used above are illustrative examples provided to illustrate this disclosure, and this disclosure is not limited to these illustrative figures.

[0119] Furthermore, the order in which the steps shown in the flowchart above are performed is illustrative for the purpose of specifically illustrating this disclosure, and other orders are acceptable as long as similar effects are achieved. Also, some of the steps above may be performed simultaneously (in parallel) with other steps. [Industrial applicability]

[0120] The technology described herein is useful as a technology for separating multiple acoustic signals from a mixed acoustic signal because it eliminates the need for complicated preparatory processing to create auxiliary information regarding the acoustic signal of the target sound source in advance, and prevents a decrease in performance when separating multiple acoustic signals from a mixed acoustic signal.

Claims

1. A mixed acoustic signal acquisition unit that acquires a mixed acoustic signal containing multiple acoustic signals, A mixed feature quantity conversion unit converts the mixed acoustic signal into a mixed feature quantity that represents the characteristics of the mixed acoustic signal, A mask estimation unit that estimates a plurality of masks corresponding to each of the plurality of acoustic signals based on the aforementioned mixed features, An acoustic signal conversion unit that uses the aforementioned multiple masks to calculate multiple separate feature quantities corresponding to each of the multiple acoustic signals from the mixed feature quantity, and converts the calculated multiple separate feature quantities into multiple separate acoustic signals, An ambient sound interval estimation unit estimates an ambient sound interval that contains only the ambient sound signals representing ambient sound in the entire input interval of the mixed ambient sound signal, based on the plurality of separated ambient sound signals. An ambient sound signal extraction unit extracts the estimated ambient sound section of the mixed acoustic signal from the mixed acoustic signal as an ambient sound signal, An ambient sound feature quantity conversion unit converts the ambient sound signal into ambient sound feature quantities that represent the characteristics of the ambient sound signal, Equipped with, The mask estimation unit weights the mixed feature quantities using the ambient sound feature quantities and estimates the plurality of masks based on the weighted mixed feature quantities. Signal processing device.

2. The mixed feature conversion unit includes a first acoustic model that outputs the mixed feature when the mixed acoustic signal is input, The mask estimation unit includes a second acoustic model that outputs the plurality of masks when the mixed feature quantity is input, The acoustic signal conversion unit includes a third acoustic model that outputs the plurality of separated acoustic signals when the calculated plurality of separated feature quantities are input. The ambient sound feature conversion unit includes a fourth acoustic model that outputs the ambient sound feature when the ambient sound signal is input. The signal processing apparatus according to claim 1.

3. A learning acoustic signal acquisition unit that acquires a learning mixed acoustic signal and a plurality of correct acoustic signals corresponding to the correct answers of a plurality of acoustic signals included in the learning mixed acoustic signal, A parameter update unit that updates the parameters of the first acoustic model, the second acoustic model, the third acoustic model, and the fourth acoustic model, Furthermore, The mixed feature conversion unit inputs the training mixed acoustic signal to the first acoustic model and acquires the mixed feature output from the first acoustic model. The ambient sound feature conversion unit inputs a correct ambient sound signal, which represents the correct ambient sound among the plurality of correct acoustic signals, to the fourth acoustic model, and acquires the ambient sound feature output from the fourth acoustic model. The mask estimation unit weights the mixed feature output from the first acoustic model using the ambient sound feature output from the fourth acoustic model, inputs the weighted mixed feature to the second acoustic model, and obtains the plurality of masks output from the second acoustic model. The acoustic signal conversion unit calculates a plurality of separation features corresponding to each of the plurality of ground truth acoustic signals from the mixed features using the plurality of masks output from the second acoustic model, inputs the calculated plurality of separation features to the third acoustic model, and acquires the plurality of separation acoustic signals output from the third acoustic model. The parameter update unit calculates the error between each of the plurality of acoustic signals output from the third acoustic model and each of the plurality of correct acoustic signals, and updates the parameters of the first acoustic model, second acoustic model, third acoustic model and fourth acoustic model based on the calculated plurality of errors. The signal processing apparatus according to claim 2.

4. The plurality of acoustic signals include an acoustic signal representing the ambient sound and an acoustic signal representing a sound other than the ambient sound. The signal processing apparatus according to any one of claims 1 to 3.

5. Other than the aforementioned ambient sounds, the aforementioned sounds are human speech. The signal processing apparatus according to claim 4.

6. Other sounds besides the aforementioned ambient sounds are sounds emitted by a specific object. The signal processing apparatus according to claim 4.

7. The aforementioned ambient sound signal extraction unit stores the extracted ambient sound signal in memory. The ambient sound feature conversion unit reads the ambient sound signal from the memory and converts the read ambient sound signal into ambient sound features. The signal processing apparatus according to any one of claims 1 to 3.

8. The system further includes an acoustic signal output unit that outputs the plurality of separated acoustic signals converted by the acoustic signal conversion unit. The signal processing apparatus according to any one of claims 1 to 3.

9. Computers A mixed acoustic signal containing multiple acoustic signals is acquired, The mixed acoustic signal is converted into a mixed feature quantity that represents the characteristics of the mixed acoustic signal, Based on the aforementioned mixed features, multiple masks corresponding to each of the multiple acoustic signals are estimated. Using the aforementioned multiple masks, multiple separation features corresponding to each of the multiple acoustic signals are calculated from the mixed feature quantity, and the calculated multiple separation features are converted into multiple separate acoustic signals. Based on the aforementioned multiple separated acoustic signals, an ambient sound section is estimated that contains only acoustic signals representing ambient sound in the entire input section of the mixed acoustic signal. From the aforementioned mixed acoustic signal, the estimated mixed acoustic signal of the ambient sound section is extracted as the ambient sound signal. The aforementioned ambient sound signal is converted into an ambient sound feature quantity that represents the characteristics of the ambient sound signal, In estimating the plurality of masks, the mixed features are weighted using the ambient sound features, and the plurality of masks are estimated based on the weighted mixed features. Signal processing method.

10. A mixed acoustic signal acquisition unit that acquires a mixed acoustic signal containing multiple acoustic signals, A mixed feature quantity conversion unit converts the mixed acoustic signal into a mixed feature quantity that represents the characteristics of the mixed acoustic signal, A mask estimation unit that estimates a plurality of masks corresponding to each of the plurality of acoustic signals based on the aforementioned mixed features, An acoustic signal conversion unit that uses the aforementioned multiple masks to calculate multiple separate feature quantities corresponding to each of the multiple acoustic signals from the mixed feature quantity, and converts the calculated multiple separate feature quantities into multiple separate acoustic signals, An ambient sound interval estimation unit estimates an ambient sound interval that contains only the ambient sound signals representing ambient sound in the entire input interval of the mixed ambient sound signal, based on the plurality of separated ambient sound signals. An ambient sound signal extraction unit extracts the estimated ambient sound section of the mixed acoustic signal from the mixed acoustic signal as an ambient sound signal, The computer functions as an ambient sound feature conversion unit that converts the ambient sound signal into ambient sound feature quantities that represent the characteristics of the ambient sound signal. The mask estimation unit weights the mixed feature quantities using the ambient sound feature quantities and estimates the plurality of masks based on the weighted mixed feature quantities. Signal processing program.

Citation Information

Patent Citations

  • Voice signal processing device, television receiver, voice signal processing method, program and recording medium

    JP2011141540A

  • Signal processing device, learning device, signal processing method, learning method and program

    JP2020134657A