Acoustic treatment method

WO2026181254A1PCT designated stage Publication Date: 2026-09-03NT T INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/007084
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2026-09-03

Smart Images

  • Figure JP2025007084_03092026_PF_FP_ABST
    Figure JP2025007084_03092026_PF_FP_ABST
Patent Text Reader

Abstract

In this acoustic treatment method, an acoustic event is extracted from an input acoustic signal representing a mixed sound, and an acoustic signal corresponding to the acoustic event is extracted from the input acoustic signal.
Need to check novelty before this filing date? Find Prior Art

Description

Sound processing method

[0001] This invention relates to acoustic processing technology, and more particularly to technology for extracting acoustic events and acoustic signals.

[0002] Techniques are known for separating the acoustic signals of a specified class of acoustic events from an input acoustic signal that represents a mixed sound composed of sounds corresponding to multiple types of acoustic events (see, for example, Non-Patent Document 1).

[0003] I. Kavalerov et al., "Universal Sound Separation," 2019 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), New Paltz, NY, USA, 2019, pp. 175-179, doi: 10.1109 / WASPAA. 2019. 8937253.

[0004] However, conventional techniques cannot extract acoustic events from input acoustic signals. Therefore, if it is unknown which acoustic events correspond to the input acoustic signal, it is necessary to specify all classes of acoustic events by brute force and separate the acoustic signals, which is computationally inefficient.

[0005] This invention provides a technology that can efficiently extract the acoustic signal corresponding to the acoustic event contained in the input acoustic signal, even when it is unknown what kind of acoustic event the input acoustic signal corresponds to.

[0006] In this invention, acoustic events are extracted from an input acoustic signal representing a mixed sound, and acoustic signals corresponding to those acoustic events are extracted from the input acoustic signal.

[0007] This allows for the extraction of the corresponding acoustic signal from the acoustic event contained within the input acoustic signal, even when it is unknown what acoustic event the input acoustic signal corresponds to.

[0008] Figure 1 is a block diagram illustrating the functional configuration of the acoustic processing system of the first to third embodiments. Figure 2 is a block diagram illustrating the functional configuration of the acoustic processing device of the first embodiment. Figure 3 is a block diagram illustrating the functional configuration of the machine learning device of the first embodiment. Figure 4 is a block diagram illustrating the functional configuration of the machine learning device of the second embodiment. Figure 5 is a block diagram illustrating the functional configuration of the machine learning device of the third embodiment. Figure 6 is a block diagram illustrating the functional configuration of the cost function calculation unit of Figure 5. Figure 7 is a block diagram illustrating the functional configuration of the acoustic processing system of the fourth embodiment. Figure 8 is a block diagram illustrating the functional configuration of the acoustic processing device of the fourth embodiment. Figure 9 is a block diagram illustrating the functional configuration of the machine learning device of the fourth embodiment. Figure 10 is a block diagram illustrating the hardware configuration of the embodiment.

[0009] Embodiments of the present invention will be described below with reference to the drawings. [Principle] First, the principle of each embodiment will be described. In each embodiment, the apparatus extracts an acoustic event from an input acoustic signal representing a mixed sound picked up in an environment in which one type of acoustic event or multiple types of acoustic events are occurring, and extracts an acoustic signal corresponding to the acoustic event from the input acoustic signal. That is, in each embodiment, an acoustic event and an acoustic signal corresponding to the acoustic event are extracted from the input acoustic signal. Note that "extraction" may be rephrased as "detection" or "detection". The acoustic event to be extracted may be one or multiple. An acoustic event represents an individual sound event occurring in a certain environment. That is, an acoustic event is a meaningful sound phenomenon that is time-limited and distinguishable from ambient background sound. Examples of acoustic events include, for example, voices, car horns, the sound of doors opening and closing, dog barking, the sound of glass breaking, etc. The acoustic signal corresponding to the extracted acoustic event may be the original sound signal of the acoustic event (original signal), an acoustic signal obtained by convolving a room transfer function onto the original signal, or an acoustic signal based on other original sounds.

[0010] In each embodiment, the input acoustic signal x is modeled by the following equation (1). Here, x is an input acoustic signal based on acoustic signals observed by M microphones. The input acoustic signal x is a time-series discrete signal. When M is an integer of 2 or more, x is a multi-channel acoustic signal, and when M is 1, x is a monaural-channel acoustic signal. That is, x at each discrete time is represented by an M-dimensional vector. Note that a one-dimensional vector is a scalar (the same applies hereinafter). For example, x is a discrete signal obtained by digitizing an acoustic signal observed by a microphone. For example, the input acoustic signal x is an Ambisonics signal, an omnidirectional signal based on a signal collected by an omnidirectional microphone array, or the like. However, these do not limit the present invention. K represents the number of acoustic events corresponding to the input acoustic signal x. That is, the input acoustic signal x includes acoustic signals based on K acoustic events. Note that K is an integer of 1 or more, and for example, K is an integer of 2 or more. c k is a class representing the k-th (k=1,...,K) acoustic event. s(c k ) represents the original signal of the acoustic event of class c k . s(c k ) is a time-series discrete signal. h k represents the room transfer function from the position (sound source) of the k-th acoustic event to the position of the microphone. h k is represented by an M-dimensional vector. h k *s(c k ) represents the convolution of h k with s(c k ). Note that when M is an integer of 2 or more, h k *s(c k ) represents convolution for each element of h k with s(c k ). n represents a noise acoustic signal (noise signal). n may be a time-series discrete signal or may be a constant.

[0011] The target task in each embodiment is to obtain original signals s(c1),...,s(c KAlternatively, one could extract the (acoustic signal), or convolve the original signal with the room transfer function h1*s(c1),...,h k *s(c K It may also be possible to extract the (acoustic signal) from the original signal s(c1),...,s(c K This may involve extracting other acoustic signals based on ).

[0012] In each embodiment, acoustic events are extracted from the input acoustic signal x representing the mixed sound, and acoustic signals corresponding to acoustic events x are extracted from the input acoustic signal. Here, the methods described in each embodiment are categorized. <Inference Method 1> In inference method 1, the classes c1,...,c of acoustic events are extracted from the input acoustic signal x. K Then, after estimating the number K of existing acoustic events, classes c1, ..., c K For each of these, the original signals s(c1), ..., s(c K ), h1*s(c1),...,h k *s(c K ), or s(c1),...,s(c K Other acoustic signals are estimated based on the above. In other words, in inference method 1, acoustic events are extracted from the input acoustic signal x, and then acoustic signals corresponding to the acoustic events extracted from the input acoustic signal x are extracted.

[0013] In inference method 1, the acoustic event detection model M SED and acoustic signal separation model M USS The acoustic event detection model M is used. SED This involves the input acoustic signal x and the parameter θ. SED (Acoustic event detection parameters) and classes c1^,...,c K This is a model that estimates ^. The acoustic event detection model M is given by equation (2) below. SED This is an example: {c1^,...,c K ^}=M SED (x;θ SED (2) Here, classes c1^,...,c K ^ represents the estimated classes c1,...,c K This represents the acoustic signal separation model M. USSThis involves input acoustic signal x and class c1^,...,c K ^ and parameter θ USS (Acoustic signal separation parameters) From the original signals s^(c1^),...,s^(c K This is a model that estimates ^). The acoustic signal separation model M is given by equation (3) below. USS Let's take an example. {s^(c1^),...,s^(c K ^)}=M USS (x, c1^, ..., c K ^;θ USS ) (3) Here, the original signals s^(c1^),...,s^(c K ^) represents the estimated original signals s^(c1^), ..., s^(c K ^) represents. In other words, in inference method 1, the acoustic event detection model M SED Using this method, the class of acoustic events c1^,...,c is determined from the input acoustic signal x. K ^ After estimating the number of existing acoustic events K, an acoustic signal separation model M is created. USS Using this, class c1^,...,c K For each of ^, the original signal s^(c1^), ..., s^(c K We estimate ^). Then s^(c1^),...,s^(c K ^) from h1*s^(c1^),...,h k *s^(c K ^) may be obtained, or s^(c1^),...,s^(c K Other acoustic signals based on ^) may be obtained. Note that the acoustic event detection model M SED and acoustic signal separation model M USS This can be constructed, for example, using a neural network. For example, the acoustic event detection model M SED This can be constructed using a neural network that takes log-mel spectrograms as input acoustic features, and is an acoustic signal separation model M USS These can be configured as described in Non-Patent Document 1. However, these are not limiting to the present invention, and they may be configured in other ways.

[0014] <Inference Method 1, Machine Learning Method 1> In Inference Method 1, Machine Learning Method 1, the apparatus uses training data D = {x', c1',...,c R ', s'(c1'),...,s'(c R Using ')}, the acoustic event detection model M SED The parameter θ SED and acoustic signal separation model M USS The parameter θ USS and are learned independently. Here, x' is a learning acoustic signal x' representing a mixed sound for learning, and c1',...,c R ' is the class of acoustic events (ground truth label of acoustic events) corresponding to the training acoustic signal x, and s'(c1'),...,s'(c R ') is class c1',...,c R This is the original signal (event sound teacher signal) corresponding to '. R is an integer greater than or equal to 1, for example, R is an integer greater than or equal to 2. r=1,...,R, and c1',...,c R Each of ' is c r ' is written as, s'(c1'),...,s'(c R Each of ') is s'(c r It is written as '). x', c1',...,c R ', s'(c1'),...,s'(c R The form ') is, for example, the aforementioned x, c1,...,c K , s(c1),...,s(c K It is the same as ).

[0015] parameter θ SED is {c1~,...,c L ~}=M SED (x';θ SED ) (Inference result) and {c1',...,c R It is trained based on a cost function for '} (training data). Here, L is an integer greater than or equal to 1, for example, L is an integer greater than or equal to 2. c1~,...,c L The format is as follows: c1,...,c K It is the same as the parameter θ. SED is {c1~,...,c L~} and {c1',...,c R '}, learning can be performed by the error backpropagation method using binary cross entropy (BCE) loss or the like as a cost function (loss function). Note that the learning method is not limited thereto. For example, {c1~,...,c L ~} (inference result) and {c1',...,c R '} (training data), any learning method may be used as long as the parameter θ SED is learned such that the distance between the inference result and the training data becomes small. That is, the acoustic event detection model M L ~}=M SED (x';θ SED ) is machine-learned based on an evaluation for extraction of acoustic events {c1~,...,c SED corresponding to the training acoustic signal x'.

[0016] The parameter θ USS is learned based on a cost function for {s~(c1~),...,s~(c L ~)}=M USS (x',c1~,...,c L ~;θ USS ) (inference result) and {s'(c1'),...,s'(c R ')} (training data). The format of s~(c1~),...,s~(c L ~) is the same as the format of s(c1),...,s(c K ) described above. For example, the parameter θ USS is for {s'(c1'),...,s'(c R ')}, {s~(c1~),...,s~(c L ~)}=M USS (x',c1~,...,c L ~;θ USS ) can be learned by error backpropagation using a signal-to-distortion ratio (SDR), Scale-invariant SDR (SI-SDR), or the like as a cost function (loss function). Note that the learning method is not limited thereto. For example, {s~(c1~),...,s~(c L ~)} (inference result) and {s'(c1'),...,s'(cR The parameter θ is set so that the distance to the training data becomes smaller. USS Any learning method may be used as long as the learning acoustic signal ' corresponds to the acoustic event {c1~,...,c L ~} corresponds to the acoustic signal {s~(c1~),...,s~(c L Based on the evaluation of the extraction of ~)}, the acoustic signal separation model M USS This is then subjected to machine learning.

[0017] Here, the acoustic event detection model M SED The appropriate acoustic event {c1~,...,c L Acoustic signal separation model M when ~} cannot be obtained USS When learning is performed, inappropriate acoustic events (labels) {c1~,...,c L Acoustic signal separation model M based on ~} USS This will involve training the acoustic event detection model M. SED Learn and adapt the acoustic event detection model M SED After obtaining the parameter θ, SED Fixed, acoustic signal separation model M USS It is desirable to study this.

[0018] <Inference Method 1, Machine Learning Method 2> As described above, parameter θ SED and parameter θ USS By sequentially optimizing the acoustic event detection model M SED and acoustic signal separation model M USS It can learn this. However, the optimized parameter θ SED It is unclear whether this is optimal for the target task. The optimal parameter θ for the target task. SED and parameter θ USS To find this, it is desirable to optimize these simultaneously. Machine learning method 2 of inference method 1 has parameter θ SED and parameter θ USS The parameters θ are optimized simultaneously. SED and parameter θ USSTo simultaneously optimize the acoustic event detection model M, we introduce a new cost function called Permutation Sensitive Signal-to-Distortion Ratio improvement (PS-SDRi). PS-SDRi is used in acoustic event detection model M SED and acoustic signal separation model M USS This serves as a criterion for simultaneously optimizing both. PS-SDRi can be expressed, for example, by the following equation (4). Here, x' W This is the omnidirectional component of the acoustic signal extracted from the learning acoustic signal x'. For example, if x' is an ambisonic signal, its 0th-order component is x'. W It can be used as follows. SDR(a,b) means the SDR of b for a. As shown in equation (4), the estimated L acoustic events are of class c1~,...,c L ~ contains the correct sound event class (correct class) c r If '(r∈{1,...,R}) is not included (Otherwise), then PS-SDRi is 0. That is, the correct answer classes c1',...,c R If a class of an acoustic event not present in ' is extracted (False positive case), or if the correct class is c1',...,c R If the class of the acoustic event within ' is not extracted (False Negative case), PS-SDRi is set to 0. In other words, unlike regular SDRi, PS-SDRi is designed to penalize failures in acoustic event detection. Also, even if the estimated class c1~,...,c L ~ contains the correct class c r Even if ' is included, c r s~(c r ~) (Inference result) is c r 'corresponding to s'(c r If the result differs significantly from the training data, PS-SDRi will be a low value, resulting in a penalty. Therefore, PS-SDRi is an index that can simultaneously evaluate the extraction of acoustic events and the extraction of acoustic signals corresponding to acoustic events, and is useful for acoustic event detection models MSED and acoustic signal separation model M USS This can be used as a cost function to simultaneously optimize the acoustic event detection model M so that PS-SDRi becomes larger. SED The parameter θ SED and acoustic signal separation model M USS The parameter θ USS It can perform machine learning on both simultaneously. For example, {c1~,...,c L ~}=M SED (x';θ SED ) and {s~(c1~),...,s~(c L ~)}=M USS (x',c1~,...,c L ~;θ USS ) (Inference result) and {c1',...,c R '} and {s'(c1'),...,s'(c R The cost function is PS-SDRi in equation (4) for ')} (training data), and the parameter θ is set such that PS-SDRi is large (for example, maximized). SED and parameter θ USS The two can be learned simultaneously through machine learning. For example, the reciprocal of PS-SDRi can be used as the loss function, or PS-SDRi can be used as the reward function. In other words, the acoustic event detection model M can be developed based on a criterion that simultaneously evaluates the extraction of acoustic events corresponding to the training acoustic signal x' and the extraction of acoustic signals corresponding to the acoustic events corresponding to the training acoustic signal x'. SED and acoustic signal separation model M USS This can be done using machine learning. Acoustic event detection model M SED and acoustic signal separation model M USS This can be described as a machine learning model based on criteria that penalize failures in extracting acoustic events corresponding to the training acoustic signal x' and failures in extracting acoustic signals corresponding to acoustic events corresponding to the training acoustic signal x'. SED and acoustic signal separation model M USSThis can be described as a machine learning model based on an evaluation of the extraction of acoustic events corresponding to the training acoustic signal x', and an evaluation of the extraction of acoustic signals corresponding to the acoustic events corresponding to the training acoustic signal x'.

[0019] <Machine learning method 3 of inference method 1> In inference method 1, the original signals s^(c1^),...,s^(c K Instead of estimating ^), h1*s^(c1^),...,h from the input acoustic signal x. K *s^(c K ^) may be estimated. That is, instead of equation (3), the acoustic signal separation model M of equation (3a) below may be used. USS The following may be used: {h1*s^(c1^),...,h K *s^(c K ^)}=M USS (x, c1^, ..., c K ^;θ USS ) (3a) In this case, acoustic signal separation model M USS This process extracts an acoustic signal by convolving the original signal with a spatial transfer function, rather than the original signal itself.

[0020] In this case, instead of equation (4), the PS-SDRi in equation (5) below may be used as the cost function. In this case, for example, D = {x', c1',...,c R ', s'(c1'),...,s'(c R Replace ')} with {x', c1',...,c R ', h1*s'(c1'),...,h R *s'(c R The data ')} is used as training data. However, the spatial transfer functions h1,...,h R If the above is known, then D = {x', c1',...,c R ', s'(c1'),...,s'(c R ')} can be used as training data as is. In this case as well, the acoustic event detection model M should be set so that PS-SDRi is large. SED The parameter θ SED and acoustic signal separation model M USS The parameter θUSS It can perform machine learning on both simultaneously. For example, {c1~,...,c L ~}=M SED (x';θ SED ) and {h1*s~(c1~),...,h L *s~(c L ~)}=M USS (x',c1~,...,c L ~;θ USS ) and {c1',...,c R '} and {h1*s'(c1'),...,h R *s'(c R Let PS-SDRi in equation (5) for ')} be the cost function, and set the parameter θ such that PS-SDRi is large (for example, to the maximum). SED and parameter θ USS This can be done by simultaneously using machine learning. Note that the spatial transfer functions h1,...,h L If is known, then replace equation (3a) with the acoustic signal separation model M of equation (3) USS Using {h1*s~(c1~),...,h L *s~(c L It is acceptable if ~)} is calculated.

[0021] <Inference Method 1 Machine Learning Method 4> The PS-SDRi with intermediate conditions between equations (4) and (5) described above may also be used as the cost function. For example, {s~(c1~),...,s~(c L ~)}=M USS (x',c1~,...,c L ~;θ USS ) and the training data {s'(c1'),...,s'(c RA phase shift between ')} may be allowed. In other words, a PS-SDRi that is unaffected by the phase difference between these may be used as the cost function. That is, the acoustic signal corresponding to each acoustic event contained in the input acoustic signal x has a delay depending on the distance between the microphone and the sound source. However, if the task is to separate the input acoustic signal x into acoustic signals corresponding to each acoustic event, estimating the delay (phase difference) of the acoustic signals is not essential. Therefore, such a relaxation can be said to be a more realistic setting. Such a relaxation is expressed in equation (4) SDR(s'(c' r ), s~(c r This can be achieved by replacing ~)) with the following equation (6). τ = argmax τ R τ (s'(c' r ), s~(c r ~)) (7) α=argmin α ||αs'(c' r )[0:T]-s~(c r ~)[τ:T+τ]|| (8) Here, T is a positive integer representing a given time length. ||ν|| is the norm of ν (e.g., the L2 norm). [θ1:θ2] is a closed time interval between θ1 and θ2. That is, the numerator of equation (6) is αs'(c') in the closed time interval [0:T]. r The function represents the norm of the elements of ) and the denominator is αs'(c') in the closed time interval [0:T]. r ) elements and s~(c in the closed time interval [τ:T+τ] r R represents the norm of the difference between elements of ~). τ (ν1,ν2) represents the cross-correlation between ν1 and ν2. argmax τ η represents the τ that maximizes η, and argmin α η represents α that minimizes η. α is also related to the scale of the extracted acoustic signal, s'(c'). r This coefficient is introduced to allow for changes from the (correct signal) and select the appropriate one. Otherwise, it is the same as machine learning method 2 in inference method 1.

[0022] <Machine learning method 5 of inference method 1> In inference method 1, the original signals s^(c1^),...,s^(cK Instead of estimating ^), the direct sound and early reflection components may be estimated from the input acoustic signal x. That is, instead of equation (3), the acoustic signal separation model M of equation (3b) below may be used. USS This may also be used. {h1 e *s^(c1^),...,h K e *s^(c K ^)}=M USS (x, c1^, ..., c K ^;θ USS ) (3b) Here, h k e The indoor transfer function h k This represents the direct sound and early reflection components. In this case, the acoustic signal separation model M USS This process extracts an acoustic signal that is not the original signal, but rather the original signal convolved with the direct sound and early reflection components of the spatial transfer function.

[0023] In this case, instead of equation (4), the PS-SDRi in equation (9) below may be used as the cost function. In this case, for example, {x', c1',...,c R ', s'(c1'),...,s'(c R Replace ')} with {x', c1',...,c R ', h1 e *s'(c1'),...,h R e *s'(c R ')} is used as training data. However, h1 e ,...,h R e If the above is known, then D = {x', c1',...,c R ', s'(c1'),...,s'(c R ')} can be used as training data as is. In this case as well, the acoustic event detection model M should be set so that PS-SDRi is large. SED The parameter θ SED and acoustic signal separation model M USS The parameter θ USS It can perform machine learning on both simultaneously. For example, {c1~,...,c L ~}=MSED (x';θ SED ) and {h1 e *s~(c1~),...,h L e *s~(c L ~)}=M USS (x',c1~,...,c L ~;θ USS ) and {c1',...,c R '} and {h1 e *s'(c1'),...,h R e *s'(c R Let PS-SDRi in equation (9) for ')} be the cost function, and set the parameter θ such that PS-SDRi is large (for example, to the maximum). SED and parameter θ USS This can be done by simultaneously using machine learning. e ,...,h L e If is known, then replace equation (3a) with the acoustic signal separation model M of equation (3) USS Using {h1 e *s~(c1~),...,h L e *s~(c L You may also calculate}.

[0024] <Inference Method 1 Machine Learning Method 6> In the case of the Otherwise method in equation (4) above, that is, the correct answer classes c1',...,c R If a class of an acoustic event not present in ' is extracted (False positive case), or if the correct class is c1',...,c R If no class of acoustic event within ' is extracted (False Negative case), instead of setting PS-SDRi to 0, PS-SDRi may be set to a negative value. For example, instead of equation (4), the PS-SDRi in equation (10) below may be used as the cost function. Here, ε is a positive real number. There are no restrictions on ε, but for example, silence or noise and s~(c rThe SDR and constant values ​​between ~) can be denoted as ε. Alternatively, ε can be a real number, and the upper limit of PS-SDRi in the Otherwise case can be set to 0 so that PS-SDRi is always less than or equal to 0 in the Otherwise case. Equation (11) below shows an example of PS-SDRi in this case. Here, min(0,-ε) represents the smaller of 0 and -ε. Similarly, in the PS-SDRi of machine learning methods 3 to 5 of inference method 1, instead of setting PS-SDRi to 0 in the Otherwise case, PS-SDRi may be set to a negative value.

[0025] <Inference Method 1, Machine Learning Method 7> Using PS-SDRi as the cost function, the parameter θ is determined by methods such as gradient descent. SED and parameter θ USS When performing machine learning, the parameters θ of PS-SDRi SED and parameter θ USS The values ​​corresponding to the partial derivatives (gradient vectors and total derivatives) with respect to are used. Here, the SDR(s'(c) in equation (4) r '),x' W The term ) is composed solely of the training data, and the parameter θ SED and parameter θ USS It does not depend on SDR(s'(c r '),x' W ) parameter θ SED and parameter θ USS When the partial derivative is taken, it becomes zero, and does not affect machine learning. Therefore, the SDR(s'(c) in equation (4) r '),x' W The term ) may be any constant Z (see equation (12) below) or 0. Similarly, for inference method 1, machine learning methods 3 to 6, the SDR(s'(c r '),x' W The term in (12) may be any constant Z or 0. Also, in the PS-SDRi of equation (12), instead of setting PS-SDRi to 0 in the Otherwise case, PS-SDRi may be set to a negative value.

[0026] <Machine Learning Method 8 of Inference Method 1> The PS-SDRi exemplified in Machine Learning Methods 2 to 7 of Inference Method 1 described above cannot distinguish between false positive (FP) cases (when a class that does not exist in the correct class is estimated) and false negative (FN) cases (when a class that exists in the correct class is not estimated) and impose penalties accordingly. In contrast, a PS-SDRi that distinguishes between FP and FN and imposes penalties accordingly may be used as the cost function. For example, a PS-SDRi represented by the following equation (12a) may be used. Here, C' is the set of correct classes included in the training data C'={c1',...,c R '} (training data), and C~ is the set of acoustic events corresponding to the training acoustic signal x' C~={c1~,...,c L ~}=M SED (x';θ SED ) (Inference result). ξ represents a class ξ∈C'∪C~ contained in the union C'∪C~ of C' and C~. ε FN ,ε FP ε is a real number representing a penalty. For example, ε FN ,ε FP These are all distinct constants. For example, ε FN ,ε FP ε is a constant greater than or equal to 0. FN ,ε FP At least one of them may be a variable. For example, ε FN This could be the function value of s'(ξ) (training data), or ε FP This could also be the function value of s~(ξ) (the inference result). For example, ε FN ,ε FP ε may be set as shown in equations (12b) and (12c) below. FN =max(10log 10 ||s'(ξ)||,0) (12b) ε FP =max(10log 10||s~(ξ)||,0) (12c) By setting it in this way, a penalty can be imposed meticulously (for example, in dB units) as in the case of ξ∈C'∧ξ∈C~. Note that max(δ,0) represents the larger of δ and 0. That is, the ε exemplified in equations (12b) and (12c) FN ,ε FP Since is greater than or equal to 0, the penalty in equation (12a) is -ε FN and -ε FP All of these values ​​are negative. The rest are the same as machine learning methods 2 through 7 of inference method 1.

[0027] Furthermore, in all of the machine learning methods 1 to 8 of inference method 1, the acoustic event detection model M is based on a criterion that simultaneously evaluates the extraction of acoustic events corresponding to the training acoustic signal x' and the extraction of acoustic signals corresponding to the acoustic events corresponding to the training acoustic signal x'. SED and acoustic signal separation model M USS This can be done using machine learning. In this case, the acoustic event detection model M SED and acoustic signal separation model M USS This can be described as a machine learning model based on criteria that penalize failures in extracting acoustic events corresponding to the training acoustic signal x' and failures in extracting acoustic signals corresponding to acoustic events corresponding to the training acoustic signal x'. SED and acoustic signal separation model M USS This can be described as a machine learning model based on an evaluation of the extraction of acoustic events corresponding to the training acoustic signal x', and an evaluation of the extraction of acoustic signals corresponding to the acoustic events corresponding to the training acoustic signal x'.

[0028] <Inference Method 2> In inference method 2, the device uses multitasking to determine the class of acoustic events c1,...,c from the input acoustic signal x. K and the original signals s(c1), ..., s(c K ) and are estimated simultaneously. Original signals s(c1),...,s(c K Instead of ), h1*s(c1),...,h k *s(c K ) or s(c1), ..., s(cK Other acoustic signals may be estimated based on the above. In other words, in inference method 2, an acoustic event and the acoustic signal corresponding to that acoustic event are simultaneously extracted from the input acoustic signal x.

[0029] In inference method 2, a multitask model M is used. The multitask model M uses the input acoustic signal x and the parameter θ to determine classes c1^,...,c K ^ and the original signals s^(c1^), ..., s^(c K This is a model that simultaneously estimates {c1^,...,c K ^,s^(c1^),...,s^(c K ^)}=M(x;θ) (13)

[0030] <Inference Method 2, Machine Learning Method 1> Acoustic Event Detection Model M SED and acoustic signal separation model M USS This is replaced by a multitasking model M, with the parameter θ SED and parameter θ USS {c1~,...,c L ~}=M SED (x';θ SED ) and {s~(c1~),...,s~(c L ~)}=M USS (x',c1~,...,c L ~;θ USS ) is {c1~,...,c L ~, s~(c1~),...,s~(c L Except for the substitution with} = M(x';θ), it is the same as machine learning method 2 for inference method 1.

[0031] <Inference Method 2 Machine Learning Method 2> In inference method 2, the original signals s^(c1^),...,s^(c K Instead of estimating ^), h1*s^(c1^),...,h from the input acoustic signal x. k *s^(c K ^) may be estimated. That is, instead of equation (13), the multitask model M of equation (3a) below may be used. {c1^,..., cK ^, h1*s^(c1^),...,h K *s^(c K ^)}=M(x;θ) (13a) In this case, the multitask model M extracts not the original signal and the class, but the acoustic signal obtained by convolving the spatial transfer function onto the original signal and the class.

[0032] For learning the parameters θ in this case, for example, {x', c1',...,c R ', s'(c1'),...,s'(c R Replace ')} with {x', c1',...,c R ', h1*s'(c1'),...,h R *s'(c R The data ')} is used as training data. However, the spatial transfer functions h1,...,h R If the above is known, then D = {x', c1',...,c R ', s'(c1'),...,s'(c R The data ')} can be used directly as training data. In this case as well, the parameters θ can be trained so that PS-SDRi in equation (5) is large. That is, the acoustic event detection model M SED and acoustic signal separation model M USS This is replaced by a multitasking model M, with the parameter θ SED and parameter θ USS {c1~,...,c L ~}=M SED (x';θ SED ) and {s~(c1~),...,s~(c L ~)}=M USS (x',c1~,...,c L ~;θ USS ) is {c1~,...,c L ~, h1*s~(c1~),...,h L *s~(c L Except for the substitution with} = M(x';θ), it is the same as machine learning method 3 of inference method 1. Note that the spatial transfer functions h1,...,h LIf is known, then instead of equation (13a), use the multitask model M in equation (13) to get h1*s'(c1'),...,h R *s'(c R ') may be calculated.

[0033] <Machine learning method 3 of inference method 2> Similar to machine learning method 4 of inference method 1, phase shift may be allowed. That is, SDR(s'(c') in equation (4) r ), s~(c r You may also replace ~)) with equation (6).

[0034] <Machine Learning Method 4 of Inference Method 2> Similar to Machine Learning Method 5 of Inference Method 1, direct sound and early reflection components may be estimated from the input acoustic signal x. Instead of equation (13), the multitask model M of equation (13b) below may be used. {c1^,..., c K ^, h1 e *s^(c1^),...,h K e *s^(c K ^)}=M(x;θ) (13b)

[0035] In this case, instead of equation (4), PS-SDRi in equation (9) may be used as the cost function. In this case, for example, {x', c1',...,c R ', s'(c1'),...,s'(c R Replace ')} with {x', c1',...,c R ', h1 e *s'(c1'),...,h R e *s'(c R ')} is used as training data. However, h1 e ,...,h R e If the above is known, then D = {x', c1',...,c R ', s'(c1'),...,s'(c R The data ')} can be used directly as training data. In this case as well, the parameters θ can be trained so that PS-SDRi is large. That is, the acoustic event detection model M SED and acoustic signal separation model MUSS This is replaced by a multitasking model M, with the parameter θ SED and parameter θ USS {c1~,...,c L ~}=M SED (x';θ SED ) and {h1 e *s~(c1~),...,h R e *s~(c L ~)}=M USS (x',c1~,...,c L ~;θ USS ) is {c1~,...,c L ~, h1 e *s~(c1~),...,h R e *s~(c L Except for the substitution with} = M(x';θ), it is the same as machine learning method 5 of inference method 1. Note that h1 e ,...,h L e If is known, then instead of equation (13a), use the multitask model M in equation (13) to h1 e *s~(c1~),...,h R e *s~(c L ~) may be calculated.

[0036] <Inference Method 2, Machine Learning Method 5> In the case of PS-SDRi, instead of setting PS-SDRi to 0 in the Otherwise case, PS-SDRi may be set to a negative value.

[0037] <Inference Method 2 Machine Learning Method 6> PS-SDRi's SDR(s'(c r '),x' W The term ) may be any constant Z or 0.

[0038] <Machine Learning Method 7 for Inference Method 2> Similar to Machine Learning Method 8 for Inference Method 1, PS-SDRi, which distinguishes between FP and FN and imposes penalties accordingly, may also be used as the cost function in Inference Method 2.

[0039] Furthermore, in any of the machine learning methods 1 to 7 of inference method 2, a multi-task model M can be machine-learned based on a criterion that simultaneously evaluates the extraction of acoustic events corresponding to the training acoustic signal x' and the extraction of acoustic signals corresponding to the acoustic events corresponding to the training acoustic signal x'. In this case, the multi-task model M can be said to be a model that has been machine-learned based on a criterion that imposes penalties for failures in extracting acoustic events corresponding to the training acoustic signal x' and for failures in extracting acoustic signals corresponding to the acoustic events corresponding to the training acoustic signal x'.

[0040] [First Embodiment] Next, the first embodiment will be described. In this embodiment, inference is performed using the inference method 1 described above. Acoustic event detection model M SED and acoustic signal separation model M USS For learning, the machine learning method 1 of the inference method 1 described above is used.

[0041] <Configuration> As illustrated in Figure 1, the acoustic processing system 1 of this embodiment includes an acoustic processing device 11 that performs inference processing and a machine learning device 12 that performs learning processing. The acoustic processing system 1 receives, for example, a parameter θ from the machine learning device 12 to the acoustic processing device 11 via a network (e.g., the Internet). SED and parameter θ USS It is configured to enable the provision (e.g., transmission) of data.

[0042] <Sound Processing Device 11> As illustrated in Figure 2, the sound processing device 11 of this embodiment includes a control unit 110, storage units 111 and 112, an acoustic event detection unit 113, an acoustic signal separation unit 114, and a memory 115. Although not explained below, the processing of the sound processing device 11 is performed under the control of the control unit 110. Information input to the sound processing device 11 and information obtained by each unit are stored in the memory 115 one by one and read out and used as needed.

[0043] <Machine Learning Device 12> As illustrated in Figure 3, the machine learning device 12 in this embodiment includes a control unit 120, storage units 121 and 122, an acoustic event detection unit 123, an acoustic signal separation unit 124, an acoustic event detection cost function calculation unit 125, parameter update units 126 and 128, an acoustic signal separation cost function calculation unit 127, and a memory 129. Although not explained below, the processing of the machine learning device 12 is performed under the control of the control unit 120. Information input to the machine learning device 12 and information obtained by each unit are stored in the memory 129 one by one and read out and used as needed.

[0044] <Inference Processing (Inference Method 1)> The acoustic processing device 11 (Figure 2) extracts acoustic events from the input acoustic signal x representing the mixed sound, and extracts the acoustic signal corresponding to the acoustic event from the input acoustic signal. This will be explained in detail below.

[0045] First, as a preprocessing step, an acoustic event detection model M obtained by machine learning in the machine learning device 12 is used. SED The parameter θ SED and acoustic signal separation model M USS The parameter θ USS The signal is sent to the sound processing device 11. Parameter θ SED The parameters θ are stored in the storage unit 111 of the sound processing device 11 (Figure 2). USS This is stored in the memory unit 112. This preprocessing is performed initially and as needed, thereby storing the parameter θ in the memory unit 111. SED or the parameter θ stored in the memory unit 112 USS It may be updated.

[0046] Based on the above preprocessing, the aforementioned input acoustic signal x is input to the acoustic processing device 11 (Equation (1)). The input acoustic signal x is sent to the acoustic event detection unit 113 and the acoustic signal separation unit 114 (Step S111).

[0047] The acoustic event detection unit 113 processes the input acoustic signal x and the parameter θ extracted from the storage unit 111. SED Using this, and according to equation (2) above, the estimated acoustic event classes c1^,...,c K^ is obtained and output (extracts an acoustic event from the input acoustic signal x). Class c1^,...,c K ^ also includes information about K, which represents the number of acoustic events. In other words, the number of existing acoustic events itself is also the target of estimation. Output classes c1^,...,c K The signal is sent to the acoustic signal separation unit 114 (step S112).

[0048] The acoustic signal separation unit 114 separates the received input acoustic signal x from class c1^,...,c K ^ and the parameter θ extracted from the memory unit 112 USS Using this, and following equation (3) above, the estimated original signals s^(c1^),...,s^(c K ^) (acoustic signal) is obtained and output (acoustic signal corresponding to the acoustic event is extracted from the input acoustic signal x). That is, the acoustic signal separation unit 114 obtains s^(c1^),...,s^(c) corresponding to the K acoustic events estimated by the acoustic event detection unit 113. K ^) is obtained and output (step S113).

[0049] The acoustic processing device 11 uses the number of acoustic events estimated as described above, K, such as s^(c1^),...,s^(c K ^) (separated sound, acoustic signal). Alternatively, the acoustic processing device 11 outputs h1*s^(c1^),...,h k *s^(c K ^), or s^(c1^),...,s^(c K Other acoustic signals based on ^) may be output (step S114).

[0050] <Learning Process (Machine Learning Method 1 of Inference Method 1)> First, the parameter θ SED The initial value of the parameter θ is stored in the memory unit 121 of the machine learning device 12 (Figure 3), USS The initial value is stored in the storage unit 122 (step S121).

[0051] The machine learning device 12 has the aforementioned training data D = {x', c1',...,c R ', s'(c1'),...,s'(c RThe input is ')}. Training data D = {x', c1',...,c R ', s'(c1'),...,s'(c R ')} may be stored in a memory unit (not shown) or sent via a network. The learning acoustic signal x' is sent to the acoustic event detection unit 123 and the acoustic signal separation unit 124, and classes c1',...,c R ' is sent to the acoustic event detection cost function calculation unit 125, and the original signals s'(c1'),...,s'(c R The value of ') is sent to the acoustic signal separation cost function calculation unit 127 (step S122).

[0052] The acoustic event detection unit 123 uses the received learning acoustic signal x' and the parameter θ extracted from the storage unit 121. SED Using this, and following equation (2) above, the estimated classes {c1~,...,c L ~}=M SED (x';θ SED The inference result is obtained and output. Estimated class {c1~,...,c L The data {c1~,...,c} is sent to the acoustic signal separation unit 124 and the acoustic event detection cost function calculation unit 125. The acoustic event detection cost function calculation unit 125 calculates {c1~,...,c L ~} (inference result) and {c1',...,c R The cost function for the training data is calculated and the cost function value cost1 is output. An example of a cost function is the BCE loss. The cost function value cost1 is sent to the parameter update unit 126. The parameter update unit 126 uses the cost function value cost1 to update the parameters θ stored in the memory unit 121. SED Update the value. For this process, for example, a backpropagation method using the cost function value cost1 can be used (step S123).

[0053] The control unit 120 controls the parameter θ SED This determines whether the update process has met the termination condition. There are no limitations on the termination condition, but for example, the parameter θ in step S123. SED The update range fell below the standard, and the parameter θ SEDThe termination condition can be that the number of updates reaches a certain threshold (step S124). If it is determined that the termination condition has not been met, the process returns to step S122. On the other hand, if it is determined that the termination condition has been met, the process proceeds to the next step S125.

[0054] In step S125, the parameter θ SED With the parameter θ fixed USS The update is performed. First, the acoustic signal separation unit 124 separates the learning acoustic signal x' and the class {c1~,...,c L ~} and the parameter θ extracted from the memory unit 122 USS Using this, and according to equation (3) above, the estimated original signal {s~(c1~),...,s~(c L ~)}=M USS (x',c1~,...,c L ~;θ USS The estimated original signal {s~(c1~),...,s~(c L The {s~(c1~),...,s~(c L ~)}=M USS (x',c1~,...,c L ~;θ USS ) (Inference result) and {s'(c1'),...,s'(c R The cost function for the training data (')} is calculated and the cost function value cost2 is output. Examples of cost functions include SDR and SI-SDR. The cost function value cost2 is sent to the parameter update unit 128. The parameter update unit 128 uses the cost function value cost2 to update the parameters θ stored in the memory unit 122. USS Update the value. For this process, for example, a backpropagation method using the cost function value cost2 can be used (step S125).

[0055] The control unit 120 controls the parameter θ USS This determines whether the update process has met the termination condition. There are no limitations on the termination condition, but for example, the parameter θ in step S125. USSThe update range fell below the standard, and the parameter θ USS The termination condition can be that the number of updates reaches a certain threshold (step S126). If it is determined that the termination condition is not met, the process returns to step S122, and the parameter θ USS The update process will continue. However, the parameter θ in this form USS The update process is performed using parameter θ. SED This is done with the parameter θ fixed. That is, in step S123, after returning from step S126 to step S122, the parameter θ SED The estimated classes {c1~,...,c L ~} is sent to the acoustic signal separation unit 124. On the other hand, if it is determined that the termination condition is met, the parameter θ SED and θ USS The learning process is completed. The rmeter θ obtained as described above is finished. SED and θ USS As mentioned above, it is sent to the sound processing device 11.

[0056] [Second Embodiment] In the second embodiment, inference is also performed using the inference method 1 described above. Acoustic event detection model M SED and acoustic signal separation model M USS For learning, we will use machine learning method 2 of inference method 1, as described above. Below, we will focus on explaining the differences from what has been explained so far, and for things that have already been explained, we will cite the same reference numbers to simplify the explanation.

[0057] <Configuration> As illustrated in Figure 1, the acoustic processing system 2 of this embodiment includes an acoustic processing device 11 that performs inference processing and a machine learning device 22 that performs learning processing. The acoustic processing system 2 receives, for example, a parameter θ via a network from the machine learning device 22 to the acoustic processing device 11. SED and parameter θ USS It is configured to be available.

[0058] <Sound processing device 11> This is the same as in the first embodiment.

[0059] <Machine Learning Device 22> As illustrated in Figure 4, the machine learning device 22 in this embodiment includes a control unit 120, storage units 121 and 122, an acoustic event detection unit 123, an acoustic signal separation unit 124, a simultaneous optimization cost function calculation unit 225, parameter update units 226 and 228, and a memory 129. Although not explained below, the processing of the machine learning device 22 is performed under the control of the control unit 120. Information input to the machine learning device 22 and information obtained by each unit are stored in the memory 129 one by one and read out and used as needed.

[0060] <Inference Processing (Inference Method 1)> This is the same as in the first embodiment.

[0061] <Learning Process (Inference Method 1, Machine Learning Method 2)> First, the parameter θ SED The initial value of the parameter θ is stored in the memory unit 121 of the machine learning device 22 (Figure 3), USS The initial value is stored in the storage unit 122 (step S221).

[0062] The machine learning device 22 has the aforementioned training data D = {x', c1',...,c R ', s'(c1'),...,s'(c R The input is ')}. The learning acoustic signal x' is sent to the acoustic event detection unit 123 and the acoustic signal separation unit 124, and classes c1',...,c R 'and original signals s'(c1'),...,s'(c R The result is sent to the simultaneous optimization cost function calculation unit 225 (step S222).

[0063] The acoustic event detection unit 123 uses the received learning acoustic signal x' and the parameter θ extracted from the storage unit 121. SED Using this, and following equation (2) above, the estimated classes {c1~,...,c L ~}=M SED (x';θ SED The inference result is obtained and output. Estimated class {c1~,...,c L The ~} is sent to the acoustic signal separation unit 124 and the simultaneous optimization cost function calculation unit 225 (step S223).

[0064] The acoustic signal separation unit 124 separates the learning acoustic signal x' and the class {c1~,...,c L ~} and the parameter θ extracted from the memory unit 122 USS Using this, and according to equation (3) above, the estimated original signal {s~(c1~),...,s~(c L ~)}=M USS (x',c1~,...,c L ~;θ USS The estimated original signal {s~(c1~),...,s~(c L The result is sent to the simultaneous optimization cost function calculation unit 225 (step S224).

[0065] The simultaneous optimization cost function calculation unit 225 calculates {c1~,...,c L ~} and {s~(c1~),...,s~(c L ~)} (inference result) and {c1',...,c R '} and {s'(c1'),...,s'(c R For the training data, calculate PS-SDRi in equation (4) as the cost function, and the cost function value cost 12 Outputs the cost function value cost. 12 This is sent to the parameter update units 226 and 228. The parameter update unit 226 processes the cost function value cost 12 The parameter θ stored in the memory unit 121 is used SED The parameter update unit 228 updates the cost function value cost 12 The parameter θ stored in the memory unit 122 is used USS Update the cost function value. This process involves, for example, the cost function value cost 12 The backpropagation method using this method can be used (step S225).

[0066] The control unit 120 controls the parameter θ SED and θ USS Step S226 determines whether the update process has met the termination condition. If it is determined that the termination condition has not been met, the process returns to step S222. On the other hand, if it is determined that the termination condition has been met, the parameter θ SED and θUSS The learning process is completed. The parameters θ obtained as described above are then finished. SED and θ USS As mentioned above, it is sent to the sound processing device 11.

[0067] [Modified Version of the Second Embodiment] In the second embodiment, inference is performed using inference method 1, and the parameter θ is set using machine learning method 2 of inference method 1. SED and θ USS The learning process was performed. However, the parameter θ was determined by one of the machine learning methods 3 to 8 of the inference method 1. SED and θ USS Learning may be performed. That is, in step S225, instead of PS-SDRi in equation (4), any of PS-SDRi in equations (5), (9), (10), (11), (12), (12a) may be used as the cost function. SDR(s'(c') in equation (4) r ), s~(c r The cost function may also be PS-SDRi obtained by replacing ~)) with equation (6). In equations (5) and (9), the cost function may also be PS-SDRi that is less than or equal to 0 or negative in the Otherwise case. SDR(s'(c r '),x' W The cost function can also be PS-SDRi, where the term ) is replaced with an arbitrary constant Z or 0. Equation (12a) SDR(s'(ξ),x' W The cost function can also be PS-SDRi, where the term in () is replaced with an arbitrary constant Z or 0. Furthermore, depending on the variation of PS-SDRi, the acoustic signal separation model M in equation (3a) or equation (3b) can be used instead of equation (3). USS The method may be used, or the training data may be modified (see Machine Learning Methods 3 and 5 of Inference Method 1 mentioned above).

[0068] [Third Embodiment] In the second embodiment or its modification, instead of using PS-SDRi as the cost function, a combination of PS-SDRi and a cost function other than PS-SDRi (e.g., BCE loss, SDR, SI-SDR, etc.) is used as the parameter θ SED and θ USSThis can also be used as a cost function for learning. For example, the weighted sum of PS-SDRi and cost functions other than PS-SDRi can be used as the parameter θ. SED and θ USS This may also be used as a cost function for learning. In this embodiment, as an example, the cost function value of PS-SDRi obtained by the simultaneous optimization cost function calculation unit 225 (Figure 4) of the second embodiment is cost 12 The weighted sum of this sum and the cost function value cost1 obtained by the acoustic event detection cost function calculation unit 125 (Figure 3) is calculated using the parameter θ. SED and θ USS Cost function value for learning cost Multi Let's explain an example of this.

[0069] <Configuration> As illustrated in Figure 1, the acoustic processing system 3 of this embodiment includes an acoustic processing device 11 that performs inference processing and a machine learning device 32 that performs learning processing. The acoustic processing system 3 receives, for example, a parameter θ via a network from the machine learning device 32 to the acoustic processing device 11. SED and parameter θ USS It is configured to be available.

[0070] <Sound processing device 11> This is the same as in the first embodiment.

[0071] <Machine Learning Device 32> As illustrated in Figure 5, the machine learning device 32 in this embodiment includes a control unit 120, storage units 121 and 122, an acoustic event detection unit 123, an acoustic signal separation unit 124, a cost function calculation unit 325, parameter update units 226 and 228, and a memory 129. As illustrated in Figure 6, the cost function calculation unit 325 in this embodiment includes an acoustic event detection cost function calculation unit 125, a simultaneous optimization cost function calculation unit 225, and a multitasking cost function calculation unit 3251. The following explanation is omitted, but the processing of the machine learning device 32 is performed under the control of the control unit 120. Information input to the machine learning device 32 and information obtained by each unit are stored in the memory 129 one by one and read out and used as needed.

[0072] <Inference Processing (Inference Method 1)> This is the same as in the first embodiment.

[0073] <Learning Process> First, the parameter θ SED The initial value of the parameter θ is stored in the memory unit 121 of the machine learning device 32 (Figure 5), USS The initial value is stored in the storage unit 122 (step S321).

[0074] The machine learning device 32 has the aforementioned training data D = {x', c1', ..., c R ', s'(c1'),...,s'(c R The input is ')}. The learning acoustic signal x' is sent to the acoustic event detection unit 123 and the acoustic signal separation unit 124, and classes c1',...,c R 'and original signals s'(c1'),...,s'(c R The ') is sent to the cost function calculation unit 325. More specifically, classes c1',...,c R ' is sent to the acoustic event detection cost function calculation unit 125 and the simultaneous optimization cost function calculation unit 225 of the cost function calculation unit 325 (Figure 6), and the original signals s'(c1'),...,s'(c R The result is sent to the simultaneous optimization cost function calculation unit 225 (step S322).

[0075] The acoustic event detection unit 123 uses the received learning acoustic signal x' and the parameter θ extracted from the storage unit 121. SED Using this, and following equation (2) above, the estimated classes {c1~,...,c L ~}=M SED (x';θ SED The inference result is obtained and output. Estimated class {c1~,...,c L The ~} is sent to the acoustic signal separation unit 124 and the acoustic event detection cost function calculation unit 125 and the simultaneous optimization cost function calculation unit 225 of the cost function calculation unit 325 (Figure 6) (step S323).

[0076] The acoustic signal separation unit 124 (Figure 5) separates the learning acoustic signal x' and the class {c1~,...,c L ~} and the parameter θ extracted from the memory unit 122 USS Using this, and according to equation (3) above, the estimated original signal {s~(c1~),...,s~(c L ~)}=M USS(x',c1~,...,c L ~;θ USS The estimated original signal {s~(c1~),...,s~(c L The result is sent to the simultaneous optimization cost function calculation unit 225 of the cost function calculation unit 325 (Figure 6) (step S324).

[0077] The acoustic event detection cost function calculation unit 125 calculates {c1~,...,c L ~} (inference result) and {c1',...,c R The cost function for {c1~,...,c} (training data) is calculated and the cost function value cost1 is output. The simultaneous optimization cost function calculation unit 225 calculates {c1~,...,c}. L ~} and {s~(c1~),...,s~(c L ~)} (inference result) and {c1',...,c R '} and {s'(c1'),...,s'(c R For the training data, calculate PS-SDRi in equation (4) as the cost function, and the cost function value cost 12 Outputs the cost function values ​​cost1 and cost 12 The cost function values ​​cost1 and cost are sent to the multitask cost function calculation unit 3251. 12 The weighted sum of the cost function value Multi Calculate and output the cost function value. Multi This is sent to the parameter update units 226 and 228 (Figure 5). The parameter update unit 226 processes the cost function value cost Multi The parameter θ stored in the memory unit 121 is used SED The parameter update unit 228 updates the cost function value cost Multi The parameter θ stored in the memory unit 122 is used USS Update (step S325).

[0078] The control unit 120 controls the parameter θ SED and θ USSStep S326 determines whether the update process has met the termination condition. If it is determined that the termination condition has not been met, the process returns to step S322. On the other hand, if it is determined that the termination condition has been met, the parameter θ SED and θ USS The learning process is completed. The parameters θ obtained as described above are then finished. SED and θ USS As mentioned above, it is sent to the sound processing device 11.

[0079] [Fourth Embodiment] In this embodiment, inference is performed using the inference method 2 described above. For training the multitask model M, the machine learning method 1 of the inference method 2 described above is used.

[0080] <Configuration> As illustrated in Figure 7, the acoustic processing system 4 of this embodiment includes an acoustic processing device 41 that performs inference processing and a machine learning device 42 that performs learning processing. The acoustic processing system 4 is configured to allow the machine learning device 42 to provide parameters θ to the acoustic processing device 41, for example, through a network.

[0081] <Sound Processing Device 41> As illustrated in Figure 8, the sound processing device 41 in this embodiment includes a control unit 110, a storage unit 411, a multitasking unit 413, and a memory 115. Although not explained below, the processing of the sound processing device 41 is performed under the control of the control unit 110. Information input to the sound processing device 41 and information obtained by each unit are stored in the memory 115 one by one and read out and used as needed.

[0082] <Machine Learning Device 42> As illustrated in Figure 9, the machine learning device 42 in this embodiment includes a control unit 120, a storage unit 421, a multitasking unit 423, a simultaneous optimization cost function calculation unit 225, a parameter update unit 426, and a memory 129. Although not explained below, the processing of the machine learning device 42 is performed under the control of the control unit 120. Information input to the machine learning device 42 and information obtained by each unit are stored in the memory 129 one by one and read out and used as needed.

[0083] <Inference Processing (Inference Method 2)> The acoustic processing device 41 (Figure 8) extracts acoustic events from the input acoustic signal x representing the mixed sound, and extracts the acoustic signal corresponding to the acoustic event from the input acoustic signal. This will be explained in detail below.

[0084] First, as a preprocessing step, the parameters θ of the multitask model M obtained by machine learning in the machine learning device 42 are sent to the sound processing device 41. The parameters θ are stored in the memory unit 411 of the sound processing device 41 (Figure 8). This preprocessing is performed initially, and also as needed, and the parameters θ stored in the memory unit 411 may be updated as a result.

[0085] Based on the above preprocessing, the aforementioned input acoustic signal x is input to the acoustic processing device 41 (Equation (1)). The input acoustic signal x is sent to the multitasking unit 413 (Step S411).

[0086] The multitasking unit 413 uses the input acoustic signal x and the parameter θ extracted from the storage unit 411 to estimate the class c1^,...,c of the acoustic event according to the equation (13) described above. K ^ and the original signals s^(c1^), ..., s^(c K ^) is obtained and output (step S412).

[0087] The acoustic processing device 41 is of the class c1^,...,c estimated as described above. K ^ and the original signals s^(c1^), ..., s^(c K ^) is output. Alternatively, the sound processing device 41 outputs h1*s^(c1^),...,h k *s^(c K ^), or s^(c1^),...,s^(c K Other acoustic signals based on ^) may be output (step S413).

[0088] <Learning Process (Machine Learning Method 1 of Inference Method 2)> First, the initial value of the parameter θ is stored in the memory unit 421 of the machine learning device 42 (Figure 9) (Step S421).

[0089] The machine learning device 42 has the aforementioned training data D = {x', c1',...,c R', s'(c1'),...,s'(c R ')} is input. The learning acoustic signal x' is sent to the multitasking unit 423, and classes c1',...,c R 'and original signals s'(c1'),...,s'(c R The result is sent to the simultaneous optimization cost function calculation unit 225 (step S422).

[0090] The multitasking unit 423 uses the input acoustic signal x and the parameters θ extracted from the storage unit 421 to estimate {c1~,...,c} according to the aforementioned equation (13). L ~, s~(c1~),...,s~(c L The estimated {c1~,...,c L ~, s~(c1~),...,s~(c L The result is sent to the simultaneous optimization cost function calculation unit 225 (step S423).

[0091] The simultaneous optimization cost function calculation unit 225 calculates {c1~,...,c L ~} and {s~(c1~),...,s~(c L ~)} (inference result) and {c1',...,c R '} and {s'(c1'),...,s'(c R For the training data, calculate PS-SDRi in equation (4) as the cost function, and the cost function value cost 12 Outputs the cost function value cost. 12 This is sent to the parameter update unit 426. The parameter update unit 426 processes the cost function value cost 12 Then, the parameter θ stored in the memory unit 421 is updated (step S424).

[0092] The control unit 120 determines whether the parameter update process has met the termination condition (step S425). If it is determined that the termination condition has not been met, the process returns to step S422. On the other hand, if it is determined that the termination condition has been met, the parameter learning process for parameter θ is terminated. The parameter θ obtained as described above is sent to the acoustic processing device 41 as described above.

[0093] [Modification of the Fourth Embodiment] In the fourth embodiment, inference was performed using inference method 2, and the parameter θ was learned using machine learning method 1 of inference method 2. However, the parameter θ may be learned using any of machine learning methods 2 to 7 of inference method 2. In addition, a multitask model M of equation (13a) or equation (13b) may be used instead of equation (13), or the training data may be changed (see machine learning methods 2, 4, etc. of inference method 2 described above). Furthermore, as explained in the third embodiment, in the fourth embodiment or its modifications, instead of using PS-SDRi as the cost function, a combination of PS-SDRi and a cost function other than PS-SDRi (e.g., BCE loss, SDR, SI-SDR, etc.) may be used as the cost function for learning the parameter θ.

[0094] [Features of Each Embodiment] As described above, the acoustic processing apparatus of each embodiment and its modified form extracts acoustic events from an input acoustic signal representing a mixed sound, and extracts acoustic signals corresponding to the acoustic events from the input acoustic signal. Therefore, even when it is not known what kind of acoustic events the input acoustic signal corresponds to, it is possible to efficiently extract (separate) the acoustic signals corresponding to the acoustic events contained in the input acoustic signal.

[0095] Alternatively, after extracting an acoustic event from the input acoustic signal, the corresponding acoustic signal may be extracted from the input acoustic signal (inference method 1). However, by simultaneously extracting both the acoustic event and the corresponding acoustic signal from the input acoustic signal (inference method 2), the corresponding acoustic signal can be extracted (separated) even more efficiently.

[0096] As mentioned above, the models used for extracting acoustic events and acoustic signals (acoustic event detection models, acoustic signal separation models, and multitask models) are machine learning models based on evaluations of the extraction of acoustic events corresponding to training acoustic signals and evaluations of the extraction of acoustic signals corresponding to acoustic events corresponding to training acoustic signals. By using models machine learning based on criteria that perform these evaluations simultaneously (e.g., PS-SDRi), it is possible to extract (separate) acoustic signals corresponding to acoustic events with even greater accuracy.

[0097] Furthermore, as mentioned above, by using a machine learning model based on criteria that penalize failures in extracting acoustic events corresponding to training acoustic signals and failures in extracting acoustic signals corresponding to acoustic events corresponding to training acoustic signals, it is possible to extract (separate) acoustic signals corresponding to acoustic events with greater accuracy.

[0098] [Hardware Configuration] The functions realized by the components described herein may be implemented in a circuitry or processing circuitry, including a general-purpose processor, an application-specific processor, an integrated circuit, an ASIC (Application Specific Integrated Circuit), a CPU (a Central Processing Unit), conventional circuits, and / or a combination thereof, programmed to realize the functions described herein. A processor includes transistors and other circuits and is considered a circuitry or processing circuitry. A processor may be a programmed processor that executes a program stored in memory.

[0099] In this specification, circuitry, unit, and means are hardware programmed to perform or execute the functions described herein. Such hardware may be any hardware disclosed herein, or any hardware known to be programmed to perform or execute the functions described herein.

[0100] If the hardware is a processor that is considered to be a type of circuitry, then the circuitry, means, or unit is a combination of hardware and software used to constitute the hardware and / or processor.

[0101] For example, the sound processing devices 11, 41 and the machine learning devices 12, 22, 32, 42 in each embodiment are devices configured by a general-purpose or dedicated computer equipped with a processor (hardware processor) such as a CPU (central processing unit) and memory such as RAM (random-access memory) and ROM (read-only memory) executing a predetermined program. That is, the sound processing devices 11, 41 and the machine learning devices 12, 22, 32, 42 in each embodiment have, for example, processing circuits configured to implement each of their respective parts. This computer may have one processor and memory, or it may have multiple processors and memories. This program may be installed on the computer, or it may be pre-recorded in ROM, etc. Furthermore, some or all of the processing units may be configured using electronic circuits that realize processing functions independently, rather than electronic circuits that realize the functional configuration by loading a program, such as a CPU. Also, the electronic circuits that constitute one device may include multiple CPUs.

[0102] Figure 6 is a block diagram illustrating the hardware configuration of the sound processing devices 11, 41 and machine learning devices 12, 22, 32, 42 in each embodiment. As illustrated in Figure 6, the sound processing devices 11, 41 and machine learning devices 12, 22, 32, 42 in this example have a CPU (Central Processing Unit) 10a, an input unit 10b, an output unit 10c, a RAM (Random Access Memory) 10d, a ROM (Read Only Memory) 10e, an auxiliary storage device 10f, a communication unit 10h, and a bus 10g. The CPU 10a in this example has a control unit 10aa, an arithmetic unit 10ab, and a register 10ac, and performs various arithmetic processing according to various programs loaded into the register 10ac. The input unit 10b is an input terminal, keyboard, mouse, touch panel, etc., to which data is input. The output unit 10c is an output terminal, display, etc., to which data is output. The communication unit 10h is a LAN card or the like controlled by the CPU 10a, which has loaded a predetermined program. The RAM 10d is an SRAM (Static Random Access Memory), DRAM (Dynamic Random Access Memory), etc., and has a program area 10da where a predetermined program is stored and a data area 10db where various data is stored. The auxiliary storage device 10f is, for example, a hard disk, MO (Magneto-Optical disc), semiconductor memory, etc., and has a program area 10fa where a predetermined program is stored and a data area 10fb where various data is stored. The bus 10g connects the CPU 10a, input unit 10b, output unit 10c, RAM 10d, ROM 10e, communication unit 10h, and auxiliary storage device 10f so that information can be exchanged. The CPU 10a writes the program stored in the program area 10fa of the auxiliary storage device 10f to the program area 10da of the RAM 10d, according to the loaded OS (Operating System) program. Similarly, the CPU 10a writes various data stored in the data area 10fb of the auxiliary storage device 10f to the data area 10db of the RAM 10d.The addresses on RAM 10d where the program and data are written are then stored in register 10ac of the CPU 10a. The control unit 10aa of the CPU 10a sequentially reads these addresses stored in register 10ac, reads the program and data from the area on RAM 10d indicated by the read addresses, sequentially has the calculation unit 10ab execute the calculations indicated by the program, and stores the calculation results in register 10ac. This configuration realizes the functional configurations of the sound processing devices 11, 41 and the machine learning devices 12, 22, 32, 42.

[0103] The program describing this process can be recorded on a computer-readable recording medium. Examples of computer-readable recording media include non-transitory recording media. Examples of such recording media include magnetic recording devices, optical discs, magneto-optical recording media, and semiconductor memory. The program describing this process (computer program) may be included in a computer program product.

[0104] Furthermore, this program may be distributed, for example, by selling, transferring, or lending portable recording media such as DVDs or CD-ROMs on which the program is recorded. Alternatively, the program may be stored in the storage device of a server computer and distributed by transferring the program from the server computer to other computers via a network.

[0105] A computer executing such a program may, for example, first store the program recorded on a portable storage medium or a program transferred from a server computer in its own storage device. Then, when processing is to be executed, the computer reads the program stored on its own storage medium and executes the processing according to the read program. Alternatively, the computer may directly read the program from the portable storage medium and execute the processing according to that program, or it may sequentially execute the processing according to the received program each time a program is transferred to it from a server computer. Furthermore, the processing may be executed using a so-called ASP (Application Service Provider) type service, where the processing function is realized only by issuing execution instructions and obtaining results, without transferring the program from the server computer to this computer.In addition, the processing may be executed using a so-called SaaS (Software as a Service) type service, where a part of the server computer is made available to the user along with the program. Furthermore, the term "program" in this form includes information used for processing by an electronic computer that is equivalent to a program (data, etc., that is not a direct instruction to the computer but has the property of defining the computer's processing).

[0106] Furthermore, in this configuration, the device is configured by executing a predetermined program on a computer, but at least a part of these processes may be implemented in hardware.

[0107] [Other Modifications] The present invention is not limited to the embodiments described above. For example, instead of PS-SDRi, the function value of PS-SDRi may be used as the cost function value. For example, the monotonically increasing function value, monotonically decreasing function value, monotonically non-decreasing function value, monotonically non-increasing function value, etc., of PS-SDRi may be used as the cost function value. Furthermore, the various processes described above may not only be executed in time series as described, but may also be executed in parallel or individually as needed, depending on the processing capacity of the device performing the process. Needless to say, other modifications can be made as appropriate without departing from the spirit of the present invention.

[0108] This invention enables the efficient extraction of acoustic signals corresponding to acoustic events contained in an input acoustic signal, even when it is unknown what acoustic events the input acoustic signal corresponds to. This allows for the separation and extraction of metadata regarding the types (classes) and number of acoustic events contained in a given environment, as well as the original signals corresponding to those acoustic events. Therefore, this invention can be applied, for example, to immersive communication that selectively reconstructs the sound space of one environment in another space. For instance, through semantic understanding and separation of acoustic scenes according to this invention, all acoustic events present in an environment can be separated and objectified, eliminating the influence of that environment (noise, reverberation, etc.). By transferring and reconstructing this objectified sound space in a desired format, desired immersive communication can be realized.

[0109] 1-4 Acoustic Processing Systems 11, 41 Acoustic Processing Devices 12, 22, 32, 42 Machine Learning Devices

Claims

1. An acoustic processing method for extracting acoustic events from an input acoustic signal representing a mixed sound, and extracting an acoustic signal corresponding to the acoustic event from the input acoustic signal.

2. An acoustic processing method according to claim 1, comprising simultaneously extracting the acoustic event and the acoustic signal corresponding to the acoustic event from the input acoustic signal.

3. An acoustic processing method according to claim 1 or 2, wherein the acoustic processing method is based on a model trained on a criterion that simultaneously performs evaluation of the extraction of acoustic events corresponding to a learning acoustic signal representing a mixed sound for learning, and evaluation of the extraction of acoustic signals corresponding to acoustic events corresponding to the learning acoustic signal.

4. An acoustic processing method according to claim 3, wherein the acoustic processing method is based on a model that has been machine-trained on criteria for imposing penalties for failure to extract an acoustic event corresponding to the training acoustic signal and for failure to extract an acoustic signal corresponding to an acoustic event corresponding to the training acoustic signal.