Information processing apparatus, output method, and recording medium

By acquiring sound source location information and mixed sound signals, and using a learned model to process sound feature quantities, a target sound signal is generated. This solves the problem of signal separation when the incident angle of the target sound and the interference sound is small, and achieves effective target sound output.

CN116997961BActive Publication Date: 2026-07-31MITSUBISHI ELECTRIC CORP
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
MITSUBISHI ELECTRIC CORP
Filing Date
2021-04-07
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

When the incident angle between the target sound and the interfering sound is small, existing technologies struggle to effectively output the target sound signal.

Method used

By acquiring the source location information and mixed sound signals of the target sound, the learned model is used to extract, enhance, mask, and estimate sound features, generate enhanced and masked sound signals of the target sound direction, and finally output the target sound signal.

Benefits of technology

When the incident angle between the target sound and the interfering sound is small, it can effectively output the target sound signal, thus improving the accuracy and quality of signal separation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116997961B_ABST
    Figure CN116997961B_ABST
Patent Text Reader

Abstract

The information processing apparatus (100) includes: an acquisition unit (120) that acquires sound source location information (111), a mixed sound signal, and a learned model (112); a sound feature extraction unit (130) that extracts multiple sound features based on the mixed sound signal; an enhancement unit (140) that enhances the sound feature of the target sound direction among the multiple sound features based on the sound source location information (111); an estimation unit (150) that estimates the target sound direction based on the multiple sound features and the sound source location information (111); a masking feature extraction unit (160) that extracts masking features based on the estimated target sound direction and the multiple sound features; a generation unit (170) that generates a target sound direction enhanced sound signal based on the enhanced sound features and generates a target sound direction masking sound signal based on the masking features; and a target sound signal output unit (180) that outputs a target sound signal using the target sound direction enhanced sound signal, the target sound direction masking sound signal, and the learned model (112).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an information processing apparatus, an output method, and a recording medium containing an output program. Background Technology

[0002] When multiple speakers speak simultaneously, the speech becomes mixed. Sometimes it is desirable to extract the speech of a target speaker from the mixed speech. For example, in the case of extracting the speech of a target speaker, a method for noise suppression is considered. Here, a method for noise suppression has been proposed (see Patent Document 1).

[0003] Existing technical documents

[0004] Patent documents

[0005] Patent Document 1: Japanese Patent Application Publication No. 2010-239424

[0006] Patent Document 2: International Publication No. 2016 / 143125

[0007] Non-patent literature

[0008] Non-patent literature 1: Yi Luo, Nima Mesgarani, “Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Separation”, 2019

[0009] Non-patent literature 2: Ashish Vaswani et al., “Attention Is All You Need”, in Proc. NIPS, 2017 Summary of the Invention

[0010] The problem that the invention aims to solve

[0011] However, when the angle between the direction in which the target sound (e.g., the voice of the target speaker) is incident on the microphone and the direction in which the interfering sound (e.g., the voice of the interfering speaker) is incident on the microphone is small, sometimes even when the device uses the above-described techniques, it is difficult to output a signal representing the target sound, i.e., the target sound signal.

[0012] The purpose of this invention is to output a target sound signal.

[0013] Methods for solving problems

[0014] An information processing apparatus according to an embodiment of the present invention is provided. The information processing apparatus includes: an acquisition unit that acquires sound source location information (i.e., sound source location information), a signal representing a mixed sound including the target sound and interfering sound (i.e., a mixed sound signal), and a learned model; a sound feature extraction unit that extracts multiple sound features based on the mixed sound signal; an enhancement unit that enhances a sound feature representing the direction of the target sound (i.e., the target sound direction) among the multiple sound features based on the sound source location information; an estimation unit that estimates the target sound direction based on the multiple sound features and the sound source location information; and a masking feature extraction unit that extracts masking features based on the estimated masking features. The system comprises: a target sound direction and a plurality of sound feature quantities; an extraction unit that extracts the feature quantity of the target sound direction being masked, i.e., a masking feature quantity; a generation unit that generates a sound signal with enhanced target sound direction, i.e., a target sound direction enhanced sound signal, based on the enhanced sound feature quantity, and a sound signal with masked target sound direction, i.e., a target sound direction masking sound signal, based on the masking feature quantity; and a target sound signal output unit that uses the target sound direction enhanced sound signal, the target sound direction masking sound signal, and the learned model to output a signal representing the target sound, i.e., a target sound signal.

[0015] Invention Effects

[0016] According to the present invention, a target sound signal can be output. Attached Figure Description

[0017] Figure 1 This is a diagram illustrating an example of the target sound signal output system of Embodiment 1.

[0018] Figure 2 This is a diagram showing the hardware of the information processing apparatus of Embodiment 1.

[0019] Figure 3 This is a block diagram illustrating the function of the information processing device in Embodiment 1.

[0020] Figure 4 This is a diagram illustrating a structural example of the learned model in Implementation 1.

[0021] Figure 5 This is a flowchart illustrating an example of the processing performed by the information processing apparatus of Embodiment 1.

[0022] Figure 6 This is a block diagram illustrating the function of the learning device in Embodiment 1.

[0023] Figure 7 This is a flowchart illustrating an example of the processing performed by the learning device in Implementation 1.

[0024] Figure 8 This is a block diagram illustrating the function of the information processing device in Embodiment 2.

[0025] Figure 9 This is a flowchart illustrating an example of the processing performed by the information processing apparatus of Embodiment 2.

[0026] Figure 10 This is a block diagram illustrating the function of the information processing device in Embodiment 3.

[0027] Figure 11 This is a flowchart illustrating an example of the processing performed by the information processing apparatus of Embodiment 3.

[0028] Figure 12 This is a block diagram illustrating the function of the information processing device in Embodiment 4.

[0029] Figure 13 This is a flowchart illustrating an example of the processing performed by the information processing apparatus of Embodiment 4. Detailed Implementation

[0030] The embodiments will now be described with reference to the accompanying drawings. These embodiments are merely examples, and various modifications can be made within the scope of this invention.

[0031] Implementation Method 1

[0032] Figure 1 This diagram illustrates an example of a target sound signal output system according to Embodiment 1. The target sound signal output system includes an information processing unit 100 and a learning unit 200. The information processing unit 100 is an apparatus for executing an output method. The information processing unit 100 outputs a target sound signal using a learned model. The learned model is generated by the learning unit 200.

[0033] The information processing device 100 will be explained using the utilization phase. The learning device 200 will be explained using the learning phase. First, the utilization phase will be explained.

[0034] <Application Phase>

[0035] Figure 2 This is a diagram illustrating the hardware of the information processing apparatus of Embodiment 1. The information processing apparatus 100 includes a processor 101, a volatile storage device 102, and a non-volatile storage device 103.

[0036] The processor 101 controls the information processing device 100 as a whole. For example, the processor 101 may be a CPU (Central Processing Unit), an FPGA (Field Programmable Gate Array), or the like. The processor 101 may also be a multi-processor. Furthermore, the information processing device 100 may also have processing circuitry. This processing circuitry may be a single circuit or a composite circuit.

[0037] Volatile storage device 102 is the main storage device of information processing device 100. For example, volatile storage device 102 is RAM (Random Access Memory). Non-volatile storage device 103 is an auxiliary storage device of information processing device 100. For example, non-volatile storage device 103 is HDD (Hard Disk Drive) or SSD (Solid State Drive).

[0038] Furthermore, the storage area secured by the volatile storage device 102 or the non-volatile storage device 103 is called the storage section.

[0039] Next, the functions of the information processing device 100 will be explained.

[0040] Figure 3 This is a block diagram illustrating the functions of the information processing apparatus of Embodiment 1. The information processing apparatus 100 includes an acquisition unit 120, a sound feature extraction unit 130, an enhancement unit 140, an estimation unit 150, a masking feature extraction unit 160, a generation unit 170, and a target sound signal output unit 180.

[0041] The acquisition unit 120, the sound feature extraction unit 130, the enhancement unit 140, the estimation unit 150, the masking feature extraction unit 160, the generation unit 170, and the target sound signal output unit 180 may also be implemented by a processing circuit. Furthermore, the acquisition unit 120, the sound feature extraction unit 130, the enhancement unit 140, the estimation unit 150, the masking feature extraction unit 160, the generation unit 170, and the target sound signal output unit 180 may also be implemented as modules of a program executed by the processor 101. For example, the program executed by the processor 101 is also called an output program. For example, the output program is recorded on a recording medium.

[0042] The storage unit can also store sound source location information 111 and the learned model 112. Sound source location information 111 is the location information of the sound source of the target sound. For example, if the target sound is speech produced by a target speaker, sound source location information 111 is the location information of the target speaker.

[0043] The acquisition unit 120 acquires the sound source location information 111. For example, the acquisition unit 120 acquires the sound source location information 111 from the storage unit. Here, the sound source location information 111 may also be stored in an external device (e.g., a cloud server). If the sound source location information 111 is stored in an external device, the acquisition unit 120 acquires the sound source location information 111 from the external device.

[0044] The acquisition unit 120 acquires the learned model 112. For example, the acquisition unit 120 acquires the learned model 112 from the storage unit. Alternatively, for example, the acquisition unit 120 acquires the learned model 112 from the learning device 200.

[0045] The acquisition unit 120 acquires the mixed sound signal. For example, the acquisition unit 120 acquires the mixed sound signal from a microphone array having N (N is an integer of 2 or more) microphones. The mixed sound signal is a signal representing a mixture of sound including the target sound and interference sound. The mixed sound signal can also be represented as N sound signals. In addition, for example, the target sound is the speech of the target speaker, the sound of an animal, etc. The interference sound is the sound that interferes with the target sound. Furthermore, noise may also be included in the mixed sound. In the following description, it is assumed that the mixed sound includes the target sound, interference sound, and noise.

[0046] The sound feature extraction unit 130 extracts multiple sound features from the mixed sound signal. For example, the sound feature extraction unit 130 extracts the time series of the power spectrum obtained by performing a short-time Fourier transform (STFT) on the mixed sound signal as multiple sound features. Alternatively, the extracted multiple sound features can also be represented as N sound features.

[0047] The enhancement unit 140 enhances the sound feature quantity in the target sound direction from among multiple sound feature quantities based on the sound source location information 111. For example, the enhancement unit 140 uses multiple sound feature quantities, the sound source location information 111, and an MVDR (Minimum Variance Distortionless Response) beamformer to enhance the sound feature quantity in the target sound direction.

[0048] The estimation unit 150 estimates the direction of the target sound based on multiple sound feature quantities and sound source location information 111. In detail, the estimation unit 150 estimates the direction of the target sound using equation (1).

[0049] l represents time. k represents frequency. x lk This represents the sound feature quantity corresponding to the sound signal obtained from the microphone closest to the sound source location of the target sound determined according to sound source location information 111.lk It can also be considered as an STFT spectrum. θ,k This represents the turning vector at a certain angle θ. H is the conjugate transpose.

[0050]

[0051] The masking feature extraction unit 160 extracts masking features based on the estimated target sound direction and multiple sound feature quantities. Masking features are features representing the state where the target sound direction's features are masked. The extraction process for masking features is described in detail. The masking feature extraction unit 160 creates a directional mask based on the target sound direction. The directional mask is a mask of the enhanced sound extracted from the target sound direction. This mask is a matrix of the same size as the sound feature quantities. When the angular range of the target sound direction is ω, the directional mask M... lk It is expressed by equation (2).

[0052]

[0053] The masking feature extraction unit 160 extracts masking features by multiplying the element-wise product of the masking matrix with multiple sound features.

[0054] The generation unit 170 generates an enhanced sound signal (hereinafter referred to as the target sound direction enhanced sound signal) based on the sound feature quantities enhanced by the enhancement unit 140. For example, the generation unit 170 generates the target sound direction enhanced sound signal using the sound feature quantities enhanced by the enhancement unit 140 and the inverse short-time Fourier transform (ISTFT).

[0055] The generation unit 170 generates a target sound direction masking sound signal (hereinafter referred to as the target sound direction masking sound signal) based on masking feature quantities. For example, the generation unit 170 generates the target sound direction masking sound signal using masking feature quantities and short-time Fourier transform.

[0056] The target sound direction enhancement sound signal and the target sound direction masking sound signal can also be input into the learning device 200 as learning signals.

[0057] The target sound signal output unit 180 outputs a target sound signal using a target sound direction enhancement sound signal, a target sound direction masking sound signal, and a learned model 112. Here, an example of the structure of the learned model 112 will be described.

[0058] Figure 4This is a diagram illustrating a structural example of the learned model in Implementation 1. The learned model 112 includes an encoder 112a, a separator 112b, and a decoder 112c.

[0059] Encoder 112a estimates the target sound direction enhancement time-frequency performance in "M dimensions × time" based on the target sound direction enhancement sound signal. Furthermore, encoder 112a estimates the target sound direction masking time-frequency performance in "M dimensions × time" based on the target sound direction masking sound signal. For example, encoder 112a can also estimate the target sound direction enhancement time-frequency performance and the target sound direction masking time-frequency performance from the power spectrum estimated by STFT. Furthermore, for example, encoder 112a can also use one-dimensional convolution operations to estimate the target sound direction enhancement time-frequency performance and the target sound direction masking time-frequency performance. In this estimation, the target sound direction enhancement time-frequency performance and the target sound direction masking time-frequency performance can be projected to the same time-frequency performance space or to different time-frequency performance spaces. Additionally, for example, this estimation is described in Non-Patent Document 1.

[0060] Separator 112b estimates an "M-dimensional × time" masking matrix based on the target sound direction enhancement time-frequency performance and the target sound direction masking time-frequency performance. Furthermore, when the target sound direction enhancement time-frequency performance and the target sound direction masking time-frequency performance are input to separator 112b, they can also be concatenated along the frequency axis. This results in a "2M-dimensional × time" performance. Alternatively, the target sound direction enhancement time-frequency performance and the target sound direction masking time-frequency performance can be concatenated along an axis different from the time and frequency axes. This results in a "M-dimensional × time × 2" performance. The target sound direction enhancement time-frequency performance and the target sound direction masking time-frequency performance can also be weighted. The weighted target sound direction enhancement time-frequency performance and the weighted target sound direction masking time-frequency performance can also be combined. The weights can also be estimated by the learned model 112.

[0061] Furthermore, the separator 112b is a neural network consisting of an input layer, an intermediate layer, and an output layer. For example, regarding the propagation between layers, a method similar to LSTM (Long Short Term Memory) and a method combining one-dimensional convolution operations can also be used.

[0062] Decoder 112c multiplies the target sound direction enhancement time-frequency representation ("M-dimensional × time") with the masking matrix ("M-dimensional × time"). Decoder 112c uses the information obtained through multiplication and a method corresponding to that used in encoder 112a to output the target sound signal. For example, if the method used in encoder 112a is STFT, decoder 112c uses the information obtained through multiplication and ISTFT to output the target sound signal. Furthermore, for example, if the method used in encoder 112a is one-dimensional convolution, decoder 112c uses the information obtained through multiplication and the inverse one-dimensional convolution operation to output the target sound signal.

[0063] The target audio signal output unit 180 can also output the target audio signal to the speaker. Thus, the target audio is output from the speaker. The speaker illustration is omitted.

[0064] Next, a flowchart will be used to explain the processes performed by the information processing device 100.

[0065] Figure 5 This is a flowchart illustrating an example of the processing performed by the information processing apparatus of Embodiment 1.

[0066] (Step S11) The acquisition unit 120 acquires the mixed sound signal.

[0067] (Step S12) The sound feature extraction unit 130 extracts multiple sound features based on the mixed sound signal.

[0068] (Step S13) The enhancement unit 140 enhances the sound feature quantity of the target sound direction according to the sound source location information 111.

[0069] (Step S14) The estimation unit 150 estimates the direction of the target sound based on multiple sound feature quantities and sound source location information 111.

[0070] (Step S15) The masking feature extraction unit 160 extracts masking features based on the estimated target sound direction and multiple sound feature quantities.

[0071] (Step S16) The generation unit 170 generates a target sound direction enhanced sound signal based on the sound feature quantity enhanced by the enhancement unit 140. In addition, the generation unit 170 generates a target sound direction masking sound signal based on the masking feature quantity.

[0072] (Step S17) The target sound signal output unit 180 outputs the target sound signal using the target sound direction enhancement sound signal, the target sound direction masking sound signal and the learned model 112.

[0073] Alternatively, steps S14 and S15 can be executed in parallel with step S13. Furthermore, steps S14 and S15 can also be executed before step S13.

[0074] Next, the learning stages will be explained.

[0075] <Learning Phase>

[0076] During the learning phase, an example of the generation of the learned model 112 will be explained.

[0077] Figure 6 This is a block diagram illustrating the functions of the learning device according to Embodiment 1. The learning device 200 includes an audio data storage unit 211, an impulse response storage unit 212, a noise storage unit 213, an impulse response application unit 220, a mixing unit 230, a processing execution unit 240, and a learning unit 250.

[0078] Furthermore, the sound data storage unit 211, the impulse response storage unit 212, and the noise storage unit 213 can also be implemented as storage areas secured by the volatile or non-volatile storage devices provided by the learning device 200.

[0079] Some or all of the impulse response application unit 220, mixing unit 230, processing execution unit 240, and learning unit 250 can also be implemented by the processing circuitry of the learning device 200. Furthermore, some or all of the impulse response application unit 220, mixing unit 230, processing execution unit 240, and learning unit 250 can also be implemented as modules of a program executed by the processor of the learning device 200.

[0080] The sound data storage unit 211 stores the target sound signal and the interference sound signal. The interference sound signal is a signal representing the interference sound. The impulse response storage unit 212 stores impulse response data. The noise storage unit 213 stores the noise signal. The noise signal is a signal representing noise.

[0081] The impulse response application unit 220 convolves the impulse response data corresponding to the position of the target sound and the position of the interference sound with one target sound signal and an arbitrary number of interference sound signals stored in the sound data storage unit 211.

[0082] The mixing unit 230 generates a mixed sound signal based on the sound signal output by the impulse response application unit 220 and the noise signal stored in the noise storage unit 213. Furthermore, the sound signal output by the impulse response application unit 220 can also be processed as a mixed sound signal. The learning device 200 can also send the mixed sound signal to the information processing device 100.

[0083] The processing execution unit 240 executes steps S11 to S16, thereby generating a target sound direction enhancement sound signal and a target sound direction masking sound signal. That is, the processing execution unit 240 generates a learning signal.

[0084] The learning unit 250 performs learning using a learning signal. Specifically, the learning unit 250 uses the target sound direction enhancement sound signal and the target sound direction masking sound signal to learn the output target sound signal. Furthermore, during learning, the parameters of the neural network, i.e., the input weight coefficients, are determined. The loss function shown in Non-Patent Document 1 can also be used during learning. Additionally, during learning, the sound signal output by the impulse response application unit 220 and the loss function can be used to calculate the error. Moreover, for example, optimization methods such as Adam are used during learning to determine the input weight coefficients of each layer of the neural network based on the backpropagation method.

[0085] In addition, the learning signal can be a learning signal generated by the processing execution unit 240 or a learning signal generated by the information processing device 100.

[0086] Next, a flowchart will be used to explain the processes performed by the learning device 200.

[0087] Figure 7 This is a flowchart illustrating an example of the processing performed by the learning device in Implementation 1.

[0088] (Step S21) The impulse response application unit 220 convolves the impulse response data with the target sound signal and the interference sound signal.

[0089] (Step S22) The mixing unit 230 generates a mixed sound signal based on the sound signal and noise signal output by the impulse response application unit 220.

[0090] (Step S23) The processing execution unit 240 executes steps S11 to S16, thereby generating a learning signal.

[0091] (Step S24) The learning unit 250 uses the learning signal to learn.

[0092] Then, the learning device 200 repeatedly learns, thereby generating the learned model 112.

[0093] According to Embodiment 1, the information processing device 100 uses a learned model 112 to output a target sound signal. The learned model 112 is a learned model generated by learning a target sound direction enhancement sound signal and a target sound direction masking sound signal for outputting the target sound signal. Specifically, the learned model 112 identifies enhanced or masked target sound components and unenhanced or unmasked target sound components, thereby outputting the target sound signal even when the angle between the target sound direction and the interference sound direction is small. Therefore, even when the angle between the target sound direction and the interference sound direction is small, the information processing device 100 also uses the learned model 112, thereby enabling it to output the target sound signal.

[0094] Implementation Method 2

[0095] Next, Embodiment 2 will be described. In Embodiment 2, the differences from Embodiment 1 will be mainly explained. Furthermore, in Embodiment 2, the description of matters common to Embodiment 1 will be omitted.

[0096] Figure 8 This is a block diagram illustrating the function of the information processing apparatus of Embodiment 2. The information processing apparatus 100 also includes a selection unit 190.

[0097] Part or all of the selection unit 190 can also be implemented by processing circuitry. Alternatively, part or all of the selection unit 190 can be implemented as a module of a program executed by the processor 101.

[0098] The selection unit 190 uses the mixed sound signal and the sound source location information 111 to select the sound signal of the channel with the target sound direction. In other words, the selection unit 190 selects the sound signal of the channel with the target sound direction from N sound signals based on the sound source location information 111.

[0099] Here, the selected sound signal, the target sound direction enhancement sound signal, and the target sound direction masking sound signal can also be input into the learning device 200 as learning signals.

[0100] The target sound signal output unit 180 outputs the target sound signal using the selected sound signal, the target sound direction enhancement sound signal, the target sound direction masking sound signal, and the learned model 112.

[0101] Next, the processing of the encoder 112a, separator 112b and decoder 112c contained in the learned model 112 will be explained.

[0102] Encoder 112a estimates the target sound direction enhancement time-frequency performance in "M dimensions × time" based on the target sound direction enhancement sound signal. Furthermore, encoder 112a estimates the target sound direction masking time-frequency performance in "M dimensions × time" based on the target sound direction masking sound signal. Moreover, encoder 112a estimates the mixed sound time-frequency performance in "M dimensions × time" based on the selected sound signal. For example, encoder 112a can also estimate the power spectrum estimated by STFT as the target sound direction enhancement time-frequency performance, the target sound direction masking time-frequency performance, and the mixed sound time-frequency performance. Furthermore, for example, encoder 112a can also use one-dimensional convolution operation to estimate the target sound direction enhancement time-frequency performance, the target sound direction masking time-frequency performance, and the mixed sound time-frequency performance. In this estimation, the target sound direction enhancement time-frequency performance, the target sound direction masking time-frequency performance, and the mixed sound time-frequency performance can be projected to the same time-frequency performance space, or they can be projected to different time-frequency performance spaces. Additionally, for example, this estimation is described in Non-Patent Document 1.

[0103] Separator 112b estimates a masking matrix of "M dimensions × time" based on the target sound direction enhancement time-frequency performance, the target sound direction masking time-frequency performance, and the mixed sound time-frequency performance. Furthermore, when the target sound direction enhancement time-frequency performance, the target sound direction masking time-frequency performance, and the mixed sound time-frequency performance are input to separator 112b, these performances can also be concatenated along the frequency axis. This results in a "3M dimensions × time" performance. Alternatively, the target sound direction enhancement time-frequency performance, the target sound direction masking time-frequency performance, and the mixed sound time-frequency performance can be concatenated along an axis different from the time and frequency axes. This results in a "M dimensions × time × 3" performance. The target sound direction enhancement time-frequency performance, the target sound direction masking time-frequency performance, and the mixed sound time-frequency performance can also be weighted. The weighted target sound direction enhancement time-frequency performance, the weighted target sound direction masking time-frequency performance, and the weighted mixed sound time-frequency performance can also be combined. The weights can also be estimated from the learned model 112.

[0104] The processing of decoder 112c is the same as in implementation method 1.

[0105] In this way, the target sound signal output unit 180 outputs the target sound signal using the selected sound signal, the target sound direction enhancement sound signal, the target sound direction masking sound signal, and the learned model 112.

[0106] Next, a flowchart will be used to explain the processes performed by the information processing device 100.

[0107] Figure 9 This is a flowchart illustrating an example of the processing performed by the information processing apparatus of Embodiment 2. Figure 9 processing and Figure 5 The difference in processing lies in the execution of steps S11a and 17a. Therefore, in Figure 9 The steps S11a and 17a will be explained in the text. Furthermore, the explanation of the processes other than steps S11a and 17a is omitted.

[0108] (Step S11a) The selection unit 190 uses the mixed sound signal and the sound source location information 111 to select the sound signal of the channel of the target sound direction.

[0109] (Step S17a) The target sound signal output unit 180 outputs the target sound signal using the selected sound signal, the target sound direction enhancement sound signal, the target sound direction masking sound signal and the learned model 112.

[0110] In addition, step S11a can be executed before step S17a, and can be executed at any time.

[0111] Here, the generation of the learned model 112 will be explained. The learning device 200 learns using a learning signal that includes the sound signal of the channel containing the target sound direction (i.e., the mixed sound signal of the target sound direction). For example, this learning signal may also be generated by the processing execution unit 240.

[0112] The learning device 200 learns the difference between the enhanced sound signal in the target sound direction and the mixed sound signal in the target sound direction. Furthermore, the learning device 200 learns the difference between the masked sound signal in the target sound direction and the mixed sound signal in the target sound direction. The learning device 200 learns that the signal with the largest difference is the target sound signal. In this way, the learning device 200 learns and generates a learned model 112.

[0113] According to Embodiment 2, the information processing device 100 uses the learned model 112 obtained through learning, thereby enabling it to output the target sound signal.

[0114] Implementation Method 3

[0115] Next, Embodiment 3 will be described. In Embodiment 3, the differences from Embodiment 1 will be mainly explained. Furthermore, in Embodiment 3, the descriptions of matters common to Embodiment 1 will be omitted.

[0116] Figure 10This is a block diagram illustrating the functions of the information processing apparatus of Embodiment 3. The information processing apparatus 100 also includes a reliability calculation unit 191.

[0117] A portion or all of the reliability calculation unit 191 may also be implemented by processing circuitry. Alternatively, a portion or all of the reliability calculation unit 191 may be implemented as a module of a program executed by the processor 101.

[0118] The reliability calculation unit 191 calculates the reliability F of the masking characteristic quantity using a pre-set method. i The reliability F of the masking feature quantity i This can also be referred to as the reliability F of directional masking. i The pre-defined method is represented by the following formula (3). ω represents the angular range of the target sound direction. θ represents the angular range of the sound generation direction.

[0119]

[0120] Reliability F i It is a matrix of the same size as the directional masking. Additionally, the reliability F... i It can also be input into the learning device 200.

[0121] The target audio signal output unit 180 uses a reliability of F. i The target sound direction enhancement sound signal, the target sound direction masking sound signal, and the target sound signal output by the learned model 112.

[0122] Next, the processing of the encoder 112a, separator 112b and decoder 112c contained in the learned model 112 will be explained.

[0123] Encoder 112a performs the following processing based on the processing in Embodiment 1. Encoder 112a will increase reliability F i The time-frequency representation (FT) is calculated by multiplying the number of frequency intervals F and the number of frames T. The number of frequency intervals F represents the number of elements along the frequency axis of the time-frequency representation. The number of frames T is the number of segments obtained by dividing the mixed audio signal into pre-defined time intervals.

[0124] When the target sound direction enhancement time-frequency representation and the time-frequency representation FT are consistent, the time-frequency representation FT is treated as the mixed sound time-frequency representation of Implementation Method 2 in subsequent processing. When the target sound direction enhancement time-frequency representation and the time-frequency representation FT are inconsistent, encoder 112a performs transformation matrix transformation processing. Specifically, encoder 112a transforms the reliability F... i The number of elements in the frequency axis direction is converted into the number of elements in the frequency axis direction for enhancing the frequency performance of the target sound direction.

[0125] When the target sound direction enhancement time-frequency performance and the time-frequency performance (FT) are consistent, the separator 112b performs the same processing as the separator 112b in Embodiment 2.

[0126] When the time-frequency response (FT) of the target sound direction enhancement is inconsistent with that of the time-frequency response (FT), the reliability F of the splitter 112b after the number of elements in the frequency axis direction is converted is... i The target sound direction enhancement time-frequency performance is integrated. For example, the separator 112b uses the Attention method shown in Non-Patent Document 2 for integration. The separator 112b estimates an "M-dimensional × time" masking matrix based on the target sound direction enhancement time-frequency performance and the target sound direction masking time-frequency performance obtained by integration.

[0127] The processing of decoder 112c is the same as in implementation method 1.

[0128] Thus, the target audio signal output unit 180 uses a reliability F i The target sound direction enhancement sound signal, the target sound direction masking sound signal, and the target sound signal output by the learned model 112.

[0129] Next, a flowchart will be used to explain the processes performed by the information processing device 100.

[0130] Figure 11 This is a flowchart illustrating an example of the processing performed by the information processing apparatus of Embodiment 3. Figure 11 processing and Figure 5 The difference in processing lies in the execution of steps S15b and 17b. Therefore, in Figure 11 The steps S15b and 17b will be explained in the text. Furthermore, the explanation of the processes other than steps S15b and 17b is omitted.

[0131] (Step S15b) The reliability calculation unit 191 calculates the reliability F of the masking characteristic quantity. i .

[0132] (Step S17b) The target audio signal output unit 180 uses reliability F i The target sound direction enhancement sound signal, the target sound direction masking sound signal, and the target sound signal output by the learned model 112.

[0133] Here, the generation of the learned model 112 will be explained. When the learning device 200 is learning, it uses a reliability F. i Learning is performed. The learning device 200 can also use the reliability F obtained from the information processing device 100. iLearning is performed. The learning device 200 can also use the reliability F stored in its volatile or non-volatile storage device. i Learning is performed. The learning device 200 uses a reliability of F. i The system determines how much of the target sound direction should be considered to mask the sound signal. The learning device 200 learns the parameters used to make this decision, thereby generating a learned model 112.

[0134] According to implementation method 3, the target sound direction enhancement sound signal and the target sound direction masking sound signal are input into the learned model 112. The target sound direction masking sound signal is generated based on masking feature quantities. The learned model 112 uses the reliability F of the masking feature quantities. i The system determines how much of the target sound direction should be considered to mask the sound signal. The learned model 112 outputs the target sound signal based on this determination. Thus, the information processing device 100 transmits the target sound signal with a reliability F. i When input into the learned model 112, it can output a more appropriate target sound signal.

[0135] Implementation Method 4

[0136] Next, Embodiment 4 will be described. In Embodiment 4, the differences from Embodiment 1 will be mainly explained. Furthermore, in Embodiment 4, the description of matters common to Embodiment 1 will be omitted.

[0137] Figure 12 This is a block diagram illustrating the function of the information processing apparatus of Embodiment 4. The information processing apparatus 100 also includes a noise range detection unit 192.

[0138] A portion or all of the noise interval detection unit 192 can also be implemented by a processing circuit. Alternatively, a portion or all of the noise interval detection unit 192 can also be implemented as a module of a program executed by the processor 101.

[0139] The noise interval detection unit 192 detects noise intervals based on the target sound direction enhancement sound signal. For example, when detecting noise intervals, the noise interval detection unit 192 uses the method described in Patent Document 2. For example, after detecting a speech interval based on the target sound direction enhancement sound signal, the noise interval detection unit 192 corrects the start time and end time of the speech interval, thereby determining the speech interval. The noise interval detection unit 192 excludes the determined speech interval from the interval representing the target sound direction enhancement sound signal, thereby detecting the noise interval. Here, the detected noise interval can also be input to the learning device 200.

[0140] The target sound signal output unit 180 outputs the target sound signal using the detected noise range, the target sound direction enhancement sound signal, the target sound direction masking sound signal, and the learned model 112.

[0141] Next, the processing of the encoder 112a, separator 112b and decoder 112c contained in the learned model 112 will be explained.

[0142] The encoder 112a performs the following processing based on the processing in Embodiment 1. The encoder 112a estimates the non-target sound time-frequency response in "M dimensions × time" based on the signal corresponding to the noise region of the enhanced sound signal in the target sound direction. For example, the encoder 112a may also estimate the non-target sound time-frequency response using the power spectrum estimated by the STFT. Furthermore, for example, the encoder 112a may also use a one-dimensional convolution operation to estimate the non-target sound time-frequency response. When this estimation is performed, the non-target sound time-frequency response can be projected to the same time-frequency response space or to a different time-frequency response space. Additionally, this estimation is described, for example, in Non-Patent Document 1.

[0143] Separator 112b integrates the time-frequency representation of the non-target sound and the time-frequency representation of the target sound's directional enhancement. For example, separator 112b uses the Attention method shown in Non-Patent Document 2 for integration. Separator 112b estimates an "M-dimensional × time" masking matrix based on the time-frequency representation of the target sound's directional enhancement and the time-frequency representation of the target sound's directional masking obtained through integration.

[0144] Additionally, for example, the separator 112b can estimate the tendency of noise based on the time-frequency performance of the non-target sound.

[0145] The processing of decoder 112c is the same as in implementation method 1.

[0146] Next, a flowchart will be used to explain the processes performed by the information processing device 100.

[0147] Figure 13 This is a flowchart illustrating an example of the processing performed by the information processing apparatus of Embodiment 4. Figure 13 processing and Figure 5 The difference in processing lies in the execution of steps S16c and 17c. Therefore, in Figure 13 The steps S16c and 17c are explained in the text. Furthermore, the explanation of the processes other than steps S16c and 17c is omitted.

[0148] (Step S16c) The noise interval detection unit 192 detects the interval representing noise, i.e., the noise interval, based on the enhanced sound signal of the target sound direction.

[0149] (Step S17c) The target sound signal output unit 180 outputs the target sound signal using the noise range, the target sound direction enhancement sound signal, the target sound direction masking sound signal and the learned model 112.

[0150] Here, the generation of the learned model 112 will be explained. When performing learning, the learning device 200 uses a noise range for learning. The learning device 200 may also use a noise range obtained from the information processing device 100. The learning device 200 may also use a noise range detected by the processing execution unit 240. The learning device 200 learns the tendency of noise based on the noise range. Taking into account the noise tendency, the learning device 200 learns to output a target sound signal based on the target sound direction enhancement sound signal and the target sound direction masking sound signal. Thus, the learning device 200 performs learning, thereby generating the learned model 112.

[0151] According to embodiment 4, a noise range is input to the learned model 112. The learned model 112 estimates the tendency of noise contained in the target sound direction enhancement sound signal and the target sound direction masking sound signal based on the noise range. Taking into account the noise tendency, the learned model 112 outputs a target sound signal based on the target sound direction enhancement sound signal and the target sound direction masking sound signal. As a result, the information processing device 100 outputs a target sound signal considering the noise tendency, and therefore, it is able to output a more appropriate target sound signal.

[0152] The features described above in the various embodiments can be appropriately combined with each other.

[0153] Label Explanation

[0154] 100: Information processing device; 101: Processor; 102: Volatile storage device; 103: Non-volatile storage device; 111: Sound source location information; 112: Learned model; 120: Acquisition unit; 130: Sound feature extraction unit; 140: Enhancement unit; 150: Estimation unit; 160: Masking feature extraction unit; 170: Generation unit; 180: Target sound signal output unit; 190: Selection unit; 191: Reliability calculation unit; 192: Noise interval detection unit; 200: Learning device; 211: Sound data storage unit; 212: Impulse response storage unit; 213: Noise storage unit; 220: Impulse response application unit; 230: Mixing unit; 240: Processing execution unit; 250: Learning unit.

Claims

1. An information processing apparatus, comprising: The acquisition unit acquires the location information of the target sound source, i.e., the sound source location information, the signal representing the mixed sound containing the target sound and the interference sound, i.e., the mixed sound signal, and the learned model; The sound feature extraction unit extracts multiple sound features based on the mixed sound signal. The multiple sound features are time series of power spectra obtained by performing a short-time Fourier transform on the mixed sound signal. The enhancement unit enhances the sound feature quantity of the direction of the target sound, i.e. the direction of the target sound, among the plurality of sound feature quantities, based on the sound source location information. The estimation unit estimates the direction of the target sound based on the plurality of sound feature quantities and the sound source location information; The masking feature extraction unit extracts the masking feature of the target sound direction based on the estimated target sound direction and the multiple sound feature quantities. The generation unit generates an enhanced sound signal (i.e., a target sound direction enhanced sound signal) based on the enhanced sound feature quantity, and generates a masked sound signal (i.e., a target sound direction masked sound signal) based on the masking feature quantity. as well as The target sound signal output unit uses the target sound direction enhancement sound signal, the target sound direction masking sound signal, and the learned model to output a signal representing the target sound, namely the target sound signal.

2. The information processing apparatus according to claim 1, wherein, The information processing device further includes a selection unit that uses the mixed sound signal and the sound source location information to select the sound signal of the channel with the target sound direction. The target sound signal output unit outputs the target sound signal using the selected sound signal, the target sound direction enhancement sound signal, the target sound direction masking sound signal, and the learned model.

3. The information processing apparatus according to claim 1, wherein, The information processing device also includes a reliability calculation unit, which calculates the reliability of the masking feature quantity using a pre-set method. The target sound signal output unit uses the reliability, the target sound direction enhancement sound signal, the target sound direction masking sound signal, and the learned model to output the target sound signal.

4. The information processing apparatus according to claim 2, wherein, The information processing device also includes a reliability calculation unit, which calculates the reliability of the masking feature quantity using a pre-set method. The target sound signal output unit uses the reliability, the target sound direction enhancement sound signal, the target sound direction masking sound signal, and the learned model to output the target sound signal.

5. The information processing apparatus according to any one of claims 1 to 4, wherein, The mixed sound contains noise.

6. The information processing apparatus according to claim 5, wherein, The information processing device further includes a noise interval detection unit, which detects the interval representing the noise, i.e., the noise interval, based on the enhanced sound signal from the direction of the target sound. The target sound signal output unit uses the noise range, the target sound direction enhancement sound signal, the target sound direction masking sound signal, and the learned model to output the target sound signal.

7. An output method, wherein, The information processing device acquires the location information of the target sound source (i.e., sound source location information), a signal representing the mixed sound containing the target sound and interference sounds (i.e., mixed sound signal), and a learned model. The information processing device extracts multiple sound feature quantities based on the mixed sound signal. These multiple sound feature quantities are time series of the power spectrum obtained by performing a short-time Fourier transform on the mixed sound signal. The information processing device enhances the sound feature quantity representing the direction of the target sound (i.e., the direction of the target sound) among the plurality of sound feature quantities based on the sound source location information. The information processing device estimates the direction of the target sound based on the multiple sound feature quantities and the sound source location information. The information processing device extracts the masking feature value, i.e., the feature value of the target sound direction when the feature value of the target sound direction is masked, based on the estimated target sound direction and the multiple sound feature values. The information processing device generates, based on the enhanced sound feature quantity, a sound signal with enhanced sound in the target sound direction, i.e., a target sound direction enhanced sound signal; and generates, based on the masking feature quantity, a sound signal with masked sound in the target sound direction, i.e., a target sound direction masked sound signal. The information processing device uses the target sound direction enhancement sound signal, the target sound direction masking sound signal, and the learned model to output a signal representing the target sound, namely the target sound signal.

8. A recording medium having an output program that causes an information processing apparatus to perform the following processing: The system obtains the location information of the target sound source (i.e., sound source location information), the signal representing the mixed sound containing the target sound and the interference sound (i.e., the mixed sound signal), and the learned model. Multiple sound features are extracted from the mixed sound signal. These multiple sound features are time series of the power spectrum obtained by performing a short-time Fourier transform on the mixed sound signal. Based on the sound source location information, the sound feature quantity representing the direction of the target sound (i.e., the direction of the target sound) among the multiple sound feature quantities is enhanced. The direction of the target sound is estimated based on the multiple sound feature quantities and the sound source location information. Based on the estimated target sound direction and the multiple sound feature quantities, the feature quantity of the target sound direction that is masked is extracted, i.e., the masking feature quantity. Based on the enhanced sound feature quantity, an enhanced sound signal is generated for the target sound direction, i.e., the target sound direction enhanced sound signal. Based on the masking feature quantity, a masked sound signal is generated for the target sound direction, i.e., the target sound direction masked sound signal. The target sound signal is generated by using the target sound direction to enhance the sound signal, the target sound direction to mask the sound signal, and the signal representing the target sound output by the learned model.