Voice section detection device, learning device, and voice section detection program
The voice section detection device enhances accuracy in diverse noise environments by calculating voice presence scores from both acoustic and non-acoustic features, effectively addressing the challenges of noise variability.
Patent Information
- Application Number
- JP2022040292
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-15
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-03-15
AI Technical Summary
Existing voice segment detection devices face challenges in maintaining accuracy in diverse and complex noise environments, where noise characteristics vary significantly and change in real time.
The proposed voice section detection device incorporates an acquisition unit for acoustic and non-acoustic signals, followed by feature calculation units for acoustic, non-acoustic, voice emphasis, and voice presence/absence features. These features are then used to calculate a voice presence score, which is compared to a threshold to detect voice sections.
This approach significantly improves the detection accuracy of voice sections in various noise environments by effectively utilizing both acoustic and non-acoustic signals, leading to robust voice presence scoring.
Smart Images

Figure 0007693593000007 
Figure 0007693593000008 
Figure 0007693593000009
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to a voice section detection device, a learning device, and a voice section detection program.
Background Art
[0002] Voice activity detection (VAD) is a technique for detecting a voice section including a user's utterance from an input signal. Voice activity detection is mainly used to improve the recognition accuracy of speech recognition, or in the field of speech coding, it is used to assist data compression in non-voice sections.
[0003] In voice activity detection, a process of detecting a voice section including predetermined voice from a time interval of an input signal is required. For example, whether a frame to be processed is a voice section including voice such as an utterance or not, a predetermined voice section is detected from an input acoustic signal using a pre-trained model.
[0004] In order to further improve the processing accuracy of voice activity detection in a processing-difficult environment such as a noisy environment, a voice activity detection device that uses both an acoustic signal and a non-acoustic signal such as a lip video signal as input signals to detect a voice section has been proposed. For example, the technique disclosed in Non-Patent Document 1 takes an acoustic signal and a lip video signal as inputs, calculates a speech score based on a deep neural network from integrated features, and detects a voice section.
Prior Art Documents
Non-Patent Documents
[0005]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] Recently, due to the full-scale commercialization of voice recognition and voice interfaces and their use in a wide variety of devices, voice segment detection devices are exposed to far more diverse and complex noise environments than before. For example, the noise characteristics collected from each device are different, and the noise characteristics change in real time during use while moving from the user by mobile devices. There are still limitations in maintaining the accuracy of voice segment detection in such a wide variety of noise environments, and there is a need to improve the detection accuracy. The problem to be solved by the present invention is to provide a voice segment detection device, a learning device, and a voice segment detection program capable of improving the detection accuracy in a wide variety of noise environments.
Means for Solving the Problems
[0007] The voice section detection device according to the embodiment includes an acquisition unit, an acoustic feature calculation unit, a non-acoustic feature calculation unit, a voice emphasis feature calculation unit, a voice presence / absence feature calculation unit, a voice presence score calculation unit, and a detection unit. The acquisition unit acquires an acoustic signal and a non-acoustic signal related to a voice generation source. The acoustic feature calculation unit calculates an acoustic feature based on the acoustic signal. The non-acoustic feature calculation unit calculates a non-acoustic feature based on the non-acoustic signal. The voice feature calculation unit calculates a voice emphasis feature based on the acoustic feature and the non-acoustic feature. The voice presence / absence feature calculation unit calculates a voice presence / absence feature based on the acoustic feature and the non-acoustic feature. The voice presence score calculation unit calculates a voice presence score based on the voice emphasis feature and the voice presence / absence feature. The detection unit detects a voice section that is a time section in which voice is emitted and / or a non-voice section that is a time section in which no voice is emitted based on a comparison with a threshold value of the voice presence score.
Brief Description of Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Embodiments for Carrying Out the Invention
[0009] Hereinafter, a voice section detection device, a learning device, and a voice section detection program according to this embodiment will be described with reference to the drawings.
[0010] (Voice section detection device) FIG. 1 is a diagram showing a configuration example of a voice section detection device 100. The voice section detection device 100 is a computer that detects a voice section of an input signal. As shown in FIG. 1, the voice section detection device 100 includes a processing circuit 11, a storage device 12, an input device 13, a communication device 14, a display device 15, and an acoustic device 16.
[0011] The processing circuit 11 includes a processor such as a CPU (Central Processing Unit) and a memory such as a RAM (Random Access Memory). The processing circuit 11 executes a voice section detection process for detecting a voice section of an input signal by executing a voice section detection program stored in the storage device 12. The voice section detection program is recorded on a non-transitory computer-readable recording medium. The processing circuit 11 reads and executes the voice section detection program from the recording medium to implement an input signal acquisition unit 111, an acoustic feature calculation unit 112, a non-acoustic feature calculation unit 113, an integrated feature calculation unit 114, a voice emphasis feature calculation unit 115, a voice presence / absence feature calculation unit 116, a voice emphasis signal calculation unit 117, a voice presence score calculation unit 118, a voice section detection unit 119, and an output control unit 120. Note that the voice section detection program may include a plurality of modules that implement the functions of the respective units 111 to 120 separately.
[0012] The hardware implementation of the processing circuit 11 is not limited to only the above-described aspects. For example, it may be configured by a circuit such as an application specific integrated circuit (ASIC) that realizes the input signal acquisition unit 111, the acoustic feature calculation unit 112, the non-acoustic feature calculation unit 113, the integrated feature calculation unit 114, the voice emphasis feature calculation unit 115, the voice presence / absence feature calculation unit 116, the voice emphasis signal calculation unit 117, the voice presence score calculation unit 118, the voice section detection unit 119, and / or the output control unit 120. The input signal acquisition unit 111, the acoustic feature calculation unit 112, the non-acoustic feature calculation unit 113, the integrated feature calculation unit 114, the voice emphasis feature calculation unit 115, the voice presence / absence feature calculation unit 116, the voice emphasis signal calculation unit 117, the voice presence score calculation unit 118, the voice section detection unit 119, and / or the output control unit 120 may be implemented in a single integrated circuit or may be individually implemented in a plurality of integrated circuits.
[0013] The input signal acquisition unit 111 acquires an acoustic signal and a non-acoustic signal related to a voice generation source. The acoustic signal and the non-acoustic signal are time series signals and are temporally synchronized in frame units. The acoustic signal is a signal related to the voice of a speaker who is the voice generation source. Specifically, the acoustic signal includes a voice signal derived from the speaker's vocalization and a noise signal derived from noise. The non-acoustic signal is a signal other than the acoustic signal related to the speaker, which is collected substantially simultaneously with the acoustic signal. For example, the non-acoustic signal is an image signal related to the speaker who is speaking, a sensor signal related to the physiological reaction of the speaker's lips and facial muscles due to speech, an electroencephalogram, etc. Although it is assumed that the acoustic signal and the non-acoustic signal are derived from the same voice generation source, they do not necessarily have to be exactly the same if there is a correlation between the acoustic signal and the non-acoustic signal.
[0014] The acoustic feature calculation unit 112 calculates a feature amount of an acoustic signal (hereinafter referred to as an acoustic feature). The acoustic feature has a value based on the acoustic signal and has a value correlated with the voice by the speaker. The acoustic feature is calculated for each frame. As an example, the acoustic feature is calculated using a first trained model. The first trained model is a neural network trained to input an acoustic signal and output an acoustic feature. The first trained model is stored in the storage device 12 or the like.
[0015] The non-acoustic feature calculation unit 113 calculates a feature amount of a non-acoustic signal (hereinafter referred to as a non-acoustic feature). The non-acoustic feature has a value based on the non-acoustic signal and has a characteristic value correlated with the voice by the speaker. The non-acoustic feature is calculated for each frame. As an example, the non-acoustic feature is calculated using a second trained model. The second trained model is a neural network trained to input a non-acoustic signal and output a non-acoustic feature. The second trained model is stored in the storage device 12 or the like.
[0016] Note that the first trained model and the second trained model are neural networks trained to reduce a first loss related to a difference between an acoustic feature and a non-acoustic feature regarding the same voice source.
[0017] The integrated feature calculation unit 114 calculates an integrated feature based on the acoustic feature and the non-acoustic feature. The integrated feature is calculated for each frame. As an example, an integrated feature calculated as the sum of the acoustic feature and the non-acoustic feature is calculated.
[0018] The voice emphasis feature calculation unit 115 calculates a feature amount (hereinafter referred to as a voice emphasis feature) regarding the emphasized voice signal based on the integrated feature. Here, the voice signal means an acoustic signal among the acoustic signals that is derived from the utterance by the speaker. The voice emphasis feature corresponds to the voice feature regarding the voice signal generated by separating and emphasizing only the voice signal by the speaker from the acoustic signal including noise. The voice emphasis feature is calculated for each frame. As an example, the voice emphasis feature is calculated using the third pre-trained model. The third pre-trained model is a neural network trained to input the integrated feature and output the voice emphasis feature. The third pre-trained model is stored in the storage device 12 or the like.
[0019] The voice presence / absence feature calculation unit 116 calculates a feature amount (hereinafter referred to as a voice presence / absence feature) regarding the presence / absence of the voice signal based on the integrated feature. The voice presence / absence feature has a value representing the presence / absence of the voice signal by the speaker from the acoustic signal including noise. The voice presence / absence feature is calculated for each frame. As an example, the voice presence / absence feature is calculated using the fourth pre-trained model. The fourth pre-trained model is a neural network trained to input the integrated feature and output the voice presence / absence feature. The fourth pre-trained model is stored in the storage device 12 or the like.
[0020] Note that it is not essential to provide the integrated feature calculation unit 114. That is, the voice emphasis feature calculation unit 115 does not necessarily have to calculate the voice emphasis feature from the integrated feature based on the acoustic feature and the non-acoustic feature. Instead, it may directly calculate the voice emphasis feature from the acoustic feature and the non-acoustic feature, or calculate the voice emphasis feature from other intermediate outputs based on the acoustic feature and the non-acoustic feature. As an example, the voice emphasis feature calculation unit 115 can calculate the voice emphasis feature from the acoustic feature and the non-acoustic feature of the processing target using a neural network trained to input the acoustic feature and the non-acoustic feature and output the voice emphasis feature. Similarly, the voice emphasis feature calculation unit 115 does not necessarily have to calculate the voice presence / absence feature from the integrated feature based on the acoustic feature and the non-acoustic feature. Instead, it may directly calculate the voice presence / absence feature from the acoustic feature and the non-acoustic feature, or calculate the voice presence / absence feature from other intermediate outputs based on the acoustic feature and the non-acoustic feature. As an example, the voice presence / absence feature calculation unit 116 can calculate the voice presence / absence feature from the acoustic feature and the non-acoustic feature of the processing target using a neural network trained to input the acoustic feature and the non-acoustic feature and output the voice presence / absence feature. In the following description, it is assumed that the integrated feature calculation unit 114 is provided.
[0021] The voice emphasis signal calculation unit 117 calculates a voice emphasis signal based on the voice emphasis feature and the voice presence / absence feature. The voice emphasis signal is a restored voice signal in which only the voice signal by the speaker in the acoustic signal including noise is separated and emphasized. The voice emphasis signal is calculated for each frame. As an example, the voice emphasis signal is calculated using the fifth trained model. The fifth trained model is a neural network trained to input the voice emphasis feature and the voice presence / absence feature and output the voice emphasis signal. The fifth trained model is stored in the storage device 12 or the like.
[0022] Note that it is not essential to provide the voice emphasis signal calculation unit 117. If the voice emphasis signal is not used in the voice section detection device 100, the voice emphasis signal calculation unit 117 may not be provided.
[0023] The voice presence score calculation unit 118 calculates a voice presence score based on the voice emphasis feature and the voice presence / absence feature. The voice presence score is used as a measure for discriminating between a voice section and a non-voice section. Note that the voice section is a time section in which voice is emitted within the time section of the input signal, and the non-voice section is a time section in which no voice is emitted within the time section of the input signal. The voice presence score is calculated for each frame. As an example, the voice presence score is defined by the probability that the target frame is a voice section and / or the probability that it is a non-voice section. As an example, the voice presence score is calculated using the sixth trained model. The sixth trained model is a neural network trained to input the voice emphasis feature and the voice presence / absence feature and output the voice presence score. The sixth trained model is stored in the storage device 12 or the like.
[0024] The third trained model and the fifth trained model are neural networks trained to reduce a second loss related to the difference between the correct voice signal and the voice emphasis signal regarding the voice signal by the speaker included in the acoustic signal.
[0025] The fourth trained model and the sixth trained model are neural networks trained to reduce a third loss related to the difference between the correct label regarding the voice section and the non-voice section and the voice presence score.
[0026] The voice section detection unit 119 detects a voice section, which is a time section in which voice is emitted, and / or a non-voice section, which is a time section in which no voice is emitted, based on a comparison with a threshold value of the voice presence score.
[0027] The output control unit 120 displays various information via the display device 15 and the audio device 16. For example, the output control unit 120 displays an image signal on the display device 15 or outputs an audio signal via the audio device 16.
[0028] The memory device 12 is composed of a ROM (Read Only Memory), HDD (Hard Disk Drive), SSD (Solid State Drive), integrated circuit memory device, etc. The memory device 12 stores various calculation results by the processing circuit 11 and the voice section detection program executed by the processing circuit 11. The memory device 12 is an example of a computer-readable recording medium.
[0029] The input device 13 inputs various commands from the user. As the input device 13, a keyboard, mouse, various switches, touch pad, touch panel display, etc. can be used. The output signal from the input device 13 is supplied to the processing circuit 11. Note that the input device 13 may be a computer connected to the processing circuit 11 via wire or wireless.
[0030] The communication device 14 is an interface for performing information communication between the voice section detection device 100 and an external device connected via a network. The communication device 14 receives, for example, acoustic signals and non-acoustic signals from a device that collects acoustic signals and non-acoustic signals, or receives the first learned model, the second learned model, the third learned model, the fourth learned model, the fifth learned model, and the sixth learned model from a learning device described later.
[0031] The display device 15 displays various information. As the display device 15, a CRT (Cathode-Ray Tube) display, liquid crystal display, organic EL (Electro Luminescence) display, LED (Light-Emitting Diode) display, plasma display, or any other display known in the art can be appropriately used. The display device 15 may be a projector.
[0032] The acoustic device 16 converts an electrical signal into sound and emits it. As the acoustic device 16, a magnetic speaker, dynamic speaker, condenser speaker, or any other speaker known in the art can be appropriately used.
[0033] Next, an example of the voice section detection process by the processing circuit 11 of the voice section detection device 100 will be described. For the sake of specific description below, it is assumed that the non-acoustic signal is an image signal.
[0034] FIG. 2 is a diagram showing an example of the flow of the voice section detection process by the processing circuit 11. FIG. 3 is a diagram schematically showing the voice section detection process. The voice section detection process is executed by the processing circuit 11 operating according to a voice section detection program stored in the storage device 12 or the like.
[0035] As shown in FIGS. 2 and 3, the input signal acquisition unit 111 acquires an input signal including an acoustic signal and an image signal (step SA1). The input signal is a video signal including an acoustic signal and an image signal related to the same voice source.
[0036] FIG. 4 is a diagram showing an example of an input signal (video signal), an acoustic signal, and an image signal. As shown in FIG. 4, the video signal is a time-series signal including a time-series acoustic signal and a time-series image signal that are temporally synchronized. The length of the time interval of the video signal is not particularly limited, but for example, it is assumed to be a frame length of about 10 seconds.
[0037] The video signal is collected by a video camera device including a microphone and an imaging device. The acoustic signal is collected by the microphone. The microphone collects the voice regarding the speech of the speaker, converts the sound pressure of the collected voice into an analog electrical signal (acoustic signal), and A / D-converts the acoustic signal into a digital electrical signal (acoustic signal) in the time domain. The acoustic signal in the time domain is acquired by the input signal acquisition unit 111 and converted into an acoustic signal in the frequency domain by short-time Fourier transform or the like. The image signal is collected substantially simultaneously with the acoustic signal. The image signal is collected by an imaging device including a plurality of imaging elements such as a CCD (Charge Coupled Device). The imaging device optically photographs the speaking speaker and generates a digital image signal (image data) in the spatial domain regarding the speaker in units of frames. The image signal is required to be correlated with the speech of the speaker. The image frame may include, as a photographing target, at least a lip region whose form is deformed according to the speech, and may also include the entire face region of the speaker. The image signal is acquired in units of frames by the input signal acquisition unit 111.
[0038] Here, the time-series acoustic signal A and the time-series image signal V are defined according to the following formula (1). The time-series acoustic signal A is an acoustic signal having a T dimension in the time domain and an F dimension in the frequency domain of the frame to be processed. The image signal V is an image signal having dimensions of time T, vertical width H, horizontal width W, and color channel C.
[0039]
Equation
[0040] When step SA1 is performed, the acoustic feature calculation unit 112 calculates an acoustic feature E A from the acoustic signal A acquired in step SA1 using the first learned model (step SA2). The acoustic feature E A is calculated based on the acoustic signal A for each frame. The acoustic feature E A is time-series data. The first learned model inputs the acoustic signal A and outputs the acoustic feature E AIt is a neural network trained to output. As such a neural network, for example, an encoder network trained to convert an acoustic signal A into an acoustic feature E A is used. The first learned model is generated by a learning device described later.
[0041] The relationship between the acoustic signal and the acoustic feature here is as follows. The acoustic signal is time-series waveform data of the sound pressure value of the voice uttered by the speaker. The acoustic signal is correlated with the voice uttered by the speaker. For example, the peak value of the acoustic signal has a relatively high value when the speaker is pronouncing, and has a relatively low value when the speaker is not pronouncing. The acoustic feature is designed to have a value correlated with the peak value of the acoustic signal, in other words, to discriminate between the voice component and the silent component included in the acoustic signal. For example, the higher the peak value of the acoustic signal, the higher the value of the acoustic feature, and the lower the peak value of the acoustic signal, the lower the value of the acoustic feature.
[0042] When step SA2 is performed, the non-acoustic feature calculation unit 113 uses the second learned model to calculate an image feature E V from the image signal V acquired in step SA1 (step SA3). The image feature E V is calculated based on the image signal V for each frame. That is, the image feature E V is time-series data. The second learned model is a neural network trained to input the image signal V and output the image feature E V As such a neural network, for example, an encoder network trained to convert the image signal V into the image feature E V is used. The second learned model is generated by a learning device described later.
[0043] The relationship between the image signal and the image features here is as follows. The image signal is correlated with the form of the face partial region when the speaker is pronouncing. The image features are designed to discriminate between the speech component and the non-speech component included in the image signal. Specifically, the lip region of the speaker represented by the image signal has different forms when the speaker is uttering a sound and when not. The image features are designed such that their values are correlated with the form of the face partial region of the speaker. For example, the higher the value of the image features when the speaker has their mouth open, and the lower the value of the image features when the speaker has their mouth closed.
[0044] The first learned model and the second learned model are neural networks trained such that the difference between the acoustic feature EA and the image feature EV is small with respect to the normal input for the same voice source. The first learned model and the second learned model are generated by the learning device described later.
[0045] The order of step SA2 and step SA3 is not particularly limited, and step SA2 may be executed after step SA3, or step SA2 and step SA3 may be executed in parallel.
[0046] When step SA2 and step SA3 are performed, the integrated feature calculation unit 114 calculates an integrated feature of the acoustic feature and the non-acoustic feature based on the acoustic feature calculated in step SA2 and the image feature calculated in step SA3 (step SA4). In step S4, the integrated feature calculation unit 114 calculates the integrated feature of the acoustic feature E A and the image feature E V . The integrated feature Z AV is calculated for each frame time t. Specifically, the integrated feature Z AV is calculated as the sum of the acoustic feature E A and the image feature E V as represented by the following formula (2).
[0047]
Equation
[0048] Here, E A (t、d) is an acoustic feature vector at frame time t ∈ {1, 2, …, T} and compressed coordinates d ∈ {1, 2, …, D} in D dimensions, and is an example of an acoustic feature. E V (t、d) is an image feature vector at frame time t ∈ {1, 2, …, T} and compressed coordinates d ∈ {1, 2, …, D} in D dimensions, and is an example of an image feature. Z AV (t、d) is an integrated feature vector at frame time t ∈ {1, 2, …, T} and compressed coordinates d ∈ {1, 2, …, D} in D dimensions, and is an example of an integrated feature.
[0049] When step SA4 is performed, the voice emphasis feature calculation unit 115 uses the third learned model to obtain the integrated feature Z obtained in step SA4 AV to calculate the voice emphasis feature E SE (step SA5). The voice emphasis feature E SE has a feature value in which only the voice signal derived from the speaker's utterance is separated and emphasized from the acoustic signal A. The voice emphasis feature E SE is calculated based on the integrated feature Z AV for the frame.
[0050] When step SA5 is performed, the voice presence / absence feature calculation unit 116 uses the fourth learned model to obtain the integrated feature Z obtained in step SA4 AV to calculate the voice presence / absence feature E VAD (step SA6). The voice presence / absence feature E VAD has a feature value representing the presence or absence of the voice signal derived from the speaker's utterance in the acoustic signal A. The voice presence / absence feature E VAD is calculated based on the integrated feature Z AV for the frame.
[0051] The order of step SA5 and step SA6 is not particularly limited, and step SA5 may be executed after step SA6, or step SA5 and step SA6 may be executed in parallel.
[0052] When step SA5 and step SA6 are performed, the voice emphasis signal calculation unit 117 uses the fifth learned model to calculate the voice emphasis signal from the voice emphasis feature E obtained in step SA5 SE and the voice presence / absence feature E obtained in step SA6 VAD (step SA7). The voice emphasis signal Y SE represents a restored voice signal in which only the voice signal derived from the speaker's utterance is separated and emphasized from the acoustic signal A including noise. The voice emphasis signal Y SE is calculated for each frame based on the voice emphasis feature E SE and the voice presence / absence feature E VAD .
[0053] When step SA5 and step SA6 are performed, the voice presence score calculation unit 118 uses the sixth learned model to calculate the voice presence score Y SE from the voice emphasis feature E obtained in step SA5 VAD and the voice presence / absence feature E obtained in step SA6 VAD (step SA8). The voice presence score Y VAD represents the score indicating the presence of voice in the frame. The voice presence score Y is calculated for each frame based on the voice emphasis feature E SE and the voice presence / absence feature E VAD .
[0054] The third learned model and the fifth learned model are neural networks trained to reduce the second loss function regarding the difference between the correct voice signal related to the voice signal and the voice emphasis signal Y SE . As the neural network, for example, an estimation network trained to estimate the voice emphasis feature E from the integrated feature Z AV and an estimation network trained to estimate the voice emphasis signal Y SE from the voice emphasis feature E SE and the voice presence / absence feature E VAD are used. The third learned model and the fifth learned model are generated by a learning device described later SE .
[0055] The fourth trained model and the sixth trained model are neural networks trained to reduce a third loss function regarding the difference between the correct label regarding the voice section and the voice presence score Y VAD and. As the neural network, for example, the integrated feature Z AV to estimate the voice presence / absence feature E VAD and an estimation network trained to estimate the voice enhancement feature E SE and the voice presence / absence feature E VAD and an estimation network trained to estimate the voice presence score Y are used. The fourth trained model and the sixth trained model are generated by a learning device described later.
[0056] When step SA8 is performed, the voice section detection unit 119 detects a voice section based on a comparison with a threshold η of the voice presence score Y calculated in step SA8 (step SA9). VAD
[0057] For each frame time, the value of the voice presence score Y VAD is compared with the threshold η. The threshold η is set at the boundary between a value corresponding to speech and a value not corresponding to speech. For example, when the voice presence score Y VAD takes a value from "0" to "1", the threshold η is set to "0.5". When the value of the voice presence score Y VAD is greater than the threshold η, the frame time is determined to be a voice section, and when the value of the voice presence score Y VAD is less than the threshold η, the frame time is determined to be a non-voice section. By performing the determination process for each frame time, the voice section and the non-voice section in the time interval corresponding to the input signal are detected. A label of a voice section or a non-voice section is assigned to each frame time of the input signal.
[0058] When step SA9 is performed, the output control unit 120 outputs the voice section and / or non-voice section detected in step SA9 (step SA10). Various forms are possible as the output mode. As an example, in step SA10, the output control unit 120 displays the voice section and / or non-voice section on the display device 15. At this time, the output control unit 120 may display the voice section and / or non-voice section in visual association with the acoustic signal and / or image signal. Also, the output control unit 120 may output the voice enhancement signal calculated in step SA7 via the acoustic device 16.
[0059] Thus, the voice section detection process by the processing circuit 11 ends. The input signal after the voice section is detected is used for processes such as voice recognition and data compression, for example.
[0060] Note that the voice section detection processes shown in FIGS. 2 and 3 are examples, and the voice section detection process according to the present embodiment is not limited to the procedures shown in FIGS. 2 and 3. As described above, the voice enhancement signal calculation unit 117 may not be provided, and in this case, step SA7 may not be performed. Also, the integrated feature calculation unit 114 may not be provided, and in this case, step SA4 may not be performed.
[0061] As described above, the voice section detection device 100 according to the present embodiment includes an acoustic feature calculation unit 112, a non-acoustic feature calculation unit 113, a voice emphasis feature calculation unit 115, a voice presence / absence feature calculation unit 116, a voice presence score calculation unit 118, and a voice section detection unit 119. The acoustic feature calculation unit 112 calculates acoustic features based on an acoustic signal. The acoustic features have values correlated with pronunciation. The non-acoustic feature calculation unit 113 calculates non-acoustic features based on a non-acoustic signal. The non-acoustic features have values correlated with pronunciation. The voice emphasis feature calculation unit 115 calculates voice emphasis features from the acoustic features and the non-acoustic features. The voice presence / absence feature calculation unit 116 calculates voice presence / absence features from the acoustic features and the non-acoustic features. The voice presence score calculation unit 118 calculates a voice presence score from the voice emphasis features and the voice presence / absence features. The voice section detection unit 119 detects a voice section, which is a time interval during which voice is uttered, and / or a non-voice section, which is a time interval during which no voice is uttered, based on a comparison with a threshold value of the voice presence score.
[0062] According to the present embodiment, voice emphasis features and voice presence / absence features are calculated from acoustic features based on an acoustic signal and non-acoustic features based on a non-acoustic signal, and a voice presence score is calculated from the voice presence / absence features taking into account the voice emphasis features. Therefore, it is possible to detect a voice section with high accuracy even in a wide variety of noise environments. Further, preferably, the voice emphasis features are calculated from integrated features using a third pre-trained model, and the third pre-trained model is generated by learning using a loss function defined to reduce a second loss related to the difference between a correct voice signal and a voice emphasis signal. By obtaining voice emphasis features using the third pre-trained model obtained by such learning and calculating a voice presence score from the voice presence / absence features taking into account the voice emphasis features, it is possible to further improve the detection accuracy of a voice section and / or a non-voice section even when there are a wide variety of noises.
[0063] (Learning device) FIG. 5 is a diagram showing a configuration example of the learning device 200. The learning device 200 is a computer that generates a first pre-trained model used for calculating acoustic features, a second pre-trained model used for calculating image features, a third pre-trained model used for calculating voice emphasis features, a fourth pre-trained model used for calculating voice presence / absence features, a fifth pre-trained model used for calculating voice emphasis signals, and a sixth pre-trained model used for detecting voice sections and / or non-voice sections. As shown in FIG. 9, the learning device 200 includes a processing circuit 21, a storage device 22, an input device 23, a communication device 24, a display device 25, and an acoustic device 26.
[0064] The processing circuit 21 includes a processor such as a CPU and a memory such as a RAM. By executing the learning program stored in the storage device 22, the processing circuit 21 executes a learning process for generating the first pre-trained model, the second pre-trained model, the third pre-trained model, the fourth pre-trained model, the fifth pre-trained model, and the sixth pre-trained model. The learning program is recorded on a non-transitory computer-readable recording medium. The processing circuit 21 reads and executes the learning program from the recording medium to implement an input signal acquisition unit 211, an acoustic feature calculation unit 212, a non-acoustic feature calculation unit 213, an integrated feature calculation unit 214, a voice emphasis feature calculation unit 215, a voice presence / absence feature calculation unit 216, a voice emphasis signal calculation unit 217, a voice presence score calculation unit 218, an update unit 219, an update end determination unit 220, and an output control unit 221. Note that the learning program may include a plurality of modules that separately implement the functions of the units 211 to 221.
[0065] The hardware implementation of the processing circuit 21 is not limited to only the above-described manner. For example, it may be configured by a circuit such as an ASIC that realizes the input signal acquisition unit 211, the acoustic feature calculation unit 212, the non-acoustic feature calculation unit 213, the integrated feature calculation unit 214, the voice emphasis feature calculation unit 215, the voice presence / absence feature calculation unit 216, the voice emphasis signal calculation unit 217, the voice presence score calculation unit 218, the update unit 219, the update end determination unit 220, and the output control unit 221. The input signal acquisition unit 211, the acoustic feature calculation unit 212, the non-acoustic feature calculation unit 213, the integrated feature calculation unit 214, the voice emphasis feature calculation unit 215, the voice presence / absence feature calculation unit 216, the voice emphasis signal calculation unit 217, the voice presence score calculation unit 218, the update unit 219, the update end determination unit 220, and / or the output control unit 221 may be implemented in a single integrated circuit or may be individually implemented in a plurality of integrated circuits.
[0066] The input signal acquisition unit 211 acquires training data having a plurality of training samples. The training sample is an input signal including a set of an acoustic signal and a non-acoustic signal. The input signal is a time-series signal and includes a time-series acoustic signal and a time-series non-acoustic signal. As described above, the non-acoustic signal is an image signal related to the speaker who is speaking, a sensor signal related to the physiological reaction of the speaker's lips or facial muscles due to speech, or the like.
[0067] The acoustic feature calculation unit 212 calculates acoustic features from the acoustic signal using a first neural network. The acoustic features calculated by the acoustic feature calculation unit 212 are the same as the acoustic features calculated by the acoustic feature calculation unit 112. A first learned model is generated by training the first neural network.
[0068] The non-acoustic feature calculation unit 213 calculates non-acoustic features from the non-acoustic signal using a second neural network. The non-acoustic features calculated by the non-acoustic feature calculation unit 213 are the same as the non-acoustic features calculated by the non-acoustic feature calculation unit 113. A second learned model is generated by training the second neural network.
[0069] The first neural network and the second neural network are trained to reduce a first loss regarding the difference between acoustic features and non-acoustic features related to the same voice source.
[0070] The integrated feature calculation unit 214 calculates an integrated feature based on the acoustic feature calculated by the acoustic feature calculation unit 212 and the non-acoustic feature calculated by the non-acoustic feature calculation unit 213. The integrated feature calculated by the integrated feature calculation unit 214 is the same as the integrated feature calculated by the integrated feature calculation unit 114.
[0071] The voice emphasis feature calculation unit 215 calculates a voice emphasis feature from the integrated feature using a third neural network. The acoustic feature calculated by the voice emphasis feature calculation unit 215 is the same as the voice emphasis feature calculated by the voice emphasis feature calculation unit 115. By training the third neural network, a third trained model is generated.
[0072] The voice presence / absence feature calculation unit 216 calculates a voice presence / absence feature from the integrated feature using a fourth neural network. The voice presence / absence feature calculated by the voice presence / absence feature calculation unit 216 is the same as the voice presence / absence feature calculated by the voice presence / absence feature calculation unit 116. By training the fourth neural network, a fourth trained model is generated.
[0073] The voice emphasis signal calculation unit 217 calculates a voice emphasis signal from the voice emphasis feature and the voice presence / absence feature using a fifth neural network. The voice emphasis signal calculated by the voice emphasis signal calculation unit 217 is the same as the voice emphasis signal calculated by the voice emphasis signal calculation unit 117. By training the fifth neural network, a fifth trained model is generated.
[0074] The voice presence score calculation unit 218 calculates a voice presence score from the voice enhancement feature and the voice presence / absence feature using a sixth neural network. The voice presence score calculated by the voice presence score calculation unit 218 is the same as the voice presence score calculated by the voice presence score calculation unit 118. By training the sixth neural network, a sixth trained model is generated.
[0075] The third neural network and the fifth neural network are trained to reduce a second loss related to the difference between the correct voice signal and the voice enhancement signal regarding the voice signal by the speaker included in the acoustic signal. The correct voice signal is an acoustic signal in which the voice signal derived from the utterance of the same voice source (speaker) as the acoustic signal acquired by the input signal acquisition unit 211 is enhanced. As an example, the correct voice signal is an acoustic signal obtained by collecting the utterance of the voice source in an environment without noise. In this case, the acoustic signal acquired by the input signal acquisition unit 211 is generated by adding a noise signal to the correct voice signal. Of course, separately from the correct voice signal, it may be obtained by collecting the utterance of the voice source in an environment with noise. As another example, the correct voice signal may be an acoustic signal obtained by removing the noise signal from the acoustic signal obtained by collecting the utterance of the voice source in an environment with noise. In this case, the acoustic signal acquired by the input signal acquisition unit 211 may be the acoustic signal before noise removal. Of course, separately from the correct voice signal, it may be obtained by collecting the utterance of the voice source in an environment with noise.
[0076] The fourth neural network and the sixth neural network are trained to reduce a third loss related to the difference between the correct label regarding the voice section and the non-voice section and the voice presence score. The correct label is one to which a label indicating that it is a voice section or a label indicating that it is a non-voice section is assigned for each frame of the acoustic signal. The correct label may be created manually based on the acoustic signal, may be created automatically using voice recognition technology or the like, or may be created by combining these manual and automatic methods.
[0077] The update unit 219 updates the first neural network, the second neural network, the third neural network, the fourth neural network, the fifth neural network, and the sixth neural network by using a fourth loss function including a first loss function related to the difference between the acoustic features and non-acoustic features of the voice generation source, a second loss function related to the difference between the correct voice signal and the voice enhancement signal, and a third loss function related to the difference between the correct label regarding the voice section and non-voice section and the voice presence score.
[0078] The update end determination unit 220 determines whether or not the stop condition of the learning process is satisfied. When it is determined that the stop condition is not satisfied, the calculation of the acoustic features by the acoustic feature calculation unit 212, the calculation of the non-acoustic features by the non-acoustic feature calculation unit 213, the calculation of the integrated features by the integrated feature calculation unit 214, the calculation of the voice enhancement features by the voice enhancement feature calculation unit 215, the calculation of the voice presence / absence features by the voice presence / absence feature calculation unit 216, the calculation of the voice enhancement signal by the voice enhancement signal calculation unit 217, the calculation of the voice presence score by the voice presence score calculation unit 218, and the update of the first neural network, the second neural network, the third neural network, the fourth neural network, the fifth neural network, and the sixth neural network by the update unit 219 are repeated. When it is determined that the stop condition is satisfied, at this point, the first neural network is output as the first learned model, the second neural network is output as the second learned model, the third neural network is output as the third learned model, the fourth neural network is output as the fourth learned model, the fifth neural network is output as the fifth learned model, and the sixth neural network is output as the sixth learned model.
[0079] The output control unit 221 displays various information via the display device 25 and the audio device 26. For example, the output control unit 221 displays an image signal on the display device 25 or outputs an audio signal via the audio device 26.
[0080] The storage device 22 is composed of a ROM, HDD, SSD, integrated circuit memory device, etc. The storage device 22 stores various calculation results by the processing circuit 21 and learning programs executed by the processing circuit 21. The storage device 22 is an example of a computer-readable recording medium.
[0081] The input device 23 inputs various commands from the user. As the input device 23, a keyboard, mouse, various switches, touch pad, touch panel display, etc. can be used. The output signal from the input device 23 is supplied to the processing circuit 21. Note that the input device 23 may be a computer connected to the processing circuit 21 via wire or wireless.
[0082] The communication device 24 is an interface for performing information communication between the learning device 200 and an external device connected via a network.
[0083] The display device 25 displays various information. As the display device 25, a CRT display, liquid crystal display, organic EL display, LED display, plasma display, or any other display known in the art can be appropriately used. The display device 25 may be a projector.
[0084] The audio device 26 converts an electrical signal into sound and emits it. As the audio device 26, a magnetic speaker, dynamic speaker, condenser speaker, or any other speaker known in the art can be appropriately used.
[0085] Next, an example of the learning process by the processing circuit 21 of the learning device 200 will be described. For the following description to be specific, it is assumed that the non-audio signal is an image signal.
[0086] FIG. 6 is a diagram showing an example of the flow of learning processing by the processing circuit 21. The learning processing is executed by the processing circuit 21 operating according to a learning program stored in the storage device 22 or the like. In the learning processing, the processing circuit 21 trains the first neural network NN1, the second neural network NN2, the third neural network NN3, the fourth neural network NN4, the fifth neural network NN5, and the sixth neural network NN6 in parallel. More specifically, the processing circuit 21 performs supervised learning on an integrated neural network including the first neural network NN1, the second neural network NN2, the third neural network NN3, the fourth neural network NN4, the fifth neural network NN5, and the sixth neural network NN6 based on training data.
[0087] FIG. 7 is a diagram showing a configuration example of the integrated neural network. As shown in FIG. 7, the integrated neural network includes the first neural network NN1, the second neural network NN2, an integrated feature calculation module NM1, the third neural network NN3, the fourth neural network NN4, the fifth neural network NN5, and the sixth neural network NN6.
[0088] The first neural network NN1 inputs an acoustic signal A and outputs an acoustic feature E A The second neural network NN2 inputs an image signal V and outputs an image feature E V The integrated feature calculation module NM1 inputs the acoustic feature E A output from the first neural network NN1 and the image feature E V output from the second neural network NN2 and outputs an integrated feature Z AV The integrated feature calculation module NM1 is a module corresponding to the function of the integrated feature calculation unit 214. The third neural network NN3 inputs the integrated feature Z AV output from the integrated feature calculation module NM1 and outputs a voice enhancement feature E SEOutputs it. The fourth neural network NN4 takes the integrated feature Z output from the integrated feature calculation module NM1 AV as input and outputs the voice presence / absence feature E VAD The fifth neural network NN5 takes the voice enhancement feature E output from the third neural network NN3 SE and the voice presence / absence feature E output from the fourth neural network NN4 VAD as input and outputs the voice enhancement signal Ŷ SE The sixth neural network NN6 takes the voice enhancement feature E output from the third neural network NN3 SE and the voice presence / absence feature E output from the fourth neural network NN4 VAD as input and outputs the voice presence score Ŷ VAD It is assumed that initial values of learning parameters and the like are assigned to the first neural network NN1, the second neural network NN2, the third neural network NN3, the fourth neural network NN4, the fifth neural network NN5, and the sixth neural network NN6. The learning parameters are weights, biases, etc. Note that the learning parameters may include any hyperparameters.
[0089] The first neural network NN1 has an architecture of an encoder network capable of calculating the acoustic feature E from the acoustic signal A
[0090] The first neural network NN1 has an architecture of an encoder network capable of calculating the acoustic feature E from the acoustic signal A A For example, the first neural network NN1 includes three 1D convolutional layers and one L2 normalization layer. The second neural network NN2 has an architecture of an encoder network capable of calculating the image feature E from the image signal V V For example, the second neural network NN2 includes three 3D convolutional layers and one L2 normalization layer. At least one of the three 3D convolutional layers may be connected to a max pooling layer and / or a global average pooling layer.
[0091] The third neural network NN3 has an architecture of a detection network capable of calculating the integrated feature Z AV to the voice emphasis feature E SE . As an example, the third neural network NN3 includes one Dense layer. The fourth neural network NN4 has an architecture of a detection network capable of calculating the voice presence / absence feature E AV from the integrated feature Z VAD . As an example, the fourth neural network NN4 includes one Dense layer.
[0092] The fifth neural network NN5 has an architecture of a detection network capable of calculating the voice emphasis signal Ŷ SE from the voice emphasis feature E VAD and the voice presence / absence feature E SE . As an example, the fifth neural network NN5 includes an Attention layer and a Dense layer. The mathematical expression of the Attention layer is shown in the following formula (3). As shown in formula (3), the Attention layer takes the voice emphasis feature E SE as the main input and uses the three memory elements W key and W qry and W val to calculate the inner product with the voice presence / absence feature E VAD , a scale factor d key , and a neural network layer defined by a softmax function. More specifically, the output value CTA SE of the Attention layer is obtained by the inner product of the softmax operation value and E SE and W val . The softmax operation value is output by applying the softmax operation to the inner product of the transpose of the inner product of E SE and W SE , the reciprocal of the square root of d key , E key , and the inner product of E VAD and W qry . The Dense layer is a network layer that calculates the voice emphasis signal Ŷ SE from the output value CTA SE .
[0093] [Number]
[0094] The sixth neural network NN6 has an architecture of a detection network capable of calculating an audio presence score Ŷ from an audio emphasis feature E SE and an audio presence / absence feature E VAD . As an example, the sixth neural network NN6 includes an attention layer and a dense layer. The mathematical expression of the attention layer is shown in the following equation (4). As shown in equation (4), the attention layer takes the audio presence / absence feature E VAD as the main input and uses three memory elements W VAD , W key , and W qry to calculate the inner product with the audio emphasis feature E val , a scale factor d SE , and a neural network layer defined by a softmax function. More specifically, the output value CTA key of the attention layer is obtained by the inner product of the softmax operation value and E VAD and W VAD . The softmax operation value is output by applying the softmax operation to the inner product of the transpose of the inner product of E val and W VAD , the reciprocal of the square root of d key , E key , and W SE . The dense layer is a network layer that calculates the audio presence score Ŷ from the output value CTA qry . VAD VAD
[0095] [Number]
[0096] For all convolutional and dense layers of the first neural network NN1, the second neural network NN2, the third neural network NN3, the fourth neural network NN4, the fifth neural network NN5, and the sixth neural network NN6, a normalization linear function unit follows, and a normalization linear function is applied to the output of the layer. Note that for the final layers of the fifth neural network NN5 and the sixth neural network NN6, a sigmoid activation function unit follows instead of a normalization linear function unit, and a sigmoid function is applied to the output of the final layer.
[0097] The integrated neural network takes an acoustic signal A and an image signal V as inputs and outputs an audio enhancement signal Y^ SE and an audio presence score Y^ VAD The learning parameters of the first neural network NN1, the second neural network NN2, the third neural network NN3, the fourth neural network NN4, the fifth neural network NN5, and the sixth neural network NN6 are trained so as to minimize a fourth loss function including a first loss function regarding the difference between the acoustic feature E A and the non-acoustic feature E V a second loss function regarding the difference between the audio enhancement signal Y^ SE and the correct audio signal (ground truth) Y SE and a third loss function regarding the difference between the audio presence score Y^ VAD and the correct label (ground truth) Y VAD Hereinafter, the learning of the first neural network NN1, the second neural network NN2, the third neural network NN3, the fourth neural network NN4, the fifth neural network NN5, and the sixth neural network NN6 will be described with reference to FIGS. 6 and 7.
[0098] As shown in FIGS. 6 and 7, the input signal acquisition unit 211 acquires an input signal including an acoustic signal A and an image signal V (step SB1). In step SB1, an input signal that is one learning sample is acquired. The frame length of the time interval of the input signal is not particularly limited, but for example, it is assumed to be about 10 frames. The acoustic signal A and the image signal V included in the input signal are temporally synchronized.
[0099] When step SB1 is performed, the acoustic feature calculation unit 212 uses the first neural network NN1 to calculate an acoustic feature E from the acoustic signal A acquired in step SB1 A (step SB2). It is assumed that the first neural network NN1 in step SB2 has not completed learning. Note that it is assumed that the acoustic signal A input to the first neural network NN1 has been converted from the time domain to the frequency domain.
[0100] When step SB2 is performed, the non-acoustic feature calculation unit 213 uses the second neural network NN2 to calculate an image feature Ev from the image signal V acquired in step SB1 (step SB3). It is assumed that the second neural network NN2 in step SB3 has not completed learning.
[0101] The order of step SB2 and step SB3 is not particularly limited, and step SB2 may be executed after step SB3, or step SB2 and step SB3 may be executed in parallel.
[0102] When step SB3 is performed, the integrated feature calculation unit 214 uses the integrated feature calculation module NM1 to calculate an integrated feature Z A between the acoustic feature E V and the image feature E AV (step SB4). In step S4, the integrated feature calculation unit 114 calculates the integrated feature between the acoustic feature E A and the image feature E V . The integrated feature Z AV is calculated for each frame time t.
[0103] When step SB4 is performed, the voice emphasis feature calculation unit 215 uses the third neural network NN3 to calculate the voice emphasis feature E from the integrated feature Z obtained in step SB4. AV from the voice emphasis feature E SE is calculated (step SB5). The third neural network NN3 in step SB5 is assumed to have not completed learning.
[0104] When step SB5 is performed, the voice presence / absence feature calculation unit 216 uses the fourth neural network NN4 to calculate the voice presence / absence feature E from the integrated feature Z obtained in step SB4. AV from the voice presence / absence feature E VAD is calculated (step SB6). The fourth neural network NN4 in step SB6 is assumed to have not completed learning.
[0105] The order of step SB5 and step SB6 is not particularly limited, and step SB5 may be executed after step SB6, or step SB5 and step SB6 may be executed in parallel.
[0106] When step SB6 is performed, the voice emphasis signal calculation unit 217 uses the fifth neural network NN5 to calculate the voice emphasis signal Ŷ from the voice emphasis feature E obtained in step SB5 and the voice presence / absence feature E obtained in step SB6. SE and the voice presence / absence feature E obtained in step SB6 VAD from the voice emphasis signal Ŷ SE is calculated (step SB7).
[0107] When step SB7 is performed, the voice presence score calculation unit 218 uses the 6 th neural network NN 6 to calculate the voice presence score Ŷ from the voice emphasis feature E obtained in step SB5 and the voice presence / absence feature E obtained in step SB6. SE and the voice presence / absence feature E obtained in step SB6 VAD from the voice presence score Ŷ VAD is calculated (step SB8).
[0108] When step SB7 is performed, the update unit 219 updates the first neural network NN1, the second neural network NN2, the third neural network NN3, the fourth neural network NN4, the fifth neural network NN5, and the sixth neural network NN6 using a fourth loss function that includes a first loss function, a second loss function, and a third loss function (step SB8). As shown in the following equation (5), the fourth loss function L TOTAL at frame time t and coordinate d is defined by the sum of the first loss function L KLD , the second loss function L SDR , and the third loss function L BCE .
[0109]
Equation
[0110] The first loss function L KLD is a loss function for imposing a penalty on the difference between the acoustic features and non-acoustic features regarding the same sound source. Specifically, as shown in the following equation (6), the first loss function L KLD is defined by the Kullback-Leibler information amount based on the acoustic feature E A and the image feature E V . The Kullback-Leibler information amount is used as a measure for evaluating the difference between the acoustic feature E A and the image feature E V . The second loss function L SDR is a loss function for imposing a penalty on the difference between the correct speech signal and the voice enhancement signal. As shown in the following equations (7) and (8), the second loss function L SDR is defined by the signal-to-distortion ratio (SDR). The signal-to-distortion ratio is used as a measure for evaluating the difference between the correct speech signal Y SE and the voice enhancement signal Ŷ SE . The third loss function L BCEis a loss function for giving a penalty for the difference between the correct label regarding the voice section and the non-voice section and the voice presence score. As shown in the following formula (9), the third loss function L BCE is defined as binary cross entropy (BCE). Binary cross entropy is used as a measure for evaluating the difference between the correct label Y VAD and the voice presence score Y^ VAD .
[0111]
Equation
[0112] The update unit 219 updates the learning parameters of the first neural network NN1, the second neural network NN2, the third neural network NN3, the fourth neural network NN4, the fifth neural network NN5, and the sixth neural network NN6 according to an arbitrary optimization method so that the fourth loss function L TOTAL is minimized. Thereby, the difference between the acoustic feature E A and the image feature E V , the difference between the correct voice signal Y SE and the voice enhancement signal Y^ SE , and the difference between the correct label Y VAD and the voice presence score Y^ VAD are comprehensively minimized, and the learning parameters of the first neural network NN1, the second neural network NN2, the third neural network NN3, the fourth neural network NN4, the fifth neural network NN5, and the sixth neural network NN6 are updated. As the optimization method, any method such as stochastic gradient descent method or Adam (adaptive moment estimation) may be used.
[0113] When step SB8 is performed, the update end determination unit 220 determines whether or not the stop condition is satisfied (step SB9). The stop condition may be set, for example, that the number of updates of the learning parameter has reached a predetermined number or that the update amount of the learning parameter is less than a threshold value. When it is determined that the stop condition is not satisfied (step SB9: NO), the input signal acquisition unit 211 acquires other acoustic signals and image signals (step SB1). Then, for the other acoustic signals and image signals, calculation of acoustic features by the acoustic feature calculation unit 212 (step SB2), calculation of non-acoustic features by the non-acoustic feature calculation unit 213 (step SB3), calculation of integrated features by the integrated feature calculation unit 214 (step SB4), calculation of voice emphasis features by the voice emphasis feature calculation unit 215 (step SB5), calculation of voice presence / absence features by the voice presence / absence feature calculation unit 216 (step SB6), calculation of voice emphasis signals by the voice emphasis signal calculation unit 217 (step SB7), calculation of voice presence scores by the voice presence score calculation unit 218 (step SB8), update of the first neural network, the second neural network, the third neural network, the fourth neural network, the fifth neural network, and the sixth neural network by the update unit 219 (step SB9), and determination of satisfaction of the stop condition by the update end determination unit 220 (step SB10) are executed in order.
[0114] Note that steps SB2 to SB9 may be repeated for one learning sample (batch learning), or steps SB2 to SB9 may be repeated for a plurality of learning samples (mini-batch learning).
[0115] When it is determined in step SB10 that the stop condition is satisfied (step SB10: YES), the update end determination unit 220 outputs the first neural network NN1 at the time when the stop condition is satisfied as the first learned model, the second neural network NN2 as the second learned model, the third neural network NN3 as the third learned model, the fourth neural network NN4 as the fourth learned model, the fifth neural network NN5 as the fifth learned model, and the sixth neural network NN6 as the sixth learned model (step SB11). The first learned model, the second learned model, the third learned model, the fourth learned model, the fifth learned model, and the sixth learned model are transmitted to the voice section detection device 100 via the communication device 24 or the like and stored in the storage device 12. Note that the update end determination unit 220 may output an integrated neural network including the first learned model, the second learned model, the third learned model, the fourth learned model, the fifth learned model, the sixth learned model, and the integrated feature calculation module NM1.
[0116] When step SB10 is performed, the learning process by the processing circuit 21 ends. As described above, the voice section detection device 100 calculates a voice presence score from the acoustic signal and the image signal using the first learned model, the second learned model, the third learned model, the fourth learned model, and the sixth learned model generated by the learning device 200. The voice section detection device 100 may calculate a voice presence score by inputting the acoustic signal and the image signal to the integrated neural network.
[0117] As described above, the learning device 200 according to the present embodiment includes an input signal acquisition unit 211, an acoustic feature calculation unit 212, a non-acoustic feature calculation unit 213, an integrated feature calculation unit 214, a voice emphasis feature calculation unit 215, a voice presence / absence feature calculation unit 216, a voice emphasis signal calculation unit 217, a voice presence score calculation unit 218, and an update unit 219. The input signal acquisition unit 211 acquires an acoustic signal and a non-acoustic signal related to the same voice source. The acoustic feature calculation unit 212 calculates acoustic features from the acoustic signal using a first neural network. The non-acoustic feature calculation unit 213 calculates non-acoustic features from the non-acoustic signal using a second neural network. The integrated feature calculation unit 214 calculates an integrated feature of the acoustic feature and the non-acoustic feature. The voice emphasis feature calculation unit 215 calculates voice emphasis features from the integrated feature using a third neural network. The voice presence / absence feature calculation unit 216 calculates voice presence / absence features from the integrated feature using a fourth neural network. The voice emphasis signal calculation unit 217 calculates a voice emphasis signal from the voice emphasis features and the voice presence / absence features using a fifth neural network. The voice presence score calculation unit 218 calculates a voice presence score from the voice emphasis features and the voice presence / absence features using a sixth neural network. The update unit 219 updates the first neural network, the second neural network, the third neural network, the fourth neural network, the fifth neural network, and the sixth neural network using a fourth loss function including a first loss function related to the difference between the acoustic feature and the non-acoustic feature, a second loss function related to the difference between the correct voice signal and the voice emphasis signal, and a third loss function related to the difference between the correct label and the voice presence score.
[0118] With the above configuration, it becomes possible to generate an integrated neural network of the first learned model, the second learned model, the third learned model, the fourth learned model, the fifth learned model, and the sixth learned model, which are high-precision voice segment detection models, through multi-task learning of the voice segment detection task and the voice emphasis task. From the voice presence feature obtained using the third learned model and the voice emphasis feature obtained using the fourth learned model, a voice presence score robust to noise is calculated. As a result, it becomes possible to detect high-precision voice segments even in a variety of noise environments. Although the voice segment detection task and the voice emphasis task are independent processing tasks with different outputs, in the common weight update of the network that increases the importance of voice feature amounts in a noise environment, the optimization problem by the attention mechanism that promotes joint learning with the voice emphasis process and its effect is restricted as an inductive bias to the voice segment detection process. As a result, the multi-task learning of the voice segment detection task and the voice emphasis task contributes to the improvement of the detection accuracy of voice segments.
[0119] (Example) Verification was conducted on the performance of the integrated neural network (hereinafter referred to as the proposed model) according to this embodiment, which includes the first neural network, the second neural network, the third neural network, the fourth neural network, the fifth neural network, and the sixth neural network generated according to the above embodiment. The Audio-only model (A-only), the baseline model (AV-baseline), and the proposed model were learned and evaluated using the GRID-AV sentence corpus (hereinafter simply referred to as the GRID corpus), which consists of AV recordings in which 34 speakers (18 males and 16 females) each spoke 1,000 sentences. The Audio-only model is a voice segment detection model that uses only acoustic signals without using non-acoustic signals. The baseline model is a voice segment detection model that does not utilize voice emphasis features and voice presence features.
[0120] For the training of the model considering noise, the background noise provided in the 4th CHiME Challenge was randomly selected and added to all the acoustic training data of the GRID corpus with SNR in the range of -5 to +20 dB. The noise of CHiME4 was recorded at four locations: bus, cafeteria, pedestrian area, and road intersection. To evaluate the performance under noise conditions, the noise of CHiME4 was added to all the test acoustic recordings of the GRID corpus at 5 dB intervals in the range of SNR from -10 to +20 dB.
[0121] The recordings of each speaker in the GRID corpus were split into a training dataset, a validation dataset, and a test dataset at a ratio of 6:2:2. The validation dataset was used to identify hyperparameters suitable for each model and each experimental condition. For the input audio, all audio recordings were resampled at a sampling rate of 16 kHz, and a spectrogram of 641 dimensions was calculated using the short-time Fourier transform with a window size of 1,280 samples (80 milliseconds) and a hop length of 640 samples (40 milliseconds). For the input images, all video recordings were cropped to the lip region using a 68-coordinate face landmark detector, converted to a resolution of H×W = 40×64 pixels at 25 frames per second (40 milliseconds), and the RGB channels were normalized between 0 and 1. All models were trained using Adam optimization with Gradient Clipping. The learning rate was initialized at 0.0001, and the batch size was set to 1.
[0122] Figure 8 is a table showing the results of validation using the CHiME4 dataset. Figure 9 uses the area under the ROC curve (AUROC (Area Under the Receiver Operating Characteristic)) as an indicator for the quantitative performance evaluation of AV-VAD. The validation results are shown for each model. The bold font indicates the best value.
[0123] The proposed model (AV-proposed) significantly outperforms the baseline model (AV-baseline) under all experimental conditions, presenting a relative improvement of up to 4.36% (SNR -5dB) and an average of 2.09% (Average). Furthermore, the proposed model presents a relative improvement of up to 16.41% (SNR -10dB) and an average of 4.62% (Average) compared to the Audio-only model (A-only).
[0124] The above embodiments are examples, and various modifications are possible. For example, although the acoustic signal is assumed to be a voice signal, it may also be a vocal cord signal or a vocal tract signal obtained by decomposing the voice signal. Also, although the acoustic signal is assumed to be waveform data of sound pressure values in the time domain or the frequency domain, it may also be data obtained by converting the waveform data into an arbitrary space.
[0125] Thus, according to some of the above embodiments, it becomes possible to further improve the detection accuracy of voice sections and / or non-voice sections even when various types of noise are present.
[0126] Although some embodiments of the present invention have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These novel embodiments can be implemented in various other forms, and various omissions, replacements, and changes can be made without departing from the gist of the invention. These embodiments and their modifications are included in the scope and gist of the invention, and are included in the invention described in the claims and its equivalent scope.
Description of Reference Numerals
[0127] 11... Processing circuit, 12... Storage device, 13... Input device, 14... Communication device, 15... Display device, 16... Acoustic device, 100... Voice section detection device, 111... Input signal acquisition unit, 112... Acoustic feature calculation unit, 113... Non-acoustic feature calculation unit, 114... Integrated feature calculation unit, 115... Voice enhancement feature calculation unit, 116... Voice presence / absence feature calculation unit, 117... Voice enhancement signal calculation unit, 118... Voice presence score calculation unit, 119... Voice section detection unit, 120... Output control unit.
Claims
1. An acquisition unit that acquires an acoustic signal and a non-acoustic signal related to a voice sound source; An acoustic feature calculation unit that calculates acoustic features based on the acoustic signal; A non-acoustic feature calculation unit that calculates non-acoustic features based on the non-acoustic signal; A voice emphasis feature calculation unit that calculates a voice emphasis feature based on the acoustic feature and the non-acoustic feature; A voice presence / absence feature calculation unit that calculates a voice presence / absence feature based on the acoustic feature and the non-acoustic feature; A voice presence score calculation unit that calculates a voice presence score based on the voice emphasis feature and the voice presence / absence feature; A detection unit that detects a voice section, which is a time section in which voice is being emitted, and / or a non-voice section, which is a time section in which voice is not being emitted, based on a comparison with a threshold value of the voice presence score; A voice section detection device comprising the above.
2. Further comprising an integrated feature calculation unit that calculates an integrated feature based on the acoustic feature and the non-acoustic feature, The voice emphasis feature calculation unit calculates the voice emphasis feature based on the integrated feature, The voice presence / absence feature calculation unit calculates the voice presence / absence feature based on the integrated feature, The voice section detection device according to Claim 1.
3. A voice emphasis signal calculation unit that calculates a voice emphasis signal representing a restored voice signal in which only the voice signal derived from the speech of the voice sound source is separated and emphasized from the acoustic signal based on the voice emphasis feature and the voice presence / absence feature; Further comprising an output control unit that outputs the voice emphasis signal via an acoustic device. The voice section detection device according to Claim 2.
4. The acoustic feature calculation unit calculates the acoustic feature from the acoustic signal using a first pre-trained model, The non-acoustic feature calculation unit calculates the non-acoustic feature from the non-acoustic signal using a second pre-trained model, The voice emphasis feature calculation unit calculates the voice emphasis feature from the integrated feature using a third pre-trained model, The voice presence / absence feature calculation unit calculates the voice presence / absence feature from the integrated feature using a fourth pre-trained model, The voice emphasis signal calculation unit calculates the voice emphasis signal from the voice emphasis feature and the voice presence / absence feature using a fifth pre-trained model, The voice presence score calculation unit calculates the voice presence score from the voice emphasis feature and the voice presence / absence feature using a sixth pre-trained model, The first pre-trained model, the second pre-trained model, the third pre-trained model, the fourth pre-trained model, the fifth pre-trained model, and the sixth pre-trained model are generated by learning to reduce an integrated loss including a first loss related to a difference between the acoustic feature and the non-acoustic feature regarding the voice generation source, a second loss related to a difference between the correct voice signal and the voice emphasis signal for the voice signal included in the acoustic signal, and a third loss related to a difference between the correct label regarding the voice section and the non-voice section and the voice presence score. The voice section detection device according to claim 3.
5. The voice section detection device according to claim 1, wherein the non-acoustic signal is an image signal temporally synchronized with the acoustic signal.
6. An acquisition unit that acquires an acoustic signal and a non-acoustic signal regarding a voice generation source, An acoustic feature calculation unit that calculates an acoustic feature from the acoustic signal using a first neural network, A non-acoustic feature calculation unit that calculates a non-acoustic feature from the non-acoustic signal using a second neural network, An integrated feature calculation unit that calculates an integrated feature based on the acoustic feature and the non-acoustic feature, A voice emphasis feature calculation unit that calculates a voice emphasis feature from the integrated feature using a third neural network, An audio presence feature calculation unit that calculates an audio presence feature from the integrated features using a fourth neural network; An audio emphasis signal calculation unit that calculates an audio emphasis signal from the audio emphasis feature and the audio presence feature using a fifth neural network; An audio presence score calculation unit that calculates an audio presence score from the audio emphasis feature and the audio presence feature using a sixth neural network; An update unit that updates the first neural network, the second neural network, the third neural network, the fourth neural network, the fifth neural network, and the sixth neural network using a fourth loss function including a first loss function regarding the difference between the acoustic feature and the non-acoustic feature, a second loss function regarding the difference between the correct audio signal regarding the audio signal included in the acoustic signal and the audio emphasis signal, and a third loss function regarding the difference between the correct label regarding the audio section and the non-audio section and the audio presence score; A learning device comprising:
7. A function of causing a processor to calculate an acoustic feature based on an acoustic signal; calculate a non-acoustic feature based on a non-acoustic signal obtained from the same sound source as the acoustic signal; calculate an audio emphasis feature based on the acoustic signal and the non-acoustic signal; calculate an audio presence feature based on the acoustic signal and the non-acoustic signal; calculate an audio presence score based on the audio emphasis feature and the audio presence feature; detect an audio section that is a time interval during which audio is emitted and / or a non-audio section that is a time interval during which no audio is emitted based on a comparison with a threshold value of the audio presence score; An audio section detection program for realizing the above.
Citation Information
Patent Citations
Voice activity detection method combined with voice enhancement
CN113113049A
Speech section detecting device and speech recognition device, program and recording medium
JP2011059186A
Utterance section detection device, voice recognition device, utterance section detection system, utterance section detection method, and utterance section detection program
JP2021162685A
Information processing device, program, and information processing method
WO2020144857A1