Voice analysis method and voice analysis device
The speech analysis method and device use the 'Dominant Spectrum Test' and 'Sequential Spectrum Test' to enhance the accuracy of fundamental frequency estimation by distinguishing between dominant frequencies and noise, addressing the limitations of conventional methods for hoarse voices.
Patent Information
- Application Number
- JP2024087318
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-29
- Publication Date
- 2025-12-11
AI Technical Summary
Conventional methods for estimating the fundamental frequency of speech have limitations in accuracy, particularly for hoarse voices, and are prone to errors due to noise components like subharmonics.
A speech analysis method and device that employs a 'Dominant Spectrum Test' and a 'Sequential Spectrum Test' to accurately estimate the fundamental frequency by distinguishing between dominant frequencies, subharmonics, and external noise, using frequency spectrum waveforms of frames to verify spectral intensity and frequency ranges.
Improves the accuracy of fundamental frequency estimation for both normal and hoarse speech by correcting for erroneous detections and noise interference, enhancing the precision of pitch detection algorithms.
Smart Images

Figure 2025180166000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a speech analysis method and a speech analysis device. [Background technology]
[0002] There are known methods for estimating the fundamental frequency of speech by analyzing the speech. Conventional methods for estimating the fundamental frequency can be broadly divided into three groups: the first uses characteristics in the time domain, the second uses characteristics in the frequency domain, and the third combines characteristics in both the time domain and the frequency domain.
[0003] Methods that utilize time-domain characteristics commonly use signal correlation, such as the autocorrelation method, cross-correlation method, and maximum detection method. Methods that utilize frequency-domain characteristics include the "SWIPE" and "SWIPE'" methods described in Non-Patent Document 1. Methods such as "SWIPE" and "SWIPE'" focus on the harmonic structure of the power spectrum and improve the accuracy of fundamental frequency estimation by reducing errors. Methods that combine both of the above characteristics include the "BaNa" method described in Non-Patent Document 2. The "BaNa" method estimates the fundamental frequency by combining a classical approach consisting of harmonic frequency ratio and cepstrum analysis. [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] A. Camacho and JG Harris, “A sawtooth waveform inspired pitch estimator for speech and music,” J. Acoust. Soc. 2008. [Non-patent document 2] (Ba H, Yang N. BaNa: a hybrid approach for noise resilient pitch detection. IEEE Statistical Signal Processing Workshop. 2012) Summary of the Invention [Problem to be solved by the invention]
[0005] Conventional methods have room for improvement in the accuracy of estimating the fundamental frequency of speech.
[0006] The present disclosure provides a voice analysis method and the like that can improve the accuracy of estimating the fundamental frequency of voice. [Means for solving the problem]
[0007] A speech analysis method according to one aspect of the present disclosure includes a waveform acquisition step of acquiring frequency spectrum waveforms of a plurality of frames, each frame representing a predetermined period, based on a signal including speech; a first derivation step of deriving a fundamental frequency in a first frame based on the frequency spectrum waveform of the first frame among the plurality of frames; and a step of estimating a fundamental frequency in a second frame based on the frequency spectrum waveform of a second frame among the plurality of frames, wherein the fundamental frequency estimation step determines whether a portion of the frequency spectrum waveform of the second frame is within a predetermined frequency range that includes the fundamental frequency of the first frame, and whether a verification point in the portion of the frequency spectrum waveform is within a predetermined spectral intensity range that includes the spectral intensity of the fundamental frequency, and estimates the fundamental frequency in the second frame based on the determination.
[0008] A speech analysis device according to one embodiment of the present disclosure includes a waveform acquisition unit that acquires frequency spectrum waveforms of a plurality of frames, each frame representing a predetermined period, based on a signal including speech; a derivation unit that derives a fundamental frequency of a first frame of the plurality of frames based on the frequency spectrum waveform of the first frame; and an estimation unit that estimates a fundamental frequency of a second frame of the plurality of frames based on the frequency spectrum waveform of the second frame, wherein the estimation unit determines whether a portion of the frequency spectrum waveform of the second frame is within a predetermined frequency range that includes the fundamental frequency of the first frame and whether a verification point in the portion of the frequency spectrum waveform is within a predetermined spectral intensity range that includes the spectral intensity of the fundamental frequency, and estimates the fundamental frequency of the second frame based on the determination. [Effects of the Invention]
[0009] According to the speech analysis method and the like of the present disclosure, it is possible to improve the accuracy of estimating the fundamental frequency of speech. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is a diagram showing a spectral waveform of a sound. [Figure 2] FIG. 10 is a diagram showing a spectral waveform in which noise components, mainly subharmonics, are added to audio. [Figure 3] 1 is a schematic diagram of a voice analysis system including a voice analysis device according to an embodiment. [Figure 4] FIG. 1 is a block diagram showing a configuration of a voice analysis device. [Figure 5] 1A and 1B are diagrams illustrating a voice waveform and a frequency spectrum waveform included in voice data. [Figure 6] FIG. 10 is a diagram showing an example of a frequency spectrum waveform of a first frame. [Figure 7] FIG. 10 is a diagram showing the maximum peak and frequency peaks in a frequency spectrum waveform. [Figure 8]FIG. 10 is a diagram illustrating an example of extracting a frequency peak with a large spectral intensity from among a plurality of frequency peaks. [Figure 9] FIG. 10 is a diagram showing an example of deriving a fundamental frequency in a first frame. [Figure 10] FIG. 10 is a diagram showing a predetermined frequency range including a fundamental frequency in a first frame, and a predetermined spectral intensity range including the spectral intensity of the fundamental frequency. [Figure 11] FIG. 10 is a diagram illustrating an example of estimating a fundamental frequency in a second frame. [Figure 12] FIG. 10 is a diagram illustrating another example of estimating the fundamental frequency in the second frame. [Figure 13] FIG. 10 is a diagram showing an example of displaying fundamental frequencies and spectral intensities. [Figure 14] 1 is a flowchart illustrating a speech analysis method according to an embodiment. [Figure 15a] FIG. 1 is a diagram showing a spectral waveform of a vocal cord original sound. [Figure 15b] FIG. 1 is a diagram showing the spectral waveform of speech filtered in the vocal tract and uttered from the mouth. [Figure 16] FIG. 1 is a diagram showing the spectral waveform of an audio signal including subharmonics. [Figure 17] This is a table summarizing the diagnoses of all 344 participants. [Figure 18] FIG. 10 is a diagram showing the distribution of voice quality evaluation levels in psychoacoustic evaluation. [Figure 19a] FIG. 1 is a diagram showing the process of searching for candidates for fo (fundamental frequency) in the dominant spectrum test. [Figure 19b] FIG. 10 is a diagram illustrating an example of dividing a low frequency region and extracting spectral peaks within each region. [Figure 19c] In this figure, the spectral peak of the lowest frequency having a spectral intensity equal to or greater than a certain level is compared with the spectral peaks found in FIG. 19a and is considered to be a candidate for fo (fundamental frequency). [Figure 20a]FIG. 10 is a diagram showing the process of selecting candidates for fo (fundamental frequency) in the sequential spectrum test. [Figure 20b] FIG. 10 is a diagram illustrating an example in which, when there is no frequency peak with a similar spectral intensity and frequency, a candidate fo (fundamental frequency) selected by a dominant spectrum test is used. [Figure 21] FIG. 10 is a diagram showing an example of a spectrogram of a concatenated sample of CS and SV. [Figure 22] FIG. 10 is a diagram illustrating an example of subharmonics errors. [Figure 23] FIG. 1 is a diagram showing a process of extracting ground truth of fo (fundamental frequency) on a spectrogram. [Figure 24] FIG. 10 is a diagram illustrating an example in which an estimated fo (fundamental frequency) and the ground truth of fo (fundamental frequency) are compared to extract a portion where they overlap. [Figure 25] FIG. 10 is a diagram showing an example of calculating the allocation of the fo (fundamental frequency) estimated match time to all samples, "% true-all," and the proportion of voiced sounds to all samples, "% voice-all." [Figure 26] FIG. 10 is a diagram showing a list of algorithms evaluated for comparison. [Figure 27] FIG. 10 is a diagram showing the results of plotting "% true" for each fo (fundamental frequency) estimation algorithm for all 454 samples in concatenated speech samples. [Figure 28] This figure compares the results of each fo (fundamental frequency) estimation algorithm for all 454 samples. [Figure 29] FIG. 10 shows results for speech samples rated as non-hoarse with Gtotal less than 0.5. [Figure 30]This figure shows the results for a voice sample that was evaluated as having hoarseness, with Gtotal, Rtotal, and Btotal each being 0.5 or more. [Figure 31] This figure shows the results of an investigation into whether the difference in "% true" was within 5% when comparing each type of group with hoarseness and that without hoarseness. DETAILED DESCRIPTION OF THE INVENTION
[0011] (Background to this disclosure) The fundamental frequency of a voice is closely related to pitch, a psychological quantity corresponding to the pitch of a sound, and is a physical quantity that roughly corresponds to the vibration frequency of the vocal cords. The fundamental frequency of a voice is used in technologies that quantify information about the vocal cords, which are the source of sound, and in technologies that require voice recognition.
[0012] Figure 1 shows the spectral waveform of speech, where (a) shows the spectral waveform of speech from the original vocal cords, and (b) shows the spectral waveform of speech actually emitted from the mouth.
[0013] The spectral waveform of the vocal cord original sound shown in Figure 1(a) has the strongest frequency characteristics corresponding to the fundamental frequency, and the harmonic structure attenuates as the frequency increases. This phenomenon is seen in the vocal cord original sound, which consists only of the vibration of the vocal cords. On the other hand, the actual sound emitted from the mouth is filtered by the influence of resonance in the vocal tract and anti-resonance in the nasal cavity, and forms formants with increased or decreased frequency components, as shown in Figure 1(b). If the fundamental frequency is simply estimated using the waveform in Figure 1(b), the lowest frequency of the harmonic complex sound is estimated to be the fundamental frequency.
[0014] However, in a real environment, noise from recording equipment and background noise in the audible range can significantly degrade the quality of the input audio signal.
[0015] FIG. 2 is a diagram showing a spectral waveform in which noise components, mainly subharmonics, are added to a voice.
[0016] In Figure 2, harmonic components are indicated by diagonal hatching, and noise components are indicated by hatched dots. When noise components, especially subharmonics, are added to audio as shown in Figure 2, the lowest frequency of the harmonic complex tone may be mistakenly estimated as the fundamental frequency.
[0017] Therefore, conventional techniques have been developed as pitch detection algorithms that are robust to noise, as shown in Non-Patent Documents 1 and 2. However, these conventional techniques have a problem in that they have low accuracy in estimating the fundamental frequency of hoarse voices, such as raspy and raspy voices.
[0018] In response to this problem, this disclosure proposes a new algorithm that includes two functions, the "Dominant Spectrum Test" and the "Sequential Spectrum Test." This new algorithm enables accurate estimation of the fundamental frequency not only for normal speech but also for hoarse speech.
[0019] Hereinafter, embodiments will be described with reference to the drawings. Note that the embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, component placement and connection configurations, steps, step order, etc. shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Furthermore, among the components in the following embodiments, components that are not described in an independent claim that represents a superordinate concept will be described as optional components.
[0020] It should be noted that the drawings are schematic diagrams and are not necessarily strict illustrations. In addition, in the drawings, substantially the same components are denoted by the same reference numerals, and duplicated explanations may be omitted or simplified.
[0021] (Embodiment) [Outline of the voice analysis device] The voice analysis device according to the embodiment will be described with reference to FIGS.
[0022] FIG. 3 is a schematic diagram of a voice analysis system 1 including a voice analysis device 10 according to an embodiment.
[0023] The voice analysis system 1 is a system that analyzes human voice. As shown in Fig. 3, the voice analysis system 1 includes a voice analysis device 10 and a sound collection device 90. The voice analysis device 10 and the sound collection device 90 are connected for communication via wired or wireless communication. The voice analysis system 1 may also include an input interface for inputting person identification information and the like, and a display device for outputting analysis results (not shown).
[0024] The sound collection device 90 is a microphone that collects sounds such as human speech. The sound collection device 90 collects, for example, human speech in a non-contact manner and outputs a signal including the collected sound to the voice analysis device 10.
[0025] The voice analysis device 10 is a device that estimates the fundamental frequency of a voice by analyzing a signal containing the voice output from a sound collection device 90. The voice analysis device 10 is configured by a computer having a CPU (Central Processing Unit) and a memory. The memory is, for example, a ROM (Read Only Memory), a RAM (Random Access Memory), a semiconductor memory, or an HDD (Hard Disk Drive). The memory stores programs that are executed to realize each function of the voice analysis device 10.
[0026] FIG. 4 is a block diagram showing the configuration of the voice analysis device 10.
[0027] The speech analysis device 10 includes a waveform acquisition unit 20, a derivation unit 30, an estimation unit 40, and a correction unit 50. In the speech analysis device 10 of this embodiment, the derivation unit 30 executes a "dominant spectrum test," and the estimation unit 40 executes a "sequential spectrum test." Each of the components included in the speech analysis device 10 will be described below.
[0028] First, a description will be given of the waveform acquisition unit 20 that acquires frequency spectrum waveforms. The waveform acquisition unit 20 acquires a plurality of frequency spectrum waveforms based on a voice waveform signal, which is a signal that includes voice.
[0029] FIG. 5 is a diagram showing a voice waveform and a frequency spectrum waveform included in the voice data.
[0030] Figure 5(a) shows an audio waveform with time on the horizontal axis and sound pressure on the vertical axis, and Figure 5(b) shows a frequency spectrum waveform with frequency on the horizontal axis and frequency spectrum intensity (unit: dB) on the vertical axis. Figure 5(b) also shows data for a certain period of time out of the entire audio waveform data represented by a plurality of frequency spectrum waveforms. The plurality of frequency spectrum waveforms are represented by frequency spectrum waveforms of a plurality of frames, with one predetermined period being one frame. In the example shown in Figure 5, the data for the certain period extracted from the audio waveform data is 0.1 seconds, and the plurality of frames in 0.1 seconds are made up of 30 frames, so the predetermined period making up one frame of data is 0.0033 seconds.
[0031] In this way, the waveform acquiring unit 20 acquires frequency spectrum waveforms of a plurality of frames, each frame being a predetermined period, based on a signal including audio. The waveform acquiring unit 20 outputs the acquired frequency spectrum waveforms of the plurality of frames to the derivation unit 30.
[0032] The waveform acquiring unit 20 may adjust the frequency spectrum waveform so that it does not depend on the recording level of the audio data, and output the adjusted frequency spectrum to the derivation unit 30. For example, when adjusting the frequency spectrum waveform, the waveform acquiring unit 20 may adjust and normalize the vertical level of the frequency spectrum waveform so that the average of multiple maximum values of spectral intensity included in the frequency spectrum waveform falls within a predetermined range (e.g., 70 dB±5 dB). Alternatively, the waveform acquiring unit 20 may adjust the frequency spectrum waveform so that the average spectral intensity of all audio data falls within a predetermined range.
[0033] Next, the derivation unit 30 that executes the “dominant spectrum test” will be described. The derivation unit 30 derives the fundamental frequency and the tentative fundamental frequency based on the frequency spectrum waveforms of multiple frames output from the waveform acquisition unit 20.
[0034] Below, we will explain how to derive the fundamental frequency in the first frame of multiple frames, how to derive the provisional fundamental frequency in the second frame following the first frame, and how to derive the provisional fundamental frequency in the (N+1)th frame (N is an integer greater than or equal to 2) after the second frame.
[0035] The derivation unit 30 derives the fundamental frequency in the first frame based on the frequency spectrum waveform of the first frame.
[0036] 6 is a diagram showing an example of a frequency spectrum waveform of the first frame. Note that the frequency spectrum waveform diagrams shown hereinafter are shown in a form that makes it easy to understand the data processing performed by the voice analysis device 10.
[0037] Figure 6 shows a frequency spectrum waveform with the horizontal axis representing frequencies from 0 Hz to 1 kHz and the vertical axis representing spectral intensity. Also in Figure 6, the frequency components of the fundamental frequency linked to vocal cord vibration are indicated by diagonal hatching. Also in Figure 6, subharmonics, which are modulation frequencies other than the fundamental frequency, and frequency components due to disturbance noise such as environmental noise and turbulence noise are indicated by hatching dots. The frequency components due to subharmonics and disturbance noise have sufficiently small spectral intensities compared to the frequency components of the fundamental frequency linked to vocal cord vibration.
[0038] The dominant spectrum test executed by the derivation unit 30 is an algorithm that uses differences in the spectral intensities of frequency components to distinguish between dominant fundamental frequencies, non-dominant subharmonics, and external noise. The derivation unit 30 derives the maximum value of the spectral intensity of the frequency spectrum waveform in the first frame.
[0039] As shown in Fig. 6, the derivation unit 30 detects the frequency of the maximum peak pm at which the spectral intensity is greatest within a frequency range of 50 Hz to 400 Hz, which can be the fundamental frequency of human speech. Note that when deriving the maximum peak pm, the derivation unit 30 may generate a speech spectrum waveform by removing data other than speech from the frequency spectrum waveform, and derive the maximum peak pm of the spectral intensity based on the speech spectrum waveform. Examples of data other than speech include noise mixed in from recording equipment and background noise in the audible range.
[0040] To derive the fundamental frequency, the derivation unit 30 extracts a plurality of frequency peaks from a frequency range located on the lower frequency side than the frequency of the maximum peak pm at which the spectrum intensity is greatest.
[0041] FIG. 7 is a diagram showing a frequency peak p in a frequency spectrum waveform.
[0042] For example, the derivation unit 30 divides the frequency region located on the lower frequency side than the frequency of the maximum peak pm into a plurality of regions, and derives the point in each region where the spectrum intensity has a maximum value as the frequency peak p.
[0043] Fig. 7 shows an example in which the frequency range from 0 Hz to the maximum peak pm is divided equally into eight regions. Fig. 7 also shows frequency peaks p located at points (maximum points) that indicate maximum values of spectral intensity in the eight divided regions. Of the divided regions, the first divided region, which has a lower limit of 0 Hz, is one in which the fundamental frequency of audio is almost absent, and the frequency peak of the eighth divided region coincides with pm. Therefore, the derivation unit 30 derives multiple frequency peaks p from the second to seventh divided regions, excluding the first and eighth divided regions.
[0044] The frequency range does not necessarily have to be divided equally into 8. For example, the derivation unit 30 may divide the frequency range into 3, and derive multiple frequency peaks p from a second divided region located between a first divided region with a width of 50 Hz and a lower limit of 0 Hz, and a third divided region with a width of 50 Hz and an upper limit of the maximum peak pm.
[0045] At this stage, the frequency peak p is derived in a state in which the subharmonics and disturbance noise described above are included. Therefore, in order to eliminate the influence of the subharmonics and disturbance noise, the derivation unit 30 extracts frequency peaks having sufficient spectral intensity relative to the spectral intensity of the maximum peak pm from among the multiple frequency peaks p corresponding to the multiple divided regions.
[0046] FIG. 8 is a diagram showing an example of extracting a frequency peak pk with a large spectral intensity from among a plurality of frequency peaks p.
[0047] As shown in Fig. 8, the derivation unit 30 extracts one or more frequency peaks pk having a maximum value greater than a predetermined spectral intensity from among the multiple frequency peaks p. The predetermined spectral intensity is, for example, a value that is 10 dB lower than the maximum peak pm. In this case, the derivation unit 30 uses the maximum peak pm as a reference and extracts one or more frequency peaks pk located in a range that is 10 dB lower than the maximum peak pm (the area indicated by hatched dots). Note that the predetermined spectral intensity relative to the maximum peak pm may be any value.
[0048] Next, the derivation unit 30 derives the fundamental frequency in the first frame based on the one or more frequency peaks pk.
[0049] FIG. 9 is a diagram showing an example of deriving the fundamental frequency in the first frame.
[0050] 9, the derivation unit 30 extracts the frequency peak pk0 located at the lowest frequency side from among one or more frequency peaks pk, and sets the frequency corresponding to the frequency peak pk0 as the fundamental frequency of the first frame. In this example, the frequency of the frequency peak pk0 located in the third divided region out of the eight divided regions is derived as the fundamental frequency of the first frame.
[0051] Similarly, the derivation unit 30 derives a tentative fundamental frequency for the second frame based on the frequency spectrum waveform of the second frame. The reason for using a tentative fundamental frequency is that the tentative fundamental frequency may not be adopted by the estimation unit 40, which will be described later. The method for deriving the tentative fundamental frequency for the second frame is the same as the method for deriving the fundamental frequency for the first frame (see FIGS. 6 to 9).
[0052] That is, the derivation unit 30 derives the maximum peak pm of the spectral intensity of the frequency spectrum waveform in the second frame. The derivation unit 30 divides a frequency region located on the lower frequency side than the frequency at which the spectral intensity reaches the maximum peak pm into multiple regions, and derives the point at which the spectral intensity reaches a maximum value in each region as a frequency peak p. The derivation unit 30 extracts one or more frequency peaks pk having a maximum value greater than a predetermined spectral intensity from the multiple frequency peaks p corresponding to the multiple regions. The derivation unit 30 then extracts a frequency peak pk0 located on the lowest frequency side from the one or more frequency peaks pk, and sets the frequency corresponding to the frequency peak pk0 as the tentative fundamental frequency in the second frame.
[0053] Similarly, the derivation unit 30 derives a tentative fundamental frequency for the (N+1)th frame (N is an integer equal to or greater than 2) based on the frequency spectrum waveform of the (N+1)th frame. The reason for using a tentative fundamental frequency is that the tentative fundamental frequency may not be adopted by the estimation unit 40, which will be described later. The method for deriving the tentative fundamental frequency for the (N+1)th frame is the same as the method for deriving the fundamental frequency for the first frame (see FIGS. 6 to 9).
[0054] That is, the derivation unit 30 derives the maximum peak pm of the spectral intensity of the frequency spectrum waveform in the (N+1)th frame. The derivation unit 30 divides a frequency region located on the lower frequency side than the frequency at which the spectral intensity reaches the maximum peak pm into multiple regions, and derives the point in each region where the spectral intensity reaches a maximum value as the frequency peak p. The derivation unit 30 extracts one or more frequency peaks pk having a maximum value greater than a predetermined spectral intensity from the multiple frequency peaks p corresponding to the multiple regions. The derivation unit 30 then extracts the frequency peak pk0 located on the lowest frequency side from the one or more frequency peaks pk, and sets the frequency corresponding to the frequency peak pk0 as the tentative fundamental frequency for the (N+1)th frame.
[0055] In this way, the derivation unit 30 derives the fundamental frequency in the first frame, the tentative fundamental frequency in the second frame, and the tentative fundamental frequency in the (N+1)th frame by the "dominant spectrum test."
[0056] Next, we will explain the estimation unit 40, which executes the "sequential spectrum test." The sequential spectrum test is an algorithm that focuses on changes in fundamental frequency over time in speech samples, and improves the accuracy of fundamental frequency estimation by correcting for erroneous fundamental frequency detections in the dominant spectrum test described above.
[0057] The sequential spectrum test is not performed in the first frame, but is performed in the second frame and subsequent frames. The following describes the estimation of the fundamental frequency in the second frame and the (N+1)th frame.
[0058] The estimation unit 40 estimates the fundamental frequency of the second frame based on the frequency spectrum waveform of the second frame. To estimate the fundamental frequency, the estimation unit 40 performs a sequential spectrum test to determine whether a portion of the frequency spectrum waveform of the second frame is within a predetermined frequency range that includes the fundamental frequency of the first frame, and whether a verification point pkR in the portion of the frequency spectrum waveform of the second frame is within a predetermined spectral intensity range that includes the spectral intensity of the fundamental frequency.
[0059] FIG. 10 is a diagram showing a predetermined frequency range Rf including the fundamental frequency in the first frame, and a predetermined spectral intensity range Rs including the spectral intensity of the fundamental frequency.
[0060] The predetermined frequency range Rf is a certain range of frequencies based on the fundamental frequency in a two-axis coordinate system represented by a frequency axis and a spectral intensity axis, for example, a frequency range of ±10 Hz based on the fundamental frequency. The predetermined spectral intensity range Rs is a certain range of spectral intensity based on the spectral intensity of the fundamental frequency in the two-axis coordinate system, for example, a spectral intensity range of ±10 dB based on the spectral intensity of the fundamental frequency. Note that the coordinate positions of the predetermined frequency range Rf and the predetermined spectral intensity range Rs in the two-axis coordinate system may vary depending on the frequency spectral waveform of each frame. In other words, the coordinate positions of the predetermined frequency range Rf and the predetermined spectral intensity range Rs in the first frame, the second frame, and the (N+1)th frame may differ.
[0061] The estimation unit 40 determines whether a verification point pkR in a part of the frequency spectrum waveform of the second frame is present in a predetermined frequency range Rf and a predetermined spectral intensity range Rs, and estimates the fundamental frequency of the second frame based on the determination. The processing performed by the estimation unit 40 will be specifically described below.
[0062] Fig. 11 is a diagram showing an example of estimating the fundamental frequency in frame 2. As described above, the frequency peak pk0 shown in Fig. 11 is the frequency peak located at the lowest frequency side among one or more frequency peaks pk extracted excluding subharmonics and disturbance noise.
[0063] FIG. 11(a) shows an example in which the frequency peak pk0 is located on the higher frequency side than the predetermined frequency range Rf, and the frequency spectrum waveform of the second frame crosses the upper limit of the predetermined frequency range Rf.
[0064] In this case, the estimation unit 40 determines that a portion of the frequency spectrum waveform of the second frame is within the predetermined frequency range Rf. Next, the estimation unit 40 determines whether a verification point pkR in the portion of the frequency spectrum waveform is within the predetermined spectral intensity range Rs. The verification point pkR is the point of the frequency spectrum waveform within the predetermined frequency range Rf where the spectral intensity is maximum. In this example, the intersection of the frequency spectrum waveform with the upper limit of the predetermined frequency range Rf is the verification point pkR, and this verification point pkR is within the predetermined spectral intensity range Rs. Therefore, the estimation unit 40 estimates the frequency f1a of the verification point pkR to be the fundamental frequency of the second frame.
[0065] In the above example, the frequency spectrum waveform of the second frame crosses the upper limit of the predetermined frequency range Rf, but this is not limiting. When the frequency spectrum waveform crosses the lower limit of the predetermined frequency range Rf, the estimation unit 40 may estimate that the frequency at the intersection of the frequency spectrum waveform with the lower limit of the predetermined frequency range Rf is the fundamental frequency of the second frame (not shown).
[0066] FIG. 11(b) shows an example in which the frequency peak pk0 is located within a predetermined frequency range Rf and a predetermined spectral intensity range Rs.
[0067] In this case, the estimation unit 40 determines that a portion of the frequency spectrum waveform of the second frame is within the predetermined frequency range Rf. Next, the estimation unit 40 determines whether a verification point pkR in the portion of the frequency spectrum waveform is within the predetermined spectral intensity range Rs. In this example, the frequency peak pk0 at which the spectral intensity is maximum is the verification point pkR, and this verification point pkR is within the predetermined spectral intensity range Rs. Therefore, the estimation unit 40 estimates the frequency f1b of the verification point pkR to be the fundamental frequency in the second frame.
[0068] In this way, when a portion of the frequency spectrum waveform of the second frame is within a predetermined frequency range Rf and the spectral intensity of the verification point pkR in the portion of the frequency spectrum waveform is within a predetermined spectral intensity range Rs, the estimation unit 40 estimates that the frequency of the verification point pkR is the fundamental frequency in the second frame.
[0069] On the other hand, if a part of the frequency spectrum waveform of the second frame does not fall within the predetermined frequency range Rf, or if the spectral intensity of the verification point pkR does not fall within the predetermined spectral intensity range Rs, the estimation unit 40 does not estimate the frequency of the verification point pkR as the fundamental frequency, and estimates the provisional fundamental frequency derived by the derivation unit 30 as the fundamental frequency of the second frame. If a part of the frequency spectrum waveform does not fall within the predetermined frequency range Rf, or if the spectral intensity of the verification point pkR does not fall within the predetermined spectral intensity range Rs, this may occur when the temporal continuity of the fundamental frequency is interrupted due to, for example, audio on / off, sudden pitch fluctuations, vibration mode conversion, etc.
[0070] Fig. 12 is a diagram showing another example of estimating the fundamental frequency in frame 2. The frequency peak pk0 shown in Fig. 12 is the frequency peak located at the lowest frequency side among one or more frequency peaks pk extracted excluding subharmonics and disturbance noise.
[0071] FIG. 12(a) shows an example in which the frequency peak pk0 is located on the higher frequency side than the predetermined frequency range Rf and on the smaller frequency side than the predetermined spectrum intensity range Rs.
[0072] The estimation unit 40 determines that a portion of the frequency spectrum waveform of the second frame is within the predetermined frequency range Rf. Next, the estimation unit 40 determines whether a verification point pkR in the portion of the frequency spectrum waveform is within the predetermined spectral intensity range Rs. In this example, the intersection of the frequency spectrum waveform with the upper limit of the predetermined frequency range Rf is the verification point pkR, but this verification point pkR does not exist within the predetermined spectral intensity range Rs. Therefore, the estimation unit 40 does not estimate the frequency of the verification point pkR as the fundamental frequency, but estimates the provisional fundamental frequency (f2a shown in FIG. 12(a)) derived by the derivation unit 30 for the second frame as the fundamental frequency of the second frame.
[0073] FIG. 12(b) shows an example in which the frequency peak pk0 is located on the higher frequency side than the predetermined frequency range Rf and on the larger frequency side than the predetermined spectrum intensity range Rs.
[0074] The estimation unit 40 determines that a portion of the frequency spectrum waveform of the second frame is within the predetermined frequency range Rf. Next, the estimation unit 40 determines whether a verification point pkR in the portion of the frequency spectrum waveform is within the predetermined spectral intensity range Rs. In this example, the intersection of the frequency spectrum waveform with the upper limit of the predetermined frequency range Rf is the verification point pkR, but this verification point pkR does not exist within the predetermined spectral intensity range Rs. Therefore, the estimation unit 40 does not estimate the frequency of the verification point pkR as the fundamental frequency, but estimates the provisional fundamental frequency (f2b shown in FIG. 12(b)) derived by the derivation unit 30 for the second frame as the fundamental frequency of the second frame.
[0075] FIG. 12(c) shows an example in which the frequency peak pk0 is located within the predetermined frequency range Rf and on the side greater than the predetermined spectrum intensity range Rs.
[0076] The estimation unit 40 determines that a portion of the frequency spectrum waveform of the second frame is within the predetermined frequency range Rf. Next, the estimation unit 40 determines whether a verification point pkR in the portion of the frequency spectrum waveform is within the predetermined spectral intensity range Rs. In this example, the frequency peak pk0 at which the spectral intensity is maximum is the verification point pkR, and this verification point pkR is not within the predetermined spectral intensity range Rs. Therefore, the estimation unit 40 does not estimate the frequency of the verification point pkR as the fundamental frequency, but estimates the provisional fundamental frequency (f2c shown in FIG. 12(c)) derived by the derivation unit 30 for the second frame as the fundamental frequency of the second frame.
[0077] In this way, if the frequency peak pk0 of the frequency spectrum waveform of the second frame is similar to the fundamental frequency and spectral intensity of the first frame, the estimation unit 40 estimates that the frequency derived by the sequential spectrum test is the fundamental frequency of the second frame. On the other hand, if the frequency peak pk0 of the frequency spectrum waveform of the second frame is not similar to the fundamental frequency and spectral intensity of the first frame, the estimation unit 40 estimates that the tentative fundamental frequency derived by the dominant spectrum test of the second frame is the fundamental frequency of the second frame.
[0078] The estimation unit 40 similarly estimates the fundamental frequency for the third frame and thereafter, i.e., the (N+1)th frame. The method for estimating the fundamental frequency for the (N+1)th frame is the same as the method for estimating the fundamental frequency for the second frame (see FIGS. 10 to 12).
[0079] That is, when estimating the fundamental frequency in the (N+1)th frame, the estimation unit 40 performs a sequential spectrum test to determine whether a part of the frequency spectrum waveform of the (N+1)th frame is within a predetermined frequency range Rf and whether the spectral intensity of a verification point pkR in the part of the frequency spectrum waveform is within a predetermined spectral intensity range Rs.
[0080] If a part of the frequency spectrum waveform of the (N+1)th frame is present in a predetermined frequency range Rf that includes the fundamental frequency of the Nth frame, and the spectral intensity of a verification point pkR, which has the maximum spectral intensity in the predetermined frequency range Rf in the frequency spectrum waveform of the (N+1)th frame, is present in a predetermined spectral intensity range Rs, the estimation unit 40 estimates that the frequency of the verification point pkR is the fundamental frequency of the (N+1)th frame.On the other hand, if a part of the frequency spectrum waveform of the (N+1)th frame is not present in the predetermined frequency range Rf or the spectral intensity of the verification point pkR is not present in the predetermined spectral intensity range Rs, the estimation unit 40 estimates that the tentative fundamental frequency derived by the derivation unit 30 for the (N+1)th frame is the fundamental frequency of the (N+1)th frame.
[0081] FIG. 13 is a diagram showing an example of displaying fundamental frequencies and spectral intensities.
[0082] 13, the speech analysis device 10 outputs a table of the fundamental frequency and the spectral intensity of the fundamental frequency for each of a plurality of frames. Alternatively, the speech analysis device 10 may output the fundamental frequency and the spectral intensity of the fundamental frequency in time series on a three-axis coordinate system consisting of a frequency axis, a spectral intensity axis, and a time axis.
[0083] As described above, the speech analysis device 10 includes a derivation unit 30 that performs a dominant spectrum test and an estimation unit 40 that performs a sequential spectrum test. For example, if a portion of the frequency spectrum waveform of the second frame is within a predetermined frequency range Rf that includes the fundamental frequency of the first frame, and if the spectral intensity of a verification point pkR, which has the maximum spectral intensity in the predetermined frequency range Rf in the frequency spectrum waveform of the second frame, is within a predetermined spectral intensity range Rs, the estimation unit 40 estimates that the frequency of the verification point pkR is the fundamental frequency of the second frame. On the other hand, if a portion of the frequency spectrum waveform of the second frame is not within the predetermined frequency range Rf or the spectral intensity of the verification point pkR is not within the predetermined spectral intensity range Rs, the estimation unit 40 estimates that the provisional fundamental frequency derived by the derivation unit 30 for the second frame is the fundamental frequency of the second frame. These configurations enable improved estimation accuracy of the fundamental frequency of speech.
[0084] If the estimation of the fundamental frequency by the estimation unit 40 is incomplete, the voice analysis device 10 may correct the fundamental frequency by the correction unit 50 described below.
[0085] For example, if the fundamental frequency estimated by the estimation unit 40 is incorrectly detected as a harmonic frequency or a subharmonic frequency, the correction unit 50 sets the frequency of the first harmonic corresponding to the harmonic as the corrected fundamental frequency. Here, the harmonic frequency is the second harmonic of the fundamental frequency, and the subharmonic frequency is the 1.5th harmonic, both of which are frequency bands that are prone to being incorrectly detected in calculations.
[0086] Furthermore, if the spectral intensity of the fundamental frequency estimated by the estimation unit 40 is smaller than a predetermined spectral intensity, the correction unit 50 invalidates the fundamental frequency estimated by the estimation unit 40. By invalidating the fundamental frequency, a frame having this frequency spectral waveform is deemed to be a silent frame in which no human voice is present. Note that the predetermined spectral intensity is, for example, a spectral intensity that is within 30 dB of the average sound pressure level of the entire sample. Note that the sound pressure level may be calibrated to within 30 dB during or immediately after collecting the sound data.
[0087] If the fundamental frequency estimated by the estimation unit 40 is less than 50 Hz, the correction unit 50 invalidates the fundamental frequency estimated by the estimation unit 40. By invalidating the fundamental frequency, a frame having this frequency spectrum waveform is deemed to be a silent frame in which no human voice is present. Note that the frequency used as the judgment criterion above is a value based on the commercial power frequency or a frequency band in which the fundamental frequency is generally unlikely to exist, and is not limited to 50 Hz but may be 60 Hz.
[0088] [Outline of speech analysis method] The voice analysis method according to the embodiment will be described with reference to FIG.
[0089] FIG. 14 is a flowchart showing a speech analysis method according to an embodiment.
[0090] As shown in FIG. 14, the speech analysis method includes a waveform acquisition step S10 for acquiring a frequency spectrum waveform, a first derivation step S11 for deriving a fundamental frequency in a first frame, and a step S20 for estimating a fundamental frequency in a second frame.
[0091] The speech analysis method also includes a second derivation step S12 of deriving a tentative fundamental frequency in the second frame based on the frequency spectrum waveform of the second frame, and an (N+1)th derivation step S13 of deriving a tentative fundamental frequency in the (N+1)th frame (N is an integer greater than or equal to 2) based on the frequency spectrum waveform of the (N+1)th frame.
[0092] The speech analysis method also includes step S30 of estimating the fundamental frequency in the (N+1)th frame based on the frequency spectrum waveform of the (N+1)th frame. Each step will be described below.
[0093] First, in a waveform acquisition step S10, the voice analysis device 10 acquires frequency spectrum waveforms of a plurality of frames, with one frame being a predetermined period, based on a signal containing speech (S10). For example, the voice analysis device 10 adjusts the frequency spectrum waveform so that the average of the maximum values of a plurality of spectral intensities included in the frequency spectrum waveform falls within a predetermined range, and acquires the adjusted frequency spectrum waveform.
[0094] Next, in a first derivation step S11, the voice analysis device 10 derives the fundamental frequency of the first frame based on the frequency spectrum waveform of the first frame among the multiple frames (S11).
[0095] The first derivation step S11 includes the steps of: deriving the maximum peak pm of the spectral intensity of the frequency spectral waveform in the first frame; extracting one or more frequency peaks pk that have maximum values greater than a predetermined spectral intensity from a frequency region located on the lower frequency side of the frequency at which the spectral intensity reaches the maximum peak pm; and extracting the frequency peak pk0 located on the lowest frequency side from the one or more frequency peaks pk, and setting the frequency corresponding to the frequency peak pk0 as the fundamental frequency in the first frame.
[0096] Furthermore, in a second derivation step S12, the voice analysis device 10 derives a tentative fundamental frequency in the second frame based on the frequency spectrum waveform of the second frame (S12).
[0097] The second derivation step S12 includes the steps of: deriving the maximum peak pm of the spectral intensity of the frequency spectral waveform in the second frame; extracting one or more frequency peaks pk that have maximum values greater than a predetermined spectral intensity from a frequency region located on the lower frequency side of the frequency at which the spectral intensity reaches the maximum peak pm; and extracting the frequency peak pk0 located on the lowest frequency side from the one or more frequency peaks pk, and setting the frequency corresponding to the frequency peak pk0 as the provisional fundamental frequency in the second frame.
[0098] Furthermore, in an (N+1)th derivation step S13, the speech analysis device 10 derives a tentative fundamental frequency in the (N+1)th frame based on the frequency spectrum waveform of the (N+1)th frame (S13).
[0099] The (N+1)th derivation step S13 includes the steps of: deriving the maximum peak pm of the spectral intensity of the frequency spectral waveform in the (N+1)th frame; extracting one or more frequency peaks pk that have maximum values greater than a predetermined spectral intensity from a frequency region located on the lower frequency side of the frequency at which the spectral intensity reaches the maximum peak pm; and extracting the frequency peak pk0 located on the lowest frequency side from among the one or more frequency peaks pk, and setting the frequency corresponding to the frequency peak pk0 as the provisional fundamental frequency in the Nth frame.
[0100] In addition, in the step of deriving the maximum peak pm of the spectral intensity included in each of the above derivation steps S11 to S13, the speech analysis device 10 may generate a speech spectral waveform by removing data that is different from the speech from the frequency spectral waveform, and derive the maximum peak pm of the spectral intensity based on the speech spectral waveform.
[0101] Next, in step S20 of estimating the fundamental frequency, the voice analysis device 10 estimates the fundamental frequency of the second frame based on the frequency spectrum waveform of the second frame among the multiple frames (S20).
[0102] In step S20, the speech analysis device 10 determines whether a portion of the frequency spectrum waveform of the second frame is within a predetermined frequency range Rf that includes the fundamental frequency of the first frame, and whether a verification point pkR in the portion of the frequency spectrum waveform is within a predetermined spectral intensity range Rs that includes the spectral intensity of the fundamental frequency. The verification point pkR is the point of the frequency spectrum waveform that is within the predetermined frequency range Rf where the spectral intensity is maximum.
[0103] The speech analysis device 10 estimates that the frequency of the verification point pkR is the fundamental frequency of the second frame when a portion of the frequency spectrum waveform of the second frame is within the predetermined frequency range Rf and the spectral intensity of the verification point pkR is within the predetermined spectral intensity range Rs. On the other hand, the speech analysis device 10 estimates that the provisional fundamental frequency derived in the second derivation step S12 is the fundamental frequency of the second frame when a portion of the frequency spectrum waveform of the second frame is not within the predetermined frequency range Rf or the spectral intensity of the verification point pkR is not within the predetermined spectral intensity range Rs.
[0104] Furthermore, in step S30 of estimating the fundamental frequency, the voice analysis device 10 estimates the fundamental frequency in the (N+1)th frame based on the frequency spectrum waveform of the (N+1)th frame (S30).
[0105] In step S30, the speech analysis device 10 determines whether a portion of the frequency spectrum waveform of the (N+1)th frame is within a predetermined frequency range Rf that includes the fundamental frequency in the Nth frame, and whether a verification point pkR in the portion of the frequency spectrum waveform is within a predetermined spectral intensity range Rs that includes the spectral intensity of the fundamental frequency.
[0106] The speech analysis device 10 estimates that the frequency of the verification point pkR is the fundamental frequency of the (N+1)th frame if a part of the frequency spectrum waveform of the (N+1)th frame is within a predetermined frequency range Rf that includes the fundamental frequency of the Nth frame, and if the spectral intensity of the verification point pkR, which has the maximum spectral intensity in the predetermined frequency range Rf of the frequency spectrum waveform, is within a predetermined spectral intensity range Rs.On the other hand, if a part of the frequency spectrum waveform of the (N+1)th frame is not within the predetermined frequency range Rf or the spectral intensity of the verification point pkR is not within the predetermined spectral intensity range Rs, the speech analysis device 10 estimates that the tentative fundamental frequency derived in the (N+1)th derivation step S13 is the fundamental frequency of the (N+1)th frame.
[0107] The predetermined frequency range Rf included in each of steps S20 and S30 is a fixed range of frequencies based on the fundamental frequency. The predetermined spectral intensity range Rs is a fixed range of spectral intensity based on the spectral intensity of the fundamental frequency. The ranges of the predetermined frequency range Rf and the predetermined spectral intensity range Rs may vary depending on the frequency spectral waveform of each frame.
[0108] As described above, the speech analysis method includes a derivation step of performing a dominant spectrum test and an estimation step of performing a sequential spectrum test. For example, in the estimation step, if a portion of the frequency spectrum waveform of the second frame is within a predetermined frequency range Rf that includes the fundamental frequency of the first frame, and if the spectral intensity of a verification point pkR, which has the maximum spectral intensity in the predetermined frequency range Rf, of the frequency spectrum waveform is within a predetermined spectral intensity range Rs, the frequency of the verification point pkR is estimated to be the fundamental frequency of the second frame. On the other hand, in the estimation step, if a portion of the frequency spectrum waveform of the second frame is not within the predetermined frequency range Rf or the spectral intensity of the verification point pkR is not within the predetermined spectral intensity range Rs, the provisional fundamental frequency derived in the derivation step is estimated to be the fundamental frequency of the second frame. These configurations enable improved estimation accuracy of the fundamental frequency of speech.
[0109] In the above example, step S20 of estimating the fundamental frequency in the second frame is performed after the (N+1)th derivation step S13, but this is not limiting. Step S20 of estimating the fundamental frequency in the second frame may be performed between the second derivation step S12 and the (N+1)th derivation step S13, or may be performed simultaneously with the (N+1)th derivation step S13.
[0110] For example, the speech analysis method of this embodiment may include deriving tentative fundamental frequencies in the (2+1)th frame, the (3+1)th frame, ... in the (N+1)th derivation step S13, and estimating the fundamental frequencies in the (2+1)th frame, the (3+1)th frame, ... in step S30. In another case, the speech analysis method may perform step S20 immediately after the second derivation step S12, deriving the tentative fundamental frequency in the (2+1)th frame after step S20, estimating the fundamental frequency in the (2+1)th frame immediately thereafter, deriving the tentative fundamental frequency in the (3+1)th frame again immediately thereafter, estimating the fundamental frequency in the (3+1)th frame.
[0111] In addition, although the above example shows the execution of the second derivation step S12 and the (N+1)th derivation step S13, the present invention is not limited to this. For example, if step S20 is provided before the second derivation step S12 and step S30 is provided before the third derivation step S13, there may be cases where the second derivation step S12 and the (N+1)th derivation step S13 do not need to be executed.
[0112] For example, when the speech analysis device 10 adopts the result of the sequential spectrum test in step S20 of estimating the fundamental frequency in the second frame, it is not necessary to derive a tentative fundamental frequency in the second frame, and therefore it is not necessary to perform the second derivation step S12. Also, when the speech analysis device 10 adopts the result of the sequential spectrum test in step S30 of estimating the fundamental frequency, it is not necessary to derive a tentative fundamental frequency in the (N+1)th frame, and therefore it is not necessary to perform the (N+1)th derivation step S13.
[0113] The speech analysis method may further include the following steps:
[0114] For example, the voice analysis method may include a step of correcting the fundamental frequency when the fundamental frequency estimated in the step of estimating the fundamental frequency is the frequency of a harmonic overtone. In the step of correcting the fundamental frequency, the frequency of the first harmonic overtone corresponding to the harmonic overtone may be set as the corrected fundamental frequency.
[0115] The voice analysis method may also include a step of invalidating the fundamental frequency estimated in the step of estimating the fundamental frequency if the spectral intensity of the fundamental frequency estimated in the step of estimating the fundamental frequency is smaller than a predetermined spectral intensity.
[0116] The voice analysis method may also include a step of invalidating the fundamental frequency estimated in the fundamental frequency estimating step if the frequency of the fundamental frequency estimated in the fundamental frequency estimating step is less than 50 Hz.
[0117] (Summary of the overview) An example of a speech analysis method according to one embodiment of the present disclosure will be described below.
[0118] The speech analysis method of Example 1 includes a waveform acquisition step of acquiring frequency spectrum waveforms of a plurality of frames, each frame representing a predetermined period, based on a signal including speech; a first derivation step of deriving a fundamental frequency in a first frame based on the frequency spectrum waveform of a first frame among the plurality of frames; and a step of estimating a fundamental frequency in a second frame based on the frequency spectrum waveform of a second frame among the plurality of frames, in which the fundamental frequency estimation step determines whether a portion of the frequency spectrum waveform of the second frame is within a predetermined frequency range Rf that includes the fundamental frequency of the first frame, and whether a verification point pkR in the portion of the frequency spectrum waveform is within a predetermined spectral intensity range Rs that includes the spectral intensity of the fundamental frequency, and estimates the fundamental frequency in the second frame based on the determination.
[0119] In this way, by determining whether a part of the frequency spectrum waveform of the second frame is within a predetermined frequency range Rf that includes the fundamental frequency of the first frame and whether a verification point pkR in the part of the frequency spectrum waveform is within a predetermined spectral intensity range Rs, the fundamental frequency can be estimated based on the change in the fundamental frequency along the time series of the speech, thereby improving the accuracy of estimating the fundamental frequency of the speech.
[0120] The voice analysis method of Example 2 is the voice analysis method described in Example 1, and the verification point pkR is the point at which the spectral intensity is maximum among the frequency spectral waveforms present in the above-mentioned predetermined frequency range.
[0121] This allows the fundamental frequency to be estimated using the verification point pkR at which the spectral intensity is maximized, thereby improving the accuracy of estimating the fundamental frequency of speech.
[0122] The speech analysis method of Example 3 is the speech analysis method of Example 1 or 2, further including a second derivation step of deriving a tentative fundamental frequency in the second frame based on the frequency spectral waveform of the second frame, wherein in the fundamental frequency estimating step, if a portion of the frequency spectral waveform of the second frame is within a predetermined frequency range Rf and the spectral intensity of the verification point pkR is within a predetermined spectral intensity range Rs, the frequency of the verification point pkR is estimated to be the fundamental frequency in the second frame, and if a portion of the frequency spectral waveform of the second frame is not within the predetermined frequency range Rf or the spectral intensity of the verification point pkR is not within the predetermined spectral intensity range Rs, the provisional fundamental frequency derived in the second derivation step is estimated to be the fundamental frequency in the second frame.
[0123] This allows the fundamental frequency to be estimated in response to changes in the fundamental frequency along the time series of the voice. Furthermore, the fundamental frequency can be estimated in response to, for example, on / off of the voice, sudden pitch fluctuations, vibration mode conversion, etc. This improves the accuracy of estimating the fundamental frequency of the voice in the second frame.
[0124] A speech analysis method of Example 4 is the speech analysis method of Example 3, further comprising: an (N+1)th derivation step of deriving a tentative fundamental frequency in the (N+1)th frame (N is an integer equal to or greater than 2) based on a frequency spectrum waveform of the (N+1)th frame among the plurality of frames; and a step of estimating the fundamental frequency in the (N+1)th frame based on the frequency spectrum waveform of the (N+1)th frame, wherein in the step of estimating the fundamental frequency in the (N+1)th frame, a part of the frequency spectrum waveform of the (N+1)th frame is a predetermined frequency including the fundamental frequency of the Nth frame. If the spectral intensity of a verification point pkR, which is present in the range Rf and has the maximum spectral intensity in the specified frequency range Rf in the frequency spectral waveform, is present in the specified spectral intensity range, the frequency of the verification point pkR is estimated to be the fundamental frequency in the (N+1)th frame.If a part of the frequency spectral waveform of the (N+1)th frame is not present in the specified frequency range Rf or the spectral intensity of the verification point pkR is not present in the specified spectral intensity range Rs, the provisional fundamental frequency derived in the (N+1)th derivation step is estimated to be the fundamental frequency in the (N+1)th frame.
[0125] This allows the fundamental frequency to be estimated in response to changes in the fundamental frequency along the time series of the voice. Furthermore, the fundamental frequency can be estimated in response to, for example, on / off of the voice, sudden pitch fluctuations, vibration mode conversion, etc. This improves the accuracy of estimating the fundamental frequency of the voice in the (N+1)th frame.
[0126] A voice analysis method of Example 5 is the voice analysis method according to any one of Examples 1 to 4, wherein the first derivation step includes the steps of: deriving the maximum peak pm of the spectral intensity of the frequency spectrum waveform in the first frame; extracting one or more frequency peaks pk that have a maximum value greater than a predetermined spectral intensity from a frequency region located on the lower frequency side than the frequency at which the spectral intensity becomes the maximum peak pm; and extracting a frequency peak pk0 located on the lowest frequency side from the one or more frequency peaks pk, and setting the frequency corresponding to the frequency peak pk0 as the fundamental frequency in the first frame.
[0127] This allows the fundamental frequency in the first frame to be derived accurately, thereby improving the accuracy of estimating the fundamental frequency in the second frame.
[0128] The voice analysis method of Example 6 is the voice analysis method described in Example 3, and the second derivation step includes the steps of: deriving the maximum peak pm of the spectral intensity of the frequency spectrum waveform in the second frame; extracting one or more frequency peaks pk that have maximum values greater than a predetermined spectral intensity from a frequency region located on the lower frequency side than the frequency at which the spectral intensity becomes the maximum peak pm; and extracting the frequency peak pk0 located on the lowest frequency side from the one or more frequency peaks pk, and setting the frequency corresponding to the frequency peak pk0 as the provisional fundamental frequency in the second frame.
[0129] This makes it possible to accurately derive the tentative fundamental frequency in the second frame, thereby improving the accuracy of estimating the fundamental frequency in the second frame.
[0130] A speech analysis method of Example 7 is the speech analysis method of Example 4, and the (N+1)-th derivation step includes the steps of: deriving the maximum peak pm of the spectral intensity of the frequency spectral waveform in the (N+1)-th frame; extracting one or more frequency peaks pk that have maximum values greater than a predetermined spectral intensity from a frequency region located on the lower frequency side than the frequency at which the spectral intensity becomes the maximum peak pm; and extracting a frequency peak pk0 located on the lowest frequency side from the one or more frequency peaks pk, and setting the frequency corresponding to the frequency peak pk0 as a provisional fundamental frequency in the N-th frame.
[0131] This allows the tentative fundamental frequency in the (N+1)th frame to be derived accurately, thereby improving the estimation accuracy of the fundamental frequency in the (N+1)th frame.
[0132] The voice analysis method of Example 8 is the voice analysis method described in any one of Examples 1 to 7, wherein the predetermined frequency range Rf is a certain range of frequencies based on the fundamental frequency, and the predetermined spectral intensity range Rs is a certain range of spectral intensity based on the spectral intensity of the fundamental frequency.
[0133] This allows the tentative fundamental frequencies in the second and (N+1) frames to be derived accurately, thereby improving the estimation accuracy of the fundamental frequencies in the second and (N+1) frames.
[0134] The voice analysis method of Example 9 is the voice analysis method according to any one of Examples 5 to 7, in which, in the step of deriving the maximum peak pm of the spectral intensity, a voice spectrum waveform is generated by removing data other than the voice from the frequency spectrum waveform, and the maximum peak pm of the spectral intensity is derived based on the voice spectrum waveform.
[0135] This allows the maximum peaks of the speech spectrum waveform in the first, second, and (N+1) frames to be accurately derived. This makes it possible to accurately derive the fundamental frequency in the first frame and the tentative fundamental frequencies in the second and (N+1) frames. This improves the estimation accuracy of the fundamental frequencies in the second and (N+1) frames.
[0136] The voice analysis method of Example 10 is the voice analysis method according to any one of Examples 1 to 9, wherein in the waveform acquisition step, the frequency spectrum waveform is adjusted so that the average of the maximum values of multiple spectral intensities contained in the frequency spectrum waveform falls within a predetermined range, and the adjusted frequency spectrum waveform is acquired.
[0137] By adjusting the frequency spectrum waveform in this way, it becomes possible to accurately derive the fundamental frequency in the first frame and the tentative fundamental frequencies in the second and (N+1) frames, thereby improving the estimation accuracy of the fundamental frequencies in the second and (N+1) frames.
[0138] The voice analysis method of Example 11 is the voice analysis method according to any one of Examples 1 to 10, further comprising a step of correcting the fundamental frequency when the fundamental frequency estimated in the step of estimating the fundamental frequency is the frequency of an overtone, and in the step of correcting the fundamental frequency, the frequency of the first overtone corresponding to the overtone is set as the corrected fundamental frequency.
[0139] This makes it possible to prevent the frequency of an overtone from being erroneously estimated as the fundamental frequency, thereby improving the accuracy of estimating the fundamental frequency of speech.
[0140] The speech analysis method of Example 12 is the speech analysis method according to any one of Examples 1 to 11, further including a step of invalidating the fundamental frequency estimated in the step of estimating the fundamental frequency when the spectral intensity of the fundamental frequency estimated in the step of estimating the fundamental frequency is smaller than a predetermined spectral intensity.
[0141] This makes it possible to prevent errors in estimating the fundamental frequency when there is no human voice, thereby improving the accuracy of estimating the fundamental frequency of voice.
[0142] The voice analysis method of Example 13 is the voice analysis method according to any one of Examples 1 to 12, further including a step of invalidating the fundamental frequency estimated in the step of estimating the fundamental frequency if the frequency of the fundamental frequency estimated in the step of estimating the fundamental frequency is less than 50 Hz.
[0143] This makes it possible to prevent errors in estimating the fundamental frequency when there is no human voice, thereby improving the accuracy of estimating the fundamental frequency of voice.
[0144] A speech analysis device 10 of Example 14 includes a waveform acquisition unit 20 that acquires frequency spectrum waveforms of multiple frames, each frame representing a predetermined period, based on a signal containing speech, a derivation unit 30 that derives a fundamental frequency of a first frame based on the frequency spectrum waveform of a first frame of the multiple frames, and an estimation unit 40 that estimates a fundamental frequency of a second frame based on the frequency spectrum waveform of a second frame of the multiple frames. The estimation unit 40 determines whether a portion of the frequency spectrum waveform of the second frame is within a predetermined frequency range Rf that includes the fundamental frequency of the first frame and whether a verification point pkR in the portion of the frequency spectrum waveform is within a predetermined spectral intensity range Rs that includes the spectral intensity of the fundamental frequency, and estimates the fundamental frequency of the second frame based on the determination.
[0145] In this way, by determining whether a part of the frequency spectrum waveform of the second frame is within a predetermined frequency range Rf that includes the fundamental frequency of the first frame, and whether a verification point pkR in the part of the frequency spectrum waveform is within a predetermined spectral intensity range Rs, the fundamental frequency can be estimated by focusing on the change in the fundamental frequency along the time series of the speech, thereby improving the accuracy of estimating the fundamental frequency of the speech. [Example]
[0146] (Detailed explanation of voice analysis method) 1. SUMMARY [the purpose] Fundamental frequency (fo) is closely related to pitch, a psychological quantity corresponding to the pitch of a sound, and is a physical quantity that roughly corresponds to the vibration frequency of the vocal cords. fo is used to quantify information about the vocal cords, the source of speech, and in important technologies such as speech recognition. However, the accuracy of fo estimation for hoarse speech remains low, and no established fo estimation algorithm has existed. In this study, we developed a new algorithm, "SFEEDS (Spectral-based fo Estimator Emphasized by Domination and Sequence)," which is an improvement on the Spectrum method, as a method that enables good fo estimation for not only normal speech but also hoarse speech, and compared it with conventional estimation methods.
[0147] [method] Using 455 speech samples, we calculated the estimated fo using conventional methods: autocorrelation, cross-correlation, SWIPE', BaNa, and SFEEDS. The true fo value was defined as the lowest frequency of the most dominant harmonic complex on the spectrogram, and we evaluated the agreement with each estimation method and the difference in accuracy depending on whether the patient had hoarseness or not.
[0148] [result] The accuracy of fo estimation by SFEEDS was significantly higher than that of conventional methods for both the entire speech sample set (n=455) and the sample set that did not include coarse hoarseness (n=237). Furthermore, for the sample set that included coarse hoarseness (n=218) compared to the sample set that did not include coarse hoarseness (n=237), the accuracy of fo estimation by conventional methods was significantly lower, whereas almost no change was observed with SFEEDS, making it possible to significantly reduce subharmonics error.
[0149] [Conclusion] SFEEDS, a newly developed fo estimation method, enables highly accurate fo estimation for not only normal speech but also hoarse speech.
[0150] 2. INTRODUCTION Pitch, the subjective highness of a sound, is defined by the perceived height of the sound as perceived by the human ear. In a perfectly periodic speech signal, pitch corresponds to the sound produced by the fundamental frequency (fo), which is the reciprocal of the speech signal's period. However, due to the complex vibration patterns of the vocal folds, pitch does not necessarily correspond to fo. Previous literature has shown that sounds can be periodic yet "outside the realm of pitch" (Ritsma 1962) (Pressnitzer et al. 2001). Conversely, some sounds evoke pitch even though they are not periodic (Miller and Taylor 1948; Yost 1996). Furthermore, when subharmonics are present, the perceived pitch can diverge from fo, as two different intervals are perceived simultaneously during production (Cavalli L, Hirson A. Diplophonia reappraised. J Voice. 1999). The method for estimating fo is often called the "Pitch Detection Algorithm (PDA)" (Hess, 1983), but for the reasons mentioned above, pitch and fo should be clearly distinguished. Therefore, in this study, we will not use the term PDA, but will refer to it as the "fo Estimation Algorithm."
[0151] [About fo] Accurate and robust fo estimation algorithms can be used in a variety of applications. For example, recognizing speech tones can improve speech recognition by distinguishing between homonyms (C. Wang, “Prosodic modeling for improved speech recognition and understanding,” Ph.D. dissertation, Massachusetts Institute of Technology, 2001). It is also useful for emotion detection based on prosodic changes (OW Kwon. Emotion recognition by speech signals. 2003). Accurate fo estimation is also essential for various speech processing applications, such as voice quality assessment, speech synthesis, speech coding, and speaker recognition.
[0152] In the simplified spectral waveform, the fo of the voice waveform (vocal cord original sound) emitted from the vocal cords appears as the frequency peak with the lowest frequency and the strongest spectral intensity, as shown in FIG. 15a.
[0153] FIG. 15a is a diagram showing the spectral waveform of a vocal cord original sound.
[0154] In the spectral waveform of a vocal fold primary sound, the frequency characteristic corresponding to fo is strongest, and the harmonic structure attenuates as the frequency increases. This phenomenon is observed in vocal fold primary sounds, which are composed solely of vocal fold vibration. In actual speech, the formants are filtered and the frequency components are increased or decreased due to resonance in the vocal tract and antiresonance in the nasal cavity (Dejonckere, 1995) (Dejonckere and Lebacq (1996)) (de Krom, 1995; Fukazawa, el-Assuooty, & Honjo, 1988) (Martin, Fitch, & Wolfe, 1995) (Titze, 1995). Finally, the lowest frequency peak (not the highest in spectral intensity) of the harmonic complex sound is considered to be fo (see Figure 15b).
[0155] FIG. 15b shows the spectral waveform of speech filtered through the vocal tract and uttered from the mouth.
[0156] However, in a real environment, the quality of the input audio signal can be significantly degraded by noise from the recording equipment or background noise in the audible range. In such cases, noise is added to components other than harmonics, so it is not robust to consider the lowest frequency peak as the fundamental frequency.
[0157] For this reason, various more practical, noise-resistant pitch detection algorithms have been developed, as listed below.
[0158] [About existing fo estimation methods] Existing fo estimation algorithms can be broadly divided into three groups. First, those that primarily utilize characteristics in the time domain. Second, those that utilize characteristics in the frequency domain. Third, hybrid methods that combine characteristics in both the time and frequency domains (Titze IR, Liang H. Comparison of Fo extraction methods for high-precision voice perturbation measurements. J Speech Hear Res. 1993).
[0159] (a) Utilizing the characteristics of the time domain Among methods that utilize time-domain characteristics (Rabiner L. A comparative performance study of several pitch detection algorithms. 1976), methods that use signal correlation (MJ Ross, Average magnitude difference function pitch extractor. 1974) are common. If the fundamental period of speech is T0, calculating the autocorrelation (AC) or cross-correlation (CC) of the speech waveform will yield high correlation values at integer multiples of T0. The AC or CC implementation in Praat is widely used, and it considers the maximum autocorrelation value of short sound segments as pitch candidates and uses the Viterbi algorithm to find the least-cost path through all segments to select the optimal pitch candidate for each segment (P. Boersma, “Accurate short-term analysis of the fundamental frequency and the harmonics-to-noise ratio of a sampled sound,” in Proc. of the Institute of Phonetic Sciences, vol. 17, no. 1193, Amsterdam, 1993). Furthermore, Praat 6.4 newly implements "filtered AC" and "filtered CC," which use a low-pass filter before performing the normal AC or CC, reducing the adverse effects on fo estimation caused by formant transitions that occur when transitioning from consonants to vowels. In conjunction with the Praat 6.4 update, Boersma et al. state that Filtered AC is more advantageous when there is intonation, such as pitch fluctuations, and that raw CC is the preferred method for speech analysis where noise should not be removed by filtering.
[0160] The peak picking method (W. Hess. Pitch Determination of Speech Signals: Algorithms and Devices. 1983) and the zero-crossing method (W. Hess. Pitch Determination of Speech Signals: Algorithms and Devices. 1983) are also methods for estimating fo by detecting events in the time domain, but it has been concluded that their fo estimation accuracy is inferior to methods that use signal correlation such as AC and CC (Titze IR, Liang H. Comparison of Fo extraction methods for high precision voice perturbation measurements. J Speech Hear Res. 1993).
[0161] YIN uses the cross-correlation function of speech as its basis, and by reducing unnecessary peaks and eliminating unnecessary calculations, it has been possible to reduce errors to less than one-third compared to conventional methods used before 2002 (A. Cheveigne and H. Kawahara, “YIN, a fundamental frequency estimator for speech and music,” J. Acoust. Soc. 2002.).
[0162] (b) Utilizing the characteristics of the frequency domain As mentioned above, the power spectrum of a highly periodic speech signal forms a harmonic complex tone with spectral peaks at integer multiples of fo. Spectral methods (Noll, A.M. "Short-term spectrum and "cepstrum" techniques for vocal pitch detection, 1964) estimate fo by extracting the frequency of the lowest peak in the power spectrum or estimating the peak interval of the harmonic structure. However, estimation methods based on the power spectrum are contaminated by effects caused by articulatory filters, such as formants, and therefore require a method to remove these effects.
[0163] The cepstrum method is a fundamental frequency estimation method developed by Noll in 1964 (Noll, A.M. "Short-term spectrum and 'cepstrum' techniques for vocal pitch detection." J. Acoust. Soc. Am. 41 (1964)) (A.M. Noll, "Cepstrum pitch determination," J. Acoust. Soc. Am. , vol.41, no.2, pp.293-309, 1967.). It uses a parameter called quefrency, which is obtained by inverse Fourier transform of the logarithmic power spectrum. The cepstrum can separate the effects of articulatory filters and is a typical parameter not only for fo estimation but also for speech analysis. fo is estimated by extracting the peak caused by fo in the higher-order quefrency region and calculating its inverse. However, the Cepstrum method has poor performance in terms of fo estimation accuracy and is said to be susceptible to noise (Ba H, Yang N. BaNa: a hybrid approach for noise-resilient pitch detection. IEEE Statistical Signal Processing Workshop. 2012).
[0164] SWIPE and SWIPE' (A. Camacho and JG Harris, "A sawtooth waveform inspired pitch estimator for speech and music," J. Acoust. Soc. 2008.), proposed in 2008, improved the accuracy of fo estimation by focusing on the harmonic structure of the power spectrum and implementing techniques to reduce errors. In particular, SWIPE' succeeded in significantly reducing subharmonic error, a problem in fo estimation algorithms, by using only the first and prime harmonics.
[0165] (c) Hybrid Method BaNa (Ba H, Yang N. BaNa: a hybrid approach for noise-resilient pitch detection. IEEE Statistical Signal Processing Workshop. 2012) is an algorithm that combines classical approaches consisting of harmonic frequency ratios (M.R. Schroeder. Period Histogram and Product Spectrum: New Methods for Fundamental-Frequency Measurement. 1968) and cepstrum analysis. A study evaluating the accuracy of several state-of-the-art fo estimation algorithms using online speech and noise databases showed that BaNa achieved the best pitch detection accuracy across all noise types and SNR values investigated (A Comparative Analysis of Pitch Detection Methods Under the Influence of Different Noise Conditions. J.voice. 2015).
[0166] [Limitations of existing fo estimation algorithms] The development of various fo estimation methods as described above has made it possible to separate harmonic structure and non-periodic noise with extremely high accuracy. However, existing fo estimation algorithms have been limited in that they have only been verified on normal speech. In particular, when pathological speech, such as roughness, contains modulation frequencies other than fo, known as subharmonics, estimating fo using existing methods becomes even more difficult.
[0167] Subharmonic structures appear as distinct peaks between two consecutive harmonic structures corresponding to the fundamental frequency (fo). Subharmonics typically divide the harmonic interval into multiple equal intervals (e.g., 1 / 2, 1 / 3, 1 / 4) (Baken RJ. Clinical Measurement of Speech and Voice. 1987.). In laryngeal aerodynamics, it has been suggested that subharmonics can arise from two or more different oscillators (multi-oscillators) (Dejonckere, P.H., An Analysis of the Diplophonia Phenomenon. 1983) (Ward, P.H., Annals of Otology, Rhinology & Laryngology. 1969) (Neubauer, J., Patio-temporal analysis of irregular vocal fold oscillations: Biphonation due to desynchronization of spatial modes. Journal of the Acoustical Society of America. 2001) (Kimura, M., Arytenoid Adduction for Correcting Vocal Fold Asymmetry. 2010). When two nonlinear oscillators are coupled, they exhibit very complex dynamics, including subharmonics, various types of modulation, and deterministic chaos (Berge, P., Order within Chaos. Wiley, New York. 1986) (Glass, L., The rhythms of life: Princeton University Press). Press; 1988.) Titze et al. also reported that in asymmetric vocal fold vibration, subharmonics appear over two to three periods when the period and amplitude alternate between two states (Titze RT. Fluctuations and perturbations in vocal output. 1994).Tokuda et al. (Tokuda. Non-linear dynamics in mammalian voice production. 2018) conducted a simulation using a two-mass model of the vocal folds and reported that in periodic vocal fold vibrations known as limit cycles in chaos theory, the spectrum consists only of a fundamental frequency, given by the inverse of the vocal fold vibration period, and higher harmonics that are integer multiples of fo. However, even if the vocal fold vibration cycle is a limit cycle, if the left and right vocal folds form slightly different vibration patterns, the power spectrum will contain not only the fundamental frequency of vocal fold vibration but also subharmonics at half or one-third of the fundamental frequency. They also reported that this state is often observed during the transition from one type of vocalization to another.
[0168] This subharmonic has a periodicity similar to the harmonic structure linked to vocal cord vibration, so conventional methods have not yet been able to completely resolve the subharmonic error of mistaking subharmonics for fo (see Figure 16).
[0169] FIG. 16 is a diagram showing the spectral waveform of a sound containing subharmonics.
[0170] In the conventional method, there was a problem of subharmonics error, where the subharmonic region was mistakenly detected as fo.
[0171] Some studies have attempted to estimate fo using speech containing subharmonics in order to address subharmonic errors, but this method is limited to estimating fo using sustained vowel samples that also contain subharmonics (SWIPE's paper: Camacho and Harris. A sawtooth waveform inspired pitch estimator for speech and music. J. Acoust. Soc. 2008). To improve ecological validity, evaluation using sentence reading, not just sustained vowels, is necessary (Maryn Y. Toward improved ecological validity in the acoustic measurement of overall voice quality. J Voice. 2010;24:540-555.) (Zraick RI. The effect of speaking task on perceptual judgment of the severity of dysphonic voice. J Voice. 2005).
[0172] The DSH measurement implemented in the Multidimensional Voice Program (KayPENTAX, USA) acoustic analysis package estimates the temporal dominance of subharmonics, but its accuracy depends on the detection of fo (Deliyski DD. Acoustic model and evaluation of pathological voice production. 1993). The Diplophonia Diagram is also a measure of the combined quality of simple and complex vibrators, but its clinical relevance is limited due to its computational complexity (Aichinger P. Towards Objective Voice Assessment: The Diplophonia Diagram. Journal of Voice. 2017). Furthermore, a validation experiment on the two-stage cepstrum analysis method proposed by Awan et al. (Awan, 2staged) showed limitations in quantifying subharmonics, highlighting the importance of fo estimation in hoarse voices (Kitayama, 2staged). Based on the above, it was necessary to consider the possibility of subharmonics and develop a robust fo estimation algorithm that uses the temporal changes in fo of the speech waveform produced when reading aloud.
[0173] [Development of a new fo estimation algorithm] Therefore, as an accurate method for estimating fo in hoarse speech, we developed an algorithm based on the spectral method called "SFEEDS (Spectral-based fo Estimator Emphasized by Domination and Sequence)" equipped with two functions named "Dominant Spectrum test" and "Sequential Spectrum Test", as a script in the free software Praat (Paul Boersma and David Weenink (Boersma P. Praat, a system for doing phonetics by computer. Glot. Int. 2001;5:341-345.); Institute of Phonetic Sciences, University of Amsterdam, the Netherlands https: / / www.praat.org / ).
[0174] The first objective of this research is to develop a new fo estimation algorithm called "SFEEDS" that enables good fo estimation not only for normal speech but also for speech containing subharmonics. The second objective is to define the ground truth of fo even for speech samples containing subharmonics, which are difficult to define using human pitch perception or laryngographs.
[0175] The third objective is to compare SFEEDS with conventional fo estimation methods using the ground-truth of the defined fo, and to investigate their validity and effectiveness.
[0176] The results of this research will provide new options for fo estimation, which is the core of speech processing, and will contribute to providing basic technologies toward elucidating acoustic characteristics, including hoarseness, which have not yet been fully understood.
[0177] 3.METHODS (proposed method) [Dataset] A total of 455 recordings were reused from a dataset used in another study (Hosokawa K, von Latoszek B, Ferrer-Riesgo CA, et al. Acoustic Breathiness Index for the Japanese-Speaking Population: Validation Study and Exploration of Affecting Factors. Journal of Speech Language and Hearing Research. 2019;62:2617-2631).
[0178] The imported data consisted of 289 audio recordings from participants with various organic and non-organic voice disorders with varying degrees of dysphonia, 55 audio recordings from participants without voice complaints, and 111 audio recordings taken more than three months after treatment.
[0179] Figure 17 is a table summarizing the diagnoses of all 344 participants. Figure 17 shows the diagnosis (diagnosis item), number of patients (participant number), and interventions (intervention details).
[0180] Participants were asked to sustain the vowel / a: / for at least three seconds and read the Japanese translation of the text "The North Wind and the Sun" at a comfortable pitch, loudness, and tempo.
[0181] The preparation procedures for the CS and SV samples were the same as those required to calculate the Japanese AVQI (Hosokawa, K., Barsties von Latoszek, B., Iwahashi, T., Iwahashi, M., Iwaki, S., Kato, C., . . Maryn, Y. (2019). The Acoustic Voice Quality Index version 03.01 for the Japanese-speaking population. Journal of Voice, 33(1), 125.e1-125.e12.). For the CS sample, 30 syllables were prepared, from the beginning of the first sentence to the eighth syllable of the second sentence ( / aruhi kitakaze to taiyo: ga chikara kurabe wo shimashita tabibito no gaito: wo / ). For the SV sample, a 3-second period of medial vowels was extracted, excluding the initial and final portions, except for patients who were unable to sustain vowels for more than 3 seconds. The CS and SV samples were concatenated and used for analysis. In this analysis, 454 concatenated samples were used, excluding one sample that did not contain any voiced sounds.
[0182] The data includes information on gender, age, and diagnosis, as well as auditory and perceptual assessments of G, R, and B scores for SV and CS by three raters whose intra- and inter-rater reliability was confirmed in a previous study (Hosokawa et al. ABI 2020).
[0183] The average of Gcs and Gsv measured by the three evaluators was defined as Gtotal, and Rtotal and Btotal were calculated in the same way. Hoarseness, roughness, and breathiness were defined as Gtotal, Rtoral, and Btotal being 0.5 or higher, respectively. The breakdown of the degree of hoarseness is shown in Figure 18.
[0184] FIG. 18 is a diagram showing the distribution of the levels of voice quality evaluation in the psychoacoustic evaluation.
[0185] This study was approved by the Institutional Review Boards of Osaka University Hospital, Osaka Police Hospital, and Kuma Hospital (numbers: 15497, 568, and 20120614-1).
[0186] Preparation of voice samples All samples were recorded in an acoustically treated room using a head-mounted microphone SE50 (Samson Technologies Corp.) and digitized at a sampling rate of 44.1 kHz with 16-bit resolution. A linear PCM recorder, H4n (Zoom Corporation), was used. All samples were verified to meet the commonly required signal-to-noise ratio (>30 dB) (Deliyski DD, Shaw HS, Evans MK. Adverse effects of environmental noise on acoustic voice quality measurements. Journal of Voice. 2005;19:15-28.) (Deliyski DD, Shaw HS, Evans MK. Regression tree approach to studying factors influencing acoustic voice analysis. Folia Phoniatrica Et Logopaedica. 2006;58:274-288.).
[0187] [SFEEDS Contents] SFEEDS, developed as an accurate fo estimation method for hoarse voices, is an algorithm equipped with two main functions named "Dominant Spectrum test" and "Sequential Spectrum Test." The outline and details of the test are given below.
[0188] Dominant Spectrum Test: We devised a dominant spectrum test as a method for detecting the dominant harmonic structure within the harmonic structure contained in a speech waveform. The fo in laryngeal primary speech is relatively easy to estimate because it "forms the peak with the highest spectral intensity and lowest frequency components." However, this premise is lost when frequency components are enhanced by formants, making it difficult to estimate fo through simple spectral analysis alone. Therefore, it is necessary to search for the dominant harmonic structure within a short-time frame analysis. First, we search for the peak with the greatest spectral intensity (peak of 50-400 Hz) within the 50-400 Hz range of the spectral waveform (see Figure 19a).
[0189] FIG. 19a is a diagram showing the process of searching for fo (fundamental frequency) candidates in the dominant spectrum test.
[0190] In Figure 19a, the spectral peaks are extracted between 50-400 Hz.
[0191] Next, the low frequency region below this peak is further divided, and spectral peaks within each range are detected (see FIG. 19b).
[0192] FIG. 19b is a diagram showing an example of dividing the low frequency region and extracting spectral peaks within each range.
[0193] Based on the assumption that subharmonics (the gray area in Figure 19b) and environmental / turbulent noise are sufficiently small compared to the frequency components of fo (black) linked to vocal cord vibration, the lowest-frequency spectral peak with sufficient spectral intensity compared to the "peak of 50-400Hz" is extracted as a candidate for fo in the spectrum of a short-time frame (see Figure 19c).
[0194] FIG. 19c is a diagram in which the spectral peak of the lowest frequency having a spectral intensity equal to or greater than a certain level is selected as a candidate for fo (fundamental frequency) in comparison with the spectral peaks obtained in FIG. 19a.
[0195] However, this dominant harmonic structure can sometimes result in false positives of subharmonics or momentary increases in chaotic noise. Therefore, the Sequential Spectrum Test described below can be used to correct false positives.
[0196] Sequential Spectrum Test: The Sequential Spectrum Test is an algorithm that focuses on the transition of fo along the time series of speech samples, and compensates for the false detection of fo in the Dominant Spectrum Test described above, and reduces the error in fo estimation within the same phrase. This algorithm focuses on the gradual transition of fo when reading aloud a sentence at rest, and maintains the consistency of the time series change of fo.
[0197] If there is a frequency peak in the next temporally adjacent frame that has a similar spectral intensity and frequency to the fo determined in a particular frame, it is preferentially selected as the fo for the next frame (see Figure 20a).
[0198] FIG. 20a is a diagram showing the process of selecting candidates for fo (fundamental frequency) in the sequential spectrum test.
[0199] As shown in FIG. 20a, if there is a frequency peak in the subsequent (N+1)th frame that has similar spectral intensity and frequency to the fo determined in the Nth frame, it is selected as fo for the (N+1)th frame.
[0200] On the other hand, where the temporal continuity of fo is interrupted due to the onset and offset of a sentence, a sudden pitch jump, or a change in vibration mode, the continuity in the Sequential Spectrum Test is reset, and the fo estimate extracted in the Dominant Spectrum Test is selected (see Figure 20b).
[0201] FIG. 20b is a diagram showing an example in which, when there is no frequency peak with a similar spectral intensity and frequency, a candidate fo (fundamental frequency) selected by the dominant spectrum test is used.
[0202] The specific details of the script are as follows:
[0203] [Details of SFEEDS] Whole sample analysis 0: Adjust the average intensity of all audio samples to 70dB, independent of the recording level of the samples. (1) Calculate the maximum spectral intensity of the entire sample = sample_all_dB (2) Frame division (default: frame length 0.1 s, time step 0.0033 s, Gaussian window) The following is an analysis of each divided frame. I: Calculate the maximum spectral intensity within the frame = frame_all_dB Dominant Spectrum Test II: Extract the strongest frequency spectrum between 50-400Hz and set it as "h_defo_Hz" III: A division search is performed in the frequency band lower than "h_defo_Hz" to extract several spectral peaks. Among the spectral peaks with a spectral intensity of -10 dB or more in "h_defo_Hz", the lowest frequency one is set as "f0(h1_Hz_'i')" in the provisional frame. Sequential Spectrum Test IV: Compared to fo in the "previous" frame that is consecutive in time, if the frequency range is similar (within ±10% Hz), the spectral intensity is similar (within ±10 dB), and the spectral intensity is greater than (frame_all_dB-20 dB), then overwrite "f0(h1_Hz_'i')" in the interim frame. V: Consider the possibility that "f0(h1_Hz_'i')" in the frame may be misdetected as 2*fo or 1.5*fo, and correct it if this is the case. (Even when the Sequential Spectrum Test is applied to the same phrase, a search is always performed to check for subharmonic errors.) VI: If the spectral intensity corresponding to "fo(h1_Hz_'i')" in the provisional frame is less than (sample_all_dB-30dB), it is considered a silent frame. VII: Since fo less than 50Hz is not assumed, if "f0(h1_Hz_'i')" in the frame is less than 50Hz, it is considered a silent frame. (3) Repeat steps I-VII until the final frame.
[0204] [ground truth of fo] In this study, we focused on the narrowband spectrogram of speech to define the ground truth of fo.
[0205] FIG. 21 is a diagram showing an example of a spectrogram of a concatenated sample of CS and SV.
[0206] The speech samples consisted of CS (continuous speech) and SV (sustained vowels).
[0207] This spectrogram shows signals with strong spectral intensity in the subharmonic frequency band at the beginning and end of phrases and in some sustained vowels during reading aloud. When these subharmonic signals are included, it is easy to distinguish between fo and subharmonics on the spectrogram. However, many fo estimation methods suffer from subharmonic errors, such as the false positive detection of subharmonics as fo (Camacho and Harris. A sawtooth waveform inspired pitch estimator for speech and music. J. Acoust. Soc. 2008) (A Comparative Analysis of Pitch Detection Methods Under the Influence of Different Noise Conditions. J. Voice. 2015), as shown in Figure 22.
[0208] FIG. 22 is a diagram showing an example of subharmonics errors.
[0209] In Figure 22, the area surrounded by the dashed line is where subharmonics are erroneously detected as fo (calculated using the Filtered AC method).
[0210] The errors caused by the estimated fo compared to the actual fo include "too low" errors (mainly subharmonic) as shown in Figure 22, as well as "too high" errors (Cheveigne, Kawahara, YIN, a fundamental frequency estimator for speech and music. J. Acoust. Soc. 2002). Therefore, using ImageJ software (National Institutes of Health, Bethesda, MD, USA), a free software for image analysis, we performed binarization of the image to visualize only the harmonic components, and extracted the frequency region where fo exists on the spectrogram as the "ground truth of fo" (see Figure 23).
[0211] FIG. 23 is a diagram showing a process of extracting the ground truth of fo (fundamental frequency) on a spectrogram.
[0212] [Accuracy validation of algorithms] The spectrogram depicting the fo calculated using the fo estimation algorithm was compared with the spectrogram consisting of the ground truth of fo, and the matching parts between the ground truth of fo and the estimated fo were extracted on the spectrogram using ImageJ (see Figure 24).
[0213] FIG. 24 is a diagram showing an example in which an estimated fo (fundamental frequency) and the ground truth of fo (fundamental frequency) are compared to extract a portion where they overlap.
[0214] When comparing accuracy validation, it is necessary to exclude silent parts from the analysis and calculate the agreement rate of the fo estimation results in the voiced parts. Therefore, the following variables were calculated as shown in Figure 25.
[0215] %true-all: the allocation of estimated match time to the entire sample - The percentage of voiced sounds in the entire sample "%voice-all"
[0216] Finally, the fo estimation coincidence "% true" in the fo estimation algorithm is calculated using the following (Equation 1).
[0217] "%true" = "%true-all" / "%voice-all" x 100...(Formula 1)
[0218] FIG. 25 is a diagram showing an example of calculating the allocation of the fo (fundamental frequency) estimated matching time to the entire sample, "% true-all," and the proportion of voiced sounds to the entire sample, "% voice-all."
[0219] [Calculation method for each fo estimation method] Using the ground-truth fo defined by the above method, accuracy validation of the fo estimation accuracy was performed using 455 speech samples consisting of sentence readings and sustained vowels.
[0220] FIG. 26 shows a list of algorithms evaluated for comparison.
[0221] FIG. 26 shows the algorithms to be compared and the URLs (Uniform Resource Locators) to be referenced.
[0222] All calculations are performed using the default settings recommended by the algorithm distributor.
[0223] These algorithms were selected because they are widely used and popular, and because they have been reported to provide higher accuracy for fo estimation compared to other methods (Camacho and Harris. A sawtooth waveform inspired pitch estimator for speech and music. J. Acoust. Soc. 2008) (Ba H, Yang N. BaNa: a hybrid approach for noise-resilient pitch detection. IEEE Statistical Signal Processing Workshop. 2012). Peak picking and zero crossing methods were excluded from the analysis because they lack established scripts and have been reported to be less accurate for AC and CC (Titze IR, Liang H. Comparison of Fo extraction methods for high-precision voice perturbation measurements. J Speech Hear Res. 1993;36:1120-33). Additionally, the Cepstrum method does not have an established algorithm, and previous research has concluded that it has low accuracy in estimating fo (A Comparative Analysis of Pitch Detection Methods Under the Influence of Different Noise Conditions. J.voice. 2015), so it was excluded from the analysis.
[0224] The f(o) estimation results obtained by each f(o) estimation algorithm were analyzed using Praat or MATLAB (registered trademark). The f(o) estimation results were visualized on a spectrogram, and the degree of agreement with the ground truth f(o) (% true) was calculated using the method described above.
[0225] [Statistical analysis] First, the Shapiro-Wilk normality test showed that the agreement between the estimated fo calculated by each fo estimation algorithm and the ground truth violated the assumption of normality (p<0.001), necessitating nonparametric testing.
[0226] Wilcoxon's signed-rank test with Bonferroni's correction was used to compare the accuracy of fo estimation among the fo estimation algorithms.
[0227] A two-sided equivalence test, in which a difference in population means within ±5% was considered equivalent, was used to evaluate the robustness of the fo estimation algorithm in the presence or absence of hoarseness. (All statistical analyses were performed using the JMP version 16.0.0 software package (SAS Institute, Cary, NC). All results were considered statistically significant at p<0.05.)
[0228] 4.RESULT [Distribution of hoarseness and fo estimation accuracy] FIG. 27 is a diagram showing the results of plotting "% true" for each fo (fundamental frequency) estimation algorithm in the concatenated speech samples for all 454 samples.
[0229] The horizontal axis indicates the degree of Gtotal, Rtotal, and Btotal for each sample, and the vertical axis indicates the match rate (% true) with the ground truth of fo.
[0230] As shown by the smoothing spline curves and nonparametric density, for all fo estimation algorithms, the "%true" tended to decrease and the variance increased as the psychoacoustic evaluation showed higher hoarseness.
[0231] [Comparison of agreement rates of fo estimates across all samples] FIG. 28 is a diagram comparing the results of each fo (fundamental frequency) estimation algorithm for all 454 samples.
[0232] The accuracy of SFEEDS's fo estimation was significantly higher than that of all other algorithms, with raw CC having the lowest accuracy (***p<0.001).
[0233] The "% true" for each fo estimation algorithm was compared using Wilcoxon's signed-rank test with Bonferroni's correction (significance above box plots).
[0234] [Estimation accuracy depending on whether hoarseness is present or not] (a) No hoarseness group Figure 29 shows the results for speech samples evaluated as non-hoarse with Gtotal less than 0.5. SFEEDS's fo estimation accuracy was significantly higher than that of all other algorithms. In the non-hoarse group, BaNa had the lowest fo estimation accuracy (**p<0.01, ***p<0.001).
[0235] (b) Hoarseness group Figure 30 shows the results for speech samples that were evaluated as hoarse, with Gtotal, Rtotal, and Btotal all exceeding 0.5. Regardless of the type of hoarseness, SFEEDS's "% true" was significantly higher than all other algorithms. In the hoarseness group, raw CC had the lowest fo estimation accuracy.
[0236] (c) Equivalence of fo agreement rates between those with and without hoarseness Figure 31 shows the results of an investigation into whether the difference in "%true" was within 5% when comparing each type of hoarseness group with and without hoarseness. The results showed that for Gtotal and Btotal, the "%true" for SWIPE' and SFEEDS was equivalent within a 5% margin of error, regardless of whether hoarseness was present or absent. In contrast, for Rtotal, only SFEEDS proved to be equivalent within a 5% margin of error, regardless of whether hoarseness was present or absent.
[0237] 5. Discussion The results of the distribution chart of the degree of hoarseness and "%true" showed that for all fo estimation algorithms, the more severe the hoarseness in the psychoacoustic evaluation, the lower the "%true" tended to be. Furthermore, the degree of roughness, in terms of the characteristics of hoarseness, tended to make fo estimation particularly difficult. Furthermore, the more severe the hoarseness, the greater the variability in "%true" tended to be.
[0238] When comparing the agreement rates of fo estimates across all samples, SFEEDS had a significantly higher "%true" than all other algorithms, while raw CC had the lowest "%true" result.
[0239] Filtered AC applies a low-pass filter during analysis, whereas raw CC does not apply a low-pass filter and is recommended for voice analysis. However, it has been suggested that applying a low-pass filter may improve the accuracy of fo estimation, especially for voices containing low-frequency noise such as subharmonics.
[0240] Even in the non-hoarse group, SFEEDS's fo estimation accuracy was significantly higher than that of all other algorithms. However, in the non-hoarse group, BaNa's fo estimation accuracy was the lowest. BaNa (Ba H, Yang N. BaNa: a hybrid approach for noise-resilient pitch detection. IEEE Statistical Signal Processing Workshop. 2012) was developed to improve fo detection accuracy in noisy environments and is a script that enables accurate fo estimation even in speech samples with low SNR. However, the evaluation corpora consisted of speech samples from fewer than 10 people (Ba H, Yang N. BaNa: a hybrid approach for noise-resilient pitch detection. IEEE Statistical Signal Processing Workshop. 2012) (A Comparative Analysis of Pitch Detection Methods Under the Influence of Different Noise Conditions. J.voice. 2015), suggesting that its robustness to various pitches and voice quality variations other than Rscore and Bscore (such as hypotonic phonation, which changes the A score) may be low.
[0241] In the hoarseness group, regardless of the type of hoarseness, SFEEDS's "%true" was significantly higher than all other algorithms, and raw CC had the lowest fo estimation accuracy.
[0242] Furthermore, we investigated the robustness of each fo estimation algorithm's "%true" in the presence or absence of hoarseness. For Gtotal and Btotal, SWIPE' and SFEEDS proved to be highly robust. For Rtotal, on the other hand, SFEEDS proved to maintain high robustness regardless of the presence or absence of hoarseness. Therefore, SFEEDS proved to be superior as the algorithm most capable of controlling subharmonic error, which is a problem in fo estimation for rough voice quality.
[0243] (Other embodiments) Although the speech analysis method and the like according to the embodiments have been described above, the present disclosure is not limited to the above-described embodiments.
[0244] For example, the steps in the speech analysis method may be executed by a computer (computer system). The present disclosure can be realized as a program for causing a computer to execute the steps included in the method. Furthermore, the present disclosure can be realized as a non-transitory computer-readable recording medium, such as a CD-ROM, on which the program is recorded.
[0245] For example, when the present disclosure is realized as a program (software), each step is performed by running the program using hardware resources such as a computer's CPU, memory, input / output circuits, etc. In other words, each step is performed by the CPU acquiring data from memory or input / output circuits, etc., performing calculations on the data, and outputting the calculation results to memory or input / output circuits, etc.
[0246] Furthermore, each component included in the voice analysis device of the above embodiment may be realized as a dedicated or general-purpose circuit. Furthermore, each component included in the voice analysis device of the above embodiment may be realized as an LSI (Large Scale Integration), which is an integrated circuit (IC). Furthermore, the integrated circuit is not limited to an LSI, and may be realized as a dedicated circuit or a general-purpose processor. A programmable FPGA (Field Programmable Gate Array) or a reconfigurable processor in which the connections and settings of circuit cells within an LSI can be reconfigured may also be used.
[0247] Furthermore, if an integrated circuit technology that can replace LSI emerges due to advances in semiconductor technology or other derivative technologies, it is natural that each component included in the device may be integrated using that technology.
[0248] In addition, this disclosure also includes forms obtained by making various modifications to the embodiments that a person skilled in the art would think of, and forms realized by arbitrarily combining the components and functions in each embodiment within the scope of the present disclosure. [Explanation of symbols]
[0249] 1. Voice analysis system 10. Voice analysis device 20 Waveform acquisition section 30 Derivation part 40 Estimation part 50 Correction unit 90 Sound collection device f1a, f1b, f2a, f2b, f2c frequencies p, pk, pk0 frequency peaks pkR verification point pm Maximum peak Rf specified frequency range Rs Spectral intensity range
Claims
1. a waveform acquisition step of acquiring frequency spectrum waveforms of a plurality of frames, each frame being a predetermined period, based on a signal including speech; a first derivation step of deriving a fundamental frequency in a first frame based on a frequency spectrum waveform of the first frame among the plurality of frames; estimating a fundamental frequency in a second frame based on a frequency spectrum waveform of the second frame among the plurality of frames; Including, In the step of estimating the fundamental frequency, it is determined whether a part of the frequency spectrum waveform of the second frame is present in a predetermined frequency range including the fundamental frequency of the first frame, and whether a verification point in the part of the frequency spectrum waveform is present in a predetermined spectral intensity range including the spectral intensity of the fundamental frequency, and the fundamental frequency of the second frame is estimated based on the determination. Voice analysis methods.
2. The verification point is a point where the spectral intensity is maximum among the frequency spectrum waveforms present in the predetermined frequency range. The speech analysis method according to claim 1 .
3. further comprising a second derivation step of deriving a tentative fundamental frequency in the second frame based on the frequency spectrum waveform of the second frame; In the step of estimating the fundamental frequency, If a part of the frequency spectrum waveform of the second frame is within the predetermined frequency range and the spectral intensity of the verification point is within the predetermined spectral intensity range, the frequency of the verification point is estimated to be a fundamental frequency in the second frame; When a part of the frequency spectrum waveform of the second frame does not exist within the predetermined frequency range, or when the spectral intensity of the verification point does not exist within the predetermined spectral intensity range, the provisional fundamental frequency derived in the second derivation step is estimated to be the fundamental frequency in the second frame. The speech analysis method according to claim 2 .
4. moreover, an (N+1)th derivation step of deriving a tentative fundamental frequency in the (N+1)th frame (N is an integer equal to or greater than 2) based on a frequency spectrum waveform of the (N+1)th frame among the plurality of frames; estimating a fundamental frequency in the (N+1)th frame based on a frequency spectrum waveform of the (N+1)th frame; Including, In the step of estimating a fundamental frequency in the (N+1)th frame, If a part of the frequency spectrum waveform of the (N+1)th frame exists in a predetermined frequency range including the fundamental frequency of the Nth frame, and the spectral intensity of a verification point in the frequency spectrum waveform where the spectral intensity is maximum in the predetermined frequency range exists in a predetermined spectral intensity range, the frequency of the verification point is estimated to be the fundamental frequency of the (N+1)th frame; If a part of the frequency spectrum waveform of the (N+1) frame does not exist within the predetermined frequency range, or if the spectral intensity of the verification point does not exist within the predetermined spectral intensity range, the provisional fundamental frequency derived in the (N+1) derivation step is estimated to be the fundamental frequency in the (N+1) frame. The speech analysis method according to claim 3 .
5. The first deriving step includes, in the first frame: deriving the maximum peak of the spectral intensity of the frequency spectrum waveform; extracting one or more frequency peaks having a maximum value greater than a predetermined spectral intensity from a frequency region located on the lower frequency side than the frequency at which the spectral intensity has a maximum peak; extracting a frequency peak located on the lowest frequency side from the one or more frequency peaks, and setting a frequency corresponding to the frequency peak as a fundamental frequency in the first frame; The speech analysis method of claim 2 , comprising:
6. The second deriving step includes, in the second frame: deriving the maximum peak of the spectral intensity of the frequency spectrum waveform; extracting one or more frequency peaks having a maximum value greater than a predetermined spectral intensity from a frequency region located on the lower frequency side than the frequency at which the spectral intensity has a maximum peak; extracting a frequency peak located at the lowest frequency side from the one or more frequency peaks, and setting a frequency corresponding to the frequency peak as a tentative fundamental frequency in the second frame; The speech analysis method of claim 3, comprising:
7. The (N+1) deriving step includes, in the (N+1) frame, deriving the maximum peak of the spectral intensity of the frequency spectrum waveform; extracting one or more frequency peaks having a maximum value greater than a predetermined spectral intensity from a frequency region located on the lower frequency side than the frequency at which the spectral intensity has a maximum peak; extracting a frequency peak located at the lowest frequency side from the one or more frequency peaks, and setting a frequency corresponding to the frequency peak as a tentative fundamental frequency in the Nth frame; The speech analysis method of claim 4, comprising:
8. the predetermined frequency range is a certain range of frequencies based on the fundamental frequency, The predetermined spectral intensity range is a certain range of spectral intensities based on the spectral intensity of the fundamental frequency. The speech analysis method according to any one of claims 1 to 7.
9. In the step of deriving the maximum peak of the spectral intensity, a speech spectrum waveform is generated by removing data different from the speech from the frequency spectrum waveform, and the maximum peak of the spectral intensity is derived based on the speech spectrum waveform. The speech analysis method according to any one of claims 5 to 7.
10. In the waveform acquisition step, the frequency spectrum waveform is adjusted so that an average of a plurality of maximum values of spectral intensities included in the frequency spectrum waveform falls within a predetermined range, and the adjusted frequency spectrum waveform is acquired. The speech analysis method according to any one of claims 1 to 7.
11. The method further includes a step of correcting the fundamental frequency when the fundamental frequency estimated in the step of estimating the fundamental frequency is a frequency of a harmonic, In the step of correcting the fundamental frequency, the frequency of the first harmonic corresponding to the harmonic is set as the fundamental frequency after correction. The speech analysis method according to any one of claims 1 to 7.
12. The method further includes a step of invalidating the fundamental frequency estimated in the step of estimating the fundamental frequency when the spectral intensity of the fundamental frequency estimated in the step of estimating the fundamental frequency is smaller than a predetermined spectral intensity. The speech analysis method according to any one of claims 1 to 7.
13. The method further includes a step of invalidating the fundamental frequency estimated in the step of estimating the fundamental frequency when the frequency of the fundamental frequency estimated in the step of estimating the fundamental frequency is less than 50 Hz. The speech analysis method according to any one of claims 1 to 7.
14. a waveform acquisition unit that acquires frequency spectrum waveforms of a plurality of frames, each frame being a predetermined period, based on a signal including speech; a derivation unit that derives a fundamental frequency in a first frame based on a frequency spectrum waveform of the first frame among the plurality of frames; an estimation unit that estimates a fundamental frequency in a second frame based on a frequency spectrum waveform of the second frame among the plurality of frames; Equipped with The estimation unit determines whether a part of the frequency spectrum waveform of the second frame is present in a predetermined frequency range including the fundamental frequency of the first frame, and whether a verification point in the part of the frequency spectrum waveform is present in a predetermined spectral intensity range including the spectral intensity of the fundamental frequency, and estimates the fundamental frequency of the second frame based on the determination. Voice analysis device.