Recognition or synthesis of overtones uttered by a person
The method addresses the inefficiencies in speech recognition and synthesis by analyzing and utilizing harmonic components within human speech, enhancing recognition and improving the fidelity of speech synthesis.
Patent Information
- Application Number
- JP2023571852
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-05-18
- Filing Date
- 2022-05-13
- Publication Date
- 2025-06-05
- Estimated Expiration
- 2042-05-13
AI Technical Summary
Existing technologies for recognizing and synthesizing human speech often overlook the importance of harmonic components, which carry linguistically relevant information, leading to inefficiencies in speech recognition and synthesis.
A computer-executed method that identifies and analyzes vocal harmonics within an electronic waveform by detecting and selecting harmonic components, using the primary cap frequency to select key tones, and synthesizing speech based on these harmonic components.
This method enhances speech recognition by utilizing harmonic components and improves speech synthesis by generating speech chords that accurately represent human vocal harmonics, thereby improving the fidelity and accuracy of speech processing.
Smart Images

Figure 0007688858000011 
Figure 0007688858000012 
Figure 0007688858000013
Abstract
Description
Technical Field
[0001] The field of the present invention relates to the recognition or synthesis of utterances spoken by a human. In particular, a method executed by a computer for recognizing or synthesizing harmonic sounds spoken by a human is disclosed.
Background Art
[0002] Some examples of devices or methods for processing or synthesizing utterances are disclosed below. - U.S. Patent No. 5,406,635, entitled "Noise attenuation system," issued to Jarvinen on April 11, 1995. - U.S. Patent Application Publication No. 2004 / 0181411, entitled "Voicing index controls for CELP speech encoding," issued in the name of Gao on September 16, 2004. - U.S. Patent No. 6,925,435, entitled "Method and apparatus for improved noise reduction in a speech encoder," issued to Gao on August 2, 2005. - U.S. Patent No. 7,516,067, entitled "Method and apparatus using harmonic-model-based front end for robust speech recognition," issued to Seltzer et al. on April 7, 2009. - U.S. Patent No. 7,634,401, entitled "Speech recognition method for determining missing speech," issued to Fukuda on December 15, 2009. - U.S. Patent No. 8,566,088, entitled "System and method for automatic speech to text conversion," issued to Pinson et al. on October 22, 2013. - U.S. Patent No. 8,606,566, entitled "Speech enhancement through partial speech reconstruction," granted to Li et al. on December 10, 2013. - U.S. Patent No. 8,812,312, entitled "System, method and program for speech processing," granted to Fukuda et al. on August 19, 2014. - U.S. Patent No. 9,087,513, entitled "Noise reduction method, program product, and apparatus," granted to Ichikawa et al. on July 21, 2015. - U.S. Patent No. 9,190,072, entitled "Local peak weighted-minimum mean square error (LPW-MMSE) estimation for robust speech," granted to Ichikawa on November 15, 2015. - U.S. Patent No. 9,570,072, entitled "System and method for noise reduction in processing speech signals by targeting speech and disregarding noise," granted to Pinson on February 14, 2017. - Dieter Maurer, "Acoustic of the Vowel: Preliminaries," Peter Lang AG, Bern 2016. - Bruno H. Repp, "Categorical Perception: Issues, Methods, Findings," SPEECH AND LANGUAGE: Advances in Basic Research and Practice, Vol. 10 p. 243 (Academic Press 1984), https: / / doi.org / 10.1016 / B978-0-12-608610-2.50012-1. - U.S. Patent No. 9,147,393, entitled "Syllable based speech processing method," issued to Fridman-Mintz (the inventor of the present invention) on September 29, 2015. - U.S. Patent No. 9,460,707, entitled "Method and apparatus for electronically recognizing a series of words based on syllable-defining beats," issued to Fridman-Mintz (the inventor of the present invention) on October 4, 2016. - U.S. Patent No. 9,747,892, entitled "Method and apparatus for electronically synthesizing acoustic waveforms representing a series of words based on syllable-defining beats," issued to Fridman-Mintz (the inventor of the present invention) on August 29, 2017.
[0003] The last three patents listed above (each issued to Fridman-Mintz) are incorporated by reference as if fully set forth herein.
Summary of the Invention
[0004] A method executed by a computer is employed to identify one or more vocal harmonics (e.g., harmonic phones) represented within an electronic waveform over time derived from human speech. In one embodiment, from a temporal sequence of acoustic spectra obtained from the waveform, each of a plurality of harmonic acoustic spectra within that temporal sequence is analyzed, and within each of those harmonic acoustic spectra, two or more fundamental or harmonic components each having an intensity exceeding a detection threshold are identified. The identified components have frequencies separated by at least one integer multiple of a fundamental acoustic frequency associated with that acoustic spectrum. For at least some of the plurality of acoustic spectra, a primary cap frequency is identified, which is greater than 410 Hz and is also the highest harmonic frequency among the identified harmonic components. For each acoustic spectrum for which a primary cap frequency is identified, that identified primary cap frequency is used to select at least one voice sound as a keynote from among a set of voice sounds. The selected keynote corresponds to a subset of the vocal harmonics within a set of vocal harmonics.
[0005] In one embodiment, the acoustic spectrum of the temporal sequence may correspond to one of a sequence of temporal sample intervals of the waveform, and in other embodiments, the acoustic spectrum corresponds to one of a sequence of discrete temporal segments where the time-dependent acoustic spectrum of the waveform remains consistent with a single vocal harmonic. In some cases, vocal harmonics can be selected based on harmonic components present in a harmonic acoustic spectrum that includes one or more of a primary band (primary band), secondary band (secondary band), base band (base band), or reduced base band (reduced base band), each described later.
[0006] A method executed by a computer is adopted to analyze the voice uttered by a human and generate spectral data that can be used to identify harmonic voices and chords in the above method. For each voice chord, the waveform obtained from the utterance of that voice chord by one or more human subjects is spectrally analyzed. The spectral analysis includes, for each electronic waveform, the estimation of the fundamental acoustic frequency and the identification of two or more fundamental or harmonic components that each have an intensity exceeding a detection threshold and have an acoustic frequency that is the fundamental acoustic frequency or its harmonic. The primary cap frequency is identified for each voice chord and stored together with the acoustic frequencies of the identified fundamental or harmonic components. The focus frequency can be estimated for the keynote common to a subset of voice chords using the observed primary cap frequency (e.g., the average value or the median). In some cases, the stored spectral data can include data of one or more of the primary band, secondary band, base band, or reduced base band.
[0007] A method executed by a computer for synthesizing a temporal segment of an electronic waveform is adopted. By applying this waveform segment to an electroacoustic transducer, a voice chord is generated. The keynote corresponding to the selected voice chord and the data indicating the focus frequency of that keynote are used to determine the primary cap frequency. The primary cap frequency is (i) an integer multiple of the selected fundamental frequency, (ii) greater than 410 Hz, and (iii) closer to the focus frequency of the corresponding keynote than the focus frequencies of other voices. The harmonic components of the primary cap frequency are included in the waveform segment. The primary cap frequency is the highest frequency among the harmonic components included in the waveform segment. The waveform segment can further include components of one or more harmonic frequencies of the primary band, secondary band, base band, or reduced base band. This method can be repeated for each voice chord in the temporal sequence that constitutes human speech, together with a plurality of different harmonic segments or hybrid segments, inharmonic segments or silent segments, and transition segments between them.
[0008] The objectives and advantages related to the recognition or synthesis of human voice will become apparent by referring to the exemplary embodiments illustrated in the drawings and disclosed in the following description or the appended claims.
[0009] This summary is provided to introduce in a simplified manner some concepts that will be described in detail later. This summary is not intended to identify the main features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.
Brief Description of the Drawings
[0010]
Figure 1
Number
Figure 2
Figure 3A
Figure 3B
Figure 4A
Figure 4B
Figure 5
Figure 6A
Figure 6B
Figure 7
Figure 8
[0011] The examples or embodiments depicted in the drawings are merely shown schematically and not all features are shown in complete detail, nor are they shown to scale. To clarify, certain features or structures may be exaggerated or reduced relative to other features or structures, or may be completely omitted. The drawings should not be considered to be to scale unless explicitly stated to be so. The illustrated embodiments are merely examples and should not be construed as limiting the scope of the present disclosure or the inventive subject matter. The same reference numerals refer to similar elements throughout the different figures.
[0012] The following detailed description should be read with reference to the drawings. The detailed description explains the principles of the subject matter by way of example and not by way of limitation.
[0013] The methods disclosed herein rely in part on the observation that for vocal harmonics (e.g., harmonic sounds such as vowels, nasal vowels, nasals, and approximants), the harmonic spectral components (i.e., components having frequencies that are integer multiples of the fundamental acoustic frequency) carry the core of the linguistically relevant information that conveys those vocal harmonics. Accordingly, the disclosed methods include the recognition or synthesis of vocal harmonics that use the detection, identification, or generation of such harmonic components. Each voice corresponds to a basic harmonic element (e.g., as in the example list of FIG. 2), and each vocal harmonic can include multiple harmonic frequencies in the context of a given voice (e.g., as in the example lists of FIGS. 3A and 3B).
[0014] As part of operating a speech recognition system, a method executed by a computer is adopted to identify one or more vocal harmonics represented within an electronic waveform over time derived from a human voice utterance. A time sequence of acoustic spectra is derived from the waveform in an appropriate manner (described later), and a portion of those acoustic spectra is identified as a harmonic acoustic spectrum, i.e., a spectrum containing frequency components that are integer multiples of one or more fundamental acoustic frequencies. For some or all of the harmonic acoustic spectra within the time sequence, two or more basic or harmonic components are identified that each have an intensity exceeding a detection threshold and are delimited by integer multiples of the corresponding fundamental acoustic frequency of the harmonic acoustic spectrum. In some examples, for various reasons, the fundamental frequency component may be lost or hidden. This reason may be, for example, interference with other nearby high-intensity frequency components, attenuation by obstacles in the acoustic path, weakness in the generation of the low-frequency range by a small speaker, etc. Even when the fundamental wave component does not exist in the harmonic spectrum, the harmonic spectrum can be characterized by the corresponding fundamental acoustic frequency, and its integer multiples separate the harmonic components of the harmonic spectrum.
[0015] The fundamental and harmonic components typically appear as peak-like spectral features that are higher than the background level (e.g., as in FIG. 1). The corresponding frequencies of the components can be defined in an appropriate manner, such as the frequency of maximum intensity or the centroid of the component. When the width of the spectral feature is not zero, there may be some uncertainty in assigning the frequencies corresponding to those features, and there is usually some margin when determining whether the spectral feature is an integer multiple of the fundamental frequency. The intensity of the component can be characterized in an appropriate manner, such as maximum intensity, peak intensity, or integrated intensity. Intensity is usually expressed in relative terms (e.g., decibels) with respect to a background noise level, an auditory threshold level, or another reference level. Regardless of how the relative intensity is defined, an appropriate detection threshold can be selected to determine whether a given spectral feature is "detected" and "identified" as a fundamental or harmonic component of a given harmonic spectrum. In some examples, a certain decibel level may be specified to exceed a selected reference level, exceed the background noise level, or be otherwise designated. In some examples, the detection threshold may be frequency-dependent (e.g., as in FIG. 1) and may thus differ for the fundamental and harmonic components of the spectrum.
[0016] For at least some of the harmonic spectra, a primary cap frequency, which is the highest harmonic frequency among the identified harmonic components, is identified. However, the highest harmonic frequency is also greater than 410 Hz. For each harmonic spectrum for which such a primary cap frequency is identified, the primary cap frequency is used to select at least one sound from a set of fundamental sounds (such as the table in FIG. 2). This selection is possible because it has been observed that the primary cap frequency most closely correlates with the particular vocal harmony emitted by recording and analyzing a large number of vocal harmonies emitted by humans. The data collected to create a table of fundamental sounds (such as FIG. 2) can include, for each sound, a characteristic primary cap frequency (i.e., the focus frequency) and the primary symbol label with which it correlates. Other sound data sets can also be employed. The generation of such sound data sets will be described later. Other speech recognition systems often rely on the fundamental frequency and, in some cases, two or three formants (temporally aligned intensity peaks with frequency values independent of the fundamental frequency), while substantially ignoring all harmonic components present in the acoustic spectrum, under the assumption that vowel perception and recognition are not based on categorical perception or feature analysis. The method disclosed herein utilizes these harmonic spectral components, which have not been utilized heretofore, to identify the corresponding vocal harmony and one or more of its constituent sounds, demonstrating that vowel identification can, in some cases, be based on prototype categorical perception.
[0017] For a given vocal harmony sounded at different pitches (i.e., different fundamental frequencies), the primary cap frequency may vary, so a definitive identification (e.g., an unambiguous selection from the data table of FIG. 2) cannot always be achieved. The tables of FIGS. 3A and 3B show examples of instantiated harmony data. Each frequency of the instantiated primary cap falls between the focus frequencies of an adjacent pair of vocal tones (e.g., Table 2), or else is higher than the highest focus frequency. When it is between a pair of vocal tones (labeled "specified" and "adjacent" harmonies in the tables of FIGS. 3A and 3B), the instantiated vocal harmony can be specified as corresponding to one of these two vocal tones (the "specified" harmony in this example). When it exceeds the highest focus frequency (e.g., exceeds 4160 Hz in the table of FIG. 2), the instantiated harmony can be specified as corresponding to TIFF0007688858000009.tif9170 the vocal tone. In some examples, the instantiated vocal harmony is specified as corresponding to the vocal tone with the focus frequency closer to the instantiated primary cap frequency among an adjacent pair of vocal tones. The confidence columns in the tables of FIGS. 3A and 3B indicate the relative closeness of the instantiated primary cap frequency to the corresponding adjacent pair of focus frequencies. Other suitable tests or algorithms can be employed to determine the different possible vocal harmonies that may match a particular instantiated harmonic spectrum. In some examples, an artificial intelligence (Al) system or neural network can be trained to make the selection. In some examples, a so-called "fuzzy logic" algorithm can be employed. In some examples, to distinguish different vocal tones having focus frequencies that match the observed primary cap frequency, secondary cap frequency, or base cap frequency (described further below), the identification of additional fundamental or harmonic spectral components of the observed harmonic spectrum (e.g., the primary band, secondary band, base band described below) can be employed.
[0018] One exemplary method may include deriving a time sequence of acoustic spectra from an electronic waveform over time. In other examples, the time sequence of acoustic spectra may already have been derived from the electronic waveform before the method is executed. In either case, the time sequence of acoustic spectra can be derived from the waveform in a suitable manner, such as by processing the electronic waveform itself using, for example, an electronic spectrum analyzer or by processing the numerical representation of the waveform using Fourier transform techniques.
[0019] In some examples, each of the acoustic spectra corresponds to one of a sequence of time sample intervals of the waveform. Such sample intervals can be of equal duration, but need not necessarily be so. At least some of the acoustic spectra in such a time sequence can be classified into only one of harmony, dissonance, hybrid, or silence. For at least some of the time sample intervals classified as harmony, the methods described above or below can be employed to identify a set of vocal harmonies (including respective characteristic combinations of sounds) corresponding to that time sample interval of the electronic waveform from among the set of vocal harmonies.
[0020] In some examples, each of the acoustic spectra corresponds to one of a series of different time segments during which its time-dependent acoustic spectrum remains consistent with a single vocal harmony. The determination that the acoustic spectrum "remains consistent with a single vocal harmony" can be made even if the harmony has not already been identified, provided that it is observed that no transition to an acoustic spectrum indicating a different harmony has occurred. At least some of the time segments can be classified into only one of harmony, dissonance, hybrid, or silence based on their acoustic spectra. For at least some of the time segments classified as harmony, the methods described above or below can be employed to identify a vocal harmony from among a set of vocal harmonies corresponding to that time segment of the electronic waveform.
[0021] In addition to the primary pitch frequency, identifying additional fundamental frequencies or harmonic frequencies can further facilitate the identification of the vocal harmony corresponding to the harmonic spectrum within the time sequence. In some cases, the identification of additional fundamental frequencies or harmonic components helps to distinguish vocal harmonies with similar primary pitch frequencies. The additional fundamental components or harmonic components may also form one or more of the primary band, the secondary band, or the fundamental band.
[0022] In some examples, the primary band of the harmonic components can be identified in at least some of the harmonic spectra of the time sequence. The primary band can include the primary pitch frequency and the harmonic components at the largest consecutive multiples of 1, 2, 3, or more of the fundamental acoustic frequency that are (i) less than the primary pitch frequency, (ii) greater than 410 Hz, and (iii) greater than the smallest integer multiple of the fundamental acoustic frequency greater than 410 Hz. The stored data of the vocal harmony set can include, in addition to the primary pitch frequency, the frequencies of the other harmonic components of the primary band. Based on the comparison between the data of the primary band and the observed frequencies of the specific harmonic spectrum obtained from the electronic waveform, the vocal harmony of the set can be selected as corresponding to the corresponding time portion of the harmonic spectrum and the waveform. FIGS. 4A / 4B, 5, 6A / 6B, 7, and 8 show some examples of recorded and synthesized waveforms and potentially distinguishable harmonic spectra based on the primary band of the harmonic components.
[0023] In some examples, in at least some of the harmonic spectra of a time sequence, the secondary band of a harmonic component can be identified. The secondary band is greater than the smallest integer multiple of the fundamental acoustic frequency that exceeds 410 Hz and can include harmonic components at one or more harmonic acoustic frequencies that are separated from the primary cap frequency by at least one intervening multiple of the fundamental acoustic frequency at which the acoustic spectrum lacks harmonic components. In other words, the secondary band is below a “harmonic gap” between the lowest frequency component of the primary band and the highest frequency component of the secondary band or one or more “missing overtones”. The frequency of the highest frequency component of the secondary band may be referred to as the secondary cap frequency.
[0024] For a set of vocal harmonics, the stored data can include, in addition to the primary cap frequency (or corresponding primary symbol), the secondary cap frequency (or corresponding secondary symbol) and the frequency (if any) of one or more other harmonic components of the secondary band. Based on a comparison of the secondary band data with the observed frequencies of a particular harmonic spectrum obtained from an electronic waveform, the set of vocal harmonics can be selected as corresponding to the corresponding temporal portion of that harmonic spectrum and waveform. For example, by observing the secondary band, it is possible to distinguish (i) a first vocal harmonic having a secondary band separated from the primary band by only one missing overtone, and (ii) a second vocal harmonic having a secondary band separated from the primary band by two or more missing overtones. In some examples, the comparison of the secondary band data with the observed components can be used in conjunction with the comparison of the primary band data with the observed components. In other examples, the secondary band data and components can be used without using the primary band data and components. FIGS. 4A / 4B, 5, 6A / 6B, 7, and 8 show some examples of recorded and synthesized waveforms and harmonic spectra that may be distinguished based on the secondary band of harmonic components.
[0025] In some examples, in at least some of the harmonic spectra of a time sequence, the baseband of the harmonic components can be identified. The baseband can include harmonic components at one or more fundamental or harmonic acoustic frequencies less than 410 Hz, and can also include harmonic components at the lowest harmonic acoustic frequency above 410 Hz (except when that harmonic frequency is the primary cap frequency). The frequency of the highest frequency component of the baseband can be referred to as the base cap frequency. In an example where the primary cap frequency is also the only harmonic frequency above 410 Hz, the harmonic spectrum includes only the primary cap component and the baseband component, and there are no other primary band components or secondary band components. The data of the stored harmonic set can include, in addition to the primary cap frequency (or the corresponding fundamental tone), the fundamental cap frequency, and the frequencies of one or more other harmonic components of the fundamental band. Based on the comparison between the fundamental band data and the observed frequencies of a particular harmonic spectrum obtained from an electronic waveform, the set of vocal harmonics can be selected as corresponding to the corresponding temporal portion of that harmonic spectrum and waveform.
[0026] In some examples, the comparison of the baseband data with the observed components can be used in combination with the comparison of the primary band data with the observed components, and in other examples, the baseband data and components can be used in combination with the comparison of the secondary band data with the observed components, and in other examples, the baseband data and components can be used in combination with the comparison of both the primary band data and the secondary band data with the observed components, and in other examples, the baseband data and components can be used without using either the primary band data or the secondary band data. FIGS. 4A / 4B, 5, 6A / 6B, 7, and 8 show some examples of recorded and synthesized waveforms and harmonic spectra that may be distinguishable based on the baseband of the harmonic components.
[0027] In some examples, the harmonic acoustic spectrum of a temporal sequence may include only baseband components having an upper frequency limit of less than 410 Hz. Such a harmonic acoustic spectrum is referred to as a reduced baseband and can correspond to a particular harmonic acoustic schema (e.g., nasal, nasalized vowel, or approximant) or hybrid acoustic schema (e.g., voiced fricative). The stored data for these acoustic schemas can include the frequencies of the harmonic components of the reduced baseband. Based on a comparison of the reduced baseband data with the observed frequencies of a particular harmonic spectrum obtained from an electronic waveform, a set of harmonic or hybrid acoustic schemas can be selected as corresponding to the corresponding temporal portions of that harmonic spectrum and waveform. The presence or absence of high-frequency inharmonic frequency components can also be used to distinguish (i) reduced baseband harmonic schemas (e.g., corresponding to the first column of FIG. 2) and (ii) hybrid schemas (e.g., voiced fricatives such as TIFF0007688858000010.tif13170).
[0028] The various methods disclosed herein for recognizing the phonetic sounds and harmonics within a spoken human utterance rely on stored data representing the harmonic spectra of each of a plurality of phonetic sounds or harmonics, including the harmonic frequencies expected for each phonetic sound or harmonic (e.g., the tables of FIGS. 2, 3A, and 3B). Methods for generating such data for a set of phonetic sounds and harmonics can include having one or more human subjects instructed to clearly pronounce a particular sound or word perform acoustic spectral analysis on multiple utterances of the phonetic sound while retaining the labeled information for the instructed pronunciation. Such a process is similar to “teaching” the system to recognize each phonetic sound as it occurs within the waveform representing the human utterance. In some examples, harmonic spectral data can be generated based on the utterances of a single human subject to train the system to recognize the utterances of that particular subject. In other examples, acoustic spectral data can be generated based on the average or other analysis of the utterances of multiple human subjects to train the system for more general speech recognition.
[0029] For each individual sound or chord, spectral analysis includes estimating the fundamental acoustic frequency and identifying two or more fundamental or harmonic components detected or identified within the spectrum. As described above, the components that are "detected" or "identified" are those having an intensity exceeding one or more appropriately defined detection thresholds. Each component has an acoustic frequency that is either the fundamental acoustic frequency or a harmonic acoustic frequency (i.e., an integer multiple of the fundamental acoustic frequency). The spectrum is characterized by the fundamental acoustic frequency, although the spectrum may or may not include a fundamental component at that fundamental acoustic frequency. Among the identified harmonic components, the highest harmonic acoustic frequency greater than 410 Hz is identified as the primary cap frequency, and that frequency is stored as part of the data of the voice code. Also, as described above, the acoustic frequencies of each identified fundamental or harmonic component are also stored. For some chords, as described above, the acoustic frequencies of one or more of the identified harmonic components in the primary band, secondary band, and fundamental band can also be included in the data.
[0030] To estimate the focus frequency of a sound, one or more subjects vocalize a subset of voice chords multiple times at multiple different fundamental frequencies and perform spectral analysis. The subset of voice chords share a dominant note. The focus frequency of the common dominant note can be estimated from the primary cap frequency of the vocalized voice chords. The focus frequency can be estimated from the observed primary cap frequency by an appropriate method such as the average value or median value of the observed primary cap frequencies.
[0031] In some harmonic acoustic schemas (e.g., nasal sounds, nasalized vowels, or approximants) or hybrid acoustic schemas (e.g., voiced fricatives), spectral analysis may identify fundamental or harmonic components only at frequencies below 410 Hz as the reduced baseband frequencies. The acoustic frequencies of these reduced baseband components can be included in the dataset that describes the corresponding harmonic or hybrid acoustic schema, and each also includes an indicator of whether or not there are non-harmonic components at higher frequencies.
[0032] By recognizing the importance of harmonic components in the recognition of speech sounds and chords in human speech, it is also possible to improve speech synthesis. The above-mentioned spectral data used to recognize chords can also be used to generate chords. To synthesize the selected speech chords, the primary cap frequency can be determined based on the corresponding fundamental tone (primary tone) and the selected fundamental frequency (i.e., pitch). In the case of harmonic speech chords, the primary cap frequency is (i) an integer multiple of the selected fundamental frequency, (ii) greater than 410 Hz, and (iii) closer to the focus frequency of the corresponding primary tone than the focus frequencies of other speech sounds. The frequency components of the frequency of the primary tone are included in the synthesized waveform segment. In some examples, when the selected speech chord includes a secondary tone, the secondary cap frequency can be determined as described above for the primary tone and the cap frequency, and the frequency components at the secondary cap frequency can be included in the synthesized waveform segment.
[0033] This method can be repeated for each set of chords and fundamental frequencies in a temporal sequence in which multiple different harmonic segments (vowels, nasalized vowels, nasals, approximants, etc.) and hybrid segments (voiced fricatives, etc.) are interspersed. The synthesized harmonic segments, together with the synthesized inharmonic segments (e.g., voiceless fricatives), the synthesized voiceless segments (e.g., stops, within trills, within flaps), and the transition segments between them, constitute the synthesized human speech. The electronic waveform generated in this way using spectral data is applied to an electroacoustic transducer (such as a speaker) to generate the synthesized speech chords. Such a sequence of chords can be generated to constitute a synthesized speech sequence.
[0034] In some examples, to generate voice harmonics for which spectral data is available, an electronic waveform corresponding to the harmonics is created using the corresponding spectral data (primary cap frequency along with the acoustic frequencies of the harmonic components). For some voice sounds and harmonics, as described above, the acoustic frequencies of one or more identified harmonic components in the primary band, secondary band, or base band can be included in the data. For some harmonic or hybrid acoustic schemas (e.g., voice harmonics lacking primary and secondary components such as nasal sounds and fricative voiced sounds), the data can include reduced base band components and an indication of the presence or absence of high frequency inharmonic components.
[0035] The systems and methods disclosed herein can be implemented as a general-purpose or special-purpose computer, server, or other programmable hardware device programmed through software, or as "programmed" hardware or equipment through hardwiring, or as a combination of the two, or together with these, or with them. A "computer" or "server" can be composed of a single machine or multiple interacting machines (in a single location or multiple remote locations). A computer program or other software code, when used, can be implemented in tangible, non-transitory, temporary or permanent storage or replaceable media, including programming in microcode, machine code, network-based or web-based or distributed software modules operating together, RAM, ROM, CD-ROM, CD-R, CD-R / W, DVD-ROM, DVD±R, DVD±R / W, hard drive, thumb drive, flash memory, optical media, magnetic media, semiconductor media, or any future computer-readable storage alternative. The electronic indicators of the data set can be read from, received from, or stored in any of the tangible and non-transitory computer-readable media referred to herein.
[0036] In addition to the above, the following exemplary embodiments are included within the present disclosure or the appended claims.
[0037] Example 1. A method for a computer to execute to identify one or more vocal harmonics represented in an electronic waveform over time derived from human vocal speech, comprising: (a) for each of a plurality of harmonic acoustic spectra in a temporal sequence of an acoustic spectrum derived from the waveform, identifying two or more fundamental or harmonic components within the harmonic acoustic spectrum, each identified component having an intensity exceeding a detection threshold, and the identified components having frequencies separated by at least one integer multiple of a fundamental acoustic frequency associated with the acoustic spectrum; (b) for at least some of the plurality of acoustic spectra, identifying the highest harmonic frequency among the identified harmonic components as a primary cap frequency, the highest harmonic frequency also being greater than 410 Hz; and (c) for each of the plurality of acoustic spectra for which the primary cap frequency is identified in part (b), using the identified primary cap frequency to select at least one tone sound from a set of tone sounds as a main tone, the selected main tone corresponding to a subset of vocal harmonics in a set of vocal harmonics, using one or more electronic processors of a programmed computer system.
[0038] Example 2. The method according to Example 1, further comprising: each acoustic spectrum corresponding to one of a sequence of temporal sample intervals of the waveform, and using one or more of the electronic processors of a programmed computer system to (A) classify at least some of the acoustic spectra in the temporal sequence as only one of harmonic, inharmonic, hybrid, or voiceless; and (B) for at least some of the temporal sample intervals classified as harmonic in part (A), performing parts (a)-(c) and identifying the selected main tone as corresponding to that temporal sample interval of the electronic waveform.
[0039] Example 3. Each of the acoustic spectra corresponds to one of a series of distinct time segments in which the time-dependent acoustic spectrum of the waveform remains consistent with a single vocal harmony, and using one or more electronic processors of a programmed computer system, (A) classify at least some of the time segments into only one of harmonic, inharmonic, or hybrid based on their acoustic spectra, (B) for at least some of the time segments classified as harmonic in part (A), perform parts (a)-(c) and identify the selected tonic as corresponding to that time segment of the electronic waveform. The method according to Example 1, further comprising this.
[0040] Example 4. The method according to any one of Examples 1 to 3, further comprising using one or more electronic processors of a programmed computer system to derive a time sequence of acoustic spectra from an electronic waveform over time.
[0041] Example 5. The method according to any one of Examples 1 to 4, wherein for one or more different fundamental wave components or harmonic components, their respective detection thresholds are different from each other according to the acoustic frequency.
[0042] Example 6. Using one or more of the electronic processors of a programmed computer system, for at least one of the harmonic acoustic spectra in the time sequence of acoustic spectra derived from the waveform, further comprising identifying, within the harmonic acoustic spectrum, a fundamental component having a fundamental acoustic frequency and an intensity exceeding the detection threshold. The method according to any one of Examples 1 to 5.
[0043] Example 7. Using one or more of the electronic processors of a programmed computer system, further comprising selecting a vocal harmony from among the subsets of vocal harmonies of part (c) for a particular harmonic acoustic spectrum of a plurality of acoustic spectra, based at least in part on a comparison of (i) stored data indicative of the harmonic frequencies expected for a subset of vocal harmonies, and (ii) the harmonic frequencies of the harmonic components of the primary band of the harmonic acoustic spectrum, the primary band including harmonic components at one, two, or three of the largest consecutive multiples of the fundamental acoustic frequency that are greater than the fundamental acoustic frequency, less than the primary cap frequency, greater than 410 Hz, and greater than the smallest integer multiple of the fundamental acoustic frequency that is greater than 410 Hz and less than the primary cap frequency, the method according to any of Examples 1 to 6.
[0044] Example 8. Using one or more of the electronic processors of a programmed computer system, further comprising selecting a vocal harmony from among the subsets of part (c) for a particular harmonic acoustic spectrum of a plurality of acoustic spectra, based at least in part on a comparison of (i) stored data that is data indicative of the harmonic frequencies expected for the set of vocal harmonies, and (ii) the harmonic frequencies of the harmonic components of the secondary band of the harmonic acoustic spectrum, the secondary band including harmonic components at one or more harmonic acoustic frequencies that are greater than the smallest integer multiple of the fundamental acoustic frequency that exceeds 410 Hz and that are separated from the primary cap frequency by at least one intervening multiple of the fundamental acoustic frequency at which the acoustic spectrum lacks harmonic components, the method according to any of Examples 1 to 7, further comprising.
[0045] Example 9. Using one or more of the electronic processors of a programmed computer system, for a particular harmonic acoustic spectrum of a plurality of acoustic spectra, making the selection of part (c) at least partially based on a comparison of (i) stored data indicating the harmonic frequencies expected for a set of vocal harmonics and (ii) the harmonic frequencies of the harmonic components of the baseband of the harmonic acoustic spectrum, the baseband including harmonic components at one or more fundamental or harmonic acoustic frequencies that are less than 410 Hz or equal to the smallest integer multiple of the fundamental acoustic frequency between 410 Hz and the primary cap frequency, the method according to any of Examples 1 to 8,
[0046] Example 10. For a selected harmonic acoustic spectrum in which the highest harmonic acoustic frequency of the identified harmonic components is less than 410 Hz, using one or more electronic processors of a programmed computer system to (i) compare (A) stored data indicating the harmonic frequencies expected for a set of acoustic schemas with (B) the harmonic frequencies of each identified fundamental or harmonic component, and (ii) select one of those schemas from a set of harmonic or hybrid acoustic schemas at least partially based on the presence or absence of anharmonic frequency components at higher frequencies, the method according to any of Examples 1 to 9.
[0047] Example 11. A method executed by a computer for generating stored data indicating predicted harmonic frequencies for respective harmonic spectra of a set of vocal harmonies, the method comprising: (a) for each vocal harmony in the set, performing spectral analysis on a plurality of electronic waveforms obtained when one or more human subjects vocalize the vocal harmony, the spectral analysis including, for each electronic waveform, estimating a fundamental acoustic frequency, intensities each exceeding a detection threshold, and identifying two or more fundamental or harmonic components having acoustic frequencies equal to the fundamental acoustic frequency or an integer multiple of the fundamental acoustic frequency; and (b) for each vocal harmony for which a primary cap frequency has been identified in this way, identifying the highest harmonic acoustic frequency greater than 410 Hz among the identified harmonic components as the primary cap frequency, and storing an electronic indicator of the primary cap frequency and the acoustic frequencies of the identified fundamental or harmonic components in a tangible and non-transitory computer-readable storage medium that is not a transient propagation signal, using one or more electronic processors of a computer system programmed therefor.
[0048] Example 12. The method according to Example 11, further comprising: (i) for a subset of vocal harmonies having a common tonic and for a plurality of utterances at a plurality of different fundamental frequencies by one or more human subjects, estimating a focus frequency of the common tonic corresponding to the subset of vocal harmonies from the primary cap frequency; and (ii) storing an electronic indicator of the focus frequency of the tonic in a tangible and non-transitory computer-readable storage medium that is not a transient propagation signal.
[0049] Example 13. The method according to any one of Examples 11 or 12, wherein the detection threshold varies as a function of the acoustic frequency.
[0050] Example 14. The stored data includes, for at least some of the set of vocal harmonics, the harmonic frequencies of the harmonic components in the primary band of the corresponding harmonic acoustic spectrum, where the primary band is the primary cap frequency and 1, 2, 3, or more of the largest consecutive multiples of the fundamental acoustic frequency that are (i) less than the primary cap frequency, (ii) greater than 410 Hz, and (iii) greater than the smallest integer multiple of the fundamental acoustic frequency that is greater than or equal to 410 Hz and less than the primary cap frequency, the method according to any of Examples 11 to 13.
[0051] Example 15. The stored data includes, for at least some of the set of vocal harmonics, the harmonic frequencies of the harmonic components in the secondary band of the corresponding harmonic acoustic spectrum, where the secondary band includes one or more harmonic components at harmonic acoustic frequencies that are greater than the smallest integer multiple of the fundamental acoustic frequency that exceeds 410 Hz and are separated from the primary cap frequency by at least one intervening multiple of the fundamental acoustic frequency at which the acoustic spectrum lacks harmonic components, the method according to any of Examples 11 to 14.
[0052] Example 16. The stored data includes, for at least some of the set of vocal harmonics, the harmonic frequencies of the harmonic components in the base band of the corresponding harmonic acoustic spectrum, where the base band is less than 410 Hz or equal to the smallest integer multiple of the fundamental acoustic frequency between 410 Hz and the primary cap frequency, the method according to any of Examples 11 to 15, including one or more harmonic components at one or more fundamental acoustic frequencies or harmonic acoustic frequencies.
[0053] Example 17. For each of one or more additional harmonic acoustic schemes or hybrid acoustic schemes, spectral analysis of a plurality of electronic waveforms obtained when one or more human subjects utter the scheme, wherein for each electronic waveform, the spectral analysis comprises an estimation of a fundamental acoustic frequency and an identification of two or more fundamental or harmonic components each having an intensity exceeding a detection threshold and having a harmonic acoustic frequency equal to the fundamental acoustic frequency or an integer multiple of the fundamental acoustic frequency, wherein each of the fundamental and harmonic acoustic frequencies is less than 410 Hz, and b) for each acoustic scheme for which a plurality of waveforms were analyzed in part (a), storing on a tangible and non-transitory computer-readable storage medium that is not a transient propagation signal, an indicator electronically indicating one or more fundamental and harmonic acoustic frequencies of each identified fundamental or harmonic component, and the presence or absence of high-frequency inharmonic components, the method according to any of Examples 11 to 16 further comprising this.
[0054] Example 18. A method that a computer executes to synthesize a temporal segment of an electronic waveform and apply the waveform segment to an electroacoustic transducer to generate a sound of a vocal harmony selected from a set of vocal harmonies, comprising: (a) determining a primary cap frequency using the focus frequency and the selected fundamental acoustic frequency, using received, acquired, or calculated data indicating a fundamental tone corresponding to the selected vocal harmony and the focus frequency of the fundamental tone, wherein the primary cap frequency is (i) an integer multiple of the selected fundamental frequency, (ii) greater than 410 Hz, and (iii) closer to the focus frequency of the corresponding fundamental tone than the focus frequencies of other vocal harmonies; and b) including in the waveform segment a harmonic component of the primary cap frequency, wherein the primary cap frequency is greater than the acoustic frequencies of all other harmonic components included in the waveform segment, using one or more electronic processors of a computer system programmed for this purpose.
[0055] Example 19. The method of Example 18, further comprising repeating the method of Example 18 for each vocal harmony in a temporal sequence of a plurality of different harmonic segments or hybrid segments, together with inharmonic segments or voiceless segments that constitute human speech together, and transition segments therebetween.
[0056] Example 20. The method according to any one of Examples 18 or 19, further comprising applying a waveform segment for generating the sounds of a selected vocal harmony to an electroacoustic transducer.
[0057] Example 21. For one or more vocal harmonies of a set, including harmonic components of a primary band in a waveform segment, the primary band being the primary cap frequency and harmonic components at one, two, three, or more of the largest consecutive multiples of a fundamental acoustic frequency that (i) is less than the primary cap frequency, (ii) is greater than 410 Hz, and (iii) is greater than the smallest integer multiple of the fundamental acoustic frequency that is greater than 410 Hz and less than the primary cap frequency, the method according to any one of Examples 18 to 20, further comprising this.
[0058] Example 22. For one or more vocal harmonies of a set, including harmonic components of a primary band in a waveform segment, the primary band being the primary cap frequency and at least three of the largest consecutive multiples of the fundamental acoustic frequency that (i) is less than the primary cap frequency, (ii) is greater than 410 Hz, and (iii) is greater than the smallest integer multiple of the fundamental acoustic frequency that is greater than 410 Hz and less than the primary cap frequency, the method according to any one of Examples 18 to 20, further comprising including harmonic components of at least three of the largest consecutive multiples.
[0059] Example 23. For one or more vocal harmonics of a set, (a) using the received, retrieved, or calculated data indicating the secondary sound wave corresponding to the selected vocal harmonic and its focus frequency, determine a secondary cap frequency using the focus frequency and the selected fundamental acoustic frequency, where the secondary cap frequency is (i) an integer multiple of the selected fundamental frequency, (ii) greater than 410 Hz, and (iii) closer to the focus frequency of the corresponding secondary sound than the focus frequencies of other vocal harmonics, and (b) include one or more harmonic components of the secondary band in the waveform segment, where the secondary band is (i) greater than the smallest integer multiple of the fundamental acoustic frequency exceeding 410 Hz and (ii) includes harmonic components of harmonic acoustic frequencies below the secondary cap frequency. The method according to any one of Examples 18 to 22, further comprising the above.
[0060] Example 24. For one or more vocal harmonics of a set, include one or more harmonic components of the base band in the waveform segment, where the base band (i) is less than 410 Hz or (ii) is equal to the smallest integer multiple of the fundamental acoustic frequency between 410 Hz and the primary cap frequency, including harmonic components at one or more fundamental acoustic frequencies or harmonic acoustic frequencies. The method according to any one of Examples 18 to 23, further comprising the above.
[0061] Example 25. For at least one additional time segment of the electronic waveform corresponding to the selected harmonic or hybrid acoustic schema of a set of such acoustic schemas, include one or more harmonic components of only the reduced base band in the additional waveform segment, where the reduced base band includes only harmonic components at one or more fundamental or harmonic acoustic frequencies that are less than 410 Hz. The method according to any one of Examples 18 to 24, further comprising the above.
[0062] Example 26. For at least one additional time segment of an electronic waveform corresponding to a selected harmonic or hybrid acoustic schema of such a set of acoustic schemas, the additional waveform segment further includes (i) including one or more harmonic components of only the reduced baseband, where the reduced baseband includes only harmonic components of one or more fundamental or harmonic acoustic frequencies less than 410 Hz, and (ii) one or more high-frequency inharmonic components. The method according to any one of Examples 18 to 25.
[0063] Example 27. A programmed computerized machine including one or more electronic processors and one or more tangible computer-readable storage media operably coupled to the one or more processors, the machine being configured and programmed to execute the method according to any one of Examples 1 to 26.
[0064] Example 28. An article including a tangible medium that is not a transient propagation signal, the medium being encoded with computer-readable instructions that, when applied to a computer system, direct the computer system to execute the method according to any one of Examples 1 to 26.
[0065] This disclosure is illustrative and not restrictive. Further modifications will be apparent to those skilled in the art in light of this disclosure and are intended to fall within the scope of this disclosure. Equivalents or modifications of the disclosed exemplary embodiments and methods are intended to be included within the scope of this disclosure.
[0066] In the foregoing detailed description, for the purposes of streamlining the present disclosure, various features may be grouped in some exemplary embodiments. This method of disclosure is not to be construed as reflecting an intention that the identified embodiments require more features than are expressly recited therein. Rather, the subject of the invention may lie in fewer features than all of the features of a single disclosed exemplary embodiment. Accordingly, the present disclosure is to be construed as implicitly disclosing any embodiment having any suitable subset of one or more features shown, described, or otherwise specified herein. The “suitable” subset of features includes only those features that are not incompatible or mutually exclusive with respect to other features of that subset. Further, it should be noted that the cumulative scope of the examples listed above may, but does not necessarily, encompass the entirety of the subject matter disclosed herein.
[0067] For the purposes of this disclosure, the following interpretations shall apply. The terms "comprising", "including", "having", and variations thereof shall be construed as open-ended terms having the same meaning as if each instance thereof were followed by the phrase "at least" wherever they appear, unless explicitly stated otherwise. Singular nouns shall be construed as one or more, unless the context clearly dictates otherwise or there is an explicit statement of "sole", "single", or other similar limitation. The conjunction "or" shall be construed inclusively, except where (i) it is explicitly stated otherwise, for example, by use of the phrase "either", "only one of", or similar language, or (ii) two or more of the listed alternatives are understood or disclosed (implicitly or explicitly) to be incompatible or mutually exclusive within a particular context. In the latter case, "or" shall be understood to include only combinations of alternatives that are not mutually exclusive. As an example, each of "a dog or a cat", "a dog or one or more of a cat", "one or more of a dog or a cat" shall be construed as one or more dogs (no cats), one or more cats (no dogs), or each as one or more. As another example, each of "a dog, a cat, or a mouse", "a dog, a cat, or one or more of a mouse", "one or more of a dog, a cat, or a mouse" shall be construed as (i) one or more dogs (no cats or mice), (ii) one or more cats (no dogs or mice), (iii) one or more mice (no dogs or cats), (iv) one or more dogs and one or more cats (no mice), (v) one or more dogs and one or more mice (no cats), (vi) one or more cats and one or more mice (no dogs), (vii) one or more dogs, one or more cats, and one or more mice.In another example, each of "two or more of one dog, one cat, or one mouse" or "two or more dogs, cats, or mice" is construed as (i) one or more dogs and one or more cats (no mice), (ii) one or more dogs and one or more mice (no cats), (iii) one or more cats and one or more mice (no dogs), or (iv) one or more dogs and one or more cats and one or more mice. The same interpretation applies to "three or more", "four or more", etc.
[0068] In the present disclosure or the appended claims, when terms such as "substantially equal to", "substantially the same", "substantially greater than", "substantially less than" are used in relation to numerical quantities, standard conventions regarding measurement accuracy and significant figures shall apply unless different interpretations are explicitly specified. For ineffective amounts described in expressions such as "substantially prevented", "substantially non-existent", "substantially excluded", "substantially equal to zero", "negligible", etc., such expressions each indicate the case where the quantity in question has been reduced or decreased to that extent, and for practical purposes in the context of the intended operation or use of the disclosed or claimed apparatus or method, the overall behavior or performance of the apparatus or method is the same as would have occurred if the ineffective amount had actually been completely removed, been exactly equal to zero, or been otherwise exactly rendered ineffective.
[0069] For purposes of this disclosure, any labeling (e.g., first, second, third, etc., (a), (b), (c), etc., or (i), (ii), (iii), etc.) of elements, steps, limitations, or other portions of an embodiment or example is used only for purposes of clarification and is not to be construed as implying any ordering or preference of the portions so labeled. Where such ordering or preference is intended, it will be explicitly recited in the embodiment or example, or will be implicit or inherent based on the specific content of the embodiment or example. If the provisions of 35 U.S.C. § 112(f), or corresponding laws related to the "means + function" or "step + function" claim formats, are desired to be invoked in the description of an apparatus, the word "means" will be recited in that description. If these provisions are desired to be invoked in the description of a method, the description will recite the phrase "step." Conversely, if the words "means" or "step" are not shown, it is not intended to invoke such provisions.
[0070] One or more disclosures are incorporated herein by reference, and if such incorporated disclosures conflict in part or in whole with this disclosure, or have a different scope, this disclosure will control with respect to conflicting scopes, broader disclosures, or broader definitions of terms. If part or all of the incorporated disclosure conflicts with this disclosure, the later-disclosed disclosure will apply to the extent of the conflict.
[0071] The abstract is provided as an aid in searching for a particular subject matter in the patent document. However, the abstract is not intended to imply that the elements, features, or limitations recited therein necessarily are included in, or are required in any way by, the particular specification.
Claims
1. A method executed by a computer for recognizing human speech in the audible range represented by an electronic waveform over time derived from human speech in the audible range, comprising: (a) For each of a plurality of harmonic acoustic spectra in a temporal sequence of acoustic spectra derived from an electronic waveform over time, identifying two or more fundamental components or harmonic components within the harmonic acoustic spectrum, each identified component having an intensity exceeding its respective detection threshold, and each identified component having a frequency separated by at least one integer multiple of a fundamental acoustic frequency associated with the acoustic spectrum; (b) For at least some of the plurality of harmonic acoustic spectra, identifying the highest harmonic frequency among the identified harmonic components as the primary cap frequency, and this highest harmonic frequency also being greater than 410 Hz; (c) For each of the harmonic acoustic spectra in which the primary cap frequency is identified in part (b), using the identified primary cap frequency to select at least one sound voice from a set of sound voices as the main sound, the selected main sound corresponding to a subset of the sound voices in a set of sound voice harmonics, and (d) Generating an electronic indicator of text representing human speech in the audible range based on the selection in part (c). A method including using one or more electronic processors of a computer system programmed for this purpose.
2. Each of the acoustic spectra corresponds to one of a sequence of temporal sample intervals of the electronic waveform over time, (A) Classifying at least some of the acoustic spectra in the temporal sequence into only one of harmonic, inharmonic, hybrid, or silent; (B) Performing parts (a) to (c) for at least some of the temporal sample intervals classified as harmonic in part (A). The method according to claim 1, further including using one or more of the electronic processors of a computer system programmed for this purpose.
3. Each of the acoustic spectra corresponds to one of a series of distinct time segments in which the time-dependent acoustic spectrum of the electronic waveform over time remains consistent with a single sound voice harmonic. (A) classifying at least some of the time segments as only one of harmonic, inharmonic, or hybrid based on its acoustic spectrum, and (B) for at least some of the time segments classified as harmonic in part (A), performing parts (a) through (c), The method according to claim 1, further comprising using one or more electronic processors of a computer system programmed therefor. **Claim 4** The method according to claim 1, further comprising using one or more electronic processors of a programmed computer system to derive a time sequence of acoustic spectra from an electronic waveform over time. **Claim 5** The method according to claim 1, wherein for the two or more different basic components or harmonic components, the respective detection thresholds are different from each other according to the acoustic frequency. **Claim 6** The method according to claim 1, further comprising using one or more of the electronic processors of the programmed computer system to identify, within the harmonic acoustic spectrum, the basic components having an intensity exceeding the respective detection thresholds for the fundamental acoustic frequency and the basic components for at least one of the harmonic acoustic spectra in the time sequence of acoustic spectra obtained from the electronic waveform over time. **Claim 7** Using one or more of the electronic processors of a programmed computer system to select, for a particular harmonic acoustic spectrum of a plurality of harmonic acoustic spectra, a vocal harmony from among a subset of vocal harmonies in part (c) based at least in part on a comparison of (i) stored data indicating the harmonic frequencies expected for the subset of vocal harmonies and (ii) the harmonic frequencies of the harmonic components in the primary band of the harmonic acoustic spectrum, the primary band including harmonic components at 1, 2, or 3 of the largest consecutive multiples of the fundamental acoustic frequency that are greater than the primary cap frequency, less than the primary cap frequency, greater than 410 Hz, and greater than the smallest integer multiple of the fundamental acoustic frequency that is greater than 410 Hz and less than the primary cap frequency. The method according to claim 1. **Claim 8** Using one or more of the electronic processors of a programmed computer system, for a particular harmonic acoustic spectrum of a plurality of harmonic acoustic spectra, further comprising selecting a vocal harmony from among a subset of part (c) based on a comparison of (i) stored data indicating the harmonic frequencies expected for the vocal harmony in a set of vocal harmonies, and (ii) the harmonic frequencies of the harmonic components in the secondary band of the harmonic acoustic spectrum, wherein the secondary band is greater than the smallest integer multiple of the fundamental acoustic frequency greater than 410 Hz and includes harmonic components at one or more harmonic acoustic frequencies separated from the primary cap frequency by at least one intervening multiple of the fundamental acoustic frequency at which the acoustic spectrum lacks harmonic components, the method of claim 1.
9. Using one or more of the electronic processors of a programmed computer system, for a particular harmonic acoustic spectrum of a plurality of harmonic acoustic spectra, further comprising making the selection of part (c) based on a comparison of (i) stored data indicating the harmonic frequencies expected for the vocal harmonies in a set of vocal harmonies, and (ii) the harmonic frequencies of the harmonic components in the base band of the harmonic acoustic spectrum, wherein the base band includes harmonic components at one or more fundamental or harmonic acoustic frequencies that are less than 410 Hz or equal to the smallest integer multiple of the fundamental acoustic frequency between 410 Hz and the primary cap frequency, the method of claim 1.
10. For a selected harmonic acoustic spectrum in which the highest harmonic acoustic frequency of the identified harmonic components is less than 410 Hz, using one or more electronic processors of a programmed computer system, (i) comparing (A) stored data indicating the harmonic frequencies expected for the acoustic schemas in a set of harmonic or hybrid acoustic schemas with (B) the harmonic frequencies of each identified fundamental or harmonic component, and (ii) based on the presence or absence of inharmonic frequency components at higher frequencies, selecting one of those schemas from the set of harmonic or hybrid acoustic schemas, further comprising the method of claim 1.
11. A method executed by a computer, wherein a fundamental tone is selected based on stored data indicating the harmonic frequencies expected for each harmonic acoustic spectrum of a set of vocal harmonies, the stored data being (A) For each vocal harmony in a set of vocal harmonies, performing spectral analysis on a plurality of electronic waveforms obtained from one or more human subjects vocalizing the vocal harmony, the spectral analysis comprising, for each electronic waveform, estimating a fundamental acoustic frequency, an intensity for each of the specified components that exceeds the respective detection threshold for that component, and estimating two or more fundamental or harmonic components having an acoustic frequency that is a harmonic acoustic frequency equal to the fundamental acoustic frequency or an integer multiple of the fundamental acoustic frequency, and, (B) For each vocal harmony in a set of vocal harmonies, identifying, among the components of part (A), the highest harmonic acoustic frequency greater than 410 Hz as the predicted primary cap frequency of stored data indicative of the harmonic frequencies expected in the respective harmonic spectra, and storing an electronic indicator of the predicted primary cap frequency for each vocal harmony in the set of vocal harmonies in a computer-readable storage medium, The method according to claim 1, generated by a method comprising using one or more electronic processors of a computer system programmed therefor. **Claim 12** For a subset of vocal harmonies having a common tonic, using a plurality of spectrally analyzed electronic waveforms obtained from a plurality of utterances at a plurality of different fundamental frequencies by one or more human subjects, (i) for each vocal harmony in the subset of vocal harmonies, among the components of part (A), the highest frequency separated from the predicted primary cap frequency for that respective vocal harmony by a multiple of at least one intervening fundamental acoustic frequency at which the acoustic spectrum lacks harmonic components, as the predicted secondary cap frequency of stored data indicative of the harmonic frequencies expected in the respective harmonic spectra, and (ii) storing in a computer-readable storage medium an electronic indicator of the predicted secondary cap frequency for each vocal harmony in the subset of vocal harmonies. The method according to claim 11, further comprising. **Claim 13** The method according to claim 11, wherein at least one of the respective detection thresholds varies as a function of acoustic frequency. **Claim 14** The stored data includes, for at least some of the voice harmonics in the set of voice harmonics, the harmonic frequencies of the harmonic components in the primary band of the corresponding harmonic acoustic spectrum, the primary band being the primary cap frequency and 1, 2, 3, or more of the largest consecutive multiples of the fundamental acoustic frequency that are (i) less than the primary cap frequency, (ii) greater than 410 Hz, and (iii) greater than the smallest integer multiple of the fundamental acoustic frequency that is greater than 410 Hz and less than the primary cap frequency. The method according to claim 11.
15. The stored data includes, for at least some of the voice harmonics in the set of voice harmonics, the harmonic frequencies of the harmonic components in the secondary band of the corresponding harmonic acoustic spectrum, the secondary band being greater than the smallest integer multiple of the fundamental acoustic frequency greater than 410 Hz and including one or more harmonic components at harmonic acoustic frequencies separated from the primary cap frequency by at least one intervening multiple of the fundamental acoustic frequency at which the acoustic spectrum lacks harmonic components. The method according to claim 11.
16. The stored data includes, for at least some of the voice harmonics in the set of voice harmonics, the harmonic frequencies of the harmonic components in the base band of the corresponding harmonic acoustic spectrum, the base band including one or more fundamental acoustic frequencies or harmonic components at harmonic acoustic frequencies equal to the smallest integer multiple of the fundamental acoustic frequency that is less than 410 Hz or between 410 Hz and the primary cap frequency. The method according to claim 11.
17. (A) Further including, for each one of the additional harmonic acoustic schema or hybrid acoustic schema, spectrally analyzing a plurality of electronic waveforms obtained when one or more human subjects uttered that schema, and for each electronic waveform, the spectral analysis including an estimation of the fundamental acoustic frequency and an estimation of two or more fundamental or harmonic components having intensities exceeding the respective detection thresholds for the respective identified components and having harmonic acoustic frequencies equal to the fundamental acoustic frequency or an integer multiple of the fundamental acoustic frequency, wherein each of the fundamental and harmonic acoustic frequencies is less than 410 Hz. (B) For each acoustic schema in which a plurality of electronic waveforms are analyzed in part (A) of this claim, on a computer-readable storage medium, store electronically information indicating the fundamental and harmonic acoustic frequencies for each fundamental or harmonic component in part (A) of this claim, and the presence or absence of inharmonic wave components at higher frequencies. The method according to claim 11, further comprising this.
18. A method executed by a computer for synthesizing speech from text representing human speech in the audible range, comprising: (a) For each of a plurality of vocal harmonics of the text, using received, retrieved, or calculated data indicating the fundamental tone corresponding to the vocal harmonic and the focus frequency of the fundamental tone, use the focus frequency and a selected fundamental acoustic frequency to determine a primary cap frequency, where the primary cap frequency is (i) an integer multiple of the selected fundamental acoustic frequency, (ii) greater than 410 Hz, and (iii) closer to the focus frequency of the corresponding fundamental tone than the focus frequencies of other voices. (b) For each of a plurality of vocal harmonics of the text, using the primary cap frequency determined in part (a), generate a corresponding harmonic electronic waveform segment, where the harmonic electronic waveform segment includes harmonic components of the primary cap frequency, and the primary cap frequency is greater than the acoustic frequencies of all other harmonic components included in the harmonic electronic waveform segment. (c) Generate an electronic waveform including a time sequence including the harmonic electronic waveform segments generated in part (b) together with electronic waveform segments corresponding to inharmonic segments, silent segments, or transition segments of the text, where the electronic waveform is configured to generate the sound of human speech corresponding to the text when applied to an electroacoustic transducer. Using one or more electronic processors of a computer system programmed for this purpose. The method includes this.
19. The method according to claim 18, comprising applying the electronic waveform to an electroacoustic transducer to generate the sound of human speech.
20. For the selected vocal harmony in part (a), further comprising including harmonic components in the primary band in the corresponding harmonic electronic waveform segment, the primary band being the primary cap frequency and 1, 2, 3, or more of the largest consecutive multiples of the fundamental acoustic frequency, where (i) it is less than the primary cap frequency, (ii) it is greater than 410 Hz, (iii) it is greater than 410 Hz and greater than the smallest integer multiple of the fundamental acoustic frequency that is less than the primary cap frequency, and including 1, 2, 3, or more of the largest consecutive multiples of harmonic components, the method according to claim 18.
21. For the selected vocal harmony in part (a), including harmonic components in the primary band in the corresponding harmonic electronic waveform segment, the primary band being the primary cap frequency and at least 3 of the largest consecutive multiples of the fundamental acoustic frequency, where (i) it is less than the primary cap frequency, (ii) it is greater than 410 Hz, (iii) it is greater than 410 Hz and greater than the smallest integer multiple of the fundamental acoustic frequency that is less than the primary cap frequency, and including at least 3 of the largest consecutive multiples of harmonic components, the method according to claim 18.
22. For the selected vocal harmony in part (a), (A) determining a secondary cap frequency using the received, retrieved, or calculated data indicating the secondary voice corresponding to the selected vocal harmony and the focus frequency of the secondary voice, using the focus frequency of the secondary voice and the selected fundamental acoustic frequency, where the secondary cap frequency is (i) an integer multiple of the selected fundamental acoustic frequency, (ii) greater than 410 Hz, (iii) closer to the focus frequency of the corresponding secondary voice than the focus frequencies of other voices, and (B) including harmonic components in one or more secondary bands in the corresponding harmonic electronic waveform segment, the secondary bands including harmonic components of harmonic acoustic frequencies that are (i) greater than the smallest integer multiple of the fundamental acoustic frequency greater than 410 Hz, (ii) less than or equal to the secondary cap frequency, further comprising the method according to claim 18.
23. For the selected vocal harmony of part (a), further including including one or more harmonic components of the fundamental band in the corresponding harmonic electronic waveform segment, where the fundamental band includes harmonic components at one or more fundamental acoustic frequencies or harmonic acoustic frequencies that are (i) less than 410 Hz, or (ii) equal to the smallest integer multiple of the fundamental acoustic frequency between 410 Hz and the primary cap frequency, the method according to claim 18.
24. For the selected vocal harmony of part (a), including one or more harmonic components only of the reduced fundamental band in the corresponding harmonic electronic waveform segment, where the reduced fundamental band includes only harmonic components at one or more fundamental or harmonic acoustic frequencies less than 410 Hz, the method according to claim 18, further comprising.
25. The method according to claim 24, further including including one or more non-harmonic wave components of higher frequencies in the corresponding harmonic electronic waveform segment for the selected vocal harmony.
Citation Information
Patent Citations
Device and method for phoneme analysis
JP1998187194A
Sound detecting device
JP1998301594A
Device and method for discriminating noise source
JP2003065836A
Personal identification system, personal identification method and storage device used for the personal identification system
JP2005091991A
Method and apparatus for electronically sythesizing acoustic waveforms representing a series of words based on syllable-defining beats
US9747892B1