Sound processing device and sound processing method
The sound processing device estimates formant frequencies from mouth images to modulate musical instrument signals, providing a non-contact performance effect and stable sound output, addressing the limitations of conventional talking modulators.
Patent Information
- Application Number
- JP2024077089
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-10
- Publication Date
- 2025-11-20
AI Technical Summary
Conventional talking modulators require a tube to be inserted into the performer's mouth, making it difficult to maintain a natural mouth shape and interfering with musical performance.
A sound processing device that estimates formant frequencies and amplitudes from mouth images to modulate musical instrument signals without physical contact, using digital filters to produce sound signals based on these estimates.
Enables a non-contact performance effect similar to talking modulators, allowing natural mouth shapes and stable sound output without interference, while avoiding noise contamination and sound quality degradation.
Smart Images

Figure 2025171592000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a sound processing device and a sound processing method. [Background technology]
[0002] Talking modulators have been known as musical effectors that produce a speech-like effect by replacing laryngeal primary sounds, which are located at the most upstream of the human vocalization process, with instrument sounds. Conventional talking modulators simulate laryngeal primary sounds by emitting acoustic signals from an instrument into the oral cavity using a tube. For example, Non-Patent Document 1 discloses a typical talking modulator. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] “MXR (registered trademark) TALK BOX”, [online], DUNLOP MANUFACTURING website, [searched February 10, 2024], Internet<URL:https: / / www.jimdunlop.com / mxr-talk-box / > [Non-patent document 2] red bears, "Singing Voice Research (Voice Training Research)", [online], Updated March 14, 2017, Free BGM Material Site, [Retrieved February 10, 2024], Internet<URL:https: / / nakano-sound.com / voicetraning / voicetraning16.html#google_vignette> Summary of the Invention [Problem to be solved by the invention]
[0004] However, with conventional talking modulators such as those disclosed in Non-Patent Document 1, a tube is inserted into the performer's mouth, which makes it difficult for the performer to assume the same mouth shape as when speaking naturally.
[0005] In consideration of the above circumstances, one aspect of the present disclosure aims to provide a sound processing device that can achieve the same performance effect as a talking modulator in a non-contact manner. [Means for solving the problem]
[0006] In order to solve the above problems, a sound processing device according to one aspect of the present disclosure includes an estimation unit that estimates the frequency and amplitude of formants related to a person's voice based on an input image showing the person's mouth, and a modulation unit that outputs a sound signal obtained by modulating a musical instrument performance signal according to the frequency and amplitude of the estimated formants.
[0007] A sound processing method according to one aspect of the present disclosure estimates the frequency and amplitude of formants related to a person's voice based on an image of the person's mouth, and outputs a sound signal obtained by modulating a musical instrument performance signal according to the estimated frequency and amplitude of the formants. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a block diagram showing an example of a performance system including a sound processing device according to a first embodiment. [Figure 2] FIG. 1 is a diagram illustrating a frequency spectrum of a human voice. [Figure 3] FIG. 10 is a diagram illustrating the degree of lip opening. [Figure 4] FIG. 2 is a block diagram showing an example of the sound processing device of FIG. 1. [Figure 5] 2 is a flowchart showing an example of a sound processing operation of the sound processing device of FIG. 1. [Figure 6] FIG. 10 is a block diagram showing an example of a performance system including a sound processing device according to a second embodiment. [Figure 7]FIG. 7 is a block diagram showing an example of the sound processing device of FIG. 6. [Figure 8] 7 is a flowchart showing an example of a sound processing operation of the sound processing device of FIG. 6. [Figure 9] FIG. 11 is a block diagram showing an example of a performance system including a sound processing device according to a third embodiment. [Figure 10] FIG. 10 is a block diagram showing an example of the sound processing device of FIG. 9. [Figure 11] 10 is a flowchart showing an example of a sound processing operation of the sound processing device of FIG. 9. [Figure 12] FIG. 10 is a block diagram showing an example of a sound processing device according to a first modified example. [Figure 13] FIG. 2 is a diagram illustrating the vertical and horizontal widths of lips. [Figure 14] FIG. 1 is a diagram showing examples of lip shapes corresponding to five vowels. [Figure 15] 13 is a flowchart showing an example of a sound processing operation of the sound processing device of FIG. 12. DETAILED DESCRIPTION OF THE INVENTION
[0009] A: First embodiment A1: Sound processing device configuration FIG. 1 is a block diagram showing an example of a performance system 1 including a sound processing device 20 according to the first embodiment. As shown in FIG. 1, the performance system 1 includes a storage device 10, a sound processing device 20, an image capturing device, an image capturing device 30, a musical instrument 40, a signal input device 50, and a playback device 60.
[0010] The storage device 10 is a computer-readable recording medium (e.g., a non-transitive recording medium readable by a computer). The storage device 10 includes a non-volatile memory and a volatile memory. Examples of the non-volatile memory include a ROM (Read Only Memory), an EPROM (Erasable Programmable Read Only Memory), and an EEPROM (Electrically Erasable Programmable Read Only Memory). Examples of the volatile memory include a RAM (Random Access Memory).
[0011] The storage device 10 stores a program pr1 and various information. The program pr1 defines the operation of the sound processing device 20. The storage device 10 may store the program pr1 read from a storage device in a server (not shown). In this case, the storage device in the server is an example of a computer-readable recording medium.
[0012] The sound processing device 20 reads the program pr1 from the storage device 10. By executing the program pr1, the sound processing device 20 functions as an estimation unit 21 and a modulation unit 22. At least one of the estimation unit 21 and the modulation unit 22 may be realized by a circuit such as a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), or an FPGA (Field Programmable Gate Array).
[0013] An input image IG captured by an imaging device 30 is input to the estimation unit 21.
[0014] The imaging device 30 outputs imaging information obtained by capturing an image of the outside world to the sound processing device 20.
[0015] The imaging device 30 is positioned so as to capture the face of the user U, particularly the mouth. The imaging device 30 captures an image including the mouth of the user U. Therefore, the input image input to the estimation unit 21 is an image showing the mouth of the user U. The user U is an example of a person.
[0016] The estimation unit 21 estimates the frequency and amplitude of formants related to the voice of the user U based on the input image. More specifically, in the first embodiment, the estimation unit 21 estimates the frequency and amplitude of formants based on the degree of lip opening of the user U. The estimation unit 21 transmits formant information FI indicating the estimated frequency and amplitude of formants to the modulation unit 22. Note that the formant frequency includes the formant center frequency, formant frequency band, etc. An example of a method for estimating the formant frequency and amplitude will be described below. First, the method for estimating the formant frequency will be described.
[0017] FIG. 2 is a diagram showing the frequency spectrum of human voice. Vibrations of the vocal cords resonate in the vocal tract, which is a closed tube. As shown in FIG. 2, several peaks, i.e., resonant frequency bands, can be seen in the envelope of the frequency spectrum of human voice. The frequency bands corresponding to the peaks are called formants, and these frequency bands are called, in order from lowest to highest, "first formant," "second formant," "third formant," "fourth formant," etc. In FIG. 2, the first formant, second formant, third formant, and fourth formant are indicated by F1, F2, F3, and F4, respectively. The quality of a voice is determined by the frequencies of each of the multiple formants.
[0018] The amplitude of the formant is estimated by determining the gain of a peaking filter (described later) according to the frequency of the estimated formant. More specifically, the estimation unit 21 calculates the gain of the peaking filter by applying the frequency of the estimated formant to a lookup table that defines the relationship between the frequency of the formant and the gain of the peaking filter.
[0019] FIG. 3 is a diagram illustrating the degree of lip opening. As shown in FIG. 3, in the first embodiment, the degree of lip opening op is defined as, for example, the distance between the lower end ULC at the left-right center of the upper lip UL and the upper end LLC at the left-right center of the lower lip LL. The degree of lip opening op is an example of the degree of opening. It is considered that the degree of lip opening op corresponds to the degree of jaw opening. In other words, the greater the degree of lip opening op, the greater the degree of jaw opening tends to be.
[0020] For example, as disclosed in Non-Patent Document 2, it is known that the frequency of the first formant F1 increases as the degree of jaw opening increases. More specifically, the first formant F1 is a resonant frequency band corresponding to the fundamental vibration of the vocal tract, which is a closed tube. It is known that of the five vowels, "a ( / a / )" has the highest first formant F1, followed by "e ( / e / )," "o ( / o / )," "i ( / i / )," and "u ( / u / )." Therefore, the frequency of the first formant F1 tends to increase as the degree of lip opening (op) increases.
[0021] In this embodiment, the lip opening degree op is considered to be correlated with the degree of jaw opening, and the lip opening degree op is measured from the input image. Note that in this embodiment, the lip opening degree op is defined as the distance between the lower end ULC at the center of the upper lip UL and the upper end LLC at the center of the lower lip LL, but this is not limiting and any measurement method may be used as long as it can accurately measure the distance relationship between the upper lip UL and the lower lip LL.
[0022] 1 again, the modulation unit 22 receives the performance signal PS from the signal input device 50 and the formant information FI from the estimation unit 21. The musical instrument 40 includes, for example, an electric guitar and a hardware synthesizer.
[0023] The signal input device 50 receives musical tones produced by the musical instrument 40 as a result of performance by the user U or another performer, and generates a performance signal PS representing the waveform of the musical tones. If the input signal relating to the musical tones is an analog signal, the signal input device 50 has an A / D converter that converts the analog signal into a digital signal. If the musical instrument 40 is a digital synthesizer, the input signal is a digital signal, and therefore no A / D converter is required.
[0024] The musical instrument 40 may be an acoustic guitar or other stringed instrument. When a stringed instrument is played, a sound collection device (not shown) is used. The sound collection device collects musical sounds produced from the musical instrument 40 by the performance of the user U or another performer, and outputs waveforms of the musical sounds to the signal input device 50. For example, an acoustic transducer such as a microphone is used as the sound collection device.
[0025] In the first embodiment, a configuration is exemplified in which the signal input device 50, which is separate from the sound processing device 20, is connected to the sound processing device 20 by wire or wirelessly, but the signal input device 50 may also be mounted on the sound processing device 20. Furthermore, the musical instrument 40 includes not only musical instruments configured with dedicated hardware such as a guitar or synthesizer, but also programs executed by a personal computer, a smartphone, etc. The signal input device 50 includes not only devices configured with dedicated hardware, but also programs executed by a personal computer, a smartphone, etc.
[0026] The modulation unit 22 modulates the performance signal PS by a digital filter whose characteristics are determined based on the first formant F1, and outputs the modulated performance signal as a sound signal SS to the playback device 60. In other words, the modulation unit 22 outputs the sound signal SS obtained by modulating the performance signal PS of the musical instrument 40 in accordance with the frequency and amplitude of the estimated formant to the playback device 60.
[0027] The playback device 60 plays back the sound signal SS input from the modulation unit 22. The playback device 60 has, for example, a D / A converter, an amplifier, and a speaker. The D / A converter converts the sound signal SS from digital to analog. The amplifier amplifies the analog sound signal SS. The speaker converts the amplified sound signal SS into air vibrations and emits the sound.
[0028] In the first embodiment, the sound signal SS is input to the playback device 60 and the sound is played back from a speaker. However, the sound signal SS may be input to another storage device and stored in the storage device as music data, or may be input to a device such as a music editing device and further processed.
[0029] Fig. 4 is a block diagram showing an example of the sound processing device 20 of Fig. 1. As shown in Fig. 4, the estimation unit 21 has a trapezoidal correction unit 211, a mouth opening measurement unit 212, and a formant estimation unit 213.
[0030] The keystone correction unit 211 determines whether the lips of the user U in the input image IG are photographed from the front, and if it determines that the lips are not photographed from the front, performs keystone correction on the input image IG. Well-known keystone correction techniques are used for the keystone correction. This allows the mouth opening degree measurement unit 212 at the subsequent stage to measure the mouth opening degree op more accurately.
[0031] The mouth opening measurement unit 212 measures the lip opening degree op from the input image after keystone correction. The mouth opening measurement unit 212 includes an image recognition engine. The image recognition engine detects the lower end ULC at the left and right center of the upper lip UL and the upper end LLC at the left and right center of the lower lip LL from the input image after keystone correction. The mouth opening measurement unit 212 measures the mouth opening degree op based on the detected lower end ULC at the left and right center of the upper lip UL and the upper end LLC at the left and right center of the lower lip LL.
[0032] The formant estimation unit 213 estimates the frequency and amplitude of the first formant F1 based on the measured lip opening op. The formant estimation unit 213 transmits formant information FI indicating the estimated frequency and amplitude of the first formant F1 to the modulation unit 22.
[0033] The modulation unit 22 has a filter characteristics determination unit 221 and a filter unit 222. The filter characteristics determination unit 221 determines the characteristics of a digital filter in the filter unit 222 based on formant information FI. The digital filter in the filter unit 222 is a peaking filter. A peaking filter amplifies signals in a specific frequency band including a center frequency and passes signal components outside the specific frequency band without changing them. Therefore, the filter characteristics determination unit 221 determines the filter characteristics of the peaking filter, such as the center frequency, bandwidth, gain, and order, based on the formant information FI. In this way, the peaking filter is configured by a digital filter. In other words, the modulation unit 22 includes a peaking filter corresponding to the formant, and outputs the signal output from the peaking filter as the sound signal SS.
[0034] The digital filter may be, for example, an FIR (Finite Impulse Response) filter. For example, the filter characteristics determination unit 221 sets the estimated frequency of the first formant F1 as the center frequency of a peaking filter. The filter characteristics determination unit 221 sets the bandwidth, gain, and order according to the set center frequency of the peaking filter. Note that the formant information FI may include the sharpness (Q value) of the peaking filter in addition to the frequency and amplitude of the first formant F1. The sharpness is calculated, for example, by applying the estimated frequency to a lookup table that defines the relationship between the frequency and sharpness of the first formant F1.
[0035] The filter unit 222 has at least one peaking filter. The filter unit 222 modulates the performance signal PS input from the signal input device 50 with a peaking filter that exhibits determined characteristics. The filter unit 222 outputs the modulated performance signal as a sound signal SS to the playback device 60.
[0036] The pitch of the sound signal SS is determined by the relationship between the fundamental frequency f0 of the performance signal PS and the center frequency fc of the peaking filter. Generally, the peaking filter performs a process of enhancing the intensity of the amplitude spectrum at frequencies higher than the fundamental frequency f0 of the performance signal PS input to the filter. Therefore, when the fundamental frequency f0 of the performance signal PS is lower than the center frequency fc of the peaking filter (f0 < fc), the pitch of the sound signal SS is determined by the fundamental frequency f0, so the pitch does not change due to modulation. That is, the pitch of the performance signal PS and the pitch of the sound signal SS are the same. On the other hand, when the fundamental frequency f0 of the performance signal PS is higher than the center frequency fc of the peaking filter (f0 > fc), the pitch of the sound signal SS can change from the pitch of the performance signal PS.
[0037] When the fundamental frequency f0 of the performance signal PS is higher than the center frequency fc of the peaking filter, it may be a problem that the pitch changes due to modulation, making it difficult for listeners including the user U to grasp the pitch. In order to suppress the difficulty of grasping the pitch, for example, it is effective to perform modulation with a peaking filter on the harmonic components of the fundamental frequency f0 of the performance signal PS. According to this aspect, since the fundamental frequency f0 is not modulated by the peaking filter, the pitch of the fundamental frequency f0 does not change. On the other hand, due to the missing fundamental phenomenon, the fundamental frequency f0 is emphasized and heard by the listener, so the difficulty of grasping the pitch is suppressed. Note that an IIR (Infinite Impulse Response) filter may be adopted as the digital filter.
[0038] Fig. 5 is a flowchart showing an example of the sound processing operation of the sound processing device 20 of Fig. 1. Hereinafter, the sound processing operation of the sound processing device 20 will be described with reference to Fig. 5. The routine of Fig. 5 is started, for example, when the sound processing device 20 is started, and is executed every time a certain time period has elapsed.
[0039] Considering the smoothness of modulation and processing delay by the sound processing device 20, the shorter the certain time, the better. For example, if the processing delay in the sound processing device 20 exceeds 30 milliseconds, it may be difficult for the user U to play. On the other hand, the imaging speed of the imaging device 30 may be set to, for example, 30 fps, 60 fps, or 120 fps. Therefore, it is preferable to set the certain time to 16.7 milliseconds or less. 16.7 milliseconds corresponds to one frame time when the imaging speed of the imaging device 30 is set to 60 fps.
[0040] In step S11 , the sound processing device 20 functions as the keystone correction unit 211 to perform keystone correction on the input image IG input from the imaging device 30 .
[0041] In step S12, the sound processing device 20 functions as the mouth opening degree measuring unit 212 to measure the mouth opening degree op of the user U based on the trapezoidally corrected image.
[0042] In step S13, the sound processing device 20 functions as the formant estimation unit 213 to estimate the frequency and amplitude of the first formant F1 based on the measured mouth opening op, and transmits formant information FI to the filter characteristics determination unit 221. The formant information FI is information indicating the frequency and amplitude of the estimated first formant F1.
[0043] In step S14, the sound processing device 20 functions as the filter characteristic determination unit 221 to determine filter characteristics such as the center frequency of the peaking filter, the filter bandwidth, the filter gain, and the filter order based on the formant information FI of the first formant F1.
[0044] In step S15, the sound processing device 20 functions as the filter unit 222 to modulate the performance signal PS input from the signal input device 50 by a peaking filter that reflects the determined characteristics. In step S15, the sound processing device 20 functions as the filter unit 222 to output the modulated performance signal PS as a sound signal SS, and then temporarily ends this routine.
[0045] A2: Summary of the first embodiment As described above, the sound processing device 20 according to the first embodiment includes an estimation unit 21 and a modulation unit 22. The estimation unit 21 estimates the frequency and amplitude of formants related to the voice of the user U based on an input image IG showing the mouth of the user U. The modulation unit 22 outputs a sound signal SS obtained by modulating a performance signal PS of the musical instrument 40 in accordance with the frequency and amplitude of the estimated formants.
[0046] According to this embodiment, the performance signal PS can be modulated based on image information showing the user U's mouth. Therefore, the same effect as a talking modulator can be achieved without the need to install a sensor in or around the user U's mouth or for the user U to hold a tube in his / her mouth as in conventional talking modulators. In other words, a device that does not contact the human body can be provided. As a result, the user U can easily assume the same mouth shape as when speaking naturally, and the inconvenience of interference with playing the musical instrument 40 is eliminated. Furthermore, because image information is used to estimate the frequency and amplitude of the formants, the user U only needs to move his / her mouth and does not need to vocalize during the pronunciation operation. Therefore, a stable sound output can be obtained regardless of the user U's pronunciation technique.
[0047] Furthermore, according to this aspect, acoustic transducers are not required for either signal input or output, thereby suppressing a decrease in the S / N ratio and deterioration of sound quality due to noise contamination. For example, in the case of a conventional talking modulator, in a performance environment including a live performance venue or a personal music production environment, not only the user U's intraoral resonance but also background noise is input to the acoustic transducer, which can result in a decrease in the S / N ratio. Furthermore, in the case of a conventional talking modulator, sound quality can be degraded due to the acoustic characteristics of the acoustic transducer, and the use of an acoustic transducer can cause feedback. Therefore, according to this aspect, the above-mentioned problems of conventional talking modulators are solved.
[0048] In the sound processing device 20 according to the first embodiment, the frequency and amplitude of the estimated formant are the frequency and amplitude of the first formant F1.
[0049] According to this embodiment, the frequency and amplitude of the first formant F1 are estimated from an image showing the mouth of the user U, and the input performance signal PS is modulated according to the estimated frequency and amplitude of the first formant F1, so that the same effect as that of a talking modulator can be easily obtained.
[0050] Furthermore, in the sound processing device 20 according to the first embodiment, the estimation unit 21 estimates the frequency and amplitude of formants based on the lip opening op of the user U shown in the input image IG.
[0051] According to this embodiment, the frequency and amplitude of the formants are estimated from the lip opening degree op of the user U, and the input performance signal PS is modulated according to the estimated frequency and amplitude of the formants, so that the same effect as that of a talking modulator can be easily obtained.
[0052] Furthermore, in the sound processing device 20 according to the first embodiment, the modulation unit 22 includes a peaking filter corresponding to formants, and the modulation unit 22 outputs the signal output from the peaking filter as the sound signal SS.
[0053] According to this embodiment, the peaking filter can be designed in accordance with the characteristics of the formants, which makes it easier to design the modulation section 22.
[0054] In the sound processing device 20 according to the first embodiment, the peaking filter is configured by a digital filter.
[0055] According to this aspect, the characteristics of the peaking filter can be easily changed.
[0056] In addition, the sound processing method of the first embodiment estimates the frequency and amplitude of formants related to the user U's voice based on an image showing the user U's mouth, and outputs a sound signal SS obtained by modulating the performance signal PS of the instrument 40 according to the frequency and amplitude of the estimated formants.
[0057] According to this embodiment, the performance signal PS can be modulated based on image information showing the user U's mouth. Therefore, the same effect as a talking modulator can be achieved without the need to install a sensor in or around the user U's mouth or for the user U to hold a tube in his / her mouth as in conventional talking modulators. In other words, a device that does not contact the human body can be provided. As a result, the user U can easily assume the same mouth shape as when speaking naturally, and the inconvenience of interference with playing the musical instrument 40 is eliminated. Furthermore, because image information is used to estimate the frequency and amplitude of the formants, the user U only needs to move his / her mouth and does not need to vocalize during the pronunciation operation. Therefore, a stable sound output can be obtained regardless of the user U's pronunciation technique.
[0058] Furthermore, according to this aspect, acoustic transducers are not required for either signal input or output, thereby suppressing a decrease in the S / N ratio and deterioration of sound quality due to noise contamination. For example, in the case of a conventional talking modulator, in a performance environment including a live performance venue or a personal music production environment, not only the user U's intraoral resonance but also background noise is input to the acoustic transducer, which can result in a decrease in the S / N ratio. Furthermore, in the case of a conventional talking modulator, sound quality can be degraded due to the acoustic characteristics of the acoustic transducer, and the use of an acoustic transducer can cause feedback. Therefore, according to this aspect, the above-mentioned problems of conventional talking modulators are solved.
[0059] B: Second embodiment B1: Sound processing device configuration Fig. 6 is a block diagram showing an example of a performance system 1A including a sound processing device 20A according to the second embodiment. As shown in Fig. 6, the performance system 1A includes a storage device 10A, a sound processing device 20A, an imaging device 30, a musical instrument 40, a signal input device 50, a playback device 60, and a depth measurement device 70. The performance system 1A differs from the performance system 1 according to the first embodiment in that the performance system 1A includes a storage device 10A and a sound processing device 20A instead of the storage device 10 and the sound processing device 20 according to the first embodiment, and in that the performance system 1A includes a depth measurement device 70.
[0060] The storage device 10A is a computer-readable recording medium. The storage device 10A includes a nonvolatile memory and a volatile memory. The nonvolatile memory is, for example, a ROM, an EPROM, or an EEPROM. The volatile memory is, for example, a RAM.
[0061] The storage device 10A stores a program pr2 and various information. The program pr2 defines the operation of the sound processing device 20A. The storage device 10A may store the program pr2 read from a storage device in a server (not shown). In this case, the storage device in the server is an example of a computer-readable recording medium.
[0062] The sound processing device 20A reads a program pr2 from the storage device 10A. By executing the program pr2, the sound processing device 20A functions as an estimation unit 21A and a modulation unit 22A. At least one of the estimation unit 21A and the modulation unit 22A may be realized by a circuit such as a DSP, an ASIC, a PLD, or an FPGA.
[0063] The estimation unit 21A receives an input image IG captured by the imaging device 30 and depth information DI relating to the depth measured by the depth measurement device .
[0064] The configurations of the imaging device 30, musical instrument 40, signal input device 50, and playback device 60 are the same as those of the imaging device 30, musical instrument 40, signal input device 50, and playback device 60 in the first embodiment, and therefore will not be described.
[0065] The depth measurement device 70 includes a depth sensor. The depth sensor measures the distance to an object based on the time it takes for an infrared laser beam to be irradiated onto the object and for the laser beam to be reflected by the object and then received. This measurement method is called the TOF (Time of Flight) method. For example, a LiDAR (Light Detection and Ranging) is used as the depth sensor.
[0066] The distance measurement method is not limited to the TOF method, and may be a method of measuring the distance to an object by irradiating a specific pattern with laser light and analyzing distortion of the pattern in the received reflected light. Also, in the second embodiment, the image capture device 30 and the depth measurement device 70 are configured as separate entities, but the image capture device 30 and the depth measurement device 70 may be configured as an integrated unit.
[0067] The depth measurement device 70 irradiates a specific area including the mouth of the user U with laser light, and recognizes the three-dimensional shape of the mouth of the user U from the reflected laser light.
[0068] The estimation unit 21A estimates the formant frequencies based on the input image IG and the depth information DI, and transmits formant information FI indicating the estimated formant frequencies to the modulation unit 22A.
[0069] The modulation unit 22A receives the performance signal PS from the signal input device 50 and the formant information FI from the estimation unit 21A. The modulation unit 22A modulates the performance signal PS using a digital filter whose characteristics are determined based on the formant information FI, and outputs the modulated performance signal as a sound signal SS to the reproduction device 60. In other words, the modulation unit 22A outputs the sound signal SS obtained by modulating the performance signal PS of the musical instrument 40 in accordance with the frequency and amplitude of the estimated formants to the reproduction device 60.
[0070] Fig. 7 is a block diagram showing an example of the sound processing device 20A of Fig. 6. As shown in Fig. 7, the estimation unit 21A has a trapezoidal correction unit 211, an opening degree measurement unit 212, a formant estimation unit 213A, and a tongue position measurement unit 214. The modulation unit 22A has a filter characteristics determination unit 221A and a filter unit 222A.
[0071] The configurations of the trapezoidal correction unit 211 and the mouth opening measurement unit 212 are similar to those of the trapezoidal correction unit 211 and the mouth opening measurement unit 212 in the first embodiment, and therefore a description thereof will be omitted.
[0072] The tongue position measurement unit 214 measures the position of the tongue in the oral cavity of the user U based on the depth information DI and the input image after trapezoidal correction. More specifically, the tongue position measurement unit 214 measures the tongue position by superimposing a depth map based on the depth information DI on the input image IG and extracting the depth of the part corresponding to the tongue of the user U. The tongue position measurement unit 214 transmits position information PI indicating the measured tongue position to the formant estimation unit 213A.
[0073] Based on the input image IG and the position information PI, the formant estimation unit 213A estimates the frequency and amplitude of the first formant F1 and the frequency and amplitude of the second formant F2 related to the voice of the user U. In other words, the formant estimation unit 213A estimates the frequency and amplitude of the first formant F1 and the frequency and amplitude of the second formant F2 based on the lip opening op of the user U and the position of the user U's tongue.
[0074] As described above, the first formant F1 is a resonant frequency band for the fundamental vibration of the vocal tract, which is a closed tube. The frequency of the first formant F1 tends to increase as the lip opening degree op increases and the tongue position decreases. Therefore, by using information about the depth of the oral cavity, the frequency of the first formant F1 can be estimated more accurately than when only the lip opening degree op is used.
[0075] The second formant F2 is a resonant frequency band for the triple vibration of the vocal tract. It is known that the second formant F2 is highest for "i (i / )," followed by "e ( / e / )," "u ( / u / )," "a ( / a / )," and "o ( / o / )" in that order. Therefore, the frequency of the second formant F2 tends to be higher the further forward the position where the tongue and palate approach. In other words, the frequency of the second formant F2 can be estimated by using information about the depth of the oral cavity.
[0076] From these facts, it is possible to estimate the frequency and amplitude of the first formant F1 and the frequency and amplitude of the second formant F2 from the lip opening and tongue position. Therefore, in the second embodiment, a sound processing device 20A is disclosed that estimates the frequency and amplitude of the formants with higher accuracy by measuring the tongue position of the user U in addition to the lip opening of the user U. The amplitude of the first formant F1 is estimated by determining a first gain g1 of a first peaking filter (described later) according to the estimated frequency of the first formant F1. The amplitude of the second formant F2 is estimated by determining a second gain g2 of a second peaking filter (described later) according to the estimated frequency of the second formant F2.
[0077] For example, formant estimation section 213A calculates the first gain g1 by applying the estimated frequency of first formant F1 to a first lookup table that defines the relationship between the frequency of first formant F1 and the first gain g1 of the first peaking filter.Formant estimation section 213A calculates the second gain g2 by applying the estimated frequency of second formant F2 to a second lookup table that defines the relationship between the frequency of second formant F2 and the second gain g2 of the second peaking filter.
[0078] The formant estimation section 213A transmits formant information FI indicating the estimated frequency and amplitude of the first formant F1 and the estimated frequency and amplitude of the second formant F2 to the filter characteristics determination section 221A.
[0079] The filter characteristics determination unit 221A determines the characteristics of a digital filter in the filter unit 222A based on the formant information FI. The filter unit 222A, which will be described later, has at least two peaking filters, including a first peaking filter and a second peaking filter. The first peaking filter is a filter corresponding to the first formant F1, and the second peaking filter is a filter corresponding to the second formant F2. The filter characteristics determination unit 221A determines the center frequency, bandwidth, gain, order, etc. of the first peaking filter and the second peaking filter based on the formant information FI input from the formant estimation unit 213A. In this way, the peaking filters are configured by digital filters. In other words, the modulation unit 22A includes peaking filters corresponding to the formants, and outputs the signals output from the peaking filters as the sound signal SS.
[0080] The filter unit 222A modulates the performance signal PS input from the signal input device 50 using a first peaking filter exhibiting the determined characteristics, and further modulates the signal that has passed through the first peaking filter using a second peaking filter exhibiting the determined characteristics. The filter unit 222A outputs the signal that has passed through the second peaking filter to the playback device 60 as a sound signal SS.
[0081] The first peaking filter and the second peaking filter are each configured by a digital filter. Alternatively, the performance signal PS may be input to the second peaking filter first, and the signal that has passed through the second peaking filter may be input to the first peaking filter.
[0082] Furthermore, the formant information FI may include a first sharpness of the first peaking filter and a second sharpness of the second peaking filter in addition to the frequency and amplitude of the first formant F1 and the frequency and amplitude of the second formant F2. The first sharpness is calculated, for example, by applying the estimated frequency of the first formant to a lookup table that defines the relationship between the frequency of the first formant F1 and the first sharpness. The second sharpness is calculated, for example, by applying the estimated frequency of the second formant to a lookup table that defines the relationship between the frequency of the second formant F2 and the second sharpness.
[0083] Fig. 8 is a flowchart showing an example of the sound processing operation of the sound processing device 20A of Fig. 6. Hereinafter, the sound processing operation of the sound processing device 20A will be described with reference to Fig. 8. The routine of Fig. 8 is started, for example, when the sound processing device 20A is started, and is executed every time a certain time period has elapsed.
[0084] In step S11, the sound processing device 20A functions as the keystone correction unit 211 to perform keystone correction on the input image IG input from the imaging device 30.
[0085] In step S12, the sound processing device 20A functions as the mouth opening degree measuring unit 212, thereby measuring the mouth opening degree op of the user U based on the keystone corrected image.
[0086] In step S21, the sound processing device 20A functions as the tongue position measurement unit 214, thereby measuring the position of the tongue based on the input image IG and the depth information DI.
[0087] In step S22, the sound processing device 20A functions as the formant estimation unit 213A to estimate the frequency and amplitude of the first formant F1 and the frequency and amplitude of the second formant F2, and transmits formant information FI to the filter characteristics determination unit 221. The formant information FI is information indicating the estimated frequency and amplitude of the first formant F1 and the frequency and amplitude of the second formant F2.
[0088] In step S23, the sound processing device 20A functions as the filter characteristics determining unit 221A to determine the filter characteristics such as the center frequency, bandwidth, gain, and order of the first and second peaking filters based on the formant information FI.
[0089] In step S24, the sound processing device 20A functions as the filter section 222A, modulating the performance signal PS input from the signal input device 50 using the first peaking filter and the second peaking filter that reflect the determined characteristics, outputting the modulated performance signal as a sound signal SS, and then temporarily terminating this routine.
[0090] B2: Summary of the second embodiment As described above, in the sound processing device 20A according to the second embodiment, the estimation unit 21A estimates the frequency and amplitude of the first formant F1 and the frequency and amplitude of the second formant F2 based on the input image IG and position information indicating the position of the tongue of the user U.
[0091] According to this aspect, the frequency and amplitude of the first formant F1 and the frequency and amplitude of the second formant F2 can be more accurately estimated from position information indicating the position of the user U's tongue, so that a modulation effect can be imparted to the performance signal PS that better reflects the user U's pronunciation actions.
[0092] C: Third embodiment C1: Configuration of sound processing device Fig. 9 is a block diagram showing an example of a performance system 1B including a sound processing device 20B according to the third embodiment. As shown in Fig. 9, the performance system 1B includes a storage device 10B, a sound processing device 20B, an imaging device 30, a musical instrument 40, a signal input device 50, and a playback device 60. The performance system 1B differs from the performance system 1 according to the first embodiment in that the performance system 1B includes a storage device 10B and a sound processing device 20B instead of the storage device 10 and the sound processing device 20 according to the first embodiment.
[0093] The storage device 10B is a computer-readable recording medium. The storage device 10B includes a nonvolatile memory and a volatile memory. The nonvolatile memory is, for example, a ROM, an EPROM, or an EEPROM. The volatile memory is, for example, a RAM.
[0094] The storage device 10B stores a program pr3, a learning model lm1, and various information. The program pr3 defines the operation of the sound processing device 20B. The learning model lm1 is a learning model that has learned the relationship between mouth images when multiple people speak and vowels contained in the speech of the multiple people. The storage device 10B may store the program pr3 and the learning model lm1 read from a storage device in a server (not shown). In this case, the storage device in the server is an example of a computer-readable recording medium.
[0095] The sound processing device 20B reads the program pr3 from the storage device 10B. By executing the program pr3, the sound processing device 20B functions as an estimation unit 21B and a modulation unit 22B. At least one of the estimation unit 21B and the modulation unit 22B may be realized by a circuit such as a DSP, an ASIC, a PLD, or an FPGA.
[0096] An input image IG captured by the imaging device 30 is input to the estimation unit 21B.
[0097] The configurations of the imaging device 30, musical instrument 40, signal input device 50, and playback device 60 are similar to those of the imaging device 30, musical instrument 40, signal input device 50, and playback device 60 in the first embodiment, and therefore descriptions thereof will be omitted.
[0098] The estimation unit 21B estimates vowels included in the voice of the user U based on the input image IG, and determines the frequency and amplitude of the first formant F1 and the frequency and amplitude of the second formant F2 from the estimated vowels. The estimation unit 21B transmits formant information FI indicating the estimated frequency and amplitude of the first formant F1 and the frequency and amplitude of the second formant F2 to the modulation unit 22B.
[0099] The modulation unit 22B receives the performance signal PS from the signal input device 50 and the formant information FI from the estimation unit 21B. The modulation unit 22B modulates the performance signal PS using a digital filter whose characteristics are determined based on the formant information FI, and outputs the modulated performance signal as a sound signal SS to the reproduction device 60. In other words, the modulation unit 22B outputs the sound signal SS obtained by modulating the performance signal PS of the musical instrument 40 in accordance with the frequency and amplitude of the estimated formants to the reproduction device 60.
[0100] Fig. 10 is a block diagram showing an example of the sound processing device 20B of Fig. 9. As shown in Fig. 10, the estimation unit 21B has a trapezoid correction unit 211, a vowel estimation unit 215, and a formant estimation unit 213B. The modulation unit 22B has a filter characteristics determination unit 221B and a filter unit 222B.
[0101] The configuration of the keystone correction unit 211 is similar to the configuration of the keystone correction unit 211 in the first embodiment, and therefore a description thereof will be omitted.
[0102] The trapezoidally corrected input image is input to the vowel estimation unit 215. The vowel estimation unit 215 executes an inference process to estimate a vowel by inputting the trapezoidally corrected input image to a learning model lm1. More specifically, the vowel estimation unit 215 estimates a vowel included in the voice of the user U by inputting the trapezoidally corrected input image to the learning model lm1, which has learned the relationship between mouth images when multiple people speak and vowels included in the voices spoken by multiple people.
[0103] It is known that there is the following correlation between the five vowels and the first and second formants F1 and F2. For example, when a user U pronounces "i ( / i / )," the F1-F2 width is the widest among the F1-F2 widths for the pronunciation of the five vowels by the user U on the frequency spectrum. The F1-F2 width is the width between the first formant F1 and the second formant F2, i.e., the frequency difference between the first formant F1 and the second formant F2. The F1-F2 width when the user U pronounces "e ( / e / )" is narrower than the F1-F2 width when the user U pronounces "i ( / i / )."
[0104] The F1-F2 range when the user U pronounces "a ( / a / )" is relatively narrow. Furthermore, the first formant F1 when the user U pronounces "a ( / a / )" is higher than the first formant F1 when the user U pronounces "i ( / i / )", and the second formant F2 when the user U pronounces "a ( / a / )" is lower than the second formant F2 when the user U pronounces "i ( / i / )".
[0105] The F1-F2 range when the user U pronounces "o ( / o / )" is relatively narrow, and the first formant F1 and the second formant F2 are lower than the first formant F1 and the second formant F2 when the user U pronounces "a ( / a / )". The F1-F2 range when the user U pronounces "u ( / u / )" is wider than the F1-F2 range when the user U pronounces "o ( / o / )".
[0106] Based on the correlation between these five vowels and the first and second formants F1 and F2, the first and second formants F1 and F2 can be determined from the vowels pronounced by the user U.
[0107] The formant estimation unit 213B determines the frequency and amplitude of the first formant F1 and the frequency and amplitude of the second formant F2 based on the estimated vowel. The formant estimation unit 213B transmits formant information FI indicating the estimated frequency and amplitude of the first formant F1 and the frequency and amplitude of the second formant F2 to the filter characteristics determination unit 221B. As in the second embodiment, the formant information FI may include a first sharpness of the first peaking filter and a second sharpness of the second peaking filter in addition to the frequency and amplitude of the first formant F1 and the frequency and amplitude of the second formant F2. The method of calculating the first sharpness and the second sharpness conforms to the second embodiment.
[0108] The filter characteristics determination unit 221B determines the characteristics of a digital filter in the filter unit 222B based on the formant information FI. The filter unit 222B, which will be described later, has at least two peaking filters, including a first peaking filter and a second peaking filter. The first peaking filter is a filter corresponding to the first formant F1, and the second peaking filter is a filter corresponding to the second formant F2. The filter characteristics determination unit 221B determines the center frequency, bandwidth, gain, order, etc. of the first peaking filter and the second peaking filter based on the formant information FI input from the formant estimation unit 213B. In this way, the peaking filters are configured by digital filters. In other words, the modulation unit 22B includes peaking filters corresponding to the formants, and outputs the signals output from the peaking filters as the sound signal SS.
[0109] The filter unit 222B modulates the performance signal PS input from the signal input device 50 using a first peaking filter exhibiting the determined characteristics, and further modulates the signal that has passed through the first peaking filter using a second peaking filter exhibiting the determined characteristics. The filter unit 222B outputs the signal that has passed through the second peaking filter to the playback device 60 as a sound signal SS.
[0110] The first peaking filter and the second peaking filter are each configured by a digital filter. Alternatively, the performance signal PS may be input to the second peaking filter first, and the signal that has passed through the second peaking filter may be input to the first peaking filter.
[0111] Fig. 11 is a flowchart showing an example of the sound processing operation of the sound processing device 20B of Fig. 9. Hereinafter, the sound processing operation of the sound processing device 20B will be described with reference to Fig. 11. The routine of Fig. 11 is started, for example, when the sound processing device 20B is started, and is executed every time a certain time period has elapsed.
[0112] In step S11, the sound processing device 20B functions as the keystone correction unit 211 to perform keystone correction on the input image IG input from the imaging device 30.
[0113] In step S31, the sound processing device 20B functions as the vowel estimation unit 215, and estimates a vowel from the input image after the trapezoidal correction using the learning model lm1.
[0114] In step S32, the sound processing device 20B functions as a formant estimation unit 213B to estimate the frequency and amplitude of the first formant F1 and the frequency and amplitude of the second formant F2, and transmits formant information FI to the filter characteristics determination unit 221B. The formant information FI is information indicating the estimated frequency and amplitude of the first formant F1 and the frequency and amplitude of the second formant F2.
[0115] In step S23, the sound processing device 20B functions as the filter characteristic determining unit 221B to determine the filter characteristics such as the center frequency, bandwidth, gain, and order of the first and second peaking filters based on the formant information FI.
[0116] In step S24, the sound processing device 20B functions as the filter section 222B, modulating the performance signal PS input from the signal input device 50 using the first peaking filter and the second peaking filter that reflect the determined characteristics, outputting the modulated performance signal as a sound signal SS, and then temporarily terminating this routine.
[0117] C2: Summary of the third embodiment As described above, in the sound processing device 20B according to the third embodiment, the estimation unit 21B estimates the vowels contained in the voice of the user U based on the input image IG, and determines the frequency and amplitude of the first formant F1 and the frequency and amplitude of the second formant F2 from the estimated vowels.
[0118] From the image showing the mouth of the user U, parameters that can be used to estimate vowels, such as lip shape, jaw opening, and cheek muscle movement, can be detected. Furthermore, it is easy to determine the frequency and amplitude of the first formant F1 and the frequency and amplitude of the second formant F2 from the vowels. Therefore, according to this aspect, the frequency and amplitude of the first formant and the frequency and amplitude of the second formant can be easily estimated from the image showing the mouth of the user U, and thereby a modulation effect that better reflects the sound production behavior of the user U can be imparted to the performance signal PS.
[0119] In addition, the estimation unit 21B estimates the vowels contained in the voice of the user U by inputting the input image IG into a learning model lm1 that has learned the relationship between a mouth image showing the mouth when multiple people speak and the vowels contained in the voices spoken by multiple people.
[0120] According to this aspect, by using machine learning to learn the relationship between a plurality of mouth images and vowels, it is possible to more accurately estimate vowels and more accurately estimate the frequency and amplitude of the first formant and the frequency and amplitude of the second formant, thereby imparting a modulation effect to the performance signal PS that better reflects the pronunciation behavior of the user U.
[0121] D: Modification The present invention is not limited to the above-described embodiment, and various modifications can be adopted within the scope of the present invention. Specific modified embodiments are exemplified below. Furthermore, two or more embodiments arbitrarily selected from the following examples may be appropriately combined within a range that does not contradict each other. In the modified embodiments exemplified below, for elements whose actions and functions are equivalent to those of the above-described embodiment, the symbols used in the above explanation will be used, and detailed explanations of each will be omitted as appropriate.
[0122] D1: First modified example The lip vertical width h and lip horizontal width w are important parameters that determine vowels, and the first formant F1 and the second formant F2 can be estimated from the lip vertical width h and lip horizontal width w of the user U. Therefore, the estimation unit 21B in the third embodiment may have a lip shape measurement unit that measures the lip vertical width h and lip horizontal width w, instead of the vowel estimation unit 215.
[0123] Fig. 12 is a block diagram showing an example of a sound processing device 20C according to a first modification. The sound processing device 20C reads a program pr4 from a storage device 10C. By executing the program pr4, the sound processing device 20C functions as an estimation unit 21C and a modulation unit 22C. As shown in Fig. 12, the estimation unit 21C includes a trapezoidal correction unit 211, a lip shape measurement unit 216, and a formant estimation unit 213C. The modulation unit 22C includes a filter characteristics determination unit 221C and a filter unit 222C.
[0124] The configuration of the keystone correction unit 211 is similar to the configuration of the keystone correction unit 211 in the first embodiment, and therefore a description thereof will be omitted.
[0125] The keystone-corrected input image is input to the lip shape measurement unit 216. The lip shape measurement unit 216 measures the vertical width h and horizontal width w of the lips of the user U from the keystone-corrected input image. Well-known image recognition technology is used for the measurement.
[0126] The formant estimation unit 213C estimates the frequency and amplitude of the first formant F1 and the frequency and amplitude of the second formant F2 based on the measured vertical width h and horizontal width w of the user U's lips.
[0127] Fig. 13 is a diagram illustrating the vertical width h and horizontal width w of the lips. As shown in Fig. 13, in the first modified example, the vertical width h of the lips is defined, for example, as the distance between the lower end ULC at the center of the upper lip UL and the upper end LLC at the center of the lower lip LL. The horizontal width w of the lips is defined, for example, as the distance between the left and right corners CML and CMR of the mouth, which are the junctions of the upper lip UL and the lower lip LL. Therefore, the vertical width h of the lips is the same as the mouth opening op in the first embodiment.
[0128] A correlation is observed between the vertical width h and horizontal width w of the lips for the five vowels "a ( / a / )," "i ( / i / )," "u ( / u / )," "e ( / e / )," and "o ( / o / )." FIG. 14 is a diagram showing an example of the shape of the lips corresponding to the five vowels. For example, the vertical width h1 of the lips when "a ( / a / )" is pronounced is the largest among the vertical widths h of the lips when the five vowels are pronounced, and the horizontal width w1 of the lips when "a (a)" is pronounced is relatively large among the horizontal widths w of the lips when the five vowels are pronounced.
[0129] The vertical width h2 of the lips when "i" ( / i / ) is pronounced is relatively small among the vertical widths h of the lips when the five vowels are pronounced, and the horizontal width w2 of the lips when "i" ( / i / ) is pronounced is larger than the horizontal width w1 of the lips when "a" ( / a / ). The vertical width h3 of the lips when "u" ( / u / ) is pronounced is the smallest among the vertical widths h of the lips when the five vowels are pronounced, and the horizontal width w3 of the lips when "u" ( / u / ) is pronounced is relatively small among the horizontal widths w of the lips when the five vowels are pronounced.
[0130] The vertical width h4 of the lips when "e ( / e / )" is pronounced is relatively large among the vertical widths h of the lips when five vowels are pronounced, and the horizontal width w4 of the lips when "e ( / e / )" is pronounced is relatively large among the horizontal widths w of the lips when five vowels are pronounced. The vertical width h5 of the lips when "o ( / o / )" is pronounced is relatively small among the vertical widths h of the lips when five vowels are pronounced, and the horizontal width w5 of the lips when "o ( / o / )" is pronounced is the smallest among the horizontal widths w of the lips when five vowels are pronounced.
[0131] In this way, the combination of the vertical width h and horizontal width w for the above five vowels differs for each vowel, so the vowel pronounced by the user U is estimated from the image showing the mouth of the user U. Furthermore, as described in the third embodiment, there is a correlation between the vowel and the first and second formants F1 and F2, so the frequency and amplitude of the first and second formants F1 and F2 are directly estimated from the vertical width h and horizontal width w of the lips.
[0132] Fig. 15 is a flowchart showing an example of the sound processing operation of the sound processing device 20C of Fig. 12. Hereinafter, the sound processing operation of the sound processing device 20C will be described with reference to Fig. 15. The routine of Fig. 15 is started, for example, when the sound processing device 20C is started, and is executed every time a certain time period has elapsed.
[0133] In step S11, the sound processing device 20C functions as the keystone correction unit 211 to perform keystone correction on the input image IG input from the imaging device 30.
[0134] In step S41, the sound processing device 20C functions as the lip shape measuring unit 216 to measure the vertical width h and horizontal width w of the user U's lips.
[0135] In step S42, the sound processing device 20C functions as a formant estimation unit 213C to estimate the frequency and amplitude of the first formant F1 and the frequency and amplitude of the second formant F2 based on the vertical width h and horizontal width w of the lips of the user U, and transmits formant information FI to the filter characteristics determination unit 221C. The formant information FI is information indicating the estimated frequency and amplitude of the first formant F1 and the frequency and amplitude of the second formant F2.
[0136] In step S23, the sound processing device 20C functions as the filter characteristics determining unit 221C to determine the filter characteristics such as the center frequency, bandwidth, gain, and order of the first and second peaking filters based on the formant information FI.
[0137] In step S24, the sound processing device 20C functions as the filter section 222C, modulating the performance signal PS input from the signal input device 50 using the first peaking filter and the second peaking filter that reflect the determined characteristics, outputting the modulated performance signal as a sound signal SS, and then temporarily terminating this routine.
[0138] In this way, the estimation unit 21C according to the first variant estimates the frequency and amplitude of the first formant F1 and the frequency and amplitude of the second formant F2 based on the vertical width h and horizontal width w of the lips of the user U shown in the input image IG.
[0139] According to this embodiment, the input performance signal is modulated according to the estimated frequency and amplitude of the first formant F1 and the frequency and amplitude of the second formant F2, so that a modulation effect that better reflects the sound production action of the user U can be imparted to the performance signal PS.
[0140] D2: Second modified example In the first embodiment, the frequency and amplitude of the first formant F1 are estimated, and the performance signal PS is modulated by one peaking filter corresponding to the frequency and amplitude of the first formant F1. In the second and third embodiments, the frequency and amplitude of the first formant F1 and the frequency and amplitude of the second formant F2 are estimated, and the performance signal PS is modulated by a first peaking filter and a second peaking filter corresponding to the first formant F1 and the second formant F2, respectively.
[0141] However, the number of peaking filters may be equal to or greater than the number of estimated formants. For example, in the first embodiment, when the frequency and amplitude of the first formant F1 are estimated, a plurality of peaking filters may be combined to modulate the performance signal PS.
[0142] D3: Third variant In the second and third embodiments, the frequencies and amplitudes of formants higher than the third formant F3 are not estimated, and the sound signal SS modulated by the modulation units 22A and 22B does not have peaks corresponding to formants higher than the third formant F3 in the frequency spectrum. However, higher-order formant information corresponding to higher-order formants higher than the third formant F3 may be provided to the modulation units 22A and 22B as known information, and the modulation units 22A and 22B may reflect the higher-order formants in the sound signal SS based on the higher-order formant information.
[0143] The frequencies and amplitudes of higher-order formants above the third formant F3 are factors that characterize voice quality, so by reflecting the estimated frequencies and amplitudes of the higher-order formants in the sound signal SS, it is possible to impart characteristics unique to the user U, including voice quality, to the performance signal PS.
[0144] The higher-order formant information can be obtained by, for example, any of the following methods. For example, the higher-order formant information may be prepared as a predetermined constant. In this case, the higher-order formant information is uniquely determined and is not individually adjusted by the user U. Furthermore, the higher-order formant information may be prepared in advance as presets corresponding to multiple types of voice qualities, such as male voice, female voice, and Vocaloid voice.
[0145] Alternatively, the higher-order formant information may be obtained by analyzing the speech of the user U using a formant estimation method well known in the field of speech and acoustic analysis. Examples of well-known formant estimation methods include cepstrum analysis and linear predictive analysis in a source-filter model. As a preparation before using the sound processing device 20, utterances by the user U, such as "A-I-U-U...", are picked up using a microphone, and the picked-up speech is analyzed using cepstrum analysis, linear predictive analysis, etc., thereby extracting the higher-order formant information.
[0146] For example, in the second embodiment, as preparation before using the sound processing device 20, if utterances such as "A-i-uh..." by user U are recorded and used to train the learning model, the learning model can be optimized for user U.
[0147] Furthermore, the higher-order formant information may be extracted from a model that defines the correspondence between several variables, including personal characteristics such as age and neck size, and the higher-order formants. The model is created in advance. The several variables may be set by the user U before using the sound processing device 20. The sound processing device 20 may be configured so that the user U can adjust the variables using an input device such as a knob or slider.
[0148] The modulation sections 22A and 22B are provided with peaking filters according to the number of formants included in the higher-order formant information. When the formants included in the higher-order formant information are the third formant F3 and the fourth formant F4, the modulation sections 22A and 22B further include a third peaking filter corresponding to the third formant F3 and a fourth peaking filter corresponding to the fourth formant F4.
[0149] The third peaking filter and the fourth peaking filter are each composed of a digital filter. The third peaking filter receives the signal that has passed through the second peaking filter. The third peaking filter modulates the signal that has passed through the second peaking filter. The fourth peaking filter receives the signal that has passed through the third peaking filter. The modulation units 22A and 22B output the signal that has passed through the fourth peaking filter to the playback device 60 as a sound signal SS.
[0150] D4: Fourth variant The sound processing device 20 may include a reception unit that receives an operation by the user U so that the processing related to the operation of the sound processing device 20 can be started or stopped at any timing by the operation of the user U.
[0151] D5: Fifth variant The sound processing device 20 may be provided with a determination unit for determining whether or not the face of the user U is detected so that processing related to the operation of the sound processing device 20 is performed during the period in which the face of the user U is continuously detected by the imaging device 30.
[0152] D6: 6th variant In each of the above embodiments, a peaking filter is used in the modulation section 22, but the modulation section 22 may also be configured to use a band pass filter (BPF). The band pass filters are connected in parallel, and the signal that has passed through each band pass filter is superimposed on the dry source (the sound of the performance signal as it is) at an appropriate magnification.
[0153] D7: 7th variant In each of the above embodiments, the estimation units 21, 21A, 21B, and 21C include the keystone correction unit 211. However, the keystone correction unit 211 is not an essential function. Therefore, the estimation units 21, 21A, 21B, and 21C do not necessarily need to include the keystone correction unit 211.
[0154] D8: 8th variant In the first embodiment, the program pr1 is executed by the sound processing device 20, but the program pr1 may be executed on a server. Similarly, the program pr2 executed by the sound processing device 20A in the second embodiment and the program pr3 executed by the sound processing device 20B in the third embodiment may be executed on a server (not shown). That is, the programs pr1, pr2, and pr3 may be executed on the cloud.
[0155] E: Notes From the above-described exemplary embodiments, the following configurations can be understood, for example.
[0156] A sound processing device according to one aspect (aspect 1) of the present disclosure includes an estimation unit that estimates the frequency and amplitude of formants related to a person's voice based on an input image showing the person's mouth, and a modulation unit that outputs a sound signal obtained by modulating a musical instrument performance signal according to the frequency and amplitude of the estimated formants.
[0157] According to this aspect, the performance signal can be modulated based on image information showing a person's mouth. Therefore, the same effect as a talking modulator can be achieved without the need to place a sensor in or around the mouth or for the person to hold a tube in their mouth as in conventional talking modulators. In other words, a device that does not contact the human body can be provided. As a result, a person can easily assume the same mouth shape as when speaking naturally, and this eliminates the inconvenience of interference with playing an instrument. Furthermore, because image information is used to estimate the frequency and amplitude of formants, the person only needs to move their mouth and does not need to speak during the pronunciation action. Therefore, a stable sound output can be obtained regardless of the person's pronunciation technique.
[0158] Furthermore, according to this aspect, acoustic transducers are not required for either signal input or output, thereby suppressing a decrease in the S / N ratio and deterioration of sound quality due to noise contamination. For example, in the case of a conventional talking modulator, in a performance environment including a live performance venue or a personal music production environment, not only the user U's intraoral resonance but also background noise is input to the acoustic transducer, which can result in a decrease in the S / N ratio. Furthermore, in the case of a conventional talking modulator, sound quality can be degraded due to the acoustic characteristics of the acoustic transducer, and the use of an acoustic transducer can cause feedback. Therefore, according to this aspect, the above-mentioned problems of conventional talking modulators are solved.
[0159] In a sound processing device according to one aspect (aspect 2) of the present disclosure, the estimated formant frequency and amplitude are the frequency and amplitude of a first formant.
[0160] According to this aspect, the frequency and amplitude of the first formant are estimated from an image showing a human mouth, and the input performance signal is modulated according to the estimated frequency and amplitude of the first formant, thereby easily achieving the same effect as a talking modulator.
[0161] In the sound processing device according to one aspect (aspect 3) of the present disclosure, the estimation unit estimates the frequency and amplitude of the formant based on the degree of lip opening of the person shown in the input image.
[0162] According to this aspect, the frequency and amplitude of the formants are estimated from the degree of opening of the person's lips, and the input performance signal is modulated according to the estimated frequency and amplitude of the formants, so that an effect similar to that of a talking modulator can be easily obtained.
[0163] In the sound processing device according to one aspect (aspect 4) of the present disclosure, the estimation unit estimates the frequency and amplitude of the formant based on the vertical and horizontal widths of the person's lips shown in the input image.
[0164] The vertical and horizontal widths of the lips are known as important parameters for determining vowels. The relationship between the first and second formants and vowels is also known. Therefore, the frequency and amplitude of the first and second formants can be estimated based on the vertical and horizontal widths of the lips. According to this aspect, the input performance signal is modulated according to the estimated frequency and amplitude of the first and second formants, thereby imparting a modulation effect to the performance signal that better reflects human vocalization.
[0165] In a sound processing device according to one aspect (aspect 5) of the present disclosure, the estimation unit estimates vowels contained in the human voice based on the input image, and determines the frequency and amplitude of the formant from the estimated vowels.
[0166] From an image showing a person's mouth, parameters that can be used to estimate vowels, such as lip shape, jaw opening, and cheek muscle movement, can be detected. Furthermore, it is easy to determine the frequency and amplitude of the first formant and the frequency and amplitude of the second formant from a vowel. Therefore, according to this aspect, the frequency and amplitude of the first formant and the frequency and amplitude of the second formant can be easily estimated from an image showing a person's mouth, thereby imparting a modulation effect to the performance signal PS that more closely reflects the person's pronunciation.
[0167] In a sound processing device according to one aspect (aspect 6) of the present disclosure, the estimation unit estimates the vowels contained in the person's voice by inputting the input image into a learning model that has learned the relationship between mouth images showing mouths when multiple people speak and the vowels contained in the voices spoken by the multiple people.
[0168] According to this aspect, by using machine learning to learn the relationship between a plurality of mouth images and vowels, it is possible to more accurately estimate vowels and more accurately estimate the frequency and amplitude of the first formant and the frequency and amplitude of the second formant, thereby imparting a modulation effect to the performance signal that more accurately reflects human pronunciation.
[0169] In the sound processing device according to one aspect (aspect 7) of the present disclosure, the estimation unit estimates the frequency and amplitude of the formant based on the input image and position information indicating the position of the person's tongue.
[0170] According to this aspect, the frequency and amplitude of the first formant and the frequency and amplitude of the second formant can be more accurately estimated from position information indicating the position of the tongue, thereby imparting a modulation effect to the performance signal that more accurately reflects the human pronunciation action.
[0171] In a sound processing device according to one aspect (aspect 8) of the present disclosure, the modulation unit includes one or more peaking filters corresponding to the formants, and the modulation unit outputs the signals output from the one or more peaking filters as the sound signal.
[0172] According to this aspect, the peaking filter can be designed in accordance with the characteristics of the formants, which makes it easier to design the modulation section.
[0173] In a sound processing device according to one aspect (aspect 9) of the present disclosure, each of the one or more peaking filters is configured by a digital filter.
[0174] According to this aspect, the characteristics of the peaking filter can be easily changed.
[0175] A sound processing method according to one aspect (aspect 10) of the present disclosure estimates the frequency and amplitude of formants related to a person's voice based on an image showing the person's mouth, and outputs a sound signal obtained by modulating a musical instrument performance signal according to the estimated frequency and amplitude of the formants.
[0176] According to this aspect, the performance signal can be modulated based on image information showing a person's mouth. Therefore, the same effect as a talking modulator can be achieved without the need to place a sensor in or around the mouth or for the person to hold a tube in their mouth as in conventional talking modulators. In other words, a device that does not contact the human body can be provided. As a result, a person can easily assume the same mouth shape as when speaking naturally, and this eliminates the inconvenience of interference with playing an instrument. Furthermore, because image information is used to estimate the frequency and amplitude of formants, the person only needs to move their mouth and does not need to speak during the pronunciation action. Therefore, a stable sound output can be obtained regardless of the person's pronunciation technique.
[0177] Furthermore, according to this aspect, acoustic transducers are not required for either signal input or output, thereby suppressing a decrease in the S / N ratio and deterioration of sound quality due to noise contamination. For example, in the case of a conventional talking modulator, in a performance environment including a live performance venue or a personal music production environment, not only the user U's intraoral resonance but also background noise is input to the acoustic transducer, which can result in a decrease in the S / N ratio. Furthermore, in the case of a conventional talking modulator, sound quality can be degraded due to the acoustic characteristics of the acoustic transducer, and the use of an acoustic transducer can cause feedback. Therefore, according to this aspect, the above-mentioned problems of conventional talking modulators are solved. [Explanation of symbols]
[0178] 1...performance system, 20, 20A, 20B, 20C...sound processing device, 21, 21A, 21B, 21C...estimation unit, 22, 22A, 22B, 22C...modulation unit, 40...musical instrument, F1...first formant, F2...second formant, F3...third formant, F4...fourth formant, IG...input image, PI...position information, PS...performance signal, SS...sound signal, U...user, h, h1, h2, h3, h4, h5...vertical width, lm1...learning model, op...mouth opening degree, w, w1, w2, w3, w4, w5...horizontal width.
Claims
1. an estimation unit that estimates formant frequencies and amplitudes related to a person's voice based on an input image showing the person's mouth; a modulation unit that outputs a sound signal obtained by modulating a performance signal of a musical instrument in accordance with the frequency and amplitude of the estimated formant; A sound processing device comprising:
2. the estimated formant frequency and amplitude are the frequency and amplitude of the first formant; The sound processing device according to claim 1 .
3. the estimation unit estimates the frequency and amplitude of the formant based on a degree of lip opening of the person shown in the input image. The sound processing device according to claim 1 .
4. the estimation unit estimates the frequency and amplitude of the formant based on the vertical and horizontal widths of the person's lips shown in the input image; The sound processing device according to claim 3 .
5. the estimation unit estimates vowels included in the human voice based on the input image, and determines the frequency and amplitude of the formants from the estimated vowels. The sound processing device according to claim 1 .
6. the estimation unit estimates the vowels included in the speech of the person by inputting the input image into a learning model that has learned the relationship between mouth images showing mouths when multiple people speak and vowels included in the speech of the multiple people; The sound processing device according to claim 5 .
7. the estimation unit estimates the frequency and amplitude of the formant based on the input image and position information indicating a position of the person's tongue. The sound processing device according to claim 1 .
8. the modulation unit includes one or more peaking filters corresponding to the formants, and the modulation unit outputs signals output from the one or more peaking filters as the sound signal. The sound processing device according to claim 1 .
9. each of the one or more peaking filters is configured by a digital filter; The sound processing device according to claim 8 .
10. estimating formant frequencies and amplitudes for a voice of a person based on an image showing the person's mouth; a sound signal obtained by modulating a performance signal of a musical instrument in accordance with the frequency and amplitude of the estimated formant is output; Sound processing methods.