Human-machine interaction method for tuning specific consonant phonemes in speech synthesis
By setting up a Cartesian coordinate system in the speech synthesis software and receiving the peak point of the pronunciation intensity of consonant phonemes dragged by the user, fine-tuning of plosives, fricatives, and affricates is achieved, improving the naturalness and artistic expressiveness of the speech.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- DREAMTONICS XUNYUJIUYIN CO LTD
- Filing Date
- 2025-01-21
- Publication Date
- 2026-07-30
AI Technical Summary
Current speech synthesis technology lacks precise tuning methods for plosives, fricatives, and affricates in consonants, which affects the naturalness and artistic effect of speech.
By setting up a Cartesian coordinate system in the speech synthesis software, the system receives the user's dragging of the peak point of the consonant phoneme pronunciation intensity in the interactive interface, and adjusts the peak time of pronunciation intensity and loudness parameters to achieve fine-tuning of plosives, fricatives, and affricates.
It enhances the naturalness and artistic expressiveness of speech, and can subtly adjust the peak point of pronunciation intensity according to changes in singing, thereby enhancing the rhythm and cadence of speech.
Smart Images

Figure CN2025073611_30072026_PF_FP_ABST
Abstract
Description
A human-computer interaction method for adjusting specific consonant phonemes in speech synthesis Technical Field
[0001] This application relates to the field of speech synthesis technology, and in particular to a human-computer interaction method for adjusting specific consonant phonemes in speech synthesis. Background Technology
[0002] In the field of speech synthesis technology, technicians need to use speech synthesis software to tune existing speech (which can be human voices or voices created by artificial intelligence). This is also known as "tuning" or "adjustment," which involves fine-tuning and adding embellishments to existing speech to make it more natural and fluent or to express a certain artistic effect. Currently, there are various tuning methods in the field of speech synthesis technology, such as adjusting pitch, vocal tone, gender, and vibrato. There are also precedents for adjusting individual phonemes. However, there is currently no method for fine-tuning the internal structure of specific consonant phonemes, such as plosives, fricatives, and affricates, based on their characteristics.
[0003] The specific consonant phonemes described in this invention share a common characteristic in their pronunciation: each specific consonant phoneme has a "peak point of pronunciation intensity" after its initial position, where the intensity rapidly increases and the sound begins to be produced intensely. This peak point is particularly noticeable in singing, and it can change significantly with variations in vocal expression. Fricatives are produced when airflow is obstructed at a certain point in the vocal organs, forming a narrow passage. The airflow passing through this narrow passage generates friction, thus producing the sound. Plosives are produced when the vocal organs form an obstruction in the oral cavity, and then the airflow breaks through this obstruction. Affricates are consonants that combine the characteristics of both plosives and fricatives. At the beginning of pronunciation, the vocal organs form a complete obstruction, similar to a plosive, and then the airflow is slowly released, producing a fricative sound. The timing and loudness of the peak points of pronunciation intensity for these specific consonant phonemes have a significant impact on the rhythm and intonation of the speech containing that phoneme. The way these peak points of pronunciation intensity are handled subtly affects the effect of the speech and the listener's auditory experience.
[0004] In the field of speech synthesis technology, we are committed to making speech more natural, vivid, pleasant to the ear, and capable of expressing certain artistic effects. Therefore, this application, based on the characteristics of specific consonant phonemes, makes internal adjustments to those specific consonant phonemes, expands the tone adjustment methods in the field of speech synthesis technology, and can achieve more refined adjustments to speech on the basis of existing technology, subtly enhancing the expressiveness of speech. Summary of the Invention
[0005] This application provides a method for adjusting plosives, fricatives, and affricates in consonants, which is applied to speech synthesis software, the speech synthesis software including an audio processing engine and a user interface.
[0006] The method includes the following steps:
[0007] Step 1: Decompose the pronunciation of the text fitted to speech into phonemes, some of which are plosives, fricatives, and affricates ("specific consonant phonemes"). This invention only adjusts these specific consonant phonemes;
[0008] Step 2: Set up a Cartesian coordinate system with loudness on the vertical axis and time on the horizontal axis, taking only the first quadrant where both the horizontal and vertical coordinates are greater than or equal to zero. In this coordinate system, the coordinates of the peak sound intensity point are (peak sound intensity time, peak sound intensity loudness). Simultaneously, assign each phoneme a region of its corresponding phoneme adjustment interface ("phoneme range").
[0009] Step 3: Receive the user's dragging of the peak point of pronunciation intensity up, down, left, and right on the interactive interface (the dragging range is the range of phonemes assigned to that phoneme), and respond to the change in the coordinate value of the peak point of pronunciation intensity by dragging it, thereby adjusting the peak time of pronunciation intensity and the peak loudness parameter value of the peak point of pronunciation intensity for that specific consonant phoneme.
[0010] Step 4: Send the changed coordinates of the peak pronunciation intensity point, formed by dragging the peak pronunciation intensity point in Step 3, to the speech synthesis software.
[0011] The method described in this application is characterized in that: the text fitted to speech is not limited to any particular language type.
[0012] Each consonant phoneme information described in step one includes: the consonant phoneme, the time when the consonant phoneme begins to be pronounced, the time when the consonant phoneme reaches its peak intensity, the peak loudness of the consonant phoneme, and the time when the consonant phoneme ends to be pronounced.
[0013] According to the method described in this application, the user's dragging of the peak point of the sound intensity in the interactive interface in step three can be one or both of the peak time of the sound intensity and the peak loudness of the sound. Attached Figure Description
[0014] Figure 1 is a flowchart illustrating a human-computer interaction method for adjusting specific consonant phonemes in speech synthesis provided in this application;
[0015] Figure 2 is a schematic diagram of the user interface before user adjustment in an embodiment of the human-computer interaction method for adjusting specific consonant phonemes in speech synthesis provided in this application, applied to singing synthesis.
[0016] Figure 3 is a schematic diagram of the user interface after user adjustment in an embodiment of the human-computer interaction method for adjusting specific consonant phonemes in speech synthesis provided in this application, applied to singing voice synthesis. Embodiments of the present invention
[0017] The technical solution of this application will be clearly and completely described below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0018] Many specific details are set forth in the following description in order to provide a full understanding of this application. However, this application may also be implemented in other ways different from those described herein, and therefore this application is not limited to the specific embodiments disclosed below.
[0019] Referring to Figure 2, this is a schematic diagram of the user interface of a voice synthesis software, illustrating an embodiment of the human-computer interaction method for adjusting specific consonant phonemes in speech synthesis provided in this application, applied to English singing voice synthesis. The specific implementation method includes steps one through four:
[0020] Step 1: Decompose the lyrics pronunciation into phonemes. This invention only adjusts the following specific consonant phonemes: plosives, fricatives, and affricates among consonants;
[0021] As one implementation method, in this embodiment, the text (lyrics) fitted to the speech is English text.
[0022] In this embodiment, the specific consonant phoneme is [dr].
[0023] Step 2: Establish a Cartesian coordinate system with loudness on the vertical axis and time on the horizontal axis, taking only the first quadrant where both the vertical and horizontal coordinates are greater than or equal to zero. In this coordinate system, the coordinates of the peak sound intensity point are (peak sound intensity time, peak sound intensity). Simultaneously, assign a phoneme range to each phoneme.
[0024] As one implementation method, Figure 2 illustrates a user interface for the vocal synthesis software according to one embodiment of this application. Figure 2 is a schematic diagram of the user interface for the vocal synthesis software applied to this application. In this embodiment, the coordinates of the peak point P of the pronunciation intensity of the specific consonant phoneme [dr] are (t, v), and the phoneme range occupied by [dr] is the rectangle enclosed by a, b, c, and d.
[0025] Step 3: Receive the user's dragging of the peak point P of the pronunciation intensity of a specific consonant phoneme [dr] in the interactive interface (the dragging range is the rectangle enclosed by the phoneme ranges a, b, c, and d assigned to that phoneme), and respond to the change in the coordinate value of the peak point of pronunciation intensity by dragging it, thereby adjusting the peak time of pronunciation intensity and the peak loudness parameter value represented by the peak point of pronunciation intensity of that specific consonant phoneme.
[0026] As one implementation, Figure 3 shows the user interface of a singing voice synthesis software after the user drags the peak intensity point P of a specific consonant phoneme [dr] in the lower left direction. In response to dragging the peak intensity point P of the specific consonant phoneme [dr], its coordinates are adjusted to P´(t´, v´). This results in an adjustment to the pronunciation of the phoneme [dr], that is, the loudness of the peak pronunciation of the phoneme [dr] is reduced, and the peak pronunciation time is advanced.
[0027] As one implementation method, the user can drag the peak point of the vocal intensity up, down, left, and right in the user interface of the vocal synthesis software. This can be one or two of the following: moving the peak moment of vocal intensity forward or backward, or changing the loudness of the peak vocal intensity. This method is not limited to the dragging method in this embodiment.
[0028] Step 4: Send the coordinates of the changed peak point P' of the vocal intensity, formed by dragging the peak point P in Step 3, to the singing voice synthesis software.
[0029] As one implementation method, the singing voice synthesis software receives and responds to the coordinates (t', v') of the peak point P' of the pronunciation intensity, and can synthesize and output the sound after the consonant phoneme is tuned by the singing voice synthesis software and the speaker device.
Claims
1. A human-computer interaction method for adjusting specific consonant phonemes in speech synthesis, applied to speech synthesis software, characterized in that, include: Step 1: Decompose the pronunciation of the text fitted to speech into phonemes, some of which are plosives, fricatives, and affricates ("specific consonant phonemes"). This invention only adjusts these specific consonant phonemes; Step 2: Set up a Cartesian coordinate system with loudness on the vertical axis and time on the horizontal axis, taking only the first quadrant where both the horizontal and vertical coordinates are greater than or equal to zero. In this coordinate system, the coordinates of the peak sound intensity point are (peak sound intensity time, peak sound intensity loudness). Simultaneously, assign each phoneme a region of its corresponding phoneme adjustment interface ("phoneme range"). Step 3: Receive the user's dragging of the peak point of pronunciation intensity up, down, left, and right on the interactive interface (the dragging range is the range of phonemes assigned to that phoneme), and respond to the change in the coordinate value caused by dragging the peak point of pronunciation intensity, thereby adjusting the peak time of pronunciation intensity and the peak loudness parameter value of the peak point of pronunciation intensity for that specific consonant phoneme. Step 4: Send the changed coordinates of the peak pronunciation intensity point, formed by dragging the peak pronunciation intensity point in Step 3, to the speech synthesis software.
2. The method according to claim 1, characterized in that: The text fitted to speech is not limited to any particular language.
3. According to the method of claim 1, any specific consonant phoneme information includes: the consonant phoneme, the start time of the consonant phoneme, the peak time of the consonant phoneme's pronunciation intensity, the peak loudness of the consonant phoneme's pronunciation, and the end time of the consonant phoneme.
4. According to the method of claim 1, the adjustment of the peak point of the phoneme's pronunciation intensity within any phoneme range by the user in step three can be one or both of the peak pronunciation intensity time and peak pronunciation loudness.