Automatic reasoning audio correction Chinese zither melody generation method
By using the PYIN algorithm and the Guzheng audio conversion rule library, combined with pitch shifting and dynamic loudness correction strategies, the problems of range and listening experience in mapping human vocal humming to Guzheng melodies were solved, and high-quality generation of Guzheng melodies was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-12
- Publication Date
- 2026-03-13
AI Technical Summary
Existing technologies struggle to directly map human humming into guzheng melodies, resulting in issues such as mismatched pitch ranges, abnormal glissando, and the inability to play the melody. In particular, the low or high frequency ranges exceed the guzheng's playing range, and the original humming's pitch fluctuations affect the melody quality.
The PYIN algorithm is used to extract the fundamental frequency. Combined with the Guzheng audio conversion rule library and the pitch shift mechanism, the humming audio is adjusted to adapt to the Guzheng range through spectrum resampling and dynamic loudness correction strategies. Glissando substitution and technique processing are introduced to ensure that the melody is consistent with the original humming in terms of pitch relationship and dynamic expression.
It enables accurate and natural playing of hummed melodies on the guzheng, improves the accuracy and artistry of melody generation, solves the problems of exceeding the range limit and unpleasant listening experience, and provides intelligent support for the digital creation of traditional musical instruments.
Smart Images

Figure CN121661999A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of audio signal processing and intelligent music composition technology, specifically a method for processing human vocal humming audio based on the PYIN (Probabilistic YIN) algorithm. It involves establishing an audio conversion rule base and combining fundamental frequency extraction, pitch rise / fall, and pitch correction strategies to automatically convert human vocal humming into a simplified musical score melody that conforms to the range and performance characteristics of the guzheng (Chinese zither) playing. Background Technology
[0002] With the rapid development of artificial intelligence, intelligent music composition, and audio processing technologies, melody recognition based on human humming and automatic instrumental performance generation have become research hotspots. The guzheng, as an important representative of traditional Chinese musical instruments, demonstrates exceptional expressiveness in modern music fusion. However, due to the limited range of the guzheng, directly mapping free humming from low or high-pitched individuals into playable guzheng melodies still faces numerous challenges. When the frequency range of some humming is mostly outside the physical playing range of the guzheng, directly using the original fundamental frequency to construct a melody often results in mismatched ranges, abnormal glissando, or even an unplayable melody. Although several automatic pitch raising / lowering methods have emerged in existing audio processing technologies, they have not yet been adapted and applied to traditional instruments like the guzheng, which have a fixed scale structure and a special playing range, making it difficult to effectively map hummed melodies into playable guzheng melodies. Furthermore, if the pitch fluctuations in the original humming are not corrected, they can easily lead to out-of-tune melodies, unstable rhythms, or pitch deviations, affecting the performance quality of the generated melody.
[0003] Therefore, there is a need for an automatic method to generate hummed melodies for the guzheng, which combines high-precision fundamental frequency extraction, pitch rise / fall, and dynamic loudness correction. This method should be adapted to the unique range and scale structure of the guzheng, allowing hummed melodies to be converted into playable guzheng melodies while preserving their original expressiveness and relative pitch relationships. This would improve the accuracy, artistry, and practicality of melody generation, providing intelligent support for the digital creation and dissemination of traditional Chinese musical instruments. Summary of the Invention
[0004] In view of this, this invention proposes an automatic inference-based audio pitch correction method for generating guzheng melodies based on a guzheng audio conversion rule base. This method first uses the PYIN algorithm to extract the fundamental frequency of the humming audio to obtain an initial pitch sequence. Since the pitches of some humming voices are mostly outside the guzheng's range, directly mapping them to the guzheng's range will result in missing melodies or make them unplayable. Furthermore, due to the inherent characteristics of the guzheng instrument, the glissando technique, which cannot be naturally played across intervals, results in local dynamic intensity that does not match the original humming performance.
[0005] To address the aforementioned issues, this invention constructs a Guzheng audio conversion rule library ∑, establishing a formula for mapping sharp / flat semitones based on the twelve-tone equal temperament. It introduces a pitch sharp / flat mechanism, first determining the minimum number of sharp / flat semitones using the mapping formula, then automatically selecting the most ideal number based on the desired effect. Combined with an improved spectral resampling algorithm, the entire audio segment undergoes high-quality pitch sharp / flat processing, ensuring that the melody remains within the playable range of the Guzheng while maintaining its relative pitch relationships, and is mapped using simplified notation. To enhance the naturalness and dynamic expressiveness of the audio, this invention further introduces a dynamic pitch correction strategy based on the RMS (Root Mean Square) loudness envelope. Compared to traditional static pitch correction, which typically adjusts the maximum or average RMS value of the entire audio segment to a single target value, ignoring dynamic changes at different points in time, the dynamic pitch correction strategy considers the RMS loudness variation characteristics frame by frame. This makes the generated Guzheng melody more closely resemble the original humming in terms of loudness variation, thereby improving auditory consistency and musicality.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] A method for generating guzheng melodies based on the PYIN algorithm for pitch shifting and tone correction in hummed audio is attached. Figure 1 As shown, it includes the following steps:
[0008] Step 1: Construct a rule base and a formula for mapping sharp / flat semitones based on the characteristics of the guzheng;
[0009] Step 2: Preprocess the humming audio by using the PYIN algorithm to extract the fundamental frequency data of each frame and obtain pitch sequence information;
[0010] Step 3: After obtaining the semitone values of the raised / lowered pitch, the pitch sequence is processed by raising / lowering the pitch overall based on the improved spectrum resampling technology to match the range of the guzheng, while extracting the rhythm information;
[0011] Step 4: Map the corrected pitch sequence to a note sequence within the range of the guzheng's pitch range;
[0012] Step 5: Using rhythmic information and frame duration information, combined with human auditory characteristics, the notes are filtered to remove excessively short, invalid notes. Rests are added to audio frequencies outside the guzheng's register to obtain a regularized guzheng note sequence.
[0013] Step Six: Based on the guzheng note sequence and the experimentally recorded guzheng source audio information, combine the sharp / flat rules between adjacent notes and the characteristics of the guzheng instrument to add feasible guzheng techniques to construct a preliminary guzheng melody, and use single notes to replace glissando techniques that are not feasible.
[0014] Step 7: To avoid sudden volume changes caused by glissando substitution and pitch adjustment, the volume data information is corrected frame by frame using the RMS loudness envelope comparison method, so that the final guzheng melody is more natural.
[0015] Step one is as follows:
[0016] Based on the semitone difference between adjacent notes, pitch direction, and glissando characteristics, the corresponding glissando playing technique is determined; for glissando techniques that cannot be performed, the corresponding single note is used as a substitute; the appropriate tremolo treatment is selected based on the duration of the note and whether it is followed by a rest; a formula for mapping sharp / flat semitones is constructed based on the principle of twelve equal temperament to obtain the minimum sharp / flat value, so that the maximum value of the fundamental frequency falls within the range of guzheng playing, and the most ideal number of sharp / flat semitones can be selected based on the effect.
[0017] The rule base ∑ is as follows:
[0018] S1:
[0019] S2:
[0020] S3:
[0021] S4:
[0022] S5:
[0023] S6:
[0024] Where p i Let i ∈ [1, 7] be the solfège notation symbol corresponding to the current note, where i corresponds to do, re, mi, fa, so, la, si, r is a rest, u is an upward glissando, d is a downward glissando, sg is a short glissando, lg is a long glissando, ly is a long tremolo, sy is a short tremolo, ti is the duration of the current note, and t is the duration of the note. avg The average duration of a musical note.
[0025] The formula for mapping semitones in rising / falling tones is as follows:
[0026]
[0027] Where f max f0 is the maximum fundamental frequency value of the audio signal, and f0 is the maximum fundamental frequency value of the guzheng.
[0028] Step two is as follows:
[0029] The humming audio is segmented into frames, and the fundamental frequency is estimated for each frame of audio data. The PYIN algorithm introduces a probability model on the basis of the traditional YIN differential period estimation method. The confidence score of the candidate period value of each frame is calculated, and the frequency value corresponding to the candidate period with the highest posterior probability is selected as the fundamental frequency estimation result of the frame. The results are stored in the pitch1 array.
[0030] Step three specifically involves:
[0031] Based on the fundamental frequency information obtained in step two above, which exceeds the range of the guzheng's pitch range, the minimum pitch shift value is derived using the formula for mapping sharp / flat semitones. Then, the most ideal pitch shift value n is selected based on auditory effect. Using the audio feature value information in the pitch1 array and the pitch shift value n, frequency domain waveform-level pitch transformation technology is used to resample the spectrum, initially obtaining the pitch-shifted audio. The fundamental frequency information is then extracted again and stored in the pitch2 array. The fundamental frequency ratio calibration function is used to correct positions in pitch2 that deviate excessively from the theoretical pitch shift ratio in the time domain, ensuring that the corrected fundamental frequency value basically meets the requirements of the original fundamental frequency. The pitch relationship is further improved to enhance the accuracy of the key frequency data in the time domain. Finally, all key frequencies are stored in the `corrected_pitch` array. Key frequencies outside the guzheng's range are detected in the `corrected_pitch` array and set to 0 to prevent note mapping failure. The initial intensity envelope of the humming audio is obtained to determine the beat position information, which is stored in the `beats_position` array. The detailed algorithm flowchart is attached. Figure 2 .
[0032] Step four is as follows:
[0033] The corrected fundamental frequency array `corrected_pitch` is mapped to a sequence of guzheng musical notation, and each frame frequency value in the array is mapped to a guzheng musical note name.
[0034] Step five is as follows:
[0035] The time interval of each frame is calculated based on the fundamental frequency frame number and the duration of the humming audio. The frame number k corresponding to the minimum note duration threshold (e.g., 0.1s) based on human hearing perception is set. Then, the number of frames k is traversed through the number of frames in the simplified musical score using the beat position information. The same note in k consecutive frames or more is identified and merged into a valid note. The note value, start time and duration are recorded. Segments that do not meet the minimum duration are set as rests. The merged simplified musical score is combined with short rests and its duration is merged with the adjacent note segments to obtain the final guzheng simplified musical score sequence.
[0036] Step six is as follows:
[0037] Based on the correspondence between the simplified guzheng score and the guzheng sound source data in the database, the single notes of the filtered simplified guzheng score are converted according to the corresponding guzheng sound source to construct the melody corresponding to the guzheng instrument. Then, the simplified guzheng score is traversed note by note, and techniques are added between adjacent single notes that meet the conditions according to the constructed guzheng feature rule library. Furthermore, a dynamic pitch correction processing method based on RMS loudness envelope is introduced to adjust the RMS loudness at the frame level. The RMS loudness of each frame is aligned sequentially, and frame-level volume compensation and adjustment are performed on the guzheng audio to ensure that the dynamics of the melody are consistent with the original vocal performance style, thereby improving the overall dynamic performance and auditory consistency of the audio.
[0038]
[0039] Where N represents the number of sampling points in each frame, and xi is the amplitude value of the i-th sampling point.
[0040] Step seven is as follows:
[0041] Frame-level RMS loudness sequences were extracted from the pitch-shifted time-domain audio signal data and the guzheng time-domain audio signal data, and frame-level energy analysis was performed. Since the lengths of the two RMS sequences may be inconsistent, a resampling method was used to resample the pitch-shifted time-domain audio signal RMS data to make its length consistent with the guzheng time-domain audio signal RMS data, resulting in RMMS loudness sequences of the same duration. Based on the RMMS loudness sequences of the same duration, gain adjustment coefficients were calculated frame by frame. Based on the obtained gain adjustment coefficients, amplitude weighting was performed on each frame to obtain the corrected guzheng audio signal. The corrected guzheng audio information was output as a .wav audio file, completing the dynamic pitch correction strategy based on RMS loudness envelope matching. The specific algorithm flowchart is attached. Figure 3 .
[0042] The beneficial effects of this invention are as follows: This invention proposes a method for generating guzheng melodies based on the PYIN algorithm, involving pitch shifting / falling and tone correction for humming. Addressing issues such as the fact that most vocal humming pitches fall outside the guzheng's range and that glissando substitution results in volume fluctuations that don't match the dynamics of the guzheng audio, this invention systematically introduces high-precision fundamental frequency extraction, improved spectral resampling for pitch shifting / falling, and a frame-level RMS loudness dynamic tone correction strategy into the guzheng melody generation process. Utilizing existing audio processing algorithms, through innovative combinations and task matching, it successfully solves the problems of pitch range exceeding limits and unpleasant listening experience in the automatic conversion of vocal humming into guzheng melodies, providing a practical and efficient technical solution for intelligent performance and digital creation of traditional Chinese musical instruments. Attached Figure Description
[0043] To illustrate the objectives and technical solutions of this invention, the following figures are provided:
[0044] Figure 1 A flowchart of a method for generating guzheng melodies based on the PYIN algorithm for pitch shifting and tone correction in human vocal humming audio;
[0045] Figure 2 Here is a flowchart of the specific algorithm for step two;
[0046] Figure 3 Here is the detailed algorithm flowchart for step six;
[0047] Figure 4 This is a graph showing the score before pitch reduction obtained in step two of this embodiment of the invention; where the horizontal axis represents the frame number and the vertical axis represents the Hz.
[0048] Figure 5 This is a graph showing the uncorrected pitch reduction score obtained in step two of this embodiment of the invention; where the horizontal axis represents the frame number and the vertical axis represents the Hz.
[0049] Figure 6 This is a graph showing the pitch-corrected score obtained in step two of this embodiment of the invention; where the horizontal axis represents the frame number and the vertical axis represents the Hz.
[0050] Figure 7 This is a diagram of the guzheng musical score structure obtained in step four of this embodiment of the invention; where the horizontal axis represents time and the vertical axis represents the note name;
[0051] Figure 8 This is a graph showing the RMS score results of the humming reference audio and the guzheng audio obtained in step six of this embodiment of the invention; where the horizontal axis represents the frame number and the vertical axis represents the loudness.
[0052] Figure 9 This is a graph showing the humming reference audio and the corrected guzheng audio RMS score obtained in step six of this embodiment of the invention; where the horizontal axis represents the frame number and the vertical axis represents the loudness. Detailed Implementation
[0053] The following will be combined with the appendix Figure 3 The preferred embodiments of the present invention will be described in detail below.
[0054] Implementation Case: Taking a segment of a human humming "Tianlu" as an example, the audio is stored as a .wav file in the path E: / bianqu / tianlu.wav. Using a Python development environment, the audio file audio_file = 'E: / bianqu / tianlu.wav' is called, and the audio feature value information is extracted to obtain the array pitch1, voiced_flag, voiced_prob = librosa.pyin(audio_data, fmin = 20, fmax = 4200, sr = sample_rate). Based on the initial fundamental frequency array pitch1, the minimum pitch drop semitone number 6 is obtained by combining the formula of the rise / fall semitone mapping relationship. Then, based on the user's listening effect, the most ideal pitch drop semitone number 8 is selected. The improved spectrum resampling method is called to first perform spectrum resampling to downsample the entire audio, so that the downsampled audio signal fully meets the playable frequency range of the guzheng in the frequency domain. Subsequently, a pitch ratio calibration operation was performed on the fundamental frequency sequence re-extracted from the down-pitched audio to ensure its relative interval structure remained consistent with the original melody in the time domain, thereby further improving accuracy and usability, and obtaining the beat position of the audio. The original fundamental frequency, the fundamental frequency before improvement, and the fundamental frequency after improvement are shown in the attached figures. Figure 4 , Figure 5 , Figure 6 As shown, after the improvement, all fundamental frequencies fall within the guzheng's tonal range, exhibiting good consistency in both frequency domain coverage and temporal structure. The guzheng notes were filtered based on beat position, frame duration, and human auditory characteristics; the corresponding note arrangement is shown in the attached figure. Figure 7 As shown. Based on the constructed rule base, techniques are added between each note; furthermore, a dynamic pitch correction strategy based on RMS loudness envelope is introduced, using frames as features to correct the changes in RMS loudness frame by frame, and dynamically adjusting the gain of the generated guzheng audio frame by frame. The effects before and after pitch correction are shown in the attached figure. Figure 8 and Figure 9 As shown in the figure. The comparison results show that the corrected RMS curve of the guzheng is more consistent with the RMS curve of the humming audio in terms of overall trend, which verifies the expressive compensation effect of RMS tone correction processing in the context of glissando substitution and the effect of improving the overall listening consistency.
[0055] Finally, it should be noted that the above preferred embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail through the above preferred embodiments, those skilled in the art should understand that various changes can be made to it in form and detail without departing from the scope defined by the claims of the present invention.
Claims
1. A method for automatically generating guzheng melodies through audio tone correction, characterized in that, Includes the following steps: S1: Construct a rule base and a formula for mapping semitones in sharp / flat tones based on the characteristics of the guzheng; S2: Preprocess the humming audio, use the PYIN algorithm to extract the fundamental frequency data of each frame, and obtain pitch sequence information; S3: After obtaining the semitone values of the raised / lowered pitch, the pitch sequence is processed by raising / lowering the pitch overall based on the improved spectrum resampling technology to match the range of the guzheng, while extracting the rhythm information; S4: Map the corrected pitch sequence to a note sequence within the range of the guzheng's pitch range: S5: Using beat information and frame duration information, combined with human auditory characteristics, the notes are filtered to remove excessively short invalid notes, and rests are added to audio outside the guzheng range to obtain a regularized guzheng note sequence. S6: Based on the guzheng note sequence and the experimentally recorded guzheng source audio information, combined with the established rule base information and the characteristics of the guzheng instrument, feasible guzheng techniques are added to construct a preliminary guzheng melody. S7: The volume data information is corrected frame by frame using the RMS loudness envelope comparison method, making the final guzheng melody more natural.
2. The method according to claim 1, characterized in that, In step S1: S101: Select the appropriate glissando technique based on the semitone difference between adjacent notes and the direction of pitch change; S102: Select the appropriate tremolo treatment method based on the duration of the note; S103: Select the appropriate glissando treatment method based on the number of consecutive rests; S104: For glissando techniques that cannot be performed, select appropriate single-note playing methods as alternatives; S105: Based on the principle of twelve equal temperament, construct a formula for mapping the semitones of rising / falling tones.
3. The method according to claim 1, characterized in that, In step S3: S301: Based on the pitch sequence data information pitch1 obtained in step S2, the minimum sharp / flat semitone value corresponding to different humming audios is obtained through the formula of sharp / flat semitone mapping relationship, and then the most ideal sharp / flat value n is selected according to the auditory effect. S302: Based on the audio feature value information and pitch rise / fall value n in the pitch1 array, the frequency spectrum is resampled using frequency domain waveform level pitch transformation technology to obtain the audio after pitch rise / fall, and the fundamental frequency information is extracted again and stored in the pitch2 array; S303: Corrects excessive deviations from the theoretical up / down modulation ratio in pitch2 using a fundamental frequency ratio calibration function, ensuring the corrected fundamental frequency value matches the original fundamental frequency. The pitch relationship is further improved to enhance the accuracy of the boost / down fundamental frequency data in the time domain, and finally all fundamental frequencies are stored in the corrected_pitch array; S304: Detect the fundamental frequencies in the corrected_pitch array that are outside the range of the guzheng's pitch range, and set them all to 0; S305: Obtain the initial intensity envelope of the humming audio, get the beat position information of the humming audio, and store it in the beats_position array.
4. The method according to claim 1, characterized in that, In step S7: S701: Frame-level RMS loudness sequences are extracted from the time-domain audio signal data after pitch shifting / falling and the guzheng time-domain audio signal data, respectively, and frame-level energy analysis is performed. S702: Based on the possibility that the lengths of the two RMS sequences may be inconsistent, the resampling method is used to resample the RMS data information of the time-domain audio signal after pitch shifting / falling, so that its length is consistent with the RMS data information of the guzheng time-domain audio signal, thus obtaining an RMS loudness sequence with the same duration. S703: Calculate the gain adjustment coefficient frame by frame based on RMS loudness sequences of the same duration; S704: Perform amplitude weighting on each frame based on the obtained gain adjustment coefficient to obtain the corrected guzheng audio signal; S705: Saves the guzheng audio information file after pitch correction, completing the dynamic pitch correction strategy based on RMS loudness envelope matching.
5. The method according to claim 1, characterized in that, To address the issues of range mismatch and volume fluctuation mismatch between the original volume and the dynamics of the guzheng audio caused by glissando substitution in traditional humming-to-guzheng melody generation, a solution integrating multi-dimensional processing strategies is proposed. This solution introduces high-precision fundamental frequency extraction, improved spectrum resampling pitch shifting, and frame-level RMS loudness dynamic pitch correction strategy into the guzheng melody generation process.