Intonation adjusting method, equipment, product and medium

By analyzing the singing repertoire and breathing waveform signals, unstable intervals with insufficient breath support are identified, pitch change rate and deviation value are calculated, and pitch compensation is performed. This solves the problem of low pitch adjustment accuracy in existing technologies and achieves more precise audio signal correction.

CN121565112APending Publication Date: 2026-02-24HUBEI UNIV OF EDUCATION

Patent Information

Application Number
CN202511780811.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing pitch correction technology cannot effectively distinguish between a singer's intentional slow pitch gliding and unintentional pitch errors, resulting in low pitch correction accuracy.

Method used

By analyzing the singer's target repertoire to obtain the standard pitch sequence, and combining the singer's audio signal and breathing waveform signal, the breath support strength is determined, unstable intervals are identified, and the pitch change rate and deviation value are calculated. Based on these parameters, pitch compensation is performed to achieve precise adjustment of the audio signal.

Benefits of technology

It improves the accuracy of pitch adjustment, avoids over-adjustment of normal pitch fluctuations, and significantly enhances the correction effect of audio signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121565112A_ABST
    Figure CN121565112A_ABST
Patent Text Reader

Abstract

The invention discloses an intonation adjusting method and device, a product and a medium, and relates to the technical field of audio signal processing. In the method, a standard pitch sequence is determined based on a target singing track; acquiring an audio signal and a breathing waveform signal; determining an actual pitch sequence based on the audio signal; determining breath support strength based on the breath waveform signal and the audio signal; determining an unstable interval based on the breath support strength; determining the variation trend of the pitch frequency in the unstable interval; calculating pitch change rates of adjacent time points; determining a pitch deviation value based on the actual pitch sequence and the standard pitch sequence; when the absolute value of the pitch deviation value exceeds a preset deviation threshold value, the change trend presents a monotone increasing or monotone decreasing trend, and the pitch change rate is smaller than a preset rate threshold value, determining a pitch compensation parameter; and adjusting the audio signal according to the pitch compensation parameter to obtain a compensated audio signal. The method has the effect of improving the accuracy of intonation adjustment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio signal processing technology, specifically to a pitch adjustment method, device, product, and medium. Background Technology

[0002] With the widespread adoption of digital audio technology, optimizing vocal audio has become a standard procedure in music production, live sound reinforcement, and mass entertainment applications. Among these processes, pitch adjustment, commonly known as Auto-Tune, is a crucial step in ensuring singing quality. It effectively corrects pitch deviations caused by various factors, enhancing the listening experience.

[0003] A widely used pitch correction technique, such as the method and apparatus for correcting pitch deviations in audio content described in Chinese patent document CN108257613A, relies on analyzing the audio signal output by the singer. This technique first extracts the actual pitch curve over time from the audio signal using an algorithm, then compares it point-by-point with a preset musical score or standard pitch to identify segments whose pitch deviates from the target value. Once a deviation is detected, the system uses digital signal processing algorithms (such as PSOLA or a phase vocoder) to precisely correct the pitch of the audio segment, bringing it back to the correct pitch.

[0004] However, in singing, some slow pitch glissando is an intentional artistic technique by the singer (such as glissando), while other slow pitch glissando with very similar forms is an unintentional pitch error (such as pitch deviation). For existing technologies that can only analyze audio signals, these two situations present almost identical characteristics on the pitch curve, making it impossible for the system to distinguish between them, thus resulting in low accuracy of pitch adjustment. Summary of the Invention

[0005] This application provides a pitch adjustment method, device, product, and medium, which improves the accuracy of pitch adjustment.

[0006] The first aspect of this application provides a method for adjusting pitch, specifically including: Analyze the singer's target repertoire to obtain the standard pitch sequence; Acquire the singer's audio signal and breathing waveform signal; based on the audio signal, determine the actual pitch sequence, which contains pitch frequencies that change over time; Based on respiratory waveform signals and audio signals, the breath support intensity is determined; the time period when the breath support intensity is lower than the preset support intensity threshold is defined as the unstable interval. Within the unstable interval, determine the trend of pitch frequency variation at multiple consecutive time points; Calculate the pitch frequency difference between any two adjacent time points within a series of consecutive time points, and divide the pitch frequency difference by the time interval between the adjacent time points to obtain the pitch change rate. The pitch deviation value is determined based on the actual pitch sequence and the standard pitch sequence; When the absolute value of the pitch deviation exceeds the preset deviation threshold, the change trend shows a monotonically increasing or monotonically decreasing trend, and the pitch change rate is less than the preset rate threshold, the pitch compensation parameter is determined based on the pitch deviation value. The audio signal is adjusted according to the pitch compensation parameters to obtain the compensated audio signal.

[0007] By adopting the above technical solution, a standard pitch sequence is first obtained by analyzing the singer's target repertoire, providing an accurate target reference value for pitch adjustment. The singer's audio signal and breathing waveform signal are acquired, and the actual pitch sequence is extracted from the audio signal, achieving a complete record of pitch changes during the performance. The breath support strength is determined by combining the acquired breathing waveform signal and audio signal. Time periods where the breath support strength is lower than a preset support strength threshold are marked as unstable intervals, enabling precise location of pitch problem areas caused by insufficient breath support. For these unstable intervals, the pitch frequency variation trend at continuous time points is analyzed, and the pitch change rate between adjacent time points is calculated, accurately reflecting the pitch change pattern under insufficient breath support. The pitch deviation value is obtained by comparing the actual pitch sequence with the standard pitch sequence. When the pitch deviation value exceeds a preset deviation threshold, exhibits a monotonous variation trend, and the pitch change rate is small, corresponding pitch compensation parameters are determined, effectively avoiding over-adjustment of normal pitch fluctuations. Finally, the audio signal is adjusted according to the determined pitch compensation parameters to obtain the compensated audio signal. By accurately identifying pitch deviations caused by insufficient breath support and performing targeted compensation, the accuracy of pitch adjustment is significantly improved.

[0008] Optionally, the analysis of the singer's target repertoire to obtain a standard pitch sequence specifically includes: Obtain the sheet music data corresponding to the target song. The sheet music data includes the duration of the notes and the pitch markings. The duration of notes in the musical score data is converted into a timestamp sequence, which represents the start and end times of each note; Convert the pitch markings of notes in the musical score data into standard frequency values; Based on the timestamp sequence and standard frequency values, a time-frequency correspondence table is constructed, which records the standard frequency corresponding to each time point. Convert the time-frequency correspondence table into a standard pitch sequence.

[0009] By adopting the above technical solution, firstly, musical score data containing note durations and pitch markers is obtained based on the target singing piece, achieving a complete record of the standard singing content. The duration of notes in the musical score data is converted into a timestamp sequence representing the start and end times of each note, establishing precise time positioning of the notes. At the same time, the pitch markers in the musical score data are converted into specific standard frequency values, realizing the numerical representation of pitch. Furthermore, a time-frequency correspondence table is constructed based on the timestamp sequence and standard frequency values, accurately recording the standard frequency corresponding to each time point. Finally, the time-frequency correspondence table is converted into a standard pitch sequence, constructing continuous pitch benchmark data containing the time dimension, realizing the refined quantification of the singing pitch evaluation standard.

[0010] Optionally, converting the note pitch markers in the musical score data into standard frequency values ​​specifically includes: Obtain the pitch markings of musical notes. The pitch markings of musical notes are represented in the form of letters and numbers. The letters represent the note names and the numbers represent the octaves they are in. Select a reference tone as the frequency conversion reference point and record the reference frequency value corresponding to the reference tone; Calculate the number of semitone intervals between the pitch markings of a note and the reference pitch. The number of semitone intervals represents the pitch difference between the note and the reference pitch. Substituting the semitone intervals into the preset frequency conversion formula yields the standard frequency value corresponding to the note pitch mark.

[0011] By adopting the above technical solution, the pitch markings of musical notes, represented by letters and numbers, are first obtained to clarify the complete pitch information of the note name and its octave. A reference note is selected as the frequency conversion reference point and its corresponding reference frequency value is recorded to establish a reference scale for pitch conversion. By calculating the number of semitone intervals between the pitch markings of musical notes and the reference note, the pitch difference between the musical notes and the reference note is quantified. Finally, the number of semitone intervals is substituted into the frequency conversion formula based on the exponential relationship to obtain the standard frequency value corresponding to the pitch markings of musical notes. This achieves a precise mapping conversion from musical note symbols to actual frequencies, ensuring the accuracy of the pitch reference data in the frequency dimension.

[0012] Optionally, determining the breath support strength based on the respiratory waveform signal and audio signal specifically includes: Perform a short-time Fourier transform on the respiratory waveform signal to obtain the respiratory rhythm spectrum that varies with time; Determine the main respiratory frequency and its corresponding harmonic energy percentage from the respiratory rhythm spectrum; Time-frequency analysis of audio signals is performed to determine acoustic quality characteristics; Based on the main breathing frequency, the variation law of acoustic quality characteristics over time is analyzed to determine the amount of change of acoustic quality characteristics within the main breathing frequency cycle. By matching the change with the proportion of harmonic energy, the consistency of acoustic change is obtained. The intensity of breath support is determined based on the consistency of acoustic changes.

[0013] By adopting the above technical solution, the respiratory waveform signal is first subjected to short-time Fourier transform to obtain the respiratory rhythm spectrum, effectively capturing the dynamic characteristics of respiratory fluctuations. The main respiratory frequency and harmonic energy ratio are determined from the respiratory rhythm spectrum, achieving an accurate characterization of the regularity of breathing. Time-frequency analysis of the audio signal is performed to determine the acoustic quality characteristics, effectively characterizing the various dimensions of sound performance. Based on the main respiratory frequency, the time-varying law of acoustic quality characteristics is analyzed and its change within the main respiratory frequency period is determined, accurately reflecting the effect of respiratory regulation during vocalization. The change is matched with the harmonic energy ratio to obtain the consistency of acoustic change, accurately measuring the synchronicity of breathing and vocalization during singing. Finally, the breath support strength is determined based on the consistency of acoustic change, achieving a precise quantitative assessment of the singer's breath support ability.

[0014] Optionally, determining the variation trend of pitch frequency at multiple consecutive time points within the unstable interval specifically includes: The audio signal within the unstable region is processed by frame segmentation to obtain an audio frame sequence of fixed time length; Pitch extraction is performed on each audio frame in the audio frame sequence to obtain the pitch frequency corresponding to each audio frame; Arrange the pitch frequencies of multiple consecutive time points within the unstable interval in chronological order to form a pitch frequency sequence. Calculate the pitch frequency difference between adjacent time points in the pitch frequency sequence to obtain the pitch change sequence; The trend of pitch frequency variation is determined by the number of positive and negative terms in the pitch change sequence.

[0015] By adopting the above technical solution, the audio signal in the unstable interval is first processed into an audio frame sequence, avoiding the pitch extraction deviation caused by improper time window division. The pitch frequency of each audio frame is extracted to eliminate the non-periodic interference components in the audio signal. The pitch frequencies of multiple consecutive time points are arranged in chronological order to form a pitch frequency sequence, which effectively reduces the influence of random fluctuations caused by discrete sampling. The pitch frequency difference between adjacent time points is calculated to obtain the pitch change sequence, eliminating the cumulative error of pitch changes at different time scales. The change trend of pitch frequency is determined based on the number of positive and negative terms in the pitch change sequence, realizing the reliable identification of small pitch fluctuations caused by breath instability.

[0016] Optionally, determining the pitch compensation parameters based on the pitch deviation value specifically includes: The pitch deviation values ​​are segmented and mapped to compensation coefficients, which represent the proportion of pitch that needs to be adjusted. Calculate the frequency change gradient between adjacent time points in a pitch frequency sequence with multiple consecutive time points; The time weight of pitch compensation is determined based on the frequency change gradient; Multiplying the compensation coefficient by the time weight yields the pitch compensation parameters for each time point.

[0017] By adopting the above technical solution, the pitch deviation value is first mapped to compensation coefficients in segments, avoiding excessive pitch adjustment caused by linear compensation. The frequency change gradient of adjacent time points in the pitch frequency sequence of multiple consecutive time points is calculated, revealing the dynamic adjustment requirements of pitch change. The time weight of pitch compensation is determined according to the frequency change gradient, solving the problem of discontinuity of pitch compensation in the time dimension. The compensation coefficient is multiplied by the time weight to obtain the pitch compensation parameters corresponding to each time point, realizing a smooth transition of pitch compensation intensity on the time axis and effectively eliminating abrupt distortion in the compensation process.

[0018] In a second aspect, this application provides an electronic device for pitch adjustment, the electronic device comprising: one or more processors and a memory; the memory being coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, the one or more processors calling the computer instructions to cause the electronic device for pitch adjustment to perform the method described in the first aspect and any possible implementation thereof.

[0019] Thirdly, this application provides a computer program product containing instructions that, when run on an electronic device for pitch adjustment, cause the electronic device to perform the method described in the first aspect and any possible implementation thereof.

[0020] Fourthly, this application provides a computer-readable storage medium including instructions that, when executed on a pitch-adjusting device, cause the electronic device to perform the method described in the first aspect and any possible implementation thereof. Attached Figure Description

[0021] Figure 1 This is a schematic diagram of the architecture of a pitch adjustment system provided in an embodiment of this application; Figure 2 This is a schematic flowchart of a pitch adjustment method provided in an embodiment of this application; Figure 3 This is a schematic diagram illustrating pitch variation within an unstable range, provided in an embodiment of this application. Figure 4This is an exemplary hardware structure diagram of an electronic device for pitch adjustment provided in an embodiment of this application. Detailed Implementation

[0022] Figure 1 An exemplary system architecture for a pitch adjustment system is shown.

[0023] like Figure 1 As shown, the system architecture may include electronic device 11, network 12, and data acquisition device 13. Network 12 serves as the medium for providing a communication link between electronic device 11 and data acquisition device 13. Network 12 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0024] The singer can use electronic device 11 to interact with data acquisition device 13 via network 12 to receive acquired audio data and breathing waveform data, etc. Various audio processing applications, such as recording applications and karaoke applications, can be installed on electronic device 11.

[0025] Electronic device 11 is hardware and can be various electronic devices with data processing capabilities, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0026] The data acquisition device 13 can be a device used to acquire the singer's audio signals and breathing waveform signals, such as including a microphone and a breathing sensor. This device can acquire the singer's singing data in real time and transmit the acquired data to the electronic device 11 for pitch adjustment processing.

[0027] The following detailed explanation uses the electronic device side as an example.

[0028] This embodiment provides a method for adjusting pitch. Figure 2 This is a schematic flowchart of a pitch adjustment method provided in an embodiment of this application, as shown below. Figure 2 As shown, the method includes steps S101 to S108: S101: Analyze the singer's target repertoire to obtain the standard pitch sequence.

[0029] In this embodiment of the application, the target song refers to the musical work that the singer is preparing to sing, which includes complete melody information and rhythm information; the standard pitch sequence refers to the ideal pitch change trajectory obtained by analyzing the target song, which is used to represent the accurate pitch frequency value that should correspond to each time point.

[0030] Specifically, the electronic device first acquires the sheet music data corresponding to the target song. The sheet music data contains the duration information and pitch marking information of each note. The electronic device converts the note duration information recorded in the sheet music data into a timestamp sequence. The timestamp sequence records the start and end times of each note on the singing timeline. At the same time, it converts the pitch markings of the notes in the sheet music data into corresponding standard frequency values. The pitch markings of the notes are calculated to obtain accurate frequency values ​​through frequency conversion formulas. Then, the timestamp sequence and the standard frequency values ​​are matched and combined to construct a time-frequency correspondence table. The time-frequency correspondence table establishes the correspondence between each time point and the corresponding standard frequency. Finally, the time-frequency correspondence table is converted into a continuous standard pitch sequence. The standard pitch sequence completely describes the ideal pitch change trajectory of the target song throughout the entire singing process.

[0031] Based on the above embodiments, as an optional embodiment, the step of analyzing the singer's target repertoire and obtaining the standard pitch sequence may include steps S201 to S204: S201: Obtain the sheet music data corresponding to the target song. The sheet music data includes the duration of the notes and the pitch markings.

[0032] In this embodiment of the application, the musical score data represents the complete musical notation information of the target song, including the duration of notes and pitch markings. The duration of a note refers to the length of time each note needs to be sustained during performance, representing the beat length and time occupied by the note; the pitch markings indicate the specific position of each note in the scale, using a standard notation method of letters and numbers to identify the pitch attribute of the note.

[0033] Specifically, the electronic device obtains the sheet music data corresponding to the target song through a music database or song file system. The sheet music data stores the complete musical information of the target song in a structured format. The note duration part of the sheet music data records the length occupied by each note on the time axis, including the duration information of different time values ​​such as whole notes, half notes, quarter notes, and eighth notes. At the same time, the note pitch marking part of the sheet music data adopts the international standard pitch name notation, using letters such as C, D, E, F, G, A, and B to represent note names, combined with numbers to represent the octave position, forming a complete pitch marking system. The electronic device parses the data structure of the sheet music data and extracts the note duration information and note pitch marking information contained therein, providing basic data support for the subsequent construction of pitch sequences.

[0034] S202: Convert the duration of notes in the musical score data into a timestamp sequence, which represents the start and end times of each note.

[0035] In this embodiment of the application, the timestamp sequence represents a set of precise time stamps obtained by converting the duration of the notes, which is used to represent the specific time position of each note on the singing timeline.

[0036] Specifically, the electronic device converts the relative duration of each note in the score data into an absolute time length based on the tempo and time signature information of the target song. The electronic device calculates the start time of each note one by one according to the order in which the notes appear in the score. The start time of the first note is set as the zero moment when the singing begins, and the start time of subsequent notes is equal to the end time of the previous note. At the same time, the electronic device adds the start time of each note to the corresponding absolute time length to calculate the end time of each note. The electronic device arranges the start and end times of all notes in chronological order to form a complete timestamp sequence. The timestamp sequence accurately records the time boundary information of each note during the singing process.

[0037] S203: Convert the pitch markers of notes in the musical score data into standard frequency values; construct a time-frequency correspondence table based on the timestamp sequence and standard frequency values, which records the standard frequency corresponding to each time point.

[0038] In this embodiment, the standard frequency value represents the physical acoustic frequency corresponding to each note; the time-frequency correspondence table is a mapping table constructed by matching the timestamp sequence with the standard frequency value, used to represent the ideal frequency that should correspond to each time point during the performance.

[0039] Specifically, the electronic device acquires the pitch mark of each note in the musical score data. The pitch mark uses a standard format of letters and numbers, where the letters indicate the note name and the numbers indicate the octave range. The electronic device selects a standard reference pitch as the reference point for frequency conversion. The reference pitch corresponds to a fixed reference frequency value. The electronic device calculates the number of semitone intervals between each note's pitch mark and the reference pitch. The semitone interval quantifies the pitch difference between the note and the reference pitch. The electronic device substitutes the number of semitone intervals into an exponential frequency conversion formula and calculates the standard frequency value corresponding to each note through exponential operations. At the same time, the electronic device matches each time period recorded in the timestamp sequence with the standard frequency value of the corresponding note. Within the time range from the start time to the end time of each note, the time-frequency correspondence table records that all time points within that time period correspond to the same standard frequency value. The electronic device integrates the frequency mapping relationships of all time periods to construct a complete time-frequency correspondence table covering the entire singing process.

[0040] For example, in the musical score data, the pitch of the first note is marked as C4, the pitch of the second note as E4, and the pitch of the third note as G4. The electronic device selects the A4 note as the reference pitch. The electronic device calculates the number of semitone intervals of the C4 note relative to the A4 reference pitch. The electronic device substitutes the number of semitone intervals into the frequency conversion formula f = f0 × 2^(n / 12) for calculation. In the frequency conversion formula, f represents the target note frequency, f0 represents the reference pitch frequency, and n represents the number of semitone intervals. The electronic device calculates the standard frequency value corresponding to the C4 note through the frequency conversion formula. The same method is used to calculate the standard frequency values ​​corresponding to the E4 and G4 notes. The electronic device matches the time period of the first note with the standard frequency value of C4, the time period of the second note with the standard frequency value of E4, and the time period of the third note with the standard frequency value of G4. Finally, a time-frequency correspondence table is constructed. The time-frequency correspondence table records the standard frequency of C4 corresponding to each time point in the first time range, the standard frequency of E4 corresponding to each time point in the second time range, and the standard frequency of G4 corresponding to each time point in the third time range.

[0041] Based on the above embodiments, as an optional embodiment, the step of converting the pitch marks of musical notes in the score data into standard frequency values ​​may include steps S301 to S304: S301: Get the pitch markings of musical notes. The pitch markings of musical notes are represented in the form of letters and numbers. The letters represent the note names and the numbers represent the octaves they are in.

[0042] In this embodiment, the pitch markings represent a combination of symbols used to identify the specific position of a note in a scale. They are represented using the International Standard Notation for Pitch Names, in a letter-plus-number format. The letters represent the note name, and the numbers represent the octave. Specifically, the letters C, D, E, F, G, A, and B are used to identify the note's position within an octave; the numbers represent the octave, indicating the pitch range of the note; the note name indicates the basic pitch level of the note in the twelve-tone equal temperament system, and the octave represents the specific high and low range of the note within the entire pitch spectrum.

[0043] Specifically, the electronic device extracts the pitch mark of each note from the musical score data. The pitch mark is encoded using a combination of letters and numbers. The electronic device analyzes the letter part of the pitch mark, which is selected from the seven basic note names C, D, E, F, G, A, and B. Each letter represents a basic pitch class in the twelve-tone equal temperament. The electronic device identifies the numerical part of the pitch mark, which indicates the octave number of the note. Different numbers correspond to different pitch frequency ranges. The electronic device combines and verifies the analyzed letter note name information and numerical octave information to ensure the integrity and accuracy of the pitch mark. The electronic device arranges all the pitch mark information of the notes according to the order of the musical score to form a complete sequence of pitch marks.

[0044] S302: Select the reference tone as the frequency conversion reference point and record the reference frequency value corresponding to the reference tone.

[0045] In the embodiments of this application, the reference tone refers to a specific note that serves as a calculation reference standard during the frequency conversion process.

[0046] Specifically, the electronic device selects notes with standard frequency definitions from the twelve-tone equal temperament system according to the international standard pitch specification. The standard pitch specification usually designates specific notes as the unified reference for frequency calculation. The selected reference note needs to have a standard frequency value that is widely recognized in music theory. The electronic device obtains the accurate frequency value corresponding to the selected reference note in the international standard. The reference frequency value serves as the mathematical basis for all subsequent note frequency calculations. The note marking information of the reference note is bound and stored with the reference frequency value to establish a fixed correspondence between the reference note and the reference frequency value. This verifies that the selection of the reference note meets the requirements of the international pitch standard and ensures the accuracy and consistency of the frequency conversion calculation. The electronic device saves the reference note information and the reference frequency value information as the core parameters of the frequency conversion module.

[0047] For example, according to the International Standard Musical Notation (ISM) standard, the A4 note is selected as the reference note from the twelve-tone equal temperament system. The A4 note is defined in the ISM standard as a reference note with a fixed frequency value. The A4 note is located at the A note name position in the fourth octave. The accurate frequency value corresponding to the A4 reference note in the ISM standard is obtained. The reference frequency value is a fixed value specified by the ISM standard. The A4 note mark is bound to the reference frequency value, establishing a fixed relationship between the A4 reference note and the frequency conversion reference point. This verifies that the selection of the A4 reference note conforms to the ISM standard, and that the reference frequency value accurately reflects the standard definition of the A4 note. The A4 reference note information and the corresponding reference frequency value are saved. The A4 reference note, as the frequency conversion reference point, provides a unified calculation basis for the frequency calculation of all other notes.

[0048] S303: Calculate the number of semitone intervals between the pitch markings of a note and the reference pitch. The number of semitone intervals represents the pitch difference between the note and the reference pitch.

[0049] In this embodiment, the semitone interval represents the pitch difference between the target note and the reference note, used to quantify the relative positional relationship between the two notes in the twelve-tone equal temperament; the pitch difference refers to the positional distance of the note in the scale, calculated and represented in semitones.

[0050] Specifically, the electronic device extracts the letter name and number octave portion from the pitch markings of the target note, while simultaneously acquiring the letter name and number octave information of the reference note. It calculates the octave difference between the target note and the reference note, obtained by subtracting the number octave portions of the two notes. Each octave difference corresponds to a distance of twelve semitones. The device also calculates the note name difference within the same octave between the target note and the reference note, obtained by analyzing the relative positions of the letter names of the two notes in the twelve-tone equal temperament. The electronic device converts the octave difference into the corresponding number of semitones and the note name difference into the corresponding number of semitones. Finally, it sums the number of semitones corresponding to the octave difference and the number of semitones corresponding to the note name difference to obtain the total number of semitone intervals between the target note and the reference note. The number of semitone intervals can be positive or negative; a positive number indicates the target note is higher than the reference note, and a negative number indicates the target note is lower than the reference note.

[0051] S304: Substitute the semitone intervals into the preset frequency conversion formula to obtain the standard frequency value corresponding to the note pitch mark.

[0052] Specifically, the electronic device invokes a preset frequency conversion formula. This formula uses an exponential function to describe the relationship between pitch and frequency. The formula uses the reference frequency value of the reference tone as its calculation basis, performing an exponential operation using the twelfth root of two. The semitone intervals calculated in the previous steps are substituted into the formula as the power parameter of the exponential operation. In the exponential operation, the twelfth root of two equals the semitone intervals. The electronic device executes the exponential calculation process, multiplying the reference frequency value by the result to obtain the standard frequency value corresponding to the target note. This standard frequency value accurately reflects the frequency characteristics of the note's pitch marking in physical acoustics, verifying the rationality of the calculation results and ensuring that the standard frequency value conforms to the expected range of musical theory.

[0053] S204: Convert the time-frequency correspondence table into a standard pitch sequence.

[0054] In this embodiment, the standard pitch sequence represents the continuous pitch change trajectory obtained after converting the time-frequency correspondence table, and is used to describe the ideal pitch distribution of the target song on the entire time axis.

[0055] Specifically, the electronic device reads the mapping relationship between all time periods and standard frequency values ​​recorded in the time-frequency correspondence table. The time-frequency correspondence table contains frequency information within the time range corresponding to each note. The time periods in the time-frequency correspondence table are sorted according to the time axis order to ensure the temporal continuity of the pitch sequence. The standard frequency value within each time period is expanded to the frequency data of all time points within that time period, forming a frequency distribution at the time point level. The expanded frequency data is rearranged according to the time order to construct a complete frequency time series from the beginning to the end of the performance. The electronic device converts the frequency time series into a standard pitch sequence data format. The standard pitch sequence organizes the data in the form of time as the horizontal axis and frequency as the vertical axis, forming a standard pitch sequence that describes the ideal pitch change trajectory of the target performance piece.

[0056] For example, a time-frequency correspondence table records the standard frequency in the low-range for the first time period, the standard frequency in the middle-range for the second time period, and the standard frequency in the high-range for the third time period. The mapping relationship between the three time periods and their corresponding standard frequency values ​​is read from the time-frequency correspondence table. The first, second, and third time periods are arranged in chronological order. The standard frequency value in the low-range of the first time period is expanded to the frequency data at each time point within that time period; the standard frequency value in the middle-range of the second time period is expanded to the frequency data at each time point within that time period; and the standard frequency value in the high-range of the third time period is expanded to the frequency data at each time point within that time period. The frequency data of the three time periods are connected in chronological order to form a complete frequency-time sequence from the beginning to the end of the performance. This frequency-time sequence is then converted to a standard pitch sequence format, which displays the continuous pitch change trajectory from the low-range to the middle-range and then to the high-range.

[0057] S102: Acquire the singer's audio signal and breathing waveform signal; based on the audio signal, determine the actual pitch sequence, which contains pitch frequencies that change over time.

[0058] In this embodiment of the application, the audio signal represents the sound signal generated by the singer during the singing process, including acoustic feature information such as pitch, timbre, and volume; the breathing waveform signal refers to the physiological signal generated by the singer's breathing activities during the singing process, which is used to reflect the changes in breathing rhythm and breathing intensity.

[0059] Specifically, the electronic device acquires the audio signals generated by the singer during the performance through a data acquisition device. The audio signals record the singer's vocal characteristics in digital form. At the same time, the data acquisition device acquires the singer's breathing waveform signals, which reflect the singer's breathing activity. The acquired audio signals are then processed for pitch extraction. The pitch extraction process uses frequency domain analysis to identify the fundamental frequency component in the audio signal and converts the identified fundamental frequency component into the corresponding pitch frequency value. The pitch frequency values ​​are arranged according to the time sequence of the audio signal to form a continuous pitch frequency change sequence. The electronic device organizes the pitch frequency change sequence into a data structure of actual pitch sequence. The actual pitch sequence completely records the singer's true pitch change trajectory throughout the entire performance process, including the actual pitch frequency information corresponding to each time point.

[0060] S103: Determine the breath support intensity based on the respiratory waveform signal and audio signal; the time period when the breath support intensity is lower than the preset support intensity threshold is defined as the unstable interval.

[0061] In this embodiment, breath support intensity represents the degree of support the singer's respiratory system provides for vocalization during singing, and is used to quantify the correlation between breathing state and sound quality; the preset support intensity threshold is a critical value used to determine whether breath support is sufficient, serving as a standard to distinguish between stable and unstable singing states; the unstable interval represents the time period when breath support intensity is insufficient, used to identify key periods when the singer is prone to pitch problems.

[0062] Specifically, the electronic device performs spectral analysis on the acquired breathing waveform signal to extract the frequency and energy distribution characteristics of the breathing rhythm. Simultaneously, it performs acoustic quality analysis on the audio signal to obtain stability indicators and sound quality parameters. It then performs correlation analysis between the breathing rhythm characteristics and the audio acoustic quality characteristics to calculate the degree of consistency between the breathing state and the sound quality. The degree of consistency between breathing and sound reflects the support effect of breath on vocalization. The degree of consistency is converted into a quantitative value of breath support strength. The higher the value of breath support strength, the more sufficient the breath supports vocalization. The electronic device compares the calculated breath support strength with a preset support strength threshold to identify time points when the breath support strength is lower than the threshold. It combines consecutive time points of low support strength into unstable intervals, which identify key periods when the singer's breath support is insufficient.

[0063] Based on the above embodiments, as an optional embodiment, the step of determining the breath support strength based on the respiratory waveform signal and the audio signal may include steps S401 to S406: S401: Perform a short-time Fourier transform on the respiratory waveform signal to obtain the respiratory rhythm spectrum that varies with time.

[0064] In the embodiments of this application, the short-time Fourier transform represents a mathematical transformation method for segmented spectral analysis of time-domain signals, used to obtain frequency component information of the signal at different time periods.

[0065] Specifically, the electronic device segments the acquired respiratory waveform signal into segments according to a fixed time window length. Each time window contains a respiratory signal segment of a certain length. A window function is applied to the respiratory signal segment within each time window. The window function is used to reduce the spectral leakage effect caused by signal truncation. Fourier transform is performed on the windowed respiratory signal segment. The Fourier transform converts the time-domain respiratory signal into frequency-domain spectral information. The amplitude and phase characteristics of each frequency component are calculated to form the respiratory spectrum corresponding to that time window. The electronic device arranges the respiratory spectra of all time windows in chronological order to construct the time-frequency analysis results of the respiratory signal. The time-frequency analysis results show the variation law of the respiratory rhythm spectrum over time, forming a respiratory rhythm spectrum that varies with time. The respiratory rhythm spectrum reflects the frequency characteristic distribution of respiratory activity in different time periods.

[0066] S402: Determine the main respiratory frequency and the corresponding harmonic energy percentage from the respiratory rhythm spectrum.

[0067] In this embodiment, the primary breathing frequency represents the frequency component with the strongest energy in the respiratory rhythm spectrum, used to identify the basic rhythmic characteristics of the singer's breathing activity; the harmonic energy ratio refers to the proportion of energy occupied by the harmonic component of the primary breathing frequency in the entire spectrum, used to quantify the regularity and stability of the respiratory rhythm.

[0068] Specifically, the electronic device analyzes the respiratory rhythm spectrum that changes over time, identifies the frequency component with the largest energy amplitude in the spectrum, and the frequency component with the largest energy is the main breathing frequency. The main breathing frequency reflects the basic rhythm of the singer's breathing activity. The position of each harmonic frequency of the main breathing frequency is calculated. The harmonic frequency is an integer multiple of the main breathing frequency. The energy amplitude value corresponding to the main breathing frequency and its harmonic frequency positions in the respiratory rhythm spectrum is extracted, and the total energy of the main breathing frequency and harmonic frequencies is calculated. At the same time, the total energy of all frequency components in the respiratory rhythm spectrum is calculated. The electronic device divides the total energy of the main breathing frequency and harmonics by the total energy of the spectrum to obtain the harmonic energy ratio. The higher the harmonic energy ratio value, the stronger the regularity of the respiratory rhythm. The lower the value, the more irregular components are contained in the breathing activity.

[0069] S403: Perform time-frequency analysis on the audio signal to determine the acoustic quality characteristics.

[0070] In this embodiment, acoustic quality characteristics refer to quantitative indicators reflecting the quality of audio signal production, determined by a weighted sum of acoustic parameters including frequency stability, timbre purity, and amplitude consistency. Frequency stability represents the smoothness of the fundamental frequency trajectory of the audio signal, used to quantify the continuity and consistency of vocal pitch; timbre purity represents the energy ratio between harmonic and noise components in the audio signal, used to measure the clarity and purity of the timbre; and amplitude consistency represents the uniformity of the volume intensity variation of the audio signal, used to evaluate the stability and sustainability of the sound intensity.

[0071] Specifically, the electronic device performs a short-time Fourier transform on the singer's audio signal, decomposing the audio signal into time-varying spectral information. The time-frequency transformation result shows the frequency component distribution of the audio signal at different time points. The fundamental frequency trajectory is extracted from the time-frequency analysis result, and the frequency change amplitude between adjacent time points of the fundamental frequency trajectory is calculated. The variance of the frequency change amplitude is used as a quantitative index of frequency stability. The ratio of the total energy of the harmonic components in the audio signal to the energy of the entire frequency band is calculated, and the harmonic energy ratio is used as a quantitative index of timbre purity. The time-domain variation characteristics of the audio signal volume intensity are analyzed, and the standard deviation of the volume intensity variation is calculated. The reciprocal of the volume intensity standard deviation is used as a quantitative index of amplitude consistency. The frequency stability index is normalized and assigned a preset first weight coefficient, the timbre purity index is normalized and assigned a preset second weight coefficient, and the amplitude consistency index is normalized and assigned a preset third weight coefficient. The electronic device multiplies the three normalized parameters by their corresponding weight coefficients and then sums them to obtain the acoustic quality characteristics.

[0072] For example, a short-time Fourier transform is performed on the audio signal of a singer performing a melody to obtain the time-frequency spectrum of the audio signal. The fundamental frequency trajectory is extracted from the time-frequency spectrum, and the frequency difference between adjacent time points in the fundamental frequency trajectory is calculated. The variance of the frequency difference is used as a frequency stability index; the smaller the variance, the more stable the frequency. The ratio of the total energy of the harmonic components in the time-frequency spectrum to the total energy of the spectrum is calculated; the higher the ratio, the purer the timbre. The volume intensity variation of the audio signal is extracted, and the standard deviation of the volume intensity variation is calculated; the smaller the standard deviation, the more consistent the amplitude. The frequency stability index is normalized and assigned a weight coefficient of 0.5, the timbre purity index is normalized and assigned a weight coefficient of 0.3, and the amplitude consistency index is normalized and assigned a weight coefficient of 0.2. The sum of the three weight coefficients is 1.0. The three normalized indices are multiplied by the weights of 0.5, 0.3, and 0.2 respectively, and then added together to obtain the acoustic quality characteristic value. The acoustic quality characteristic value comprehensively reflects the overall quality level of the singer's voice.

[0073] S404: Based on the main breathing frequency, analyze the change law of acoustic quality characteristics over time and determine the amount of change of acoustic quality characteristics within the main breathing frequency cycle.

[0074] Specifically, the electronic device calculates the corresponding cycle length based on the dominant breathing frequency determined in the preceding steps. The cycle length of the dominant breathing frequency is equal to the reciprocal of the dominant breathing frequency value. The singing process is divided into continuous dominant breathing frequency cycle time periods, with the length of each time period equal to a complete dominant breathing frequency cycle. Time series data of acoustic quality characteristics are extracted within each dominant breathing frequency cycle time period. The numerical change pattern of acoustic quality characteristics within a single dominant breathing frequency cycle is analyzed, and the peak, trough, and trend of acoustic quality characteristics within the dominant breathing frequency cycle are identified. The difference between the maximum and minimum values ​​of acoustic quality characteristics within each dominant breathing frequency cycle is calculated, and the difference between the maximum and minimum values ​​is taken as the change amount within that dominant breathing frequency cycle. The electronic device statistically analyzes the change amount data of multiple dominant breathing frequency cycles, analyzes the distribution characteristics and statistical laws of the change amount, and determines the typical change amount of acoustic quality characteristics within the dominant breathing frequency cycle.

[0075] S405: Match the change with the proportion of harmonic energy to obtain the consistency of acoustic change.

[0076] In the embodiments of this application, acoustic variation consistency represents the degree of coordination between changes in vocal quality and respiratory rhythm, and is used to reflect the supporting effect of respiratory state on vocal stability.

[0077] Specifically, the electronic device acquires the changes in acoustic quality characteristics within multiple main breathing frequency cycles calculated in step S404, arranges the change values ​​of each cycle in the order of performance time to form a change data sequence, and simultaneously acquires the harmonic energy percentage values ​​corresponding to each main breathing frequency cycle, arranges them in the order of performance time to form a harmonic energy percentage data sequence, calculates the increase / decrease direction between adjacent values ​​in the change data sequence, calculates the increase / decrease direction between adjacent values ​​in the harmonic energy percentage data sequence, counts the number of times the two data sequences have the same increase / decrease direction at the same time position, divides the number of times the direction is consistent by the total number of comparisons to obtain the change direction matching degree, calculates the Pearson correlation coefficient between the change data sequence and the harmonic energy percentage data sequence, assigns weight coefficients to the absolute value of the correlation coefficient and the change direction matching degree respectively, and adds the weighted absolute value of the correlation coefficient and the weighted change direction matching degree to obtain the acoustic change consistency.

[0078] S406: Determine the breath support strength based on the consistency of acoustic changes.

[0079] In this embodiment of the application, breath support strength refers to the level of support provided by the respiratory system for vocal production, which is quantitatively evaluated through the consistency of acoustic changes.

[0080] Specifically, the electronic device uses the acoustic consistency value as the main input parameter for calculating breath support strength, establishing a mapping relationship between acoustic consistency and breath support strength. This mapping relationship is described by a mathematical function, which describes the conversion between the two. A higher acoustic consistency value indicates good coordination between breathing and vocalization, corresponding to stronger breath support strength. Conversely, a lower acoustic consistency value indicates poor coordination between breathing and vocalization, corresponding to weaker breath support strength. The electronic device substitutes the acoustic consistency value into the mapping function to calculate the corresponding breath support strength value. This breath support strength value quantifies the singer's respiratory system's ability to support vocalization within a specific time period.

[0081] Based on the above embodiments, as an optional embodiment, the step of determining the breath support strength based on the consistency of acoustic changes may include steps S501 to S505: S501: Obtain the acoustic change consistency within a preset time period to obtain a numerical sequence. Calculate the average value of the numerical sequence as the consistency benchmark value.

[0082] Specifically, the electronic device is set to a preset time period for calculating breath support strength. The preset time period covers the continuous time interval during the singing process. Within the preset time period, acoustic change consistency values ​​are collected at fixed time intervals. All the collected acoustic change consistency values ​​are arranged in chronological order to form a complete numerical sequence. The numerical sequence contains the acoustic change consistency values ​​corresponding to each sampling time point within the preset time period. All values ​​in the numerical sequence are summed, and the summation result is divided by the total number of values ​​in the numerical sequence to obtain the arithmetic mean of the numerical sequence. The arithmetic mean is used as the consistency benchmark value.

[0083] S502: The foundation support strength is obtained based on the preset proportional relationship between the consistency benchmark value and the foundation support strength.

[0084] Specifically, the electronic device acquires the consistency benchmark value calculated in the preceding steps, calls the preset proportionality parameter in the system, which includes a proportionality coefficient and a conversion function. The proportionality coefficient determines the numerical conversion ratio between the consistency benchmark value and the basic support strength. The consistency benchmark value and the proportionality coefficient are multiplied, and the result of the multiplication operation is the corresponding basic support strength value. The basic support strength value quantifies the singer's basic breathing support ability within the preset time period. The electronic device verifies the rationality of the basic support strength value to ensure that the value range meets the expected standard.

[0085] S503: Calculate the degree of deviation of each numerical point in the numerical sequence from the benchmark consistency level to obtain the consistency volatility.

[0086] Specifically, the electronic device calculates the difference between each numerical point in the numerical sequence and the consistency benchmark value. The difference calculation yields the deviation of each numerical point from the benchmark consistency level. The square value of all deviations is calculated, and all squared deviations are summed. The summation result is divided by the total number of numerical points in the numerical sequence to obtain the average value of the squared deviations. The electronic device then performs a square root operation on the average value of the squared deviations. The result of the square root operation is the consistency fluctuation. The larger the consistency fluctuation value, the more drastic the fluctuation of acoustic consistency. The smaller the value, the more stable the acoustic consistency.

[0087] S504: The correction coefficient is determined based on the consistency volatility, and the correction coefficient is inversely proportional to the consistency volatility.

[0088] Specifically, the electronic equipment uses an inverse proportional function form where the correction coefficient is equal to a preset constant divided by the consistency fluctuation. The preset constant is a pre-determined adjustment parameter of the system. The consistency fluctuation value is substituted into the inverse proportional function as the denominator for division, and the result of the division is the corresponding correction coefficient.

[0089] S505: Multiply the correction factor and the basic support strength to obtain the air support strength.

[0090] Specifically, the electronic device multiplies the correction coefficient value with the basic support strength value. The multiplication operation adjusts the basic support strength according to the proportion of the correction coefficient. When the correction coefficient value is greater than 1, the result of the multiplication operation is greater than the basic support strength, indicating that the breath support strength is enhanced. When the correction coefficient value is less than 1, the result of the multiplication operation is less than the basic support strength, indicating that the breath support strength is weakened. When the correction coefficient value is equal to 1, the result of the multiplication operation is equal to the basic support strength, indicating that the breath support strength remains unchanged. The electronic device uses the result of the multiplication operation as the final breath support strength. The breath support strength comprehensively reflects the singer's actual ability to support vocalization with the respiratory system within the preset time period.

[0091] S104: Within the unstable interval, determine the trend of pitch frequency variation at multiple consecutive time points.

[0092] Specifically, the electronic device extracts pitch frequency data at fixed time intervals within the unstable interval, forming a pitch frequency sampling sequence corresponding to multiple consecutive time points. The pitch frequency sampling sequence is arranged in chronological order, and the pitch frequency difference between adjacent time points is calculated. The pitch frequency difference reflects the amplitude and direction of pitch change between adjacent time points. The positive and negative distribution of all pitch frequency differences is statistically analyzed. When the difference is positive, it indicates that the pitch is rising; when the difference is negative, it indicates that the pitch is falling. The electronic device analyzes the ratio of positive to negative differences. When positive differences dominate, the trend is determined to be monotonically increasing; when negative differences dominate, the trend is determined to be monotonically decreasing; when the number of positive and negative differences is similar, the trend is determined to be fluctuating. This forms the result of judging the trend of pitch frequency change within the unstable interval.

[0093] Based on the above embodiments, as an optional embodiment, the step of determining the variation trend of pitch frequency at multiple consecutive time points within the unstable interval may include steps S601 to S605: S601: Perform frame segmentation on the audio signal within the unstable interval to obtain an audio frame sequence of fixed time length.

[0094] Specifically, the electronic device determines the start and end times of the unstable interval, extracts the complete audio signal data within that time range, sets a fixed time length parameter for the audio frames, which determines the amount of audio data contained in each audio frame, and starts cutting the audio signal sequentially according to the fixed time length from the start position of the unstable interval audio signal. The first audio frame contains audio data of a fixed time length starting from the start time, the second audio frame contains audio data of a fixed time length immediately following the first audio frame, and so on until the unstable interval audio signal ends. The electronic device numbers and arranges all the cut audio frames in chronological order to form a complete audio frame sequence, which contains audio data segments corresponding to all time periods within the unstable interval.

[0095] S602: Extract the pitch of each audio frame in the audio frame sequence to obtain the pitch frequency corresponding to each audio frame.

[0096] Specifically, the electronic device processes each audio frame in the audio frame sequence one by one, performing a Fast Fourier Transform (FFT) operation on a single audio frame. The FFT converts the time-domain audio frame signal into frequency-domain spectral information. The spectral information displays the energy distribution of each frequency component in the audio frame. The frequency component with the maximum energy amplitude is identified from the spectral information, and the frequency corresponding to the maximum energy amplitude is the fundamental frequency of the audio frame. The fundamental frequency value is recorded as the pitch frequency corresponding to the audio frame. The above processing is repeated until the pitch of all audio frames in the audio frame sequence has been extracted. The electronic device arranges the pitch frequencies of each audio frame according to the time order of the audio frames, forming a pitch frequency data set corresponding to the audio frame sequence. The pitch frequency data set records the pitch change information of each time period within the unstable interval.

[0097] S604: Calculate the pitch frequency difference between adjacent time points in the pitch frequency sequence to obtain the pitch change sequence.

[0098] Specifically, the electronic device calculates the difference starting from the second data point in the pitch frequency sequence. It subtracts the pitch frequency of the first time point from the pitch frequency of the second time point to obtain the first pitch frequency difference. It then subtracts the pitch frequency of the second time point from the pitch frequency of the third time point to obtain the second pitch frequency difference, and so on, calculating the pitch frequency difference between all adjacent time points. A positive difference indicates that the pitch has increased, a negative difference indicates that the pitch has decreased, and a zero difference indicates that the pitch remains unchanged. The electronic device arranges all the calculated pitch frequency differences in chronological order to form a pitch change sequence. The pitch change sequence records the point-by-point changes in pitch frequency within the unstable interval.

[0099] S605: Determine the trend of pitch frequency variation based on the number of positive and negative terms in the pitch change sequence.

[0100] Specifically, the electronic device categorizes data items with a difference greater than zero in the pitch change sequence as positive items, data items with a difference less than zero as negative items, and data items with a difference equal to zero as zero items. It counts the total number of occurrences of positive items and negative items in the pitch change sequence, comparing the number of positive and negative items. When the number of positive items is greater than the number of negative items, the pitch frequency is determined to show a monotonically increasing trend; when the number of negative items is greater than the number of positive items, the pitch frequency is determined to show a monotonically decreasing trend. When the number of positive and negative items is similar, the pitch frequency is determined to show a fluctuating trend, forming a judgment result on the pitch frequency change trend within the unstable interval.

[0101] S105: Calculate the pitch frequency difference between any two adjacent time points within a series of consecutive time points, and divide the pitch frequency difference by the time interval between the adjacent time points to obtain the pitch change rate.

[0102] Specifically, the electronic device identifies pitch frequency data from multiple consecutive time points within an unstable interval. It selects any two adjacent time points from the time series, calculates the pitch frequency of the latter time point minus the pitch frequency of the former time point to obtain the pitch frequency difference between these two adjacent time points, calculates the timestamp of the latter time point minus the timestamp of the former time point to obtain the time interval between adjacent time points, and divides the pitch frequency difference by the time interval. The result of the division operation is the pitch change rate between the adjacent time point pairs. The above calculation process is repeated for all adjacent time point pairs. The electronic device arranges the pitch change rates of all adjacent time point pairs in chronological order to form a complete pitch change rate sequence. The pitch change rate sequence records the speed distribution of pitch change within the unstable interval.

[0103] S106: Determine the pitch deviation value based on the actual pitch sequence and the standard pitch sequence.

[0104] Specifically, the electronic device aligns the actual pitch sequence with the standard pitch sequence along the time axis to ensure synchronization of the two sequences in the time dimension. It extracts the actual pitch frequency value and the standard pitch frequency value at each corresponding time point, calculates the difference between the actual pitch frequency and the standard pitch frequency, and the difference is the pitch deviation value at that time point. When the difference is positive, it means that the actual pitch is higher than the standard pitch; when the difference is negative, it means that the actual pitch is lower than the standard pitch; when the difference is zero, it means that the actual pitch is the same as the standard pitch. The above calculation process is repeated for all corresponding time points. The electronic device arranges the pitch deviation values ​​of all time points in chronological order to form a complete pitch deviation value sequence. The pitch deviation value sequence records the pitch deviation at each time point during the singing process.

[0105] S107: When the absolute value of the pitch deviation exceeds the preset deviation threshold, the change trend shows a monotonically increasing or monotonically decreasing trend, and the pitch change rate is less than the preset rate threshold, the pitch compensation parameter is determined based on the pitch deviation value.

[0106] Specifically, the electronic device checks whether the absolute value of the pitch deviation exceeds a preset deviation threshold. When the absolute value exceeds the threshold, it indicates that the pitch deviation has reached a level that requires compensation. At the same time, it checks whether the trend of change is monotonically increasing or monotonically decreasing. A monotonous trend indicates that there is a continuous pitch shift rather than normal musical techniques such as vibrato or glissando. It also checks whether the rate of pitch change is less than a preset rate threshold. A smaller rate indicates that the pitch change is slow rather than a fast musical ornament effect. When all three conditions are met, it indicates that there is a continuous pitch deviation problem caused by unstable breath, and pitch compensation calculation needs to be initiated. The pitch deviation value is used as the input parameter for the compensation calculation. The electronic device uses a proportional coefficient to convert the pitch deviation value into a pitch compensation parameter. The proportional coefficient determines the intensity and direction of the compensation. When the pitch deviation value is positive, the compensation parameter is negative to lower the pitch. When the pitch deviation value is negative, the compensation parameter is positive to raise the pitch. The absolute value of the compensation parameter is directly proportional to the absolute value of the deviation value.

[0107] Based on the above embodiments, as an optional embodiment, determining the pitch compensation parameter based on the pitch deviation value within the unstable range may include steps S701 to S704: S701: Maps pitch deviation values ​​into segments as compensation coefficients, where each compensation coefficient represents the percentage of pitch that needs to be adjusted.

[0108] Specifically, the electronic device presets multiple numerical ranges for pitch deviation values. Each range covers a specific range of deviation values. A corresponding compensation coefficient is assigned to each deviation value range. The magnitude of the compensation coefficient is directly proportional to the severity of the deviation value. The calculated pitch deviation value is compared with each preset range to determine the specific range to which the pitch deviation value belongs. The corresponding compensation coefficient is then found based on the range to which the pitch deviation value belongs. When the pitch deviation value is positive, the compensation coefficient is negative, which is used to reduce the pitch frequency. When the pitch deviation value is negative, the compensation coefficient is positive, which is used to increase the pitch frequency. The electronic device uses the found compensation coefficient as the pitch adjustment parameter at that point in time. The compensation coefficient directly reflects the direction and magnitude of the pitch adjustment that needs to be made.

[0109] S702: Calculate the frequency change gradient between adjacent time points in a pitch frequency sequence with multiple consecutive time points.

[0110] Specifically, the electronic device selects two adjacent time points from a series of consecutive pitch frequency points. It calculates the pitch frequency of the next time point minus the pitch frequency of the previous time point to obtain the pitch frequency change. It also calculates the timestamp of the next time point minus the timestamp of the previous time point to obtain the time interval. The pitch frequency change is then divided by the time interval to obtain the frequency change gradient of the adjacent time point pair. A positive gradient indicates the rate of increase of the pitch frequency, while a negative gradient indicates the rate of decrease of the pitch frequency. The larger the absolute value of the gradient, the faster the pitch frequency changes. The above calculation process is repeated for all adjacent time point pairs. The electronic device then arranges the frequency change gradients of all adjacent time point pairs in chronological order to form a complete frequency change gradient sequence. The frequency change gradient sequence records the distribution of the rate of change of pitch frequency in each time period.

[0111] S703: Determine the time weight of pitch compensation based on the frequency change gradient.

[0112] Specifically, the electronic device reads the frequency change gradient sequence calculated in the previous steps, calculates the absolute value of each frequency change gradient, eliminates the influence of positive and negative signs on the weight calculation, and performs an inverse proportional transformation by dividing the absolute value of the frequency change gradient by a preset constant. The inverse proportional transformation formula is that the basic value of the time weight is equal to the preset constant divided by the absolute value of the frequency change gradient. When the absolute value of the frequency change gradient increases, the division result decreases, and when the absolute value of the frequency change gradient decreases, the division result increases. The sum of the basic values ​​of the time weight at all time points is calculated, and the basic value of the time weight at each time point is divided by the sum for normalization calculation. The normalization formula is that the time weight is equal to the basic value of the time weight divided by the sum of all basic values. The electronic device ensures that the sum of all time weights is equal to 1 through normalization calculation, and uses the normalized value as the time weight corresponding to each time point.

[0113] Specifically, the electronic device calculates the absolute value of each frequency change gradient, eliminates the influence of positive and negative signs on the weight calculation, and performs an inverse proportional transformation by dividing a preset constant by the absolute value of the frequency change gradient. The inverse proportional transformation formula is that the basic value of the time weight equals the preset constant divided by the absolute value of the frequency change gradient. When the absolute value of the frequency change gradient is large, the division result is small, and when the absolute value of the frequency change gradient is small, the division result is large. The sum of the basic values ​​of the time weights at all time points is calculated, and the basic value of the time weight at each time point is divided by the sum for normalization calculation. The normalization formula is that the time weight equals the basic value of the time weight divided by the sum of all basic values. The electronic device ensures that the sum of all time weights equals 1 through normalization calculation, and uses the normalized value as the final time weight corresponding to each time point.

[0114] S704: Multiply the compensation coefficient by the time weight to obtain the pitch compensation parameters corresponding to each time point.

[0115] Specifically, the electronic device multiplies the compensation coefficient value and the time weight value at the corresponding time point. The multiplication operation adjusts the compensation coefficient according to the time weight ratio. When the time weight value is 1, the result of the multiplication operation is equal to the compensation coefficient, keeping the compensation intensity unchanged. When the time weight value is less than 1, the result of the multiplication operation is less than the compensation coefficient, weakening the compensation intensity. When the time weight value is greater than 1, the result of the multiplication operation is greater than the compensation coefficient, strengthening the compensation intensity. The electronic device uses the result of the multiplication operation at each time point as the pitch compensation parameter corresponding to that time point. The pitch compensation parameter comprehensively considers both the degree of pitch deviation and the smoothness of frequency change, providing accurate correction values ​​for subsequent audio signal frequency adjustment.

[0116] S108: Adjust the audio signal according to the pitch compensation parameters to obtain the compensated audio signal.

[0117] Specifically, the electronic device acquires audio signal segments that require pitch compensation in unstable regions. It performs a short-time Fourier transform on the audio signal segment, converting the time-domain signal into a frequency-domain spectral representation. This spectral representation displays the amplitude and phase information of each frequency component in the audio signal. It identifies the fundamental frequency component and harmonic components in the spectrum. The fundamental frequency component corresponds to the main pitch of the audio signal. The frequency offset is calculated based on the pitch compensation parameters, which are equal to the original frequency multiplied by the pitch compensation parameters. The fundamental frequency component and harmonic components are then frequency-shifted according to the frequency offset. This frequency shift changes the pitch characteristics of the audio signal. The electronic device then performs an inverse short-time Fourier transform on the corrected spectrum, converting the frequency-domain signal back into a time-domain audio signal, resulting in the compensated audio signal. The compensated audio signal possesses frequency characteristics that have been pitch-corrected.

[0118] After elaborating on the core idea of ​​this application—identifying unstable regions by introducing respiratory signals and determining unintentional pitch errors within these regions through multi-dimensional feature analysis—the following will combine the appendix... Figure 3 This will provide a more intuitive and detailed explanation of this judgment and adjustment process.

[0119] Figure 3 This is a schematic diagram of pitch variation within an unstable range provided by an embodiment of this application. The diagram uses time as the horizontal axis, pitch frequency as the primary vertical axis (left side), and breath support strength as the secondary vertical axis (right side) to visually illustrate the correlation between pitch variation and breath support strength variation, and thereby demonstrate how this application accurately identifies and handles unconscious pitch deviations caused by insufficient breath support.

[0120] Reference Figure 3 The diagram includes the following key visual elements: Standard pitch sequence (a): Represented by a horizontal solid line in the figure, it represents the constant target pitch frequency that the singer should maintain during a certain duration of a sustained note, as determined by the analysis of the target song.

[0121] Actual pitch sequence (b): The figure is represented by a continuously fluctuating curve, which represents the pitch frequency that changes over time as obtained by processing the singer's audio signal.

[0122] Breath support strength (c): Represented by another changing curve in the figure, this strength is calculated based on the singer's audio signal and breathing waveform signal, and is used to quantify the degree of stable support of breathing for vocalization.

[0123] Preset support strength threshold (d): Represented by a horizontal dashed line in the figure, this is a key judgment benchmark used to distinguish whether the breath support is sufficient or insufficient.

[0124] In a specific implementation scenario of this application, such as Figure 3 As shown, the entire judgment and adjustment process is as follows: First, the system detected that at time point t1, the curve of breath support strength (c) began to decline and crossed the preset support strength threshold (d), only recovering above the threshold at time point t2. Based on this, the system precisely marked the time period from t1 to t2 as the unstable interval.

[0125] Next, this application focuses on analyzing the variation characteristics of the actual pitch sequence (b) only within the marked unstable interval. It can be clearly observed from the figure that within this interval, the actual pitch sequence (b) exhibits a specific variation pattern that differs from normal singing techniques: Significance of the deviation: The vertical difference between the actual pitch sequence (b) and the standard pitch sequence (a), i.e., the pitch deviation value, shows a continuously increasing absolute value and significantly exceeds the preset deviation threshold within the interval. This indicates that the pitch problem has reached a level requiring intervention.

[0126] Monotonicity of the trend: The actual pitch sequence (b) changes in a monotonically decreasing trend (or, in other cases, monotonically increasing), rather than oscillating rapidly up and down the standard pitch. This unidirectional slip or deviation is a key feature that distinguishes it from conscious artistic manipulation (such as vibrato).

[0127] Controllable rate: The absolute value of the slope of the actual pitch sequence (b) is generally small, which means that its pitch change rate is slow and smooth, and the instantaneous values ​​of the change rate are all less than the preset rate threshold. This feature distinguishes it from the singer's intentional, usually faster, portamento.

[0128] When the system detects pitch changes within the unstable range that simultaneously meet the three core conditions of significant deviation, monotonous trend, and controlled rate, it will make a high-confidence judgment: the pitch change is due to an unconscious pitch error caused by insufficient breath support. Based on this judgment, the system will activate a compensation mechanism, determine the corresponding pitch compensation parameters according to the pitch deviation value, and precisely adjust the audio signal to bring it back to the standard pitch.

[0129] In summary, Figure 3 This vividly demonstrates the core idea of ​​the technical solution presented in this application: by introducing the analysis of breath support strength, a defined unstable range is created, and within this range, through multi-dimensional (deviation, trend, rate) feature analysis, the accurate identification of unintentional pitch deviations is achieved. This fundamentally solves the problem that existing technologies cannot distinguish between intentional artistic treatment and unintentional mistakes, thereby improving the accuracy and intelligence level of automatic pitch adjustment while ensuring the naturalness of the singing.

[0130] The following describes an exemplary electronic device for pitch adjustment provided in the embodiments of this application. Figure 4 This is an exemplary hardware structure diagram of an electronic device for pitch adjustment provided in an embodiment of this application.

[0131] In some embodiments, the pitch-adjusting electronic device is a computer device or includes a computer device. The computer device includes a processor, memory, and a network interface connected via a system bus. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the computer device stores data. The network interface of the computer device is used to communicate with other external terminals or servers via a network connection. In some embodiments, the network interface can be a wired network interface; in some embodiments, the network interface can also be a wireless network interface. When the computer program is executed by the processor, it implements the methods described in the embodiments of this application.

[0132] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0133] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A method for adjusting pitch, characterized in that, The method includes: Analyze the singer's target repertoire to obtain the standard pitch sequence; Acquire the singer's audio signal and breathing waveform signal; based on the audio signal, determine the actual pitch sequence, which contains pitch frequencies that change over time; Based on respiratory waveform signals and audio signals, the breath support intensity is determined; the time period when the breath support intensity is lower than the preset support intensity threshold is defined as the unstable interval. Within the unstable interval, determine the trend of pitch frequency variation at multiple consecutive time points; Calculate the pitch frequency difference between any two adjacent time points within a series of consecutive time points, and divide the pitch frequency difference by the time interval between the adjacent time points to obtain the pitch change rate. The pitch deviation value is determined based on the actual pitch sequence and the standard pitch sequence; When the absolute value of the pitch deviation exceeds the preset deviation threshold, the change trend shows a monotonically increasing or monotonically decreasing trend, and the pitch change rate is less than the preset rate threshold, the pitch compensation parameter is determined based on the pitch deviation value. The audio signal is adjusted according to the pitch compensation parameters to obtain the compensated audio signal.

2. The pitch adjustment method according to claim 1, characterized in that, The analysis of the singer's target repertoire yields a standard pitch sequence, specifically including: Obtain the sheet music data corresponding to the target song. The sheet music data includes the duration of the notes and the pitch markings. The duration of notes in the musical score data is converted into a timestamp sequence, which represents the start and end times of each note; Convert the pitch markings of notes in the musical score data into standard frequency values; Based on the timestamp sequence and standard frequency values, a time-frequency correspondence table is constructed, which records the standard frequency corresponding to each time point. Convert the time-frequency correspondence table into a standard pitch sequence.

3. The pitch adjustment method according to claim 2, characterized in that, The process of converting the pitch markers in the musical score data into standard frequency values ​​specifically includes: Obtain the pitch markings of musical notes. The pitch markings of musical notes are represented in the form of letters and numbers. The letters represent the note names and the numbers represent the octaves they are in. Select a reference tone as the frequency conversion reference point and record the reference frequency value corresponding to the reference tone; Calculate the number of semitone intervals between the pitch markings of a note and the reference pitch. The number of semitone intervals represents the pitch difference between the note and the reference pitch. Substituting the semitone intervals into the preset frequency conversion formula yields the standard frequency value corresponding to the note pitch mark.

4. The pitch adjustment method according to claim 1, characterized in that, The determination of breath support strength based on respiratory waveform signals and audio signals specifically includes: Perform a short-time Fourier transform on the respiratory waveform signal to obtain the respiratory rhythm spectrum that varies with time; Determine the main respiratory frequency and its corresponding harmonic energy percentage from the respiratory rhythm spectrum; Time-frequency analysis of audio signals is performed to determine acoustic quality characteristics; Based on the main breathing frequency, the variation law of acoustic quality characteristics over time is analyzed to determine the amount of change of acoustic quality characteristics within the main breathing frequency cycle. By matching the change with the proportion of harmonic energy, the consistency of acoustic change is obtained. The intensity of breath support is determined based on the consistency of acoustic changes.

5. The pitch adjustment method according to claim 4, characterized in that, The determination of breath support strength based on the consistency of acoustic changes specifically includes: Obtain the consistency of acoustic changes within a preset time period to obtain a numerical sequence; Calculate the average value of the numerical sequence as a consistency benchmark. The foundation support strength is obtained based on the preset proportional relationship between the consistency benchmark value and the foundation support strength. The degree of deviation of each numerical point in the numerical sequence from the benchmark consistency level is calculated to obtain the consistency volatility. The correction coefficient is determined based on the uniform volatility, and the correction coefficient is inversely proportional to the uniform volatility. Multiply the correction factor by the basic support strength to obtain the air support strength.

6. The pitch adjustment method according to claim 1, characterized in that, The determination of the variation trend of pitch frequency at multiple consecutive time points within the unstable interval specifically includes: The audio signal within the unstable region is processed by frame segmentation to obtain an audio frame sequence of fixed time length; Pitch extraction is performed on each audio frame in the audio frame sequence to obtain the pitch frequency corresponding to each audio frame; Arrange the pitch frequencies of multiple consecutive time points within the unstable interval in chronological order to form a pitch frequency sequence. Calculate the pitch frequency difference between adjacent time points in the pitch frequency sequence to obtain the pitch change sequence; The trend of pitch frequency variation is determined by the number of positive and negative terms in the pitch change sequence.

7. The pitch adjustment method according to claim 6, characterized in that, The determination of pitch compensation parameters based on pitch deviation values ​​specifically includes: The pitch deviation values ​​are segmented and mapped to compensation coefficients, which represent the proportion of pitch that needs to be adjusted. Calculate the frequency change gradient between adjacent time points in a pitch frequency sequence with multiple consecutive time points; The time weight of pitch compensation is determined based on the frequency change gradient; Multiplying the compensation coefficient by the time weight yields the pitch compensation parameters for each time point.

8. An electronic device for pitch adjustment, characterized in that, The electronic device includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code including computer instructions, and the one or more processors call the computer instructions to cause the electronic device to perform the method as described in any one of claims 1-7.

9. A computer program product containing instructions, characterized in that, When the computer program product is run on an electronic device for pitch adjustment, the electronic device performs the method as described in any one of claims 1-7.

10. A computer-readable storage medium comprising instructions, characterized in that, When the instruction is executed on a pitch-adjusting electronic device, the electronic device performs the method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Method and device for correcting pitch deviation of audio content

    CN108257613A

Cited By

  • Fitting system and method of original singing audio information, and computer storage medium

    CN121838696A