Live human voice tone adaptive adjustment method based on a network model

The network-based model addresses audio restoration challenges in live streaming by precisely identifying packet loss and adapting base frequency and resonance peak features for natural audio repair, improving quality and user experience.

CN120126494BActive Publication Date: 2025-07-15SHENZHEN MEIJIA PHOTOELECTRIC TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510603601.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-07-15
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

The prior art cannot effectively solve the packet loss problem of audio signals in live broadcast scenarios, resulting in poor discontinuity and naturalness of voice signals, and unstable repair effect, which cannot meet the real-time processing requirements of low latency.

Method used

Through a network model-based method, audio data is collected for frame analysis, packet loss boundaries are identified, and the fundamental frequency trajectory characteristics and formant peak characteristics are used for adaptive adjustment. Cubic spline interpolation and sliding window calculation are used to achieve high-precision repair with low latency.

Benefits of technology

High-precision and low-latency adaptive repair in high packet loss scenarios are achieved, and the sound quality and user experience in live broadcast scenarios are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126494B_ABST
    Figure CN120126494B_ABST
Patent Text Reader

Abstract

The present invention provides a method for adaptively adjusting the pitch of live human voices based on a network model, which relates to the technical field of audio signal processing. It determines the exact start position and end position of packet loss through the mutation points of the short-time energy function and the abnormal change points of the zero-crossing rate, and then specifically utilizes the fundamental frequency trajectory features and formant features to repair the packet loss interval. The method in the present invention can solve the problem that in a high packet loss rate scenario, the traditional voice repair method has a contradiction between the decoupled repair of acoustic parameters and the real-time processing constraints, resulting in broken fundamental frequency trajectories and unnatural voices, so as to achieve high-precision and low-latency adaptive repair of packet loss voices, and significantly improve the sound quality and user experience in the live broadcast scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio signal processing, and more specifically, to a method for adaptively adjusting the pitch of a live human voice based on a network model. Background Art

[0002] With the rapid development of Internet technology and multimedia communication technology, real-time audio and video live streaming has become one of the important content dissemination and social interaction methods. In the live streaming scenario, the audio quality is directly related to the user experience, and the integrity and clarity of the audio signal are the key factors to ensure the sound quality. However, the instability of network transmission often leads to the occurrence of audio packet loss. This packet loss problem will cause the discontinuity of the voice signal, thus significantly affecting the fluency and naturalness of the human voice pitch. To address this problem, existing technologies generally adopt audio data recovery algorithms for packet loss repair, including voice compensation algorithms based on linear predictive coding (LPC), compensation algorithms based on time-domain waveform interpolation, and voice reconstruction algorithms based on frequency-domain feature analysis, etc. These methods can repair the lost audio data to a certain extent. However, due to the complexity and real-time requirements of the audio signal, existing methods often have problems such as unstable repair effects, high latency, and inability to fully preserve the naturalness of the voice when dealing with packet-loss audio.

[0003] Currently, the existing technologies mainly face the following deficiencies: First, most traditional methods rely on fixed models or rules for audio repair and cannot adaptively adjust according to the characteristics of different voice signals, resulting in unnatural repaired audio signals, especially with poor effects when dealing with complex voices or voice segments with high dynamic changes; Second, existing methods lack accuracy in reconstructing the fundamental frequency and formants in the packet loss interval and cannot effectively capture the harmonic characteristics and spectral envelope features of the voice signal, resulting in a low timbre matching degree between the repaired voice and the original voice; Third, due to the high computational complexity of the voice signal recovery algorithm, traditional methods are difficult to meet the real-time processing requirements of low latency in practical applications. Especially in the live streaming scenario, the latency problem will significantly weaken the user's interaction experience. The existence of these problems poses higher requirements for the voice signal repair technology, that is, to achieve high-precision and adaptive repair of the audio signal while ensuring low latency. Summary of the Invention

[0004] To solve the above technical problems, the present invention provides a method for adaptively adjusting the pitch of a live human voice based on a network model, which can, to a certain extent, solve the problem of the fundamental frequency trajectory break caused by discontinuous packet loss in the high packet loss rate scenario of mobile live streaming (such as mountainous / metro environments) due to the contradiction between the decoupled repair of acoustic parameters and the real-time processing constraints of the traditional fundamental frequency linear interpolation method.

[0005] According to one aspect of the present invention, there is provided a method for adaptively adjusting the tone of a live person's voice based on a network model, which includes:

[0006] Collect the live person's voice audio data, frame the audio data with a frame length of 32 ms and detect the packet loss situation. Identify the packet loss boundary position through the energy envelope analysis of three consecutive audio frames, and mark the packet loss interval as a sequence to be repaired;

[0007] Perform a fast Fourier transform on the complete audio frames 64 m before and after the sequence to be repaired, and extract the fundamental frequency trajectory features and formant features;

[0008] Input the fundamental frequency trajectory features and the formant features into a pre-trained feature correlation network model to obtain the fundamental frequency-formant mapping relationship parameters;

[0009] Based on the mapping relationship parameters, calculate the fundamental frequency values of the sequence to be repaired using cubic spline interpolation, and synchronously adjust the formant bandwidth according to the fundamental frequency-formant mapping relationship parameters. Achieve low-latency processing through the incremental calculation method of a 60-ms sliding window.

[0010] Further, the identification of the packet loss boundary position determines the exact start position and end position of the packet loss through the mutation points of the short-time energy function and the abnormal change points of the zero-crossing rate.

[0011] Further, the determination of the mutation points includes:

[0012] Perform wavelet decomposition on the short-time energy function values, and select the third-layer approximation coefficients and detail coefficients to construct the energy contour features;

[0013] Calculate the local variance of the energy contour features, and adaptively divide the energy mutation interval according to the variance distribution using the OTSU algorithm;

[0014] When the changes in the energy contour features at certain moments fall into the energy mutation interval, these moments are regarded as the mutation points of the short-time energy function.

[0015] Further, the energy mutation interval smooths the variance curve through a sliding window and combines it with the variance distribution histogram. Determine the optimal segmentation threshold according to the between-class variance criterion, divide the variance curve into a mutation interval and a stable interval, and finally mark the effective energy mutation interval through region growing and morphological processing.

[0016] Further, the extraction of the fundamental frequency trajectory features and the formant features includes:

[0017] Perform a fast Fourier transform on the complete audio frames to obtain the spectral features;

[0018] In the first frequency band range of the spectral features, the cepstrum analysis method is used to extract the fundamental frequency candidate points, and the dynamic programming algorithm is used to connect the trajectories of the fundamental frequency candidate points;

[0019] Calculate the similarity cost of the fundamental frequency change between adjacent frames, and select the path with the minimum total cost as the fundamental frequency trajectory;

[0020] In the second frequency band range of the spectral features, the formant frequency positions are obtained by solving the characteristic equation of the linear prediction coefficients.

[0021] Furthermore, the fundamental frequency candidate points are extracted by preprocessing the spectral features and performing cepstrum analysis, and the fundamental frequency candidate points are selected. Then, based on the prominence, sharpness, and harmonic structure verification, the fundamental frequency candidate values with harmonic characteristics are selected.

[0022] Furthermore, the similarity cost includes the continuity cost of the fundamental frequency change and the similarity cost of the spectral envelope;

[0023] The continuity cost of the fundamental frequency change is dynamically evaluated by calculating the logarithmic difference between the fundamental frequency candidate points of the current frame and the previous frame and non-linearly mapping the fundamental frequency change rate, combined with the sliding-updated standard deviation of the fundamental frequency;

[0024] The similarity cost of the spectral envelope is calculated by sampling the harmonic sequences of the fundamental frequency candidate points of the current frame and the previous frame and reconstructing the spectral envelope curve, calculating the Pearson correlation coefficient within the specified frequency band range, and normalizing the arccosine value of the correlation coefficient to between 0 and 1 as the similarity cost of the spectral envelope.

[0025] Furthermore, the feature correlation network model represents the fundamental frequency trajectory and formant features as time series, extracts the local change patterns of the fundamental frequency through convolution, fully connects and maps the formant features, and constructs a correlation matrix by combining the fundamental frequency mutation points and formant parameter differences;

[0026] Use singular value decomposition to extract the main patterns, generate a piecewise linear mapping function of the fundamental frequency - formant, and predict the formant parameter adjustment amount according to the fundamental frequency change.

[0027] Furthermore, the piecewise linear mapping function of the fundamental frequency - formant dynamically calculates the formant parameter adjustment amount according to the fundamental frequency change amount, and smooths the transition and limits the saturation value through linear or non-linear mapping. At the same time, it constrains the adjusted parameters to meet the physical validity to achieve the collaborative dynamic update of the fundamental frequency and formant features.

[0028] Furthermore, a cubic spline interpolation function is constructed by constraining the boundary fundamental frequency values and their slopes, and combined with the weighted average fundamental frequency values of the intermediate control points, the fundamental frequency values of the sequence to be repaired are calculated point by point.

[0029] Compared with the prior art, the live human voice pitch adaptive adjustment method based on a network model provided by the present invention determines the exact start position and end position of packet loss through the mutation points of the short-time energy function and the abnormal change points of the zero-crossing rate, and thus specifically utilizes the fundamental frequency trajectory features and formant features to repair the packet loss interval. In this way, it can solve the problem that the traditional voice repair method has contradictions between acoustic parameter decoupling repair and real-time processing constraints in high packet loss rate scenarios, resulting in broken fundamental frequency trajectories and unnatural voices, thereby achieving high-precision and low-latency adaptive repair of packet loss voices and significantly improving the voice quality and user experience in live broadcast scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. In the drawings:

[0031] Figure 1 It is a flowchart of the live human voice pitch adaptive adjustment method based on a network model according to an embodiment of the present invention.

[0032] Figure 2 It is a flowchart of determining the packet loss position in the live human voice pitch adaptive adjustment method based on a network model according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] Next, exemplary embodiments of the present invention will be described in detail with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments of the present invention. It should be understood that the present invention is not limited by the exemplary embodiments described herein.

[0034] Figure 1 It is a flowchart of the live human voice pitch adaptive adjustment method based on a network model according to an embodiment of the present invention. As Figure 1 shown, in the live human voice pitch adaptive adjustment method based on a network model, it includes:

[0035] S1: Collect live human voice audio data, frame the audio data with a frame length of 32 ms, detect the packet loss situation, identify the packet loss boundary position through the energy envelope analysis of three consecutive audio frames, and mark the packet loss interval as a sequence to be repaired;

[0036] Collect live human voice audio data with a sampling rate of 48 kHz and a sampling precision of 16 bit for the audio data; perform frame splitting on the audio data, where the frame splitting uses a frame length of 32 ms and a frame shift of 16 ms, so that there is a 50% overlapping interval between adjacent frames to ensure signal continuity; perform packet integrity detection on each audio frame, and the integrity detection includes verifying the packet header identification bit and the packet tail synchronization bit; when a packet loss is detected, perform energy envelope analysis on the 3 frames forward and 3 frames backward of the audio frame, and the energy envelope analysis includes calculating the short-time energy function and the zero-crossing rate, and determining the exact start position and end position of the packet loss through the mutation points of the short-time energy function and the abnormal change points of the zero-crossing rate; mark the audio data between the start position and the end position as the sequence to be repaired, and the time length range of the sequence to be repaired is 50 ms to 300 ms; add repair mark information to the audio data according to the start position and the end position, and the mark information includes the time stamp of the packet loss, the duration, and the numbers of adjacent complete frames.

[0037] Determine the exact start position and end position of the packet loss through the mutation points of the short-time energy function and the abnormal change points of the zero-crossing rate as Figure 2 shown:

[0038] Perform windowing processing on the audio frame using a Hamming window of 20 ms, calculate the short-time energy function value, perform wavelet decomposition on the short-time energy function value, and select the 3rd layer approximation coefficients and detail coefficients to construct the energy profile feature; calculate the local variance of the energy profile feature, adaptively divide the energy mutation interval using the OTSU algorithm according to the variance distribution, and when the change of the energy profile feature falls into the energy mutation interval, these moments are regarded as the mutation points of the short-time energy function; extract the zero-crossing rate sequence of the audio frame, and perform Hilbert transform on the zero-crossing rate sequence to obtain the instantaneous phase information; calculate the continuity index of the instantaneous phase information, and the continuity index is characterized by the eigenvalue of the covariance matrix of the phase angles between adjacent frames; when it is detected that the energy profile feature has a mutation and the continuity index is lower than the dynamic reference value in the current acoustic environment, mark this moment as a candidate boundary point; construct the probability distribution of a three-frame time window for the candidate boundary point, and determine the exact packet loss boundary position within the time window based on the maximum a posteriori probability criterion; the dynamic reference value is adaptively updated according to the statistical characteristics of the previous 100 frames to eliminate the interference of environmental noise and acoustic scene changes; the area between the start position and the termination position is the packet loss interval that needs to be repaired.

[0039] Among them, the steps of dividing the energy mutation interval are as follows:

[0040] Calculate the sliding variance sequence of the energy profile features, perform mean smoothing on the variance sequence using a sliding window of 64 ms to obtain a smoothed variance curve; equally divide the numerical range of the variance curve into 256 discrete levels, and count the number of pixels at each level to obtain a variance distribution histogram; traverse the level values in the variance distribution histogram, and use the traversed points as segmentation thresholds to divide the variance distribution into a foreground region and a background region; calculate the probability and mean of the foreground region, where the foreground region corresponds to a mutation interval with a larger variance value, calculate the probability and mean of the background region, where the background region corresponds to a stable interval with a smaller variance value; based on the between-class variance criterion of the foreground region and the background region, calculate the discrimination index under the current segmentation threshold, and the discrimination index is the ratio of the between-class variance to the total variance; when the discrimination index reaches the maximum value, determine the corresponding segmentation threshold as the optimal segmentation threshold; divide the variance curve into a mutation interval and a stable interval according to the optimal segmentation threshold, and the sampling points within the mutation interval are potential packet loss positions; perform region growing on the sampling points within the mutation interval, and when the variance difference between adjacent sampling points is less than 20% of the variance mean of the current mutation interval, merge these sampling points into the same mutation event; perform morphological processing on the mutation event, and if the duration of the mutation event is within the range of 50 ms to 300 ms and its variance peak value is not less than 3 times the mean of the background region, mark the mutation event as a valid energy mutation interval.

[0041] It should be noted that the zero-crossing rate sequence serves as a supplementary verification mechanism for energy features in audio signal packet loss detection, and its function is based on the following principle: when the audio signal is transmitted normally, the zero-crossing rate sequence exhibits a continuous change law that conforms to the pronunciation characteristics of the speech signal; when a data packet is lost, the zero-crossing rate sequence will show a jump characteristic that does not conform to natural speech; the zero-crossing rate sequence and the energy mutation interval form a cross-verification relationship, specifically manifested as: when an energy mutation interval is detected, the zero-crossing rate sequence in the corresponding time period will necessarily show significant abnormalities. If there is only an energy mutation while the zero-crossing rate sequence remains continuous, it may be due to normal speech behaviors such as changes in the speaker's pronunciation intensity; conversely, when the zero-crossing rate sequence shows abnormalities, by checking the energy characteristics in the corresponding time period, the zero-crossing rate fluctuations caused by noise or electrical signal interference can be filtered out; in signal boundary localization, the change of the zero-crossing rate sequence often appears earlier than the energy change, and this characteristic can be used to improve the time accuracy of detecting the start position of packet loss; in the signal recovery stage, the cross-verification result of the zero-crossing rate sequence and the energy mutation interval determines the signal repair processing strategy: when both indicate packet loss, a complete fundamental frequency-formant collaborative repair scheme is adopted, and when only the energy feature is abnormal, only the signal amplitude is smoothly transitioned.

[0042] S2: Perform fast Fourier transform on the complete audio frames of 64 ms before and after the sequence to be repaired, and extract the fundamental frequency trajectory features and formant features. The fundamental frequency trajectory features include the fundamental frequency period length and phase information, and the formant features include the formant center frequency and bandwidth parameters;

[0043] Select complete audio frames of 64 ms each for the forward and backward directions of the sequence to be repaired. Perform 1024-point fast Fourier transform on the complete audio frames to obtain spectral features, and the frequency resolution of the spectral features is 46.875 Hz; perform logarithmic amplitude normalization on the spectral features and map the amplitude values to the range of 0-1; in the frequency band range of 50 Hz to 400 Hz of the spectral features, use cepstrum analysis method to extract fundamental frequency candidate points, and use dynamic programming algorithm to connect the trajectories of the fundamental frequency candidate points, calculate the continuity cost of the fundamental frequency change between adjacent frames and the similarity cost of the spectral envelope, and select the path with the minimum total cost as the fundamental frequency trajectory; in the frequency band range of 300 Hz to 3400 Hz of the spectral features, use linear prediction analysis to extract formant features, and the order of the linear prediction analysis is 12; obtain the formant frequency positions by solving the characteristic equation of the linear prediction coefficients, and determine the bandwidth and amplitude of each formant based on spectral peak search; sort the obtained formants according to the frequency size, and extract the first three formants as the vocal tract resonance features; calculate the change trends of the frequencies, bandwidths and amplitudes of the formants between the forward frames and the backward frames, and the change trends are characterized by the first-order difference sequences of the characteristic parameters.

[0044] Among them, the method of using cepstrum analysis to extract fundamental frequency candidate points includes:

[0045] Preprocess the spectral characteristics; calculate the logarithmic magnitude of the spectrum to compress the spectral energy into a similar dynamic range; perform an inverse fast Fourier transform on the logarithmic spectrum to obtain a complex cepstrum sequence, where the independent variable of the complex cepstrum sequence is the cepstral coefficient; take the real part of the complex cepstrum sequence to obtain a real cepstrum; search for peak points within the time delay range of 2.5 ms to 20 ms in the real cepstrum, and this time delay range corresponds to the fundamental frequency range of 50 Hz to 400 Hz; when a cepstrum peak is detected, calculate the prominence of this peak point, and the prominence is characterized by the sum of the amplitude differences between this point and its three sampling points on the left and right; if the prominence is greater than twice the cepstrum mean of this frame, record this peak point as a fundamental frequency candidate point; when multiple peak points meeting the prominence requirement are detected within the time delay range, calculate the second-order derivative of the cepstrum at each peak point; sort the candidate points according to the product of the prominence and sharpness of the peak, and select the three candidate points with the largest product value as the fundamental frequency candidate values for the current frame; at the frequency positions corresponding to the fundamental frequency candidate values, verify whether there is an obvious harmonic structure in the original spectrum. When the amplitudes of the second and third harmonics of a certain fundamental frequency candidate value in the spectrum are both greater than 1.5 times the average amplitude of this frequency band, preferentially retain this candidate value.

[0046] The specific steps for calculating the continuity cost of the fundamental frequency change and the similarity cost of the spectral envelope between adjacent frames are as follows:

[0047] Calculation of the continuity cost of the fundamental frequency change: Take the frequency value Fi of the i-th fundamental frequency candidate point in the current frame and the frequency value Fj of the j-th fundamental frequency candidate point in the previous frame, and calculate the logarithmic difference between the two frequency values; divide the logarithmic difference by the time interval between the two frames to obtain the fundamental frequency change rate; perform a non-linear mapping on the fundamental frequency change rate using the hyperbolic tangent function. The input of the mapping function is the fundamental frequency change rate divided by the fundamental frequency standard deviation, and the output range is between 0 and 1; use the output value of the mapping function as the continuity cost of the fundamental frequency change. When the fundamental frequency change rate is close to 0, the continuity cost is small, and when the fundamental frequency change rate increases, the continuity cost increases non-linearly; perform a sliding update on the fundamental frequency standard deviation for 10 consecutive frames to enable the calculation of the continuity cost to adapt to the changing characteristics of the speech signal.

[0048] Calculation of the similarity cost of the spectral envelope: Taking the frequency value Fi of the i-th fundamental frequency candidate point in the current frame as a reference, sample the original spectrum at integer multiples of this frequency to obtain the harmonic sequence Hi; perform cubic spline interpolation on the harmonic sequence to reconstruct the spectral envelope curve Ei of the current frame; use the same method to obtain the spectral envelope curve Ej corresponding to the j-th fundamental frequency candidate point in the previous frame; logarithmically equally sample the spectral envelope curve in the frequency band from 300 Hz to 3400 Hz to obtain the envelope sampling sequence; calculate the Pearson correlation coefficient of the two envelope sampling sequences, and the correlation coefficient reflects the similarity degree of the envelope shape; divide the arccosine value of the correlation coefficient by π as the similarity cost of the spectral envelope, so that the cost value is mapped between 0 and 1.

[0049] Constructing the total cost function: Combining the continuity cost and the similarity cost using an adaptive weight mechanism, and the adaptive weight is determined by the logarithmic function of the ratio of the spectral energy of the current frame to that of the previous frame; when the ratio of spectral energy is close to 1, the weights of the continuity cost and the similarity cost are close to equal; when the spectral energy changes greatly, increase the weight of the similarity cost and decrease the weight of the continuity cost; perform exponential smoothing on the adaptive weight to avoid discontinuity of the fundamental frequency trajectory caused by sudden changes in the weight.

[0050] More specifically, the total cost function can be expressed by the following formula:

[0051] where,

[0052]

[0053] where, is the frequency value of the i-th fundamental frequency candidate point in the current frame; is the frequency value of the j-th fundamental frequency candidate point in the previous frame; is the standard deviation of the fundamental frequency at the current moment, calculated through a 10-frame sliding window; is the spectral envelope and is the Pearson correlation coefficient of; is the spectral energy of the current frame; is the spectral energy of the previous frame; is the adaptive weight at time t; is the exponential smoothing coefficient, with a value of 0.9; is the serial number of the current frame; is the serial number of the starting frame.

[0054] It should be noted that although the above description outlines the general process of calculating the cost between adjacent frames, the specific implementation details need to be adjusted according to different speech characteristics and application scenarios. For example:

[0055] In the calculation of continuity cost, the speaker's speech features will significantly affect the selection of weight coefficients. For read speech, since the speech rate is relatively stable and the fundamental frequency changes are relatively smooth, more attention can be paid to the weight of continuity cost; while for natural conversations, where the speech rate and tone change greatly, the weight of continuity constraint needs to be appropriately reduced to avoid erroneously smoothing out normal fundamental frequency jumps. When calculating the spectral envelope similarity, the speech acquisition environment will affect the feature extraction strategy. When processing speech recorded in a quiet environment, the spectral envelope features are clear and reliable and can be directly used for similarity calculation; while in the presence of background noise, the spectral envelope needs to be smoothed first to highlight the main formant features before similarity calculation. In the optimization of fundamental frequency tracking, the speaker's pronunciation habits will also affect the parameter settings of the algorithm. For professional announcers with standardized pronunciation, whose fundamental frequency changes are regular, relatively strict tracking criteria can be adopted; while for ordinary speakers, especially in informal conversations, the tracking constraints need to be appropriately relaxed to adapt to more natural speech features.

[0056] These adjustments and optimizations are all to ensure the best fundamental frequency tracking effect in different application scenarios. For example, in a speech synthesis system, a smooth fundamental frequency trajectory is crucial for generating natural speech; while in an emotion recognition system, the detailed features of fundamental frequency changes need to be retained to accurately analyze the tone. Therefore, flexibly adjusting the cost calculation strategy according to specific application requirements is the key to improving the quality of speech processing.

[0057] S3: Input the fundamental frequency trajectory features and the formant features into a pre-trained feature correlation network model to obtain fundamental frequency-formant mapping relationship parameters, where the mapping relationship parameters characterize the modulation effect of fundamental frequency changes on formant features;

[0058] First, represent the input fundamental frequency trajectory features as a sequence of frequencies varying with time; represent the formant features as a sequence of three-dimensional feature vectors varying with time, including the fundamental frequency value and its first-order difference features at each time point, and including the center frequency, bandwidth, and amplitude information of each formant; when the input features enter the network, the fundamental frequency features extract local frequency change patterns through a forward convolutional layer with a window size of 5 and a stride of 1; the formant features are mapped to the same feature space as the fundamental frequency features through a fully connected layer; for the fundamental frequency feature sequence, calculate the frequency change gradient between adjacent time points, and mark it as a mutation point when the frequency change gradient is greater than a threshold; for the formant features, calculate the degree of difference of each formant parameter before and after the mutation point; based on the frequency mutation point and the formant parameter difference, construct a feature correlation matrix, where the rows of the matrix represent the fundamental frequency change intervals and the columns represent the formant parameter types; when significant changes occur in the fundamental frequency features, analyze the response characteristics of the formant parameters. If the changes in the formant parameters are consistent with the fundamental frequency changes, record a positive correlation coefficient at the corresponding position in the correlation matrix; if the changes in the formant parameters are in the reverse relationship with the fundamental frequency changes, record a negative correlation coefficient; perform singular value decomposition on the correlation matrix to extract the main correlation patterns, where the main patterns correspond to the eigenvectors with larger singular values; construct a fundamental frequency-formant mapping function based on the eigenvectors, and the mapping function adopts a piecewise linear interpolation form; when the fundamental frequency changes, predict the adjustment amount of the formant parameters according to the correlation matrix, and the adjustment amount is proportional to the product of the fundamental frequency change amount and the correlation coefficient; through the mapping function, the modulation effect of the fundamental frequency change on the formant parameters can be obtained for the subsequent signal reconstruction process.

[0059] The specific steps for constructing the fundamental frequency-formant mapping function based on the eigenvectors are as follows:

[0060] Sort the feature vectors according to the singular value magnitudes, and select the first three feature vectors as the main correlation patterns; Recombine the feature vectors into a mapping matrix, where the rows of the matrix represent the fundamental frequency change intervals and the columns represent the formant parameters; Normalize each row of the mapping matrix so that the sum of the mapping coefficients corresponding to each fundamental frequency change interval is 1; Construct a piecewise linear mapping function based on the mapping matrix, with each piece corresponding to a fundamental frequency change interval; When the fundamental frequency change falls within a certain interval, calculate the adjustment amount of the formant parameters through the mapping coefficients corresponding to that interval; If the fundamental frequency change is at the junction of adjacent intervals, use linear interpolation for smooth transition to avoid sudden changes in parameter adjustment; For each formant parameter, establish an independent mapping function, where the independent variable of the mapping function is the fundamental frequency change amount and the dependent variable is the adjustment amount of the parameter; When the fundamental frequency change amount is small, use a linear mapping relationship where the adjustment amount is proportional to the fundamental frequency change amount; When the fundamental frequency change amount exceeds a preset threshold, use a non-linear mapping relationship and limit the saturation upper limit of the adjustment amount through a sigmoid function; Constrain the output of the mapping function to ensure that the adjusted formant parameters satisfy physical validity, such as frequencies greater than zero and bandwidths being positive; If the adjustment of a certain formant parameter results in an unreasonable result, correct it based on the historical mean of that parameter; Apply the mapping function to each time point of the fundamental frequency trajectory to dynamically update the formant parameters and achieve the co-variation of the fundamental frequency and formant features. More specifically, the mapping function can be expressed by the following formula:

[0061]

[0062] where,

[0063] where, is the fundamental frequency change amount, is the linear mapping coefficient of the i-th interval, is the preset threshold, A is the amplitude coefficient of the sigmoid function, is the non-linear modulation coefficient, is the smooth transition coefficient, and are the boundary values of adjacent intervals, is the historical mean correction term for this interval.

[0064] S4: Based on the mapping relationship parameters, use cubic spline interpolation to calculate the fundamental frequency values of the sequence to be repaired, and synchronously adjust the formant bandwidth according to the fundamental frequency-formant mapping relationship parameters, and achieve low-latency processing through the incremental calculation method of a 60ms sliding window.

[0065] Detect the fundamental frequency values of the forward and backward boundary points of the sequence to be repaired, and use these two fundamental frequency values as the endpoint constraints for cubic spline interpolation; calculate the first derivative of the fundamental frequency value at the endpoints as the slope constraint condition for cubic spline interpolation; when the length of the sequence to be repaired exceeds 100 ms, add control points at the middle position of the sequence, and the fundamental frequency value of the control points is determined by the weighted average of the front and back fundamental frequencies; construct a cubic spline interpolation function, which is a cubic polynomial within each sub-interval and ensures the continuity of the function value and the first derivative at the control points; segment the sequence to be repaired using a 60-ms sliding window, and the sliding window slides forward in steps of 10 ms; when the sliding window moves to a new position, only calculate the fundamental frequency values of the newly added part within the window, and reuse the results of the already calculated part to achieve incremental calculation; for each sampling point within the window, calculate its fundamental frequency value through the cubic spline interpolation function; according to the change amount of the fundamental frequency value, query the fundamental frequency-formant mapping relationship parameters; if the fundamental frequency value changes relative to the previous sampling point, calculate the adjustment amount of the formant bandwidth according to the mapping relationship parameters; when the fundamental frequency change amount is positive, increase the formant bandwidth, and the increase amplitude is proportional to the product of the fundamental frequency change amount and the mapping coefficient; when the fundamental frequency change amount is negative, decrease the formant bandwidth, and the decrease amplitude is also determined by the fundamental frequency change amount and the mapping coefficient; smooth the bandwidth adjustment amount, and use a 5-point moving average filter to eliminate mutations; if the adjusted bandwidth value exceeds the preset range, truncate it to ensure the rationality of the bandwidth value; in the overlapping area of the sliding window, linearly weight and fuse the bandwidth values to avoid discontinuity caused by window switching; transfer the fundamental frequency value and the adjusted bandwidth value to the subsequent signal reconstruction module for generating the repaired speech signal.

[0066] In summary, the live human voice pitch adaptive adjustment method based on the network model according to the embodiments of the present invention is elucidated. It determines the exact start position and end position of the packet loss through the mutation points of the short-time energy function and the abnormal change points of the zero-crossing rate, and thus specifically uses the fundamental frequency trajectory characteristics and formant characteristics to repair the packet loss interval. In this way, it can solve the problem that the traditional voice repair method has a contradiction between the decoupled repair of acoustic parameters and the real-time processing constraints in high packet loss rate scenarios, resulting in broken fundamental frequency trajectories and unnatural voices, so as to achieve high-precision and low-latency adaptive repair of packet loss voices, and significantly improve the sound quality and user experience in the live broadcast scenario.

[0067] Here, those skilled in the art can understand that the specific operations of each step in the above live human voice pitch adaptive adjustment method based on the network model have been introduced in detail in the description of the live human voice pitch adaptive adjustment method based on the network model with reference to Figure 1 and Figure 2 and thus, the repeated description thereof will be omitted.

Claims

1. A method for adaptively adjusting the tone of a live voice based on a network model, characterized in that, Including: Collecting live human voice audio data, performing frame segmentation on the audio data with a frame length of 32 ms and detecting packet loss situations, identifying the packet loss boundary positions through the energy envelope analysis of three consecutive audio frames, and marking the packet loss intervals as sequences to be repaired; Performing fast Fourier transform on the complete audio frames 64 ms before and after the sequence to be repaired, and extracting fundamental frequency trajectory features and formant features; Inputting the fundamental frequency trajectory features and the formant features into a pre-trained feature correlation network model to obtain fundamental frequency-formant mapping relationship parameters; Based on the mapping relationship parameters, calculating the fundamental frequency values of the sequence to be repaired using cubic spline interpolation, and synchronously adjusting the formant bandwidth according to the fundamental frequency-formant mapping relationship parameters, and realizing low-latency processing through the incremental calculation method of a 60-ms sliding window.

2. The method for adaptively adjusting the live human voice tone based on a network model according to claim 1, wherein The identification of the packet loss boundary positions determines the exact start position and end position of the packet loss through the mutation points of the short-time energy function and the abnormal change points of the zero-crossing rate.

3. The live human voice tone adaptive adjustment method based on a network model according to claim 2, characterized in that The determination of the mutation points includes: Performing wavelet decomposition on the short-time energy function values, and selecting the third-layer approximation coefficients and detail coefficients to construct energy profile features; Calculating the local variance of the energy profile features, and adaptively dividing the energy mutation intervals using the OTSU algorithm according to the variance distribution; When the changes in the energy profile features at certain moments fall into the energy mutation intervals, these moments are regarded as the mutation points of the short-time energy function.

4. The live human voice tone adaptive adjustment method based on a network model according to claim 3, characterized in that The energy mutation intervals smooth the variance curve through a sliding window and combine the variance distribution histogram, determine the optimal segmentation threshold according to the between-class variance criterion, divide the variance curve into mutation intervals and stable intervals, and finally mark the effective energy mutation intervals through region growing and morphological processing.

5. The live human voice pitch adaptive adjustment method based on a network model according to claim 1, characterized in that Extracting the fundamental frequency trajectory features and the formant features includes: Performing fast Fourier transform on the complete audio frames to obtain spectral features; In the first frequency band range of the spectral features, extracting fundamental frequency candidate points using cepstrum analysis method, and connecting the fundamental frequency candidate points using dynamic programming algorithm; Calculating the similarity cost of the fundamental frequency change between adjacent frames, and selecting the path with the minimum total cost as the fundamental frequency trajectory; In the second frequency band range of the spectral features, obtaining the formant frequency positions by solving the characteristic equation of the linear prediction coefficients.

6. The live human voice tone adaptive adjustment method based on a network model according to claim 5, wherein Extracting the fundamental frequency candidate points through preprocessing and cepstrum analysis of the spectral features, extracting fundamental frequency candidate points, and verifying according to prominence, sharpness and harmonic structure to select the fundamental frequency candidate values with harmonic characteristics.

7. The method for adaptively adjusting the live human voice tone based on a network model according to claim 5, wherein The similarity cost includes the continuity cost of the fundamental frequency change and the similarity cost of the spectral envelope; The continuity cost of the fundamental frequency change is obtained by calculating the logarithmic difference between the fundamental frequency candidate points of the current frame and the previous frame, dividing the logarithmic difference by the time interval between the two frames to obtain the fundamental frequency change rate, and non-linearly mapping the fundamental frequency change rate, and dynamically evaluating the continuity cost of the fundamental frequency change in combination with the sliding-updated fundamental frequency standard deviation. The similarity cost of the spectral envelope is obtained by sampling the harmonic sequences of the fundamental frequency candidate points of the current frame and the previous frame and reconstructing the spectral envelope curve, calculating the Pearson correlation coefficient within the specified frequency band range, and normalizing the arccosine value of the correlation coefficient to be between 0 and 1 as the similarity cost of the spectral envelope.

8. The live human voice pitch adaptive adjustment method based on a network model according to claim 1, characterized in that, The feature correlation network model represents the fundamental frequency trajectory and formant features as time series, extracts the local change patterns of the fundamental frequency through convolution, maps the formant features through fully connected layers, and constructs a correlation matrix by combining the fundamental frequency mutation points and the formant parameter differences; Extracts the main patterns using singular value decomposition, generates a piecewise linear mapping function of the fundamental frequency - formant, and predicts the formant parameter adjustment amount according to the fundamental frequency change.

9. The method for adaptively adjusting the live human voice tone based on a network model according to claim 8, characterized in that, The piecewise linear mapping function of the fundamental frequency - formant dynamically calculates the formant parameter adjustment amount according to the fundamental frequency change amount, and smooths the transition and limits the saturation value through linear or nonlinear mapping. At the same time, it constrains the adjusted parameters to meet the physical validity to achieve the collaborative dynamic update of the fundamental frequency and formant features.

10. The method for adaptively adjusting the live human voice tone based on a network model according to claim 1, wherein Constructs a cubic spline interpolation function through the boundary fundamental frequency value and its slope constraint, combines the weighted average fundamental frequency value of the intermediate control points, and calculates the fundamental frequency value of the sequence to be repaired point by point.

Citation Information

Patent Citations

  • A voice data mapping method and apparatus

    CN103366735A

  • Voice processing method and device, computer readable storage medium and electronic equipment

    CN111326166A