Audio signal processing method, computer device and computer program product

By identifying peak points and calculating the probability of periodic points in the time domain sampling sequence of the audio signal, the target peak points are screened out, which solves the problem of inaccurate determination of the fundamental frequency period in the existing technology and realizes efficient and accurate identification of the fundamental frequency period of the audio signal.

CN114694681BActive Publication Date: 2025-09-16TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210333267.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-31
Publication Date
2025-09-16
Estimated Expiration
2042-03-31

AI Technical Summary

Technical Problem

Existing audio pitch shifting technologies are difficult to efficiently and accurately determine the fundamental frequency period of an audio signal, which affects the pitch shifting effect.

Method used

By obtaining the peak points in the time domain sampling sequence of the audio signal, determining the candidate peak points associated with the current peak point to be analyzed, calculating the periodic point probability of the candidate peak points, screening out the target peak points that meet the preset probability conditions, and determining the fundamental frequency period based on the target peak point interval.

Benefits of technology

The positioning of the periodic points of the fundamental frequency period is simplified, the amount of calculation is reduced, the accuracy and efficiency of identifying the target peak point are improved, and the efficient and accurate determination of the fundamental frequency period of the audio signal is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114694681B_ABST
    Figure CN114694681B_ABST
Patent Text Reader

Abstract

The present application relates to an audio signal processing method, computer device, and computer program product. The method includes: obtaining multiple peak points in a time domain sampling sequence of an audio signal, determining a current peak point to be analyzed from the multiple peak points; determining multiple candidate peak points associated with the current peak point to be analyzed; obtaining the periodic point probability of each candidate peak point, and determining a target peak point whose periodic point probability meets a preset probability condition from the multiple candidate peak points; using the target peak point as the current peak point to be analyzed, returning to execute the step of determining multiple candidate peak points corresponding to the current peak point to be analyzed, until all peak points in the time domain sampling sequence are traversed to obtain at least one determined target peak point; determining the fundamental frequency period of the audio signal based on the interval between the at least one target peak point, realizing accurate identification of the periodic points, and being able to efficiently and accurately determine the fundamental frequency period of the audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of audio technology, and in particular to an audio signal processing method, a computer device, and a computer program product. Background Art

[0002] With the development of multimedia technology, people's leisure activities are becoming increasingly rich and colorful. The demand for audio materials is also showing a trend of diversification and high quality. As a result, audio pitch shifting technology has emerged. Audio pitch shifting technology specifically refers to adjusting the pitch of a specified audio.

[0003] In related technologies, audio can be pitch-shifted using a neural network vocoder, or it can be achieved through a combination of speed change and resampling. However, these pitch-shifting methods often find it difficult to efficiently and accurately determine the fundamental frequency period of the audio signal, which significantly affects the final pitch-shifting effect of the audio. Summary of the Invention

[0004] Based on this, it is necessary to provide an audio signal processing method, a computer device and a computer program product to address the above technical problems.

[0005] In a first aspect, the present application provides an audio signal processing method. The method comprises:

[0006] Acquire multiple peak points in a time domain sampling sequence of an audio signal, and determine a current peak point to be analyzed from the multiple peak points;

[0007] Determine a plurality of candidate peak points associated with a current peak point to be analyzed; the plurality of candidate peak points include the current peak point to be analyzed and at least one peak point adjacent to the current peak point to be analyzed and not previously used as a candidate peak point;

[0008] Obtaining a periodic point probability for each candidate peak point, and determining a target peak point from the plurality of candidate peak points whose periodic point probability satisfies a preset probability condition; the periodic point probability represents a probability that the candidate peak point is a periodic point in the fundamental frequency period of the audio signal;

[0009] Taking the target peak point as the current peak point to be analyzed, returning to the step of determining multiple candidate peak points corresponding to the current peak point to be analyzed, until all peak points in the time domain sampling sequence are traversed to obtain at least one determined target peak point;

[0010] The fundamental frequency period of the audio signal is determined based on an interval between the at least one target peak point.

[0011] In one embodiment, obtaining the periodic point probability of each candidate peak point includes:

[0012] Obtaining an observation probability of the candidate peak point; the observation probability represents a probability that the candidate peak point is a peak point within a preset sequence range of the time domain sampling sequence;

[0013] Obtaining a state transition probability of the candidate peak point; the state transition probability represents a probability that the periodic point position determined as the fundamental frequency period is transferred from other candidate peak points to the candidate peak point, wherein the other candidate peak points are candidate peak points before the candidate peak point among the multiple candidate peak points;

[0014] Based on the observation probability and the state transition probability, the periodic point probability of the candidate peak point is determined.

[0015] In one embodiment, obtaining the observation probability of the candidate peak point includes:

[0016] Obtaining a first sampling moment of the candidate peak point, and obtaining a first time range in the time domain sampling sequence that includes the first sampling moment; the first time range is determined based on a preset reference fundamental frequency period;

[0017] determining a signal amplitude fluctuation range based on a minimum value and a maximum value of the audio signal in the first time range;

[0018] A signal amplitude difference between the candidate peak point and the minimum value is obtained, and an observation probability of the candidate peak point is determined based on a ratio of the signal amplitude difference to the signal amplitude fluctuation range.

[0019] In one embodiment, the step of obtaining a plurality of peak points in a time domain sampling sequence of an audio signal includes:

[0020] Acquire multiple candidate periodic points in a time domain sampling sequence of an audio signal; the candidate periodic points are sampling points with periodic characteristics in the time domain sampling sequence;

[0021] Determine a target periodic point among the multiple candidate periodic points; the target periodic point is a candidate periodic point in the time domain sampling sequence corresponding to a complete signal period;

[0022] A peak point in the signal period of each target period point is obtained to obtain multiple peak points of the audio signal in the time domain sampling sequence.

[0023] In one embodiment, determining a target cycle point from among the multiple candidate cycle points includes:

[0024] Obtaining a second sampling moment of the candidate periodic point, and obtaining a second time range including the second sampling moment; the second time range is determined based on a preset reference fundamental frequency period;

[0025] If the second time range is within the time domain sampling sequence, the candidate periodic point is determined as the target periodic point.

[0026] In one embodiment, the step of obtaining a plurality of candidate periodic points in a time domain sampling sequence of an audio signal includes:

[0027] Obtaining a zero-crossing point in a time-domain sampling sequence of an audio signal;

[0028] Based on each zero-crossing point, a plurality of candidate periodic points in the time-domain sampling sequence of the audio signal are obtained.

[0029] In one embodiment, after determining the fundamental frequency period of the audio signal based on the interval between the target peak points, the method further includes:

[0030] framing the audio signal based on the fundamental frequency period, and obtaining a plurality of signal frames based on the framing result;

[0031] Acquire a modulation parameter sequence, where the modulation parameter sequence includes at least two modulation parameters;

[0032] For each signal frame, mapping the modulation parameters in the modulation parameter sequence to target peak points in the signal frame to obtain a modulation-processed signal frame;

[0033] A plurality of pitch-shifted signal frames are synthesized to obtain a pitch-shifted audio signal.

[0034] In one embodiment, framing the audio signal based on the fundamental frequency period, and obtaining a plurality of signal frames based on the framing result, includes:

[0035] For each target peak point in the audio signal, obtaining a target fundamental frequency period corresponding to the target peak point;

[0036] Taking the target peak point as the center and the target fundamental frequency period as the intercept, intercepting the audio signal corresponding to the target peak point from the audio signal to obtain a corresponding audio frame;

[0037] Windowing is performed on each audio frame to obtain multiple signal frames.

[0038] In one embodiment, before acquiring a plurality of peak points in the time domain sampling sequence of the audio signal, the method further includes:

[0039] Acquire an audio signal to be modulated, and perform DC component removal processing on the audio signal to obtain a DC-free signal;

[0040] performing low-pass filtering on the DC-removed signal, and performing trend elimination on the signal after the low-pass filtering;

[0041] Based on the audio signal after trend elimination, a time domain sampling sequence corresponding to the audio signal is obtained.

[0042] In a second aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.

[0043] In a third aspect, the present application also provides a computer program product, comprising a computer program, which implements the steps of the above method when executed by a processor.

[0044] The above-mentioned audio signal processing method, computer device, and computer program product can obtain multiple peak points in a time-domain sampling sequence of an audio signal, determine a current peak point to be analyzed from the multiple peak points, and determine multiple candidate peak points associated with the current peak point to be analyzed, wherein the multiple candidate peak points include the current peak point to be analyzed and at least one peak point adjacent to the current peak point to be analyzed and not previously selected as a candidate peak point. Furthermore, the periodicity probability of each candidate peak point can be obtained, and a target peak point whose periodicity probability satisfies a preset probability condition can be determined from the multiple candidate peak points. After the target peak point is selected as the current peak point to be analyzed, the step of determining multiple candidate peak points corresponding to the current peak point to be analyzed can be returned to, until all peak points in the time-domain sampling sequence are traversed and at least one target peak point is determined. Then, the fundamental frequency period of the audio signal can be determined based on the interval between the at least one target peak point. The present application can simplify the periodic point positioning of the entire fundamental frequency period into the identification of target peak points in each local range of the time domain sampling sequence, and can quickly identify reliable target peak points while reducing the amount of calculation. In addition, the target peak point in the next local range is determined based on the current target peak point, and the correlation between the periodic points is fully considered, thereby improving the accuracy and efficiency of the multiple target peak points finally screened out, realizing accurate identification of the periodic points, and being able to efficiently and accurately determine the fundamental frequency period of the audio signal. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 1 is a flow chart of an audio signal processing method according to an embodiment;

[0046] Figure 2 Schematic diagram of a process for determining the probability of a periodic point in one embodiment;

[0047] Figure 3 is a schematic diagram of candidate periodic points in a time domain sampling sequence in one embodiment;

[0048] Figure 4 is a schematic diagram of a target periodic point in a time domain sampling sequence in one embodiment;

[0049] Figure 5 is a schematic diagram of a target peak point in a time domain sampling sequence in one embodiment;

[0050] Figure 6 A schematic diagram of a process for obtaining a target cycle point in one embodiment;

[0051] Figure 7 1 is a flow chart of performing pitch shifting processing on an audio signal in one embodiment;

[0052] Figure 8 is a schematic diagram of a modulation parameter sequence in one embodiment;

[0053] Figure 9 is a time domain sampling sequence comparison diagram of an audio signal in one embodiment;

[0054] Figure 10 is a spectrum comparison diagram of an audio signal in one embodiment;

[0055] Figure 11-a is an original audio signal in one embodiment;

[0056] Figure 11-b The audio signal after DC removal processing in one embodiment;

[0057] Figure 11-c is an audio signal after low-pass filtering in one embodiment;

[0058] Figure 11-d An audio signal after trend elimination processing in one embodiment;

[0059] Figure 12 is a flowchart of an audio signal processing method according to another embodiment;

[0060] Figure 13 is a structural block diagram of an audio signal processing device in one embodiment;

[0061] Figure 14 is a diagram of the internal structure of a computer device in one embodiment;

[0062] Figure 15 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0063] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0064] In one embodiment, Figure 1 As shown, a method for processing audio signals is provided. This embodiment is illustrated using a terminal as an example. It is understood that the method can also be applied to a server, or to a system including a terminal and a server, and implemented through interaction between the terminal and the server. The terminal may be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. The portable wearable devices may be smart watches, smart bracelets, head-mounted devices, etc. The server may be implemented as a standalone server or a server cluster consisting of multiple servers.

[0065] In this embodiment, the method may include the following steps:

[0066] Step S110 , obtaining a plurality of peak points in a time domain sampling sequence of an audio signal, and determining a peak point to be currently analyzed among the plurality of peak points.

[0067] As an example, the audio signal in step S110 can be an audio signal that needs to determine the fundamental frequency or pitch of the signal, and specifically can be an audio signal to be modulated, such as a speaking voice or a singing voice. Specifically, sound is composed of a series of vibrations with different frequencies and amplitudes emitted by a sound-producing body. Among these vibrations, there is a vibration with the lowest frequency. The sound generated by this vibration can be called the fundamental tone, and the sounds generated by the remaining vibrations can be called overtones. In real life, the sound generated by the overall vibration of the sound-producing body can be called the fundamental tone, which determines the pitch of the sound, and the sound generated by the partial vibration of the sound-producing body can be called the overtone, which determines the timbre of the sound. In other words, the fundamental tone plays a dominant role in the pitch or tone of the sound. By determining the fundamental tone in the audio signal and modifying it, the sound modulation effect can be achieved.

[0068] In practical applications, audio signals can be expressed using time domain or frequency domain information. This embodiment analyzes and processes audio signals using their time domain expression. Specifically, after acquiring the audio signal, a time domain sampling sequence corresponding to the audio signal can be obtained. This time domain sampling sequence can be a sequence describing the relationship between the sound intensity and time of the audio signal. That is, the horizontal axis of the time domain sampling sequence can be the sample point number or sampling time, and the vertical axis can be decibels.

[0069] After obtaining the time domain sampling sequence, the time domain sampling sequence includes multiple sampling points whose sound intensity changes over time. Based on the changes of the multiple sampling points in the time domain sampling sequence, multiple peak points in the time domain sampling sequence can be obtained, and one peak point among the multiple peak points can be determined as the current peak point to be analyzed.

[0070] Step S120: determining a plurality of candidate peak points associated with the current peak point to be analyzed.

[0071] The multiple candidate peak points include the current peak point to be analyzed and at least one peak point that is adjacent to the current peak point to be analyzed and has not been a candidate peak point.

[0072] Specifically, at least one peak point adjacent to the current peak point to be analyzed may be a preset number of peak points that are sorted after the current peak point to be analyzed and are within a preset range. For example, if the multiple peak points include peak points A1, A2, A3, ..., A m If peak point A1 is determined as the current peak point to be analyzed, at least one peak point adjacent to the current peak point to be analyzed can be determined from k (k ≥ 1) peak points after peak point A1, where k is a preset range. The peak points in the preset range do not include peak points that have been candidate peak points. Each peak point in the preset range can be determined as adjacent to the current peak point to be analyzed, or some of the peak points in the preset range can be determined as adjacent peak points.

[0073] In a specific implementation, after determining the current peak point to be analyzed, multiple candidate peak points associated with the current peak point to be analyzed can be further determined from the time-domain sampling sequence. The multiple candidate peak points include not only at least one peak point adjacent to the current peak point to be analyzed, but also the current peak point itself. In one example, three candidate peak points can be obtained each time multiple candidate peak points are determined. That is, in addition to the current peak point to be analyzed, two peak points adjacent to the current peak point to be analyzed can also be selected.

[0074] Step S130 , obtaining the periodic point probability of each candidate peak point, and determining a target peak point whose periodic point probability satisfies a preset probability condition from multiple candidate peak points; the periodic point probability represents the probability that the candidate peak point is a periodic point in the fundamental frequency period of the audio signal.

[0075] As an example, the preset probability condition may be n (n≥1) candidate peak points with the highest periodic point probability.

[0076] In practical applications, after determining multiple candidate peak points, the periodic probability corresponding to each candidate peak point can be obtained. Specifically, peak points in a signal generally have periodic or nearly periodic characteristics. However, an audio signal can be a superposition of a fundamental frequency and overtones. Furthermore, audio signals may contain interference signals, which can affect the peak period and the accuracy of the subsequently determined fundamental frequency period.

[0077] In this embodiment, after obtaining the periodic point probability corresponding to each candidate peak point, multiple candidate peak points can be compared based on the corresponding periodic point probabilities, and then the target peak point whose periodic point probability meets the preset probability condition can be determined from the multiple candidate peak points.

[0078] In step S140, the target peak point is used as the current peak point to be analyzed, and the process returns to the step of determining multiple candidate peak points associated with the current peak point to be analyzed, until all peak points in the time domain sampling sequence are traversed to obtain at least one determined target peak point.

[0079] After determining the target peak point from multiple candidate peak points, the target peak point can be stored and used as the current peak point to be analyzed. The process then returns to step S120 to determine multiple candidate peak points associated with the current peak point to be analyzed. This process is repeated until all peak points in the time domain sampling sequence are traversed. This means that when the target peak point is used as the current peak point to be analyzed, no candidate peak points that meet the criteria are found when searching for candidate peak points associated with the peak point to be analyzed. Furthermore, based on the stored target peak point, at least one target peak point determined for the audio signal can be obtained. For example, when storing the target peak point, the target peak point can be recorded in a preset sequence for subsequent timely retrieval. Based on the concept of dynamic programming, the present application can recursively determine the next target peak point from multiple subsequent candidate peak points based on the currently determined target peak point, effectively improving the accuracy and efficiency of the selected target peak points.

[0080] Step S150: determining a fundamental frequency period of the audio signal based on an interval between at least one target peak point.

[0081] After obtaining the target peak point in the time domain sampling sequence, the fundamental frequency period of the audio signal can be determined based on the interval between at least one target peak point. For example, if the identified target peak point is one, the fundamental frequency period can be determined based on the interval between the target peak point and the starting sample point of the audio signal. If there are multiple target peak points, the fundamental frequency period of the audio signal can be determined based on the interval between the target peak points. Specifically, the intervals between different adjacent target peak points may be different. For example, the interval between target peak points T1 and T2 is not equal to the interval between target peak points T2 and T3. When determining the fundamental frequency period, the fundamental frequency period of the audio signal can be determined based on the intervals between multiple groups of adjacent target peak points in the time domain sampling sequence.

[0082] In the above-mentioned audio signal processing method, multiple peak points in a time-domain sampling sequence of an audio signal can be obtained, a current peak point to be analyzed can be determined from the multiple peak points, and multiple candidate peak points associated with the current peak point to be analyzed can be determined, wherein the multiple candidate peak points include the current peak point to be analyzed and at least one peak point adjacent to the current peak point to be analyzed and not previously used as a candidate peak point. Furthermore, a periodic point probability of each candidate peak point can be obtained, and a target peak point whose periodic point probability satisfies a preset probability condition can be determined from the multiple candidate peak points. After the target peak point is used as the current peak point to be analyzed, the step of determining multiple candidate peak points associated with the current peak point to be analyzed can be returned to, until all peak points in the time-domain sampling sequence are traversed and at least one target peak point is determined. Then, the fundamental frequency period of the audio signal can be determined based on the interval between the at least one target peak point. In this embodiment, the periodic point positioning of the entire fundamental frequency period can be simplified to the identification of target peak points in each local range of the time domain sampling sequence, which can quickly identify reliable target peak points while reducing the amount of calculation. Moreover, the target peak point in the next local range is determined based on the current target peak point, and the correlation between the periodic points is fully considered, thereby improving the accuracy and efficiency of the multiple target peak points finally screened out, realizing accurate identification of the periodic points, and being able to efficiently and accurately determine the fundamental frequency period of the audio signal.

[0083] In one embodiment, Figure 2 As shown, in step S130, obtaining the periodic point probability of each candidate peak point may include the following steps:

[0084] Step S131 , obtaining the observation probability of the candidate peak point; the observation probability represents the probability that the candidate peak point is a peak point within a preset sequence range of the time domain sampling sequence.

[0085] When determining a target peak point from a plurality of candidate peak points, the target peak point may be obtained based on dynamic programming (DP). In this embodiment, dynamic programming may be implemented based on a criterion of maximizing the best probability.

[0086] In practical applications, for each candidate peak point, a preset sequence range containing the candidate peak point can be determined, and the probability that the candidate peak point is a peak point within the preset sequence range can be estimated, that is, the probability that the amplitude of the candidate peak point is the maximum value within the preset sequence range is obtained.

[0087] Step S132, obtaining the state transition probability corresponding to the candidate peak point; the state transition probability represents the probability that the periodic point position determined as the fundamental frequency period transfers from other candidate peak points to the candidate peak point, and the other candidate peak points are candidate peak points before the candidate peak point among multiple candidate peak points.

[0088] In actual applications, each candidate peak point may or may be used as a periodic point position for determining the fundamental frequency period, but the accuracy of the fundamental frequency period determined thereby may vary. For example, when a candidate peak point A is used as a periodic point position for determining the fundamental frequency period, the accuracy of the fundamental frequency period calculated based on the candidate peak point A may be lower than the accuracy of the fundamental frequency period calculated when other candidate peak points (such as the candidate peak point B) are used as the periodic point position, and there is a probability that the periodic point position will shift.

[0089] Based on this, this embodiment can obtain the state transition probability of the candidate peak point. Specifically, for the candidate peak point whose periodic point probability is currently to be calculated, other candidate peak points before the candidate peak point can be determined from multiple candidate peak points. For example, if the multiple candidate peak points include A, B, and C, then when determining the periodic point probability of candidate peak point B, candidate peak point A can be determined as another candidate peak point of candidate peak point B. Furthermore, the state transition probability of the periodic point position of the fundamental frequency period shifting from the other candidate peak points to the candidate peak point can be obtained.

[0090] In one example, the state transition probability corresponding to the candidate peak point can be calculated in the following manner.

[0091] First, obtain a predefined state transition function, which can be as follows:

[0092]

[0093] Among them, P i (j) represents the periodic point position from the candidate peak point t i Move to the next candidate peak point t jThe possibility of state transition, δ, can be determined by the preset reference fundamental frequency period T0, for example, δ=0.4·T0.

[0094] Step S133: determining the periodic point probability of the candidate peak point based on the observation probability and the state transition probability.

[0095] After obtaining the observation probability and state transition probability of the candidate peak point, the period point probability of the candidate peak point can be determined by combining the observation probability and the state transition probability.

[0096] For example, if the candidate peak point t is determined j The observation probability is O(j), and the state transition probability is P i (j), a cost function (also called a price function) can be constructed. For example, the cost function Cost(j) can be as follows:

[0097] Cost(j)=O(j)·max{Cost(m)P m (j)}|m=j-1,…,jJ

[0098] Among them, O(j) is the candidate peak point t j The observation probability of Cost(m) is the probability of observation at the candidate peak point t i Other candidate peak points before t m The cost function result of (m=j-1,...,jJ), P m (j) is the periodic point position from other candidate peak points t m Transfer to candidate peak point t j The state transition probability of . For example, J=3.

[0099] To prevent data from overflowing to infinity or infinitesimal, the cost function can be deformed and the next step of calculation can be performed in logarithmic form. The logarithmic cost function is as follows:

[0100]

[0101] Where C(j)=log(Cost(j)),

[0102] Based on the above-mentioned cost function associated with the observation probability and state transition probability, the determined observation probability and state transition probability can be substituted into the cost function. Based on the processing result of the cost function, the periodic point probability of the candidate peak point can be determined. Then, for multiple candidate peak points sorted in sequence, multiple target peak points in the audio signal can be determined by recursively performing dynamic programming.

[0103] In this embodiment, by obtaining the observation probability of the candidate peak point and the state transition probability of the candidate peak point, the periodic point probability of the candidate peak point is determined based on the observation probability and the state transition probability. The periodic point probability of each candidate peak point as a periodic point position can be accurately obtained, providing a screening basis for determining the accurate fundamental frequency period.

[0104] In one embodiment, in step S131, obtaining the observation probability of the candidate peak point may include the following steps:

[0105] Obtain a first sampling moment of a candidate peak point, and obtain a first time range in a time domain sampling sequence that includes the first sampling moment; determine a signal amplitude fluctuation range based on a minimum value and a maximum value of the audio signal in the first time range; obtain a signal amplitude difference between the candidate peak point and the minimum value, and determine an observation probability of the candidate peak point based on a ratio of the signal amplitude difference to the signal amplitude fluctuation range.

[0106] The first time range is determined based on a preset reference fundamental frequency period. For example, the reference fundamental frequency period T0 may be 20 ms. The first sampling moment is the sampling moment of the candidate peak point in the time domain sampling sequence.

[0107] Specifically, when determining the target epoch marker, it can be considered that the position of the target epoch marker is positively correlated with the peak value of the candidate epoch marker, that is, the larger the signal peak value of the candidate epoch marker, the more likely the candidate epoch marker is to be the target epoch marker.

[0108] In this embodiment, after the candidate peak point is determined, the first sampling moment of the candidate peak point can be obtained. For example, if the horizontal coordinate of the time domain sampling sequence is the sampling time of the sample point, the first sampling time can be obtained based on the horizontal coordinate corresponding to the candidate peak point. If the horizontal coordinate is the sample point sequence number, the first sampling moment can be determined based on the sampling interval and the horizontal coordinate corresponding to the candidate peak point.

[0109] After determining the first sampling moment, the candidate peak points can be evaluated within a time range that includes the first sampling moment. Specifically, a reference fundamental frequency period can be obtained and based on the period size of the reference fundamental frequency period, a time range size of the reference fundamental frequency period and a first time range that includes the first sampling moment can be obtained in the time domain sampling sequence. For example, the first sampling moment can be used as the starting point, and a range of a reference fundamental frequency period can be intercepted forward and backward respectively to obtain the first time range. For example, for the first sampling time j, the corresponding first time range can be [j-T0, j+T0]. Of course, based on actual conditions, the size of the reference fundamental frequency period corresponding to the sampling multiple can also be used to obtain the first time range.

[0110] After determining the first time range, the minimum and maximum values ​​corresponding to the audio signal at each sample point within the first time range can be determined based on the time domain sampling sequence. The amplitude fluctuation range of the audio signal can then be determined based on the minimum and maximum values. Furthermore, the amplitude difference between the candidate peak point and the minimum value can be obtained, and the observation probability corresponding to the candidate peak point can be determined based on the ratio of the amplitude difference to the amplitude fluctuation range. In one example, the observation probability can be obtained as shown in the following formula:

[0111]

[0112] Among them, h min , h max They represent the minimum and maximum values ​​of the signal at sample point j within the first time range, respectively, and h(j) is the signal value at sample point j.

[0113] Accordingly, the cost function in logarithmic form can be shown as follows:

[0114]

[0115] In this embodiment, by obtaining the signal amplitude difference between the candidate peak point and the minimum value, and based on the ratio of the signal amplitude difference to the signal amplitude fluctuation range, the probability that the candidate peak point is a peak point in its adjacent area can be quickly determined to obtain the observation probability.

[0116] In one embodiment, the step of obtaining a plurality of peak points in a time domain sampling sequence of an audio signal includes:

[0117] Acquire multiple candidate periodic points in a time domain sampling sequence of an audio signal; determine a target periodic point among the multiple candidate periodic points; obtain a peak point in a signal period of each target periodic point, and obtain multiple peak points of the audio signal in the time domain sampling sequence.

[0118] As an example, a candidate periodic point is a sample point with periodic characteristics in a time domain sampling sequence, wherein the periodic characteristics may refer to approximately periodic appearance or periodic appearance, such as a peak point, a trough point or a zero-crossing point in an audio signal.

[0119] The target periodic point is a candidate periodic point corresponding to a complete signal period in the time domain sampling sequence, ie, an audio signal having a complete signal period within a preset range including the target periodic point.

[0120] In practical applications, multiple candidate periodic points can be obtained in the time domain sampling sequence of the audio signal, and the multiple candidate periodic points are periodic points of the same type. For example, the candidate periodic points are determined based on multiple peak points in the time domain sampling sequence, or the candidate periodic points are determined based on multiple zero-crossing points in the time domain sampling sequence.

[0121] After determining multiple candidate periodic points, it can be determined whether each candidate periodic point corresponds to a complete signal cycle. In other words, it can be determined whether a sample sequence within a preset time range that includes the candidate periodic point can be determined in the time domain sampling sequence. If it is determined that the candidate periodic point corresponds to a complete signal cycle, the candidate periodic point can be determined as a target periodic point.

[0122] After determining the target periodic point, since the sample point reliability of the peak position is higher when screening periodic sample points, the peak point in the signal cycle can be further obtained within the complete signal cycle corresponding to the target periodic point, thereby obtaining multiple peak points of the audio signal in the time domain sampling sequence.

[0123] In this embodiment, by obtaining multiple candidate periodic points in the time domain sampling sequence of the audio signal, determining the target periodic point among the multiple candidate periodic points, obtaining the peak point in the signal period of each target periodic point, and obtaining multiple peak points of the audio signal in the time domain sampling sequence, it is possible to first obtain candidate periodic points that may be periodic, and then obtain multiple more reliable peak points through peak adsorption, thereby providing a basis for screening out accurate target peak points.

[0124] In one embodiment, obtaining multiple candidate periodic points in the time domain sampling sequence of the audio signal may include the following steps:

[0125] A zero-crossing point in a time-domain sampling sequence of an audio signal is obtained; and based on each zero-crossing point, a plurality of candidate periodic points in the time-domain sampling sequence corresponding to the audio signal are obtained.

[0126] In practical applications, some zero-crossing points in audio signals are periodic. Based on this, after obtaining the time domain sampling sequence of the audio signal, the zero-crossing points in the time domain sampling sequence can be obtained. For example, the positive zero-crossing points in the time domain sampling sequence can be obtained. In one example, the positive zero-crossing points can be screened out as candidate periodic points using the following formula.

[0127]

[0128] Where y(n) is the time domain sampling sequence, n is the sample point index in the time domain sampling sequence y(n), and when the value of epoch(n) is 1, it is determined to be a positive zero crossing point. Figure 3 As shown, the positive zero crossing point in the audio signal corresponds to the position of epoch=1 in the figure.

[0129] In this embodiment, multiple candidate periodic points in the time domain sampling sequence corresponding to the audio signal can be quickly screened out based on the positive zero-crossing points in the time domain sampling sequence, thereby effectively improving the processing efficiency of the audio signal.

[0130] In one embodiment, determining a target periodic point from among a plurality of candidate periodic points may include the following steps:

[0131] A second sampling moment of the candidate periodic point is obtained, and a second time range including the second sampling moment is obtained; if the second time range is within the time domain sampling sequence, the candidate periodic point is determined as the target periodic point.

[0132] The second time range is determined based on a preset reference fundamental frequency period; and the second sampling moment is a sampling moment of the candidate period point in the time domain sampling sequence.

[0133] In practical applications, the second sampling moment of the candidate periodic point can be obtained. After determining the second sampling moment, a second time range including the second sampling moment can be obtained. When obtaining the second time range, the range size of the second time range is determined based on the reference fundamental frequency period.

[0134] For example, the second time range may be an integer multiple of the reference fundamental frequency period. For example, if the second time range is a reference fundamental frequency period T0, then after determining the second sampling time n, the second sampling time may be used as the center and T0 may be intercepted forward and backward. half (i.e., half of the reference fundamental frequency period), thereby obtaining a second time range [n-T0] with a range size of one reference fundamental frequency period. half , n+T0 half ].

[0135] After determining the second time range, it can be determined whether the second time range exceeds the time range of the time domain sampling sequence. In other words, it can be determined whether a reference fundamental frequency period including the second sampling moment has a corresponding audio signal in the time domain sampling sequence. If the second time range is within the time domain sampling sequence, the candidate periodic point can be determined as the target periodic point; if the second time range exceeds the time domain sampling sequence, for example, the start time of the second time range is less than the start time of the time domain sampling sequence or the end time of the second time range is greater than the end time of the time domain sampling sequence, it is determined that the periodic signal range of the candidate periodic point has crossed the boundary, the current candidate periodic point is not used as the target periodic point, and the next candidate periodic point is obtained to continue the judgment until all candidate periodic points are traversed and multiple target periodic points are obtained.

[0136] like Figure 4 As shown in the figure, multiple target periodic points are obtained by screening the specified period of the audio signal, which corresponds to the "epoch adsorbed to the peak position" in the figure. After taking multiple target periodic points as the peak points of the audio signal, the following can be finally obtained: Figure 5 The target peak points are shown as follows, wherein the target peak points correspond to Figure 5 The epoch marker position in .

[0137] In this embodiment, a second sampling moment of a candidate periodic point is obtained, and a second time range containing the second sampling moment is obtained. If the second time range is within the time domain sampling sequence, the candidate periodic point is determined as the target periodic point. By determining whether the second time range of the candidate periodic point is within the time domain sampling sequence, it is possible to quickly determine whether the candidate periodic point actually corresponds to a complete periodic signal, thereby screening out reliable target periodic points with periodicity.

[0138] In order to enable those skilled in the art to better understand the above steps, the embodiment of the present application is illustrated below by using an example, but it should be understood that the embodiment of the present application is not limited to this.

[0139] like Figure 6 As shown, the candidate periodic point can be obtained by searching for the positive zero-crossing position of the time-domain sampling sequence y(n), where the candidate periodic point is also called the initial epoch marker position. During the search process, the positive zero-crossing point in the time-domain sampling sequence can be marked. If the sample point with the sample point index n in the time-domain sampling sequence is a positive zero-crossing point, epoch(n)=1 can be set; otherwise, epoch(n)=0.

[0140] After the search is completed, each sample point in the time domain sampling sequence can be traversed in turn to determine whether the epoch (n) of the sample point is 1. If it is equal to 1, then further determine whether the time point corresponding to the first half of the reference base frequency cycle starting from the sample point index n is in the time domain sampling sequence (i.e., determine n-T0 in the figure). half Is it greater than or equal to 0), when it is determined that it is in the time domain sampling sequence, it is determined whether the time point corresponding to the second half of the reference base frequency period starting from the sample index n does not exceed the time domain sampling sequence (i.e., the time point n+T0 in the figure is determined to be within the time domain sampling sequence). half Is it less than lx, where lx is the sequence length of the time domain sampling sequence).

[0141] If the time domain sampling sequence is not exceeded, the sample point with the maximum signal value can be obtained within a reference base frequency period corresponding to the sample point, and the peak point is recorded as [vp, ip] (corresponding to [vp, ip] = max{s(n+[-T0 half :T0 half ])},n+[-T0 half :T0 half ] is to take n as the starting point and intercept T0 forward and backward respectively half The second time range, s(n+[-T0 half :T0 half]) is the sampling sequence of the time domain sampling sequence in the second time range). Among them, vp is the peak value of the peak point, ip is the peak point in [-T0 half :T0 half ], so that the peak value of the peak point and its position in the time domain sampling sequence can be quickly determined based on [vp, ip]. At the same time, the sequence number corresponding to the peak point can be determined, that is, idx+1. The corresponding horizontal coordinate is n-T0 half In the above manner, each sampling point in the time domain sampling sequence is traversed in sequence until multiple peak points are obtained.

[0142] In one embodiment, Figure 7 As shown, after determining the fundamental frequency period of the audio signal based on the interval between target peak points, the following steps S160-S190 may be further included:

[0143] Step S160 : Frame the audio signal based on the fundamental frequency period, and obtain a plurality of signal frames based on the frame division result.

[0144] As an example, the fundamental tone signal may be an audio signal corresponding to the fundamental tone.

[0145] In a specific implementation, after obtaining the fundamental frequency period corresponding to the audio signal, the audio signal can be framed based on the fundamental frequency period. For example, a signal frame is determined based on one fundamental frequency period, and multiple signal frames are obtained based on the framing result.

[0146] Step S170: Acquire a modulation parameter sequence, where the modulation parameter sequence includes at least two modulation parameters.

[0147] In practical applications, a pitch shift parameter sequence can be obtained in response to an audio pitch shift request. Specifically, the pitch shift parameter sequence can be included in the audio pitch shift request, and then, after the audio pitch shift request is obtained, the pitch shift parameter sequence can be read from the request. Alternatively, the audio pitch shift request can include pitch shift configuration information, and after receiving the audio pitch shift request, the pitch shift parameter sequence corresponding to the pitch shift configuration information can be obtained. The pitch shift configuration information can include at least one of the following: a target pitch, an audio signal range to be pitch shifted, and an audio morpheme. The target pitch can be a specific musical pitch, such as the key of G.

[0148] Compared with the fixed pitch shift in the traditional technology, that is, the pitch shift of the input audio is fixed constant, in this embodiment, the acquired pitch shift parameter sequence can include at least two pitch shift parameters, that is, when the audio signal is pitch shifted, the audio signal can be pitch shifted based on different pitch shift parameters, for example, a single tone pitch can be changed to a vibrato, and the accuracy can reach the frame shift (tens of milliseconds) level. For example, the pitch shift parameter sequence can be as follows: Figure 8 shown.

[0149] Step S180 : For each signal frame, mapping the modulation parameters in the modulation parameter sequence to target peak points in the signal frame to obtain a modulation-processed signal frame.

[0150] After obtaining the modulation parameter sequence, for each signal frame, the modulation parameters in the modulation parameter sequence can be mapped to the target peak point in the signal frame by interpolation (interp1) or resampling (resample), thereby obtaining the modulation-processed signal frame.

[0151] Step S190 , synthesizing a plurality of tone-shifted signal frames to obtain a tone-shifted audio signal.

[0152] After obtaining multiple signal frames after the pitch shifting process, the multiple signal frames after the pitch shifting process can be synthesized, and based on the synthesis result, the pitch shifted audio signal can be obtained. Figure 9 、 10 As shown, it shows the waveform comparison of the audio signal in the time domain sampling sequence before and after the pitch shift optimization, as well as the corresponding spectrum comparison.

[0153] In practical applications, after obtaining multiple signal frames after tone shifting, the synthesis time point corresponding to each signal frame can be determined. Specifically, the sample sequence corresponding to the signal frame after tone shifting can be recorded as β(t j ), the synthesis time point in the signal frame is The next synthesis time point It can be determined as:

[0154]

[0155] Among them, P Marker (t j ) is the time domain sampling sequence t j The sampling point at the moment, in order to distinguish the framing process and synthesis process of the signal frame, can be used Marker (t j ) represents the target peak point used in the signal frame synthesis process, and P Marker (t i ) represents the target peak point used in the audio signal framing process, that is, the subscripts i and j are used to distinguish the points in the audio signal segmentation and synthesis processes.

[0156] By determining multiple synthesis time points in the above manner, the resulting pitch-shifted audio signal y(n) can be expressed as:

[0157]

[0158] Among them, yj is the jth signal frame to be synthesized after pitch shifting, by splicing multiple pitch shifted signal frames y j The audio signal y(n) can be obtained, t s (k) is the signal frame y j The target peak point in the signal frame y j The sample index in nt s (k) represents signal frame y j The middle sample point n and the target peak point t s (k) between the signal points.

[0159] Furthermore, in order to maintain the stability of the signal output amplitude waveform, the Hanning window weight coefficient can be added to perform OLA (Overlap-Add algorithm, Overlap-Add algorithm is based on frame splicing and regularizes the speech duration in the time domain) synthesis. At the same time, the waveform envelope jitter caused by windowing is removed from the output audio signal, and the current frame is output after modulation. The final modulated audio signal can be expressed as:

[0160]

[0161] Wherein, w(j, k) represents the window function, j represents the jth signal frame after pitch shifting, and k represents the sample point index in the signal frame.

[0162] In this embodiment, the audio signal can be framed based on the fundamental frequency period, and multiple signal frames can be obtained based on the framing results to obtain a modulation parameter sequence. For each signal frame, the modulation parameter in the modulation parameter sequence is mapped to the target peak point in the signal frame to obtain a modulation-processed signal frame. Then, multiple modulation-processed signal frames can be synthesized to obtain a modulated audio signal. When the audio signal is modulated, the audio signal can be modulated by multiple modulation parameters to avoid a single and fixed modulation method, making the modulation form of the audio signal more natural and optimizing the modulation effect finally obtained. Moreover, since the audio signal can be segmented and synthesized based on a precise fundamental frequency period, the phase discontinuity of the audio signal can be effectively avoided, and a high-fidelity and highly natural modulation effect of the singing can be achieved.

[0163] In one embodiment, in step S160, framing the audio signal based on the fundamental frequency period and obtaining a plurality of signal frames based on the framing result includes:

[0164] For each target peak point in the audio signal, the target fundamental frequency period corresponding to the target peak point is obtained; with the target peak point as the center and the target fundamental frequency period as the intercept, the audio signal corresponding to the target peak point is intercepted from the audio signal to obtain the corresponding audio frame; each audio frame is windowed to obtain multiple signal frames.

[0165] In a specific implementation, after obtaining the fundamental frequency period, framing and synthesizing the audio signal based on the periodic points corresponding to the fundamental frequency period can achieve a smoother and more natural signal processing effect. Taking a sinusoidal signal as an example, if the signal is segmented based on the peak points or other periodic points in the sinusoidal signal, three consecutive audio frames A, B, and C are obtained. Because the segmentation is performed at the peak points or other periodic points, the splicing of the two audio frames A and C can achieve a smooth and natural splicing effect. In this embodiment, after determining the fundamental frequency period corresponding to each target peak point in the audio signal, the audio signal can be segmented and synthesized based on the fundamental frequency period.

[0166] For each target peak point in the audio signal, the fundamental frequency period corresponding to the target peak point can be obtained. In order to distinguish it from other fundamental frequency periods, the fundamental frequency period corresponding to the target peak point can be called the target fundamental frequency period. Specifically, for the current target peak point, the target fundamental frequency period Period(i) corresponding to the target peak point can be obtained based on the position interval between the previous target peak point adjacent to the target peak point and the current target peak point. For example, the target peak point can be recorded as P Marker (t i ), represents the time domain sampling sequence t i When calculating the fundamental frequency period of the target peak point, the current target peak point to be processed can be recorded as P Marker (t a (i)), the corresponding fundamental frequency period can be expressed as:

[0167] P eriod (i) = P Marker (t a (i))-P Marker (t a (i-1))

[0168] Among them, P eriod (i) is the target peak point P Marker (t a (i)) corresponds to the fundamental frequency period, P Marker (t a (i-1)) is the peak point P of the target wave Marker (t a (i)) The adjacent previous target peak point.

[0169] After obtaining the target fundamental frequency period, the target peak point can be used as the center and the target fundamental frequency period as the intercept to intercept the audio signal corresponding to the target peak point from the audio signal to obtain the corresponding audio frame. In other words, the target peak point can be used as the center and the length Period(i) can be intercepted before and after to obtain the corresponding audio frame. The audio frame is based on P Marker (ti ) as the center, the window length is L i 2P eriod (i)+1 audio frame. Based on multiple audio frames, the audio signal x(n) can be expressed as:

[0170]

[0171] Among them, x i The i-th audio frame is obtained by splicing multiple audio frames x i The audio signal x(n) can be restored, t a (i) is the audio frame x i The target peak point in the audio frame is n, i The sample index in nt a (i) represents the audio frame x i The middle sample point n and the target peak point t a (i) The signal points between them.

[0172] After obtaining the audio frame, the audio frame can be windowed to obtain the corresponding signal frame. The window function used in the windowing process can be any of the following: Hanning window, rectangular window, triangular window, Hamming window, Gaussian window. By windowing the audio frame x(i, k), the signal frame x after windowing is obtained. w (i, k) can be expressed as:

[0173] x w (i, k) = x(i, k) w(i, k)

[0174] Where i represents the i-th audio frame or signal frame, k represents the sample point index within the audio frame or signal frame, and w(i, k) is the window function. Taking the Hanning window as an example, w(i, k) is expressed as:

[0175]

[0176] In this embodiment, for each target peak point in the audio signal, a target fundamental frequency period corresponding to the target peak point is obtained, and the audio signal corresponding to the target peak point is intercepted from the audio signal with the target peak point as the center and the target fundamental frequency period as the intercept to obtain a corresponding audio frame, and each audio frame is windowed to obtain multiple signal frames. The audio signal can be segmented based on accurate and reliable target peak points, effectively avoiding phase discontinuity and fundamental frequency distortion in the spliced ​​audio signal, providing a basis for obtaining a natural and smooth signal frame splicing effect and pitch shifting effect. In addition, by windowing the audio frames, the cross-fade effect in the signal processing process can be further guaranteed to avoid abrupt changes in the audio signal.

[0177] In one embodiment, before step S110, the method may further include the following steps:

[0178] An audio signal to be modulated is obtained, and a DC component is removed from the audio signal to obtain a DC-free signal; the DC-free signal is low-pass filtered, and a trend is eliminated on the low-pass filtered signal; based on the trend-eliminated audio signal, a time domain sampling sequence corresponding to the audio signal is obtained.

[0179] In practical applications, an audio signal to be modulated can be obtained. The audio signal to be modulated can also be called an original signal or original input. Its corresponding time domain representation can be as follows: Figure 11-a In one example, in order to ensure the stability of the audio signal processing process and prevent the data from crossing the boundary due to the data being too small, the input signal can be normalized first, that is:

[0180]

[0181] Among them, s normmd (n) is the normalized audio signal, and s(n) is the original signal.

[0182] For the received audio signal to be modulated, the DC component of the audio signal to be modulated can be removed first, and the processed audio signal can be Figure 11-b In one example, the time-varying low-frequency offset can be differentially removed to obtain a DC-free signal, which can be expressed as:

[0183] x dcr (n) = x(n) - x(n-1)

[0184] Among them, x dcr (n) is the DC-free signal, x represents the audio signal to be modulated, that is, the original signal, and n is the sample point index.

[0185] After obtaining the DC-free signal, the DC-free signal can be low-pass filtered. By low-pass filtering the DC-free signal, the interference signal in the DC-free signal can be reduced and the periodic signal can be roughly extracted from it. The signal obtained after filtering is as follows: Figure 11-c shown.

[0186] Exemplarily, filtering can be performed by zero-frequency filtering (ZFF), which can be a second-order zero-frequency filter with the following filter structure:

[0187]

[0188] Among them, z is the coefficient of the filter. When filtering, the coefficient of the filter can be adjusted according to actual conditions. For example, the ZFF coefficient can be increased to prevent data overflow.

[0189] Correspondingly, the signal after low-pass filtering can be expressed as:

[0190] y ZFF (n) = x dcr (n)+2y ZFF (n-1)-y ZFF (n-2)

[0191] For the signal after low-pass filtering, there may be obvious trend items that cause signal distortion, so trend elimination can be performed to improve signal quality. The audio signal after trend elimination can be as follows: Figure 11-d In one example, the audio signal after trend elimination can be expressed as:

[0192]

[0193] The 2M+1 value can be set to an interval of 10ms. When performing detrending, multiple iterations can be performed to ensure stable results, for example, 2 to 3 times.

[0194] Furthermore, based on the audio signal after trend elimination, a time domain sampling sequence corresponding to the audio signal to be modulated can be obtained.

[0195] In this embodiment, by obtaining the audio signal to be modulated and performing DC component removal processing on the audio signal to obtain a DC-free signal, then performing low-pass filtering on the DC-free signal, and performing trend elimination on the signal after low-pass filtering, and then obtaining a time domain sampling sequence corresponding to the audio signal based on the audio signal after trend elimination, the interference signal in the original signal can be effectively removed, providing a basis for the subsequent accurate extraction of periodic peak points.

[0196] In order to enable those skilled in the art to better understand the above steps, the embodiment of the present application is illustrated below by using an example, but it should be understood that the embodiment of the present application is not limited to this.

[0197] like Figure 12 As shown, after obtaining the audio signal corresponding to the original singing voice, the time domain sampling sequence corresponding to the audio signal can be preprocessed, including but not limited to DC removal processing, low-pass filtering processing and trend elimination processing, and then multiple peak points can be obtained from the preprocessed time domain sampling sequence as alternative epochs.

[0198] After identifying candidate epochs in the time-domain sampling sequence, multiple candidate peak points are obtained from these candidate epochs. Using a dynamic programming algorithm, the probability of a period corresponding to each candidate peak point is determined. Based on the period probability, a target peak point is identified from the candidate peak points as the confirmed epoch marker. After traversing the peak points in the time-domain sampling sequence, an epoch marker sequence consisting of multiple target peak points is obtained.

[0199] After determining multiple target peak points, the audio signal can be modulated in combination with the FE-SOLA (synchronous overlap-addition based on fuzzy (glottal closure) moments) algorithm. Specifically, the fundamental frequency period of the corresponding audio signal can be determined based on the interval between two adjacent target peak points, and based on the fundamental frequency period, the time domain sampling sequence is segmented to obtain multiple signal frames. For each signal frame, the modulation parameters in the modulation parameter sequence can be mapped to the target peak points in the signal frame by interpolation or resampling, and then the multiple signal frames are synthesized to obtain the singing voice after free modulation.

[0200] The audio signal processing in this embodiment can be applied according to actual scenarios. For example, it can be applied to data set enhancement, intelligent sound correction, sound beautification, intelligent harmony, intelligent accompaniment, etc. during neural network training. For intelligent sound correction, it can also expand and achieve high-quality dynamic speed change while changing the pitch.

[0201] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0202] Based on the same inventive concept, embodiments of the present application further provide an audio signal processing device for implementing the aforementioned audio signal processing method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more of the following audio signal processing device embodiments can be found in the aforementioned limitations of the audio signal processing method and are not further elaborated here.

[0203] In one embodiment, Figure 13 As shown, an audio signal processing device is provided, comprising:

[0204] A peak point acquisition module 1301 is configured to acquire a plurality of peak points in a time domain sampling sequence of an audio signal, and determine a peak point to be currently analyzed from among the plurality of peak points;

[0205] The candidate peak point determination module 1302 is configured to determine a plurality of candidate peak points associated with the current peak point to be analyzed; the plurality of candidate peak points includes the current peak point to be analyzed and at least one peak point adjacent to the current peak point to be analyzed and not previously used as a candidate peak point;

[0206] The periodic point probability determination module 1303 is configured to obtain a periodic point probability for each candidate peak point and determine a target peak point from the plurality of candidate peak points whose periodic point probability satisfies a preset probability condition; the periodic point probability represents a probability that the candidate peak point is a periodic point in the fundamental frequency period of the audio signal;

[0207] a target peak point determining module 1304, configured to use the target peak point as the current peak point to be analyzed, and return to the step of determining multiple candidate peak points corresponding to the current peak point to be analyzed, until all peak points in the time domain sampling sequence are traversed to obtain at least one target peak point;

[0208] The fundamental frequency period determination module 1305 is configured to determine the fundamental frequency period of the audio signal based on the interval between the at least one target peak point.

[0209] In one embodiment, the periodic point probability determination module 1303 includes:

[0210] An observation probability determination submodule is configured to obtain the observation probability of the candidate peak point; the observation probability represents the probability that the candidate peak point is a peak point within a preset sequence range of the time domain sampling sequence;

[0211] a state transition probability determination submodule, configured to obtain a state transition probability of the candidate peak point; the state transition probability represents a probability that the position of a periodic point determined as a fundamental frequency period is transferred from another candidate peak point to the candidate peak point, where the other candidate peak point is a candidate peak point before the candidate peak point among the multiple candidate peak points;

[0212] The periodic point probability acquisition submodule is used to determine the periodic point probability of the candidate peak point based on the observation probability and the state transition probability.

[0213] In one embodiment, the observation probability determination submodule is specifically configured to:

[0214] Obtaining a first sampling moment of the candidate peak point, and obtaining a first time range in the time domain sampling sequence that includes the first sampling moment; the first time range is determined based on a preset reference fundamental frequency period;

[0215] determining a signal amplitude fluctuation range based on a minimum value and a maximum value of the audio signal in the first time range;

[0216] A signal amplitude difference between the candidate peak point and the minimum value is obtained, and an observation probability of the candidate peak point is determined based on a ratio of the signal amplitude difference to the signal amplitude fluctuation range.

[0217] In one embodiment, the peak point acquisition module 1301 includes:

[0218] A candidate periodic point acquisition submodule is used to acquire multiple candidate periodic points in the time domain sampling sequence of the audio signal; the candidate periodic points are sampling points with periodic characteristics in the time domain sampling sequence;

[0219] A target periodic point determination submodule is configured to determine a target periodic point among the plurality of candidate periodic points; the target periodic point is a candidate periodic point in the time domain sampling sequence corresponding to a complete signal period;

[0220] The peak point determination submodule is used to obtain the peak point in the signal period of each target period point, and obtain multiple peak points of the audio signal in the time domain sampling sequence.

[0221] In one embodiment, the target period point determination submodule is specifically configured to:

[0222] Obtaining a second sampling moment of the candidate periodic point, and obtaining a second time range including the second sampling moment; the second time range is determined based on a preset reference fundamental frequency period;

[0223] If the second time range is within the time domain sampling sequence, the candidate periodic point is determined as the target periodic point.

[0224] In one embodiment, the candidate periodic point acquisition submodule is specifically configured to:

[0225] Obtaining a zero-crossing point in a time-domain sampling sequence of an audio signal;

[0226] Based on each zero-crossing point, a plurality of candidate periodic points in the time-domain sampling sequence of the audio signal are obtained.

[0227] In one embodiment, the apparatus further comprises:

[0228] a signal frame acquisition module, configured to divide the audio signal into frames based on the fundamental frequency period, and obtain a plurality of signal frames based on the framing result;

[0229] A modulation parameter acquisition module, configured to acquire a modulation parameter sequence, wherein the modulation parameter sequence comprises at least two modulation parameters;

[0230] a modulation processing module for mapping the modulation parameters in the modulation parameter sequence to target peak points in the signal frame for each signal frame, thereby obtaining a signal frame after modulation processing;

[0231] The synthesis module is used to synthesize multiple signal frames after pitch shifting to obtain a pitch shifted audio signal.

[0232] In one embodiment, the signal frame acquisition module is specifically configured to:

[0233] For each target peak point in the audio signal, obtaining a target fundamental frequency period corresponding to the target peak point;

[0234] Taking the target peak point as the center and the target fundamental frequency period as the intercept, intercepting the audio signal corresponding to the target peak point from the audio signal to obtain a corresponding audio frame;

[0235] Windowing is performed on each audio frame to obtain multiple signal frames.

[0236] In one embodiment, the apparatus further comprises:

[0237] a DC removal module, configured to obtain an audio signal to be modulated and remove a DC component from the audio signal to obtain a DC-free signal;

[0238] A low-pass filtering module, configured to perform low-pass filtering on the DC-removed signal and perform trend elimination on the signal after the low-pass filtering;

[0239] The trend elimination signal processing module is used to obtain a time domain sampling sequence corresponding to the audio signal based on the trend eliminated audio signal.

[0240] Each module in the aforementioned audio signal processing device may be implemented in whole or in part through software, hardware, or a combination thereof. Each module may be embedded in or independent of a processor in a computer device in the form of hardware, or may be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0241] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 14As shown. The computer device includes a processor, a memory, and a network interface connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store audio data. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, an audio signal processing method is implemented.

[0242] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 15 As shown. The computer device includes a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, an audio signal processing method is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad provided on the computer device housing, or an external keyboard, touchpad or mouse.

[0243] Those skilled in the art will understand that Figure 14 、 15 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0244] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented:

[0245] Acquire multiple peak points in a time domain sampling sequence of an audio signal, and determine a current peak point to be analyzed from the multiple peak points;

[0246] Determine a plurality of candidate peak points associated with a current peak point to be analyzed; the plurality of candidate peak points include the current peak point to be analyzed and at least one peak point adjacent to the current peak point to be analyzed and not previously used as a candidate peak point;

[0247] Obtaining a periodic point probability for each candidate peak point, and determining a target peak point from the plurality of candidate peak points whose periodic point probability satisfies a preset probability condition; the periodic point probability represents a probability that the candidate peak point is a periodic point in the fundamental frequency period of the audio signal;

[0248] Taking the target peak point as the current peak point to be analyzed, returning to the step of determining multiple candidate peak points associated with the current peak point to be analyzed, until all peak points in the time domain sampling sequence are traversed to obtain at least one determined target peak point;

[0249] The fundamental frequency period of the audio signal is determined based on an interval between the at least one target peak point.

[0250] In one embodiment, when the processor executes the computer program, the steps in the other embodiments described above are also implemented.

[0251] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0252] Acquire multiple peak points in a time domain sampling sequence of an audio signal, and determine a current peak point to be analyzed from the multiple peak points;

[0253] Determine a plurality of candidate peak points associated with a current peak point to be analyzed; the plurality of candidate peak points include the current peak point to be analyzed and at least one peak point adjacent to the current peak point to be analyzed and not previously used as a candidate peak point;

[0254] Obtaining a periodic point probability for each candidate peak point, and determining a target peak point from the plurality of candidate peak points whose periodic point probability satisfies a preset probability condition; the periodic point probability represents a probability that the candidate peak point is a periodic point in the fundamental frequency period of the audio signal;

[0255] Taking the target peak point as the current peak point to be analyzed, returning to the step of determining multiple candidate peak points associated with the current peak point to be analyzed, until all peak points in the time domain sampling sequence are traversed to obtain at least one determined target peak point;

[0256] The fundamental frequency period of the audio signal is determined based on an interval between the at least one target peak point.

[0257] In one embodiment, when the computer program is executed by a processor, the steps in the other embodiments described above are also implemented.

[0258] In one embodiment, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the following steps:

[0259] Acquire multiple peak points in a time domain sampling sequence of an audio signal, and determine a current peak point to be analyzed from the multiple peak points;

[0260] Determine a plurality of candidate peak points associated with a current peak point to be analyzed; the plurality of candidate peak points include the current peak point to be analyzed and at least one peak point adjacent to the current peak point to be analyzed and not previously used as a candidate peak point;

[0261] Obtaining a periodic point probability for each candidate peak point, and determining a target peak point from the plurality of candidate peak points whose periodic point probability satisfies a preset probability condition; the periodic point probability represents a probability that the candidate peak point is a periodic point in the fundamental frequency period of the audio signal;

[0262] Taking the target peak point as the current peak point to be analyzed, returning to the step of determining multiple candidate peak points associated with the current peak point to be analyzed, until all peak points in the time domain sampling sequence are traversed to obtain at least one determined target peak point;

[0263] The fundamental frequency period of the audio signal is determined based on an interval between the at least one target peak point.

[0264] In one embodiment, when the computer program is executed by a processor, the steps in the other embodiments described above are also implemented.

[0265] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0266] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.

[0267] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0268] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A method for processing an audio signal, characterized in that: The method comprises: Acquire multiple peak points in a time domain sampling sequence of an audio signal, and determine a current peak point to be analyzed from the multiple peak points; Determining a plurality of candidate peak points associated with a current peak point to be analyzed; the plurality of candidate peak points being the current peak point to be analyzed and at least one peak point that is ranked after the current peak point to be analyzed, is adjacent to the current peak point to be analyzed, and has not been a candidate peak point; Obtaining a periodic point probability for each candidate peak point, and determining a target peak point from the plurality of candidate peak points whose periodic point probability satisfies a preset probability condition; the periodic point probability represents a probability that the candidate peak point is a periodic point in the fundamental frequency period of the audio signal; Taking the target peak point as the current peak point to be analyzed, returning to the step of determining multiple candidate peak points associated with the current peak point to be analyzed, and obtaining at least one target peak point in the time domain sampling sequence; The fundamental frequency period of the audio signal is determined based on an interval between the at least one target peak point.

2. The method according to claim 1, characterized in that The obtaining of the periodic point probability of each candidate peak point includes: Obtaining an observation probability of the candidate peak point; the observation probability represents a probability that the candidate peak point is a peak point within a preset sequence range of the time domain sampling sequence; Obtaining a state transition probability of the candidate peak point; the state transition probability represents a probability that the periodic point position determined as the fundamental frequency period is transferred from other candidate peak points to the candidate peak point, wherein the other candidate peak points are candidate peak points before the candidate peak point among the multiple candidate peak points; Based on the observation probability and the state transition probability, the periodic point probability of the candidate peak point is determined.

3. The method according to claim 2, characterized in that The obtaining of the observation probability of the candidate peak point includes: Obtaining a first sampling moment of the candidate peak point, and obtaining a first time range in the time domain sampling sequence that includes the first sampling moment; the first time range is determined based on a preset reference fundamental frequency period; determining a signal amplitude fluctuation range based on a minimum value and a maximum value of the audio signal in the first time range; A signal amplitude difference between the candidate peak point and the minimum value is obtained, and an observation probability of the candidate peak point is determined based on a ratio of the signal amplitude difference to the signal amplitude fluctuation range.

4. The method according to claim 1, wherein The step of obtaining a plurality of peak points in a time domain sampling sequence of an audio signal includes: Acquire multiple candidate periodic points in a time domain sampling sequence of an audio signal; the candidate periodic points are sampling points with periodic characteristics in the time domain sampling sequence; Determine a target periodic point among the multiple candidate periodic points; the target periodic point is a candidate periodic point in the time domain sampling sequence corresponding to a complete signal period; A peak point in the signal period of each target period point is obtained to obtain multiple peak points of the audio signal in the time domain sampling sequence.

5. The method according to claim 4, characterized in that The determining of a target periodic point from the plurality of candidate periodic points includes: Obtaining a second sampling moment of the candidate periodic point, and obtaining a second time range including the second sampling moment; the second time range is determined based on a preset reference fundamental frequency period; If the second time range is within the time domain sampling sequence, the candidate periodic point is determined as the target periodic point.

6. The method according to claim 4, characterized in that The step of obtaining multiple candidate periodic points in the time domain sampling sequence of the audio signal includes: Obtaining a zero-crossing point in a time-domain sampling sequence of an audio signal; Based on each zero-crossing point, a plurality of candidate periodic points in the time-domain sampling sequence of the audio signal are obtained.

7. The method according to any one of claims 1 to 6, characterized in that After determining the fundamental frequency period of the audio signal based on the interval between the target peak points, the method further includes: framing the audio signal based on the fundamental frequency period, and obtaining a plurality of signal frames based on the framing result; Acquire a modulation parameter sequence, where the modulation parameter sequence includes at least two modulation parameters; For each signal frame, mapping the modulation parameters in the modulation parameter sequence to target peak points in the signal frame to obtain a modulation-processed signal frame; A plurality of pitch-shifted signal frames are synthesized to obtain a pitch-shifted audio signal.

8. The method according to claim 7, characterized in that The framing of the audio signal based on the fundamental frequency period and obtaining a plurality of signal frames based on the framing result includes: For each target peak point in the audio signal, obtaining a target fundamental frequency period corresponding to the target peak point; Taking the target peak point as the center and the target fundamental frequency period as the intercept, intercepting the audio signal corresponding to the target peak point from the audio signal to obtain a corresponding audio frame; Windowing is performed on each audio frame to obtain multiple signal frames.

9. The method according to claim 1, characterized in that Before acquiring a plurality of peak points in the time domain sampling sequence of the audio signal, the method further includes: Acquire an audio signal to be modulated, and perform DC component removal processing on the audio signal to obtain a DC-free signal; performing low-pass filtering on the DC-removed signal, and performing trend elimination on the signal after the low-pass filtering; Based on the audio signal after trend elimination, a time domain sampling sequence corresponding to the audio signal is obtained.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 are implemented.

11. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Method for estimating fractional base sound of code excitation linear predicted speech encoder

    CN101030380A

  • Voice signal processing method and apparatus

    CN105845146A