Audio processing method, computer device and storage medium
Patent Information
- Application Number
- CN202410013744.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-04
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2044-01-04
AI Technical Summary
然而,通过多种乐器多次录制得到不同音色的歌曲音频,会导致不同音色的歌曲音频的生成效率降低
[0039]上述音频处理方法、计算机设备、存储介质和计算机程序产品,通过基于音频的旋律的声源信号,得到旋律的音高序列,确定目标音色与音程的各音高对应的音频数据,并将音高序列与音频数据对应的各音高进行音高匹配,根据音高匹配结果确定音高序列中各音高对应的音色片段音频数据,根据音色片段音频数据得到具有所述目标音色的目标音频。相较于传统的使用多种乐器多次录制的方式,本方案通过结合音频中主体声源信号的音高,与具有目标音色的音频数据进行匹配,从而基于匹配得到的音色片段音频数据生成具有目标音色的目标音频,提高了对歌曲音频音色处理的处理效率。
Smart Images

Figure CN117789679B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing technology, and in particular to an audio processing method, computer device, storage medium, and computer program product. Background Technology
[0002] With social and cultural progress, musical aesthetics have become increasingly diverse, with various musical styles and genres emerging and attracting a large audience. Musical style is a high-dimensional expression of timbre, requiring musicians to produce audio tracks with multiple timbres for each song. Currently, altering a song's timbre typically involves multiple recordings using different instruments. However, this method of obtaining different timbres through multiple recordings with various instruments reduces the efficiency of generating audio tracks with diverse timbres.
[0003] Therefore, current methods for processing the timbre of song audio suffer from low processing efficiency. Summary of the Invention
[0004] Therefore, it is necessary to provide an audio processing method, computer device, computer-readable storage medium, and computer program product that can improve processing efficiency in response to the above-mentioned technical problems.
[0005] In a first aspect, this application provides an audio processing method, the method comprising:
[0006] Based on the sound source signal of the melody in the audio, the pitch sequence of the melody is obtained;
[0007] Determine the audio data corresponding to each pitch of the target timbre and interval;
[0008] Based on each pitch in the pitch sequence of the melody, the audio data segment corresponding to each pitch in the pitch sequence is determined in the audio data as timbre segment audio data;
[0009] Based on the audio data of the timbre segments corresponding to each pitch in the pitch sequence, a target audio with the target timbre is obtained.
[0010] In one embodiment, the audio data corresponding to each pitch of the target timbre and interval includes:
[0011] For each pitch in the interval for the target timbre, obtain the original audio data of the pitch;
[0012] Based on the temporal dynamic information of pitch, audio data corresponding to the preset temporal dynamic is extracted from the original audio data to obtain the audio data corresponding to each pitch.
[0013] In one embodiment, obtaining the target audio with the target timbre based on the timbre segment audio data corresponding to each pitch in the pitch sequence includes:
[0014] Obtain the duration of each pitch in the pitch sequence;
[0015] Based on the duration of each pitch in the pitch sequence, the audio data of the timbre segment corresponding to each pitch in the pitch sequence is processed by changing the speed without changing the pitch.
[0016] Based on the audio data of each timbre segment after speed-changing but pitch-unchanged processing, a target audio with the target timbre is obtained.
[0017] In one embodiment, the method further includes:
[0018] Determine multiple preset chord combinations;
[0019] For each preset chord combination, extract the segment audio data combination corresponding to the preset chord combination from the audio data corresponding to each pitch; the pitch of the segment audio data in the segment audio data combination matches the pitch of the preset chord combination.
[0020] The step of determining the segment audio data corresponding to each pitch in the pitch sequence of the melody as timbre segment audio data in the audio data, based on each pitch in the pitch sequence of the melody, includes:
[0021] The pitch sequence is grouped and matched with the multiple preset chord combinations. Based on the pitch matching results, the segment audio data combination corresponding to each pitch sequence group in the pitch sequence is determined from the multiple segments of audio data combination, and used as the timbre segment audio data.
[0022] In one embodiment, the step of performing grouped pitch matching between the pitch sequence and the plurality of preset chord combinations, and determining the segment audio data combination corresponding to each pitch sequence group in the pitch sequence from the plurality of segment audio data combinations based on the pitch matching results, includes:
[0023] The pitch sequence is grouped according to a preset duration to obtain multiple pitch sequence groups;
[0024] For each pitch sequence group, each pitch in the pitch sequence group is matched with each pitch in each preset chord combination to obtain the number of pitch matches between the pitch sequence group and each preset chord combination.
[0025] Based on the number of pitch matches, determine the target preset chord combination that matches the pitch sequence group in each group of preset chord combinations;
[0026] The audio data combination corresponding to the target preset chord combination is determined from multiple sets of audio data combinations.
[0027] In one embodiment, obtaining the target audio with the target timbre based on the timbre segment audio data corresponding to each pitch in the pitch sequence includes:
[0028] Based on the corresponding preset duration, the audio data of the timbre segments corresponding to each pitch sequence group are processed to change speed without changing pitch.
[0029] Based on the audio data of each timbre segment after speed-changing but pitch-unchanged processing, a target audio with the target timbre is obtained.
[0030] In one embodiment, obtaining the target audio with the target timbre based on the timbre segment audio data corresponding to each pitch in the pitch sequence includes:
[0031] The target audio is obtained by superimposing the audio data of adjacent timbre segments according to their corresponding delay durations.
[0032] In one embodiment, the target timbre includes at least two target timbres;
[0033] The step of obtaining the target audio with the target timbre based on the timbre segment audio data corresponding to each pitch in the pitch sequence includes:
[0034] Based on the audio data of each timbre segment corresponding to the at least two target timbres, at least two audio files are obtained;
[0035] Based on the linear addition of the at least two audio samples, a target audio sample with the target timbre is obtained.
[0036] Secondly, this application provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method.
[0037] Thirdly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.
[0038] Fourthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method.
[0039] The aforementioned audio processing method, computer equipment, storage medium, and computer program product obtain a melody pitch sequence from the sound source signal of the melody, determine the audio data corresponding to each pitch of the target timbre and interval, match the pitch sequence with each pitch corresponding to the audio data, determine the timbre segment audio data corresponding to each pitch in the pitch sequence based on the pitch matching result, and obtain the target audio with the target timbre based on the timbre segment audio data. Compared with the traditional method of recording multiple times using multiple instruments, this solution improves the processing efficiency of song audio timbre processing by combining the pitch of the main sound source signal in the audio with the audio data with the target timbre and matching it with the audio data with the target timbre. Attached Figure Description
[0040] Figure 1 This is a flowchart illustrating an audio processing method in one embodiment;
[0041] Figure 2 This is a schematic diagram of the sound source separation step in one embodiment;
[0042] Figure 3 This is a schematic diagram of the baseband information in one embodiment;
[0043] Figure 4 This is a schematic diagram of audio data in one embodiment;
[0044] Figure 5 This is a schematic diagram of audio data in another embodiment;
[0045] Figure 6 This is a flowchart illustrating the steps for determining audio data of a timbre segment in one embodiment;
[0046] Figure 7 This is a schematic diagram illustrating duration scaling in one embodiment;
[0047] Figure 8 This is a schematic diagram of duration scaling in another embodiment;
[0048] Figure 9 This is a schematic diagram of pitch sequence group matching in one embodiment;
[0049] Figure 10 This is a schematic diagram of audio data overlay in one embodiment;
[0050] Figure 11 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0051] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0052] In one embodiment, such as Figure 1 As shown, an audio processing method is provided. This embodiment illustrates the method by applying it to a terminal. It is understood that the method can also be applied to a server, or to a system including both a terminal and a server, and is implemented through the interaction between the terminal and the server. The method includes the following steps:
[0053] Step S202: Obtain the pitch sequence of the melody based on the sound source signal of the audio melody.
[0054] The audio can be a type of audio to be processed, specifically stereo audio. Stereo audio can have multiple channels, such as left and right channels. Furthermore, the audio can be composed of various sound source signals. These sound source signals can be signals generated by sound-producing devices, and there is a one-to-one correspondence between sound source signals and sound-producing devices. That is, one sound-producing device can emit one type of sound source signal. To change the timbre of the audio, the terminal can determine the melody's sound source signal from the various sound source signals in the audio. For example, the terminal can pre-train a sound source separation model. The terminal inputs the audio containing multiple channels into the trained sound source separation model, which then separates the audio signals into sound sources and outputs the corresponding sound source signals for each sound source contained in the audio.
[0055] Specifically, such as Figure 2 As shown, Figure 2 This is a schematic diagram of the sound source separation step in one embodiment. The sound source separation model described above can be a neural network model, and the sound sources of the audio can include, but are not limited to, human voice, guitar, piano, bass, drums, etc. The terminal can then obtain multiple sound source signals from the audio through the sound source separation model. Furthermore, the separated sound source signals can be linearly superimposed to reconstruct the audio.
[0056] The terminal can identify the melody's sound source signal from the various sound source signals mentioned above. The melody represents the most prominent and easily perceived main sound source in the audio. For example, in a popular song, the melody's sound source signal can be represented by the vocal sound source signal. It should be noted that the sound source corresponding to the melody can be different for different audio tracks; for instance, the melody's sound source signal in a piano piece can be the piano's sound source signal.
[0057] To ensure that the audio after timbre transformation retains its original characteristics, the terminal needs to extract the melody's source signal. After extracting the melody's source signal, the terminal can use it for timbre conversion. This timbre conversion process can be based on pitch conversion. The terminal can then obtain the melody's pitch sequence based on the melody's source signal. This pitch sequence can include various information, such as the pitch's sign value, duration, and energy value. The pitches in the sequence can be ordered according to their position and duration within the source signal.
[0058] Pitch is determined by the fundamental frequency of sound, and any vocalization involving regular vibrations has a fundamental frequency. Therefore, the terminal can extract the fundamental frequency from the sound source signal of the melody and determine the pitch sequence based on the fundamental frequency. Taking the time-domain method as an example, since frequency and period are inversely proportional, if it is necessary to determine the fundamental frequency from the sound source signal of the melody, the terminal can do so by finding the minimum positive period of the waveform of the sound source signal. That is, the terminal can find how much the signal needs to be shifted to achieve the highest degree of overlap with the original signal.
[0059] For example, in one embodiment, the terminal can perform frame-by-frame processing on the sound source signal of the aforementioned melody. For each frame of the sound source signal, the terminal can obtain the peak shift amount corresponding to the sound source signal in that frame's waveform, and obtain a preset number of sampling points in the waveform of that frame's sound source signal. The terminal can shift the signals corresponding to the preset number of sampling points according to the peak shift amount to obtain the shifted signal. Then, based on the product of the signals corresponding to the preset number of sampling points and the shifted signal, and the sum of the product and the peak shift amount, the terminal obtains the minimum positive period corresponding to that frame of the melody signal. Thus, the terminal can obtain the frequency corresponding to that frame of the sound source signal based on the minimum positive period, and obtain the fundamental frequency information corresponding to the sound source signal based on the multiple frequencies corresponding to multiple frames of sound source signals.
[0060] Specifically, such as Figure 3 As shown, Figure 3 This is a schematic diagram of fundamental frequency information in one embodiment. The minimum positive period of the sound source signal of the melody, as determined by the terminal, can be represented by the following function: .in, Represents the minimum positive period of the sound source signal. The shift corresponding to the peak value of the autocorrelation function of the sound source signal of the above melody can represent the period of the signal at point t. In the above formula, W is the number of sampling points in one frame. iLet represent the sound source signal at time i. The terminal multiplies the shifted signal with the original signal, then sums the result with the period within one frame. The terminal then obtains the period by taking the peak value of the autocorrelation function. Furthermore, the terminal obtains the fundamental frequency corresponding to one frame of the sound source signal by inversely proportionaling the aforementioned minimum positive period. By integrating the fundamental frequencies of multiple frames of sound source signals, the terminal can obtain, as shown below... Figure 3 The image shown is a fundamental frequency image. It should be noted that the terminal can use any fundamental frequency extraction method to extract the fundamental frequency from the sound source signal of the melody.
[0061] Furthermore, after determining the fundamental frequency information, the terminal can obtain the pitch sequence of the melody based on the fundamental frequency information. The function by which the terminal converts the fundamental frequency information into pitch can be specifically expressed as: p = 69 + 12 * log2(f / 440). Here, p is the pitch, and f is the frequency, i.e., the frequency corresponding to each frame in the fundamental frequency. Pitch is a crucial fundamental feature describing musical melody, so the terminal can determine the pitch of each frame based on the frequency of each frame in the fundamental frequency. Moreover, the terminal can record each extracted pitch, the duration of the pitch, and the audio energy value within that duration using a data structure to form the aforementioned pitch sequence. This data structure includes, but is not limited to, MIDI format, JSON format, or multidimensional arrays.
[0062] Step S204: Determine the audio data corresponding to each pitch of the target timbre and interval.
[0063] The target timbre can be the timbre that the audio needs to be converted to, and the interval can refer to an octave in music. For example, converting a pop song with vocals as the main melody into a piano timbre. The target timbre can be input by the user of the terminal or determined by the terminal according to a certain strategy. The terminal can pre-determine the audio data segments corresponding to the target timbre. These audio segments can be sound files representing pitches, and there can be a one-to-one correspondence between audio segments and pitches. Multiple audio segments can form audio data corresponding to the pitches of each interval; that is, the audio data can be a collection of multiple audio segments. After determining the target timbre, the terminal can extract multiple audio segments under that timbre. Each audio segment represents the audio data segment of each pitch presented by the target timbre within the interval (to distinguish it from the pitches in the pitch sequence, the pitch of the target timbre within the interval can be called the first pitch, and the pitch in the pitch sequence can be called the second pitch). Taking a piano as an example, the audio data segments of the piano can be represented as the pitches emitted by each key on the piano, that is, the pitches corresponding to the sounds produced by the piano timbre. The terminal can acquire audio data of the target timbre at multiple pitches and generate multiple files.
[0064] The target timbre can include a variety of timbres. For popular songs, in addition to the vocals as the main melody, the terminal can also add accompaniment to the audio after the timbre is finally converted, such as adding chords. In this case, the terminal also needs to determine the timbre of the chord as the target timbre, and further determine the audio data of each segment corresponding to the chord timbre, such as chord combinations.
[0065] Step S206: Based on each pitch in the pitch sequence of the melody, determine the segment audio data corresponding to each pitch in the pitch sequence as timbre segment audio data in the audio data.
[0066] To convert the pitches in the aforementioned melody's pitch sequence into pitches that satisfy the target timbre, the terminal can perform pitch matching between the melody's pitch sequence and the audio data corresponding to the target timbre. Based on the matching result, the terminal determines the corresponding audio segment data as timbre segment audio data. Timbre segment audio data refers to the segment audio data matched from the target timbre's audio data (audio data corresponding to each pitch of the interval). Each segment audio data can be an audio file with a corresponding pitch. The terminal can match each second pitch in the pitch sequence with the first pitch in each segment audio data. The melody can be a main part separated from the audio, such as a human voice. The terminal can then match the second pitch in the human voice's pitch sequence with each of the first pitches to obtain the timbre segment audio data.
[0067] Step S208: Based on the audio data of the timbre segments corresponding to each pitch in the pitch sequence, obtain the target audio with the target timbre.
[0068] The aforementioned timbre segment audio data can be segment audio data corresponding to the pitch of the melody under the target timbre. The aforementioned melody can be the main part separated from the audio, such as a human voice. The terminal can combine the target timbre segment audio data that matches the second pitch in the pitch sequence of the human voice to form a sequence of timbre segment audio data. For example, the terminal extracts the timbre segment audio data corresponding to each second pitch from the audio data, combines the timbre segment audio data to obtain a sequence of timbre segment audio data, i.e., a sequence of melody timbre segment audio data. That is, this sequence of timbre segment audio data can be a sequence composed of timbre segment audio data that satisfies the target timbre, and each segment audio data matches each pitch in the pitch sequence.
[0069] In addition, the terminal can also perform pitch matching between the pitch sequence of the melody and the audio data of each segment corresponding to the target timbre of the chord. For example, the terminal can match the pitch sequence with the pitch in each chord combination corresponding to the chord to obtain the audio data of the timbre segment corresponding to each pitch sequence group in the pitch sequence.
[0070] After obtaining the timbre fragment data of the target timbre, the terminal can generate target audio with the target timbre based on the timbre fragment audio data corresponding to each pitch in the pitch sequence. For example, if there is only one set of timbre fragment audio data corresponding to each pitch in the pitch sequence, taking the timbre conversion of vocals in a popular song as an example, if the timbre fragment audio data corresponding to each pitch in the pitch sequence only contains the timbre fragment audio data corresponding to the melody, then the terminal can generate audio based on the timbre fragment audio data corresponding to each pitch in the pitch sequence as the target audio.
[0071] In this embodiment, the aforementioned timbre segment audio data can represent segment audio data corresponding to each pitch in the pitch sequence under the target timbre. This timbre segment audio data can include various types, each type being obtained by matching different numbers of pitches in the pitch sequence. For example, timbre segment audio data matched to each pitch in the melody, and timbre segment audio data matched to each group of pitches in the melody (such as timbre segment audio data corresponding to chords). A song audio can be separated into multiple sound sources using a neural network, including the melody sound source representing the main part of the audio, and the sound sources of the auxiliary parts of the audio. The melody represents the most prominent main sound source in the audio. Taking a popular song as an example, the melody sound source signal can be represented by the vocal sound source signal in the audio. Therefore, the aforementioned timbre segment audio data corresponding to the melody can be obtained by matching each pitch in the vocal pitch sequence, or by matching each group of vocal pitches. Therefore, the terminal can superimpose audio data of various timbre segments corresponding to a pitch sequence to obtain target audio with the target timbre. The target audio can represent the audio obtained after superimposing all the timbre segment audio data. Additionally, in some embodiments, the terminal can also mix the target audio, such as by adding effects to the target audio or performing volume balancing.
[0072] In the aforementioned audio processing method, a pitch sequence of the melody is obtained from the sound source signal based on the audio melody. The target timbre and the corresponding audio data for each pitch of the interval are determined. The pitch sequence and the corresponding pitches of the audio data are then matched. Based on the pitch matching results, audio data of timbre segments corresponding to each pitch in the pitch sequence are determined. Finally, a target audio with the target timbre is obtained based on the audio data of the timbre segments corresponding to each pitch in the pitch sequence. Compared to the traditional method of recording multiple times using multiple instruments, this solution improves the processing efficiency of song audio timbre processing by combining the pitch of the main sound source signal in the audio with the audio data containing the target timbre and matching them.
[0073] In one embodiment, determining the audio data corresponding to each pitch of the target timbre and the interval includes: acquiring the original audio data of each pitch of the target timbre in the interval; and extracting the audio data corresponding to the preset temporal dynamics from the original audio data based on the temporal dynamics information of the pitch to obtain the audio data corresponding to each pitch.
[0074] In this embodiment, the terminal can obtain the corresponding audio data based on the waveforms generated when the target timbre is emitted at each pitch, that is, the terminal can obtain audio data in the time domain. For example, the terminal can determine each pitch (first pitch) of the target timbre in the interval. For each first pitch, the original audio data of the first pitch is obtained, and based on the time-domain dynamic information of the pitch, the audio data corresponding to the preset time-domain dynamic information in the original audio data is extracted to obtain the audio data corresponding to each pitch of the interval for the target timbre. Here, the aforementioned time-domain dynamic information represents the dynamic change information of the pitch in the time domain. Taking the target timbre as the timbre required for the melody as an example, specifically such as a piano, for example... Figure 4 As shown, Figure 4 This is a schematic diagram of audio data in one embodiment. The terminal can collect audio data of the target timbre within multiple intervals, such as collecting audio data within four octaves, and save the audio segments from the audio data as separate audio files, such as saving each semitone as a lossless format like WAV or FLAC. Figure 4 The CDEFGABC in the diagram represent pitch values within an octave. Each semitone corresponds to a unique fundamental frequency value. In the time domain, the dynamic information of each pitch includes multiple dynamic ranges, as detailed below. Figure 5 As shown, Figure 5 This is a schematic diagram of audio data in another embodiment. Figure 5The image shows audio data in the time domain. This audio data includes four dynamic states: attack, decay, sustain, and release. These states can be collectively referred to as wave seals, which are parameters that roughly outline the waveform of a timbre to represent its characteristics in terms of volume changes.
[0075] like Figure 6 As shown, Figure 6 This is a flowchart illustrating the steps for determining timbre segment audio data in one embodiment. The terminal can use the pitch sequence of the melody described above to match the corresponding timbre segment audio data from the audio data. The terminal can pre-construct an empty array as the basis for combining the timbre segment audio data. Taking the pitch sequence [..., C4:2, A5:5, E3:6,...] as an example, C4, A5, and E3 in the sequence represent pitches at different intervals, and the numbers carried by the pitches indicate the duration of that pitch in the original melody. The terminal, through a mapping relationship, matches the timbre segment audio data corresponding to the pitches C4, A5, and E3 from the target timbre audio data, reads the corresponding audio file, and fills it into the empty array according to the order of the corresponding pitches in the pitch sequence, thus obtaining the timbre segment audio data corresponding to each pitch in the pitch sequence.
[0076] Since the duration of each timbre segment audio data may not match the duration of the corresponding pitch in the pitch sequence, the terminal can also scale the timbre segment audio data over time to match the duration of the corresponding pitch in the pitch sequence.
[0077] In one embodiment, obtaining target audio with a target timbre based on the timbre segment audio data corresponding to each pitch in the pitch sequence may include: obtaining the duration of each pitch in the pitch sequence; performing speed-changing and pitch-invariant processing on the timbre segment audio data corresponding to each pitch in the pitch sequence based on the duration of each pitch in the pitch sequence; and obtaining target audio with a target timbre based on the speed-changing and pitch-invariant processed timbre segment audio data.
[0078] In this embodiment, the terminal can perform speed-variable and pitch-invariant processing on the audio data of the timbre segments corresponding to each pitch (second pitch) in the pitch sequence based on the duration of each pitch (second pitch) in the pitch sequence, and obtain the target audio with the target timbre based on the speed-variable and pitch-invariant audio data of each timbre segment after the speed-variable and pitch-invariant processing. Here, speed-variable and pitch-invariant processing means scaling the duration of the timbre segment audio data. Various methods can be used to perform speed-variable and pitch-invariant processing on the above-mentioned timbre segment audio data, such as... Figure 7 As shown, for example, it can be done through the principle of waveform similarity superposition, etc. Figure 7 The terminal takes an input audio signal x, divides the input signal into frames and applies a window to obtain x'. mThe first frame signal x' m The signal is directly output to the output signal y to form y m and through the window Read each frame of the x signal in a loop and find the frame that is most similar to the second frame, for example... Figure 7 x' in m +1. The terminal superimposes this onto the second frame of the output signal y, and so on, until each frame of the signal that requires duration scaling has completed the speed-variable pitch-invariant processing.
[0079] In some embodiments, the above-described speed-changing algorithm can also be implemented using a frequency-domain-based speed-changing algorithm with invariant pitch. For example... Figure 8 As shown, Figure 8 This is a schematic diagram of duration scaling in another embodiment. The terminal can first perform a Fourier transform on the time-domain audio signal to obtain the audio signal in the frequency domain. In the frequency domain, the terminal takes the first frame of the audio signal as the first frame signal x(m,k). Then, it selects the spectrum of the second frame signal according to the speed of the transformation. For example, if the spectrum of the first frame is φ1=φ(m,k), the terminal obtains the spectrum of the second frame signal as φ2=φ(m+1,k) by changing the speed. Then, the terminal can reassemble a frame according to the phase corresponding to the time of the second frame signal. For example, if the first frame is X(m,k), the second frame can be composed as X(m+1,k). The terminal can repeat the above process until each frame of the audio signal to be scaled has been processed with speed-changing and pitch-invariant methods. Based on the audio data of each timbre segment after speed-changing and pitch-invariant processing, the duration-scaled target audio is obtained. The target audio can be represented by various types of audio, such as the timbre-transformed and duration-scaled target audio corresponding to the melody of the above-mentioned human voice, or the timbre-transformed and duration-scaled target audio corresponding to the melody of a piano piece.
[0080] Through the above embodiments, and by performing pitch matching and speed-changing without pitch-changing on the pitch sequence and audio data of each timbre segment, audio data of each timbre segment that satisfies the target timbre and the duration of each pitch in the pitch sequence are formed, thereby improving the processing effect of song audio timbre processing.
[0081] In one embodiment, the method of this application may further include: determining multiple sets of preset chord combinations; for each set of preset chord combinations, extracting a combination of segment audio data corresponding to the preset chord combination from the audio data corresponding to each pitch; and matching the pitch of the segment audio data in the segment audio data combination with the pitch of the set of preset chord combinations.
[0082] In this embodiment, the terminal can utilize the audio data combinations of chord segments corresponding to chord combinations to form audio of chords with a target timbre. Specifically, the terminal can obtain each audio data combination of segments based on the aforementioned preset chord combinations. For example, the terminal first determines the target timbre, grouping the pitches in the preset chord combinations into sets; each set of pitches under the target timbre corresponds to a set of audio data segments. The terminal can predetermine multiple preset chord combinations, where each chord combination contains multiple pitches. Based on these multiple preset chord combinations, the terminal can match the audio data segments corresponding to each preset chord combination from the audio data corresponding to each pitch (first pitch) of the target timbre and interval. The pitches corresponding to the audio data segments in the audio data segments match the pitches of that preset chord combination. For example, the terminal can pre-set multiple chord combinations based on the key and tonality of the song audio, such as major triads, minor triads, augmented triads, etc., specifically represented as: [0, 4, 7], m: [0, 3, 7], +: [0, 4, 8], dim: [0, 3, 6], 7: [0, 4, 7, 10], maj7: [0, 4, 7, 11], m7: [0, 3, 7, 10], m7b5: [0, 3, 6, 10]. Here, m, +, dim, 7, maj7, m7, and m7b5 represent the names of the chord combinations. The numbers in each chord combination represent the pitch. It should be noted that the terminal can also select other chord combinations that conform to the above tonality and pitch as preset chord combinations.
[0083] Therefore, in one embodiment, determining the segment audio data corresponding to each pitch in the pitch sequence as timbre segment audio data in the audio data based on each pitch in the pitch sequence of the melody may include: performing pitch matching between the pitch sequence and multiple preset chord combinations, and determining the segment audio data combination corresponding to each pitch sequence group in the pitch sequence from the multiple segments of audio data combinations based on the pitch matching results, as timbre segment audio data.
[0084] In this embodiment, the timbre segment audio data refers to the combination of segment audio data that matches the pitches in the preset chord combinations. For example, the terminal can determine the timbre segment audio data based on the number of pitch matches in the pitch matching result. Preset chord combinations can be grouped together, and each group of preset chord combinations (containing multiple pitches) can be matched with multiple pitches in a pitch sequence (grouped pitch matching). Thus, under each grouped pitch matching, a pitch matching result can be obtained. This pitch matching result can include the number of pitch matches, which refers to the number of matches between multiple pitches in the pitch sequence and multiple pitches in the preset chord combinations in each round of matching. One round of matching refers to the matching of a group of pitch sequences (containing multiple pitches, such as three consecutive pitches) in a pitch sequence with multiple pitches in a preset chord combination. This allows us to obtain the pitch matching results for each preset chord combination in each round of matching. Then, based on the number of pitch matches in each preset chord combination's pitch matching results, we can determine the segment audio data combination corresponding to that pitch sequence group from multiple segments of audio data combinations, thus obtaining the timbre segment audio data corresponding to that pitch sequence group. After multiple rounds of matching, we can obtain the timbre segment audio data corresponding to each pitch (in units of pitch sequence groups) in the pitch sequence.
[0085] In one embodiment, pitch sequence is grouped and pitch matched with multiple preset chord combinations. Based on the pitch matching results, the segment audio data combination corresponding to each pitch sequence group in the multiple segments of audio data combinations is determined. This may include:
[0086] The pitch sequence is grouped according to a preset duration to obtain multiple pitch sequence groups. For each pitch sequence group, each pitch in the pitch sequence group is matched with each pitch in each preset chord combination to obtain the number of pitch matches between the pitch sequence combination and each preset chord combination. Based on the number of pitch matches, the target preset chord combination that matches the pitch sequence group is determined. The segment audio data combination corresponding to the target preset chord combination is determined from multiple segments of audio data combination.
[0087] The terminal can group pitch sequences according to a preset duration, resulting in multiple pitch sequence groups. The preset duration can be a unit duration for dividing the pitch sequence into equal time intervals; this preset duration can be pre-set or determined according to the audio's beat information. Each pitch sequence group contains one or more pitches, and each pitch sequence group corresponds to a specific pitch duration.
[0088] For each pitch sequence group, the terminal can perform pitch matching between each pitch (second pitch) in the pitch sequence group and each pitch (third pitch) in each preset chord combination, obtaining the number of pitch matches between the pitch sequence group and each preset chord combination. The number of pitch matches indicates how many second pitches in the pitch sequence group match the third pitches in the preset chord combinations. Based on the number of pitch matches, the terminal can determine the preset chord combination (target preset chord combination) that matches the pitch sequence group in each preset chord combination. From multiple sets of audio segment combinations, the terminal can determine the audio segment combination corresponding to the target preset chord combination, thus obtaining the timbre segment audio data corresponding to the pitch sequence group. For the number of pitch matches, the terminal selects the preset chord combination with the largest number of pitch matches as the target preset chord combination for that pitch sequence group. Therefore, the terminal can obtain the timbre segment audio data corresponding to each pitch in the pitch sequence (based on the pitch sequence group) based on the timbre segment audio data corresponding to each pitch sequence group.
[0089] As an example, such as Figure 9 As shown, Figure 9 This is a schematic diagram of pitch sequence group matching in one embodiment. The preset duration can be represented as ΔT. The terminal can use ΔT to group the pitch sequence to obtain multiple pitch sequence groups. Furthermore, the terminal can normalize the pitches in the above pitch sequence to within one octave by traversing. Taking a piano as an example, the 88 pitches of a piano are derived from a seven-tone scale. The 1 in one group and the 1 in the next group are octaves apart, that is, their fundamental frequencies are twice each other. The terminal can normalize these 88 notes to one octave, that is, 12 pitches. Thus, the terminal matches each pitch sequence group with each preset chord combination according to a fixed period duration, calculates the number of times each pitch (represented by a note) in the pitch sequence group falls into each pitch in each preset chord combination, and selects the segment audio data combination corresponding to the preset chord combination with the most falling notes (target preset chord combination) as the segment audio data combination for matching that pitch sequence group. The terminal then matches each pitch sequence group in the pitch sequence to obtain a combination of audio data for each matched segment.
[0090] Since the duration of each pitch in the pitch sequence is different, the length of the aforementioned period ΔT is independent of the number of notes within that period ΔT. The terminal can determine the target timbre corresponding to the chord, match the audio data of the timbre segments corresponding to each pitch in the pitch sequence under that target timbre, read the corresponding audio files, and thus combine the corresponding audio files to obtain the target audio with the target timbre corresponding to the chord.
[0091] Since the duration of the timbre segment audio data corresponding to the above chord combinations may differ from the determined period ΔT, in some embodiments, obtaining the target audio with the target timbre based on the timbre segment audio data corresponding to each pitch in the pitch sequence may include:
[0092] Based on the corresponding preset duration, the audio data of the timbre segments corresponding to each pitch sequence group are processed by speed variation without pitch change; based on the audio data of each timbre segment after speed variation without pitch change processing, the target audio with the target timbre is obtained.
[0093] In this embodiment, the terminal can perform speed-variable and pitch-unchanging processing on the audio data of the timbre segments according to a corresponding preset duration, and obtain a target audio with a target timbre based on the audio data of each timbre segment after speed-variable and pitch-unchanging processing. This target audio can be a target audio with a chordal target timbre, and can serve as accompaniment for a target audio with the aforementioned melody. Furthermore, the target timbre of the chords and the timbre of the melody can be the same or different; for example, the melody timbre can be a piano timbre, while the chord timbre can be a bass timbre, etc.
[0094] Through the above embodiments, the terminal can match the audio data formed by the pitch sequence group and the preset chord combination under the target timbre, and after processing the audio data of the timbre segment by changing the speed without changing the pitch, it can form audio data of each timbre segment with the target timbre, and finally form audio with chord timbre, thereby improving the processing effect of timbre processing of song audio.
[0095] In one embodiment, obtaining target audio with a target timbre based on the timbre segment audio data corresponding to each pitch in the pitch sequence includes: superimposing adjacent timbre segment audio data with corresponding delay durations to obtain target audio with the target timbre.
[0096] In this embodiment, the terminal can superimpose adjacent timbre segment audio data using a partial overlapping method. This partial overlapping can simulate the lingering sustain of the previous note during instrument performance. The terminal can acquire sustain information, which is determined based on the characteristics of the instrument corresponding to the target timbre. This sustain information can include the sustain duration and the position in the pitch sequence where sustain is required for each pitch. Based on the sustain information, the terminal can superimpose a certain number (which can be determined based on the required sustain position) of adjacent timbre segment audio data corresponding to each pitch in the pitch sequence for the corresponding sustain duration, thereby obtaining the target audio with the target timbre.
[0097] Specifically, such as Figure 10 As shown, Figure 10This is a schematic diagram of audio data superposition in one embodiment. Each audio file can represent each segment of audio data, and ΔT1, ΔT2, etc., represent the duration of each segment of audio data. The duration of the overlapping portion between each audio file can be represented by t. The terminal can use ΔT of each audio file as a reference value, and the duration of the duration can be specifically expressed as: t = (1 / 5) * ΔT. The 1 / 5 can be changed according to the characteristics (type and playing method) of the instrument corresponding to the target timbre. For example, for instruments with long sustains and relatively independent sound production mechanisms for each note, this value will be larger, such as the harp and piano; while for transient burst-type instruments with short sustains, this value will be smaller, such as the marimba and flute. Thus, the terminal can superimpose the audio data of adjacent timbre segments according to their corresponding characteristics, and based on the superimposed audio data of each timbre segment, obtain the target audio with the target timbre.
[0098] Through this embodiment, the terminal can superimpose a certain number (the position of the sustain can be determined as needed) of adjacent timbre segment audio data in the timbre segment audio data corresponding to each pitch in the pitch sequence based on the sustain information, thereby improving the coherence of the audio after timbre conversion and improving the processing effect of song audio timbre processing.
[0099] In one embodiment, obtaining target audio with a target timbre based on audio data of timbre segments corresponding to each pitch in a pitch sequence includes: obtaining at least two audios based on audio data of timbre segments corresponding to at least two target timbres respectively; and obtaining target audio with a target timbre based on linear addition of the at least two audios.
[0100] In this embodiment, the target timbre may include at least two timbres, such as the timbre assigned to a melody or the timbre assigned to a chord. These at least two timbres may be the same or different. The terminal can then obtain at least two audio clips based on the audio data of each timbre segment corresponding to the at least two target timbres, and perform linear addition on the at least two audio clips to obtain the target audio with the target timbre. Taking a target timbre that includes both a melody timbre and a chord timbre as an example, after obtaining the audio data of each timbre segment corresponding to the main melody timbre and the chord timbre, the terminal can perform linear addition on the audio data of each timbre segment of the two timbres to obtain the target audio with the target timbre. Furthermore, in some embodiments, the terminal can also perform effects layering, volume balancing, and other processing on the target audio to improve its listening experience.
[0101] Through this embodiment, the terminal can convert the sound source signal of the main melody in the audio into audio data of each timbre segment corresponding to at least two timbres, thereby obtaining the target audio that satisfies the target timbre and improving the processing efficiency of song audio timbre processing. Furthermore, it can also provide listeners with a way to adapt familiar melodies according to their preferred musical style.
[0102] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0103] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 11 As shown, the computer device includes a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When executed by the processor, the computer program implements an audio processing method. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink display screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device casing, or an external keyboard, touchpad, or mouse.
[0104] Those skilled in the art will understand that Figure 11The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0105] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the audio processing method described above.
[0106] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the audio processing method described above.
[0107] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the audio processing method described above.
[0108] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0109] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0110] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0111] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. An audio processing method, characterized in that, The method includes: Based on the sound source signal of the melody in the audio, the pitch sequence of the melody is obtained; Determine the audio data corresponding to each pitch of the target timbre and interval; The pitch sequence is grouped and pitch matched with multiple preset chord combinations. Based on the pitch matching results, the segment audio data combination corresponding to each pitch sequence group in the pitch sequence is determined from the multiple segments of audio data combinations, and is used as the timbre segment audio data. The multiple segments of audio data combinations are obtained by extracting the segment audio data combinations corresponding to each preset chord combination from the audio data corresponding to each pitch. The pitch of the segment audio data in the segment audio data combination matches the pitch of the preset chord combination. Obtaining a target audio with the target timbre based on the timbre segment audio data corresponding to each pitch in the pitch sequence includes: obtaining the target audio with the target timbre based on the timbre segment audio data after speed-changing and pitch-invariant processing; the speed-changing and pitch-invariant processing step includes: performing speed-changing and pitch-invariant processing on the timbre segment audio data corresponding to each pitch in the pitch sequence based on the duration of each pitch; and performing speed-changing and pitch-invariant processing on the timbre segment audio data corresponding to each pitch sequence group based on the corresponding preset duration. It also includes: superimposing adjacent timbre segment audio data with corresponding delay durations to obtain target audio with the target timbre.
2. The method according to claim 1, characterized in that, The audio data corresponding to each pitch of the target timbre and interval includes: For each pitch in the interval for the target timbre, obtain the original audio data of the pitch; Based on the temporal dynamic information of pitch, audio data corresponding to the preset temporal dynamic is extracted from the original audio data to obtain the audio data corresponding to each pitch.
3. The method according to claim 1, characterized in that, The step of performing pitch matching between the pitch sequence and multiple preset chord combinations, and determining the segment audio data combination corresponding to each pitch sequence group in the pitch sequence from multiple segments of audio data combinations based on the pitch matching results, includes: The pitch sequence is grouped according to a preset duration to obtain multiple pitch sequence groups; For each pitch sequence group, each pitch in the pitch sequence group is matched with each pitch in each preset chord combination to obtain the number of pitch matches between the pitch sequence group and each preset chord combination. Based on the number of pitch matches, determine the target preset chord combination that matches the pitch sequence group in each group of preset chord combinations; The audio data combination corresponding to the target preset chord combination is determined from multiple sets of audio data combinations.
4. The method according to any one of claims 1 to 3, characterized in that, The target timbre includes at least two target timbres; obtaining the target audio with the target timbre based on the audio data of the timbre segments corresponding to each pitch in the pitch sequence includes: Based on the audio data of each timbre segment corresponding to the at least two target timbres, at least two audio files are obtained; Based on the linear addition of the at least two audio samples, a target audio sample with the target timbre is obtained.
5. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 4.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Method and device for image transmission and terminal device
CN104918059A
Audio adjustment method, computer device and computer program product
CN114743526A