Audio processing method, computer device and computer program product

By aligning animal sound signals with musical beats and adjusting frequencies to match the melody, the method efficiently integrates animal vocalizations into music, enhancing the musical experience.

CN115862587BActive Publication Date: 2025-07-15TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211509155.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-29
Publication Date
2025-07-15
Estimated Expiration
2042-11-29

AI Technical Summary

Technical Problem

In the prior art, it is inefficient to integrate animal calls into musical melodies and is difficult to adapt quickly.

Method used

By obtaining the initial audio signal and animal call signal, the start playback position of the sound signal is determined based on the beat point of the music melody, and the pitch adjustment process is performed according to the pitch of the melody segment, and finally the mixing process is performed to generate the target audio signal.

Benefits of technology

The rhythm and pitch of animal calls and musical melody are adapted to improve processing efficiency and generate beautiful musical melody containing animal calls.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115862587B_ABST
    Figure CN115862587B_ABST
Patent Text Reader

Abstract

This application relates to audio processing technology and provides an audio processing method, a computer device, and a computer program product, which can quickly generate a pleasant melody containing animal calls. The method includes: obtaining an initial audio signal and sound signals of at least one animal call, where the initial audio signal has a music melody; determining the start playback positions of the sound signals in the initial audio signal based on the beat points of the music melody; determining the melody segments corresponding to the sound signals in the initial audio signal based on the start playback positions and signal durations of the sound signals, and respectively performing pitch adjustment processing on the sound signals based on the pitches of the melody segments corresponding to the sound signals to obtain respective matching sound signals; and mixing the respective matching sound signals with the initial audio signal based on the start playback positions corresponding to the respective matching sound signals to obtain a target audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of audio processing, and in particular, to an audio processing method, a computer device, and a computer program product. Background Art

[0002] In the process of audio processing, in order to increase the interest of music or create a specific atmosphere, animal calls, such as bird songs or tiger roars, are sometimes incorporated into the music melody.

[0003] In the related art, professional audio processing personnel can incorporate animal calls into the music melody according to personal experience to obtain a music melody with animal calls. However, this method has low processing efficiency and it is difficult to quickly and appropriately incorporate animal calls into the music melody. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide an audio processing method, a computer device, and a computer program product that can quickly generate a pleasant melody containing animal calls.

[0005] In a first aspect, the present application provides an audio processing method. The method includes:

[0006] Obtain an initial audio signal and sound signals of at least one animal call, where the initial audio signal has a music melody;

[0007] Based on the beat points of the music melody, determine the start playback positions of the respective sound signals in the initial audio signal;

[0008] Based on the start playback position and signal duration of each sound signal, determine the melody segment corresponding to each sound signal in the initial audio signal, and based on the pitch of the melody segments corresponding to the respective sound signals, perform pitch adjustment processing on each sound signal to obtain respective matching sound signals;

[0009] Based on the start playback positions corresponding to the respective matching sound signals, mix the respective matching sound signals with the initial audio signal to obtain a target audio signal.

[0010] In a second aspect, the present application further provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0011] Obtain an initial audio signal and sound signals of at least one animal call, where the initial audio signal has a music melody;

[0012] Determine the start playback positions of the respective sound signals in the initial audio signal based on the beat points of the music melody;

[0013] Based on the start playback positions and signal durations of the respective sound signals, determine the melody segments corresponding to the respective sound signals in the initial audio signal, and respectively perform pitch adjustment processing on the respective sound signals based on the pitches of the melody segments corresponding to the respective sound signals to obtain respective matching sound signals;

[0014] Based on the start playback positions corresponding to the respective matching sound signals, mix the respective matching sound signals with the initial audio signal to obtain a target audio signal.

[0015] In a third aspect, the present application also provides a computer program product. The computer program product includes a computer program which, when executed by a processor, implements the following steps:

[0016] Obtain an initial audio signal and sound signals of at least one animal call, where the initial audio signal has a music melody;

[0017] Determine the start playback positions of the respective sound signals in the initial audio signal based on the beat points of the music melody;

[0018] Based on the start playback positions and signal durations of the respective sound signals, determine the melody segments corresponding to the respective sound signals in the initial audio signal, and respectively perform pitch adjustment processing on the respective sound signals based on the pitches of the melody segments corresponding to the respective sound signals to obtain respective matching sound signals;

[0019] Based on the start playback positions corresponding to the respective matching sound signals, mix the respective matching sound signals with the initial audio signal to obtain a target audio signal.

[0020] The above audio processing method, computer device, and computer program product can obtain an initial audio signal and the sound signals of at least one animal call. The initial audio signal has a music melody. Then, the start playback positions of each sound signal in the initial audio signal can be determined based on the beat points of the music melody. Based on the start playback positions and signal durations of each sound signal, the melody segments corresponding to each sound signal in the initial audio signal can be determined, and the pitch of each sound signal can be adjusted based on the pitch of the melody segments corresponding to each sound signal to obtain each matched sound signal. Furthermore, based on the start playback positions corresponding to each matched sound signal, each matched sound signal can be mixed with the initial audio signal to obtain a target audio signal. In this solution, on the one hand, the start playback position of the animal call sound signal can be determined according to the beat points of the music melody, so that the playback rhythm of the animal call matches the rhythm of the music melody. On the other hand, after obtaining the matched sound signal based on the pitch of the melody segment corresponding to the sound signal, the matched sound signal can be mixed with the initial audio signal, so that the animal call naturally blends with the music melody, thereby enabling an appealing melody containing animal calls to be obtained quickly and effectively improving the processing efficiency of obtaining a music melody containing animal calls. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 FIG. is a schematic flowchart of an audio processing method in an embodiment;

[0022] Figure 2 FIG. is a schematic flowchart of steps for determining the frequency of a sound signal in an embodiment;

[0023] Figure 3a FIG. is a spectrogram of an animal call in an embodiment;

[0024] Figure 3b FIG. is a spectrogram of another animal call in an embodiment;

[0025] Figure 4a FIG. is a spectrogram of another animal call in an embodiment;

[0026] Figure 4b FIG. is a spectrogram of another animal call in an embodiment;

[0027] Figure 5 FIG. is a schematic flowchart of steps for determining the frequency of a melody segment in another embodiment;

[0028] Figure 6 FIG. is a schematic flowchart of another audio processing method in an embodiment;

[0029] Figure 7 FIG. is a schematic diagram of the beat alignment of an animal call in an embodiment;

[0030] Figure 8 The internal structure diagram of a computer device in an embodiment;

[0031] Figure 9 The internal structure diagram of another computer device in an embodiment. Specific implementation manners

[0032] In order to make the purpose, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0033] In one embodiment, as Figure 1 shown, an audio processing method is provided. In this embodiment, this method is illustrated by taking its application to a server as an example. It can be understood that this method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0034] S101, obtain an initial audio signal and sound signals of at least one animal call, where the initial audio signal has a music melody.

[0035] As an example, a music melody can be composed of basic music elements such as beats and pitches.

[0036] In practical applications, the sound signals of animal calls can be mixed into the initial audio signal with a music melody. Before starting the mixing, at least one animal call can be obtained in advance, and the sound signals of the animal calls are collected. The sound signals of each animal call can be processed as one path of sound signals; the initial audio signal to which the sound signals of the animal calls are to be added can be a pure music audio signal or a song audio signal containing lyrics.

[0037] S102, determine the start playback positions of each path of sound signals in the initial audio signal based on the beat points of the music melody.

[0038] Among them, the beat that constitutes the music melody is also called a beat, which can be used to represent a preset unit time value. In other words, the equally divided unit time value can be called "one beat", "one measure" or "one bar". The time value of a beat can be represented by the time value of a note. For example, the time value of one beat can be a quarter note (i.e., one beat is a quarter note), or a half note (one beat is a half note) or an eighth note (one beat is an eighth note).

[0039] The beat point can be understood as the switching point between two adjacent beats in a musical melody. Taking a musical melody with a time signature of "4 / 4 time" as an example, the musical melody can include multiple bars. With a quarter note as one beat, each bar contains four beats, and the beat point can be the switching point (or connection point) between two adjacent quarter notes.

[0040] In this step, the beat points in the musical melody corresponding to the initial audio signal can be determined, and based on the beat points in the musical melody, the start playback positions of the sound signals of each animal call can be determined, so that the start playback moments of the sound signals of the animal calls can be aligned with the beat points of the musical melody, making the appearance moments of the sound signals of the animal calls match the rhythm of the musical melody, and enhancing the sense of rhythm of the animal calls in the musical melody.

[0041] For example, if the musical melody includes 8 beats and correspondingly has 7 beat points, and if there are 3 sound signals of animal calls to be added, the 1st, 2nd, and 3rd beat points can be used as the start playback positions of the sound signals of each animal call respectively, or 2 of the beat points can be selected as the start playback positions of the 3 sound signals, that is, the start playback positions of the sound signals of different animal calls can be the same or different.

[0042] S103, based on the start playback position and signal duration of each sound signal, determine the melody segment corresponding to each sound signal in the initial audio signal, and based on the pitch of the melody segments corresponding to each sound signal, perform pitch adjustment processing on each sound signal respectively to obtain each matching sound signal.

[0043] Among them, the signal duration can be the time consumed for playing the sound signal at a preset playback rate, for example, the time consumed for playing the sound signal of the animal call at a speed of one time speed.

[0044] In specific implementation, the pitch of the sound signal of the animal call obtained in advance may not be consistent or mismatched with the pitch of the musical melody. For example, the pitch of the sound signal is different from the pitch of the musical melody or differs by multiple octaves. If the sound signal of the animal call is directly mixed with the initial audio signal, the situation of "being out of tune" will occur, resulting in the finally mixed audio signal not being pleasant or melodious.

[0045] And the pitch at different positions in the musical melody can be different, that is, the pitch in the musical melody can be variable, and the signal durations and start playback positions of the sound signals of different animal calls can also vary. In this step, based on the start playback position and signal duration of each sound signal, the melody segment corresponding to each sound signal in the initial audio signal can be determined.

[0046] Specifically, for the sound signal of each animal call, the end playback position of the sound signal in the initial audio signal can be determined according to the start playback position and signal duration of the sound signal. By combining the start playback position and end playback position of the sound signal in the initial audio signal, the melody segment corresponding to the sound signal in the music melody corresponding to the initial audio signal can be determined.

[0047] After determining the melody segment corresponding to the sound signal, the pitch of the melody segment can be obtained. Furthermore, the pitch of the obtained sound signal of the animal call can be adjusted to obtain a sound signal that matches the pitch of the melody segment as the matching sound signal. For example, the sound signal can be directly processed into the matching sound signal, or the matching sound signal can be generated using the sound signal.

[0048] Among them, pitch matching can mean that the pitch of the sound signal is the same as the pitch of at least part of the melody segment (such as the first n seconds or the entire melody segment of the melody segment start), or it can mean that the pitch of the sound signal is within the key range corresponding to the pitch of the melody segment. For example, if the key range corresponding to the pitch of the melody segment is C major, and the pitch of the sound signal is a pitch in C major, then the sound signal can also be used as the matching sound signal.

[0049] S104. Based on the start playback positions corresponding to each path of matching sound signals, mix each path of matching sound signals with the initial audio signal to obtain the target audio signal.

[0050] Among them, the target sound signal can be a sound signal that is played simultaneously with the melody segment.

[0051] After obtaining the matching sound signal, since the matching sound signal is obtained by performing pitch adjustment processing on the sound signal, the start playback position of the sound signal in the initial audio signal can be used as the start playback position corresponding to the matching sound signal. Based on the start playback positions of each path of matching sound signals, mix each path of matching sound signals and the initial audio signal to obtain the target audio signal. During the playback of the target audio signal, when the playback progress reaches the start playback position, the matching sound signal of the corresponding animal call can start to be played.

[0052] In this embodiment, an initial audio signal and sound signals of at least one animal call can be obtained. The initial audio signal has a music melody. Then, the start playback positions of each sound signal in the initial audio signal can be determined based on the beat points of the music melody. Based on the start playback positions and signal durations of each sound signal, the melody segments corresponding to each sound signal in the initial audio signal can be determined, and based on the pitches of the melody segments corresponding to each sound signal, pitch adjustment processing can be performed on each sound signal respectively to obtain each matching sound signal. Furthermore, based on the start playback positions corresponding to each matching sound signal, each matching sound signal and the initial audio signal can be mixed to obtain a target audio signal. In this solution, on the one hand, the start playback position of the animal call sound signal can be determined according to the beat points of the music melody, so that the playback rhythm of the animal call is adapted to the rhythm of the music melody. On the other hand, after obtaining the matching sound signal based on the pitch of the melody segment corresponding to the sound signal, the matching sound signal and the initial audio signal can be mixed, so that the animal call naturally blends with the music melody, thereby a pleasant melody containing animal calls can be quickly obtained, effectively improving the processing efficiency of obtaining a music melody containing animal calls.

[0053] In one embodiment, performing pitch adjustment processing on each sound signal respectively based on the pitch of the melody segment corresponding to each sound signal to obtain each matching sound signal may include:

[0054] If the frequency of the sound signal matches the frequency corresponding to the pitch of the melody segment, the sound signal is determined as the matching sound signal; if the frequency of the sound signal does not match the frequency corresponding to the pitch of the melody segment, the frequency of the sound signal is adjusted to obtain the matching sound signal.

[0055] Wherein, the pitch of the frequency-converted sound signal matches the pitch of the melody segment.

[0056] The essence of sound is a mechanical wave, and the pitch of the sound is determined by the frequency and wavelength of the mechanical wave. When the wavelength is constant, a high frequency results in a high pitch, and a low frequency results in a low pitch.

[0057] Based on this, after obtaining the sound signal of the animal call and the corresponding melody segment of the sound signal, the frequency of the animal call sound signal and the frequency of the melody segment can be determined respectively and compared. By judging whether the frequency of the sound signal matches the frequency corresponding to the pitch of the melody segment, it can be determined whether the pitch of the sound signal matches the pitch of the melody segment. Wherein, the frequency of the sound signal and the frequency of the melody segment can be a frequency value or a frequency range (i.e., frequency band) composed of multiple frequency values.

[0058] If the frequency of the sound signal matches the frequency of the melody segment, such as when the two frequency values match or the frequency value of the sound signal falls within the frequency range corresponding to the pitch of the melody segment, the sound signal can be determined as a matching sound signal that matches the pitch of the melody segment.

[0059] If the frequency of the sound signal does not match the frequency of the melody segment, the frequency of the sound signal can be frequency-converted, the frequency of the sound signal can be adjusted, and the frequency of the adjusted sound signal can be made to match the frequency corresponding to the pitch of the melody segment, so that a matching sound signal that matches the pitch of the melody segment can be obtained. Exemplarily, the sound signal can be frequency-converted through a preset frequency conversion application (such as Phasevocoder) to quickly obtain the adjusted sound signal.

[0060] In this embodiment, by determining whether the frequency of the sound signal matches the frequency corresponding to the pitch of the melody segment, and by adjusting the frequency of the sound signal in the case of non-matching melody, the pitch of the animal call is made to match the pitch of the melody segment more closely, making the animal call in the finally synthesized audio more musical and rhythmical, and making the animal call more naturally and adaptively integrated into the music melody.

[0061] To compare the frequency of the sound signal with the frequency corresponding to the pitch of the melody segment, as Figure 2 shown, in one embodiment, the frequency of the sound signal can be obtained through the following steps:

[0062] S201, obtain the power spectrum of each of multiple audio frames of the sound signal, and based on the multiple power spectra, determine the power values of the sound signal at different frequencies to obtain the spectral distribution of the sound signal.

[0063] As an example, the power spectrum can indicate the power values of the signal at different frequencies, and the power spectrum can represent the change of the signal power with frequency, that is, the distribution of the signal power in the frequency domain.

[0064] In practical applications, multiple audio frames of the animal call sound signal can be obtained, and the power spectrum of each audio frame can be determined.

[0065] Specifically, the sound signal in the time domain of the animal call can be obtained, and the sound signal in the time domain can be frame-divided to obtain multiple audio frames; taking the sound signal x(i) as an example, the sound signal can be segmented into multiple frames according to a frame shift of 10 ms and a frame length of 30 ms, where i is a positive integer used to represent the sample index of each sampling point in the sound signal in the time domain. For the multiple frames obtained by segmentation, further processing can be performed using a window function, and multiple audio frames can be obtained based on the processing results. In one example, the Hann window can be used to process the sound signal, for example, the Hann window shown below can be adopted:

[0066]

[0067] Among them, N is the window length. Correspondingly, the audio frame processed by the window function can be expressed as:

[0068] x w (n, i) = x(L·n + i)·w hann (i),

[0069] Among them, L represents the frame shift, and n is the frame index of the audio frame.

[0070] After obtaining multiple audio frames of the time-domain signal, the Fourier transform can be performed on each audio frame to obtain the spectrogram of each audio frame, and the corresponding power spectrum can be obtained according to the spectrogram. Specifically, for example, X(k, n) is the spectrogram of the audio frame, where k is the frequency point index in the spectrogram, then the power spectrum P(k, n) can be determined by the following method:

[0071] P(k, n) = ||X(k, n)|| 2

[0072] The power spectrum of the audio frame can record the power of the sound signal at different frequencies, and then the frequency spectrum distribution of the sound signal can be obtained based on the power spectra of multiple audio frames.

[0073] Specifically, for the power spectra corresponding to the audio frames at n moments, the power values of k frequencies at each moment of each power spectrum can be obtained; then, the multiple power values can be sorted according to the frequency and the moment corresponding to the power spectrum to obtain a k*n matrix. For example, for the first frequency, the power values at n moments of this frequency can be obtained. For each frequency, based on the power values at n moments, a power value that meets the preset condition can be obtained as the target power value of this frequency, such as the median, maximum value, minimum value, or average value, etc. In one example, if the median is taken as the target power value of this frequency, it can be determined by the following method:

[0074]

[0075] Furthermore, the frequency spectrum distribution of the sound signal can be obtained according to the target power values of each frequency.

[0076] S202, determine at least one frequency whose energy concentration degree of the frequency spectrum distribution meets the preset concentration condition as the frequency of the sound signal.

[0077] After obtaining the frequency spectrum distribution of the sound signal, based on the frequency spectrum distribution, at least one frequency whose energy concentration degree meets the preset concentration condition can be determined from the power values at multiple frequencies. In practical applications, one or more frequencies with the largest power values can be determined as the frequencies whose energy concentration degree meets the preset concentration condition.

[0078] If the fundamental frequency of the animal call is obvious, for example Figure 3a the frequency spectrum of the fox call shown in Figure 3b and the frequency spectrum of the bird call shown in Figure 4a the tiger call shown in Figure 4b and the lion call shown in

[0079] Exemplarily, the position κ where the energy is concentrated in the spectrum distribution can be determined by the following method:

[0080]

[0081] Then, further estimation is performed on the frequency at this position, and the frequency of the sound signal is obtained based on the estimation result. The estimation method can be as follows:

[0082]

[0083] where fs is the sampling rate of the sound signal of the animal call.

[0084] In this embodiment, the frequency of the sound signal of the animal call can be quickly determined according to the spectrum distribution of the sound signal.

[0085] In one embodiment, as Figure 5 shown, the frequency corresponding to the pitch of the melody segment can be obtained through the following steps:

[0086] S501, obtain a pitch class label in the numbered musical notation of the melody segment and the reference pitch of the music melody.

[0087] Specifically, the numbered musical notation may include multiple pitch class labels, and each pitch class label corresponds to a pitch class. A pitch class can be understood as a unit for dividing the intervals between the tones in a scale. In one example, the pitch class labels may include 1, 2, 3, 4, 5, 6, 7, which respectively correspond to 7 pitch classes, and the solfège syllables are do, re, mi, sol, la, si in sequence. Taking the numbered musical notation "4|3|2|1||2|0|0|0" with a time signature of 4 / 4 as an example, "|" represents the beat interval, each number corresponds to a beat, "||" represents the number of bar intervals, 0 represents a rest beat where no sound is required, and the remaining numbers respectively correspond to a pitch class.

[0088] The reference pitch can also be called the key name, and each reference pitch corresponds to a key domain. The pitch of the same pitch class is different in different key domains. For example, the pitch of C key 1 and D key 1 is different, while the pitch of D key 1 is the same as the pitch of C key 2.

[0089] In this step, the numbered musical notation of the melody segment and the reference pitch of the entire music melody can be obtained. Specifically, the numbered musical notation of the music melody can be pre-stored. This numbered musical notation can also be called a melody template, which can record the time signature, reference pitch, and multiple pitch-class identifiers corresponding to the music melody. Then, the numbered musical notation of the melody segment and the reference pitch of the music melody can be obtained from the numbered musical notation of the music melody.

[0090] And after obtaining the numbered musical notation of the melody segment, a pitch-class identifier can be obtained from the melody segment to identify the audio of the melody segment. Exemplarily, the first pitch-class identifier of the melody segment can be obtained, so that the animal sound can match the pitch of the corresponding beat when it starts to play; or, if the melody segment includes the climax part of the music melody, a pitch-class identifier of the climax part in this melody segment can also be obtained, so that the animal sound matches the pitch of the climax part of the music melody.

[0091] S502. Determine the actual pitch of the pitch-class corresponding to the pitch-class identifier under the reference pitch, and obtain the pitch difference between the actual pitch and the preset reference pitch.

[0092] After obtaining the pitch-class corresponding to the pitch-class identifier and the reference pitch of the music melody, the pitch of this pitch-class can be determined by combining the pitch-class and the reference pitch, that is, the actual pitch of this pitch-class under the current reference pitch. Furthermore, the pitch difference between the actual pitch and the preset reference pitch can be obtained.

[0093] In one example, the note numbers of multiple reference pitches can be provided in advance, as shown in Table 1 below:

[0094] Table 1

[0095] Note number 60 62 64 65 67 69 71 Tonality C D E F G A B

[0096] Among them, the note numbers are set according to the preset standard. Each note number corresponds to a pitch. The difference between the note numbers of two adjacent whole tones is 2, and the difference between the note numbers of two adjacent semitones is 1. In other words, the pitch difference can be determined based on the note number difference.

[0097] Under each reference pitch, the note numbers corresponding to the pitch-class can be determined according to the mapping relationships shown in Table 2 (Note Number Resolution), Table 3 (Note Offset Resolution), and Table 4 (Octave Offset Resolution) below:

[0098] Table 2

[0099] Pitch class 1 2 3 4 5 6 7 Semitone interval f(n) 0 2 4 5 7 9 11

[0100] Table 3

[0101]

[0102]

[0103] Table 4

[0104] Identification + - (blank) Semitone interval oct 1 -1 0

[0105] Among them, the "+" and "-" in numbered musical notation can represent octave raising and lowering, and the "#" and "b" represent the raising and lowering between semitones.

[0106] Then the note number note corresponding to the actual pitch of the corresponding pitch level in numbered musical notation under the reference melody can be expressed as:

[0107] note = N base + f(n) + Shift + 12 * oct

[0108] Among them, N base is the note number of the reference pitch.

[0109] After obtaining the note number corresponding to the actual pitch, the preset standard pitch (such as A4) can be obtained. The standard pitch can be a musical tone with a determined corresponding frequency. Furthermore, the difference between the note number corresponding to the standard pitch and the note number corresponding to the actual pitch can be determined, and the pitch difference can be determined based on this difference.

[0110] S503. Determine the frequency of the melody segment based on the frequency corresponding to the reference pitch, the frequency difference corresponding to the unit pitch difference, and the pitch difference.

[0111] After obtaining the pitch difference, the frequency corresponding to the reference pitch and the frequency difference corresponding to the unit pitch difference can be obtained. In specific implementation, for the pitch difference between every two adjacent semitones, the corresponding frequency difference is That is, the frequency difference corresponding to the unit pitch difference can be determined as

[0112] Then, based on the pitch difference between the actual pitch and the reference pitch, and the frequency difference corresponding to the unit pitch difference, the frequency corresponding to the actual pitch can be determined, and this frequency can be used as the frequency of the melody segment. Exemplarily, when the reference pitch is A4 (note number is 69), the frequency freq corresponding to the actual pitch can be determined based on the following method:

[0113]

[0114] In this embodiment, the frequency of the melody segment can be quickly determined based on the numbered musical notation of the melody segment, providing a basis for subsequent pitch matching with the sound signal of animal calls.

[0115] In one embodiment, based on the start playback positions corresponding to the respective matched voice signals, mixing the respective matched voice signals with the initial audio signal may include the following steps:

[0116] Perform a fade-out process on each of the matched voice signals to obtain target voice signals; based on the start playback positions corresponding to the respective target voice signals, mix the respective target voice signals with the initial audio signal.

[0117] In practical applications, if the obtained matched voice signals are directly mixed with the initial audio signal, when the matched voice signals finish playing, the animal calls in the audio will suddenly disappear, resulting in unnatural connection between the voice signals of the animal calls and the initial audio signal.

[0118] In this step, a fade-out process may be performed on each of the matched voice signals to obtain target voice signals. Exemplarily, the fade-out process may include adjusting the signal value of some of the matched voice signals, or adjusting the signal loudness of some of the matched voice signals, so that the animal calls gradually disappear during playback. Then, based on the start playback positions corresponding to the respective target voice signals, the respective target voice signals may be mixed with the initial audio signal to obtain a target audio signal.

[0119] In this embodiment, by performing a fade-out process on each of the matched voice signals, the connection between the voice signals and the initial audio signal can be made more natural during the mixing process.

[0120] In one embodiment, performing a fade-out process on each of the matched voice signals to obtain target voice signals may include the following steps:

[0121] For each of the matched voice signals, if the melody segment corresponding to the matched voice signal includes multiple beats and the pitch of the melody segment is determined based on the first beat of the melody segment, then determine the voice signals other than the first beat as the voice signals to be adjusted; perform a fade-out process on the voice signals to be adjusted in each of the matched voice signals to obtain target voice signals.

[0122] Specifically, when the melody segment includes multiple beats, the actual pitches corresponding to the multiple beats can be the same or different. If the pitch of the melody segment is determined based on the first beat in the melody segment, when performing pitch matching between the sound signal of the animal call and the melody segment, it can be understood that the pitch of the sound signal is matched with the pitch of the first beat of the melody segment. For the other beats in the melody segment except the first beat, the pitch of the sound signal may not match the pitch of the other beats. Exemplarily, by comparing the frequency corresponding to the sound signal with the frequency corresponding to the pitch of the other beats, it can be determined whether the pitch of the sound signal matches the pitch of the other beats. If the frequency of the sound signal is different from the frequency corresponding to the pitch of the other beats, it can be determined that the pitches do not match.

[0123] In this step, the matching sound signals other than the first beat can be determined as the sound signals to be adjusted. After performing a fade-out process, the obtained sound signal after processing can be used as the target sound signal corresponding to the melody segment. For example, for the sound signal of any animal call It can be processed in the following manner:

[0124]

[0125] where the value range of j is [0, J m , J m represents the length of the sound signal of the m-th animal call; the value range of i is [0, L], and L represents the total length of the signal after splicing the sound signals of multiple animal calls; T represents the duration of one beat; w m (j) is a fade-out processing function. In one example, w m (j) can be as follows:

[0126]

[0127] In this embodiment, by performing a fade-out process on the matching sound signals beyond the first beat, the situation where the pitch of the sound signal to be adjusted does not match the pitch of the other beats in the melody segment can be weakened, and the rhythm of the finally synthesized target audio signal can be increased.

[0128] In one embodiment, the music melody of the initial audio signal can be obtained through the following steps:

[0129] Obtain the music score associated with the initial audio signal and the rhythm information of the music score; based on the music score and the rhythm information of the music score, obtain the music melody of the initial audio signal.

[0130] Specifically, the initial audio signal can be an audio signal corresponding to a song, accompaniment, or background music. After determining the initial audio signal, a musical score associated with the song, accompaniment, or background music (BGM) can be obtained from a music library as the musical score associated with the initial audio signal, and the rhythm information recorded in the musical score can be obtained. Among them, the musical score can be a numbered musical notation or a staff notation; the rhythm information is used to indicate the number of beats played per unit time when playing the musical score.

[0131] It can be understood that the musical score can contain multiple beats, and each beat can have a time value, and the actual duration corresponding to the time value can vary, that is, under different rhythms, the actual duration of the same beat is different. For example, if the rhythm is 89 beats per minute, the actual duration of each beat is 1 / 89 second, and if the rhythm is 60 beats per minute, the actual duration of each beat is 1 second. In other words, under different rhythms, the playing duration of the same musical score is different.

[0132] Based on this, after obtaining the rhythm information, the playing duration of each note in the musical score can be determined according to the rhythm information, and the music melody of the initial audio signal can be obtained based on the order of the notes in the musical score and the playing duration of each note.

[0133] In this embodiment, based on the musical score and the rhythm information, determining the music melody of the initial audio signal can obtain melody information with duration information, providing a basis for subsequently determining the corresponding melody segment of the animal call sound signal.

[0134] In one embodiment, based on the start playback positions of each path of matching sound signals, mixing each path of matching sound signals with the initial audio signal to obtain a target audio signal may include the following steps:

[0135] Based on the start playback positions of each path of matching sound signals, splice each path of matching sound signals; mix the animal sound signal and the initial audio signal to obtain a target audio signal.

[0136] Specifically, each path of matching sound signals corresponding to each animal call can be spliced based on the start playback position. In other words, when the time progress reaches the start playback position, the corresponding target sound signal can be added.

[0137] For example, when the time progress reaches 0.5 seconds, add a path of target sound signal, and the duration of the target sound signal is determined based on the duration of the target sound signal. Thus, by splicing each path of target sound signals, the spliced animal sound signal can be obtained.

[0138] After obtaining the spliced animal sound signal, the animal sound signal can be mixed with the initial audio signal to obtain a target audio signal.

[0139] In this embodiment, by splicing each target sound signal based on the start playback position, an animal sound signal after splicing is obtained, and then the spliced animal sound signal and the initial audio signal are subjected to a mixing process, so that each target sound signal can be mixed into the initial audio signal at one time, avoiding multiple mixing processes and effectively improving the mixing efficiency.

[0140] In one embodiment, performing a mixing process on the animal sound signal and the initial audio signal to obtain a target audio signal may include:

[0141] Determining the target loudness of each of the animal sound signal and the initial audio signal; based on the target loudness of each of the animal sound signal and the initial audio signal, mixing the animal sound signal and the initial audio signal to obtain a target audio signal.

[0142] In practical applications, there may be differences in the generation method or signal acquisition method between the sound signal of the animal call and the initial audio signal. For example, the animal call can be recorded live through a recording device, and then the silent segments in the recording are removed through Voice Activity Detection (VAD) to obtain an effective segment containing the animal call, and the sound signal of the animal call is obtained based on the signal in this segment; while the initial audio signal can be obtained after being recorded through a dedicated audio processing tool or a recording studio.

[0143] Correspondingly, there may be a situation where the initial loudness of the animal sound signal and the initial loudness of the initial audio signal do not match. For example, the initial loudness of the animal sound signal is much smaller than the initial loudness of the initial audio signal, or the initial loudness of the accompaniment audio signal is smaller than the initial loudness of the animal sound signal, resulting in the animal sound signal or the accompaniment sound signal being masked by the other signal.

[0144] In this step, after obtaining the animal sound signal and the initial audio signal, the target loudness of the animal sound signal and the target loudness of the initial audio signal can be determined first, and the target loudness is the loudness when the signal is played.

[0145] In a specific implementation, the target loudness of the animal sound signal and the initial audio signal can be set through weight coefficients. The weight coefficient is positively correlated with the loudness of the signal, that is, the larger the weight coefficient, the greater the signal loudness. In some examples, before configuring the loudness of the animal sound signal and the accompaniment audio signal using the weight coefficient, loudness equalization processing can be performed on the animal sound signal and the initial audio signal, that is, adjusting the signal loudness of the animal sound signal and the initial audio signal to the same benchmark, such as making the loudness of the animal sound signal the same as that of the initial audio signal or the loudness difference between the two within a preset range, and then configuring through the weight coefficient to obtain their respective target loudness. Exemplarily, if it is necessary to highlight the animal sound signal in the target audio signal, the target loudness of the animal sound signal can be made greater than the target loudness of the initial audio signal, and the loudness of the initial audio signal can be suppressed.

[0146] Furthermore, the animal sound signal and the initial audio signal can be mixed according to the target loudness of the animal sound signal and the initial audio signal, and the target audio signal can be obtained based on the mixing result. In some embodiments, when mixing the animal sound signal and the initial audio signal, the signals can be divided into left-channel and right-channel signals for mixing, so as to achieve the effect of two-channel stereo. For example, the mixing can be performed in the following manner:

[0147]

[0148] where y L () is the target audio signal of the left channel, y R () is the target audio signal of the right channel, c L () is the accompaniment audio signal of the left channel, c R () is the initial audio signal of the right channel, α L is the weight coefficient of the accompaniment audio signal of the left channel, α R is the weight coefficient of the initial audio signal of the right channel, x(i) is the animal sound signal. Exemplarily, m is the number of types of animal call sound signals in each path.

[0149] Of course, in some other embodiments, one or more of fade-out processing, splicing processing, or loudness adjustment processing can be performed on the matching sound signal to achieve the mixing process of the matching sound signal and the initial audio signal.

[0150] In this embodiment, by mixing the animal sound signal and the accompaniment audio signal according to their respective target loudness, the animal sound and the accompaniment of the music melody can be played simultaneously at an appropriate loudness, avoiding any signal being masked by another signal during subsequent audio playback.

[0151] To enable those skilled in the art to better understand the above steps, the following provides an exemplary illustration of the embodiments of the present application through an example. However, it should be understood that the embodiments of the present application are not limited thereto.

[0152] As Figure 6 shown, in practical applications, an audio processing request can be obtained. For example, a user can send an audio processing request carrying a melody template and an initial audio signal to a server through a terminal. After receiving the audio processing request, the server can read and parse the melody template from the audio processing request to obtain the music melody corresponding to the initial audio signal. Exemplarily, the melody template can be description information of the music melody, and the description information can include the beats of the melody, the syllables under each beat, the rhythm of the melody, and the reference pitch of the melody.

[0153] Moreover, animal sound materials can also be obtained. By performing voice activity detection on the animal sound materials to remove the silent segments in the recording, effective segments containing animal call sound signals are obtained, and the fundamental frequency or the energy concentration frequency band of the segments is estimated. According to the estimation result, the frequency of the animal call sound signal is determined.

[0154] Among them, the animal call sound signal and its frequency can be pre-stored in a material library. When the user sends an audio processing request, the identifier of the animal call to be added to the music melody can be sent accordingly. Thus, the server can find the corresponding animal call sound signal and its frequency in the preset material library according to the identifier. Alternatively, the audio processing request sent by the user can carry animal sound materials. After receiving the animal sound materials, the server can real-time identify the segments containing animal call sound signals therein and determine the corresponding frequency.

[0155] Then, the sound signals of each animal call can be aligned with the beat points of the music melody, that is, some beat points in the music melody are determined as the start playback positions of the animal call sound signals, and the melody segments corresponding to each sound signal after the beat point alignment are determined.

[0156] Taking the music melody "4|3|2|1||2|0|0|0" as an example, its corresponding frequency sequence can be as Figure 7 shown. If the animal call sound signals to be added to the initial audio signal include 4 signals s1, s2, s3, and s4, and their signal durations are 2.5 s, 0.5 s, 1 s, and 1.3 s respectively, the sound signals after the beat point alignment and the melody segments of each sound signal can be as Figure 7 shown.

[0157] After determining the melody segments corresponding to each voice signal, the frequencies corresponding to the melody segments can be obtained, and the frequencies of the melody segments and the frequencies of the corresponding voice signals are compared. If the two match, the frequency of the voice signal can be left unchanged, and the voice signal outside the first beat of the melody segment is faded out (faded out). If the two do not match, the frequency of the voice signal is adjusted so that the frequency of the adjusted voice signal matches the frequency of the melody segment, and then the fade-out process is performed.

[0158] Furthermore, the current voice signals of each path can be spliced to obtain an animal voice signal, and the initial audio signal corresponding to the music melody is mixed with the animal voice signal to obtain a target audio signal, which is used as the animal chorus music.

[0159] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limitation, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.

[0160] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as Figure 8 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store audio data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements an audio processing method.

[0161] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as Figure 9As shown in the figure. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it realizes an audio processing method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0162] Those skilled in the art can understand that Figure 8 and Figure 9 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0163] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are realized:

[0164] Obtain an initial audio signal and the sound signals of at least one animal call, and the initial audio signal has a music melody;

[0165] Based on the beat points of the music melody, determine the start playback positions of the respective sound signals in the initial audio signal;

[0166] Based on the start playback position and signal duration of each sound signal, determine the melody segments corresponding to each sound signal in the initial audio signal, and respectively perform pitch adjustment processing on each sound signal based on the pitch of the melody segments corresponding to each sound signal to obtain each matching sound signal;

[0167] Based on the start playback positions corresponding to each matching sound signal, mix each matching sound signal with the initial audio signal to obtain a target audio signal.

[0168] In one embodiment, when the processor executes the computer program, it also realizes the steps in the above other embodiments.

[0169] In one embodiment, a computer program product is provided, including a computer program which, when executed by a processor, implements the following steps:

[0170] Obtain an initial audio signal and sound signals of at least one animal call, where the initial audio signal has a music melody;

[0171] Based on the beat points of the music melody, determine the start playback positions of the respective sound signals in the initial audio signal;

[0172] Based on the start playback positions and signal durations of the respective sound signals, determine the melody segments corresponding to the respective sound signals in the initial audio signal, and based on the pitches of the melody segments corresponding to the respective sound signals, perform pitch adjustment processing on the respective sound signals to obtain respective matching sound signals;

[0173] Based on the start playback positions corresponding to the respective matching sound signals, mix the respective matching sound signals with the initial audio signal to obtain a target audio signal.

[0174] In one embodiment, when the computer program is executed by a processor, it also implements the steps in the above-mentioned other embodiments.

[0175] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0176] Those of ordinary skill in the art can understand that all or part of the processes in the above-described method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the method embodiments as described above. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0177] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0178] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. An audio processing method, characterized in that The method includes: Obtaining an initial audio signal and a sound signal of an animal call, where the initial audio signal has a music melody; Determining the start playback position of each of the sound signals in the initial audio signal based on the beat points of the music melody; Based on the start playback position and signal duration of each of the sound signals, determining the corresponding melody segment of each of the sound signals in the initial audio signal, determining the spectral distribution of the sound signal based on the power spectrum of each of the multiple audio frames of the sound signal, and taking at least one frequency whose energy concentration in the spectral distribution meets a preset concentration condition as the frequency of the sound signal; determining the frequency corresponding to the pitch of the melody segment based on the pitch level identifier in the simple musical notation of the melody segment and the reference pitch of the music melody, and respectively performing pitch adjustment processing on each of the sound signals based on the frequency of the sound signal and the frequency corresponding to the pitch of the melody segment to obtain each matching sound signal; the pitch of the matching sound signal matches the pitch of the melody segment; Based on the start playback position corresponding to each of the matching sound signals, mixing each of the matching sound signals with the initial audio signal to obtain a target audio signal; the target audio signal plays the corresponding matching sound signal when the playback progress reaches the start playback position.

2. The method according to claim 1, wherein The performing pitch adjustment processing on each of the sound signals based on the frequency of the sound signal and the frequency corresponding to the pitch of the melody segment to obtain each matching sound signal includes: If the frequency of the sound signal matches the frequency corresponding to the pitch of the melody segment, determining the sound signal as a matching sound signal; If the frequency of the sound signal does not match the frequency corresponding to the pitch of the melody segment, adjusting the frequency of the sound signal to obtain a matching sound signal; the pitch of the sound signal after frequency adjustment matches the pitch of the melody segment.

3. The method according to claim 1, characterized in that, The determining the spectral distribution of the sound signal based on the power spectrum of each of the multiple audio frames of the sound signal includes: Obtaining the power spectrum of each of the multiple audio frames of the sound signal, and determining the power value of the sound signal at different frequencies based on the multiple power spectra to obtain the spectral distribution of the sound signal.

4. The method according to claim 1, characterized in that The determining the frequency corresponding to the pitch of the melody segment based on the pitch level identifier in the simple musical notation of the melody segment and the reference pitch of the music melody includes: Determining the actual pitch of the pitch level corresponding to the pitch level identifier under the reference pitch, and obtaining the pitch difference between the actual pitch and a preset reference pitch; Based on the frequency corresponding to the reference pitch, the frequency difference corresponding to a unit pitch difference, and the pitch difference, determining the frequency of the melody segment.

5. The method according to claim 1, wherein The mixing each of the matching sound signals with the initial audio signal based on the start playback position corresponding to each of the matching sound signals includes: Performing a fade-out process on each of the matching sound signals to obtain a target sound signal; Based on the start playback position corresponding to each of the target sound signals, mixing each of the target sound signals with the initial audio signal.

6. The method according to claim 5, wherein Performing a fade-out process on each of the matched voice signals to obtain a target voice signal includes: For each of the matched voice signals, if the melody segment corresponding to the matched voice signal includes multiple beats and the pitch of the melody segment is determined based on the first beat of the melody segment, then the matched voice signals other than the first beat are determined as voice signals to be adjusted; Performing a fade-out process on the voice signals to be adjusted in each of the matched voice signals to obtain a target voice signal.

7. The method according to claim 1, characterized in that, The music melody obtaining step of the initial audio signal includes: Obtaining the musical score associated with the initial audio signal and the rhythm information of the musical score; the rhythm information is used to indicate the number of beats played per unit time when playing the musical score; Based on the musical score and the rhythm information, obtaining the music melody of the initial audio signal.

8. The method according to any one of claims 1-7, characterized in that, Based on the start playback positions corresponding to each of the matched voice signals, mixing each of the matched voice signals with the initial audio signal to obtain a target audio signal includes: Based on the start playback positions corresponding to each of the matched voice signals, splicing each of the matched voice signals to obtain a spliced animal voice signal; Mixing the animal voice signal and the initial audio signal to obtain a target audio signal.

9. The method according to claim 8, wherein Mixing the animal voice signal and the initial audio signal to obtain a target audio signal includes: Determining the target loudness of each of the animal voice signal and the initial audio signal; Based on the target loudness of each of the animal voice signal and the initial audio signal, mixing the animal voice signal and the initial audio signal to obtain a target audio signal.

10. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 are implemented.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Audio processing method and device and storage medium

    CN109346044A