Audio processing method and related device
By calculating the change in the beat by beat, frequency domain filtering and time domain weighting, the song playback speed is automatically adjusted, and the problem of inconsistent rhythms when splicing multiple songs is solved, achieving efficient and low-cost playback consistency and listening consistency.
Patent Information
- Application Number
- CN202211666811.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-23
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-12-23
AI Technical Summary
When the prior art splices multiple songs, it is difficult to automatically adjust the playback speed to achieve unified rhythm, resulting in incoherent playback, affecting the user's listening experience, and manual speed regulation is low efficiency and high cost.
By calculating the amount of beat-by-beat change between audio bands, the audio data playback speed within each beat is automatically adjusted, so that the time interval between each beat is consistent, and combined with frequency domain filtering and time domain weighting processing, a smooth transition of the audio band is achieved.
It realizes the uniform and smooth and coherent rhythm of multiple songs when playing, improves the user's listening experience and reduces labor costs and operational complexity.
Smart Images

Figure CN115985272B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of audio technology, and in particular to audio processing methods and related devices. Background Art
[0002] In daily life, many users need to coordinate the playing rhythms of multiple audio segments to produce playing effects such as song medleys or overlays. However, when splicing multiple songs into a compilation, it is inevitable to face the abruptness caused by inconsistent beats per minute (BPM) of the playing speed, and the transition is not smooth, resulting in discontinuous playing between songs, thus affecting the user's listening experience.
[0003] The existing method is to import the songs to be spliced into a digital audio workstation (DAW) software and manually adjust the playing speed of one or even two songs to make their rhythms as unified as possible. However, this manual speed adjustment method has limited number of tracks that can be adjusted, and is prone to errors such as discontinuous playing, with low production efficiency and high labor costs, making it difficult to meet the production requirements of the majority of users.
[0004] In view of this, it is necessary to provide a more effective solution to solve the above problems. Summary of the Invention
[0005] The embodiments of the present application provide an audio processing method and related devices, which are used to automatically unify the playing speeds of multiple audio segments and improve the signal coherence between spliced audios.
[0006] A first aspect of the embodiments of the present application provides an audio processing method, including:
[0007] Determine a first audio segment of a first audio and a second audio segment of a second audio, where the first audio is played before the second audio and has a different playing speed, and the beat length of the first audio segment is the same as the beat length of the second audio segment;
[0008] Obtain first BPM information of the first audio segment and second BPM information of the second audio segment, where the BPM information is the number of beats of an audio segment per unit time;
[0009] Based on the first BPM information and the second BPM information, calculate the beat-by-beat change amount between the first audio segment and the second audio segment, where the beat-by-beat change amount is used to represent the change between the first audio segment and the second audio segment at the beat length;
[0010] Use the per-beat variation amount to adjust the playback speed of the audio data within each beat of the first audio segment and the second audio segment respectively, so as to correspondingly obtain a third audio segment and a fourth audio segment, wherein the time intervals between the beats of the third audio segment are the same as the time intervals between the beats of the fourth audio segment.
[0011] Optionally, calculating the per-beat variation amount between the first audio segment and the second audio segment includes:
[0012] Calculate the BPM difference between the first BPM information and the second BPM information, and calculate the per-beat variation amount based on the BPM difference and the beat length.
[0013] Optionally, using the per-beat variation amount to adjust the playback speed of the audio data within each beat of the first audio segment and the second audio segment respectively includes:
[0014] For each beat among the multiple beats of the first audio segment, calculate the speed change parameter corresponding to the sequence number of the beat according to the proportional relationship between the second BPM information and the beat difference, where the beat difference is the difference between the second BPM information and n times the per-beat variation amount, and n refers to the sequence number value;
[0015] For each beat among the multiple beats of the second audio segment, calculate the speed change parameter corresponding to the sequence number of each beat according to the proportional relationship between the first BPM information and the beat difference;
[0016] Adjust the audio data within each beat of the first audio segment and the second audio segment according to the speed change parameters of the corresponding sequence numbers.
[0017] Optionally, using the per-beat variation amount to adjust the playback speed of the audio data within each beat of the first audio segment and the second audio segment respectively includes:
[0018] According to the preset forward cut-off frequency range, filter out the frequency data below the preset cut-off frequency at each moment in the first audio segment to obtain the output audio segment of the first audio segment;
[0019] According to the preset reverse cut-off frequency range, filter out the frequency data below the preset cut-off frequency at each moment in the second audio segment to obtain the output audio segment of the second audio segment;
[0020] Use the per-beat variation amount to adjust the playback speed of the audio data within each beat of the output audio segment of the first audio segment and the output audio segment of the second audio segment respectively.
[0021] Optionally, after using the per-beat change amount to respectively adjust the playback speed of the audio data within each beat in the first audio segment and the second audio segment to correspondingly obtain a third audio segment and a fourth audio segment, the method further includes:
[0022] Filter out the frequency data below the preset cut-off frequency range at each moment in the third audio segment according to the preset forward cut-off frequency range to obtain the output audio segment of the third audio segment;
[0023] Filter out the frequency data below the preset cut-off frequency range at each moment in the fourth audio segment according to the preset reverse cut-off frequency range to obtain the output audio segment of the fourth audio segment.
[0024] Optionally, the method further includes:
[0025] Use the dependent variable of the decreasing part of the first window function as a weighting factor to weight the output audio segment filtered according to the forward range to reduce the playback volume of the output audio segment;
[0026] Use the dependent variable of the increasing part of the second window function as a weighting factor to weight the output audio segment filtered according to the reverse range to increase the playback volume of the output audio segment.
[0027] Optionally, after determining the first audio segment of the first audio and the second audio segment of the second audio, the method further includes:
[0028] Determine the initial target position point in the first audio segment and the final target position point in the second audio segment according to the number-of-sections information of the audio segment;
[0029] For the target segment between the initial target position point and the final target position point, add target rhythm data to the audio data within each beat of the target segment.
[0030] Optionally, after determining the first audio segment of the first audio and the second audio segment of the second audio, the method further includes:
[0031] Concatenate the second audio segment to the tail of the first audio segment, or concatenate the fourth audio segment to the tail of the third audio segment.
[0032] The second aspect of the embodiments of the present application provides an electronic device, including:
[0033] A central processing unit, a memory, and an input / output interface;
[0034] The memory is a transient storage memory or a persistent storage memory;
[0035] The central processing unit is configured to communicate with the memory and execute the instruction operations in the memory to execute the method described in the first aspect or any specific implementation manner of the first aspect of the embodiments of the present application.
[0036] A third aspect of the embodiments of the present application provides a computer-readable storage medium including instructions, which when running on a computer, cause the computer to execute the method described in the first aspect or any specific implementation manner of the first aspect of the embodiments of the present application.
[0037] A fourth aspect of the embodiments of the present application provides a computer program product including instructions or a computer program, which when running on a computer, cause the computer to execute the method described in the first aspect or any specific implementation manner of the first aspect of the embodiments of the present application.
[0038] From the above technical solutions, it can be seen that the embodiments of the present application have at least the following advantages:
[0039] Calculating and applying the beat-by-beat change amount can effectively and relatively adjust the playback speeds of the audio data included in the first audio segment and the second audio segment based on an investigation basis, so that the time intervals between the beats of the third audio segment obtained after speed adjustment are the same as those between the beats of the fourth audio segment. Furthermore, it effectively ensures that the beat points of the third audio segment and the fourth audio segment present a one-to-one symmetry or strict alignment at the time points (which can be regarded as beats) when they appear. Therefore, the embodiments of the present application help to make multiple audio segments with originally conflicting rhythm information such as beats or drumbeats become harmonious, coherent, and highly consistent in pace when played, thereby improving the sense of rhythm unity and smooth coherence felt by the listener in a multi-song playing scenario. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0041] Figure 1 It is a schematic diagram of the system architecture of the embodiments of the present application;
[0042] Figure 2 It is a flowchart of an audio processing method according to an embodiment of the present application;
[0043] Figure 3 It is another flowchart of the audio processing method according to an embodiment of the present application;
[0044] Figure 4This is a schematic structural diagram of an electronic device according to an embodiment of the present application. Detailed implementation manners
[0045] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0046] Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification, claims and drawings of the present application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order different from that shown or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0047] In the following description, expressions such as "a specific implementation manner" or "an embodiment" are involved, which describe a subset of all possible embodiments. However, it can be understood that "a specific implementation manner" or "an embodiment" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict. In the following description, the term "a plurality" refers to at least two.
[0048] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application. The nouns and terms involved in the embodiments of the present application will be described below, and the nouns and terms involved in the embodiments of the present application are applicable to the following explanations.
[0049] 1) BPM (Beats per minute): It represents the total number of beats within one minute and can be understood as the playing speed of a song.
[0050] 2) beat: It represents the beat information of a song, that is, the time point when a beat (or called a measure) appears, and can be understood as the time interval between adjacent beat points.
[0051] 3) Time Signature: The symbol of the beat. There is a time signature at the beginning of each sheet of music. If the rhythm changes in the middle, the changed time signature will be marked. The time signature is often marked in the form of a fraction, such as 2 / 4, 3 / 4, etc. Among them, the denominator represents the duration of the beat, that is, which kind of note is used as one beat. For example, 2 / 4 means that a quarter note represents one beat, and there are two beats in each measure; the numerator represents how many beats there are in each measure. As mentioned before, in 2 / 4 time, a quarter note is one beat, and there are two beats in a measure. In 3 / 4 time, a quarter note is one beat, and there are three beats in each measure. The reading method is to read the denominator first, and then the numerator. For example, 2 / 4 is called two-four time, 3 / 4 is called three-four time, and 6 / 8 is called six-eight time.
[0052] The audio processing method provided by the embodiments of the present application can be applied to an application environment as Figure 1 shown. Taking songs as an example for the audio of the present application, among them, the terminal 102 communicates with the server 101 through the network, and the data storage system can store the data that the server 101 needs to process; the data storage system 100 can be integrated on the server 101, or can be placed in the cloud or other network servers. The terminal 102 can obtain the song to be adjusted input by the user and send the song to be adjusted to the server 101. The server 101 can analyze information such as the musical score structure information and song representation information based on the obtained song to be adjusted, and perform corresponding strategy adjustments on the song to be adjusted, such as speed adjustment processing, based on this information. In addition, the server 101 can also send the adjusted song to the terminal 102 for playing. Among them, the terminal 102 can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc. The server 101 can be implemented by an independent server or a server cluster composed of multiple servers. It should be noted that the method provided by the embodiments of the present application can be jointly implemented by the terminal device and the server as described above, or can be all implemented on the server side, or can also be all implemented on the terminal device side, which can be specifically determined according to the actual application scenario and is not limited here.
[0053] The method of the present application will be further described in detail below.
[0054] Please refer to Figure 2 , an embodiment of the audio processing method provided by the first aspect of the present application includes the following operation steps 21 to 23:
[0055] 21. Determine the first audio segment of the first audio and the second audio segment of the second audio.
[0056] Among them, the first audio is played before the second audio and at different playing speeds, and the beat lengths of the first audio segment and the second audio segment are the same.
[0057] In music theory knowledge, the musical score structure information includes beats, the number of measures, etc. Through the musical score structure information, the song segments (audio segments) can be specifically delimited. Therefore, taking the medley song scenario as an example, in the face of the application requirement of splicing the second song at the end of the first song (i.e., the first audio), the audio segment composed of the last 8 beat lengths in the first song can be selected as the first audio segment, and the audio segment composed of the first 8 beat lengths in the second song (i.e., the second audio) can be selected as the second audio segment. In addition, since multiple beats can form a measure, it can be understood that, taking 4 / 4 time as an example, the last two measures of the first song can also be correspondingly selected as the first audio segment, and the first two measures of the second song can be selected as the second audio segment. It can be seen that the data lengths (or called beat lengths) of the selected first audio segment and second audio segment can be delimited by the musical score structure information such as the number of beats or the number of measures. Of course, in some user requirement scenarios, the entire first song can also be directly used as the first audio segment, and the entire second song can be used as the second audio segment.
[0058] 22. Obtain the first BPM information of the first audio segment and the second BPM information of the second audio segment, and calculate the beat-by-beat change amount between the first audio segment and the second audio segment.
[0059] Specifically, the BPM information is the number of beats of the audio segment per unit time (which can be regarded as the playing speed). Based on the first BPM information and the second BPM information, calculate the beat-by-beat change amount, which is used to represent the change between the first audio segment and the second audio segment under the beat length.
[0060] In practical applications, calculating the beat-by-beat change amount helps to examine and analyze the feature distribution between the first audio segment and the second audio segment, so as to adjust the playing speeds of the audio data of the two subsequently, and prompt a strict matching relationship between the beat points of the two, such as equal time intervals.
[0061] 23. Use the beat-by-beat change amount to adjust the speeds of the first audio segment and the second audio segment respectively to obtain the third audio segment and the fourth audio segment.
[0062] Use the beat-by-beat change amount to adjust the playing speeds of the audio data within each beat of the first audio segment and the second audio segment respectively to correspondingly obtain the third audio segment and the fourth audio segment, where the time intervals between the beats of the third audio segment and the time intervals between the beats of the fourth audio segment are the same.
[0063] In summary, by referring to the beat-by-beat variation, the playback speed of the audio data within each beat in the first audio segment and the second audio segment can be effectively adjusted in a well-documented manner, ensuring that the beat points of the third audio segment and the fourth audio segment are symmetrically paired or strictly aligned in terms of the occurrence time points. This enables the embodiments of the present application to truly achieve data alignment of the audio at the signal level compared to the traditional methods of directly connecting the heads and tails of multiple songs without any additional processing or manually adjusting the speed using DAW software. It promotes a smooth transition of the playback speed between the spliced songs without jitter, thus automatically and efficiently meeting the needs of the majority of users for rhythm-unified mixed music production at low cost and enhancing the user's coordinated listening experience.
[0064] Based on the above example illustration, some specific possible implementation examples will be provided below. In actual applications, the implementation contents among these examples can be combined as needed according to the corresponding functional principles and application logics.
[0065] Please refer to Figure 3 , another embodiment of the audio processing method provided by the present application, which includes the following operation steps 31 to 35:
[0066] 31. Determine the first audio segment of the first audio and the second audio segment of the second audio.
[0067] In actual applications, the automatic implementation of this solution requires the assistance of various music information retrieval (MIR, Music Information Retrieval) contents, such as beat, BPM, small nodes of the musical score structure, or segmentation points of the song audio paragraphs and other musical score structure information. These contents can be detected by detection methods such as deep neural networks and signal processing. Among them, two different songs, A and B, can obtain corresponding BPM values after detection, denoted as BPM_A and BPM_B respectively; when BPM_A is not equal to BPM_B, it is a common situation where the playback speeds of the two songs are different.
[0068] In some specific scenarios, before or during the use of the following steps, the blank data segments at the beginning and end of the song can be automatically removed to avoid wasting operation resources. This removal method can be based on the loudness of each frame of audio data as the removal feature. For example, the loudness conversion result of each frame of data is 20 * log10(RMS(a frame of data)), where RMS(·) is the root mean square for calculating the loudness of a certain frame of data. If the loudness conversion result is less than a certain threshold, it is considered that the data frame is a blank data frame and should be removed to avoid occupying subsequent operation resources; in addition, when removing the blank data at the beginning, it can be traversed and compared from the first frame backward, and when removing the blank data at the end, it can be traversed and compared from the last frame forward.
[0069] 32. Obtain the first BPM information of the first audio segment and the second BPM information of the second audio segment, and calculate the beat-by-beat change amount between the first audio segment and the second audio segment.
[0070] In some specific examples, the specific implementation process of the operation of "calculating the beat-by-beat change amount between the first audio segment and the second audio segment" in step 32 may include: calculating the BPM difference between the first BPM information and the second BPM information, and calculating the beat-by-beat change amount based on the BPM difference and the beat length of the audio segment.
[0071] It should be noted that the time signatures of the first song (or called song A) and the second song (or called song B) can be inconsistent, as long as the selected first audio segment and second audio segment are song segments with the same beat length, which helps to strictly align each beat data; for example, when song A is in 3 / 4 time and song B is in 6 / 8 time, at this time, the determined first audio segment can be a segment composed of the last 6 beat lengths (i.e., the last two measures) of song A, and the determined second audio segment can be a segment composed of the first 6 beat lengths (i.e., the first measure) of song B; of course, in actual situations, the beat length of the above audio segments can be determined according to the musical score structure information of the specific song as the case may be.
[0072] For the convenience of description and understanding, the following will take the first song and the second song with the same time signature (such as 4 / 4 time) as an operation example. In this case, the last two measures of song A and the first two measures of song B respectively contain audio data with 8 beat lengths. Among them, the playing speeds of song A and song B can be respectively recorded as BPM_A (i.e., the first BPM) and BPM_B (i.e., the second BPM). Then, the calculation formula for the beat-by-beat change amount (which can be recorded as Delta) between the two songs at this time can be as follows:
[0073] Delta = (BPM_B – BPM_A) / the beat length 8 of the audio segment, where BPM_B can be greater than or less than BPM_A.
[0074] 33. Use the beat-by-beat change amount to adjust the speeds of the first audio segment and the second audio segment respectively to obtain the third audio segment and the fourth audio segment.
[0075] Based on the example description of the above step 32, the specific implementation process of step 33 may include the following BPM adjustment operations:
[0076] For each beat in the multiple beats of the first audio segment, according to the proportional relationship between the second BPM information and the beat difference, calculate the speed change parameter corresponding to the sequence number of the beat. The beat difference is the difference between the second BPM information and n times the per-beat change amount, where n refers to the sequence number value. For each beat in the multiple beats of the second audio segment, calculate the speed change parameter corresponding to the sequence number of each beat according to the proportional relationship between the first BPM information and the beat difference. Adjust the audio data within each beat in the first audio segment and the second audio segment according to the speed change parameter of the corresponding sequence number.
[0077] Still taking the above-mentioned first song and second song, both in 4 / 4 time signature, as an example. Specifically, for the first audio segment, the speed change parameter V_A corresponding to the nth beat (or the nth measure) with beat sequence number n can be calculated by the formula V_An = BPM_B / (BPM_B - n * Delta), where n ∈ [1, the beat length 8 of the audio segment], and BPM_B - n * Delta is the beat difference. Then, the audio data within the nth beat in the last 8 beat lengths (i.e., the last two measures) of the first audio segment, which is song A, can be processed with speed change without pitch change according to the speed change factor V_An = BPM_B / (BPM_B - n * Delta). For example, the audio data within the first beat can be processed with speed change without pitch change according to the speed change factor V_A1 = BPM_B / (BPM_B - Delta). Here, the so-called without pitch change in the context can specifically refer to not changing the vibration frequency of the audio data within the current beat.
[0078] Similarly, for the second audio segment, the speed change parameter V_B corresponding to the nth beat (or the nth measure) with beat sequence number n can be calculated by the formula V_Bn = BPM_A / (BPM_B - n * Delta), where n ∈ [1, the beat length 8 of the audio segment]. Then, the audio data within the nth beat in the first 8 beats (i.e., the first two measures) of the second audio segment, which is song B, can be processed with speed change without pitch change according to the speed change factor V_Bn = BPM_A / (BPM_B - n * Delta). For example, the audio data within the first beat can be processed with speed change without pitch change according to the speed change factor V_B1 = BPM_A / (BPM_B - Delta).
[0079] Thus, for the first audio segment and the second audio segment whose data durations are not exactly equal due to different playing speeds and which will not be completely aligned end to end when overlapped, after the above speed adjustment process, the audio data of each beat of the first audio segment and the second audio segment are strictly aligned, achieving a truly high degree of unity in playing speed. That is, the time points (i.e., beats) of the 8 beats of Songs A and B will be aligned one by one. Subsequently, when multiple songs with originally conflicting rhythm information such as beats or drumbeats are played, they become more harmonious, coherent, and highly consistent in pace, improving the sense of rhythm unity and coherence felt by the listener in a multi-song playing scenario. Taking the measure of the first audio segment as an example, the data duration of the aforementioned first audio segment can be regarded as the duration formed by the time period from the starting moment of the first measure (which can be denoted as TA1) to the ending moment of the last measure (which can be denoted as TA2) of the first audio segment, i.e., TA2 - TA1. Similarly, the data duration of the second audio segment can be regarded as the duration formed by the time period from the starting moment of the first measure (which can be denoted as TB1) to the ending moment of the last measure (which can be denoted as TB2) of the second audio segment, i.e., TB2 - TB1.
[0080] Of course, for the case where the playing speeds of Song A and Song B are the same, that is, when BPM_A is equal to BPM_B, the above speed change factor is constantly equal to 1. And since the durations of TA2 - TA1 and TB2 - TB1 are equal, there is no need to perform the above speed change without pitch change process. It can be seen that the speed adjustment process in the embodiments of this application can be used to verify whether the playing speeds of two songs are the same, and the verification basis can be whether the above speed change factor is constantly equal to 1.
[0081] Based on the above example description, to meet the requirements for the fade-in and fade-out of tones in a mixed-song scenario, the specific implementation process of step 33 above may include the following steps 331 to 333:
[0082] 331. Filter the frequency data below the preset cut-off frequency range at each moment in the first audio segment according to the preset positive cut-off frequency range to obtain the output audio segment of the first audio segment. For example, this positive cut-off frequency range may refer to the cut-off frequency trend formed by the preset cut-off frequencies at each moment during the period from the starting moment of the first measure to the ending moment of the last measure of the first high-pass filter for the first audio segment, where the cut-off frequency at a later moment is higher than that at the previous adjacent moment. When the first audio segment is input into the first high-pass filter, the frequency data below the corresponding preset cut-off frequency at each moment in the first audio segment can be filtered out (which can be regarded as filtering) to obtain the output audio segment of the first audio segment.
[0083] 332. Filter the frequency data below the preset cut-off frequency at each moment in the second audio segment according to the preset cut-off frequency reverse range to obtain the output audio segment of the second audio segment. For example, this cut-off frequency reverse range may refer to the cut-off frequency trend formed by the preset cut-off frequencies at each moment during the period from the starting moment of the first section to the ending moment of the last section of the second audio segment by the second high-pass filter, where the cut-off frequency at the later moment is lower than that at the previous adjacent moment; inputting the second audio segment into the second high-pass filter can filter the frequency data below the corresponding preset cut-off frequency at each moment in the second audio segment (which can be regarded as filtering) to obtain the output audio segment of the second audio segment.
[0084] 333 (belonging to the speed adjustment step). Use the beat-by-beat change amount to separately adjust the playback speed of the audio data within each beat in the output audio segment of the first audio segment and the output audio segment of the second audio segment. The specific operation process of step 333 can be as described in the above operation content for adjusting the BPM, which will not be elaborated here.
[0085] Exemplarily, the first audio segment and the second audio segment are respectively song segments with an integer number of bar lengths from song A and song B (or called song A and song B), such as the last 2 bars at the end of song A and the first 2 bars at the beginning of song B. Of course, the specific number of bars of the first audio segment and the second audio segment can be determined according to the situation, and it is not necessarily a 1:1 bar number ratio. For example, the number of bars of the two here can be determined according to the beat lengths of the two described in the above speed adjustment process, and no specific limitation is made. Generally, in the actual song splicing scenario, one of the situations is that the audio content at the tail of the first song (which can be called the fade-out area) is to be overlapped with the audio content at the head of the second song (which can be called the fade-in area) to generate an overlapping area and a mixed audio signal between the two songs; in this case, if the time lengths of the fade-out area and the fade-in area are not equal, the start time point of the fade-out area and the start time point of the fade-in can be adjusted to be aligned, that is, the start of the first audio segment and the start of the second audio segment are aligned to ensure that the required overlapping area and mixed audio signal can be finally generated; of course, if the time lengths of the fade-out area and the fade-in area are equal, then the start time of the fade-out is the start time of the fade-in, and the end time of the fade-out is the end time of the fade-in. During this time from start to end, the signals of the two songs are overlapped together, and there is no need to align their starts. Another situation is that the audio content of the first song stops playing after a certain moment and directly switches to start playing the audio content of the second song, that is, there is no overlapping part between song A and song B.
[0086] For the sake of illustration and understanding, by way of example, still taking the last 2 bars of Song A as the fade-out part, i.e., the first audio segment, and the first 2 bars of Song B as the fade-in part, i.e., the second audio segment. Also, since the spectral components representing the rhythm part of the song (which are often low-frequency components compared to the verse and chorus parts) are mainly concentrated below 250 Hz, the cut-off frequency of the high-pass filter can be set between 20 - 250 Hz. Of course, the specific cut-off frequency value can be adjusted according to the specific situation of the song. In addition, the following described frequency-domain filtering process can be applied to meet the above two situations:
[0087] For the first audio segment, the low-frequency part needs to be faded out, that is, gradually filter out the low-frequency component data of the first audio segment. Within the last 2 bars of Song A, the time period from the start time of the first bar to the end time of the last bar can be denoted as TA1 to TA2. The cut-off frequency (or called the cut-off frequency) of the first high-pass filter can gradually increase from 20 Hz (at time TA1) to 250 Hz (at time TA2). For example, the cut-off frequency at time TA1.5 is an intermediate value such as 135 Hz. Then correspondingly, input the first audio segment into this first high-pass filter to filter out the frequency data in the first audio segment below the corresponding preset cut-off frequency at each moment from TA1 to TA2, and obtain the output audio segment of the first audio segment, so as to filter out the frequency data below the corresponding preset cut-off frequency in the audio data at each moment of the first audio segment, and achieve the fade-out effect. It can be seen that the high-pass filter in the embodiment of the present application has a gradually changing cut-off frequency. The advantage of the gradually changing cut-off frequency is that it makes the spectrum (which can be understood as the vibration frequency) of the first audio segment smoothly transition to the spectrum of the second segment, rather than suddenly changing from one low frequency to another low frequency at a certain moment, which is equivalent to helping to avoid the incoherence of the playback signal caused by the spectral mutation.
[0088] Similarly, for the second audio segment, the low-frequency part needs to be faded in, that is, gradually introduce the low-frequency component data of the second audio segment. Within the first 2 bars of Song B, the time period from the start time of the first bar to the end time of the last bar can be denoted as TB1 to TB2. The cut-off frequency of the gradually changing second high-pass filter can gradually decrease from 250 Hz (at time TB1) to 20 Hz (at time TB2), so that the low-frequency component data of the second audio segment gradually continues to enter, so as to avoid audio content conflict during the transition playback from Song A to Song B due to the radical pitch, which affects the listening experience.
[0089] In summary, for the requirements of one of the above situations, the first high-pass filter will filter out the low-frequency rhythm data components in each moment of the first audio segment that are lower than the cut-off frequency, and the positions of the audio component data filtered out therein can be filled with the low-frequency rhythm data components output by the second high-pass filter at the corresponding moments; for the requirements of the second of the above situations, after the TA2 moment, the audio data content of song A will be replaced by the audio data content of the second audio segment or the subsequent audio segments, so that the second audio segment will be smoothly directly broadcast after the TA2 moment.
[0090] It should be noted that the above first high-pass filter and second high-pass filter can be the same high-pass filter or two high-pass filters; the order of execution of the above steps 331 and 332 is not limited and can also be executed simultaneously; the operation contents of steps 31 to 33 and steps 21 to 23 are similar and will not be elaborated here.
[0091] It can be seen that the above steps 331 and 332 are executed before the speed adjustment step 333. In this case, the objects to be filtered are the first audio segment and the second audio segment, and the filtering effect is achieved before the speed adjustment effect; in some specific examples, differently, operations similar to the above steps 331 to 332 can also be executed after the speed adjustment step 333. Therefore, in this case, the objects to be filtered are the third audio segment and the fourth audio segment, rather than the first audio segment and the second audio segment, and the filtering effect is achieved later than the speed adjustment effect. Because the speed adjustment process changes the data playback speed of the first audio segment and the second audio segment and does not add new spectral content to their spectral signals, that is, it does not change the vibration frequency of the signals. Therefore, looking back at this case, it can also be understood as a filtering and pitch adjustment process for the first audio segment and the second audio segment. Or, in terms of the musical score structure information (such as bar segmentation points) of the audio segment, the audio segments before and after speed adjustment are the same audio segment in terms of the bar data length level. Therefore, in a possible implementation where the filtering effect is achieved later than the speed adjustment effect, after obtaining the third audio segment and the fourth audio segment by speed adjustment, the method described in this application may further include step 34 (such as using a high-pass filter to perform frequency-domain filtering processing on the third audio segment and the fourth audio segment), and this step 34 may specifically include the following steps 341 to 342:
[0092] 341. Filter out the frequency data lower than the preset cut-off frequency at each moment in the third audio segment according to the positive range of the preset cut-off frequency to obtain the output audio segment of the third audio segment. For example, this positive range of the cut-off frequency may refer to the cut-off frequency trend formed by the preset cut-off frequencies at each moment when the first high-pass filter is from the starting moment of the first section to the ending moment of the last section of the third audio segment, where the cut-off frequency at the later moment is higher than the cut-off frequency at the previous adjacent moment; input the third audio segment into the first high-pass filter, then the frequency data lower than the corresponding preset cut-off frequency at each moment in the third audio segment can be filtered out to obtain the output audio segment of the third audio segment.
[0093] 342. Filter the frequency data below the preset cut-off frequency at each moment in the fourth audio segment according to the preset reverse cut-off frequency range to obtain the output audio segment of the fourth audio segment. For example, this cut-off frequency reverse range may refer to the cut-off frequency trend formed by the preset cut-off frequencies at each moment during which the second high-pass filter operates from the starting moment of the first section to the ending moment of the last section of the fourth audio segment, where the cut-off frequency at a later moment is lower than that at the previous adjacent moment; inputting the fourth audio segment into the second high-pass filter can filter out the frequency data below the corresponding preset cut-off frequency at each moment in the fourth audio segment to obtain the output audio segment of the fourth audio segment.
[0094] The operations and application effects in steps 341 to 342 and steps 331 to 332 are similar, and will not be elaborated here specifically; the order of execution of the above steps 341 and 342 is not limited, and they can also be executed simultaneously.
[0095] In some specific examples, after completing the above filtering process to meet the fade-in and fade-out requirements of the pitch of the audio segment, to further meet the fade-in and fade-out requirements of the audio data of the audio segment at the volume level, the audio processing method of the present application may further include step 35 (such as performing a weighting process on the filtered output audio segment), and this step 35 may specifically include the following steps 351 to 352:
[0096] 351. Fade-out processing in the time domain. Use the dependent variable of the decreasing part of the first window function as the weighting factor to weight the output audio segment filtered according to the above forward range (specifically, the output of the first high-pass filter) to reduce the playback volume of the output audio segment; for example, the window length of the first window function is more than twice the corresponding duration from the starting moment of the first section to the ending moment of the last section of the audio segment input to the first high-pass filter.
[0097] 352. Fade-in processing in the time domain. Use the dependent variable of the increasing part of the second window function as the weighting factor to weight the output audio segment filtered according to the above reverse range (specifically, the output of the second high-pass filter) to increase the playback volume of the output audio segment; for example, the window length of the second window function is more than twice the corresponding duration from the starting moment of the first section to the ending moment of the last section of the audio segment input to the second high-pass filter.
[0098] As described above, performing fade - out or fade - in processing on the output audio segment signal of the high - pass filter in the time domain can adjust the volume playback effect of the output audio segment and improve the user's experience of data gradual change in hearing. Exemplarily, for the output audio segment of the high - pass filter, such as the first audio segment or the third audio segment, the weighting factor required for fade - out processing in the time domain can be served by the data in the second half of the first window function, such as the Hann window, that is, the dependent variable of the decreasing part. Among them, the length of the first window is equal to twice the length corresponding to the time period TA2 - TA1; specifically, since the Hann window function consists of a set of data that gradually changes from 0 to 1 and then from 1 to 0, taking the data in its second half means taking the values of the part that gradually changes from 1 to 0 as the weighting factor. Then, multiplying the weighting factors at different times with the corresponding data frames in the first audio segment point - by - point (weighting) can make the data frame finally form a fade - out effect, which is equivalent to making the audio data content of song A no longer appear after time TA2.
[0099] Similarly, for the second audio segment or the fourth audio segment, the weighting factor required for fade - in processing in the time domain can be served by the data in the first half of the second window function, such as the Hann window, that is, the dependent variable of the increasing part. Multiplying the weighting factors at different times with the corresponding data frames in the second audio segment point - by - point (weighting) can make the data frame finally form a fade - in effect, enabling song B to be smoothly played in a gradual change at the end of song A, thereby, to a certain extent, avoiding the problem of the audio rhythms of the two songs conflicting or fighting with each other during playback. Among them, the length of the second window is equal to twice the length corresponding to the time period TB2 - TB1.
[0100] It should be noted that the order of execution of steps 351 and 352 above is not limited. The first window function and the second window function can specifically be the same window function (when TA2 - TA1 = TB2 - TB1), and conversely, they can also be two window functions with different window lengths.
[0101] In addition, since the last few bars of a song (such as the last 2 bars) belong to the coda part in the musical score structure information and are often in the fading stage in terms of emotional expression; the first few bars (such as the first 2 bars) belong to the prelude part in the musical score structure information and are often in the brewing and rising stage in terms of emotional expression. Therefore, in order to make the emotion in the transition stage from song A to song B also become smooth and coherent, target rhythm data can be mixed in this transition stage, such as mixing loop sound elements related to the drumbeat rhythm; thus, based on the above - mentioned examples, after step 31, the audio processing method of the present application can further include step 36, that is, adding target rhythm data. The specific operation process of step 36 includes:
[0102] Determine the initial target position point in the first audio segment and the final target position point in the second audio segment according to the number-of-sections information of the audio segments; for the target segment between the initial target position point and the final target position point, add target rhythm data to the audio data within each beat of the target segment.
[0103] Exemplarily, for the last two bars of Song A as the first audio segment, the time node of the last song segment can be traced back from the tail to serve as the initial target position point, such as the start position point of the coda, which can be denoted as TA_e; for the first two bars of Song B as the second audio segment, the time node of the second song segment can be found starting from the head to serve as the final target position point, such as the end position of the intro, i.e., the switching point in terms of the appearance time between the intro and the interlude, which can be denoted as TB_s. Then, the segment between TA_e and TB_s can be regarded as a complete paragraph structure. Still taking a 4 / 4-time song as an example, add a kick drum loop element to the first beat of each bar within this paragraph, a hi-hat loop element to the second beat, a snare drum loop element to the third beat, and a hi-hat loop element to the fourth beat, so as to unify the overall mood of this paragraph and cover the transition playback gap between Song A and Song B; of course, background audio data such as gradually increasing rain sounds can also be selected to replace the above-mentioned drum loop elements. It should be noted that as mentioned above, the audio segments before and after speed adjustment are the same audio segment in terms of the bar data length, so the operation object of step 36 can also be the third audio segment and the fourth audio segment, that is, the target segment can be specifically composed of the third audio segment and the fourth audio segment at this time.
[0104] After step 31, the audio processing method of the present application may further include a splicing step: splicing the second audio segment to the tail of the first audio segment, or splicing the fourth audio segment to the tail of the third audio segment; in this way, a spliced audio segment as in one of the above situations or the second situation can be provided for the user, and the spliced audio segment can finally present a playback effect with a highly unified playback speed and a smoother rhythm perception. It should be noted that in some specific application scenarios, the execution order of this splicing step and any one of the above filtering operations (such as steps 341 to 342) or speed adjustment operations (such as step 33) or weighting processes (steps 351 and 352) may not be limited, and can be specifically determined according to the user's needs.
[0105] In summary, the audio processing method of the present application can, based on the dimension of the song beat point, start from dimensions such as the time-frequency domain of the audio signal and / or the emotion of the song segmentation structure, efficiently and automatically produce spliced audio segments with highly unified playback rhythm requirements, and avoid signal conflicts and incoherence during the spliced playback of multiple audio segments, which affect the listening experience. In addition, this process does not require the intervention of a professional sound mixer. For the majority of users, there is no additional usage cost and waiting time, and the cost performance is high. Among them, the operation of adjusting the playback speed of the data segment beat by beat (level by level) in combination with the song score structure information can achieve a gradual playback speed adjustment effect, and at the same time ensure the strict alignment of each beat point of the two audio segments, thereby completing the data alignment at the song signal level, enabling multiple songs to achieve a smooth transition in playback speed without sudden changes. In addition, since the low-frequency drum beats in music are the main factors that bring a strong sense of rhythm to users, the combined application of frequency domain filtering and time domain weighting can maximize the avoidance of the problem of mutual conflict and fighting of the low-frequency rhythm information of different songs. Of course, the aforementioned beat-by-beat speed adjustment process can also avoid the occurrence of this problem to a certain extent. Finally, by combining the segmentation information of the song and using relevant loop elements such as drum beats to cover the emotion of the gap between songs, different songs can be connected with a similar paragraph emotion in the transition stage, improving the smooth transition experience of the user's listening sense.
[0106] Please refer to Figure 4 , the electronic device 400 of the embodiment of the present application may include one or more central processing units CPU (CPU, central processing units) 401 and a memory 405, and one or more application programs or data are stored in the memory 405.
[0107] Among them, the memory 405 may be volatile storage or persistent storage. The program stored in the memory 405 may include one or more modules, and each module may include a series of instruction operations on the electronic device. Further, the central processing unit 401 may be set to communicate with the memory 405 and execute a series of instruction operations in the memory 405 on the electronic device 400.
[0108] The electronic device 400 may further include one or more power supplies 402, one or more wired or wireless network interfaces 403, one or more input / output interfaces 404, and / or one or more operating systems, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0109] The central processing unit 401 may perform the operations executed by the foregoing first aspect or any specific method embodiment of the first aspect, which will not be elaborated herein.
[0110] A computer-readable storage medium provided by the present application includes instructions that, when run on a computer, cause the computer to execute the method described in the foregoing first aspect or any specific implementation manner of the first aspect (such as Figure 2 or Figure 3 ).
[0111] A computer program product provided by the present application includes instructions or a computer program that, when run on a computer, cause the computer to execute the method described in the foregoing first aspect or any specific implementation manner of the first aspect.
[0112] It can be understood that in various embodiments of the present application, the sequence numbers of the steps do not indicate the order of execution. The execution order of each step should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0113] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system (if any) and device can refer to the corresponding processes in the foregoing method embodiments and will not be described herein again.
[0114] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system or device, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical or other forms.
[0115] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0116] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0117] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product (computer program product) is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a business server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
Claims
1. An audio processing method, characterized in that, Including: Determine a first audio segment of a first audio and a second audio segment of a second audio, where the first audio is played before the second audio and at different playing speeds, and the beat lengths of the first audio segment and the second audio segment are the same; Obtain first BPM information of the first audio segment and second BPM information of the second audio segment, where the BPM information is the number of beats in the audio segment per unit time; Based on the first BPM information and the second BPM information, calculate the beat-by-beat change amount between the first audio segment and the second audio segment, where the beat-by-beat change amount is used to represent the change between the first audio segment and the second audio segment at the beat length; Use the beat-by-beat change amount to adjust the playing speeds of the audio data within each beat in the first audio segment and the second audio segment respectively, so as to correspondingly obtain a third audio segment and a fourth audio segment, where the time intervals between the beats in the third audio segment are the same as the time intervals between the beats in the fourth audio segment.
2. The audio processing method according to claim 1, wherein The calculating the beat-by-beat change amount between the first audio segment and the second audio segment includes: Calculate the BPM difference between the first BPM information and the second BPM information, and calculate the beat-by-beat change amount based on the BPM difference and the beat length.
3. The audio processing method according to claim 2, wherein Using the beat-by-beat change amount to adjust the playing speeds of the audio data within each beat in the first audio segment and the second audio segment respectively includes: For each beat in the multiple beats of the first audio segment, calculate the speed change parameter corresponding to the serial number of the beat according to the proportional relationship between the second BPM information and the beat difference, where the beat difference is the difference between the second BPM information and n times the beat-by-beat change amount, and n refers to the beat serial number; For each beat in the multiple beats of the second audio segment, calculate the speed change parameter corresponding to the serial number of each beat according to the proportional relationship between the first BPM information and the beat difference; Adjust the audio data within each beat in the first audio segment and the second audio segment according to the speed change parameters of the corresponding serial numbers.
4. The audio processing method according to claim 1, wherein Using the beat-by-beat change amount to adjust the playing speeds of the audio data within each beat in the first audio segment and the second audio segment respectively includes: Filter out the frequency data lower than the preset cut-off frequency range at each moment in the first audio segment according to the preset positive cut-off frequency range to obtain the output audio segment of the first audio segment; Filter out the frequency data lower than the preset cut-off frequency range at each moment in the second audio segment according to the preset reverse cut-off frequency range to obtain the output audio segment of the second audio segment; Use the beat-by-beat change amount to adjust the playing speeds of the audio data within each beat in the output audio segment of the first audio segment and the output audio segment of the second audio segment respectively.
5. The audio processing method according to claim 1, characterized in that After using the beat-by-beat change amount to adjust the playing speeds of the audio data within each beat in the first audio segment and the second audio segment respectively to correspondingly obtain a third audio segment and a fourth audio segment, the method further includes: Filter out the frequency data lower than the preset cut-off frequency range at each moment in the third audio segment according to the preset positive cut-off frequency range to obtain the output audio segment of the third audio segment; Filter the frequency data below the preset cut-off frequency at each moment in the fourth audio segment according to the preset reverse range of the cut-off frequency to obtain the output audio segment of the fourth audio segment.
6. The audio processing method according to claim 4 or 5, characterized in that The method further includes: Using the dependent variable of the decreasing part of the first window function as a weighting factor to weight the output audio segment filtered according to the forward range, so as to reduce the playback volume of the output audio segment; Using the dependent variable of the increasing part of the second window function as a weighting factor to weight the output audio segment filtered according to the reverse range, so as to increase the playback volume of the output audio segment.
7. The audio processing method according to claim 1, wherein After determining the first audio segment of the first audio and the second audio segment of the second audio, the method further includes: Determine the initial target position point in the first audio segment and the final target position point in the second audio segment according to the number-of-sections information of the audio segment; For the target segment between the initial target position point and the final target position point, add target rhythm data to the audio data within each beat of the target segment.
8. The audio processing method according to claim 1 or 7, characterized in that After determining the first audio segment of the first audio and the second audio segment of the second audio, the method further includes: Splice the second audio segment to the tail of the first audio segment, or splice the fourth audio segment to the tail of the third audio segment.
9. An electronic device, characterized in that, Comprising: A central processing unit, a memory, and an input / output interface; The memory is a transient storage memory or a persistent storage memory; The central processing unit is configured to communicate with the memory and execute the instruction operations in the memory to execute the method according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, Comprising instructions, when the instructions are run on a computer, causing the computer to execute the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Contents reproducer and reproduction method
CN101371311A
System and method for automatically beat mixing a plurality of songs using an electronic equipment
CN101689392A