Audio rhythm extraction method and device, and storage medium
Through audio track separation and fusion rhythm information processing, the problem of poor universality of audio rhythm extraction in the prior art is solved, and accurate rhythm information extraction of different types and rhythm audio is achieved.
Patent Information
- Application Number
- CN202510263169.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is poor in extracting audio rhythm point information, making it difficult to process audio of different types and rhythms.
By separating the audio track, the fused rhythm information of at least one track is extracted, and the final rhythm information is generated by combining the basic rhythm information.
It improves the versatility of audio rhythm extraction, can analyze different types and rhythm audio, and extract more accurate rhythm information.
Smart Images

Figure CN120111301A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method, device and storage medium for extracting audio rhythm. Background Art
[0002] A card-point video is a short and powerful video work produced by adding special effects at specific time points and coordinating with the rhythm of music. Its core lies in the perfect matching of audio and picture, so that the music and picture form logical emotional sections, allowing the audience to better feel the rhythm and emotion of the music.
[0003] In related technologies, when extracting the rhythm point information of audio, the audio signal is short-time Fourier transformed to generate a time-frequency representation, and then the peaks are detected in the amplitude spectrum of each time frame, which usually correspond to the rhythm points. Combined with other relevant parameters (such as pitch, timbre, etc.), the basic rhythm structure or duration of the audio is estimated, and finally an output containing specific beat information is generated, including beat position, duration, etc.
[0004] However, the above methods can only process specific audio and have poor versatility. Summary of the invention
[0005] The embodiment of the present application provides a method, device and storage medium for extracting audio rhythm. The technical solution provided by the embodiment of the present application is as follows:
[0006] According to one aspect of an embodiment of the present application, a method for extracting audio rhythm is provided, the method comprising:
[0007] Get the first audio;
[0008] Performing audio track separation on the first audio to obtain at least one audio track of the first audio;
[0009] According to the at least one audio track, fusion rhythm information of the at least one audio track is obtained, and according to the first audio, basic rhythm information of the first audio is obtained, wherein the fusion rhythm information is rhythm information at the audio track level in the first audio, and the basic rhythm information is rhythm information at the audio level of the first audio;
[0010] Final rhythm information of the first audio is obtained according to the fused rhythm information and the basic rhythm information.
[0011] According to one aspect of an embodiment of the present application, a device for extracting audio rhythm is provided, the device comprising:
[0012] An audio acquisition module, used to acquire a first audio;
[0013] An audio track separation module, configured to perform audio track separation on the first audio to obtain at least one audio track of the first audio;
[0014] A rhythm fusion module, configured to obtain fusion rhythm information of the at least one audio track according to the at least one audio track, and obtain basic rhythm information of the first audio according to the first audio, wherein the fusion rhythm information is rhythm information at the audio track level in the first audio, and the basic rhythm information is rhythm information at the audio level of the first audio;
[0015] The rhythm determination module is used to obtain final rhythm information of the first audio according to the fused rhythm information and the basic rhythm information.
[0016] According to one aspect of an embodiment of the present application, a computer device is provided, comprising a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the above-mentioned audio rhythm extraction method.
[0017] According to one aspect of an embodiment of the present application, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is loaded and executed by a processor to implement the above-mentioned audio rhythm extraction method.
[0018] According to one aspect of an embodiment of the present application, a computer program product is provided, the computer program product comprising a computer program, the computer program being loaded and executed by a processor to implement the above-mentioned audio rhythm extraction method.
[0019] The technical solution provided in the embodiments of the present application can bring the following beneficial effects:
[0020] By performing track separation on the first audio and extracting the fused rhythm information of at least one track, it is possible to analyze the rhythm information at the track level in the first audio without being limited to the audio type and audio rhythm of the first audio, thereby improving the versatility of audio rhythm extraction and refining the method for extracting rhythm information, which is helpful to extract more accurate rhythm information of the first audio. When the fused rhythm information of at least one track is obtained, the basic rhythm information of the first audio is extracted, and the fused rhythm information and the basic rhythm information are fused to obtain the final rhythm information of the first audio, so that the final rhythm information includes the rhythm information at the track level and the audio level, further improving the accuracy of the extracted final rhythm information. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a schematic diagram of an implementation environment of a solution provided by an embodiment of the present application;
[0022] Figure 2 is a flow chart of an audio rhythm extraction method provided by an embodiment of the present application;
[0023] Figure 3 is a schematic diagram of a process for determining the i-th suppressed data point provided by an embodiment of the present application;
[0024] Figure 4 is a schematic diagram of a process for determining the i-th smoothed data point provided by an embodiment of the present application;
[0025] Figure 5 is a schematic diagram of a process for determining an i-th enhanced data point provided by an embodiment of the present application;
[0026] Figure 6 is a schematic diagram of rhythm information of each audio track provided by an embodiment of the present application;
[0027] Figure 7 is a schematic diagram of a process for determining an interpolation step of an mth final data point provided by an embodiment of the present application;
[0028] Figure 8 is a schematic diagram of an overall audio rhythm extraction process provided by an embodiment of the present application;
[0029] Fig. 9 is a block diagram of an audio rhythm extraction device provided by an embodiment of the present application;
[0030] Fig.10 It is a structural block diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0031] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0032] Please refer to Figure 1 , which shows a schematic diagram of a solution implementation environment provided by an embodiment of the present application. The solution implementation environment can be implemented as an audio rhythm extraction system. The solution implementation environment may include: a terminal device 10 and a server 20.
[0033] The number of terminal devices 10 can be one or more. The terminal device 10 can be an electronic device such as a mobile phone, a tablet computer, a laptop computer, a desktop computer, a game console, an e-book reader, a multimedia player, a wearable device, an intelligent voice interaction device, a smart home appliance, a vehicle terminal, an aircraft, etc.
[0034] The terminal device 10 may be installed with a client of a target application, and the target application has an audio processing function. When a user inputs audio in the target application, the final rhythm information of the audio is obtained. The present application does not limit the type of the target application, including but not limited to an audio processing program, a music application, a video application, a game application, etc. Optionally, the target application may be an application that needs to be downloaded and installed, or may be a click-to-use application, and the present application does not limit this.
[0035] The server 20 is used to provide background services for the client of the target application installed and running in the terminal device 10. For example, the server 20 can be the background server of the above-mentioned target application. The server 20 can be an independent physical server, or a server cluster composed of multiple servers, or a cloud computing service center. Optionally, the server 20 provides background services for the target applications in multiple terminal devices 10 at the same time. The terminal device 10 and the server 20 can communicate with each other through the network.
[0036] In an embodiment of the present application, after obtaining the first audio, the first audio is subjected to track separation to obtain at least one track of the first audio. Based on the at least one track, fused rhythm information of the at least one track is obtained, and the fused rhythm information is the rhythm information at the track level in the first audio. Based on the first audio, basic rhythm information of the first audio is obtained, and the basic rhythm information is the rhythm information at the audio level of the first audio. Based on the fused rhythm information and the basic rhythm information, the final rhythm information of the first audio is obtained.
[0037] Please refer to Figure 2 , which shows a flow chart of an audio rhythm extraction method provided by an embodiment of the present application. The execution subject of each step of the method may be a computer device. The method may include at least one of the following steps 210 to 240:
[0038] Step 210: Acquire first audio.
[0039] The first audio may be any audio. Optionally, the first audio may be an audio including a vocal track and an accompaniment track, for example, the first audio may be an audio of a singer singing a song following the accompaniment. Optionally, the first audio may also be an audio including only an accompaniment track without a vocal track, for example, the first audio may be an audio of pure music. Optionally, the first audio may also be an audio including only a vocal track without an accompaniment track, for example, the first audio may be an audio track of a singer singing a cappella. If the first audio includes a vocal track, the vocal track may be a vocal track obtained based on real human voice singing, or a vocal track obtained based on mechanically synthesized human voice. If the first audio includes an accompaniment track, the accompaniment track may be an accompaniment track obtained based on real musical instrument performance, or a track obtained based on mechanically synthesized accompaniment.
[0040] Step 220: Perform audio track separation on the first audio to obtain at least one audio track of the first audio.
[0041] Exemplarily, the first audio track is input into the audio track separation model, and the audio track separation model outputs at least one audio track of the first audio.
[0042] The first audio track includes at least one of a vocal track and an accompaniment audio track. The vocal track may include at least one vocal track corresponding to a human voice. For example, if the vocal track is a track obtained by singing only by user 1, the vocal track includes the vocal track corresponding to user 1. If the vocal track is a track obtained by singing by user 1 and user 2, the vocal track includes the human voice corresponding to user 1 and the vocal track corresponding to user 2. The vocal tracks corresponding to the human voices in the vocal track include but are not limited to tracks sung by different groups of people, such as male tracks, female tracks, and children's tracks. The accompaniment audio track may include at least one accompaniment audio track corresponding to a musical instrument. For example, if the accompaniment audio track is a track obtained based only on musical instrument 1, the accompaniment audio track includes the accompaniment audio track corresponding to musical instrument 1. If the accompaniment audio track is a track obtained based on musical instrument 1 and musical instrument 2, the accompaniment audio track includes the accompaniment audio track corresponding to musical instrument 1 and the accompaniment audio track corresponding to musical instrument 2. The accompaniment tracks corresponding to the various instruments in the accompaniment track include but are not limited to tracks generated by different instruments such as drum tracks, piano tracks, guitar tracks, and violin tracks.
[0043] Therefore, at least one audio track of the first audio may include at least one vocal track corresponding to at least one human voice, and at least one accompaniment track corresponding to at least one musical instrument, and the present application does not limit this.
[0044] Step 230, obtains fused rhythm information of at least one audio track based on at least one audio track, and obtains basic rhythm information of the first audio based on the first audio, wherein the fused rhythm information is the rhythm information at the track level in the first audio, and the basic rhythm information is the rhythm information at the audio level of the first audio.
[0045] According to at least one audio track, rhythm information corresponding to at least one audio track is obtained, and the rhythm information corresponding to at least one audio track is fused to obtain fused rhythm information of at least one audio track. Therefore, the fused rhythm information is the rhythm information obtained by fusion of the rhythm information corresponding to at least one audio track, and is the rhythm information at the audio track level in the first audio. The basic rhythm information is obtained by directly extracting the rhythm information from the first audio, and is the rhythm information at the audio level of the first audio.
[0046] The specific process of determining the rhythm information may refer to the following embodiment, which will not be described in detail here.
[0047] Step 240: Obtain final rhythm information of the first audio according to the fused rhythm information and the basic rhythm information.
[0048] Exemplarily, the fused rhythm information and the basic rhythm information are weighted to obtain the final rhythm information of the first audio. The weight of the fused rhythm information and the weight of the basic rhythm information are not limited in this application. For example, the weight of the fused rhythm information and the weight of the basic rhythm information can be the same or different.
[0049] The technical solution provided by the embodiment of the present application performs track separation on the first audio and extracts the fused rhythm information of at least one track, so that the rhythm information at the track level in the first audio can be analyzed without being limited to the audio type and audio rhythm of the first audio, thereby improving the versatility of audio rhythm extraction and refining the method for extracting rhythm information, which helps to extract more accurate rhythm information of the first audio. When the fused rhythm information of at least one track is obtained, the basic rhythm information of the first audio is extracted, and the fused rhythm information and the basic rhythm information are fused to obtain the final rhythm information of the first audio, so that the final rhythm information includes the rhythm information at the track level and the audio level, further improving the accuracy of the extracted final rhythm information.
[0050] Next, the process of determining the fused rhythm information is introduced.
[0051] In some embodiments, step 230 includes at least one of sub-steps 231 - 234 .
[0052] Sub-step 231, for each of the at least one audio track, performing windowing processing on the audio track to obtain N data points corresponding to the audio track, where the data points are used to indicate energy information within each window of the audio track, and N is an integer greater than 1.
[0053] The windowing process is used to segment the audio track into N windows using a sliding window audio segmentation method, obtain the audio information in the N windows, and then obtain N data points corresponding to the audio track based on the audio information in the N windows. Each data point is used to indicate the energy information in the window corresponding to the data point, and the energy information in the window corresponding to each data point is determined by the audio information in the window corresponding to each data point. The specific process of obtaining N data points can refer to the following embodiment, which will not be described in detail here.
[0054] Sub-step 232, performing smoothing processing on the N data points corresponding to the audio track to obtain K smoothed data points corresponding to the audio track.
[0055] Smoothing is used to reduce low-frequency noise and irregular fluctuations in the N data points corresponding to the audio track to improve the overall quality of each data point corresponding to the audio track, as well as to improve the continuity and stability of the audio signal in the audio track, making the sound quality smoother and more natural.
[0056] In some embodiments, the smoothing process includes a suppression process and a removal process, wherein the suppression process is used to suppress low-frequency noise in the audio track, and the removal process is used to remove data points corresponding to low-frequency information in the audio track.
[0057] The suppression process needs to compare the energy intensity of two adjacent data points and calculate the difference in energy intensity between the two adjacent data points. The larger the difference, the stronger the rhythm change between the two adjacent data points in the audio track, and the smaller the difference, the weaker the rhythm change between the two adjacent data points in the audio track. Therefore, the data points with stronger rhythm changes in the audio track need to be protected, and the data points with weaker rhythm changes in the audio track need to be suppressed, so as to achieve protection of high-frequency jitter, ensure the rhythm of the audio signal in the audio track, and suppress low-frequency jitter, and improve the stability and continuity of the audio signal in the audio track.
[0058] The removal process requires comparing three consecutive data points, calculating the difference in energy intensity between two adjacent data points among the three data points, and judging whether the current data point should be retained based on the two differences. The focus is on removing the data points corresponding to the low-frequency information in the audio track to reduce the low-frequency noise in the audio track, suppress the low-frequency jitter in the audio track, and improve the stability and continuity of the audio signal in the audio track.
[0059] The specific process of performing smoothing processing on the N data points corresponding to the audio track may refer to the following embodiment, which will not be described in detail here.
[0060] Sub-step 233, performing enhancement processing on the K smoothed data points corresponding to the audio track to obtain K enhanced data points corresponding to the audio track.
[0061] Enhancement processing is used to enhance the data points with obvious energy intensity changes among the K smoothed data points corresponding to the audio track, and restore them to the energy intensity before smoothing processing, so as to avoid the smoothing processing suppressing high-frequency information while suppressing low-frequency noise, so that the signal vibration in the audio track can be restored by enhancement processing, and the rhythm of the audio signal in the audio track can be enhanced. For data points with insignificant energy intensity changes, enhancement processing can be omitted.
[0062] The specific process of performing enhancement processing on the K smoothed data points corresponding to the audio track can be referred to the following embodiment, which will not be described in detail here.
[0063] Sub-step 234, performing fusion processing on the K enhanced data points corresponding to the at least one audio track, to obtain fused rhythm information of the at least one audio track.
[0064] The fusion process is used to fuse K enhanced data points corresponding to at least one audio track into data points in one audio track, and the data points in one audio track after fusion are fused rhythm information of at least one audio track. The fused rhythm information is used to comprehensively represent the K enhanced data points corresponding to at least one audio track.
[0065] For the specific process of performing the fusion processing on the K enhanced data points corresponding to at least one audio track, reference may be made to the following embodiment, which will not be described in detail here.
[0066] By obtaining N data points corresponding to each track, and performing smoothing and enhancement processing on the N data points corresponding to each track, the high-frequency information in the track is retained while suppressing the low-frequency information in the track, thereby enhancing the rhythm of the audio signal in the track. Finally, by performing fusion processing, the rhythm information at the track level is comprehensively represented by the fused rhythm information, making the representation of the fused rhythm information more accurate and sufficient.
[0067] In some embodiments, sub-step 231 includes at least one of sub-steps 2311 - 2314 .
[0068] Sub-step 2311, using the first sliding window to perform audio segmentation on the audio track, to obtain audio information in N first windows corresponding to the audio track.
[0069] The first sliding window refers to a sliding window with a window length of the first length. Using the first sliding window to perform audio segmentation on the audio track means sequentially intercepting the audio of the first length from the audio track to obtain audio information in N first windows corresponding to the audio track. The first window refers to the audio track window obtained by using the first sliding window to perform audio segmentation on the audio track, and the audio information in the first window refers to the audio information at the corresponding position of the first window in the audio track.
[0070] Among the N first windows, the window lengths of the first N-1 first windows are the first length. However, since the Nth first window obtained after intercepting the N-1th first window is the remaining audio track window, the window length of the Nth first window can be the first length or less than the first length. This application does not limit this.
[0071] The first length corresponding to the above-mentioned first sliding window is set by the technician according to the requirements for extracting data points, and this application does not limit it.
[0072] Sub-step 2312, for each of the audio information in the N first windows corresponding to the audio track, calculate the average value of the sum of squares of the audio amplitudes of at least one sampling point in the first window to obtain the average energy of the audio information in the first window.
[0073] At least one sampling point is set on the audio information in the first window, and the sampling frequency of the audio information in each first window is the same. This application does not limit the sampling frequency of the sampling point.
[0074] The average energy of the audio information in the first window = the average value of the sum of the squares of the audio amplitudes of at least one sampling point in the first window. Exemplarily, the average energy of the audio information in the first window can be expressed as:
[0075]
[0076] Wherein, P represents the number of sampling points in the first window, x(p) represents the audio amplitude of the p-th sampling point in the first window, and E represents the average energy of the audio information in the first window.
[0077] Sub-step 2313, performing normalization processing on the average energy of the audio information in the first window to obtain data points of the audio information in the first window.
[0078] The maximum value of the average energy of the audio information in the N first windows corresponding to the audio track is obtained, and the average energy of the audio information in the first window is divided by the maximum value, so as to realize the normalization processing of the average energy of the audio information in the first window and obtain the data point of the audio information in the first window. For example, the value of the data point of the audio information in the first window can be expressed as E / Emax , E represents the average energy of the audio information in the first window corresponding to the audio track, E max Represents the maximum value of the average energy of the audio information in the N first windows corresponding to the audio track.
[0079] Sub-step 2314, obtaining N data points corresponding to the audio track according to the data points of the audio information in the N first windows corresponding to the audio track.
[0080] The above steps 2312 to 2313 are respectively performed on the audio information in the N first windows corresponding to the audio track to obtain data points of the audio information in the N first windows corresponding to the audio track, that is, to obtain N data points corresponding to the audio track.
[0081] Through the above steps, each audio track is processed into N data points respectively, and the energy situation in the audio track is represented in the form of data points. The energy changes at different moments in the audio track can be more intuitively seen, so that the original rhythm changes in each audio track can be more intuitively felt based on the energy changes, which facilitates the execution of subsequent processing steps for the data points and enhances the sense of rhythm in the audio track.
[0082] In some embodiments, the smoothing process includes suppression and removal, wherein the suppression process is used to suppress low-frequency noise in the audio track, and the removal process is used to remove data points corresponding to low-frequency information in the audio track. Sub-step 232 includes at least one of sub-steps 2321-2322.
[0083] Sub-step 2321, performing suppression processing on the N data points corresponding to the audio track to obtain N suppressed data points corresponding to the audio track.
[0084] In some embodiments, sub-step 2321 includes at least one of sub-steps A1 to A3.
[0085] Sub-step A1, calculating the difference between the i-th data point and the i-1-th suppressed data point, where i is an integer greater than 1.
[0086] For example, reference may be made to Figure 3 As shown, the difference between the i-th data point and the i-1-th suppressed data point can be expressed as:
[0087] d(i)=y(i)-y ′ (i-1)
[0088] Where d(i) represents the difference between the i-th data point and the i-1-th suppressed data point, y(i) represents the value of the i-th data point, and y ′ (i-1) represents the value of the i-1th data point after the suppression process, and i is an integer greater than 1 and less than or equal to N.
[0089] Sub-step A2, obtaining the suppression coefficient of the ith data point according to the difference between the ith data point and the i-1th data point after the suppression process, wherein the suppression coefficient of the ith data point is negatively correlated with the difference between the ith data point and the i-1th data point after the suppression process.
[0090] First, according to the difference between the i-th data point and the i-1-th suppressed data point, the first suppression coefficient of the i-th data point is obtained, and the first suppression coefficient of the i-th data point is positively correlated with the difference between the i-th data point and the i-1-th suppressed data point. The larger the difference between the i-th data point and the i-1-th suppressed data point, the larger the first suppression coefficient of the i-th data point.
[0091] For example, reference may be made to Figure 3 As shown, the first suppression coefficient of the i-th data point can be expressed as:
[0092] b1(i)=d(i)×scale
[0093] Wherein, b1(i) represents the first suppression coefficient of the i-th data point, scale represents the first preset coefficient, and the first preset coefficient is a preset parameter.
[0094] Secondly, according to the first suppression coefficient of the ith data point, the second suppression coefficient of the ith data point is obtained, and the second suppression coefficient of the ith data point is negatively correlated with the difference between the ith data point and the i-1th suppression-processed data point. The larger the difference between the ith data point and the i-1th suppression-processed data point, the smaller the second suppression coefficient of the ith data point.
[0095] For example, reference may be made to Figure 3 As shown, the second suppression coefficient of the i-th data point can be expressed as:
[0096] b2(i)=-b1(i) ∧ 2
[0097] Wherein, b2(i) represents the second suppression coefficient of the i-th data point.
[0098] Finally, the suppression coefficient of the i-th data point is obtained according to the second suppression coefficient of the i-th data point. The suppression coefficient of the i-th data point is negatively correlated with the difference between the i-th data point and the i-1-th suppression-processed data point. The larger the difference between the i-th data point and the i-1-th suppression-processed data point, the smaller the suppression coefficient of the i-th data point.
[0099] For example, reference may be made to Figure 3 As shown, the suppression coefficient of the i-th data point can be expressed as:
[0100] b(i)=exp(b2(i))×smoothin
[0101] Wherein, b(i) represents the suppression coefficient of the i-th data point, and b2(i) is subjected to exponential calculation with base e and multiplied by a preset coefficient to obtain b. smoothing represents the second preset coefficient, which is a preset parameter.
[0102] Sub-step A3, obtaining the i-th suppressed data point according to the suppression coefficient of the i-th data point, the i-th data point and the i-1-th suppressed data point.
[0103] For example, reference may be made to Figure 3 As shown, the value of the i-th suppressed data point can be expressed as:
[0104] y ′ (i) = y ′ (i-1)×b(i)+y(i)×(1-b(i))
[0105] Among them, y ′ (i) represents the data point after the i-th suppression process.
[0106] According to the above steps, the data points after the second suppression process to the Nth suppression process can be calculated. For the data point after the first suppression process, no suppression process is performed on it, which means that the value of the data point after the first suppression process is equal to the original value of the first data point. ′ (i-1) is the value of the first data point.
[0107] The above steps realize the suppression processing of N data points corresponding to the audio track, so that when the value of each data point changes greatly compared with the value of the previous data point, the change of the data point is controlled to be stable, thereby suppressing the low-frequency jitter in the audio track and improving the stability and continuity of the audio signal in the audio track.
[0108] Sub-step 2322, performing removal processing on the N suppressed data points corresponding to the audio track to obtain K smoothed data points corresponding to the audio track.
[0109] In some embodiments, sub-step 2322 includes at least one of sub-steps A4 to A6.
[0110] Sub-step A4, performing beat extraction on the first audio to obtain beat information of the first audio, wherein the beat information of the first audio includes beat information corresponding to N data points respectively, and the beat information corresponding to the data points is used to indicate the beat information of the first window corresponding to the data point in the first audio.
[0111] The first sliding window is used to perform audio segmentation on the first audio to obtain audio information in N third windows corresponding to the first audio, and beat extraction is performed on the audio information in the N third windows to obtain beat information corresponding to N data points. Performing beat extraction on the first audio can be understood as performing beat extraction on the audio information in the N third windows.
[0112] The third window refers to an audio window obtained by performing audio segmentation on the first audio using the first sliding window, and the audio information in the third window refers to the audio information at the corresponding position of the third window in the first audio. The beat information corresponding to each data point is used to indicate the beat information of the third window corresponding to the data point in the first audio, specifically refers to the number of beats in the third window corresponding to each data point.
[0113] The window lengths of the N third windows are the same as the window lengths of the above-mentioned N first windows, that is, among the N third windows, the window lengths of the first N-1 third windows are the first length, but since the Nth third window obtained after intercepting the N-1th third window is the remaining audio window, the window length of the Nth third window can be the first length or less than the first length, and the present application does not limit this.
[0114] Sub-step A5, obtaining the removal threshold of the i-th data point according to the beat information of the i-th data point and the maximum and minimum values of the beat information corresponding to the N data points respectively, wherein the removal threshold of the i-th data point is negatively correlated with the beat information of the i-th data point, and i is an integer greater than 1.
[0115] The beat information of the i-th data point refers to the number of beats in the third window corresponding to the i-th data point.
[0116] First, the beat coefficient of the i-th data point is obtained according to the beat information of the i-th data point and the maximum and minimum values of the beat information corresponding to the N data points. The beat coefficient of the i-th data point is positively correlated with the beat information of the i-th data point. The larger the value of the beat information of the i-th data point, the larger the beat coefficient of the i-th data point.
[0117] For example, reference may be made to Figure 4 As shown, the beat coefficient of the i-th data point can be expressed as:
[0118] br(i)=(bpm(i)-min_bpm) / (max_bpm-min_bpm)
[0119] Among them, br(i) represents the beat coefficient of the i-th data point, bpm(i) represents the beat information of the i-th data point, min_bpm represents the minimum value of the beat information corresponding to the N data points, and max_bpm represents the maximum value of the beat information corresponding to the N data points.
[0120] Secondly, according to the beat coefficient of the ith data point and the preset beat threshold, the removal threshold of the ith data point is obtained, and the removal threshold of the ith data point is negatively correlated with the beat information of the ith data point. The larger the value of the beat information of the ith data point, the smaller the removal threshold of the ith data point.
[0121] For example, reference may be made to Figure 4 As shown, the removal threshold of the i-th data point can be expressed as:
[0122] r(i)=thres×(1-br(i))
[0123] Wherein, r(i) represents the removal threshold of the i-th data point, and thres represents the preset beat threshold.
[0124] Sub-step A6, performing removal processing on the N suppressed data points corresponding to the audio track according to the removal thresholds respectively corresponding to the N data points, to obtain K smoothed data points corresponding to the audio track.
[0125] By combining the beat information corresponding to N data points, the N suppressed data points corresponding to the audio track are removed, so that the smoothed data points can match the beat information. When the beat is faster, the low-frequency information in the audio track is protected, and when the beat is slower, the high-frequency information in the audio track is protected, thereby enhancing the rhythm of the audio signal in the audio track.
[0126] In some embodiments, sub-step A6 includes at least one of sub-steps A61 to A68.
[0127] Sub-step A61, calculate the first difference and the second difference, the first difference refers to the difference between the i-th suppressed data point and the i-1-th suppressed data point, and the second difference refers to the difference between the i+1-th suppressed data point and the i-th suppressed data point.
[0128] The difference between the i-th suppressed data point and the i-1-th suppressed data point is calculated to obtain a first difference. The difference between the i+1-th suppressed data point and the i-th suppressed data point is calculated to obtain a second difference.
[0129] For example, reference may be made to Figure 4 As shown, the first difference and the second difference can be expressed as:
[0130] d1(i)=y ′ (i)-y ′ (i-1)
[0131] d2(i)=y ′ (i+1)-y ′ (i)
[0132] Where d1(i) represents the first difference, d2(i) represents the second difference, and y ′ (i) represents the data point after the i-th suppression process, y ′ (i-1) represents the i-1th suppressed data point, y ′ (i+1) represents the i+1th data point after suppression processing.
[0133] Sub-step A62, calculating the product of the first difference and the second difference.
[0134] For example, reference may be made to Figure 4 As shown, the first difference and the second difference can be expressed as d1(i)×d2(i).
[0135] If the product of the first difference and the second difference is greater than 0, it means that the first difference and the second difference are both greater than 0, or the first difference and the second difference are both less than 0. If the first difference and the second difference are both greater than 0, it means that the energy information shows a continuous upward trend from the i-1th data point to the i+1th data point. If the first difference and the second difference are both less than 0, it means that the energy information shows a continuous downward trend from the i-1th data point to the i+1th data point.
[0136] If the product of the first difference and the second difference is equal to 0, it means that at least one of the first difference and the second difference is 0. If the first difference is 0, it means that the energy information has not changed from the i-1th data point to the i-th data point. If the second difference is 0, it means that the energy information has not changed from the i-th data point to the i+1th data point.
[0137] If the product of the first difference and the second difference is less than 0, it means that one of the first difference and the second difference is greater than 0 and the other is less than 0. If the first difference is greater than 0 and the second difference is less than 0, it means that from the i-1th data point to the i+1th data point, the energy information first increases and then decreases, and the value of the i-th data point is the largest. If the first difference is less than 0 and the second difference is greater than 0, it means that from the i-1th data point to the i+1th data point, the energy information first decreases and then increases, and the value of the i-th data point is the smallest.
[0138] Sub-step A63, when the product of the first difference and the second difference is greater than 0, determine the removal result of the i-th suppressed data point as removing the i-th suppressed data point; and perform removal processing on the i+2-th suppressed data point.
[0139] For example, reference may be made to Figure 4 As shown in the figure, when the removal process is performed on the i-th suppressed data point, when d1(i)×d2(i)>0, k(i)=y ′ (i+1), i=i+2. Among them, k(i)=y ′ (i+1) means removing the i-th suppressed data point and retaining the i+1-th suppressed data point. i=i+2 means after performing the removal process on the i-th suppressed data point, the removal process is performed on the i+2-th suppressed data point.
[0140] Sub-step A64, when the product of the first difference and the second difference is less than or equal to 0, and the absolute value of the first difference is greater than or equal to the removal threshold of the i-th data point, and the absolute value of the second difference is less than or equal to the removal threshold of the i-th data point, determines that the removal processing result of the i-th suppressed data point is to retain the i-th suppressed data point; and performs removal processing on the i+2-th suppressed data point.
[0141] For example, reference may be made to Figure 4 As shown in the figure, when the removal process is performed on the i-th suppressed data point, when d1(i)×d2(i)≤0, |d1(i)|≥r(i), |d2(i)|≤r(i), k(i)=y ′ (i), i=i+2. Among them, k(i)=y ′ (i) indicates that the i-th suppressed data point is retained, and i=i+2 indicates that after the i-th suppressed data point is removed, the i+2-th suppressed data point is removed.
[0142] Sub-step A65, when the product of the first difference and the second difference is less than or equal to 0, and the absolute value of the first difference is less than or equal to the removal threshold of the i-th data point, and the absolute value of the second difference is less than or equal to the removal threshold of the i-th data point, determines that the removal processing result of the i-th suppressed data point is to remove the i-th suppressed data point; and performs removal processing on the i+2-th suppressed data point.
[0143] For example, reference may be made to Figure 4As shown, when performing the removal process on the data point after the i-th suppression process, in the case where d1(i)×d2(i)≤0, and |d1(i)|≤r(i), |d2(i)|≥r(i), k(i) = y ′ (i + 1), i = i + 2. Where k(i) = y ′ (i + 1) means removing the data point after the i-th suppression process and retaining the data point after the (i + 1)-th suppression process, and i = i + 2 means that after performing the removal process on the data point after the i-th suppression process, performing the removal process on the data point after the (i + 2)-th suppression process.
[0144] Sub-step A66, in the case where the product of the first difference and the second difference is less than or equal to 0, and the absolute value of the first difference is greater than or equal to the removal threshold of the i-th data point, and the absolute value of the second difference is greater than or equal to the removal threshold of the i-th data point, determine that the removal result of the data point after the i-th suppression process is to retain the data point after the i-th suppression process; perform the removal process on the data point after the (i + 1)-th suppression process.
[0145] Exemplarily, reference can be made to Figure 4 As shown, when performing the removal process on the data point after the i-th suppression process, in the case where d1(i)×d2(i)≤0, and |d1(i)|≥r(i), |d2(i)|≥r(i), k(i) = y ′ (i), i = i + 1. Where k(i) = y ′ (i) means retaining the data point after the i-th suppression process, and i = i + 1 means that after performing the removal process on the data point after the i-th suppression process, performing the removal process on the data point after the (i + 1)-th suppression process.
[0146] Sub-step A67, in the case where the product of the first difference and the second difference is less than or equal to 0, and the absolute value of the first difference is less than the removal threshold of the i-th data point, and the absolute value of the second difference is less than the removal threshold of the i-th data point, obtain the data point after the i-th smoothing process according to the removal results of the data point after the i-th suppression process and the data point after the (i - 1)-th suppression process; perform the removal process on the data point after the (i + 2)-th suppression process.
[0147] Exemplarily, reference can be made to Figure 4 As shown, when performing the removal process on the data point after the i-th suppression process, in the case where d1(i)×d2(i)≤0, and |d1(i)|<r(i), |d2(i)|<r(i), k(i) = y ′ (i) / 2 + k(i - 1) / 2, i = i + 2. Where k(i) = y ′(i) / 2+k(i-1) / 2 represents the value of the i-th data point after smoothing, and i=i+2 represents that after performing removal processing on the i-th data point after suppression processing, removal processing is performed on the i+2-th data point after suppression processing.
[0148] Sub-step A68, obtaining K smoothing-processed data points corresponding to the audio track according to the removal processing results of the N suppressed data points corresponding to the audio track.
[0149] According to the above steps, K smoothed data points corresponding to the audio track are obtained. Since some data points may be removed during the removal process, the number of K smoothed data points corresponding to the audio track may be less than N, that is, K is less than or equal to N.
[0150] Through the above steps, the N suppressed data points corresponding to the audio track are removed respectively, and the data points corresponding to some low-frequency information in the N suppressed data points are removed, thereby reducing the low-frequency noise in the audio track, suppressing the low-frequency jitter in the audio track, and improving the stability and continuity of the audio signal in the audio track.
[0151] In some embodiments, sub-step 233 includes at least one of sub-steps 2331 - 2334 .
[0152] Sub-step 2331, calculate the third difference and the fourth difference, the third difference refers to the difference between the j-th smoothed data point and the j-1-th smoothed data point, and the fourth difference refers to the difference between the j+1-th smoothed data point and the j-th smoothed data point, where j is an integer greater than 1.
[0153] The difference between the j-th smoothed data point and the j-1-th smoothed data point is calculated to obtain a third difference. The difference between the j+1-th smoothed data point and the j-th smoothed data point is calculated to obtain a fourth difference, where j is an integer greater than 1 and less than or equal to K.
[0154] For example, reference may be made to Figure 5 As shown, the third difference and the fourth difference can be expressed as:
[0155] d1(j)=k(j)-k(j-1)
[0156] d2(j)=k(j+1)-k(j)
[0157] Among them, d1(j) represents the third difference, d2(j) represents the fourth difference, k(j) represents the j-th smoothed data point, k(j-1) represents the j-1-th smoothed data point, and k(j+1) represents the j+1-th smoothed data point.
[0158] Sub-step 2332, calculate the product of the third difference and the fourth difference.
[0159] For example, reference may be made to Figure 5 As shown, the third difference and the fourth difference can be expressed as d1(j)×d2(j).
[0160] Sub-step 2333, when the j-th smoothed data point meets the enhancement condition, enhancement processing is performed on the j-th smoothed data point to obtain the j-th enhanced data point.
[0161] The enhancement condition is that the product of the third difference and the second difference is less than or equal to 0, the absolute value of the third difference is greater than the removal threshold of the i-th data point, and the absolute value of the fourth difference is greater than the removal threshold of the i-th data point.
[0162] For example, reference may be made to Figure 5 As shown in the figure, when d1(j)×d2(j)<0, |d1(j)|>r(i), |d2(j)|>r(i), s(j) is returned. Returning s(j) means that the jth smoothed data point is enhanced to obtain the jth enhanced data point, and s(j) means the jth enhanced data point.
[0163] Sub-step 2334, when the j-th smoothed data point does not meet the enhancement condition, determine the j-th smoothed data point as the j-th enhanced data point.
[0164] For example, reference may be made to Figure 5 As shown, when one of the enhancement conditions is not satisfied, k(j) is returned. Returning k(j) indicates that the enhancement process is not performed on the jth smoothed data point, which can also be understood as determining the jth smoothed data point as the jth enhanced data point, that is, s(j) = k(j).
[0165] In some embodiments, the jth enhanced data point is restored to the data point before smoothing, that is, the jth enhanced data point is a normalized point, s(j)=y(t), the position of the jth enhanced data point among the N data points is the tth, and the tth data point is determined as the jth enhanced data point.
[0166] The K enhanced data points corresponding to the audio track obtained above can be considered as the rhythm information of the audio track.
[0167] Through the above steps, the K smoothed data points corresponding to the audio track are enhanced, and the maximum and minimum values in the K smoothed data points corresponding to the audio track are enhanced, thereby enhancing the rhythm of the audio signal in the audio track.
[0168] In some embodiments, sub-step 234 includes at least one of sub-steps 2341 - 2342 .
[0169] Sub-step 2341, for each audio track in at least one audio track, perform partitioning processing on K enhanced data points corresponding to the audio track to obtain flat areas and non-flat areas corresponding to the audio track, and the numerical change of the data points in the flat area is smaller than the numerical change of the data points in the non-flat area.
[0170] The partitioning process is used to divide the K enhanced data points corresponding to the audio track into flat areas and non-flat areas, so that the overall numerical changes of the K enhanced data points corresponding to the audio track can be represented by the division of flat areas and non-flat areas.
[0171] In some embodiments, sub-step 2341 includes at least one of sub-steps B1 to B3.
[0172] Sub-step B1, using a second sliding window to perform data point segmentation on K enhanced data points corresponding to the audio track, to obtain a data point set within at least one second window, wherein each data point set within the second window includes at least two enhanced data points.
[0173] The second sliding window refers to a sliding window with a window length of the second length. Using the second sliding window to perform audio segmentation on the audio track means sequentially intercepting the audio of the second length from the audio track to obtain audio information in at least one second window corresponding to the audio track. The second window refers to a data point window obtained by performing data point segmentation on the audio track using the second sliding window, and the data point set in the second window includes at least one data point at a corresponding position of the second window among the K enhanced data points corresponding to the audio track.
[0174] In at least one second window, except for the last second window, the window length of the previous second window is the second length. However, since the last second window obtained after intercepting the previous second window is the remaining data point window, the window length of the last second window can be the second length or less than the second length. This application does not limit this.
[0175] The second length corresponding to the second sliding window is set by the technicians according to the partitioning requirements of the data points, and this application does not limit it.
[0176] Sub-step B2, when the difference between the maximum value and the minimum value in the data point set in the second window is less than or equal to the first value, the data point set in the second window is determined as a flat area of the audio track.
[0177] Compare the maximum and minimum values of the data point set in the second window. If the maximum and minimum values of the data point set in the second window are less than or equal to the first value, it can be considered that the value change of the data point in the second window is small, and the data point set in the second window can be determined as a flat area of the audio track. The first value is set by the technician according to the partition requirements of the data point, and this application does not limit it.
[0178] Sub-step B3, when the difference between the maximum value and the minimum value in the data point set in the second window is greater than the first value, the data point set in the second window is determined as a non-flat area of the audio track.
[0179] When the difference between the maximum and minimum values in the data point set in the second window is greater than the first value, it can be considered that the value variation of the data points in the second window is large, and the data point set in the second window can be determined as a non-flat area of the audio track.
[0180] The above method can be used to obtain the respective flat areas and non-flat areas in at least one audio track, thereby facilitating the subsequent fusion processing of the rhythm information of at least one audio track based on the flat areas and non-flat areas, improving the efficiency of the fusion processing, avoiding fusion processing based on individual data points, and increasing the fusion time.
[0181] Sub-step 2342, according to the pre-set processing priority of at least one audio track, performing point filling processing on the flat area of the first audio track to obtain the fused rhythm information of the at least one audio track, the point filling processing is used to enhance the numerical change of the data points in the flat area of the first audio track, and the first audio track is the audio track with the first processing priority in the at least one audio track.
[0182] Different types of tracks are pre-set with processing priorities, which refer to the priority of performing point filling processing, which is used to replace the non-flat area in the track whose processing priority is after the first track with the flat area in the first track, so as to enhance the numerical change of the data points in the flat area of the first track. The higher the processing priority, the higher the priority of the processing, the higher the priority of the processing, the more preferred it is to replace the non-flat area in the track with the flat area in the first track.
[0183] For example, if at least one audio track includes a drum track, a piano track, a guitar track, and a vocal track, the processing priority of the at least one audio track may be arranged in the following order: drum track, piano track, guitar track, vocal track.
[0184] In some embodiments, sub-step 2342 includes at least one of sub-steps B4 to B6.
[0185] Sub-step B4, according to the processing priority of at least one audio track, obtaining non-flat areas corresponding to other audio tracks located after the first audio track.
[0186] Exemplarily, if at least one audio track includes a drum track, a piano track, a guitar track, and a vocal track, the drum track is the first audio track, and non-flat areas corresponding to the piano track, the guitar track, and the vocal track are obtained.
[0187] Sub-step B5, according to the processing priorities of other audio tracks, the smooth area of the first audio track is replaced by the non-smooth area of other audio tracks with higher processing priorities, so as to obtain the fused smooth area of the first audio track.
[0188] The smooth area of the first track after blending is the non-smooth area of the other tracks that are replaced.
[0189] For example, if the first track includes the flat area 1-1, flat area 1-2, non-flat area 1-1, non-flat area 1-2 and flat area 1-3 in sequence, the piano track includes the flat area 2-1, non-flat area 2-1, non-flat area 2-2, flat area 2-2 and flat area 2-3 in sequence, the guitar track includes the non-flat area 3-1, non-flat area 3-2, non-flat area 3-3, flat area 3-1 and flat area 3-2 in sequence, and the vocal track includes the flat area 4-1, flat area 4-2, non-flat area 4-1, non-flat area 4-2 and non-flat area 4-3 in sequence, then it is necessary to perform point filling processing on the flat area 1-1, flat area 1-2 and flat area 1-3 in the first track. The smooth region 1-1 in the first track is replaced by the non-smooth region 3-1 in the guitar track, the smooth region 1-2 in the first track is replaced by the non-smooth region 2-1 in the piano track, and the smooth region 1-3 in the first track is replaced by the non-smooth region 4-2 in the vocal track. The smooth region after fusion of the first track is obtained, that is, the smooth region after fusion of the first track includes the non-smooth region 3-1 in the guitar track, the non-smooth region 2-1 in the piano track, and the non-smooth region 4-2 in the vocal track.
[0190] Sub-step B6, obtaining fusion rhythm information of at least one audio track according to the fused smooth area of the first audio track and the non-smooth area of the first audio track.
[0191] The fusion information of at least one audio track refers to a data point set of the fused first audio track, including a fused flat region of the first audio track and a non-flat region of the first audio track.
[0192] For example, the fused rhythm information may include a non-flat area 3-1, a non-flat area 2-1, a non-flat area 1-1, a non-flat area 1-2, and a non-flat area 4-2 in sequence.
[0193] In some embodiments, if the flat area of the first audio track does not have a non-flat area of other audio tracks at a corresponding position, the flat area of the first audio track is retained. For example, if the human voice track in the above example includes flat area 4-1, flat area 4-2, non-flat area 4-1, non-flat area 4-2 and flat area 4-3 in sequence, the fused rhythm information may include non-flat area 3-1, non-flat area 2-1, non-flat area 1-1, non-flat area 1-2 and flat area 1-3 in sequence.
[0194] By using the above method, the flat area in the first audio track is supplemented with points, and the numerical changes of the data points in the flat area of the first audio track are enhanced, so that the numerical changes of the final fused rhythm information are more obvious, further enhancing the rhythm of the audio signal in the audio track.
[0195] In some embodiments, step 230 includes at least one of sub-steps 235 - 238 .
[0196] Sub-step 235: Use the first sliding window to perform audio segmentation on the first audio to obtain audio information in N third windows corresponding to the first audio, where N is an integer greater than 1.
[0197] Using the first sliding window to perform audio segmentation on the first audio refers to sequentially intercepting audio of the first length from the first audio to obtain audio information in N third windows corresponding to the first audio. The third window refers to an audio window obtained by using the first sliding window to perform audio segmentation on the first audio, and the audio information in the third window refers to the audio information at the corresponding position of the third window in the first audio.
[0198] The window lengths of the N third windows are the same as the window lengths of the above-mentioned N first windows, that is, among the N third windows, the window lengths of the first N-1 third windows are the first length, but since the Nth third window obtained after intercepting the N-1th third window is the remaining audio window, the window length of the Nth third window can be the first length or less than the first length, and the present application does not limit this.
[0199] Sub-step 236, for each of the audio information in the N third windows corresponding to the first audio, calculate the average value of the sum of squares of the audio amplitudes of at least one sampling point in the third window to obtain the average energy of the audio information in the third window.
[0200] At least one sampling point is set on the audio information in the third window, and the sampling frequency of the audio information in each third window is the same. This application does not limit the sampling frequency of the sampling point.
[0201] The average energy of the audio information in the third window=the average value of the sum of the squares of the audio amplitudes of at least one sampling point in the third window. Exemplarily, the average energy of the audio information in the third window can be expressed as:
[0202]
[0203] Wherein, Q represents the number of sampling points in the third window, x(q) represents the audio amplitude of the qth sampling point in the third window, and E′ represents the average energy of the audio information in the third window. Here, Q and P may be the same or different, and this application does not limit this.
[0204] Sub-step 237, performing normalization processing on the average energy of the audio information in the third window to obtain data points of the audio information in the third window.
[0205] The maximum value of the average energy of the audio information in the N third windows corresponding to the first audio is obtained, that is, the global maximum average energy in the first audio is obtained, and the average energy of the audio information in the third window is divided by the maximum value, so as to realize the normalization processing of the average energy of the audio information in the third window and obtain the data point of the audio information in the third window. Exemplarily, the value of the data point of the audio information in the third window can be expressed as E' / E' max .
[0206] Sub-step 238, obtaining basic rhythm information of the first audio according to the data points of the audio information in the N third windows corresponding to the first audio.
[0207] The basic rhythm information of the first audio includes data points of audio information within N third windows corresponding to the first audio.
[0208] Through the above steps, the first audio is processed into N data points, and the energy situation in the first audio is represented in the form of data points, which facilitates weighted processing of the basic rhythm information and the fused rhythm information to obtain the final rhythm information.
[0209] In some embodiments, the above method further includes step 240 , and step 240 includes at least one step among sub-steps 241 - 242 .
[0210] Sub-step 241: extract at least one peak point from the final rhythm information of the first audio.
[0211] Exemplarily, a peak detection algorithm is used to perform peak extraction on the final rhythm information of the first audio to obtain at least one peak point.
[0212] You can refer to Figure 6 As shown, after the above steps, according to the first audio, the rhythm information of the drum track (K enhanced data points corresponding to the drum track), the rhythm information of the piano track (K enhanced data points corresponding to the drum track), the rhythm information of the guitar track (K enhanced data points corresponding to the drum track), and the rhythm information of the vocal track (K enhanced data points corresponding to the drum track) are obtained in turn. According to the rhythm information of the drum track, the rhythm information of the piano track, the rhythm information of the guitar track and the rhythm information of the vocal track, the fused rhythm information is obtained, and then combined with the basic rhythm information, the final rhythm information of the first audio is obtained. From Figure 6 It can be seen that the final rhythm information contains multiple peak points.
[0213] Sub-step 242, based on at least one peak point, add separate animation effects at time points corresponding to at least one peak point in the first video to generate a processed first video.
[0214] Individual animation effects refer to animation effects added at the time point corresponding to the peak point in the first video. Individual animation effects match the peak point and appear only at the time point corresponding to the peak point. Individual animation effects include but are not limited to flash effects, transition effects, subtitle effects, sound effects, etc.
[0215] Optionally, the duration of the first audio may be shorter than the duration of the first video, in which case it is not possible to add separate animation effects to the first video after the first audio ends. Optionally, the duration of the first audio may be longer than the duration of the first video, in which case it is not possible to add separate animation effects corresponding to the remaining peak points in the first audio after the first video ends. Optionally, the duration of the first audio may be equal to the duration of the first video, in which case it is possible to perfectly match and add separate animation effects based on at least one peak point and the time point corresponding to at least one peak point in the first video.
[0216] The processed first video includes at least one individual animation effect, each individual animation effect can be different, or two individual animation effects can be the same. Exemplarily, peak point 1, peak point 2, and peak point 3 are extracted from the first audio. If each individual animation effect is different, the individual animation effect added corresponding to peak point 1 can be a flash effect, the individual animation effect added corresponding to peak point 2 can be a transition effect, and the individual animation effect added corresponding to peak point 3 can be a subtitle effect. If two individual animation effects are the same in at least one individual animation effect, the individual animation effects added corresponding to peak point 1 and peak point 2 can be a flash effect, and the individual animation effect added corresponding to peak point 3 can be a subtitle effect.
[0217] By obtaining the peak point in the final rhythm information of the first audio, when the first audio is used to produce a video, a separate animation effect can be added at the time point corresponding to the peak point, so that the separate animation effect matches the position with a stronger sense of rhythm in the first audio, thereby enhancing the video effect of the processed first video and meeting the user's auditory and visual perception.
[0218] In some embodiments, the final rhythm information of the first audio includes M final data points, where M is an integer greater than 1. The above method further includes step 250, and step 250 includes at least one step of sub-steps 251-255.
[0219] Sub-step 251, performing beat extraction on the first audio to obtain beat information of the first audio, where the beat information of the first audio includes beat information corresponding to N data points respectively.
[0220] Sub-step 252, obtaining a removal threshold of the mth final data point according to the beat information of the mth final data point and the maximum value and the minimum value of the beat information respectively corresponding to at least one data point, where m is a positive integer.
[0221] According to the time period corresponding to the mth final data point in the first audio and the time periods corresponding to the N data points in the first audio, the beat information of the mth final data point is determined to obtain the removal threshold of the mth final data point. Figure 7 As shown, the removal threshold of the mth final data point can be expressed as: r(m)=thres×(1-br(m)).
[0222] Sub-step 253 , multiplying the removal threshold of the m th final data point by the preset window to obtain the interpolation step of the m th final data point.
[0223] For example, reference may be made to Figure 7 As shown, the interpolation step of the mth final data point can be expressed as: step = window × r (m).
[0224] Sub-step 254, interpolating at least one final data point between at least one peak point according to the interpolation step corresponding to at least one peak point, to obtain at least one final data point after interpolation.
[0225] The interpolation step corresponding to at least one peak point is obtained, and at least one final data point is inserted between the current peak point and the next peak point according to the interpolation step corresponding to the current peak point to obtain at least one final data point after interpolation of the current peak point. After the interpolation processing of the current peak point is completed, the interpolation processing is performed on the next peak point, and after the interpolation processing is performed on the at least one peak point, at least one final data point after interpolation is obtained.
[0226] Exemplarily, if peak point 1, peak point 2 and peak point 3 are extracted from the first audio, the interpolation step corresponding to peak point 1 is 2, the interpolation step corresponding to peak point 2 is 3, and the interpolation step corresponding to peak point 3 is 2, then between peak point 1 and peak point 2, a final data point is inserted every 2 final data points, between peak point 2 and peak point 3, a final data point is inserted every 3 final data points, and after peak point 3, a final data point is inserted every 2 final data points. After performing interpolation processing on peak point 3, at least one final data point after interpolation is obtained.
[0227] Sub-step 255, adding a continuous animation effect to the second video according to the at least one final data point after interpolation to generate a processed second video, wherein the rhythm of the continuous animation effect matches the at least one final data point after interpolation.
[0228] Continuous animation effects refer to animation effects added to the second video, and the rhythm of the continuous animation effects matches at least one final data point after interpolation, that is, the rhythm point of the continuous animation effects matches the time point of at least one final data point after interpolation. Continuous animation effects include but are not limited to transformation effects, animation effects, subtitle effects, sound effects, etc.
[0229] Optionally, the duration of the first audio may be less than the duration of the first video, and the continuous animation effect is added only in the time period corresponding to the first audio in the first video. Optionally, the duration of the first audio may be greater than the duration of the first video, and the continuous animation effect is added only in the time period of the first video. Optionally, the duration of the first audio may be equal to the duration of the first video, and at least one final data point after interpolation may be added to the second video.
[0230] The continuous animation special effect may occupy the final data points after multiple differences in at least one final data point after interpolation, and the continuous animation special effect may be added in any time period of the first video.
[0231] After obtaining the peak points in the final rhythm information of the first audio, some final data points are inserted between the peak points to form continuous final data points, thereby facilitating the addition of continuous animation effects to meet different video production needs of users.
[0232] Figure 8 A schematic diagram of the overall extraction process of audio rhythm is shown. Track separation is performed on the first audio to obtain at least one track of the first audio, and then audio segmentation is performed on each track to obtain audio information in multiple first windows corresponding to each track. The average energy of the audio information in the multiple first windows is normalized to obtain data points corresponding to each track. Smoothing is performed on the data points corresponding to each track in combination with the beat information of the first audio. After obtaining the smoothed data points corresponding to each track, enhancement processing and fusion processing are performed in sequence to obtain fused rhythm information at the track level, and combined with the basic rhythm information of the first audio to obtain the final rhythm information. Combined with the beat information of the first audio, the final rhythm information is applied to extract the peak points in the final rhythm information to obtain data points for adding separate animation effects. Interpolation processing is performed on the peak points to obtain data points for adding continuous animation effects.
[0233] The following are device embodiments of the present application, which can be used to execute the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.
[0234] Please refer to Fig. 9 , which shows a block diagram of an audio rhythm extraction device provided by an embodiment of the present application. The device has the function of implementing the above-mentioned audio rhythm extraction method, and the function can be implemented by hardware, or by hardware executing corresponding software. The device can be the computer device introduced above, or it can be set in a computer device. Fig. 9 As shown, the device 900 may include: an audio acquisition module 910 , a track separation module 920 , a rhythm fusion module 930 and a rhythm determination module 940 .
[0235] The audio acquisition module 910 is used to acquire the first audio.
[0236] The audio track separation module 920 is used to perform audio track separation on the first audio to obtain at least one audio track of the first audio.
[0237] The rhythm fusion module 930 is used to obtain the fused rhythm information of the at least one audio track based on the at least one audio track, and to obtain the basic rhythm information of the first audio based on the first audio. The fused rhythm information is the rhythm information at the track level in the first audio, and the basic rhythm information is the rhythm information at the audio level of the first audio.
[0238] The rhythm determination module 940 is used to obtain final rhythm information of the first audio according to the fused rhythm information and the basic rhythm information.
[0239] In some embodiments, the rhythm fusion module 930 includes:
[0240] A windowing processing unit is used to perform windowing processing on each of the at least one audio track to obtain N data points corresponding to the audio track, wherein the data points are used to indicate energy information within each window of the audio track, and N is an integer greater than 1.
[0241] A smoothing processing unit is used to perform smoothing processing on the N data points corresponding to the audio track to obtain K smoothed data points corresponding to the audio track.
[0242] The enhancement processing unit is used to perform enhancement processing on the K smoothed data points corresponding to the audio track to obtain K enhanced data points corresponding to the audio track.
[0243] The fusion processing unit is used to perform fusion processing on the K enhanced data points corresponding to the at least one audio track to obtain fusion rhythm information of the at least one audio track.
[0244] In some embodiments, the windowing processing unit is used to:
[0245] Performing audio segmentation on the audio track using a first sliding window to obtain audio information in N first windows corresponding to the audio track;
[0246] For each piece of audio information in the N first windows corresponding to the audio track, calculate an average of the sum of squares of the audio amplitudes of at least one sampling point in the first window to obtain average energy of the audio information in the first window;
[0247] Performing normalization processing on average energy of the audio information in the first window to obtain data points of the audio information in the first window;
[0248] According to the data points of the audio information in the N first windows corresponding to the audio track, the N data points corresponding to the audio track are obtained.
[0249] In some embodiments, the smoothing process includes suppression process and removal process, wherein the suppression process is used to suppress low-frequency noise in the audio track, and the removal process is used to remove data points corresponding to low-frequency information in the audio track; the smoothing processing unit is used to:
[0250] Performing the suppression process on the N data points corresponding to the audio track to obtain N suppressed data points corresponding to the audio track;
[0251] The removal process is performed on the N suppressed data points corresponding to the audio track to obtain K smoothed data points corresponding to the audio track.
[0252] In some embodiments, the smoothing processing unit is used to:
[0253] Calculate the difference between the i-th data point and the i-1-th suppressed data point, where i is an integer greater than 1;
[0254] Obtaining a suppression coefficient of the i-th data point according to a difference between the i-th data point and the i-1-th suppression-processed data point, wherein the suppression coefficient of the i-th data point is negatively correlated with the difference between the i-th data point and the i-1-th suppression-processed data point;
[0255] The i-th suppressed data point is obtained according to the suppression coefficient of the i-th data point, the i-th data point and the i-1-th suppressed data point.
[0256] In some embodiments, the smoothing processing unit is used to:
[0257] Performing beat extraction on the first audio to obtain beat information of the first audio, wherein the beat information of the first audio includes beat information corresponding to the N data points respectively, and the beat information corresponding to the data points is used to indicate beat information of a first window corresponding to the data point in the first audio;
[0258] According to the beat information of the i-th data point and the maximum and minimum values of the beat information respectively corresponding to the N data points, a removal threshold of the i-th data point is obtained, wherein the removal threshold of the i-th data point is negatively correlated with the beat information of the i-th data point, and i is an integer greater than 1;
[0259] According to the removal thresholds respectively corresponding to the N data points, removal processing is performed on the N suppressed data points corresponding to the audio track to obtain K smoothed data points corresponding to the audio track.
[0260] In some embodiments, the smoothing processing unit is used to:
[0261] Calculating a first difference and a second difference, wherein the first difference refers to the difference between the i-th suppressed data point and the i-1-th suppressed data point, and the second difference refers to the difference between the i+1-th suppressed data point and the i-th suppressed data point;
[0262] Calculating the product of the first difference and the second difference;
[0263] When the product of the first difference and the second difference is greater than 0, determining the result of removing the i-th suppressed data point is to remove the i-th suppressed data point; and performing the removal process on the i+2-th suppressed data point;
[0264] When the product of the first difference and the second difference is less than or equal to 0, the absolute value of the first difference is greater than or equal to the removal threshold of the i-th data point, and the absolute value of the second difference is less than or equal to the removal threshold of the i-th data point, determining the removal result of the i-th suppressed data point is to retain the i-th suppressed data point; and performing removal processing on the i+2-th suppressed data point;
[0265] When the product of the first difference and the second difference is less than or equal to 0, the absolute value of the first difference is less than or equal to the removal threshold of the i-th data point, and the absolute value of the second difference is less than or equal to the removal threshold of the i-th data point, determining the removal result of the i-th suppressed data point is to remove the i-th suppressed data point; and performing removal processing on the i+2-th suppressed data point;
[0266] When the product of the first difference and the second difference is less than or equal to 0, the absolute value of the first difference is greater than or equal to the removal threshold of the i-th data point, and the absolute value of the second difference is greater than or equal to the removal threshold of the i-th data point, determining the removal result of the i-th suppressed data point is to retain the i-th suppressed data point; and performing removal processing on the i+1-th suppressed data point;
[0267] When the product of the first difference and the second difference is less than or equal to 0, and the absolute value of the first difference is less than the removal threshold of the i-th data point, and the absolute value of the second difference is less than the removal threshold of the i-th data point, the i-th smoothed data point is obtained according to the removal processing results of the i-th suppressed data point and the i-1-th suppressed data point; and the removal processing is performed on the i+2-th suppressed data point;
[0268] According to the removal processing results of the N suppressed data points corresponding to the audio track, K smoothed data points corresponding to the audio track are obtained.
[0269] In some embodiments, the enhancement processing unit is used to:
[0270] Calculate a third difference and a fourth difference, wherein the third difference refers to the difference between the j-th smoothed data point and the j-1-th smoothed data point, and the fourth difference refers to the difference between the j+1-th smoothed data point and the j-th smoothed data point, where j is an integer greater than 1;
[0271] Calculating the product of the third difference and the fourth difference;
[0272] When the j-th smoothed data point satisfies the enhancement condition, performing enhancement processing on the j-th smoothed data point to obtain the j-th enhanced data point;
[0273] When the j-th smoothed data point does not satisfy the enhancement condition, determining the j-th smoothed data point as the j-th enhanced data point;
[0274] Among them, the enhancement condition is that the product of the third difference and the second difference is less than or equal to 0, the absolute value of the third difference is greater than the removal threshold of the i-th data point, and the absolute value of the fourth difference is greater than the removal threshold of the i-th data point.
[0275] In some embodiments, the fusion processing unit is used to:
[0276] For each of the at least one audio track, performing partitioning processing on K enhanced data points corresponding to the audio track to obtain a flat area and a non-flat area corresponding to the audio track, wherein a value change of the data points in the flat area is smaller than a value change of the data points in the non-flat area;
[0277] According to the preset processing priority of the at least one audio track, point filling processing is performed on the flat area of the first audio track to obtain the fused rhythm information of the at least one audio track, wherein the point filling processing is used to enhance the numerical changes of data points in the flat area of the first audio track, and the first audio track is the audio track ranked first in the processing priority among the at least one audio track.
[0278] In some embodiments, the fusion processing unit is used to:
[0279] Using a second sliding window to perform data point segmentation on the K enhanced data points corresponding to the audio track to obtain at least one data point set in a second window, wherein each data point set in the second window includes at least two enhanced data points;
[0280] When the difference between the maximum value and the minimum value in the data point set in the second window is less than or equal to the first value, determining the data point set in the second window as a flat area of the audio track;
[0281] When the difference between the maximum value and the minimum value in the data point set in the second window is greater than the first value, the data point set in the second window is determined as a non-flat area of the audio track.
[0282] In some embodiments, the fusion processing unit is used to:
[0283] According to the processing priority of the at least one audio track, obtaining non-flat areas respectively corresponding to other audio tracks located after the first audio track;
[0284] According to the processing priorities of the other audio tracks, the flat region of the first audio track is replaced by the non-flat region of the other audio tracks with higher processing priorities, so as to obtain the fused flat region of the first audio track;
[0285] The fusion rhythm information of the at least one audio track is obtained according to the fused smooth area of the first audio track and the non-soft area of the first audio track.
[0286] In some embodiments, the rhythm fusion module 930 is used to:
[0287] Performing audio segmentation on the first audio by using a first sliding window to obtain audio information in N third windows corresponding to the first audio, where N is an integer greater than 1;
[0288] For each piece of audio information in the N third windows corresponding to the first audio, calculating an average of the sum of squares of audio amplitudes of at least one sampling point in the third window to obtain average energy of the audio information in the third window;
[0289] Performing normalization processing on the average energy of the audio information in the third window to obtain data points of the audio information in the third window;
[0290] The basic rhythm information of the first audio is obtained according to the data points of the audio information in the N third windows corresponding to the first audio.
[0291] In some embodiments, the apparatus 900 further includes a rhythm application module, wherein the rhythm application module is configured to:
[0292] Extracting at least one peak point from the final rhythm information of the first audio;
[0293] According to the at least one peak point, separate animation effects are respectively added at time points corresponding to the at least one peak point in the first video to generate a processed first video.
[0294] In some embodiments, the final rhythm information of the first audio includes M final data points, where M is an integer greater than 1; and the rhythm application module is used to:
[0295] Performing beat extraction on the first audio to obtain beat information of the first audio, where the beat information of the first audio includes beat information corresponding to the N data points respectively;
[0296] Obtaining a removal threshold of the mth final data point according to the beat information of the mth final data point and the maximum and minimum values of the beat information respectively corresponding to the at least one data point, where m is a positive integer;
[0297] Multiplying the removal threshold of the mth final data point by the preset window to obtain the interpolation step of the mth final data point;
[0298] According to the interpolation steps respectively corresponding to the at least one peak point, at least one final data point is respectively inserted between the at least one peak point to obtain at least one final data point after interpolation;
[0299] According to the at least one final data point after interpolation, a continuous animation effect is added to the second video to generate a processed second video, wherein the rhythm of the continuous animation effect matches the at least one final data point after interpolation.
[0300] It should be noted that the device provided in the above embodiment, when implementing its functions, is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the content structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0301] Please refer to Fig.10, which shows a block diagram of a computer device 1000 provided in one embodiment of the present application. The computer device 1000 may be any electronic device with data calculation, processing and storage functions. The computer device 1000 may be used to implement the audio rhythm extraction method provided in the above embodiment.
[0302] Typically, the computer device 1000 includes a processor 1001 and a memory 1002 .
[0303] The processor 1001 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1001 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), and PLA (Programmable Logic Array). The processor 1001 may also include a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor 1001 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1001 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.
[0304] The memory 1002 may include one or more computer-readable storage media, which may be non-transitory. The memory 1002 may also include a high-speed random access memory and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1002 is used to store a computer program, which is configured to be executed by one or more processors to implement the above-mentioned audio rhythm extraction method.
[0305] Those skilled in the art will understand that Fig.10 The structure shown in the figure does not constitute a limitation on the computer device 1000, and the computer device 1000 may include more or less components than those shown in the figure, or combine some components, or adopt a different component arrangement.
[0306] In an exemplary embodiment, a computer-readable storage medium is also provided, wherein a computer program is stored in the storage medium, and when the computer program is executed by a processor of a computer device, the above-mentioned audio rhythm extraction method is implemented. Optionally, the above-mentioned computer-readable storage medium can be a ROM (Read-Only Memory), a RAM (Random Access Memory), a CD-ROM (Compact Disc Read-Only Memory), a magnetic tape, a floppy disk, an optical data storage device, etc.
[0307] In an exemplary embodiment, a computer program product is also provided, the computer program product includes a computer program, the computer program is stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device performs the above-mentioned audio rhythm extraction method.
[0308] It should be noted that this application can display a prompt interface, pop-up window or output voice prompt information before collecting relevant data of users and during the process of collecting relevant data of users. The prompt interface, pop-up window or voice prompt information is used to prompt the user that relevant data is currently being collected, so that this application only starts to execute the relevant steps of obtaining relevant data of users after obtaining the confirmation operation issued by the user to the prompt interface or pop-up window, otherwise (that is, when the confirmation operation issued by the user to the prompt interface or pop-up window is not obtained), the relevant steps of obtaining relevant data of users are terminated, that is, the relevant data of users are not obtained. In other words, all user data collected by this application are processed strictly in accordance with the requirements of relevant national laws and regulations, and the informed consent or separate consent of the subject of personal information is obtained with the consent and authorization of the user. The subsequent data use and processing behavior is carried out within the scope of authorization of laws and regulations and the subject of personal information, and the collection, use and processing of relevant user data shall comply with the relevant laws, regulations and standards of relevant countries and regions.
[0309] It should be understood that the "multiple" mentioned in this article refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. In addition, the step numbers described in this article only illustrate a possible execution sequence between the steps. In some other embodiments, the above steps may not be executed in the order of the numbers, such as two steps with different numbers are executed at the same time, or two steps with different numbers are executed in the opposite order to the diagram. The embodiments of the present application are not limited to this.
[0310] The above description is only an exemplary embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method for extracting audio rhythm, characterized in that: The method comprises: Get the first audio; Performing audio track separation on the first audio to obtain at least one audio track of the first audio; According to the at least one audio track, fusion rhythm information of the at least one audio track is obtained, and according to the first audio, basic rhythm information of the first audio is obtained, wherein the fusion rhythm information is rhythm information at the audio track level in the first audio, and the basic rhythm information is rhythm information at the audio level of the first audio; Final rhythm information of the first audio is obtained according to the fused rhythm information and the basic rhythm information.
2. The method according to claim 1, characterized in that: The obtaining, according to the at least one audio track, fusion rhythm information of the at least one audio track comprises: For each of the at least one audio track, performing windowing processing on the audio track to obtain N data points corresponding to the audio track, wherein the data points are used to indicate energy information within each window of the audio track, where N is an integer greater than 1; Performing smoothing on the N data points corresponding to the audio track to obtain K smoothed data points corresponding to the audio track, where K is a positive integer; Performing enhancement processing on the K smoothed data points corresponding to the audio track to obtain K enhanced data points corresponding to the audio track; Fusion processing is performed on the K enhanced data points corresponding to the at least one audio track to obtain fused rhythm information of the at least one audio track.
3. The method according to claim 2, characterized in that The performing windowing processing on the audio track to obtain N data points corresponding to the audio track includes: Performing audio segmentation on the audio track using a first sliding window to obtain audio information in N first windows corresponding to the audio track; For each piece of audio information in the N first windows corresponding to the audio track, calculate an average of the sum of squares of the audio amplitudes of at least one sampling point in the first window to obtain average energy of the audio information in the first window; Performing normalization processing on average energy of the audio information in the first window to obtain data points of the audio information in the first window; According to the data points of the audio information in the N first windows corresponding to the audio track, the N data points corresponding to the audio track are obtained.
4. The method according to claim 2, characterized in that: The smoothing process includes a suppression process and a removal process, wherein the suppression process is used to suppress low-frequency noise in the audio track, and the removal process is used to remove data points corresponding to low-frequency information in the audio track; The performing smoothing on the N data points corresponding to the audio track to obtain K smoothed data points corresponding to the audio track includes: Performing the suppression process on the N data points corresponding to the audio track to obtain N suppressed data points corresponding to the audio track; The removal process is performed on the N suppressed data points corresponding to the audio track to obtain K smoothed data points corresponding to the audio track.
5. The method according to claim 4, characterized in that The performing the suppression process on the N data points corresponding to the audio track to obtain the N suppressed data points corresponding to the audio track includes: Calculate the difference between the i-th data point and the i-1-th suppressed data point, where i is an integer greater than 1; Obtaining a suppression coefficient of the i-th data point according to a difference between the i-th data point and the i-1-th suppression-processed data point, wherein the suppression coefficient of the i-th data point is negatively correlated with the difference between the i-th data point and the i-1-th suppression-processed data point; The i-th suppressed data point is obtained according to the suppression coefficient of the i-th data point, the i-th data point and the i-1-th suppressed data point.
6. The method according to claim 4, characterized in that The performing the removal process on the N suppressed data points corresponding to the audio track to obtain K smoothed data points corresponding to the audio track includes: Performing beat extraction on the first audio to obtain beat information of the first audio, wherein the beat information of the first audio includes beat information corresponding to the N data points respectively, and the beat information corresponding to the data points is used to indicate beat information of a first window corresponding to the data point in the first audio; According to the beat information of the i-th data point and the maximum and minimum values of the beat information respectively corresponding to the N data points, a removal threshold of the i-th data point is obtained, wherein the removal threshold of the i-th data point is negatively correlated with the beat information of the i-th data point, and i is an integer greater than 1; According to the removal thresholds respectively corresponding to the N data points, removal processing is performed on the N suppressed data points corresponding to the audio track to obtain K smoothed data points corresponding to the audio track.
7. The method according to claim 6, characterized in that The performing removal processing on the N suppressed data points corresponding to the audio track according to the removal thresholds respectively corresponding to the N data points to obtain K smoothed data points corresponding to the audio track comprises: Calculating a first difference and a second difference, wherein the first difference refers to the difference between the i-th suppressed data point and the i-1-th suppressed data point, and the second difference refers to the difference between the i+1-th suppressed data point and the i-th suppressed data point; Calculating the product of the first difference and the second difference; When the product of the first difference and the second difference is greater than 0, determining the result of removing the i-th suppressed data point is to remove the i-th suppressed data point; and performing the removal process on the i+2-th suppressed data point; When the product of the first difference and the second difference is less than or equal to 0, the absolute value of the first difference is greater than or equal to the removal threshold of the i-th data point, and the absolute value of the second difference is less than or equal to the removal threshold of the i-th data point, determining the removal result of the i-th suppressed data point is to retain the i-th suppressed data point; and performing removal processing on the i+2-th suppressed data point; When the product of the first difference and the second difference is less than or equal to 0, the absolute value of the first difference is less than or equal to the removal threshold of the i-th data point, and the absolute value of the second difference is less than or equal to the removal threshold of the i-th data point, determining the removal result of the i-th suppressed data point is to remove the i-th suppressed data point; and performing removal processing on the i+2-th suppressed data point; When the product of the first difference and the second difference is less than or equal to 0, the absolute value of the first difference is greater than or equal to the removal threshold of the i-th data point, and the absolute value of the second difference is greater than or equal to the removal threshold of the i-th data point, determining the removal result of the i-th suppressed data point is to retain the i-th suppressed data point; and performing removal processing on the i+1-th suppressed data point; When the product of the first difference and the second difference is less than or equal to 0, and the absolute value of the first difference is less than the removal threshold of the i-th data point, and the absolute value of the second difference is less than the removal threshold of the i-th data point, the i-th smoothed data point is obtained according to the removal processing results of the i-th suppressed data point and the i-1-th suppressed data point; and the removal processing is performed on the i+2-th suppressed data point; According to the removal processing results of the N suppressed data points corresponding to the audio track, K smoothed data points corresponding to the audio track are obtained.
8. The method according to claim 6 or 7, characterized in that: The performing enhancement processing on the K smoothed data points corresponding to the audio track to obtain the K enhanced data points corresponding to the audio track includes: Calculate a third difference and a fourth difference, wherein the third difference refers to the difference between the j-th smoothed data point and the j-1-th smoothed data point, and the fourth difference refers to the difference between the j+1-th smoothed data point and the j-th smoothed data point, where j is an integer greater than 1; Calculating the product of the third difference and the fourth difference; When the j-th smoothed data point satisfies the enhancement condition, performing enhancement processing on the j-th smoothed data point to obtain the j-th enhanced data point; When the j-th smoothed data point does not satisfy the enhancement condition, determining the j-th smoothed data point as the j-th enhanced data point; Among them, the enhancement condition is that the product of the third difference and the second difference is less than or equal to 0, the absolute value of the third difference is greater than the removal threshold of the i-th data point, and the absolute value of the fourth difference is greater than the removal threshold of the i-th data point.
9. The method according to claim 2, characterized in that: The performing fusion processing on the K enhanced data points respectively corresponding to the at least one audio track to obtain fused rhythm information of the at least one audio track includes: For each of the at least one audio track, performing partitioning processing on K enhanced data points corresponding to the audio track to obtain a flat area and a non-flat area corresponding to the audio track, wherein a value change of the data points in the flat area is smaller than a value change of the data points in the non-flat area; According to the preset processing priority of the at least one audio track, point filling processing is performed on the flat area of the first audio track to obtain the fused rhythm information of the at least one audio track, wherein the point filling processing is used to enhance the numerical changes of data points in the flat area of the first audio track, and the first audio track is the audio track ranked first in the processing priority among the at least one audio track.
10. The method according to claim 9, characterized in that The performing of partition processing on the K enhanced data points corresponding to the audio track to obtain the flat area and the non-flat area corresponding to the audio track comprises: Using a second sliding window to perform data point segmentation on the K enhanced data points corresponding to the audio track to obtain at least one data point set in a second window, wherein each data point set in the second window includes at least two enhanced data points; When the difference between the maximum value and the minimum value in the data point set in the second window is less than or equal to the first value, determining the data point set in the second window as a flat area of the audio track; When the difference between the maximum value and the minimum value in the data point set in the second window is greater than the first value, the data point set in the second window is determined as a non-flat area of the audio track.
11. The method according to claim 9 or 10, characterized in that: The step of performing fill-in processing on the flat region of the first audio track according to the preset processing priority of the at least one audio track to obtain the fusion rhythm information of the at least one audio track includes: According to the processing priority of the at least one audio track, obtaining non-flat areas respectively corresponding to other audio tracks located after the first audio track; According to the processing priorities of the other audio tracks, the flat region of the first audio track is replaced by the non-flat region of the other audio tracks with higher processing priorities, so as to obtain the fused flat region of the first audio track; The fusion rhythm information of the at least one audio track is obtained according to the fused smooth area of the first audio track and the non-soft area of the first audio track.
12. The method according to claim 1, characterized in that The obtaining basic rhythm information of the first audio according to the first audio includes: Performing audio segmentation on the first audio by using a first sliding window to obtain audio information in N third windows corresponding to the first audio, where N is an integer greater than 1; For each piece of audio information in the N third windows corresponding to the first audio, calculating an average of the sum of squares of audio amplitudes of at least one sampling point in the third window to obtain average energy of the audio information in the third window; Performing normalization processing on the average energy of the audio information in the third window to obtain data points of the audio information in the third window; The basic rhythm information of the first audio is obtained according to the data points of the audio information in the N third windows corresponding to the first audio.
13. The method according to claim 1, characterized in that The method further comprises: Extracting at least one peak point from the final rhythm information of the first audio; According to the at least one peak point, separate animation effects are respectively added at time points corresponding to the at least one peak point in the first video to generate a processed first video.
14. The method according to claim 13, characterized in that The final rhythm information of the first audio includes M final data points, where M is an integer greater than 1; the method further includes: Performing beat extraction on the first audio to obtain beat information of the first audio, where the beat information of the first audio includes beat information corresponding to the N data points respectively; Obtaining a removal threshold of the mth final data point according to the beat information of the mth final data point and the maximum and minimum values of the beat information respectively corresponding to the at least one data point, where m is a positive integer; Multiplying the removal threshold of the mth final data point by the preset window to obtain the interpolation step of the mth final data point; According to the interpolation steps respectively corresponding to the at least one peak point, at least one final data point is respectively inserted between the at least one peak point to obtain at least one final data point after interpolation; According to the at least one final data point after interpolation, a continuous animation effect is added to the second video to generate a processed second video, wherein the rhythm of the continuous animation effect matches the at least one final data point after interpolation.
15. A computer device, characterized in that: The computer device includes a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the audio rhythm extraction method according to any one of claims 1 to 14.
16. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which is loaded and executed by a processor to implement the audio rhythm extraction method according to any one of claims 1 to 14.
17. A computer program product, characterized in that The computer program product comprises a computer program, which is loaded and executed by a processor to implement the audio rhythm extraction method according to any one of claims 1 to 14.