Method, apparatus, device, and storage medium for detecting accompaniment re-sampling
By extracting and comparing the characteristics of the accompaniment audio clips and dry sound audio clips in the song, the accompaniment retrieval detection problem is solved during the recording of songs, the auxiliary judgment of the interface standards of the song recording equipment is realized, and the audio quality of the song recording is improved.
Patent Information
- Application Number
- CN202211185216.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-27
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-09-27
AI Technical Summary
The prior art is difficult to detect whether there is a problem of accompaniment retrieval when recording songs, especially when the mobile phone is connected to wired headphones that are not supported.
By obtaining the accompaniment audio and dry sound audio of the target song, extract the accompaniment audio clip of the preset duration, and obtain the corresponding audio clip to be compared from the dry sound audio. Then, the audio feature similarity between the two is calculated. If the similarity satisfies the preset condition, it is determined that there is accompaniment retrieval.
Effectively detect whether there is accompaniment when recording songs, assist in determining whether the phone supports the interface standards of the headphones used to record songs, and improve the audio quality and reliability during the recording process.
Smart Images

Figure CN115631766B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of audio technology, and particularly to a method, device, equipment and storage medium for detecting accompaniment feedback. Background Art
[0002] Currently, the common connector standards for wired earphones are the OMTP (Open Mobile Terminal Platform) standard and the CTIA (Cellular Telecommunications Industry Association) standard.
[0003] Most mobile phones on the current market can support both the OMTP standard and the CTIA standard. However, some models of mobile phones only support one of the standards. If a mobile phone is connected to a wired earphone with a connector standard that it does not support, there may be a problem of accompaniment feedback when the mobile phone uses the wired earphone to record songs. Based on this, it is possible to assist in determining whether the mobile phone supports the interface standard of the earphone used for recording by detecting whether there is accompaniment feedback during recording.
[0004] Therefore, there is an urgent need for a solution that can detect whether there is accompaniment feedback during recording. Summary of the Invention
[0005] Embodiments of the present application provide a method, device, equipment and storage medium for detecting accompaniment feedback. The technical solution is as follows:
[0006] In a first aspect, a method for detecting accompaniment feedback is provided. The method includes:
[0007] Obtain the accompaniment audio and dry audio of the target song, where the dry audio is the human voice audio collected when playing the accompaniment audio;
[0008] In the accompaniment audio, obtain an accompaniment audio segment of a preset duration, where there are no lyrics during the playing time corresponding to the accompaniment audio segment in the lyrics file of the target song; and in the dry audio, obtain a to-be-compared audio segment corresponding to the playing time of the accompaniment audio segment;
[0009] Obtain a first audio feature of the accompaniment audio segment, and obtain a second audio feature of the to-be-compared audio segment;
[0010] Determine the similarity between the first audio feature and the second audio feature, and determine whether there is accompaniment feedback according to whether the similarity meets a preset similarity condition.
[0011] In a possible implementation, obtaining an accompaniment audio segment of a preset duration from the accompaniment audio includes:
[0012] Obtaining the lyrics file of the target song, where the lyrics file includes the lyrics of the target song, as well as the singing start time and singing end time of each word in the lyrics;
[0013] In the lyrics file, if it is determined that the singing time interval between two adjacent words is greater than the preset duration, then in the accompaniment audio corresponding to the two adjacent words, obtain the accompaniment audio segment of the preset duration, where the singing time interval between the two adjacent words is the time interval from the singing end time of the previous word to the singing start time of the next word among the two adjacent words.
[0014] In a possible implementation, obtaining the first audio feature of the accompaniment audio segment and obtaining the second audio feature of the audio segment to be compared includes:
[0015] Performing frame division processing on the accompaniment audio segment to obtain a plurality of first audio frames corresponding to the accompaniment audio segment;
[0016] Determining the first feature vector of each first audio frame, and combining the first feature vectors of each first audio frame to obtain the first audio feature of the accompaniment audio segment;
[0017] Performing frame division processing on the audio segment to be compared to obtain a plurality of second audio frames corresponding to the audio segment to be compared;
[0018] Determining the second feature vector of each second audio frame, and combining the second feature vectors of each second audio frame to obtain the first audio feature of the accompaniment audio segment.
[0019] In a possible implementation, determining the first feature vector of each first audio frame includes:
[0020] Obtaining the first frequency domain signal corresponding to each first audio frame, and obtaining the energy of each first audio frame in N frequency bands according to each first frequency domain signal, where N is a preset positive integer;
[0021] For each first audio frame, determining the first feature vector of the first audio frame according to the energy of the first audio frame in the N frequency bands and the energy of the previous first audio frame of the first audio frame in the N frequency bands;
[0022] Determining the second feature vector of each second audio frame includes:
[0023] Obtain the second frequency-domain signal corresponding to each second audio frame, and based on each second frequency-domain signal, obtain the energy of each second audio frame in the N frequency bands;
[0024] For each second audio frame, determine the second feature vector of the second audio frame according to the energy of the second audio frame in the N frequency bands and the energy of the previous second audio frame of the second audio frame in the N frequency bands.
[0025] In a possible implementation manner, the determining the first feature vector of the first audio frame according to the energy of the first audio frame in the N frequency bands and the energy of the previous first audio frame of the first audio frame in the N frequency bands includes:
[0026] Determine the value of the nth element in the first feature vector of the first audio frame according to the energy of the first audio frame in the nth frequency band among the N frequency bands, the energy of the first audio frame in the (n + 1)th frequency band among the N frequency bands, the energy of the previous first audio frame of the first audio frame in the nth frequency band, and the energy of the previous first audio frame in the (n + 1)th frequency band, where 1 ≤ n ≤ N - 1;
[0027] Obtain the first feature vector according to the value of each element in the first feature vector of the first audio frame;
[0028] Determine the value of the nth element in the second feature vector of the second audio frame according to the energy of the second audio frame in the nth frequency band among the N frequency bands, the energy of the second audio frame in the (n + 1)th frequency band among the N frequency bands, the energy of the previous second audio frame of the second audio frame in the nth frequency band, and the energy of the previous second audio frame in the (n + 1)th frequency band;
[0029] Obtain the second feature vector according to the value of each element in the second feature vector of the second audio frame.
[0030] In a possible implementation manner, the determining the value of the nth element in the first feature vector of the first audio frame according to the energy of the first audio frame in the nth frequency band among the N frequency bands, the energy of the first audio frame in the (n + 1)th frequency band among the N frequency bands, the energy of the previous first audio frame of the first audio frame in the nth frequency band, and the energy of the previous first audio frame in the (n + 1)th frequency band includes:
[0031] Calculate the first difference between the energy of the first audio frame in the nth frequency band among the N frequency bands and the energy of the first audio frame in the (n + 1)th frequency band among the N frequency bands;
[0032] Calculate a second difference between the energy of the n-th frequency band and the energy of the (n + 1)-th frequency band in the N frequency bands of the previous first audio frame of the first audio frame;
[0033] If the first difference is greater than the second difference, set the value of the n-th element in the first feature vector of the first audio frame to 1;
[0034] If the first difference is not greater than the second difference, set the value of the n-th element in the first feature vector of the first audio frame to 0;
[0035] Determining the value of the n-th element in the second feature vector of the second audio frame based on the energy of the n-th frequency band in the N frequency bands of the second audio frame, the energy of the (n + 1)-th frequency band in the N frequency bands of the second audio frame, the energy of the n-th frequency band of the previous second audio frame of the second audio frame, and the energy of the (n + 1)-th frequency band of the previous second audio frame, includes:
[0036] Calculate a third difference between the energy of the n-th frequency band and the energy of the (n + 1)-th frequency band in the N frequency bands of the second audio frame;
[0037] Calculate a fourth difference between the energy of the n-th frequency band and the energy of the (n + 1)-th frequency band in the N frequency bands of the previous second audio frame of the second audio frame;
[0038] If the third difference is greater than the fourth difference, set the value of the n-th element in the second feature vector of the second audio frame to 1;
[0039] If the third difference is not greater than the fourth difference, set the value of the n-th element in the second feature vector of the second audio frame to 0.
[0040] In a possible implementation, determining the similarity between the first audio feature and the second audio feature includes:
[0041] Starting from the first first feature vector in the first audio feature, select first feature vectors as first starting feature vectors in sequence according to a first step size until the M-th first feature vector is selected, and then stop selecting the first starting feature vectors, where the first step size is the number of first feature vectors offset each time, and M is a preset positive integer;
[0042] For each selected first starting feature vector, in the first audio feature, the current selected first starting feature vector to the last first feature vector are used as the first comparison audio feature of the accompaniment audio segment, and in the second audio feature, a plurality of second feature vectors are continuously selected starting from the first second feature vector as the second comparison audio feature of the audio segment to be compared, where the number of first feature vectors included in the first comparison audio feature is the same as the number of second feature vectors included in the second comparison audio feature;
[0043] Starting from the first second feature vector in the second audio feature, according to the second step size, second feature vectors are sequentially selected as the second starting feature vectors until the Mth second feature vector is selected, and then the selection of the second starting feature vectors stops, where the second step size is the number of second feature vectors offset each time;
[0044] For each selected second starting feature vector, in the second audio feature, the current selected second starting feature vector to the last second feature vector are used as the second comparison audio feature of the audio segment to be compared, and in the first audio feature, a plurality of first feature vectors are continuously selected starting from the first first feature vector as the first comparison audio feature of the accompaniment audio segment;
[0045] For the first comparison audio feature and the second comparison audio feature obtained within the same loop, calculate the number of elements with the same alignment, and take the ratio of the number of elements with the same alignment to the total number of elements of the first comparison audio feature as the reference similarity;
[0046] Select the maximum reference similarity as the similarity between the first audio feature and the second audio feature.
[0047] In a possible implementation manner, the determining whether there is accompaniment resampling according to whether the similarity meets a preset similarity condition includes:
[0048] If the similarity between the first audio feature and the second audio feature is greater than a preset similarity threshold, and the range of the reference similarity is greater than a preset range threshold, it is determined that there is accompaniment resampling;
[0049] If the similarity between the first audio feature and the second audio feature is not greater than the preset similarity threshold, and / or the range of the reference similarity is not greater than the preset range threshold, it is determined that there is no accompaniment resampling.
[0050] In a possible implementation manner, after determining that there is accompaniment resampling, the method further includes:
[0051] If the dry audio is collected through a wired headset, a prompt message is displayed, where the prompt message is used to prompt the possibility of non - support for the wired headset connector standard.
[0052] In a second aspect, a device for detecting accompaniment feedback is provided. The device includes:
[0053] An acquisition module, configured to acquire the accompaniment audio and dry audio of a target song, where the dry audio is the human voice audio collected when playing the accompaniment audio; in the accompaniment audio, acquire an accompaniment audio segment of a preset duration, where there are no lyrics during the playing time corresponding to the accompaniment audio segment in the lyrics file of the target song; and in the dry audio, acquire a to - be - compared audio segment corresponding to the playing time of the accompaniment audio segment; acquire a first audio feature of the accompaniment audio segment, and acquire a second audio feature of the to - be - compared audio segment;
[0054] A judgment module, configured to determine the similarity between the first audio feature and the second audio feature, and determine whether there is accompaniment feedback according to whether the similarity meets a preset similarity condition.
[0055] In a third aspect, a terminal is provided. The terminal includes a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the method for detecting accompaniment feedback as described in the first aspect above.
[0056] In a fourth aspect, a computer - readable storage medium is provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by the processor to implement the method for detecting accompaniment feedback as described in the first aspect above.
[0057] In a fifth aspect, a computer program product is provided. At least one instruction is included in the computer program product, and the at least one instruction is loaded and executed by the processor to implement the method for detecting accompaniment feedback as described in the first aspect above.
[0058] The beneficial effects brought by the technical solutions provided in the embodiments of the present application at least include:
[0059] In the embodiments of the present application, first obtain the accompaniment audio and dry voice audio of the target song, and then obtain an accompaniment audio segment of a preset duration from the accompaniment audio. There are no lyrics during the playing time corresponding to the accompaniment audio segment in the lyrics file of the target song. Then, obtain a to-be-compared audio segment corresponding to the playing time of the accompaniment audio segment from the dry voice audio, and then determine the similarity between the first audio feature of the accompaniment audio segment and the second audio feature of the to-be-compared audio segment. If the similarity meets the preset similarity condition, it indicates that the to-be-compared audio segment is mixed with the recollected accompaniment audio, thereby detecting the problem of accompaniment recollection in the target song. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0061] Figure 1 is a flowchart of a method for detecting accompaniment recollection provided by an embodiment of the present application;
[0062] Figure 2 is a schematic diagram of offset comparison provided by an embodiment of the present application;
[0063] Figure 3 is a schematic structural diagram of a device for detecting accompaniment recollection provided by an embodiment of the present application;
[0064] Figure 4 is a schematic structural diagram of a terminal provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0065] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the drawings.
[0066] The embodiment of the present application provides a method for detecting the recapture of accompaniment, which can be implemented by a terminal. Among them, the terminal can be a mobile phone, a tablet computer, etc. The method can be applied in a song recording scenario. In this scenario, the terminal can be connected to a wired headset, and a song recording application can be installed in the terminal. When the user wants to record a song, he can select the song recording application and enter the song selection interface. Then, the user can select the song he wants to record. The terminal obtains the accompaniment audio of the song from the cloud or locally, and starts playing the accompaniment audio when the user clicks the start option. The user can sing according to the accompaniment audio, and the wired headset connected to the terminal can collect the dry sound audio when the user sings and transmit it to the terminal. After the user finishes singing, the terminal can obtain the dry sound audio and the accompaniment audio, and use the method for checking the recapture of accompaniment provided in the embodiment of the present application to detect whether there is a problem of recapture of accompaniment in this song recording.
[0067] It should be noted that the above is only to facilitate the understanding of the present application, and an exemplary possible application scenario is provided. The method provided in the embodiment of the present application is not limited to the configuration in the above application scenario, nor is it limited to the above application scenario. For example, the device for collecting dry sound audio in the above application scenario is an audio headset, which can also be a wireless headset, other external audio receiving devices, a microphone built into a mobile phone, etc. For another example, the above application scenario is a song recording scenario, and it can also be a karaoke scenario, etc.
[0068] The following is a brief introduction to the method for detecting accompaniment recovery provided in an embodiment of the present application.
[0069] In this method, the accompaniment audio and dry audio of the target song are obtained. In the accompaniment audio, an accompaniment audio clip of preset duration is obtained, wherein in the lyrics file of the target song, there are no lyrics in the target singing time corresponding to the audio clip. Then, in the dry audio, an audio clip to be compared corresponding to the target singing time is obtained. The first audio feature of the accompaniment audio clip is obtained, and the second audio feature of the audio clip to be compared is obtained. Finally, the similarity between the first audio feature and the second audio feature is determined. If the similarity meets the preset similarity condition, it means that the audio clip to be compared is mixed with the accompaniment audio that is picked up, resulting in a high similarity between the audio clip to be compared and the accompaniment audio clip.
[0070] The following describes the method for detecting accompaniment replay provided by the embodiment of the present application in conjunction with the accompanying drawings. The method can be implemented by a terminal. Figure 1 As shown, the processing flow of the method may include the following steps:
[0071] Step 101: Obtain the accompaniment audio and dry audio of the target song.
[0072] Among them, the ways for the terminal to obtain the accompaniment audio and the dry audio can be different. For example, the accompaniment audio can be obtained by the terminal from local or other devices, and the dry audio can be collected by the terminal from the external environment through a pickup module. Specifically, for example, the dry audio is the human voice audio collected from the external environment by the terminal when playing the accompaniment audio. The durations of the dry audio and the accompaniment audio are the same, that is, the terminal obtains the dry audio when the accompaniment audio starts to play, and stops obtaining the dry audio when the accompaniment audio finishes playing.
[0073] In implementation, when a user wants to record a song (or sing karaoke), the user can open the song recording application (or karaoke application) installed in the terminal. The user can select the song for which they want to record a song (or sing karaoke). The terminal obtains the accompaniment audio of the song from the cloud or locally. Among them, the cloud can be the background server of the song recording application. When the user clicks the start option, the terminal starts playing the accompaniment audio. The user can sing according to the accompaniment audio, and the wired earphone connected to the terminal can collect the dry audio of the user's singing and transmit it to the terminal. After the user finishes recording and singing, the terminal obtains the accompaniment audio of the target song for this time and the collected dry audio.
[0074] Step 102: In the accompaniment audio, obtain an accompaniment audio segment of a preset duration, where, in the lyrics file of the target song, there are no lyrics during the playing time of the accompaniment audio segment.
[0075] In implementation, the terminal obtains the lyrics file of the target song. Among them, the lyrics file of the target song includes the lyrics of the target song, as well as the singing start time and singing end time of each word in the lyrics.
[0076] The terminal calculates the singing time interval between two adjacent words in the lyrics file. If it is determined that the singing time interval between the first lyric and the second lyric is greater than the preset duration, then in the accompaniment audio corresponding to the singing time between the first lyric and the second lyric, obtain the accompaniment audio segment of the preset duration. Among them, the first lyric and the second lyric are two adjacent words in the lyric text, and the singing start time of the first lyric is earlier than the singing start time of the second lyric, and the singing time between the first lyric and the second lyric is the singing time from the singing end time of the first lyric to the singing start time of the second lyric.
[0077] For example, the preset duration is 8s. In the lyrics file, the singing end time of the first lyric is the 8th second (i.e., the singing of the first lyric ends at the 8th second), and the singing start time of the second lyric is the 18th second (i.e., the singing of the second lyric starts at the 18th second). Then, the singing time interval between the first lyric and the second lyric is 10s. Then, the accompaniment audio segment of the middle 8s of this 10s can be obtained in the accompaniment audio, that is, the accompaniment audio segment from the 9th to the 16th second (a total of 8s) in the accompaniment audio is obtained. The performance time of this accompaniment audio segment does not have lyrics in the lyrics file. Here, the purpose of obtaining the middle 8s is to try to avoid the problem that the accompaniment audio segment is not a pure accompaniment audio due to the inaccurate marking of the singing time of the lyrics in the lyrics file.
[0078] Step 103: Obtain, in the dry audio, the audio segment to be compared corresponding to the playing time of the accompaniment audio segment.
[0079] In implementation, the dry audio and the accompaniment audio have the same duration. After obtaining the accompaniment audio clip, the playing time of the accompaniment audio clip in the accompaniment audio is used as the target singing time, and the audio clip to be compared corresponding to the target singing time is obtained in the dry audio.
[0080] For example, in combination with the example in step 102 above, the performance time of the acquired accompaniment audio clip is from 9s to 16s, then 9s to 16s is taken as the target singing time, and in the dry audio, the audio clip from 9s to 17s is obtained as the audio clip to be compared.
[0081] In a possible implementation, after a large number of experimental analyses, it was found that there was a loss in the recovered signal, but there was an obvious residual in the mid- and low-frequency recovered signal. Based on this, in order to improve the processing efficiency in subsequent comparisons, after obtaining the accompaniment audio segment and the audio segment to be compared, the accompaniment audio segment can be downsampled, for example, only the part of the accompaniment audio segment with a frequency below 4khz is retained. Similarly, the audio segment to be compared can also be downsampled, for example, only the part of the audio segment to be compared with a frequency below 4khz is retained.
[0082] Step 104: Obtain a first audio feature of the accompaniment audio segment, and obtain a second audio feature of the audio segment to be compared.
[0083] It should be noted that the first audio feature is the audio feature of the accompaniment audio segment, and the second audio feature is the audio feature of the audio segment to be compared. The two audio features are respectively referred to as the first audio feature and the second audio feature only for the purpose of distinction. Additionally, if downsampling is performed on the accompaniment audio segment and the audio segment to be compared before step 104, then, in step 104 and subsequent steps, the accompaniment audio segment and the audio segment to be compared mentioned are both the downsampled accompaniment audio segment and the audio segment to be compared.
[0084] To obtain the audio feature of the accompaniment audio segment, first, frame division processing is performed on the accompaniment audio segment to obtain multiple audio frames corresponding to the accompaniment audio segment. To distinguish from the audio frames of the audio segment to be compared, the audio frames of the accompaniment audio segment can be referred to as the first audio frames. Then, the first frequency domain signal corresponding to each first audio frame is obtained.
[0085] The following explains the acquisition of the first frequency domain signal.
[0086] First, windowing processing is performed on the first audio frame. The signal obtained after windowing can be expressed as follows:
[0087] x 1n w n (Ln + i) = x 1n (i)·w(i)
[0088]
[0089] Among them, x 1n w n (Ln + i) represents the signal obtained after windowing the nth first audio frame. L is the frame shift during frame division, i is the index value starting from 0 for N sampling points during windowing, and N is the window length. For example, N is taken as N = 512. x 1n (i) is the time domain signal of the sampling point with index value i in the nth first audio frame before windowing.
[0090] Then, Fourier transform is performed on the signal obtained after windowing. The Fourier transform can be expressed as follows:
[0091]
[0092] Among them, X 1 (n, k) is the frequency domain signal obtained after Fourier transform corresponding to the nth first audio frame (referred to as the first frequency domain signal for easy distinction). k represents the kth frequency point, and j is the imaginary unit.
[0093] After obtaining the first frequency domain signal corresponding to each first audio frame, the energy of each first frequency domain signal is calculated. The energy corresponding to the first frequency domain signal can be expressed as follows:
[0094] P 1 (n, k) = |X 1 (n, k)| 2
[0095] Among them, P 1 (n, k) represents the energy corresponding to the first frequency-domain signal.
[0096] Then, input the first frequency-domain signal into bark filters (bark filter bank) to map the energy corresponding to the first frequency-domain signal into the bark domain, obtaining an A-dimensional energy vector, which is used to describe the energy of the first audio frame in A frequency bands. Among them, according to the setting of the bark domain, A can take the value of A = 33. The energy vector corresponding to the nth first audio frame can be expressed as follows:
[0097] (E 1 (n, 1), E 1 (n, 2), …, E 1 (n, A))
[0098] Among them, E 1 (n, 1) represents the energy of the nth first audio frame in the first frequency band, and E 1 (n, 2) represents the energy of the nth first audio frame in the second frequency band, and so on.
[0099] For each first audio frame, perform differential processing on the energy vector corresponding to the first audio frame and the energy vector corresponding to the previous first audio frame of the first audio frame to determine the first feature vector of the first audio frame. Among them, the first feature vector is an A - 1 dimensional vector.
[0100] The following is an explanation of this differential processing.
[0101] For each first audio frame, calculate the first difference between the energy of the a-th frequency band in the A frequency bands of the first audio frame and the energy of the (a + 1)-th frequency band in the A frequency bands, where 1 ≤ a ≤ A - 1. Calculate the second difference between the energy of the a-th frequency band in the A frequency bands of the previous first audio frame of the first audio frame and the energy of the (a + 1)-th frequency band in the A frequency bands. If the first difference is greater than the second difference, then set the value of the a-th element in the first feature vector of the first audio frame to 1. If the first difference is not greater than the second difference, then set the value of the a-th element in the first feature vector of the first audio frame to 0. According to the value of each element in the first feature vector of the first audio frame, the first feature vector is obtained.
[0102] The above differential processing can be expressed by the following formula.
[0103]
[0104] Among them, F 1 (n, m) is the m-th element in the first feature vector of the n-th first audio frame, and if (if) represents the condition for taking values.
[0105] According to the sorting of each first audio frame in the accompaniment audio segment, the first feature vectors of each first audio frame are combined to obtain the first audio feature of the accompaniment audio segment. Among them, the first audio feature includes the first feature vector of each first audio frame, and the sorting of each first feature vector in the first audio feature is the same as the sorting of the first audio frame to which it belongs in the accompaniment audio segment.
[0106] Similarly, the second audio feature of the audio segment to be compared also needs to be obtained.
[0107] The audio segment to be compared is frame-divided to obtain a plurality of second audio frames corresponding to the audio segment to be compared. Then, the second frequency domain signal corresponding to each second audio frame is obtained.
[0108] The acquisition of the second frequency domain signal is described below.
[0109] First, windowing processing is performed on the second audio frame. The signal obtained after windowing can be expressed as follows:
[0110] x 2n w n (Ln + i) = x 2n (i)·w(i)
[0111]
[0112] Among them, x 2n w n (Ln + i) represents the signal obtained after windowing the n-th second audio frame. L is the frame shift during frame division, i is the index value starting from 0 for N sampling points during windowing, and N is the window length. For example, the value is N = 512, and x 2n (i) is the time domain signal of the sampling point with index value i in the n-th second audio frame before windowing.
[0113] Then, Fourier transform is performed on the signal obtained after windowing. The Fourier transform can be expressed as follows:
[0114]
[0115] Among them, X 2 (n, k) is the second frequency domain signal obtained after Fourier transform corresponding to the n-th second audio frame, k represents the k-th frequency point, and j is the imaginary unit.
[0116] After obtaining the second frequency domain signals corresponding to each second audio frame, calculate the energy for each obtained second frequency domain signal. The energy corresponding to the second frequency domain signal can be expressed as follows:
[0117] P 2 (n, k) = |X 2 (n, k)| 2
[0118] Among them, P 2 (n, k) represents the energy corresponding to the second frequency domain signal.
[0119] Then, input the second frequency domain signal into bark filters to map the energy corresponding to the second frequency domain signal into the bark domain, obtaining an A-dimensional energy vector, which is used to describe the energy of the second audio frame in A frequency bands. Among them, according to the setting of the bark domain, A can take the value of A = 33. The energy vector corresponding to the nth second audio frame can be expressed as follows:
[0120] (E 2 (n, 1), E 2 (n, 2), …, E 2 (n, A))
[0121] Among them, E 2 (n, 1) represents the energy of the nth second audio frame in the first frequency band, E 2 (n, 2) represents the energy of the nth second audio frame in the second frequency band, and so on.
[0122] For each second audio frame, perform a difference process on the energy vector corresponding to the second audio frame and the energy vector corresponding to the previous second audio frame of the second audio frame to determine the second feature vector of the second audio frame. Among them, the second feature vector is an A - 1 dimensional vector.
[0123] The following explains this difference process.
[0124] For each second audio frame, calculate the first difference between the energy of the second audio frame in the ath frequency band among the A frequency bands and the energy of the second audio frame in the (a + 1)th frequency band among the A frequency bands, where 1 ≤ a ≤ A - 1. Calculate the second difference between the energy of the previous second audio frame of the second audio frame in the ath frequency band among the A frequency bands and the energy of the previous second audio frame of the second audio frame in the (a + 1)th frequency band among the A frequency bands. If the first difference is greater than the second difference, then set the ath element in the second feature vector of the second audio frame to 1. If the first difference is not greater than the second difference, then set the ath element in the second feature vector of the second audio frame to 0. Obtain the second feature vector according to the value of each element in the second feature vector of the second audio frame.
[0125] The above differential processing can be expressed by the following formula.
[0126]
[0127] Among them, F 2 (n, m) is the m-th element in the second feature vector of the n-th second audio frame, and if represents the condition for taking values.
[0128] According to the sorting of each second audio frame in the audio segment to be compared, the second feature vectors of each second audio frame are combined to obtain the first audio feature of the audio segment to be compared. Among them, the second audio feature includes the second feature vectors of each second audio frame, and the sorting of each second feature vector in the second audio feature is the same as the sorting of the corresponding second audio frame in the audio segment to be compared.
[0129] Step 105: Determine the similarity between the first audio feature and the second audio feature, and determine whether there is an accompaniment re-acquisition according to whether the similarity meets the preset similarity condition.
[0130] In implementation, there are various methods to determine the similarity between the first audio feature and the second audio feature. Several of them are listed below for illustration:
[0131] Method 1: Calculate the Euclidean distance, cosine distance, etc. between the first audio feature and the second audio feature as their similarity.
[0132] Method 2: For each element in the first audio feature, compare whether it is the same as the element in the same position in the second audio feature. If they are the same, add 1 to the number of identical elements (the initial value of the number of identical elements is 0). Divide the number of identical elements by the total number of elements in the first audio feature to obtain the similarity between the first audio feature and the second audio feature.
[0133] Method 3: Considering that the collected dry audio needs to go through loop processing, thus, due to a certain time delay in the loop processing, there may be a time deviation between the dry audio and the accompaniment audio. To compensate for the influence of this deviation on the calculation of similarity, combined with Method 2 above, the first audio feature and the second audio feature can be offset and compared multiple times at a certain step size. Each offset comparison can obtain a calculated reference similarity. Finally, select the maximum reference similarity as the similarity between the first audio feature and the second audio feature.
[0134] There are many ways of offset comparison. One of them is listed below for illustration.
[0135] In this way, it can be divided into two rounds of loops. In the first round of loop, the first audio feature is offset first, and in the second round of loop, the second audio feature is offset.
[0136] In the first round of loop, offset the first audio feature:
[0137] Starting from the first first feature vector in the first audio feature, select the first feature vector as the first starting feature vector in sequence according to the first step length until the Mth first feature vector is selected, and then stop selecting the first starting feature vector. Here, the first step length is the number of first feature vectors offset each time, and M is a preset positive integer, and the value of M can be set by technicians according to the loop processing time delay. Each time a first starting feature vector is selected, in the first audio feature, the currently selected first starting feature vector to the last first feature vector are used as the comparison audio feature of the accompaniment audio segment (for easy distinction, the comparison audio feature obtained from the first audio feature can be called the first comparison audio feature), and in the second audio feature, starting from the first second feature vector, continuously select multiple second feature vectors as the comparison audio feature of the audio segment to be compared (for easy distinction, the comparison audio feature obtained from the audio segment to be compared can be called the second comparison audio feature). Among them, the number of first feature vectors included in the first comparison audio feature obtained within the same loop is the same as the number of second feature vectors included in the second comparison audio feature.
[0138] For the convenience of description, denote the first step length as a. Then, in the first round of loop, M / a loops are performed. In each loop, a first comparison audio feature and a second comparison audio feature can be obtained. For a first comparison audio feature and a second comparison audio feature obtained in the same loop, they can be called a group of first comparison audio feature and second comparison audio feature. Then, M / a groups of first comparison audio feature and second comparison audio feature can be obtained in M / a loops.
[0139] In the second round of loop, offset the second audio feature:
[0140] Starting from the first second feature vector in the second audio feature, select the second feature vector as the second starting feature vector in sequence according to the second step length until the Mth second feature vector is selected, and then stop selecting the second starting feature vector. Here, the second step length is the number of second feature vectors offset each time. The second step length can be the same as the first step length. Each time a second starting feature vector is selected, in the second audio feature, the currently selected second starting feature vector to the last second feature vector are used as the comparison audio feature of the audio segment to be compared (for easy distinction, the comparison audio feature obtained from the audio segment to be compared can be called the second comparison audio feature), and in the first audio feature, starting from the first first feature vector, continuously select multiple first feature vectors as the comparison audio feature of the accompaniment audio segment (for easy distinction, the comparison audio feature obtained from the first audio feature can be called the first comparison audio feature).
[0141] For ease of description, denote the first step length as b. Then, in the first round of loop, M / b loops are performed. In each loop, a first comparison audio feature and a second comparison audio feature can be obtained. For a first comparison audio feature and a second comparison audio feature obtained in the same loop, they can be called a set of first comparison audio feature and second comparison audio feature. Then, M / b sets of first comparison audio feature and second comparison audio feature can be obtained in M / b loops.
[0142] In this way, after the above two rounds of loops, (M / a + M / b) sets of first comparison audio feature and second comparison audio feature can be obtained. For each set of first comparison audio feature and second comparison audio feature, the reference similarity between the first comparison audio feature and the second comparison audio feature in this set can be calculated. In this way, (M / a + M / b) reference similarities can be obtained. Then, among these (M / a + M / b) reference similarities, select the maximum reference similarity as the similarity between the first audio feature and the second audio feature.
[0143] The method for calculating the reference similarity can be as follows:
[0144] For each set of first comparison audio feature and second comparison audio feature, calculate the number of elements with the same position. Take the ratio of the number of elements with the same position to the total number of elements of the first comparison audio feature as the reference similarity. Select the maximum reference similarity as the similarity between the first audio feature and the second audio feature. The following explains how to calculate the number of elements with the same position:
[0145] For each element in the first comparison audio feature, compare whether the element is the same as the element at the same position in the second comparison audio feature. If they are the same, increment the number of elements with the same position by 1 (the initial value of the number of elements with the same position is 0). Divide the number of elements with the same position by the total number of elements of the first comparison audio feature to obtain the similarity between the first comparison audio feature and the second comparison audio feature.
[0146] The following combines Figure 2 to give an example of the two rounds of loops in the above offset comparison.
[0147] As Figure 2 shown, the first audio feature consists of y first feature vectors, and the y first feature vectors are respectively denoted as: S 1 、S 2 、…、S y , and the second audio feature consists of y second feature vectors, and the y second feature vectors are respectively denoted as: D 1 、D 2 、…、D yAssume that both the first step length and the second step length are 1, that is, each time it offsets by 1 eigenvector, and M takes the value of 100.
[0148] First, perform the first round of loop.
[0149] The first loop: Take S in the first audio feature 1 as the first starting eigenvector, take S 1 to S y as the first comparison audio feature, and take D in the second audio feature 1 to D y as the second comparison audio feature. Determine that S 1 is not the 100th first audio feature, and continue to execute.
[0150] The second loop: Offset by 1 first eigenvector, take S in the first audio feature 2 as the first starting eigenvector, take S 2 to S y as the first comparison audio feature, and take D in the second audio feature 1 to D y-1 as the second comparison audio feature. Determine that S 2 is not the 100th first audio feature, and continue to execute.
[0151] The third loop: Offset by 2 first eigenvectors, take S in the first audio feature 3 as the first starting eigenvector, take S 3 to S y as the first comparison audio feature, and take D in the second audio feature 1 to D y-2 as the second comparison audio feature. Determine that S 3 is not the 100th first audio feature, and continue to execute.
[0152] And so on, until the 100th loop, offset by 100 first eigenvectors, take S in the first audio feature 100 as the first starting eigenvector, take S 100 to S y as the first comparison audio feature, and take D in the second audio feature 1 to D y-99 as the second comparison audio feature. Determine that S 100 is the 100th first audio feature, and stop the loop.
[0153] Next, perform the second round of loop.
[0154] The first loop: Take D in the second audio feature 1 as the second starting eigenvector, take D 1 to Dy As the second comparison audio feature, and take S in the first audio feature 1 to S y as the first comparison audio feature. Determine that D 1 is not the 100th second audio feature, and continue to execute.
[0155] Second loop: Offset by 1 second feature vector, and take D in the second audio feature 2 as the second starting feature vector, and take D 2 to D y as the second comparison audio feature, and take S in the first audio feature 1 to S y-1 as the first comparison audio feature. Determine that D 2 is not the 100th second audio feature, and continue to execute.
[0156] Third loop: Offset by 2 second feature vectors, and take D in the second audio feature 3 as the second starting feature vector, and take D 3 to D y as the second comparison audio feature, and take S in the first audio feature 1 to S y-2 as the first comparison audio feature. Determine that D 3 is not the 100th second audio feature, and continue to execute.
[0157] And so on, until the 100th loop, offset by 100 second feature vectors, and take D in the second audio feature 100 as the second starting feature vector, and take D 100 to D y as the second comparison audio feature, and take S in the first audio feature 1 to S y-99 as the first comparison audio feature. Determine that D 100 is the 100th second audio feature, and stop the loop.
[0158] During the above loop process, within each loop, the reference similarity between the first comparison audio feature and the second comparison audio feature obtained in this loop can be calculated.
[0159] In addition, it should also be noted that the above second loop can start not from the first second feature vector in the second audio feature, but from the second second feature vector in the second audio feature. Additionally, the loop logic above is only an example, and there can be multiple loop logics that can be adopted during specific implementation. For example, it can also start offsetting from the end of the first audio feature or the second audio feature. The specific offset method is not limited in the embodiments of this application, as long as the purpose of the final offset comparison can be achieved.
[0160] After calculating the similarity between the first audio feature and the second audio feature, it is determined whether the similarity is greater than a preset similarity threshold. If it is greater, it can be determined that there is an accompaniment re-recording in the terminal for collecting the dry audio.
[0161] In a possible implementation, when calculating the similarity using the above method three, if it is determined that the similarity between the first audio feature and the second audio feature is greater than the preset similarity threshold, it is further determined whether the range difference of each reference similarity is greater than the preset range difference threshold. If the range difference of each reference similarity is greater than the preset range difference threshold, it is further determined that there is an accompaniment re-recording in the terminal for collecting the dry audio.
[0162] In a possible implementation, when it is determined that there is an accompaniment re-recording in the terminal for collecting the dry audio, if the dry audio is collected through a wired headset, the terminal can display a target prompt message, where the target prompt message is used to prompt the possibility that this terminal does not support the connector standard of the wired headset.
[0163] In yet another possible implementation, in a KTV singing scenario, when it is determined that there is an accompaniment re-recording in the terminal for collecting the dry audio, a notification can be sent to the scoring module of the KTV singing application. The notification is used to inform the scoring module that there is a re-recorded accompaniment in the collected dry audio. The specific notification method is not limited in the embodiments of the present application. Considering that when there is a re-recorded accompaniment in the dry audio, the score given by the scoring module for the user's KTV singing is relatively high. In the present application, after receiving the notification, the scoring module subtracts a preset score from the score of the user's current KTV singing to obtain the final score of the user's current KTV singing.
[0164] Based on the same technical concept, the embodiments of the present application further provide a device for detecting accompaniment re-recording. The device can be the terminal in the above embodiments, such as Figure 3 As shown, the device includes: an acquisition module 310 and a judgment module 320.
[0165] The acquisition module 310 is configured to acquire the accompaniment audio and the dry audio of the target song, where the dry audio is the human voice audio collected when playing the accompaniment audio; in the accompaniment audio, acquire an accompaniment audio segment of a preset duration, where the accompaniment audio segment has no lyrics during the playing time corresponding to it in the lyrics file of the target song; and in the dry audio, acquire a to-be-compared audio segment corresponding to the playing time of the accompaniment audio segment; acquire the first audio feature of the accompaniment audio segment, and acquire the second audio feature of the to-be-compared audio segment;
[0166] The judgment module 320 is configured to determine the similarity between the first audio feature and the second audio feature, and determine whether there is an accompaniment re-recording according to whether the similarity meets the preset similarity condition.
[0167] In a possible implementation, the obtaining module 310 is further configured to:
[0168] Obtain the lyrics file of the target song, where the lyrics file includes the lyrics of the target song, and the singing start time and singing end time of each word in the lyrics;
[0169] In the lyrics file, if it is determined that the singing time interval between two adjacent words is greater than a preset duration, obtain an accompaniment audio segment of the preset duration in the accompaniment audio corresponding to the two adjacent words, where the singing time interval between the two adjacent words is the time interval from the singing end time of the previous word to the singing start time of the next word among the two adjacent words.
[0170] In a possible implementation, the obtaining module 310 is configured to:
[0171] Perform frame splitting on the accompaniment audio segment to obtain a plurality of first audio frames corresponding to the accompaniment audio segment;
[0172] Determine the first feature vector of each first audio frame, and combine the first feature vectors of each first audio frame to obtain the first audio feature of the accompaniment audio segment;
[0173] Perform frame splitting on the audio segment to be compared to obtain a plurality of second audio frames corresponding to the audio segment to be compared;
[0174] Determine the second feature vector of each second audio frame, and combine the second feature vectors of each second audio frame to obtain the first audio feature of the accompaniment audio segment.
[0175] In a possible implementation, the obtaining module 310 is configured to:
[0176] Obtain the first frequency domain signal corresponding to each first audio frame, and based on each first frequency domain signal, obtain the energy of each first audio frame in N frequency bands, where N is a preset positive integer;
[0177] For each first audio frame, determine the first feature vector of the first audio frame based on the energy of the first audio frame in the N frequency bands and the energy of the previous first audio frame of the first audio frame in the N frequency bands;
[0178] The obtaining module 310 is configured to:
[0179] Obtain the second frequency domain signal corresponding to each second audio frame, and based on each second frequency domain signal, obtain the energy of each second audio frame in the N frequency bands;
[0180] For each second audio frame, determine a second feature vector of the second audio frame according to the energy of the second audio frame in the N frequency bands and the energy of the previous second audio frame of the second audio frame in the N frequency bands.
[0181] In a possible implementation, the obtaining module 310 is configured to:
[0182] Determine the value of the nth element in the first feature vector of the first audio frame according to the energy of the first audio frame in the nth frequency band among the N frequency bands, the energy of the first audio frame in the (n + 1)th frequency band among the N frequency bands, the energy of the previous first audio frame of the first audio frame in the nth frequency band, and the energy of the previous first audio frame in the (n + 1)th frequency band, where 1 ≤ n ≤ N - 1;
[0183] Obtain the first feature vector according to the value of each element in the first feature vector of the first audio frame;
[0184] Determine the value of the nth element in the second feature vector of the second audio frame according to the energy of the second audio frame in the nth frequency band among the N frequency bands, the energy of the second audio frame in the (n + 1)th frequency band among the N frequency bands, the energy of the previous second audio frame of the second audio frame in the nth frequency band, and the energy of the previous second audio frame in the (n + 1)th frequency band;
[0185] Obtain the second feature vector according to the value of each element in the second feature vector of the second audio frame.
[0186] In a possible implementation, the determining module 320 is configured to:
[0187] Calculate a first difference between the energy of the first audio frame in the nth frequency band among the N frequency bands and the energy of the first audio frame in the (n + 1)th frequency band among the N frequency bands;
[0188] Calculate a second difference between the energy of the previous first audio frame of the first audio frame in the nth frequency band among the N frequency bands and the energy of the previous first audio frame in the (n + 1)th frequency band among the N frequency bands;
[0189] If the first difference is greater than the second difference, set the value of the nth element in the first feature vector of the first audio frame to 1;
[0190] If the first difference is not greater than the second difference, set the value of the nth element in the first feature vector of the first audio frame to 0;
[0191] The obtaining module 310 is configured to:
[0192] Calculate a third difference between the energy of the second audio frame in the n-th frequency band among the N frequency bands and the energy of the second audio frame in the (n + 1)-th frequency band among the N frequency bands;
[0193] Calculate a fourth difference between the energy of the previous second audio frame of the second audio frame in the n-th frequency band among the N frequency bands and the energy of the second audio frame in the (n + 1)-th frequency band among the N frequency bands;
[0194] If the third difference is greater than the fourth difference, set the n-th element in the second feature vector of the second audio frame to 1;
[0195] If the third difference is not greater than the fourth difference, set the n-th element in the second feature vector of the second audio frame to 0.
[0196] In a possible implementation manner, the determination module 320 is configured to:
[0197] Starting from the first first feature vector in the first audio feature, select first feature vectors as first starting feature vectors in sequence according to a first step size until the M-th first feature vector is selected, and then stop selecting the first starting feature vectors, where the first step size is the number of first feature vectors offset each time, and M is a preset positive integer;
[0198] For each selected first starting feature vector, in the first audio feature, use the currently selected first starting feature vector to the last first feature vector as the first comparison audio feature of the accompaniment audio segment, and in the second audio feature, continuously select multiple second feature vectors starting from the first second feature vector as the second comparison audio feature of the audio segment to be compared, where the number of first feature vectors included in the first comparison audio feature is the same as the number of second feature vectors included in the second comparison audio feature;
[0199] Starting from the first second feature vector in the second audio feature, select second feature vectors as second starting feature vectors in sequence according to a second step size until the M-th second feature vector is selected, and then stop selecting the second starting feature vectors, where the second step size is the number of second feature vectors offset each time;
[0200] For each selected second starting feature vector, in the second audio feature, use the currently selected second starting feature vector to the last second feature vector as the second comparison audio feature of the audio segment to be compared, and in the first audio feature, select multiple first feature vectors continuously starting from the first first feature vector as the first comparison audio feature of the accompaniment audio segment;
[0201] For the first comparison audio feature and the second comparison audio feature obtained within the same loop, calculate the number of elements with the same position. Take the ratio of the number of elements with the same position to the total number of elements of the first comparison audio feature as the reference similarity.
[0202] Select the maximum reference similarity as the similarity between the first audio feature and the second audio feature.
[0203] In a possible implementation manner, the determination module 320 is configured to:
[0204] If the similarity between the first audio feature and the second audio feature is greater than a preset similarity threshold, and the range of the reference similarity is greater than a preset range threshold, it is determined that there is an accompaniment resampling.
[0205] If the similarity between the first audio feature and the second audio feature is not greater than the preset similarity threshold, and / or the range of the reference similarity is not greater than the preset range threshold, it is determined that there is no accompaniment resampling.
[0206] In a possible implementation manner, the device further includes a display module, configured to:
[0207] If the dry audio is collected through a wired headset, display a prompt message, where the prompt message is used to prompt the possibility of not supporting the wired headset jack standard.
[0208] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0209] It should be noted that when the device for detecting accompaniment resampling provided in the above embodiments detects accompaniment resampling, only the above division of each functional module is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device for detecting accompaniment resampling provided in the above embodiments and the method embodiments for detecting accompaniment resampling belong to the same concept, and the specific implementation process can be seen in the method embodiments, which will not be repeated here.
[0210] Figure 4The block diagram of the electronic device 400 provided by an exemplary embodiment of the present application is shown. The electronic device 400 may be a portable mobile terminal, such as: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer or a desktop computer. The electronic device 400 may also be referred to by other names such as a terminal, a user equipment, a portable terminal, a laptop terminal, a desktop terminal, etc.
[0211] Generally, the electronic device 400 includes: a processor 401 and a memory 402.
[0212] The processor 401 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 401 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), PLA (Programmable Logic Array). The processor 401 may also include a main processor and a co-processor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the co-processor is a low-power processor for processing data in the standby state. In some embodiments, the processor 401 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 401 may also include an AI (Artificial Intelligence) processor, and the AI processor is used to process the computational operations related to machine learning.
[0213] The memory 402 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 402 may also include high-speed random access memory, as well as non-volatile memory, such as one or more disk storage devices, flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 402 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 401 to implement the method for detecting and retrieving the accompaniment provided by the method embodiments in the present application.
[0214] In some embodiments, the electronic device 400 may further optionally include: a peripheral device interface 403 and at least one peripheral device. The processor 401, the memory 402, and the peripheral device interface 403 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 403 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of: a radio frequency circuit 404, a display screen 405, a camera assembly 406, an audio circuit 407, a positioning assembly 408, and a power supply 409.
[0215] The peripheral device interface 403 may be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 401 and the memory 402. In some embodiments, the processor 401, the memory 402, and the peripheral device interface 403 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 401, the memory 402, and the peripheral device interface 403 may be implemented on a separate chip or circuit board, and this embodiment does not limit this.
[0216] The radio frequency circuit 404 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 404 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 404 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 404 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 404 may communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, generations of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 404 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.
[0217] The display screen 405 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 405 is a touch display screen, the display screen 405 also has the ability to collect touch signals on or above the surface of the display screen 405. The touch signals can be input to the processor 401 as control signals for processing. At this time, the display screen 405 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there can be one display screen 405, which is provided on the front panel of the electronic device 400; in other embodiments, there can be at least two display screens 405, which are respectively provided on different surfaces of the electronic device 400 or are in a foldable design; in other embodiments, the display screen 405 can be a flexible display screen, which is provided on the curved surface or the folding surface of the electronic device 400. Even more, the display screen 405 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 405 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0218] The camera module 406 is used to collect images or videos. Optionally, the camera module 406 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the terminal, and the rear camera is provided on the back of the terminal. In some embodiments, there are at least two rear cameras, which are respectively any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, so as to implement the function of background blurring by fusing the main camera and the depth-of-field camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting function or other fused shooting functions. In some embodiments, the camera module 406 can also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.
[0219] The audio circuit 407 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 401 for processing, or input to the radio frequency circuit 404 to enable voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the electronic device 400. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 401 or the radio frequency circuit 404 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 407 may also include a headphone jack.
[0220] The positioning component 408 is used to locate the current geographical location of the electronic device 400 to achieve navigation or LBS (Location Based Service). The positioning component 408 may be a positioning component based on the GPS (Global Positioning System) of the United States, the Beidou system of China, or the Galileo system of Russia.
[0221] The power supply 409 is used to supply power to each component in the electronic device 400. The power supply 409 may be alternating current, direct current, a disposable battery, or a rechargeable battery. When the power supply 409 includes a rechargeable battery, the rechargeable battery may be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0222] In some embodiments, the electronic device 400 further includes one or more sensors 410. The one or more sensors 410 include but are not limited to: an acceleration sensor 411, a gyroscope sensor 412, a pressure sensor 413, a fingerprint sensor 414, an optical sensor 415, and a proximity sensor 416.
[0223] The acceleration sensor 411 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established with the electronic device 400. For example, the acceleration sensor 411 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 401 can control the display screen 405 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 411. The acceleration sensor 411 can also be used for game or user motion data collection.
[0224] The gyroscope sensor 412 can detect the body direction and rotation angle of the electronic device 400. The gyroscope sensor 412 can cooperate with the acceleration sensor 411 to collect the 3D actions of the user on the electronic device 400. Based on the data collected by the gyroscope sensor 412, the processor 401 can implement the following functions: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.
[0225] The pressure sensor 413 can be disposed on the side frame of the electronic device 400 and / or the lower layer of the display screen 405. When the pressure sensor 413 is disposed on the side frame of the electronic device 400, it can detect the holding signal of the user on the electronic device 400, and the processor 401 can perform left / right hand recognition or shortcut operations according to the holding signal collected by the pressure sensor 413. When the pressure sensor 413 is disposed on the lower layer of the display screen 405, the processor 401 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 405. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0226] The fingerprint sensor 414 is used to collect the fingerprints of the user. The processor 401 can identify the user's identity according to the fingerprints collected by the fingerprint sensor 414, or the fingerprint sensor 414 can identify the user's identity according to the collected fingerprints. When the identity of the user is identified as a trusted identity, the processor 401 authorizes the user to perform relevant sensitive operations, and the sensitive operations include unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings, etc. The fingerprint sensor 414 can be disposed on the front, back, or side of the electronic device 400. When there are physical buttons or manufacturer logos on the electronic device 400, the fingerprint sensor 414 can be integrated with the physical buttons or manufacturer logos.
[0227] The optical sensor 415 is used to collect the ambient light intensity. In one embodiment, the processor 401 can control the display brightness of the display screen 405 according to the ambient light intensity collected by the optical sensor 415. Specifically, when the ambient light intensity is high, the display brightness of the display screen 405 is increased; when the ambient light intensity is low, the display brightness of the display screen 405 is decreased. In another embodiment, the processor 401 can also dynamically adjust the shooting parameters of the camera module 406 according to the ambient light intensity collected by the optical sensor 415.
[0228] The proximity sensor 416, also known as a distance sensor, is typically disposed on the front panel of the electronic device 400. The proximity sensor 416 is used to collect the distance between the user and the front of the electronic device 400. In one embodiment, when the proximity sensor 416 detects that the distance between the user and the front of the electronic device 400 is gradually decreasing, the processor 401 controls the display screen 405 to switch from the lit state to the off state; when the proximity sensor 416 detects that the distance between the user and the front of the electronic device 400 is gradually increasing, the processor 401 controls the display screen 405 to switch from the off state to the lit state.
[0229] Those skilled in the art can understand that Figure 4 the structure shown in does not constitute a limitation on the electronic device 400, and it may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.
[0230] In an exemplary embodiment, a computer-readable storage medium is further provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the method for detecting accompaniment back collection in the above embodiment. For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0231] In an exemplary embodiment, a computer program product is further provided. The computer program product includes at least one instruction, and the at least one instruction is loaded and executed by a processor to implement the method for detecting accompaniment back collection as described above.
[0232] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiment can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, or an optical disc, etc.
[0233] The foregoing is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
[0234] It should be noted that the information involved in this application (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals (including but not limited to signals transmitted between user terminals and other devices, etc.) are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. For example, the dry audio involved in this application is obtained under full authorization.
Claims
1. A method for detecting accompaniment re - extraction, characterized in that, the method includes: Obtain the accompaniment audio and dry - voice audio of the target song, where the dry - voice audio is the human - voice audio collected when playing the accompaniment audio; In the accompaniment audio, obtain an accompaniment audio segment with a preset duration, where there are no lyrics within the playing time corresponding to the accompaniment audio segment in the lyrics file of the target song; and in the dry - voice audio, obtain a to - be - compared audio segment corresponding to the playing time of the accompaniment audio segment; Obtain the first audio feature of the accompaniment audio segment and obtain the second audio feature of the to - be - compared audio segment; Determine the similarity between the first audio feature and the second audio feature, and determine whether there is accompaniment re - extraction according to whether the similarity meets the preset similarity condition.
2. The method according to claim 1, characterized in that, the obtaining an accompaniment audio segment with a preset duration in the accompaniment audio includes: Obtain the lyrics file of the target song, where the lyrics file includes the lyrics of the target song and the singing start time and singing end time of each word in the lyrics; In the lyrics file, if it is determined that the singing time interval between two adjacent words is greater than the preset duration, then in the accompaniment audio corresponding to the two adjacent words, obtain the accompaniment audio segment with the preset duration, where the singing time interval between the two adjacent words is the time interval from the singing end time of the previous word to the singing start time of the next word among the two adjacent words.
3. The method according to claim 1, characterized in that, the obtaining the first audio feature of the accompaniment audio segment and obtaining the second audio feature of the to - be - compared audio segment includes: Perform frame - splitting processing on the accompaniment audio segment to obtain multiple first audio frames corresponding to the accompaniment audio segment; Determine the first feature vector of each first audio frame, and combine the first feature vectors of each first audio frame to obtain the first audio feature of the accompaniment audio segment; Perform frame - splitting processing on the to - be - compared audio segment to obtain multiple second audio frames corresponding to the to - be - compared audio segment; Determine the second feature vector of each second audio frame, and combine the second feature vectors of each second audio frame to obtain the first audio feature of the accompaniment audio segment.
4. The method according to claim 3, characterized in that, the determining the first feature vector of each first audio frame includes: Obtain the first frequency - domain signal corresponding to each first audio frame, and according to each first frequency - domain signal, obtain the energy of each first audio frame in N frequency bands, where N is a preset positive integer; For each first audio frame, determine the first feature vector of the first audio frame according to the energy of the first audio frame in the N frequency bands and the energy of the previous first audio frame of the first audio frame in the N frequency bands; the determining the second feature vector of each second audio frame includes: Obtain the second frequency - domain signal corresponding to each second audio frame, and according to each second frequency - domain signal, obtain the energy of each second audio frame in the N frequency bands; For each second audio frame, determine a second feature vector of the second audio frame according to the energy of the second audio frame in the N frequency bands and the energy of the previous second audio frame of the second audio frame in the N frequency bands.
5. The method according to claim 4, wherein, the determining the first feature vector of the first audio frame according to the energy of the first audio frame in the N frequency bands and the energy of the previous first audio frame of the first audio frame in the N frequency bands includes: determining the value of the n-th element in the first feature vector of the first audio frame according to the energy of the first audio frame in the n-th frequency band among the N frequency bands, the energy of the first audio frame in the (n + 1)-th frequency band among the N frequency bands, the energy of the previous first audio frame of the first audio frame in the n-th frequency band, and the energy of the previous first audio frame in the (n + 1)-th frequency band, where 1 ≤ n ≤ N - 1; obtaining the first feature vector according to the value of each element in the first feature vector of the first audio frame; determining the value of the n-th element in the second feature vector of the second audio frame according to the energy of the second audio frame in the n-th frequency band among the N frequency bands, the energy of the second audio frame in the (n + 1)-th frequency band among the N frequency bands, the energy of the previous second audio frame of the second audio frame in the n-th frequency band, and the energy of the previous second audio frame in the (n + 1)-th frequency band; obtaining the second feature vector according to the value of each element in the second feature vector of the second audio frame.
6. The method according to claim 5, wherein, the determining the value of the n-th element in the first feature vector of the first audio frame according to the energy of the first audio frame in the n-th frequency band among the N frequency bands, the energy of the first audio frame in the (n + 1)-th frequency band among the N frequency bands, the energy of the previous first audio frame of the first audio frame in the n-th frequency band, and the energy of the previous first audio frame in the (n + 1)-th frequency band includes: calculating a first difference between the energy of the first audio frame in the n-th frequency band among the N frequency bands and the energy of the first audio frame in the (n + 1)-th frequency band among the N frequency bands; calculating a second difference between the energy of the previous first audio frame of the first audio frame in the n-th frequency band among the N frequency bands and the energy of the previous first audio frame in the (n + 1)-th frequency band among the N frequency bands; if the first difference is greater than the second difference, then setting the value of the n-th element in the first feature vector of the first audio frame to 1; if the first difference is not greater than the second difference, then setting the value of the n-th element in the first feature vector of the first audio frame to 0. Determining the value of the n-th element in the second feature vector of the second audio frame based on the energy of the second audio frame in the n-th frequency band among the N frequency bands, the energy of the second audio frame in the (n + 1)-th frequency band among the N frequency bands, the energy of the previous second audio frame of the second audio frame in the n-th frequency band, and the energy of the previous second audio frame in the (n + 1)-th frequency band includes: Calculating a third difference between the energy of the second audio frame in the n-th frequency band among the N frequency bands and the energy of the second audio frame in the (n + 1)-th frequency band among the N frequency bands; Calculating a fourth difference between the energy of the previous second audio frame of the second audio frame in the n-th frequency band among the N frequency bands and the energy of the previous second audio frame in the (n + 1)-th frequency band among the N frequency bands; If the third difference is greater than the fourth difference, setting the value of the n-th element in the second feature vector of the second audio frame to 1; If the third difference is not greater than the fourth difference, setting the value of the n-th element in the second feature vector of the second audio frame to 0.
7. The method according to claim 6, wherein, determining the similarity between the first audio feature and the second audio feature includes: Starting from the first first feature vector in the first audio feature, successively selecting first feature vectors as first starting feature vectors according to a first step length until the M-th first feature vector is selected, and then stopping selecting the first starting feature vectors, where the first step length is the number of first feature vectors offset each time, and M is a preset positive integer; For each selected first starting feature vector, in the first audio feature, taking the currently selected first starting feature vector to the last first feature vector as the first comparison audio feature of the accompaniment audio segment, and in the second audio feature, successively selecting a plurality of second feature vectors starting from the first second feature vector as the second comparison audio feature of the audio segment to be compared, where the number of first feature vectors included in the first comparison audio feature is the same as the number of second feature vectors included in the second comparison audio feature; Starting from the first second feature vector in the second audio feature, successively selecting second feature vectors as second starting feature vectors according to a second step length until the M-th second feature vector is selected, and then stopping selecting the second starting feature vectors, where the second step length is the number of second feature vectors offset each time; For each selected second starting feature vector, in the second audio feature, taking the currently selected second starting feature vector to the last second feature vector as the second comparison audio feature of the audio segment to be compared, and in the first audio feature, successively selecting a plurality of first feature vectors starting from the first first feature vector as the first comparison audio feature of the accompaniment audio segment; For the first comparison audio feature and the second comparison audio feature obtained in the same cycle, calculating the number of elements with the same position, and taking the ratio of the number of elements with the same position to the total number of elements of the first comparison audio feature as the reference similarity; Select the maximum reference similarity as the similarity between the first audio feature and the second audio feature.
8. The method according to claim 7, wherein, the determining whether there is accompaniment resampling according to whether the similarity meets a preset similarity condition includes: if the similarity between the first audio feature and the second audio feature is greater than a preset similarity threshold, and the range of the reference similarity is greater than a preset range threshold, it is determined that there is accompaniment resampling; if the similarity between the first audio feature and the second audio feature is not greater than the preset similarity threshold, and / or the range of the reference similarity is not greater than the preset range threshold, it is determined that there is no accompaniment resampling.
9. The method according to any one of claims 1-8, wherein, after determining that there is accompaniment resampling, the method further includes: if the dry audio is collected through a wired headset, display a prompt message, wherein the prompt message is used to prompt the possibility of not supporting the wired headset connector standard.
10. An electronic device, wherein, the electronic device includes a processor and a memory, and at least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the method for detecting accompaniment resampling according to any one of claims 1 to 9.
11. A computer-readable storage medium, wherein, at least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the method for detecting accompaniment resampling according to any one of claims 1 to 9.
Citation Information
Patent Citations
Audio file grading method and device
CN106782600A
Audio processing method and device, equipment and storage medium
CN113963707A