A singing detection method, device, equipment and storage medium

By calculating the difference in duration, pitch similarity and rhythm similarity of users when singing, and using them as target features, the problem of low accuracy in the detection of the degree of matching between users and standard songs in the prior art is solved, and higher matching accuracy and business quality are achieved.

CN114550676BActive Publication Date: 2025-05-16BIGO TECH PTE LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210171896.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-24
Publication Date
2025-05-16
Estimated Expiration
2042-02-24

AI Technical Summary

Technical Problem

The prior art is relatively low in the accuracy of testing the degree of matching between a user when singing and a standard song, especially in scenarios such as free singing, accompaniment, silence, and shouting.

Method used

By collecting the user's audio data, compute the duration difference, pitch similarity, and rhythm similarity between them and the standard song, and splicing these features into target features to improve the accuracy of the degree of matching.

Benefits of technology

This method can improve the accuracy of the degree of matching between the user's singing and the standard song in a variety of scenarios, and enhance the quality of business, especially when the user's tone and rhythm are offset.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114550676B_ABST
    Figure CN114550676B_ABST
Patent Text Reader

Abstract

The present application discloses a singing detection method, device, equipment and storage medium, the method comprising: collecting audio data of the user when the user imitates the singing of a song; calculating the deviation between the singing time of the audio data and the song as the time difference; calculating the similarity between the audio data and the song in high pitch as the pitch similarity; calculating the similarity between the audio data and the song in rhythm as the rhythm similarity; splicing the time difference, pitch similarity and rhythm similarity as target features; detecting the matching degree between the audio data and the song according to the target features. The influence of non-singing operations such as silence, speaking, shouting and interaction on the matching degree is suppressed, and the robustness is improved when the audio data sung by the user is overall offset in pitch and rhythm, and the accuracy of the matching degree between the audio data and the song is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of audio processing, and in particular to a singing detection method, device, equipment and storage medium. Background Art

[0002] Singing is a popular entertainment medium that is widely used online and offline. Accurately detecting the degree of match between a user's singing and a standard song plays a significant role in singing scoring, improving singing skills, and other businesses.

[0003] The current method for evaluating the degree of match between a user's singing and a standard song is mainly to require that the lyrics of the song be strictly aligned with the user's singing voice, that is, to extract the pitch curve of the song sung by the user, and then compare it with the pitch curve of the song, and calculate the error area between the two.

[0004] However, with the rapid popularization of mobile terminals, singing scenarios such as singing programs and singing games have become increasingly diverse, including free singing, a cappella, silence, and shouting. In these scenarios, the lyrics of the song cannot be strictly aligned with the user's singing voice, and the accuracy of detecting the degree of match between the user's singing and the standard song is low, which affects the quality of the service. Summary of the invention

[0005] The present application provides a singing detection method, device, equipment and storage medium to solve the problem of how to improve the accuracy of detecting the degree of matching between a user's singing and a standard song.

[0006] According to one aspect of the present application, a singing detection method is provided, comprising:

[0007] When the user imitates singing a song, collecting audio data of the user;

[0008] Calculating the deviation in singing duration between the audio data and the song as a duration difference;

[0009] Calculating the similarity in pitch between the audio data and the song as pitch similarity;

[0010] Calculating the similarity in rhythm between the audio data and the song as the rhythm similarity;

[0011] splicing the duration difference, the pitch similarity and the rhythm similarity into a target feature;

[0012] The matching degree between the audio data and the song is detected according to the target feature.

[0013] According to another aspect of the present application, a singing detection device is provided, comprising:

[0014] An audio data collection module, used to collect audio data from the user when the user imitates singing a song;

[0015] A duration difference calculation module, used to calculate the deviation in singing duration between the audio data and the song as the duration difference;

[0016] A pitch similarity calculation module, used to calculate the similarity in pitch between the audio data and the song as the pitch similarity;

[0017] A rhythm similarity calculation module, used for calculating the similarity in rhythm between the audio data and the song as the rhythm similarity;

[0018] A target feature splicing module, used for splicing the duration difference, the pitch similarity and the rhythm similarity into a target feature;

[0019] A matching degree calculation module is used to detect the matching degree between the audio data and the song according to the target feature.

[0020] According to another aspect of the present application, a singing detection device is provided, the singing detection device comprising:

[0021] at least one processor; and

[0022] a memory communicatively connected to the at least one processor; wherein,

[0023] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the singing detection method described in any embodiment of the present application.

[0024] According to another aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the singing detection method described in any embodiment of the present application when executed.

[0025] In this embodiment, when the user imitates the singing of a song, audio data of the user is collected; the deviation between the singing time of the audio data and the song is calculated as the time difference; the similarity between the pitch of the audio data and the song is calculated as the pitch similarity; the similarity between the rhythm of the audio data and the song is calculated as the rhythm similarity; the time difference, pitch similarity and rhythm similarity are spliced ​​as the target feature; the matching degree between the audio data and the song is detected according to the target feature. By calculating the deviation between the audio data and the song by the singing time, singing and non-singing can be distinguished, and the influence of non-singing operations such as silence, speaking, and shouting on the matching degree can be suppressed. The similarity between the audio data and the song in pitch and rhythm can be calculated, and the robustness can be improved when the audio data sung by the user has an overall deviation in pitch and rhythm. The features of these three aspects are integrated into the target feature. This multimodal target feature can improve the accuracy of the matching degree between the audio data and the song.

[0026] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0028] Figure 1 is a flow chart of a singing detection method provided according to Embodiment 1 of the present application;

[0029] Figure 2 is a structural diagram of a singing detection model provided according to Embodiment 1 of the present application;

[0030] Figure 3 is a flow chart of a singing detection method provided according to Embodiment 2 of the present application;

[0031] Figure 4 is a structural schematic diagram of a singing detection device provided according to Embodiment 3 of the present application;

[0032] Figure 5 It is a structural schematic diagram of a singing detection device for implementing the singing detection method of an embodiment of the present application. DETAILED DESCRIPTION

[0033] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0035] Embodiment 1

[0036] Figure 1 This is a flow chart of a singing detection method provided in the first embodiment of the present application. This embodiment can be used to detect the matching degree between the audio data sung by the user and the song by comprehensively considering the three factors of singing duration, pitch and rhythm. The method can be executed by a singing detection device, which can be implemented in the form of hardware and / or software, and can be configured in a singing detection device. Figure 1 As shown, the method includes:

[0037] Step 101: When the user imitates the singing of a song, audio data of the user is collected.

[0038] In this embodiment, in scenarios such as karaoke, games, competitions, and practice, the user will select a song as a standard template and imitate the song to sing. During the user's imitation of the song, the microphone in the singing detection device can be called to collect audio data to obtain the sound of the user's singing.

[0039] The song serving as the standard template is usually audio data sung by the original author (singer), or it may be audio data sung with high quality by an imitator (singer) other than the original author, or it may be a MIDI (Musical Instrument Digital Interface) file, etc. This embodiment does not impose any restrictions on this.

[0040] In some cases, while the user is imitating the singing of a song, the singing detection device will play the background accompaniment of these songs. Then, the audio data collected from the user will include the background accompaniment of the song. At this time, the audio data can be separated into tracks, and the background accompaniment and the remaining user's singing voice can be filtered out from the audio data.

[0041] In other cases, when the user imitates the singing of a song, the singing detection device will not play the background accompaniment of these songs. In this case, the user is singing freely, and the audio data collected from the user does not include the background accompaniment of the song, but the user's singing. At this time, the audio data is not separated into tracks, and the audio data is directly used.

[0042] Different from karaoke, games, competitions, practice and other scenarios usually lack the function of scrolling the lyrics of songs by word, and even lack accompaniment. Therefore, the following problems are often faced:

[0043] (1) The audio data freely sung by the user usually has an overall rhythm shift. Compared with the song itself, the rhythm of the user's singing may be faster or slower overall.

[0044] (2) When there is no background accompaniment, users usually sing using a pitch that suits them, which makes it difficult for the pitch of the audio data to be consistent with the pitch of the song.

[0045] (3) In scenarios with a heavy entertainment nature, such as live broadcasts and mini-games, the audio data of users' free singing often contains non-singing audio signals such as silence, talking, shouting, and mic interaction.

[0046] In order to make the lyrics of the song and the singing voice of the user not strictly aligned, the matching degree between the audio data sung by the user and the song can be reasonably detected. If the DTW (Dynamic Time Warp) algorithm is used to align and regularize the pitch curves and then calculate the difference between the pitch curves, however, this method depends on the fundamental frequency. When the fundamental frequency extraction is inaccurate, it will significantly affect the matching degree between the audio data sung by the user and the song, thereby affecting the service quality such as scoring. If the alignment between the audio data sung by the user and the song is achieved through spectral segmentation, however, this method requires the production of a song resource library and the training of a large number of spectral models, which is costly and also has the problem of dependence on fundamental frequency.

[0047] Step 102: Calculate the deviation in singing duration between the audio data and the song as the duration difference.

[0048] Taking into account that the audio data of users' free singing often contains non-singing audio signals such as silence, speaking, and rap interaction, in this embodiment, non-singing audio signals such as silence, speaking, and rap interaction can be suppressed, and the duration of the user's singing in the audio data is counted and compared with the singing duration of the song to obtain the overall deviation between the two, which is recorded as the duration difference.

[0049] In one embodiment of the present application, step 102 may include the following steps:

[0050] Step 1021: Divide the audio data into multiple audio segments.

[0051] In order to accurately detect the duration of the user's singing, the audio data may be divided into multiple segments of the same length, recorded as audio segments, each of which contains one or more frames of audio signals.

[0052] Exemplarily, a window h may be added to the audio data from the initial position of the audio data. win , according to the preset step size h step Moving window h win , until the end of the audio data is reached, extract the audio data in the window (one or more frames of audio signal) as a series of audio clips

[0053] X=[x1,x2,…,x i ,…,x M ], where M is the number of audio clips.

[0054] In general, the step length h step Less than or equal to window h win Length.

[0055] Step 1022: Calculate the confidence of the user's singing for each audio clip.

[0056] In this embodiment, the structure of the singing detection model can be pre-selected and the parameters of the singing detection model can be trained so that the singing detection model can be used to distinguish between singing and non-singing.

[0057] Generally speaking, the singing detection model is a deep learning model, and the structure of the singing detection model is not limited to the artificially designed neural network. It can also be a neural network optimized by a model quantization method, a neural network searched for the characteristics of the user's singing songs by the NAS (Neural Architecture Search) method, and so on. This embodiment does not impose any restrictions on this.

[0058] When training a singing detection model, a singing data set (a collection of audio data) and a non-singing data set (a collection of audio data) can be constructed. The singing data set is a positive sample, and the non-singing data set is a negative sample. For the singing data set, since it does not rely on high-quality professional singers to sing songs, a large amount of audio data of different singing levels is relatively easy to obtain. For the non-singing data set, audio data such as speech data, background music, silence, noise, etc. are also relatively easy to obtain, which makes the cost of training the singing detection model lower.

[0059] Each audio clip is input into the singing detection model. The singing detection model processes the audio clip according to its structure and outputs the confidence that the audio clip is sung by the user. This process can be expressed as f:X→P, where f is the singing detection model, X is the audio clip, P is the confidence, and P=[p1,p2,…,p i ,…,p M ], p i ∈[0,1], represents the confidence of the i-th audio clip. The higher the confidence, the more likely the user is singing normally, and the lower the confidence, the lower the possibility of the user singing normally.

[0060] In one example, if Figure 2 As shown, the trained singing detection model can be loaded into the memory for running. The singing detection model includes an encoder and a decoder. The encoder has a first convolution block ConvBlock_1, a second convolution block ConvBlock_2, a third convolution block ConvBlock_3, and a fourth convolution block ConvBlock_4.

[0061] Among them, the first convolution block ConvBlock_1, the second convolution block ConvBlock_2, the third convolution block ConvBlock_3, and the fourth convolution block ConvBlock_4 are all convolution blocks (Convolution block), and the convolution block is an encapsulation of some structures containing convolution layers.

[0062] Furthermore, the structures of the first convolution block ConvBlock_1, the second convolution block ConvBlock_2, the third convolution block ConvBlock_3, and the fourth convolution block ConvBlock_4 can be set according to the needs of services such as singing scoring and singing skill practice. The structures of the first convolution block ConvBlock_1, the second convolution block ConvBlock_2, the third convolution block ConvBlock_3, and the fourth convolution block ConvBlock_4 can be the same or different, and this embodiment does not limit this.

[0063] Since the encoder Encoder is used to encode the Mel spectrum features and extract high-dimensional features from the Mel spectrum features, the number of channels of the features output by the first convolution block ConvBlock_1, the second convolution block ConvBlock_2, the third convolution block ConvBlock_3, and the fourth convolution block ConvBlock_4 increases in sequence.

[0064] For example, the first convolution block ConvBlock_1 contains a convolution layer, the number of channels of the feature output by the convolution layer is 64, the second convolution block ConvBlock_2 contains a convolution layer, the number of channels of the feature output by the convolution layer is 128, the third convolution block ConvBlock_3 contains two convolution layers, the number of channels of the features output by these two convolution layers is 256, and the fourth convolution block ConvBlock_4 contains two convolution layers, the number of channels of the features output by these two convolution layers is 512.

[0065] For another example, the first convolution block ConvBlock_1 contains two convolution layers, and the number of channels of the features output by these two convolution layers is 64. The second convolution block ConvBlock_2 contains two convolution layers, and the number of channels of the features output by these two convolution layers is 128. The third convolution block ConvBlock_3 contains two convolution layers, and the number of channels of the features output by these two convolution layers is 256. The fourth convolution block ConvBlock_4 contains two convolution layers, and the number of channels of the features output by these two convolution layers is 512.

[0066] In the preprocessing stage, the audio clip can be converted into a spectrum signal, and the Mel spectrum feature is extracted from the spectrum signal to achieve the extraction of Mel spectrum features from the audio clip.

[0067] In the encoder stage, the Mel spectrum feature is input into the first convolution block ConvBlock_1 for convolution processing to obtain the first audio feature.

[0068] The first audio feature is input into the second convolution block ConvBlock_2 for convolution processing to obtain the second audio feature.

[0069] The second audio feature is input into the third convolution block ConvBlock_3 for convolution processing to obtain the third audio feature.

[0070] The third audio feature is input into the fourth convolution block ConvBlock_4 for convolution processing to obtain a fourth audio feature.

[0071] The average value of the fourth audio feature is calculated along the dimension of the channel by using methods such as torch.mean to obtain a feature sequence.

[0072] In the decoder stage, the feature sequence is input into the decoder, and the multi-head attention mechanism is used to fuse the temporal context information of the feature sequence to obtain the target feature.

[0073] The target features are activated through methods such as sigmoid to obtain the confidence of the user singing in the audio clip.

[0074] Step 1023: Calculate the duration of the user's singing in the audio data according to the confidence level as the duration of the user's singing.

[0075] For each audio clip, its confidence represents the possibility of the user singing normally. For all audio clips, its overall confidence can be mapped to the duration of the user singing in the audio data, which is recorded as the user singing duration.

[0076] Generally speaking, the duration of a user's singing is positively correlated with the overall confidence, that is, the higher the overall confidence, the longer the user's singing duration; conversely, the lower the overall confidence, the shorter the user's singing duration.

[0077] In one mapping method, a threshold is taken for the confidence. If the confidence is greater than the preset threshold, it means that the confidence is high, and it can be determined that the user is singing in the audio clip. If the confidence is less than or equal to the preset threshold, it means that the confidence is low, and it can be determined that the user is not singing in the audio clip, thereby counting the number of audio clips in which the user is singing, as shown below:

[0078]

[0079] Among them, n sing is the number of audio clips that the user is singing, M is the number of all audio clips, and p i is the confidence of the ith audio segment, τ is the threshold, [p i -τ] + Indicates that in p i >τ, the result is 1, at p i When ≤τ, the result is 0.

[0080] Compare the quantity to zero.

[0081] If the number is greater than zero, calculate the difference between the number and the preset constant, calculate the product of the difference and the step size, calculate the sum of the product and the window, and assign the sum to the duration of the user's singing in the audio data as the user's singing duration.

[0082] If the quantity is equal to zero, zero is assigned as the duration of the user's singing in the audio data as the duration of the user's singing.

[0083] The above process can be expressed as follows:

[0084]

[0085] Among them, t is the duration of the user singing, n sing is the number of audio clips that the user is singing, α is a constant, such as 1, h win is the window (length), h step is the step length.

[0086] Step 1024, query the duration of the singer's singing in the song, and use it as a reference singing duration.

[0087] In this embodiment, the singing time of the singer can be calculated in advance for the song serving as the standard template and recorded as the reference singing time. A mapping relationship between the song and the reference singing time can be established in the server. When the user imitates the singing of the song, the reference singing time of the song can be queried from the server based on the song's ID, name and other information.

[0088] Furthermore, the method of calculating the singing time of the user based on the audio data recorded when the user sings is the same as the method of calculating the reference singing time of the song recorded when the singer sings.

[0089] Step 1025: Calculate the deviation between the user's singing duration and the reference singing duration as the duration difference.

[0090] In this embodiment, the user's singing duration can be compared with a reference singing duration, and the deviation between the user's singing duration and the reference singing duration can be calculated and recorded as the duration difference.

[0091] Exemplarily, the absolute value of the difference between the user's singing duration and the reference singing duration can be taken, and the ratio between the absolute value and the reference singing duration can be calculated as the duration difference, which is expressed as follows:

[0092] feat=abs(t ref -t) / t eef

[0093] Among them, feat is the duration difference, abs() is the absolute value, t ref is the reference singing time, and t is the user's singing time.

[0094] Step 103: Calculate the similarity in pitch between the audio data and the song as the pitch similarity.

[0095] In view of the situation that users adjust the pitch according to their own habits when imitating the singing of a song, this embodiment compares the audio data sung by the user with the song in terms of pitch to obtain the similarity between the two, which is recorded as pitch similarity, thereby making the pitch adjustment robust.

[0096] In one embodiment of the present application, step 103 may include the following steps:

[0097] Step 1031: Calculate the fundamental frequency of each frame of audio signal in the audio data.

[0098] In this embodiment, the audio data can be processed by framing to obtain each frame of audio signal. For each frame of audio signal, the fundamental frequency [f1, f2, ..., f i ,…,f M ], where M is the number of frames, f i is the fundamental frequency of the audio signal of the i-th frame.

[0099] Step 1032: Convert the fundamental frequency into a pitch feature of the audio data according to a preset tuning system as a user pitch feature.

[0100] Generally speaking, songs conform to standard musical temperament, and the audio data of users imitating songs also conform to standard musical temperament. Therefore, according to the preset temperament system (the mathematical method that specifies the origin of each note in the scale and its precise pitch), the fundamental frequency can be converted into the pitch feature of the audio data according to a certain function, which is recorded as the user pitch feature P = [p1, p2, ..., p M ], where M is the number of frames.

[0101] Taking the twelve-tone equal temperament as an example of the tuning system, the twelve-tone equal temperament means that the octave interval is divided into twelve equal parts according to the wavelength ratio, and each part is called a semitone (minor second).

[0102] In this example, the mapped function is as follows:

[0103]

[0104] Among them, p is the user's pitch feature and f is the fundamental frequency.

[0105] Step 1033: Calculate the degree to which each user's pitch feature deviates from the overall user's pitch feature as the user deviation feature.

[0106] In the scenario where users sing freely, they often choose a comfortable tone to sing. There is an offset between the actual singing pitch and the pitch of the song. Directly comparing the absolute pitches of the two is not comprehensive. In this regard, in this embodiment, the degree to which each user's pitch feature deviates from the overall user pitch feature can be calculated and recorded as the user deviation feature.

[0107] For example, the median value can be taken from the overall user pitch feature, and the median value can be subtracted from each pitch feature to obtain the user deviation feature. In this case, the user deviation feature can also be called the zero median pitch feature. This process is expressed as follows:

[0108] P zeroMedian =P-median(P)

[0109] Among them, P zeroMedian is the user deviation feature, P is the user pitch feature, and median() means taking the median value.

[0110] The user deviation feature uses relative pitch to avoid inconsistent matching due to different starting pitches. Compared with using the mean, the median is not easily affected by the outlier value of the individual user pitch feature (the overall user pitch feature).

[0111] Step 1034: perform differential processing on the user pitch feature to obtain the user differential feature.

[0112] In addition to the user deviation feature, in this embodiment, the user pitch feature may be differentially processed and recorded as the user differential feature.

[0113] Taking the first-order difference as an example, the ranking between user pitch features can be determined. For the pitch features of users ranked from the second to the last (the last), the pitch features of the users ranked in the previous position are subtracted to obtain the user differential features. At this time, the user differential features can also be called zero-difference pitch features. This process is expressed as follows:

[0114] P derivation =P[2:N]-P[1:N-1]

[0115] Among them, P derivation represents user differential features, P[2:N] represents the user pitch features from the second to the last (the first to the last), and P[1:N-1] represents the user pitch features from the first to the second to the last (the second to the last).

[0116] Of course, in addition to the first-order difference, higher-order differences such as the second-order difference, the third-order difference and their combinations may also be used to calculate user differential features, etc. This embodiment does not impose any limitation on this.

[0117] The user differential feature focuses on the amplitude and duration of the changes in the two pitch values, which also makes it less susceptible to the overall tone adjustment and more suitable for the scenario of users singing freely.

[0118] Step 1035: query the pitch feature of the song as the song pitch feature.

[0119] In this embodiment, the pitch feature of the song used as the standard template can be calculated in advance and recorded as the song pitch feature. Wherein, N is the number of frames. A mapping relationship between songs and song pitch features is established in the server. When a user imitates the song, the song pitch features of the song are queried from the server based on the song ID, name and other information.

[0120] Furthermore, the method of calculating the user's pitch features for the audio data recorded when the user sings is the same as the method of calculating the song's pitch features for the song recorded when the singer sings.

[0121] Step 1036, query the degree to which the pitch feature of each song deviates from the overall song pitch feature, as the song deviation feature.

[0122] In this embodiment, the degree to which the pitch feature of each song deviates from the pitch feature of the entire song can be calculated in advance for the song used as the standard template, and recorded as the song deviation feature. A mapping relationship between songs and song deviation features is established in the server. When a user imitates the singing of the song, the server is queried for the song deviation features based on the song ID, name and other information.

[0123] Furthermore, the method of calculating the user deviation feature for the audio data recorded while the user is singing is the same as the method of calculating the song deviation feature for the song recorded while the singer is singing.

[0124] Step 1037, query the song differential features obtained by performing differential processing on the song pitch features.

[0125] In this embodiment, the song pitch feature of the song used as the standard template can be differentially processed in advance and recorded as the song differential feature A mapping relationship between songs and song differential features is established in the server. When a user imitates the singing of the song, the server is queried for the song differential features based on the song ID, name and other information.

[0126] Furthermore, the method of calculating user differential features for audio data recorded when a user sings is the same as the method of calculating song differential features for a song recorded when a singer sings.

[0127] Step 1038, respectively calculate the similarity between the user pitch feature and the song pitch feature, the similarity between the user deviation feature and the song deviation feature, and the similarity between the user differential feature and the song differential feature as pitch similarity.

[0128] In this embodiment, the pitch similarity includes three dimensions, namely, the similarity between the user pitch feature and the song pitch feature, the similarity between the user deviation feature and the song deviation feature, and the similarity between the user differential feature and the song differential feature.

[0129] For user pitch feature P and song pitch feature P ref The similarity between the user pitch feature P and the song pitch feature P can be calculated ref The distance matrix Dist between Among them, p i is the pitch feature of the ith user, is the pitch feature of the jth song. Nonlinear regularization algorithms such as DTW and its variants and optimized versions are used to match the user pitch feature P of the two sequences with the song pitch feature P ref , the DTW distance obtained is the user pitch feature P and the song pitch feature P ref The similarity between .

[0130] For user deviation feature P zeroMedian Deviation from the song's characteristics The similarity between them can be used to calculate the user deviation feature P zeroMedian Deviation from the song's characteristics The distance matrix between them is used to match the user deviation feature P using nonlinear regularization algorithms such as DTW and its variants and optimized versions zeroMedian Deviation from the song's characteristics The DTW distance obtained is the user's deviation from feature P zeroMedian Deviation from the song's characteristics The similarity between .

[0131] For user differential features P derivation Differential features from songs The similarity between the similarities can be used to calculate the user differential feature P derivation Differential features from songs The distance matrix between them is used to match the user differential features P using nonlinear regularization algorithms such as DTW and its variants and optimized versions derivation Differential features from songs The DTW distance obtained is the user differential feature P derivation Differential features from songs The similarity between .

[0132] Step 104: Calculate the similarity in rhythm between the audio data and the song as the rhythm similarity.

[0133] In music, rhythm refers to regular pitch changes. When a user imitates singing a song, a slight overall faster or slower rhythm will not affect the judgment of whether the singing is good or bad. Therefore, in view of the fact that the user adjusts the rhythm according to his own habits when imitating the singing of a song, causing the overall rhythm to shift, this embodiment compares the audio data sung by the user with the song in terms of rhythm, obtains the similarity between the two, and records it as rhythm similarity, so as to make the rhythm adjustment robust, so that the user can allow a certain degree of overall rhythm change when imitating the singing of a song.

[0134] In one embodiment of the present application, step 104 may include the following steps:

[0135] Step 1041, query the user pitch characteristics of the audio data and the song pitch characteristics of the song.

[0136] In this embodiment, the fundamental frequency of each frame of audio signal in the audio data is calculated in advance, and the fundamental frequency is converted into a pitch feature of the audio data according to a preset law as a user pitch feature, and the user pitch feature is cached. At this time, the user pitch feature of the audio data can be queried in the cache.

[0137] And, the pitch features of the song as the standard template are calculated in advance and recorded as the song pitch features. A mapping relationship between the song and the song pitch features is established in the server. At this time, the song pitch features of the song can be queried from the server based on the song ID, name and other information.

[0138] Of course, if the song pitch features of the song have been queried from the server based on the song ID, name and other information in advance, and the song pitch features have been cached, the song pitch features in pitch can be queried in the cache.

[0139] Step 1042: Fit the user pitch feature as the first path and fit the song pitch feature as the second path.

[0140] In this embodiment, a nonlinear regularization algorithm such as DTW and its variants and optimized versions can be used to fit the user's pitch features into a path, recorded as the first path, and the first path can be used to express the rhythm of the audio data. In addition, a nonlinear regularization algorithm such as DTW and its variants and optimized versions can be used to fit the song's pitch features into a path, recorded as the second path, and the second path can be used to express the rhythm of the song.

[0141] Exemplarily, the first path and the second path are both straight lines.

[0142] Step 1043: Calculate the overall error between the first path and the second path as the rhythm similarity.

[0143] In this embodiment, the first path and the second path may be compared as a whole, and the overall error between the first path and the second path may be calculated and recorded as the rhythm similarity.

[0144] In a specific implementation, a plurality of first points are extracted from the first path and a plurality of second points are extracted from the second path, the first points are matched with the second points, and a first distance (such as Euclidean distance) between the first points and the second points that match each other is calculated, a first square is taken for each first distance, a first average value is calculated for all first squares, and the square root of the first average value is taken to obtain the overall error between the first path and the second path as the rhythm similarity, which is expressed as follows:

[0145]

[0146] Where, ∈ is the overall error between the first path and the second path, K represents the number of first points and second points that match each other, ∈ i Represents the first distance between the i-th matching first point and second point.

[0147] Step 1044: extract a first power spectrum from the audio data and a second power spectrum from the song.

[0148] In this embodiment, the preset feature operator can be used to calculate the power spectrum in the audio data, especially the short-time power spectrum, such as MFCC (Mel Frequency Cepstrum Coefficient), linear frequency cepstrum coefficients, Bark frequency cepstrum coefficients, etc., which is recorded as the first power spectrum. Similarly, the preset feature operator can be used to calculate the power spectrum in the song (the data itself), especially the short-time power spectrum, such as MFCC, linear frequency cepstrum coefficients, Bark frequency cepstrum coefficients, etc., which is recorded as the second power spectrum.

[0149] Step 1045: Fit the first power spectrum to the third path and fit the second power spectrum to the fourth path.

[0150] In this embodiment, the first power spectrum can be fit into a path using nonlinear regularization algorithms such as DTW and its variants and optimized versions, recorded as a third path, and the third path can be used to express the rhythm of the audio data. In addition, the second power spectrum can be fit into a path using nonlinear regularization algorithms such as DTW and its variants and optimized versions, recorded as a fourth path, and the fourth path can be used to express the rhythm of the song.

[0151] Exemplarily, the third path and the fourth path are both straight lines.

[0152] Step 1046: Calculate the overall error between the third path and the fourth path as the rhythm similarity.

[0153] In this embodiment, the third path and the fourth path may be compared as a whole, and the overall error between the third path and the fourth path may be calculated and recorded as the rhythm similarity.

[0154] In a specific implementation, multiple third points may be extracted from the third path and multiple fourth points may be extracted from the fourth path, and the third point and the fourth point may be matched. For the third point and the fourth point that match each other, a second distance (such as Euclidean distance) between the third point and the fourth point may be calculated. A second square may be taken for each second distance, a second average value may be calculated for all second squares, and the second average value may be squared to obtain an overall error between the third path and the fourth path as the rhythm similarity, which may be expressed as follows:

[0155]

[0156] Where, ∈ is the overall error between the third path and the fourth path, K represents the number of third points and fourth points that match each other, ∈ i represents the second distance between the i-th matching third point and fourth point.

[0157] If the audio data sung by the user is exactly the same as the rhythm of the song, the paths (i.e., the first path, the second path, the third path, and the fourth path) calculated under nonlinear regularization algorithms such as DTW and its variants and optimized versions are close to the diagonal of the distance matrix. At this time, the fitting error is small; if there is an overall deviation between the audio data sung by the user and the rhythm of the song, the paths (i.e., the first path, the second path, the third path, and the fourth path) under nonlinear regularization algorithms such as DTW and its variants and optimized versions are at a certain angle to the diagonal of the distance matrix, but because the overall speed is too fast or too slow, the fitting error is also relatively small.

[0158] The rhythm similarity that depends on pitch is related to the accuracy of extracting pitch features, while the rhythm similarity that depends on power spectrum is not affected by pitch and has the ability to express the audio envelope.

[0159] Step 105: concatenate the duration difference, pitch similarity and rhythm similarity into target features.

[0160] In this embodiment, the duration difference, pitch similarity and rhythm similarity may be fused, such as concatenated, so that the duration difference, pitch similarity and rhythm similarity are combined into an overall feature, which is recorded as the target feature.

[0161] Step 106: Detect the matching degree between the audio data and the song according to the target features.

[0162] For multimodal target features, the target features can be mapped to the matching degree between the audio data and the song.

[0163] In the specific implementation, a decision tree (DT) can be generated in advance, and its generation algorithm includes ID3, C4.5 and C5.0. A decision tree is a tree structure, in which each internal node represents a judgment on an attribute, each branch represents the output of a judgment result, and finally each leaf node represents a classification result (i.e., the matching degree between the target feature detection audio data and the song).

[0164] The decision tree is loaded into memory for execution, and the target features are input into the decision tree for processing to output the degree of match between the audio data and the song.

[0165] In this embodiment, when the user imitates the singing of a song, audio data of the user is collected; the deviation between the singing time of the audio data and the song is calculated as the time difference; the similarity between the pitch of the audio data and the song is calculated as the pitch similarity; the similarity between the rhythm of the audio data and the song is calculated as the rhythm similarity; the time difference, pitch similarity and rhythm similarity are spliced ​​as the target feature; the matching degree between the audio data and the song is detected according to the target feature. By calculating the deviation between the audio data and the song by the singing time, singing and non-singing can be distinguished, and the influence of non-singing operations such as silence, speaking, and shouting on the matching degree can be suppressed. The similarity between the audio data and the song in pitch and rhythm can be calculated, and the robustness can be improved when the audio data sung by the user has an overall deviation in pitch and rhythm. The features of these three aspects are integrated into the target feature. This multimodal target feature can improve the accuracy of the matching degree between the audio data and the song.

[0166] Embodiment 2

[0167] Figure 3 This is a flow chart of a singing detection method provided in Example 2 of this application. This embodiment adds a scoring business operation based on the above embodiment. Figure 3 As shown, the method includes:

[0168] Step 301: When the user imitates singing a song, audio data is collected from the user.

[0169] Step 302: Calculate the deviation in singing duration between the audio data and the song as the duration difference.

[0170] Step 303: Calculate the similarity in pitch between the audio data and the song as the pitch similarity.

[0171] Step 304: Calculate the similarity in rhythm between the audio data and the song as the rhythm similarity.

[0172] Step 305: concatenate the duration difference, pitch similarity and rhythm similarity into target features.

[0173] Step 306: Detect the matching degree between the audio data and the song according to the target features.

[0174] Step 307: Map the matching degree to a score of the user's imitation of the song.

[0175] In scenarios such as karaoke and games, the degree of match between the audio data sung by the user and the song can be mapped to a score of the user's imitation of the song, so that the score can be applied to business operations, such as displaying the score, sorting users by score, analyzing the user's singing disadvantages based on the score, and so on.

[0176] Generally speaking, the score is positively correlated with the degree of match, that is, the higher the degree of match, the higher the score, and conversely, the lower the degree of match, the lower the score.

[0177] For some simple scenarios, the degree of match can be directly assigned as the score of the user's imitation of the song, or the degree of match can be enlarged to a numerical value within a specified range (such as [0, 100]) as the score of the user's imitation of the song, and so on. This embodiment does not limit this.

[0178] For some complex scenarios, in addition to the degree of match between the audio data sung by the user and the song, other factors may also be taken into account to generate a score for the user's imitation of the song through weighted summation, for example, the degree of match between the user's movements and the singer's movements, etc. This embodiment does not limit this.

[0179] Embodiment 3

[0180] Figure 4 This is a schematic diagram of the structure of a singing detection device provided in Example 3 of the present application. Figure 4 As shown, the device comprises:

[0181] The audio data collection module 401 is used to collect audio data from the user when the user imitates the singing of the song;

[0182] The duration difference calculation module 402 is used to calculate the deviation in singing duration between the audio data and the song as the duration difference;

[0183] A pitch similarity calculation module 403 is used to calculate the similarity in pitch between the audio data and the song as pitch similarity;

[0184] A rhythm similarity calculation module 404 is used to calculate the similarity in rhythm between the audio data and the song as the rhythm similarity;

[0185] A target feature splicing module 405 is used to splice the duration difference, the pitch similarity and the rhythm similarity into a target feature;

[0186] The matching degree calculation module 406 is used to detect the matching degree between the audio data and the song according to the target feature.

[0187] In one embodiment of the present application, the duration difference calculation module 402 includes:

[0188] An audio segment segmentation module, used for segmenting the audio data into multiple audio segments;

[0189] A confidence calculation module, used to calculate the confidence of the user's singing for each of the audio clips;

[0190] A user singing duration calculation module, used to calculate the duration of the user singing in the audio data according to the confidence level as the user singing duration;

[0191] A reference singing duration query module is used to query the duration of singing by the singer in the song as a reference singing duration;

[0192] The duration deviation calculation module is used to calculate the deviation between the user's singing duration and the reference singing duration as the duration difference.

[0193] In one embodiment of the present application, the audio segmentation module includes:

[0194] A window adding module, used for adding a window to the audio data and moving the window according to a preset step size;

[0195] The window extraction module is used to extract the audio data in the window as an audio segment.

[0196] In one embodiment of the present application, the confidence calculation module includes:

[0197] A singing detection model loading module is used to load a singing detection model, wherein the singing detection model includes an encoder and a decoder, and the encoder has a first convolution block, a second convolution block, a third convolution block, and a fourth convolution block;

[0198] A Mel spectrum feature extraction module, used to extract the Mel spectrum feature from the audio clip;

[0199] A first convolution processing module, used for inputting the Mel frequency spectrum feature into the first convolution block for convolution processing to obtain a first audio feature;

[0200] A second convolution processing module, used for inputting the first audio feature into the second convolution block for convolution processing to obtain a second audio feature;

[0201] a third convolution processing module, configured to input the second audio feature into the third convolution block for convolution processing to obtain a third audio feature;

[0202] a fourth convolution processing module, configured to input the third audio feature into the fourth convolution block for convolution processing to obtain a fourth audio feature;

[0203] A sequence processing module, used for calculating an average value of the fourth audio feature along the dimension of the channel to obtain a feature sequence;

[0204] A feature fusion module, used to input the feature sequence into the decoder, and fuse the temporal context information of the feature sequence using a multi-head attention mechanism to obtain a target feature;

[0205] The feature activation module is used to activate the target feature to obtain the confidence of the user singing in the audio clip.

[0206] In one embodiment of the present application, the user singing duration calculation module includes:

[0207] A singing determination module, configured to determine that the user is singing in the audio clip if the confidence level is greater than a preset threshold;

[0208] A quantity counting module, used for counting the number of the audio clips being sung by the user;

[0209] A quantity calculation module, configured to calculate, if the quantity is greater than zero, a difference of the quantity minus a preset constant, calculate a product of the difference multiplied by the step length, calculate a sum of the product and the window, and assign the sum to the duration of the user singing in the audio data as the duration of the user singing;

[0210] The duration zeroing module is used to assign zero as the duration of the user's singing in the audio data if the number is equal to zero, as the user's singing duration.

[0211] In one embodiment of the present application, the duration deviation calculation module includes:

[0212] A difference calculation module, used for taking an absolute value of the difference between the singing time of the user and the reference singing time;

[0213] The ratio calculation module is used to calculate the ratio between the absolute value and the reference singing duration as the duration difference.

[0214] In one embodiment of the present application, the pitch similarity calculation module 403 includes:

[0215] A fundamental frequency calculation module, used to calculate the fundamental frequency of each frame of audio signal in the audio data respectively;

[0216] A user pitch feature conversion module, used to convert the fundamental frequency into a feature of the audio data in pitch according to a preset temperament system as a user pitch feature;

[0217] A user deviation feature calculation module, used to respectively calculate the degree to which each user pitch feature deviates from the overall user pitch feature as a user deviation feature;

[0218] A user differential feature calculation module, used for performing differential processing on the user pitch feature to obtain a user differential feature;

[0219] A song pitch feature query module, used to query the pitch feature of the song as the song pitch feature;

[0220] A song deviation feature query module, used to query the degree to which each song pitch feature deviates from the overall song pitch feature as a song deviation feature;

[0221] A song differential feature query module, used to query the song differential features obtained by performing differential processing on the song pitch features;

[0222] The multi-feature similarity calculation module is used to respectively calculate the similarity between the user pitch feature and the song pitch feature, the similarity between the user deviation feature and the song deviation feature, and the similarity between the user differential feature and the song differential feature as pitch similarity.

[0223] In one embodiment of the present application, the user deviation feature calculation module includes:

[0224] A median extraction module, used to extract the median from the overall user pitch features;

[0225] The median deviation calculation module is used to subtract the median from each pitch feature to obtain a user deviation feature.

[0226] In one embodiment of the present application, the user differential feature calculation module includes:

[0227] A pitch feature sorting module, used to determine the sorting between the user's pitch features;

[0228] The first-order difference processing module is used to subtract the pitch features of the user ranked in the previous position from the pitch features of the user ranked in the second position to the last position, respectively, to obtain user difference features.

[0229] In one embodiment of the present application, the rhythm similarity calculation module 404 includes:

[0230] A pitch feature query module, used to query the user pitch feature of the audio data in pitch and the song pitch feature of the song in pitch;

[0231] A pitch path fitting module, used for fitting the user pitch feature into a first path and fitting the song pitch feature into a second path respectively;

[0232] A pitch path error calculation module, used to calculate the overall error between the first path and the second path as rhythm similarity;

[0233] A power spectrum extraction module, used to extract a first power spectrum from the audio data and a second power spectrum from the song;

[0234] A power spectrum path fitting module, used for fitting the first power spectrum into a third path and fitting the second power spectrum into a fourth path respectively;

[0235] The power spectrum path error calculation module is used to calculate the overall error between the third path and the fourth path as the rhythm similarity.

[0236] In one embodiment of the present application, the pitch path error calculation module includes:

[0237] A pitch point extraction module, used to extract a plurality of first points from the first path and a plurality of second points from the second path respectively;

[0238] A pitch point distance calculation module, used for calculating a first distance between the first point and the second point that match each other;

[0239] The pitch point processing calculation module is used to take a first square for each of the first distances, calculate a first average value for all the first squares, take the square root of the first average value, and obtain an overall error between the first path and the second path as rhythm similarity.

[0240] In one embodiment of the present application, the power spectrum path error calculation module includes:

[0241] A power spectrum point extraction module, used to extract a plurality of third points from the third path and a plurality of fourth points from the fourth path respectively;

[0242] A power spectrum point distance calculation module, used for calculating a second distance between the third point and the fourth point that match each other;

[0243] The power spectrum point processing and calculation module is used to take a second square for each of the second distances, calculate a second average value for all the second squares, take the square root of the second average value, and obtain an overall error between the third path and the fourth path as rhythm similarity.

[0244] In one embodiment of the present application, the matching degree calculation module 406 includes:

[0245] Decision tree loading module, used to load decision trees;

[0246] A decision tree processing module is used to input the target feature into the decision tree for processing to output the degree of matching between the audio data and the song.

[0247] In one embodiment of the present application, it also includes:

[0248] A scoring mapping module is used to map the matching degree into a score of the user's imitation of the song.

[0249] The singing detection device provided in the embodiments of the present application can execute the singing detection method provided in any embodiment of the present application, and has the corresponding functional modules and beneficial effects for executing the singing detection method.

[0250] Embodiment 4

[0251] Figure 5 A structural schematic diagram of a singing detection device 10 that can be used to implement an embodiment of the present application is shown. The singing detection device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The singing detection device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.

[0252] like Figure 5As shown, the singing detection device 10 includes at least one processor 11, and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the singing detection device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0253] A number of components in the singing detection device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the singing detection device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0254] The processor 11 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs the various methods and processes described above, such as a singing detection method.

[0255] In some embodiments, the singing detection method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the singing detection device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the singing detection method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to perform the singing detection method in any other suitable manner (e.g., by means of firmware).

[0256] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0257] The computer programs for implementing the methods of the present application may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer programs are executed by the processor, the functions / operations specified in the flow charts and / or block diagrams are implemented. The computer programs may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0258] In the context of the present application, a computer readable storage medium may be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, device or equipment. A computer readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer readable storage medium may be a machine readable signal medium. A more specific example of a machine readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0259] In order to provide interaction with the user, the systems and techniques described herein can be implemented on a singing detection device, which has: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball), through which the user can provide input to the singing detection device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0260] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0261] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.

[0262] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this application can be executed in parallel, sequentially or in different orders, as long as the expected results of the technical solution of this application can be achieved, and this document is not limited here.

Claims

1. A singing detection method, characterized in that: include: When the user imitates singing a song, collecting audio data of the user; Calculating the deviation in singing duration between the audio data and the song as a duration difference; Calculating the similarity in pitch between the audio data and the song as pitch similarity; Calculating the similarity in rhythm between the audio data and the song as the rhythm similarity; splicing the duration difference, the pitch similarity and the rhythm similarity into a target feature; Detecting the degree of matching between the audio data and the song according to the target feature; The calculating the deviation in singing duration between the audio data and the song as the duration difference comprises: The audio data is divided into multiple audio segments; the confidence of the user's singing is calculated for each audio segment; the duration of the user's singing in the audio data is calculated according to the confidence, as the user's singing duration; the duration of the singer's singing in the song is queried, as a reference singing duration; the deviation between the user's singing duration and the reference singing duration is calculated, as the duration difference.

2. The method according to claim 1, characterized in that The step of dividing the audio data into a plurality of audio segments comprises: Adding a window to the audio data, and moving the window according to a preset step size; The audio data in the window is extracted as an audio segment.

3. The method according to claim 1, characterized in that The calculating the confidence of the user's singing for each of the audio clips comprises: Loading a singing detection model, the singing detection model includes an encoder and a decoder, the encoder has a first convolution block, a second convolution block, a third convolution block, and a fourth convolution block; Extracting Mel-spectrogram features from the audio clip; Inputting the Mel frequency spectrum feature into the first convolution block for convolution processing to obtain a first audio feature; Inputting the first audio feature into the second convolution block for convolution processing to obtain a second audio feature; Inputting the second audio feature into the third convolution block for convolution processing to obtain a third audio feature; Inputting the third audio feature into the fourth convolution block for convolution processing to obtain a fourth audio feature; Calculating an average value of the fourth audio feature along the dimension of the channel to obtain a feature sequence; Input the feature sequence into the decoder, and use a multi-head attention mechanism to fuse the temporal context information of the feature sequence to obtain the target feature; The target feature is activated to obtain the confidence level of the user singing in the audio clip.

4. The method according to claim 2, characterized in that: The calculating, according to the confidence level, the duration of the singing of the user in the audio data as the singing duration of the user includes: If the confidence level is greater than a preset threshold, it is determined that the user is singing in the audio clip; Counting the number of the audio clips being sung by the user; If the number is greater than zero, then calculating the difference of the number minus a preset constant, calculating the product of the difference multiplied by the step length, calculating the sum of the product and the window, and assigning the sum to the duration of the user singing in the audio data as the duration of the user singing; If the number is equal to zero, zero is assigned as the duration of the user's singing in the audio data as the user's singing duration.

5. The method according to claim 1, characterized in that The calculating the deviation between the singing duration of the user and the reference singing duration as the duration difference includes: Taking the absolute value of the difference between the singing time of the user and the reference singing time; The ratio between the absolute value and the reference singing duration is calculated as the duration difference.

6. The method according to any one of claims 1 to 5, characterized in that The calculating the similarity in pitch between the audio data and the song as pitch similarity includes: Calculating the fundamental frequency of each frame of audio signal in the audio data respectively; Convert the fundamental frequency into a pitch feature of the audio data according to a preset pitch system as a user pitch feature; Calculating the degree to which each user pitch feature deviates from the overall user pitch feature as a user deviation feature; Performing differential processing on the user pitch feature to obtain a user differential feature; Querying the pitch feature of the song as the song pitch feature; Querying the degree to which each of the song pitch features deviates from the overall song pitch feature as a song deviation feature; Querying a song differential feature obtained by performing differential processing on the song pitch feature; The similarity between the user pitch feature and the song pitch feature, the similarity between the user deviation feature and the song deviation feature, and the similarity between the user differential feature and the song differential feature are calculated respectively as pitch similarity.

7. The method according to claim 6, characterized in that The respectively calculating the degree to which each user pitch feature deviates from the overall user pitch feature as the user deviation feature comprises: Taking the median value from the overall pitch features of the user; The median value is subtracted from each of the user pitch features to obtain a user deviation feature.

8. The method according to claim 6, characterized in that The performing differential processing on the user pitch feature to obtain the user differential feature includes: Determining a ranking between the user pitch features; For the user pitch features ranked from the second to the last, the pitch features of the user ranked in the previous position are respectively subtracted to obtain user differential features.

9. The method according to any one of claims 1 to 5, characterized in that The calculating the similarity in rhythm between the audio data and the song as the rhythm similarity includes: Calculating the fundamental frequency of each frame of the audio signal in the audio data respectively, and converting the fundamental frequency into a pitch feature of the audio data according to a preset law as a user pitch feature; Calculating the pitch feature of the song as a standard template, and recording it as the song pitch feature; The user pitch feature is respectively fitted as a first path, and the song pitch feature is respectively fitted as a second path; calculating an overall error between the first path and the second path as a rhythm similarity; Extracting a first power spectrum from the audio data and a second power spectrum from the song respectively; respectively fitting the first power spectrum into a third path and fitting the second power spectrum into a fourth path; An overall error between the third path and the fourth path is calculated as the rhythm similarity.

10. The method according to claim 9, characterized in that The calculating the overall error between the first path and the second path as the rhythm similarity includes: extracting a plurality of first points from the first path and a plurality of second points from the second path respectively; For the first point and the second point that match each other, calculating a first distance between the first point and the second point; Taking a first square for each of the first distances, calculating a first average value for all the first squares, taking the square root of the first average value, and obtaining an overall error between the first path and the second path as the rhythm similarity; The calculating the overall error between the third path and the fourth path as the rhythm similarity includes: extracting a plurality of third points from the third path and a plurality of fourth points from the fourth path respectively; For the third point and the fourth point that match each other, calculating a second distance between the third point and the fourth point; A second square is taken for each of the second distances, a second average value is calculated for all the second squares, and the second average value is squared to obtain an overall error between the third path and the fourth path as the rhythm similarity.

11. The method according to any one of claims 1-5, 7-8, and 10, characterized in that: The detecting the matching degree between the audio data and the song according to the target feature comprises: Load the decision tree; The target feature is input into the decision tree for processing to output the matching degree between the audio data and the song.

12. A singing detection device, characterized in that: include: An audio data collection module, used to collect audio data from the user when the user imitates singing a song; A duration difference calculation module, used to calculate the deviation in singing duration between the audio data and the song as the duration difference; A pitch similarity calculation module, used to calculate the similarity in pitch between the audio data and the song as the pitch similarity; A rhythm similarity calculation module, used for calculating the similarity in rhythm between the audio data and the song as the rhythm similarity; A target feature splicing module, used for splicing the duration difference, the pitch similarity and the rhythm similarity into a target feature; A matching degree calculation module, used for detecting the matching degree between the audio data and the song according to the target feature; The duration difference calculation module includes: An audio segment segmentation module, used for segmenting the audio data into multiple audio segments; A confidence calculation module, used to calculate the confidence of the user's singing for each of the audio clips; A user singing duration calculation module, used to calculate the duration of the user singing in the audio data according to the confidence level as the user singing duration; A reference singing duration query module is used to query the duration of singing by the singer in the song as a reference singing duration; The duration deviation calculation module is used to calculate the deviation between the user's singing duration and the reference singing duration as the duration difference.

13. A singing detection device, characterized in that: The singing detection device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the singing detection method according to any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the singing detection method described in any one of claims 1-11 when executed.

Citation Information

Patent Citations

  • Vehicle-mounted singing scoring system, method and equipment and storage medium

    CN109003623A

  • Singing scoring method based on lyric and voice alignment

    CN110660383A