Song detection method and device
By converting song text data into phoneme representation vectors and performing acoustic encoding, and combining the similarity calculation of phonemes and acoustic representation vectors, the problem of poor accuracy in identifying duplicate songs in existing technologies is solved, and more efficient duplicate song identification is achieved.
Patent Information
- Application Number
- CN202511060530.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-07
AI Technical Summary
In existing technologies, song detection methods, especially when performing cluster analysis on different adapted versions of the same song or different singers performing the same song, have unstable results and a high error rate, resulting in poor accuracy in identifying duplicate songs.
By converting the text data of a song into a phoneme representation vector and encoding it using a phoneme processing model, an acoustic representation vector is generated. By combining the phonemes and acoustic representation vectors, similarity calculations are performed to classify songs and identify duplicate songs.
It improves the accuracy of duplicate song identification, enabling more accurate identification of different adaptations of the same song and different singers performing the same song, and reduces the clustering error rate.
Smart Images

Figure CN120910302A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular to a song detection method and device. BACKGROUND
[0002] When the existing song detection method is used to perform clustering analysis on songs, the songs are usually taken as a whole for clustering analysis, and the precision of the clustering result is limited. Especially, the clustering analysis result of different versions of the same song or different singers performing the same song is unstable and has a high error rate. SUMMARY
[0003] Therefore, the present application aims to provide a song detection method and device, which converts the text data of a first to-be-detected song into a phoneme representation vector and performs song coding on the first to-be-detected song to obtain an acoustic representation vector, so as to divide the first to-be-detected song into a first song category by performing similarity calculation on the first to-be-detected song through the phoneme representation vector, and divide a second to-be-detected song in the first song category into a second song category according to the acoustic representation vector, so as to identify the songs in the second song category as repeated songs, thereby solving the technical problem of low accuracy of repeated song identification in the prior art and achieving the technical effect of increasing the accuracy of repeated song identification.
[0004] The present application mainly includes the following aspects: In a first aspect, the present application provides a repeated song detection method, which includes: obtaining at least one first to-be-detected song and text data corresponding to the first to-be-detected song; converting the text data into a phoneme representation vector and coding the first to-be-detected song into an acoustic representation vector; performing first similarity calculation on each phoneme representation vector and dividing the first to-be-detected song into at least one first song category; for each first song category, dividing a second to-be-detected song in the first song category into at least one second song category based on the acoustic representation vector corresponding to the second to-be-detected song, and the songs belonging to the same second song category are repeated songs.
[0005] Optionally, converting the text data into a phoneme representation vector includes: converting the text data of each first to-be-detected song into phoneme data; obtaining a phoneme processing model, mapping the phoneme data to a preset dimension space, performing coding processing on each phoneme data according to the context information and phoneme dependency relationship of the phoneme data, and summarizing the coding results of all phoneme data of each first to-be-detected song to obtain at least one phoneme embedding representation sequence of each first to-be-detected song; performing fuzzy processing on each phoneme embedding representation sequence to obtain the phoneme representation vector of each first to-be-detected song.
[0006] Optionally, the fuzzy processing is performed on each phoneme embedding representation sequence to obtain a phoneme representation vector of each first to-be-detected song, including: performing attention mechanism calculation on the at least one phoneme embedding representation sequence, calculating a corresponding attention weight according to an attention score of each phoneme embedding representation sequence, performing fusion of the phoneme embedding representation sequences according to the attention weight, and decoding and restoring the fused sequence to obtain the phoneme representation vector of each first to-be-detected song.
[0007] Optionally, the decoding and restoring of the fused sequence to obtain the phoneme representation vector of each first to-be-detected song includes: for a first to-be-detected song, the phoneme processing model takes position information of all phoneme data as a query index, generates a phoneme probability distribution of the position information corresponding to each phoneme data according to the attention score and the fused sequence, and takes a phoneme feature with the maximum probability as a phoneme feature of each phoneme data to constitute a phoneme representation vector of the first to-be-detected song.
[0008] Optionally, the phoneme processing model is trained by: obtaining a song training sample set, the song training sample set including at least one sample song and text sample data corresponding to each sample song; converting each text sample data into phoneme sample data, and performing random phoneme processing on the phoneme sample data of each sample song to obtain second phoneme data; causing the phoneme processing model to predict phoneme data of a processed part, and iteratively training at least one model parameter in the phoneme processing model according to a difference between a predicted phoneme vector and an actual phoneme vector of a processed phoneme in the second phoneme data.
[0009] Optionally, the phoneme processing model is trained by: obtaining a song training sample set, the song training sample set including at least one sample song and text sample data corresponding to each sample song; converting each text sample data into phoneme sample data, and performing random phoneme processing on the phoneme sample data of each sample song to obtain second phoneme data; causing the phoneme processing model to predict phoneme data of a processed part, and iteratively training at least one model parameter in the phoneme processing model according to a difference between a predicted phoneme vector and an actual phoneme vector of a processed phoneme in the second phoneme data.
[0010] Optionally, the phoneme sample data of each sample song is randomly processed to obtain second phoneme data, including: randomly selecting a plurality of candidate phonemes in the phoneme sample data according to a preset processing ratio; selecting a first preset proportion of covered phonemes and a second preset proportion of replaced phonemes from the plurality of candidate phonemes; covering the covered phonemes and replacing the replaced phonemes in the phoneme sample data of each sample song as the second phoneme data.
[0011] Optionally, the plurality of candidate phonemes further includes a third preset proportion of original phoneme data, the original phoneme data representing data without any processing, and the original phoneme data, the covered phonemes and the replaced phonemes together constitute the second phoneme data.
[0012] Optionally, for each first song category, the second to-be-detected songs are divided into at least one second song category based on the acoustic representation vectors corresponding to the second to-be-detected songs in the first song category, including: performing second similarity calculation on the acoustic representation vectors corresponding to each second to-be-detected song in the first song category to divide the second to-be-detected songs in the first song category into at least one second song category.
[0013] In a second aspect, the embodiments of the present application also provide a song duplicate detection device, the device comprising: an acquisition module, acquiring at least one first to-be-detected song and text data corresponding to the first to-be-detected song; a conversion module, converting the text data into a phoneme representation vector and encoding the first to-be-detected song into an acoustic representation vector; a first classification module, performing first similarity calculation on each phoneme representation vector and dividing the first to-be-detected song into at least one first song category; a second classification module, for each first song category, dividing the second to-be-detected songs into at least one second song category based on the acoustic representation vectors corresponding to the second to-be-detected songs in the first song category, and the songs belonging to the same second song category are duplicate songs.
[0014] The song detection method and device provided by the embodiment of the present application, the method comprises: obtaining at least one first to-be-detected song and text data corresponding to the first to-be-detected song; converting the text data into a phoneme representation vector, and encoding the first to-be-detected song into an acoustic representation vector; performing first similarity calculation on each phoneme representation vector, and dividing the first to-be-detected song into at least one first song category; for each first song category, based on the acoustic representation vector corresponding to the second to-be-detected song in the first song category, dividing the second to-be-detected song into at least one second song category, and the songs belonging to the same second song category are repeated songs. The present application converts the text data of the first to-be-detected song into a phoneme representation vector, and encodes the first to-be-detected song into an acoustic representation vector, so as to perform similarity calculation on the first to-be-detected song through the phoneme representation vector to divide it into a first song category, and then divide the second to-be-detected song in the first song category into a second song category according to the acoustic representation vector, so as to identify the songs in the second song category as repeated songs, thereby solving the technical problem of poor accuracy of repeated song identification in the prior art, and achieving the technical effect of increasing the accuracy of repeated song identification.
[0015] In order to make the above objectives, characteristics and advantages of the present application more apparent, clear and easy to understand, the following will specifically describe the preferred embodiments in conjunction with the accompanying drawings, and make a detailed description as follows. BRIEF DESCRIPTION OF DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments, and it should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as a limitation to the scope, and for those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.
[0017] Figure 1 A flow chart of a song repetition detection method provided by the embodiment of the present application is shown.
[0018] Figure 2 A flow chart of the steps of training a phoneme processing model provided by the embodiment of the present application is shown.
[0019] Figure 3 A functional module diagram of a song repetition detection device provided by the embodiment of the present application is shown.
[0020] Figure 4 A structural schematic diagram of an electronic device provided by the embodiment of the present application is shown. DETAILED DESCRIPTION
[0021] In order to make the purposes, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. It should be understood that the drawings in the present application only serve the purpose of illustration and description, and are not used to limit the scope of protection of the present application. In addition, it should be understood that the schematic drawings are not drawn according to the actual proportions. The flowcharts used in the present application show the operations implemented according to some embodiments of the present application. It should be understood that the operations of the flowcharts can not be implemented in sequence, and the steps without logical context relationship can be reversed in sequence or implemented simultaneously. In addition, one or more other operations can be added to the flowcharts or one or more operations can be removed from the flowcharts under the guidance of the content of the present application.
[0022] In addition, the described embodiments are only some of the embodiments of the present application, not all the embodiments. The components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0023] In the prior art, a music model is trained to enable the music model to automatically generate music, and a large number of complex and different quality songs need to be obtained as training data in the model training process. If a large number of repeated songs or different cover versions of the same song are contained in the training data, the repetition rate of the same song in the training data is too high, it is difficult to obtain diverse data in the training process, and the finally generated song is biased towards repeated songs. In addition, music copyright parties need to monitor unlicensed music on network platforms to protect the intellectual property rights of the copyright parties. Therefore, it is necessary to identify repeated songs in the training samples of the music model and identify pirated songs on music playing software. In the prior art, it is difficult to identify different versions of the same song through clustering analysis of the songs, and it is also difficult to identify the same song performed by different singers, resulting in a high error rate.
[0024] Based on this, the embodiment of the present application provides a song detection method and device, by converting the text data of the first to-be-detected song into a phoneme representation vector and performing song coding on the first to-be-detected song to obtain an acoustic representation vector, to divide the first to-be-detected song into a first song category by similarity calculation of the first to-be-detected song through the phoneme representation vector, and then divide the second to-be-detected song in the first song category into a second song category according to the acoustic representation vector, to identify the song in the second song category as a duplicate song, solving the technical problem of poor duplicate song identification accuracy in the prior art, and achieving the technical effect of increasing duplicate song identification accuracy. Specifically as follows: Please refer to Figure 1 , Figure 1 The flow chart of a song duplicate detection method provided by the embodiment of the present application. As shown in Figure 1 The song duplicate detection method provided by the embodiment of the present application includes the following steps: S101: Obtain at least one first to-be-detected song and text data corresponding to the first to-be-detected song.
[0025] The first to-be-detected song can be sample data for training a music model, or a song on a music playing software. For example, if a pirated song on a music playing software needs to be identified, the at least one first to-be-detected song can be obtained by searching for the song name on the music playing software. The present application does not limit the selection method of the first to-be-detected song.
[0026] Specifically, the text data corresponding to the first to-be-detected song is obtained by the following method: determining the pre-uploaded official lyrics text data corresponding to the first to-be-detected song as the text data of the first to-be-detected song; or determining the recognized lyrics text data extracted from the first to-be-detected song as the text data of the first to-be-detected song.
[0027] That is, the official lyrics text data corresponding to the first to-be-detected song pre-uploaded to the music playing software can be directly used as the text data, or the lyrics recognition is performed on the first to-be-detected song to use the recognized lyrics text data as the text data.
[0028] Exemplarily, the text data should at least include the lyrics text of the song. The lyrics text can be officially uploaded synchronously when the first to-be-detected song is uploaded, or obtained by recognizing the first to-be-detected song. Before obtaining the text data corresponding to the first to-be-detected song, it is determined whether the official lyrics text data of the first to-be-detected song is uploaded synchronously when the first to-be-detected song is uploaded. If the official lyrics text data of the first to-be-detected song is uploaded synchronously, the official lyrics text data is taken as the text data. If the official lyrics text data of the first to-be-detected song is not uploaded synchronously, text recognition is performed on the first to-be-detected song, so as to take the recognized lyrics text data as the text data. Therefore, the text data of the first to-be-detected song is the official lyrics text data or the recognized lyrics text data.
[0029] Further, since the official lyrics text data is officially uploaded by the official, the accuracy of the official lyrics text data is high. However, the official lyrics text data can contain data irrelevant to the lyrics, such as the lyricist and the composer, which affects subsequent processing. The recognized lyrics text data can be extracted from the first to-be-detected song by a speech-to-text (STT) model. The recognized lyrics text data does not contain data irrelevant to the lyrics, such as the lyricist and the composer. However, the extraction accuracy of the recognized lyrics text data depends on the audio quality, the clarity of the singer's voice, etc., and errors in recognition are prone to occur, such as recognition of similar or identical characters.
[0030] Therefore, in order to prevent lyrics recognition errors, the present application selects to convert the text data into phoneme data and then determine the phoneme representation vector. Further, by obtaining the text data of each first to-be-detected song, the lyrics of each first to-be-detected song are obtained, so as to facilitate subsequent division of the song from the phoneme angle of the lyrics.
[0031] Exemplarily, before the first to-be-detected song is extracted by the speech-to-text model to obtain the recognized lyrics text data, the audio file of the first to-be-detected song can also be preprocessed, such as noise reduction, music standardization, etc., so as to increase the accuracy of recognition and improve the recognition efficiency.
[0032] S102: convert the text data into a phoneme representation vector, and encode the first to-be-detected song into an acoustic representation vector.
[0033] Specifically, the converting the text data into the phoneme representation vector comprises: converting the text data of each first to-be-detected song into a plurality of phoneme data; obtaining a phoneme processing model, mapping the phoneme data to a preset dimensional space, for each phoneme data, performing encoding processing on the phoneme data according to context information and phoneme dependency relationship of the phoneme data, and summarizing the encoding results of all phoneme data of each first to-be-detected song to obtain at least one phoneme embedding representation sequence of each first to-be-detected song; and performing fuzzy processing on each phoneme embedding representation sequence to obtain a phoneme representation vector of each first to-be-detected song.
[0034] Specifically, the converting the text data into the phoneme representation vector comprises: converting the text data of each first to-be-detected song into a plurality of phoneme data; obtaining a phoneme processing model, mapping the phoneme data to a preset dimensional space, for each phoneme data, performing encoding processing on the phoneme data according to context information and phoneme dependency relationship of the phoneme data, and summarizing the encoding results of all phoneme data of each first to-be-detected song to obtain at least one phoneme embedding representation sequence of each first to-be-detected song; and performing fuzzy processing on each phoneme embedding representation sequence to obtain a phoneme representation vector of each first to-be-detected song.
[0035] Specifically, the converting the text data into the phoneme representation vector comprises: converting the text data of each first to-be-detected song into a plurality of phoneme data; obtaining a phoneme processing model, mapping the phoneme data to a preset dimensional space, for each phoneme data, performing encoding processing on the phoneme data according to context information and phoneme dependency relationship of the phoneme data, and summarizing the encoding results of all phoneme data of each first to-be-detected song to obtain at least one phoneme embedding representation sequence of each first to-be-detected song; and performing fuzzy processing on each phoneme embedding representation sequence to obtain a phoneme representation vector of each first to-be-detected song.
[0036] Specifically, the converting the text data into the phoneme representation vector comprises: converting the text data of each first to-be-detected song into a plurality of phoneme data; obtaining a phoneme processing model, mapping the phoneme data to a preset dimensional space, for each phoneme data, performing encoding processing on the phoneme data according to context information and phoneme dependency relationship of the phoneme data, and summarizing the encoding results of all phoneme data of each first to-be-detected song to obtain at least one phoneme embedding representation sequence of each first to-be-detected song; and performing fuzzy processing on each phoneme embedding representation sequence to obtain a phoneme representation vector of each first to-be-detected song.
[0037] The pre-trained phoneme processing model includes at least a phoneme embedding layer, a context encoder, and an attention module. For each first to-be-detected song, each phoneme data of the first to-be-detected song is input into the phoneme embedding layer, and the phoneme embedding layer maps each phoneme data in a preset dimension space to obtain an initial phoneme vector of the first to-be-detected song. The preset dimension space is a fixed low-dimensional vector space, which converts discrete phonemes in phoneme data into continuous and learnable low-dimensional vectors, and phonemes with similar pronunciation are close to each other. The initial phoneme vector of the first to-be-detected song is input into the context encoder, and the context information of each phoneme data corresponding to the position in the initial phoneme vector is encoded to determine the phoneme embedding representation sequence of each first to-be-detected song, which covers the context information and long-distance dependence of each phoneme data. Finally, the phoneme embedding representation sequence is input into the attention module, and the attention mechanism calculation is performed through the attention module to compress the phoneme embedding representation sequence into a phoneme representation vector covering global context information.
[0038] The attention mechanism calculation is performed on each phoneme embedding representation sequence to obtain an attention score corresponding to each phoneme embedding representation sequence. The attention scores of each phoneme embedding representation sequence are normalized by a normalization layer to obtain an attention weight corresponding to each phoneme embedding representation sequence. The weighted average of each phoneme embedding representation sequence is calculated according to the attention weight of each phoneme embedding representation sequence to obtain a fusion sequence. The fusion sequence is decoded to obtain a phoneme representation vector of a first to-be-detected song.
[0039] Specifically, the fusion sequence is decoded to obtain a phoneme representation vector of each first to-be-detected song, including: for a first to-be-detected song, the phoneme processing model takes the position information of all phoneme data as a query index, generates a phoneme probability distribution of the position information corresponding to each phoneme data according to the attention score and the fusion sequence, and takes the phoneme feature with the maximum probability as the phoneme feature of each phoneme data to form a phoneme representation vector of the first to-be-detected song.
[0040] That is, for each first to-be-detected song, the position encoding or the phoneme order of each phoneme data of the first to-be-detected song is taken as the position information of each phoneme data, the fusion sequence is input into a full connection layer, the phoneme probability distribution under each position information is determined, and the phoneme embedding representation sequence corresponding to the phoneme with the maximum probability is taken as the final representation of the position information. In this way, the final representations of each phoneme data are fused into a phoneme representation vector of the first to-be-detected song.
[0041] Please refer to Figure 2 , Figure 2A flowchart of the steps of training the phoneme processing model provided by the embodiments of the present application. As shown in Figure 2 The embodiments of the present application train the phoneme processing model in the following manner: S201: Obtain a set of song training samples.
[0042] S202: Convert each text sample data into phoneme sample data, and randomly process the phoneme sample data of each sample song to obtain second phoneme data.
[0043] S203: Let the phoneme processing model predict the phoneme data of the processed part, and iteratively train at least one model parameter in the phoneme processing model according to the difference between the predicted phoneme vector and the actual phoneme vector of the processed phoneme in the second phoneme data.
[0044] The set of song training samples includes at least one sample song and text sample data corresponding to each sample song. The text sample data corresponding to the sample song refers to the official lyrics text data or the recognized lyrics text data mentioned in the foregoing. That is, if the official lyrics text data of the sample song is pre-uploaded, the official lyrics text data of the sample song is taken as the text sample data of the sample song; if the official lyrics text data of the sample song is not pre-uploaded, the recognized lyrics text data is recognized from the sample song by the speech-to-text model, and the recognized lyrics text data of the sample song is taken as the text sample data of the sample song.
[0045] In this way, the text sample data of the sample song is input into the G2P model to obtain the phoneme data of the sample song, and each phoneme in the phoneme data of the sample song is randomly covered to obtain covered phoneme sample data.
[0046] The random phoneme processing of the phoneme sample data of each sample song means that the second phoneme data to be processed is randomly extracted from all the phoneme sample data of the sample song. Moreover, the randomly extracted second phoneme data is replaced or covered, so as to predict the second phoneme data to obtain a predicted phoneme vector, and train the phoneme processing model by calculating the difference between the actual phoneme vector and the predicted phoneme vector of the covered phoneme and the replaced phoneme.
[0047] Specifically, the random phoneme processing of the phoneme sample data of each sample song to obtain the second phoneme data includes: randomly selecting a plurality of candidate phonemes in the phoneme sample data according to a preset processing ratio; selecting a first preset proportion of covered phonemes and a second preset proportion of replaced phonemes from the plurality of candidate phonemes; covering the covered phonemes in the phoneme sample data of each sample song, and replacing the replaced phonemes, as the second phoneme data.
[0048] The plurality of candidate phonemes further includes third preset proportion of original phoneme data, the original phoneme data represents data without any processing on the phoneme data, and together with the covered phoneme and the replaced phoneme, forms the second phoneme data.
[0049] That is, in the training process of the phoneme processing model, a plurality of candidate phonemes of a preset processing proportion are randomly selected from the phoneme sample data of each sample song, the covered phoneme and the replaced phoneme are screened out from the plurality of candidate phonemes, and the original phoneme data other than the covered phoneme and the replaced phoneme in the plurality of candidate phonemes is kept unchanged, so as to take the covered covered phoneme, the replaced replaced phoneme and the original phoneme data as the second phoneme data after random phoneme processing.
[0050] Further, the first preset proportion of candidate phonemes is randomly selected from the plurality of candidate phonemes of the preset processing proportion as the covered phoneme and the second preset proportion of candidate phonemes as the replaced phoneme. For example, the preset processing proportion can be controlled between 10% and 15%, the first preset proportion can be selected as 80%, and the second preset proportion can be selected as 10%, so that there are 10% of the candidate phonemes remaining in the plurality of candidate phonemes of the preset processing proportion, thereby taking the phonemes outside the plurality of candidate phonemes of the preset processing proportion and the remaining candidate phonemes in the plurality of candidate phonemes as the original phoneme data kept unchanged.
[0051] That is, the product of the preset processing proportion and the first preset proportion is taken as the covering proportion of the covered phoneme, and the product of the preset processing proportion and the second preset proportion is taken as the replacement proportion of the replaced phoneme. Moreover, the covering proportion and the replacement proportion should not be too high, and too high covering proportion and replacement proportion can easily lead to model prediction difficulty, so the preset processing proportion should not be too high. The covered phoneme can be covered by a special symbol [mask], and the phoneme processing model should predict the covered phoneme during the training process, and guess whether each phoneme not marked with [mask] is a transformed phoneme, so as to enhance the generalization ability of the model.
[0052] Further, the processed phoneme should include the covered phoneme and the replaced phoneme in the second phoneme data, and the processed phoneme does not include the original phoneme data since the original phoneme data is not covered or replaced. Therefore, the phoneme processing model should train the ability to recognize the replaced phoneme during the training process, and predict the original phoneme sample data corresponding to the replaced phoneme.
[0053] Specifically, the phoneme processing model is caused to predict phoneme data of the processed part, and at least one model parameter in the phoneme processing model is iteratively trained according to a difference between a predicted phoneme vector and an actual phoneme vector of a processed phoneme in the second phoneme data, including: mapping each phoneme in the second phoneme data to a preset dimensional space to determine an initial phoneme vector corresponding to each sample song; predicting the predicted phoneme vector according to the initial phoneme vector; establishing a first loss function between the predicted phoneme vector and the actual phoneme vector; adjusting the model parameter of the phoneme processing model according to the loss value of the first loss function until the phoneme processing model meets a preset training completion condition.
[0054] That is, after the phoneme sample data of each sample song is randomly processed to obtain the second phoneme data, each phoneme in the second phoneme data is mapped to a preset dimensional space to obtain an initial phoneme vector. And in order for the model to correctly predict the covered phonemes and the replaced phonemes subsequently, the phoneme embedding layer needs to make the phonemes with similar pronunciation closer in the vector space, and thus must learn the vector representation that can effectively distinguish different phonemes. At the same time, the mapping to the vector space also needs to contain sufficient semantic and pronunciation information so that the subsequent context encoder can understand.
[0055] Further, the initial phoneme vector of each sample song output by the phoneme embedding layer is input to the context encoder, and the context editor generates a hidden sequence representation of each sample song according to the context information before and after each phoneme in the initial phoneme vector of each sample song. The hidden sequence representation includes the hidden state of the position of each phoneme in the sample song, and the context editor learns long-distance dependence in the training process, so that the hidden sequence representation carries sufficient context information to facilitate subsequent inference of the processed phonemes. Therefore, the context editor needs to learn how to capture long-distance context information from unmasked phonemes in the training process to infer the processed phonemes through these context information. A good context encoder can generate high-quality hidden sequence representation containing rich semantics.
[0056] Among them, the phoneme processing model further includes a decoder and an output layer. Exemplarily, the decoder and the output layer are used in the training process, responsible for generating a predicted phoneme vector of the processed phoneme, and mapping the predicted phoneme vector into a phoneme probability distribution to calculate a cross-entropy loss, so as to modify each model parameter of the model through back propagation.
[0057] Further, the implicit sequence representation of each sample song is input into an attention module, and the attention module calculates a global context vector corresponding to each implicit sequence representation according to attention mechanism weighting aggregation. And the attention module should gradually modify the model parameters during the training process, learning to macroscopically focus on the key phonemes that are most helpful to prediction accuracy from the context information corresponding to the phonemes. That is, higher weights are given to the parts that are iconic and play a core role in identifying the pronunciation of the entire song, and lower weights are given to non-core parts, and fuzzy matching is achieved through attention weighting. Then the global context vector is input into the decoder for decoding, and the decoder maps the global context vector to the vector space of the dimension of the predicted phoneme vector to gradually generate the predicted phoneme vector of the covered phoneme. The output layer calculates the loss value by using a loss function that supports phoneme fuzzy matching, such as cross-entropy loss, between the predicted phoneme vector and the actual phoneme vector, so as to adjust the parameters of each module in reverse order through the back propagation mechanism, and stop training the phoneme processing model when the loss value meets the preset training completion condition, and obtain the trained phoneme processing model.
[0058] Wherein, adjusting the model parameters of the phoneme processing model through the back propagation mechanism refers to all learnable parameters of the model. All learnable parameters of the model include the vector space of the phoneme embedding layer, the weights of the context encoder, the weight matrix for calculating the attention score, the decoder weights, the output layer weights, etc. For example, by modifying the vector space to pull the phonemes in the correct direction, i.e., to pull similar phonemes closer and to pull dissimilar phonemes farther. That is, when the prediction is wrong, a larger loss value will be calculated, which is propagated to each module through the back propagation mechanism, and the predicted phoneme vector is gradually closer to the actual phoneme vector in the training process by adjusting the model parameters of the phoneme processing model. Thus, in an optimization cycle, each module in the entire model learns together to generate high-quality lyric phoneme representation vectors that can realize context awareness and fuzzy training.
[0059] Return Figure 1 S103: Perform first similarity calculation on each phoneme representation vector, and divide the first to-be-detected songs into at least one first song category.
[0060] Wherein, the first similarity calculation is performed by calculating the distance between the phoneme representation vectors of each two first to-be-detected songs to determine the similarity matrix of all first to-be-detected songs; and dividing all first to-be-detected songs into at least one first song category according to the similarity matrix.
[0061] That is, for all the first to-be-detected songs, the distance between each two first to-be-detected songs is calculated according to the phoneme representation vectors, so as to form a similarity matrix (N x N, N refers to the total number of the first to-be-detected songs). For example, the distance between each two first to-be-detected songs can be calculated by cosine similarity, and the present application does not limit the distance algorithm. Thus, the first to-be-detected songs are divided into at least one song category according to the similarity matrix, and hierarchical clustering or density clustering (DBSCAN, Density-Based Spatial Clustering of Applications with Noise) can be selected for clustering, and the present application does not limit the clustering algorithm.
[0062] Thus, the division of all the first to-be-detected songs by the phoneme representation vectors can divide all the first to-be-detected songs into at least one first song category from the lyrics. Each first song category includes at least one second to-be-detected song.
[0063] S104: For each first song category, based on the acoustic representation vectors corresponding to the second to-be-detected songs in the first song category, the second to-be-detected songs are divided into at least one second song category, and the songs belonging to the same second song category are duplicate songs.
[0064] For each first song category, based on the acoustic representation vectors corresponding to the second to-be-detected songs in the first song category, the second to-be-detected songs are divided into at least one second song category, including: performing second similarity calculation on the acoustic representation vectors corresponding to each second to-be-detected song of the first song category, so as to divide the second to-be-detected songs in the first song category into at least one second song category.
[0065] Specifically, the first song category is divided into at least one second song category by the following method: the distance between each two second to-be-detected songs in the first song category is calculated according to the acoustic representation vectors of the two songs, so as to determine the similarity matrix of the second to-be-detected songs in the first song category; and the first song category is divided according to the similarity matrix of the second to-be-detected songs.
[0066] That is, after dividing all the first to-be-detected songs into at least one first song category, for each first song category, the distance between each two second to-be-detected songs in the first song category is calculated according to the acoustic representation vector of each second to-be-detected song in the first song category to obtain a similarity matrix of the first song category, and then all the second to-be-detected songs in the first song category are divided by the similarity matrix of the first song category to divide the first song category into at least one second song category.
[0067] For example, the cosine similarity can be used to calculate the distance between the acoustic representation vectors of each two second to-be-detected songs, and hierarchical clustering or density clustering can be selected for the second division, and the application does not limit the distance algorithm and clustering algorithm of the second clustering.
[0068] Further, each second song category includes at least one third to-be-detected song, and if a second song category contains multiple third to-be-detected songs, the multiple third to-be-detected songs belonging to the same second song category are repeated songs. Thus, after the first division from the lyrics, the second division from the music is performed on the result of the first division, instead of directly performing the music-level grouping on all the first to-be-detected songs, which can reduce the calculation amount and improve the calculation efficiency and robustness of the clustering result.
[0069] For example, the acoustic representation vector of each first to-be-detected song can be generated according to fixed-dimensional music features, including rhythm, rhythm, timbre, voiceprint, chord, etc. Normalization processing can be performed after identifying each music feature to obtain the acoustic representation vector of the first to-be-detected song. Generally, each music feature can be measured by calculating the statistical quantity of each music feature, such as mean, variance, peak, valley, peak-valley difference, etc., and the application does not limit this.
[0070] For example, before determining the acoustic representation vector of the first to-be-detected song, the first to-be-detected song can also be preprocessed to improve the audio quality and increase the accuracy of music feature extraction.
[0071] Further, the technical solution of the application can perform deduplication processing on the training samples of the music model, and can also identify pirate songs on the music playing software, solving the technical problem of poor accuracy of repeated song identification in the prior art, and achieving the technical effect of increasing the accuracy of repeated song identification.
[0072] Based on the same application concept, the embodiment of the present application also provides a song repetition detection device corresponding to the song repetition detection method provided by the above-embodiment. Since the principle of solving problems in the device of the embodiment of the present application is similar to the song repetition detection method of the above-embodiment of the present application, the implementation of the device can be referred to the implementation of the method, and the repeated parts will not be described here.
[0073] As shown in Figure 3 , Figure 3 A functional module diagram of a song repetition detection device provided by the embodiment of the present application. The song repetition detection device 10 comprises: an acquisition module 101, acquiring at least one first to-be-detected song and text data corresponding to the first to-be-detected song; a conversion module 102, converting the text data into a phoneme representation vector and encoding the first to-be-detected song into an acoustic representation vector; a first classification module 103, performing a first similarity calculation on each phoneme representation vector and dividing the first to-be-detected song into at least one first song category; a second classification module 104, for each first song category, based on the acoustic representation vector corresponding to the second to-be-detected song in the first song category, dividing the second to-be-detected song into at least one second song category, and the songs belonging to the same second song category are repeated songs.
[0074] Based on the same application concept, referring to Figure 4 , a structural schematic diagram of an electronic device provided by the embodiment of the present application, the electronic device 20 comprises: a processor 201, a memory 202 and a bus 203, the memory 202 stores machine readable instructions executable by the processor 201, when the electronic device 20 runs, the processor 201 and the memory 202 communicate through the bus 203, the machine readable instructions are executed by the processor 201 to execute the steps of the song repetition detection method as described in any of the above embodiments.
[0075] Specifically, the machine readable instructions executed by the processor 201 can perform the following processing: acquiring at least one first to-be-detected song and text data corresponding to the first to-be-detected song; converting the text data into a phoneme representation vector and encoding the first to-be-detected song into an acoustic representation vector; performing a first similarity calculation on each phoneme representation vector and dividing the first to-be-detected song into at least one first song category; for each first song category, based on the acoustic representation vector corresponding to the second to-be-detected song in the first song category, dividing the second to-be-detected song into at least one second song category, and the songs belonging to the same second song category are repeated songs.
[0076] Based on the same application concept, the embodiment of the present application also provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is run by a processor to execute the steps of the song repetition detection method provided in the above embodiment.
[0077] Specifically, the storage medium can be a general storage medium, such as a mobile disk, a hard disk, etc., and the computer program stored on the storage medium can be run to execute the above song repetition detection method. The text data of the first to-be-detected song is converted into a phoneme representation vector, and acoustic representation vectors are obtained by performing song coding on the first to-be-detected song, so as to divide the first to-be-detected song into a first song category by performing similarity calculation on the first to-be-detected song through the phoneme representation vector, and then divide the second to-be-detected song in the first song category into a second song category according to the acoustic representation vector, so as to identify the songs in the second song category as repeated songs. The technical problem of poor repetition song identification accuracy in the prior art is solved, and the technical effect of increasing the repetition song identification accuracy is achieved.
[0078] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system and device can refer to the corresponding process in the foregoing method embodiment, which will not be described here. In the several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented by other ways. The device embodiments described above are only schematic, for example, the division of the units is only a logical function division, and there can be another division manner in actual implementation, and for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some communication interfaces, devices or units, and can be electrical, mechanical or other forms.
[0079] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0080] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit.
[0081] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a nonvolatile computer readable storage medium executable by a processor. Based on this understanding, the technical solutions of the present application or the parts of the present application that essentially contribute to the prior art or the parts of the technical solutions can be embodied in the form of software products. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0082] The above merely describes the specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A song detection method characterized by comprising: The method comprises: acquiring at least one first to-be-detected song and text data corresponding to the first to-be-detected song; converting the text data into a phoneme representation vector and encoding the first to-be-detected song into an acoustic representation vector; performing first similarity calculation on each phoneme representation vector and dividing the first to-be-detected song into at least one first song category; for each first song category, dividing the second to-be-detected song into at least one second song category based on the acoustic representation vector corresponding to the second to-be-detected song in the first song category, and the songs belonging to the same second song category are repeated songs.
2. The method of claim 1, wherein, The conversion of the text data into a phoneme representation vector comprises: converting the text data of each first to-be-detected song into phoneme data; acquiring a phoneme processing model, mapping the phoneme data to a preset dimensional space, for each phoneme data, performing encoding processing on the phoneme data according to the context information and phoneme dependency relationship of the phoneme data, and summarizing the encoding results of all phoneme data of each first to-be-detected song to obtain at least one phoneme embedding representation sequence of each first to-be-detected song; performing fuzzy processing on each phoneme embedding representation sequence to obtain a phoneme representation vector of each first to-be-detected song.
3. The method of claim 2, wherein, The fuzzy processing on each phoneme embedding representation sequence to obtain a phoneme representation vector of each first to-be-detected song comprises: performing attention mechanism calculation on the at least one phoneme embedding representation sequence, calculating the corresponding attention weight according to the attention score of each phoneme embedding representation sequence, performing fusion of the phoneme embedding representation sequence according to the attention weight, and decoding the fused sequence to obtain the phoneme representation vector of each first to-be-detected song.
4. The method of claim 3, wherein, The decoding of the fused sequence to obtain the phoneme representation vector of each first to-be-detected song comprises: for a first to-be-detected song, the phoneme processing model takes the position information of all phoneme data as a query index, generates a phoneme probability distribution of the position information corresponding to each phoneme data according to the attention score and the fused sequence, and takes the phoneme embedding representation sequence corresponding to the phoneme with the maximum probability as the phoneme feature of each phoneme data to constitute a phoneme representation vector of a first to-be-detected song.
5. The method of claim 2, wherein, The phoneme processing model is trained in the following manner: acquiring a song training sample set, the song training sample set comprising at least one sample song and text sample data corresponding to each sample song; converting each text sample data into phoneme sample data and performing random phoneme processing on the phoneme sample data of each sample song to obtain second phoneme data; making the phoneme processing model predict the phoneme data of the processed part, and iteratively training at least one model parameter in the phoneme processing model according to the difference between the predicted phoneme vector and the actual phoneme vector of the processed phoneme in the second phoneme data.
6. The method of claim 5, wherein, making the phoneme processing model predict the phoneme data of the processed part, and iteratively training at least one model parameter in the phoneme processing model according to the difference between the predicted phoneme vector and the actual phoneme vector of the processed phoneme in the second phoneme data comprises: Map each phoneme in the second phoneme data to a preset dimensional space to determine an initial phoneme vector corresponding to each sample song; According to the initial phoneme vector, a predicted phoneme vector is predicted; A first loss function is established between the predicted phoneme vector and the actual phoneme vector; The model parameters of the phoneme processing model are adjusted according to the loss value of the first loss function until the phoneme processing model meets a preset training completion condition.
7. The method of claim 5, wherein, Random phoneme processing is performed on the phoneme sample data of each sample song to obtain second phoneme data, including: Randomly selecting a plurality of candidate phonemes in the phoneme sample data according to a preset processing ratio; Selecting a first preset proportion of covered phonemes and a second preset proportion of replaced phonemes from the plurality of candidate phonemes; Covering the covered phonemes in the phoneme sample data of each sample song and replacing the replaced phonemes as the second phoneme data.
8. The method of claim 7, wherein, The plurality of candidate phonemes also include a third preset proportion of original phoneme data, which represents data that is not processed, and together with the covered phonemes and the replaced phonemes, forms the second phoneme data.
9. The method of claim 1, wherein, For each first song category, based on the acoustic representation vector corresponding to the second to-be-detected song in the first song category, the second to-be-detected song is divided into at least one second song category, including: Performing a second similarity calculation on the acoustic representation vector corresponding to each second to-be-detected song in the first song category to divide the second to-be-detected song in the first song category into at least one second song category.
10. A song duplicate detection apparatus characterized by comprising: The device comprises: An acquisition module acquires at least one first to-be-detected song and text data corresponding to the first to-be-detected song; A conversion module converts the text data into a phoneme representation vector and encodes the first to-be-detected song into an acoustic representation vector; A first classification module performs a first similarity calculation on each phoneme representation vector and divides the first to-be-detected song into at least one first song category; A second classification module divides the second to-be-detected song into at least one second song category based on the acoustic representation vector corresponding to the second to-be-detected song in the first song category for each first song category, and the songs belonging to the same second song category are repeated songs.
Citation Information
Patent Citations
Repetitive audio detection method, apparatus and device, and storage medium
CN111583963A
Video detection method and device and computer readable storage medium
CN113407779A
Song synthesis method and device, equipment, medium and product
CN113808555A
Synthetic speech processing by representing text by phonemes exhibiting predicted volume and pitch using neural networks
US11978431B1
System And Method For Podcast Repetitive Content Detection
US20220115029A1