Piracy Singer Detection Method, Computer Device, and Computer Storage Medium
By calculating the similarity of the tone feature vectors of the songs under the name of the singer, we will automatically identify pirated singers, solving the problem of inefficient manual recognition and achieving fast and efficient singer recognition.
Patent Information
- Application Number
- CN202211501200.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-28
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-11-28
AI Technical Summary
In the prior art, manual review and identification of pirated singers is inefficient and cannot quickly process massive and incremental singer data, resulting in the incorrect investment of resources and funds.
By obtaining the spectral charts of multiple songs under the name of the target singer, the tone feature extraction model is used to calculate the average value of the tone feature vector similarity between songs, and automatically identify whether the singer is a pirated singer.
It realizes the rapid and efficient identification of pirated singers, improves the singer recognition efficiency, and reduces the time and waste of resources for manual review.
Smart Images

Figure CN115762454B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of audio processing, and specifically to a method for detecting pirated singers, a computer device, and a computer storage medium. Background Art
[0002] Pirated singers are singers whose songs are mostly pirated. With the rise of short videos in recent years, a large number of pirated songs have been added, which continue to pollute the music library, showing a large-scale and organized piracy trend, which seriously affects user reputation. In particular, short video platforms need to invest a lot of resources and funds to support musicians. If they cannot effectively identify whether singers are pirated singers, it is very likely that the platform will mistakenly invest resources and funds in supporting pirated singers, which seriously damages the interests of the platform.
[0003] In the relevant scheme, manual review is mainly used to manually review and identify whether the singer is a pirated singer. However, this method consumes a huge amount of manpower and has extremely low recognition efficiency. It cannot cope with the massive amount of existing singer data and daily incremental singer data. Summary of the invention
[0004] The embodiments of the present application provide a method for detecting pirated singers, a computer device, and a computer storage medium, which are used to detect and identify pirated singers to improve the efficiency of identifying pirated singers.
[0005] The first aspect of the embodiment of the present application provides a method for detecting pirated singers, the method comprising:
[0006] Obtain N target songs under the name of the same target singer, and process each of the target songs to obtain a spectrogram of each of the target songs, wherein N is a positive integer greater than 1;
[0007] Inputting the spectrogram of the target song into a target timbre feature extraction model to obtain a timbre feature vector of the target song output by the target timbre feature extraction model;
[0008] For each target song, taking the average of the similarities between the target song and the timbre feature vectors of each other target song as the average similarity of the target song;
[0009] If the average similarity of the N target songs meets the preset characteristics, the target singer is determined to be the original singer;
[0010] If the average similarity of the N target songs does not meet the preset characteristics, the target singer is determined to be a pirated singer.
[0011] A second aspect of an embodiment of the present application provides a computer device, the computer device comprising:
[0012] An acquisition unit, configured to acquire N target songs under the same target singer, and process each of the target songs to obtain a spectrogram of each of the target songs, where N is a positive integer greater than 1;
[0013] A feature extraction unit, configured to input the spectrogram of the target song into a target timbre feature extraction model to obtain a timbre feature vector of the target song output by the target timbre feature extraction model;
[0014] A calculation unit, configured to, for each of the target songs, take the average value of the similarities between the timbre feature vector of the target song and the timbre feature vectors of other target songs as the average similarity value of the target song;
[0015] A detection unit, configured to determine that the target singer is an original singer if the average similarity value of the N target songs meets a preset feature;
[0016] The detection unit is further configured to determine that the target singer is a pirated singer if the average similarity value of the N target songs does not meet the preset feature.
[0017] A third aspect of the embodiments of the present application provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the method in the foregoing first aspect is implemented.
[0018] A fourth aspect of the embodiments of the present application provides a computer storage medium, where instructions are stored in the computer storage medium, and when the instructions are executed on a computer, the computer is caused to execute the method in the foregoing first aspect.
[0019] As can be seen from the above technical solutions, the embodiments of the present application have the following advantages:
[0020] In this embodiment, spectrograms of each target song are obtained by processing N target songs under the target singer, and a target timbre feature extraction model is used to extract timbre features from the spectrograms of each target song to obtain timbre feature vectors of each target song. For each target song, the average value of the similarities between the timbre feature vector of the target song and the timbre feature vectors of other target songs is taken as the average similarity value of the target song, and it is determined whether the target singer is a pirated singer according to the similarities of the timbre feature vectors of each target song. Therefore, the computer device automatically identifies and detects pirated singers based on the timbre features of the songs, solves the problem of low recognition efficiency caused by manual identification of pirated singers, can quickly and efficiently process a large amount of singer recognition work, and improves the recognition efficiency of pirated singers. Description of the Drawings
[0021] Figure 1It is a schematic flowchart of a pirated singer detection method in an embodiment of the present application;
[0022] Figure 2 It is another schematic flowchart of a pirated singer detection method in an embodiment of the present application;
[0023] Figure 3 It is a schematic diagram of a display effect of the distribution state of coordinate points corresponding to each song under the name of an original singer in an embodiment of the present application;
[0024] Figure 4 It is a schematic diagram of a display effect of the distribution state of coordinate points corresponding to each song under the name of a pirated singer in an embodiment of the present application;
[0025] Figure 5 It is a schematic structural diagram of a computer device in an embodiment of the present application;
[0026] Figure 6 It is another schematic structural diagram of a computer device in an embodiment of the present application. Detailed implementation manners
[0027] The embodiments of the present application provide a pirated singer detection method, a computer device, and a computer storage medium, which are used to detect and identify pirated singers to improve the identification efficiency of pirated singers.
[0028] In the scenarios of related solutions, with the wide application of the Internet, the dissemination of audiovisual information such as videos and songs is becoming more and more rapid. Among them, a large number of pirated songs are also generated. Pirated songs continuously contaminate the video library and music library, showing a large-scale and organized pirated trend, which seriously affects the user reputation.
[0029] A singer whose songs are mostly pirated songs is called a pirated singer. For example, songs such as Song 1, Song 2, and Song 3 are all original songs under the name of singer A, and songs such as Song 4, Song 5, and Song 6 are all original songs under the name of singer B. However, singer C puts the original singers of multiple songs under the names of singer A and singer B under his own name, making the public mistakenly think that the original singers of songs 1 to 6 are all singer C. That is, most of the songs under the name of singer C are pirated, and singer C is a pirated singer.
[0030] Internet platforms such as music platforms or video platforms need to detect and identify pirated singers in the platform. Because, the above-mentioned Internet platforms need to invest a large amount of resources and funds in the support of musicians. If they cannot effectively identify whether a singer is a pirated singer, it is very likely that the platform will misinvest resources and funds in the support of pirated singers, seriously damaging the interests of the platform.
[0031] At present, the detection and identification of pirated singers on Internet platforms is usually done manually, that is, manually reviewing and identifying whether a singer is a pirated singer. The reviewer identifies whether the songs under the singer's name are pirated songs and whether the singer is a pirated singer based on the song copyright information, his own listening experience or other reliable information. However, this manual review method requires personnel to spend a lot of time to obtain information from multiple parties in order to accurately identify pirated singers and pirated songs based on the information obtained. Obviously, it is impossible to achieve fast and efficient identification. When faced with massive stock singer data and daily incremental singer data, it is even more impossible to quickly and efficiently handle the identification of a large number of singers.
[0032] In response to the technical problems and defects existing in the above-mentioned scenarios, this application proposes a method for detecting pirated singers, which can solve the problem of slow and inefficient manual identification of pirated singers in the above-mentioned scenarios, and can improve the detection and identification efficiency of pirated singers.
[0033] The following is a description of the pirated singer detection method in the embodiment of the present application:
[0034] See also Figure 1 In the embodiment of the present application, one embodiment of the method for detecting pirated singers includes:
[0035] 101. Obtain N target songs under the name of the same target singer, and process each of the target songs to obtain a spectrogram of each of the target songs, wherein N is a positive integer greater than 1;
[0036] The method of this embodiment can be applied to a computer device, which can be a terminal device or a server device, etc. When the computer device is a terminal, it can be a terminal device such as a personal computer (PC) or a desktop computer; when the computer device is a server, it can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud databases, cloud services, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0037] The computer device can obtain N target songs under the name of the target singer to be detected, and process each target song to obtain a spectrogram of each target song. The spectrogram is used to extract the timbre feature vector corresponding to the target song in subsequent steps.
[0038] 102. Inputting the spectrogram of the target song into a target timbre feature extraction model to obtain a timbre feature vector of the target song output by the target timbre feature extraction model;
[0039] The computer device can deploy a pre-trained target timbre feature extraction model to use this model to extract the timbre feature vectors of each target song. That is, the computer device can input the spectrograms of each target song into the target timbre feature extraction model. The target timbre feature extraction model performs feature extraction on the spectrograms of each target song and outputs the timbre feature vectors of each target song when the feature extraction is completed. The timbre feature vectors can represent the timbre features of each target song.
[0040] 103. For each of the target songs, take the average of the similarities between the timbre feature vectors of the target song and the timbre feature vectors of each of the other target songs as the average similarity of the target song.
[0041] After obtaining the timbre feature vectors of each target song, the similarities between the timbre feature vectors of any two of the N target songs of the target singer can be calculated, and the average of the similarities between the timbre feature vectors of each target song and the timbre feature vectors of the other target songs among the N target songs of the target singer can be calculated. This average is used as the basis for determining whether the target singer is a pirated singer.
[0042] 104. If the average similarity of the N target songs meets the preset characteristics, determine that the target singer is an original singer.
[0043] When the average similarity corresponding to each target song meets the preset characteristics, determine that the target singer is an original singer.
[0044] 105. If the average similarity of the N target songs does not meet the preset characteristics, determine that the target singer is a pirated singer.
[0045] If the average similarity of each target song does not meet the preset characteristics, determine that the target singer is a pirated singer.
[0046] In this embodiment, the spectrograms of each target song are obtained by processing the N target songs under the name of the target singer. The target timbre feature extraction model is used to perform timbre feature extraction on the spectrograms of each target song to obtain the timbre feature vectors of each target song. For each target song, the average of the similarities between the timbre feature vectors of the target song and the timbre feature vectors of each of the other target songs is used as the average similarity of the target song. Whether the target singer is a pirated singer is determined according to the similarities of the timbre feature vectors of each target song. Therefore, the computer device automatically identifies and detects pirated singers based on the timbre features of the songs, solves the problem of low recognition efficiency caused by manual identification of pirated singers, and can quickly and efficiently process a large amount of singer recognition work, improving the recognition efficiency of pirated singers.
[0047] The following will be described above Figure 1Based on the illustrated embodiments, the embodiments of the present application will be described in further detail. Please refer to Figure 2 , another embodiment of the pirated singer detection method in the embodiments of the present application includes:
[0048] 201. Obtain N target songs under the name of the same target singer, and process each of the target songs to obtain the spectrogram of each of the target songs, where N is a positive integer greater than 1;
[0049] As a spectrum analysis view, the spectrogram has time on the abscissa, frequency on the ordinate, and the coordinate point value is the energy of the speech data. Since three-dimensional information is expressed in a two-dimensional plane, the magnitude of the energy value is represented by color. The darker the color of a point, the stronger the speech energy of that point. On the contrary, the lighter the color of a point, the weaker the speech energy of that point.
[0050] The way to process the target song to obtain the spectrogram of the target song can be that the computer device can transform the target song from the time domain to the frequency domain through a subband decomposition algorithm, and then divide it into several subbands. Further, the several subbands can be processed to obtain the spectrogram. Optionally, the subband decomposition algorithm can be the short-time Fourier transform, but it is not limited thereto.
[0051] Optionally, each subband includes: time information, frequency information, energy, and phase information. The computer device can remove the phase information of the subband, take the square of the modulus of the energy to obtain the energy value, and finally, the time information and frequency information can be used as the abscissa and ordinate of the spectrogram respectively, and the energy value of each point can be used as the coordinate value of that point.
[0052] 202. Input the spectrogram of the target song into the target timbre feature extraction model to obtain the timbre feature vector of the target song output by the target timbre feature extraction model;
[0053] In this embodiment, the target timbre feature extraction model can be any feature extraction model. For example, the target timbre feature extraction model can be a residual neural network ResNet model, a convolutional neural network, or a deep complex convolutional recurrent network DCCRN, etc. For example, if the target timbre feature extraction model is a ResNet34 model, the spectrogram of the target song can be input into the ResNet34 model, and the ResNet34 model can extract the timbre features of the spectrogram of the target song to obtain the timbre feature vector of the target song.
[0054] In this embodiment, a metric learning framework is adopted as the training framework of the timbre feature extraction model to train the model. Metric learning optimizes the network model by reducing the distance between classes and increasing the intra-class separability, and it is a practical machine learning method for comparing and measuring the similarity between data. At the same time, in this embodiment, a triplet network is adopted as the metric learning framework, and the timbre feature extraction model is trained based on this triplet network, which includes the following steps:
[0055] The computer device obtains at least one song triplet data, where each song triplet data includes a first song and a second song under the name of a first singer, and a third song under the name of a second singer, and the first singer is different from the second singer;
[0056] Process each song in each song triplet data respectively to obtain the spectrogram of each song in each song triplet data;
[0057] Obtain an initial timbre feature extraction model, and input the spectrogram of each song in the song triplet data into the initial timbre feature extraction model to obtain the timbre feature vector of each song in the song triplet data output by the initial timbre feature extraction model;
[0058] Determine the value of the triplet loss function of the initial timbre feature extraction model according to the timbre feature vector of each song in the song triplet data;
[0059] Update the network parameters of the initial timbre feature extraction model according to the value of the triplet loss function, and obtain the target timbre feature extraction model when the update is completed.
[0060] Among them, the method of processing each song in the song triplet data to obtain the spectrogram of each song has been described above, and will not be repeated here. The initial timbre feature extraction model can specifically be a ResNet model, a convolutional neural network, or a DCCRN network, etc.
[0061] The dimension of the timbre feature vector of each song in the song triplet data can be 1×40 dimensions. Use the triplet loss function to calculate the Euclidean distance between the timbre feature vector x a of the first song and the timbre feature vector x + of the second song in the song triplet data, and calculate the Euclidean distance between the timbre feature vector x a of the first song and the timbre feature vector x -The Euclidean distance between them makes the former timbre feature vectors (i.e. the timbre feature vectors of the first song of the first singer and the timbre feature vectors of the second song) close to each other, and the latter timbre feature vectors (i.e. the timbre feature vectors of the first song of the first singer and the timbre feature vectors of the third song of the second singer) distant from each other during the training of the initial timbre feature extraction model. In essence, the initial timbre feature extraction model learns the timbre commonalities between multiple songs under the name of the same singer, and at the same time learns the significant differences in timbre between songs under the names of different singers. When the value of the ternary loss function decreases to a stable state as the training continues, it represents the end of the training phase, and the network parameters of the initial timbre feature extraction model are updated. The initial timbre feature extraction model that has completed the network parameter update is used as the target timbre feature extraction model.
[0062] It can be understood that the timbre feature vector can be an embedding vector, and the distance between the above-mentioned timbre feature vectors can be not only Euclidean distance, but also Manhattan distance, Chebyshev distance, etc., which can be set according to the actual application scenario and is not limited here.
[0063] Among them, the ternary loss function can be expressed as:
[0064] L = max(0, || x a -x + ||-||x a -x - ||+α);
[0065] Where L represents the value of the ternary loss function, x a represents the timbre feature vector of the first song, x + represents the timbre feature vector of the second song, x - represents the timbre feature vector of the third song, and α represents the minimum gap.
[0066] 203. For each target song, taking the average value of the similarity between the target song and the timbre feature vector of each other target song as the average value of the similarity of the target song;
[0067] In this embodiment, the similarity of the timbre feature vectors between two of the N target songs under the name of the target singer is calculated by calculating the cosine similarity of the timbre feature vectors between two of the N target songs; or calculating the Manhattan distance of the timbre feature vectors between two of the N target songs; or calculating the Euclidean distance of the timbre feature vectors between two of the N target songs. This embodiment does not limit the method of calculating the similarity of the timbre feature vectors between two of the target songs, and any method that can measure the distance between vectors can be used in this embodiment.
[0068] Among them, the formula for calculating the cosine similarity of the timbre feature vectors between pairwise target songs can be:
[0069]
[0070] Among them, cos(θ) i,j represents the cosine similarity of the timbre feature vectors between the i-th song and the j-th song among the N target songs under the name of the target singer, and E i represents the timbre feature vector of the i-th song, and E j represents the timbre feature vector of the j-th song.
[0071] Therefore, when using the cosine similarity to represent the similarity of the timbre feature vectors between target songs, the larger the value of the cosine similarity, the more similar the timbre feature vectors are; on the contrary, it means that the timbre feature vectors are less similar.
[0072] After obtaining multiple similarities of the timbre feature vectors between each target song and other target songs, add the values of the multiple similarities and take the average to obtain the average similarity value. For example, if there are 10 target songs under the name of the target singer, the similarities of the timbre feature vectors between each target song and other target songs can be calculated respectively. For example, calculate the similarities of the timbre feature vector of the first target song with the timbre feature vectors of each of the remaining 9 target songs, that is, obtain 9 similarities of the timbre feature vectors between the first target song and the remaining 9 target songs, add these 9 similarities and take the average to obtain the average similarity value corresponding to the first target song. And so on, the average similarity values corresponding to the remaining target songs can be calculated respectively.
[0073] The formula for calculating the average similarity value can be expressed as:
[0074]
[0075] Or,
[0076]
[0077] Among them, score i represents the average similarity value corresponding to the i-th song among the N target songs under the name of the target singer, and cos(θ) i,j represents the cosine similarity of the timbre feature vectors between the i-th song and the j-th song. In the above two formulas, j can be equal to i. When j is equal to i, the value of cos(θ) i,j is 1.
[0078] In addition, the formula for calculating the average similarity value can also be expressed as:
[0079]
[0080] Or,
[0081]
[0082] where score i represents the average similarity corresponding to the i-th song among the N target songs under the name of the target singer, and cos(θ) i,j represents the cosine similarity of the timbre feature vectors between the i-th song and the j-th song. In these two formulas, j is not equal to i.
[0083] 204. If the average similarity of the N target songs meets the preset features, determine that the target singer is the original singer;
[0084] 205. If the average similarity of the N target songs does not meet the preset features, determine that the target singer is a pirated singer;
[0085] In this embodiment, the method of determining whether the target singer is a pirated singer based on the average similarity corresponding to each target song can be as follows. Since each of the N target songs under the name of the target singer has a corresponding average similarity, there are N average similarities corresponding to the N target songs under the name of the target singer. The average value of the N average similarities of the N target songs under the name of the target singer can be calculated, and it is judged whether the average value of the N average similarities of the N target songs under the name of the target singer is less than a preset threshold. If so, determine that the target singer is the original singer; if not, determine that the target singer is a pirated singer.
[0086] The calculation formula can be expressed as:
[0087]
[0088] where score A represents the average value of the N average similarities of the N target songs under the name of the target singer, and score i represents the average similarity corresponding to the i-th song among the N target songs under the name of the target singer.
[0089] In addition, the method of determining whether the target singer is a pirated singer can also be to establish a coordinate system, determine the coordinate points corresponding to the timbre feature vectors of each target song under the name of the target singer in this coordinate system, and judge whether the distribution state of the coordinate points in this coordinate system meets the preset distribution features. If so, determine that the target singer is the original singer; if not, determine that the target singer is a pirated singer.
[0090] Another way could be to use the distribution states of the coordinate points corresponding to each song under the original singer and the distribution states of the coordinate points corresponding to each song under the pirated singer as training data to input into a neural network model. Train this neural network model according to machine learning algorithms. Essentially, it is to let this neural network model learn the distribution states of the coordinate points corresponding to each song under the original singer and the distribution states of the coordinate points corresponding to each song under the pirated singer. After that, this trained neural network model can be used to identify the distribution state of the coordinate points of N target songs under the target singer, determine whether the target singer is a pirated singer or an original singer based on this distribution state, and obtain the recognition result output by this neural network model. The recognition result indicates whether the target singer is a pirated singer.
[0091] In addition, it is also possible to determine whether the target singer is a pirated singer by comparing with the distribution state of the coordinate points corresponding to each song known to be under the original singer. For example, it is known that singer A is an original singer. A rectangular coordinate system is established, and the coordinate points corresponding to the timbre feature vectors of each song are determined in the coordinate system to obtain the distribution state of the coordinate points corresponding to each song under singer A, as Figure 3 shown; a rectangular coordinate system is established, and the coordinate points corresponding to the timbre feature vectors of each song are determined in the coordinate system to obtain the distribution state of the coordinate points corresponding to each song under the target singer, assuming as Figure 4 shown. Comparing Figure 3 and Figure 4 it can be seen that Figure 3 the distribution of each coordinate point is relatively concentrated. It can be deduced that Figure 3 the concentrated distribution state shown or a distribution state more concentrated than the distribution state shown in Figure 3 should correspond to the original singer, while Figure 4 the distribution state of the coordinate points shown is extremely scattered, completely different from the distribution state shown in Figure 3 . Therefore, it can be deduced that the target singer is not an original singer and should be classified as a pirated singer.
[0092] The pirated singer detection method in the embodiments of the present application has been described above. Next, the computer device in the embodiments of the present application will be described. Please refer to Figure 5 One embodiment of the computer device in the embodiments of the present application includes:
[0093] An acquisition unit 501, configured to acquire N target songs under the same target singer, and process each of the target songs to obtain a spectrogram of each of the target songs, where N is a positive integer greater than 1;
[0094] A feature extraction unit 502, configured to input the spectrogram of the target song into a target timbre feature extraction model to obtain the timbre feature vector of the target song output by the target timbre feature extraction model;
[0095] A calculation unit 503, configured to, for each of the target songs, use the average value of the similarities between the target song and the timbre feature vectors of each of the other target songs as the average similarity value of the target song;
[0096] A detection unit 504, configured to determine that the target singer is the original singer if the average similarity value of the N target songs meets a preset feature;
[0097] The detection unit 504 is further configured to determine that the target singer is a pirated singer if the average similarity value of the N target songs does not meet the preset feature.
[0098] In a preferred implementation manner of this embodiment, the detection unit 504 is specifically configured to calculate the average value of the N average similarity values of the N target songs; determine whether the average value of the N average similarity values of the N target songs is less than a preset threshold; if so, determine that the target singer is the original singer; if not, determine that the target singer is a pirated singer.
[0099] In a preferred implementation manner of this embodiment, the computer device further includes:
[0100] A model training unit 505, configured to execute the training steps of the target timbre feature extraction model, and the steps include:
[0101] Obtain at least one song triple data, where each song triple data includes a first song and a second song under a first singer, and a third song under a second singer, and the first singer is different from the second singer;
[0102] Process each song in each of the song triple data respectively to obtain the spectrogram of each song in each of the song triple data;
[0103] Obtain an initial timbre feature extraction model, and input the spectrograms of each song in the song triple data into the initial timbre feature extraction model to obtain the timbre feature vectors of each song in the song triple data output by the initial timbre feature extraction model;
[0104] Determine the value of the triple loss function of the initial timbre feature extraction model according to the timbre feature vectors of each song in the song triple data;
[0105] Update the network parameters of the initial timbre feature extraction model according to the value of the triple loss function, and obtain the target timbre feature extraction model when the update is completed.
[0106] In a preferred implementation manner of this embodiment, the triple loss function satisfies:
[0107] L = max(0, ||x a - x + || - ||x a - x - || + α);
[0108] Wherein, L represents the value of the ternary loss function, x a represents the timbre feature vector of the first song, x + represents the timbre feature vector of the second song, x - represents the timbre feature vector of the third song, and α represents the minimum gap.
[0109] In a preferred implementation manner of this embodiment, the calculation unit 503 is specifically configured to calculate the cosine similarity of the timbre feature vectors between any two of the N target songs; alternatively, calculate the Manhattan distance of the timbre feature vectors between any two of the N target songs; alternatively, calculate the Euclidean distance of the timbre feature vectors between any two of the N target songs.
[0110] In a preferred implementation manner of this embodiment, the calculation formula of the cosine similarity of the timbre feature vectors between any two of the N target songs satisfies:
[0111]
[0112] Wherein, cos(θ) i,j represents the cosine similarity of the timbre feature vectors between the i-th song and the j-th song among the N target songs, E i represents the timbre feature vector of the i-th song, E j represents the timbre feature vector of the j-th song.
[0113] In this embodiment, the operations performed by each unit in the computer device are similar to those described in the foregoing Figures 1 to 2 illustrated embodiment, and will not be described herein again.
[0114] In this embodiment, spectrograms of each target song are obtained by processing N target songs under the name of the target singer, the timbre feature extraction model is used to extract the timbre features of the spectrograms of each target song to obtain the timbre feature vectors of each target song, the similarity of the timbre feature vectors between each target song is calculated, and it is determined whether the target singer is a pirated singer according to the similarity of the timbre feature vectors of each target song. Therefore, the computer device automatically identifies and detects pirated singers according to the timbre features of the songs, solves the problem of low recognition efficiency caused by manual identification of pirated singers, can realize fast and efficient processing of a large number of singer recognition tasks, and improves the recognition efficiency of pirated singers.
[0115] The computer device in the embodiments of the present application will be described below. Please refer to Figure 6 In one embodiment of the computer device in the embodiments of the present application, it includes:
[0116] The computer device 600 may include one or more central processing units (CPUs) 601 and a memory 605, and one or more applications or data are stored in the memory 605.
[0117] Among them, the memory 605 may be volatile storage or persistent storage. The programs stored in the memory 605 may include one or more modules, and each module may include a series of instruction operations on the computer device. Further, the central processing unit 601 may be configured to communicate with the memory 605 and execute a series of instruction operations in the memory 605 on the computer device 600.
[0118] The computer device 600 may further include one or more power supplies 602, one or more wired or wireless network interfaces 603, one or more input / output interfaces 604, and / or one or more operating systems, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0119] The central processing unit 601 may perform the operations executed by the computer device in the foregoing Figures 1 to 2 illustrated embodiments, and details are not described herein again.
[0120] The embodiments of the present application also provide a computer storage medium. In one embodiment, it includes: instructions are stored in the computer storage medium, and when the instructions are executed on the computer, the computer is caused to execute the operations executed by the computer device in the foregoing Figures 1 to 2 illustrated embodiments.
[0121] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above may refer to the corresponding processes in the foregoing method embodiments, and details are not described herein again.
[0122] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0123] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0124] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0125] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, read-only memory), random access memories (RAM, random access memory), magnetic disks, or optical disks and other various media that can store program codes.
Claims
1. A method for detecting pirated singers, characterized in that, The method includes: Obtaining N target songs under the same target singer, and processing each of the target songs to obtain the spectrogram of each of the target songs, where N is a positive integer greater than 1; Inputting the spectrogram of the target song into a target timbre feature extraction model to obtain the timbre feature vector of the target song output by the target timbre feature extraction model; For each of the target songs, taking the average of the similarities between the timbre feature vectors of the target song and each of the other target songs as the average similarity of the target song; If the average similarity of the N target songs meets a preset feature, determining that the target singer is the original singer; If the average similarity of the N target songs does not meet the preset feature, determining that the target singer is a pirated singer; The training steps of the target timbre feature extraction model include: Obtaining at least one song triple data, where each song triple data includes a first song and a second song under a first singer, and a third song under a second singer, and the first singer is different from the second singer; Processing each song in each of the song triple data to obtain the spectrogram of each song in each of the song triple data; Obtaining an initial timbre feature extraction model, and inputting the spectrogram of each song in the song triple data into the initial timbre feature extraction model to obtain the timbre feature vector of each song in the song triple data output by the initial timbre feature extraction model; Determining the value of the triple loss function of the initial timbre feature extraction model according to the timbre feature vectors of each song in the song triple data; Updating the network parameters of the initial timbre feature extraction model according to the value of the triple loss function, and obtaining the target timbre feature extraction model when the update is completed.
2. The method according to claim 1, wherein Judging whether the average similarity of the N target songs meets the preset feature includes: Calculating the average of the N average similarities of the N target songs; Judging whether the average of the N average similarities of the N target songs is less than a preset threshold; If so, determining that the target singer is the original singer; If not, determining that the target singer is a pirated singer.
3. The method according to claim 1, wherein The triple loss function satisfies: L = max(0, ||x a - x + || - ||x a - x - || + α); where L represents the value of the ternary loss function, x a represents the timbre feature vector of the first song, x + represents the timbre feature vector of the second song, x - represents the timbre feature vector of the third song, and α represents the minimum gap.
4. The method according to claim 1, wherein Calculating the similarity between the timbre feature vectors of every two of the N target songs includes: Calculating the cosine similarity between the timbre feature vectors of every two of the N target songs; Or, Calculating the Manhattan distance between the timbre feature vectors of every two of the N target songs; Or, Calculating the Euclidean distance between the timbre feature vectors of every two of the N target songs.
5. The method according to claim 4, characterized in that, The calculation formula of the cosine similarity between the timbre feature vectors of every two of the N target songs satisfies: where cos(θ) i,j represents the cosine similarity of the timbre feature vectors between the i-th song and the j-th song among the N target songs, E i represents the timbre feature vector of the i-th song, E j represents the timbre feature vector of the j-th song.
6. A computer device, characterized in that, The computer device includes: An acquisition unit for obtaining N target songs under the same target singer, and processing each of the target songs to obtain the spectrogram of each of the target songs, where N is a positive integer greater than 1; A feature extraction unit, configured to input the spectrogram of the target song into a target timbre feature extraction model to obtain a timbre feature vector of the target song output by the target timbre feature extraction model; A calculation unit, configured to, for each of the target songs, use the average value of the similarities between the timbre feature vector of the target song and the timbre feature vectors of every other target song as the average similarity of the target song; A detection unit, configured to determine that the target singer is the original singer if the average similarity of the N target songs meets a preset feature; The detection unit is further configured to determine that the target singer is a pirated singer if the average similarity of the N target songs does not meet the preset feature; The computer device further includes: A model training unit, configured to execute a training step of the target timbre feature extraction model, and the step includes: Obtaining at least one song triple data, where each song triple data includes a first song and a second song under a first singer, and a third song under a second singer, and the first singer is different from the second singer; Processing each song in each of the song triple data respectively to obtain a spectrogram of each song in each of the song triple data; Obtaining an initial timbre feature extraction model, inputting the spectrogram of each song in the song triple data into the initial timbre feature extraction model to obtain a timbre feature vector of each song in the song triple data output by the initial timbre feature extraction model; Determining a value of a triple loss function of the initial timbre feature extraction model according to the timbre feature vector of each song in the song triple data; Updating network parameters of the initial timbre feature extraction model according to the value of the triple loss function, and obtaining the target timbre feature extraction model when the update is completed.
7. The computer device according to claim 6, wherein The detection unit is specifically configured to calculate an average value of the N average similarities of the N target songs; determine whether the average value of the N average similarities of the N target songs is less than a preset threshold; if so, determine that the target singer is the original singer; if not, determine that the target singer is a pirated singer.
8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, The processor, when executing the computer program, implements the method according to any one of claims 1 to 5.
9. A computer storage medium, characterized in that, Instructions are stored in the computer storage medium, and when the instructions are executed on the computer, the computer is caused to execute the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Song turning and singing synthesis and execution method and device, equipment, medium and product
CN114708843A