Method, device, and storage medium for classifying singers

By performing multiple clustering and similarity analyses on song attribute information, the singer classification was optimized, solving the problem of inaccurate classification caused by random seeds and achieving more accurate singer classification and recommendation.

CN117609892BActive Publication Date: 2026-07-21TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
Filing Date
2023-12-06
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

In existing technologies, the clustering results for singer classification rely on random seeds, leading to inaccurate classification and fluctuations.

Method used

By performing multiple clustering processes on the attribute information of multiple songs, multiple initial classification information is obtained, forming a classification information set. Based on the corrected classification information of the songs, the final classification information of the singers is determined. The classification results are optimized using a song type recognition model and similarity analysis.

Benefits of technology

It improves the accuracy of song and artist classification, reduces the impact of fluctuations caused by random seeds, and achieves more accurate artist recommendations and searches.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117609892B_ABST
    Figure CN117609892B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, device and storage medium for classifying singers, and belongs to the technical field of computers. In the present disclosure, a plurality of clustering processes are performed on songs to determine a plurality of initial classification information corresponding to each song. Each initial classification information is from one clustering process. In the case that the clustering process is affected by the fluctuation of a random seed, the plurality of initial classification information corresponding to the song may be different. At this time, the final classification information, i.e. the corrected classification information, of the song can be determined based on the plurality of initial classification information. In this way, the fluctuation of the random seed can be eliminated to a certain extent, and the classification accuracy of the song can be improved. Furthermore, the singer classification can be determined based on the song classification, and more accurate singer classification information can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a method, apparatus and storage medium for classifying singers. Background Technology

[0002] As the number of singers increases, it becomes necessary to categorize them. Based on the categorization results, singers that match user preferences can be accurately recommended to users.

[0003] In related technologies, clustering is used to process the singer's attribute information to obtain the singer's classification information.

[0004] Because the clustering results obtained by clustering processing depend on the random seed, which is an internal parameter of the clustering algorithm, it can be randomly generated by the computer or arbitrarily set by the user. If different random seeds are used for the same task, the clustering results may fluctuate. In other words, the singer classification calculated by this technique is inaccurate. Summary of the Invention

[0005] To address the related technical problems, this disclosure provides a method, apparatus, and storage medium for classifying singers. The technical solution is as follows:

[0006] Firstly, a method for classifying singers is provided, the method comprising:

[0007] Multiple clustering processes are performed on the attribute information of multiple songs to obtain multiple clustering results, wherein each clustering result includes the initial classification information corresponding to the multiple songs respectively;

[0008] In multiple clustering results, multiple initial classification information corresponding to each song is obtained, and these are used to form a set of classification information corresponding to each song.

[0009] Based on the classification information set corresponding to each song, determine the corrected classification information corresponding to each song;

[0010] Based on the corrected classification information corresponding to at least one song by the same singer, the classification information corresponding to the singer is determined.

[0011] In one possible implementation, the attribute information includes at least one of the following: song name, album name, instrument information, pitch information, and rhythm information.

[0012] In one possible implementation, determining the corrected classification information for each song based on the classification information set corresponding to each song includes:

[0013] For each song, multiple initial classification information sets are combined to form a classification vector.

[0014] The classification vectors of multiple songs are clustered to obtain clustering results. Based on the clustering results, the corrected classification information corresponding to each song is determined.

[0015] In one possible implementation, determining the corrected classification information for each song based on the clustering results includes:

[0016] For each cluster in the clustering results, display the attribute information of at least one song corresponding to the cluster. In response to the classification information output instruction corresponding to the cluster, obtain the input corrected classification information and determine it as the corrected classification information for each song corresponding to the cluster.

[0017] In one possible implementation, determining the corrected classification information for each song based on the clustering results includes:

[0018] For each cluster in the clustering results, the attribute information of at least one song corresponding to the class is input into the song type recognition model to obtain the type corresponding to the at least one song. The type that appears most frequently among the types corresponding to the at least one song is determined as the corrected classification information for each song corresponding to the class.

[0019] In one possible implementation, determining the classification information corresponding to the singer based on the corrected classification information corresponding to at least one song by the same singer includes:

[0020] In the corrected classification information corresponding to at least one song by the same singer, the corrected classification information that appears most frequently is determined as the classification information corresponding to that singer; or,

[0021] In the corrected classification information corresponding to at least one song by the same singer, the corrected classification information that appears more than a certain number of times is determined as the classification information corresponding to the singer; or,

[0022] In the corrected classification information corresponding to at least one song by the same singer, the corrected classification information that appears most frequently, and the corrected classification information that appears more than a certain number of times, are jointly determined as the classification information corresponding to the singer; or,

[0023] The corrected classification information corresponding to at least one song by the same singer is input into the singer type recognition model to obtain the classification information corresponding to the singer.

[0024] In one possible implementation, the method further includes:

[0025] Obtain attribute information of the target song other than the multiple songs mentioned above;

[0026] Determine the similarity between the attribute information of the target song and the attribute information of the multiple songs;

[0027] Among the multiple songs, at least one song with the highest similarity to the attribute information of the target song is identified, and the classification information of the target song is determined based on the corrected classification information of the at least one song.

[0028] In one possible implementation, the method further includes:

[0029] Identify the target singer to whom the target song belongs;

[0030] Based on the classification information of the target song and the corrected classification information of other songs by the target singer, update the classification information corresponding to the target singer.

[0031] Secondly, a device for classifying singers is provided, characterized in that the device comprises:

[0032] The initial song clustering module is used to perform multiple clustering processes on the attribute information of multiple songs to obtain multiple clustering results. Each clustering result includes the initial classification information corresponding to the multiple songs.

[0033] The song classification set module is used to obtain multiple initial classification information for each song from multiple clustering results, and form a classification information set for each song respectively;

[0034] The song correction and classification module is used to determine the correction and classification information for each song based on the classification information set corresponding to each song.

[0035] The singer classification module is used to determine the classification information corresponding to the singer based on the corrected classification information corresponding to at least one song of the same singer.

[0036] In one possible implementation, the attribute information includes at least one of the following: song name, album name, instrument information, pitch information, and rhythm information.

[0037] In one possible implementation, the song correction and classification module is used for:

[0038] For each song, multiple initial classification information sets are combined to form a classification vector.

[0039] The classification vectors of multiple songs are clustered to obtain clustering results. Based on the clustering results, the corrected classification information corresponding to each song is determined.

[0040] In one possible implementation, the song correction and classification module is used for:

[0041] For each cluster in the clustering results, display the attribute information of at least one song corresponding to the cluster. In response to the classification information output instruction corresponding to the cluster, obtain the input corrected classification information and determine it as the corrected classification information for each song corresponding to the cluster.

[0042] In one possible implementation, the song correction and classification module is used for:

[0043] For each cluster in the clustering results, the attribute information of at least one song corresponding to the class is input into the song type recognition model to obtain the type corresponding to the at least one song. The type that appears most frequently among the types corresponding to the at least one song is determined as the corrected classification information for each song corresponding to the class.

[0044] In one possible implementation, the singer classification module is used for:

[0045] In the corrected classification information corresponding to at least one song by the same singer, the corrected classification information that appears most frequently is determined as the classification information corresponding to that singer; or,

[0046] In the corrected classification information corresponding to at least one song by the same singer, the corrected classification information that appears more than a certain number of times is determined as the classification information corresponding to the singer; or,

[0047] In the corrected classification information corresponding to at least one song by the same singer, the corrected classification information that appears most frequently, and the corrected classification information that appears more than a certain number of times, are jointly determined as the classification information corresponding to the singer; or,

[0048] The corrected classification information corresponding to at least one song by the same singer is input into the singer type recognition model to obtain the classification information corresponding to the singer.

[0049] In one possible implementation, the device further includes a target song classification module, used for:

[0050] Obtain attribute information of the target song other than the multiple songs mentioned above;

[0051] Determine the similarity between the attribute information of the target song and the attribute information of the multiple songs;

[0052] Among the multiple songs, at least one song with the highest similarity to the attribute information of the target song is identified, and the classification information of the target song is determined based on the corrected classification information of the at least one song.

[0053] In one possible implementation, the singer classification module is further used for:

[0054] Identify the target singer to whom the target song belongs;

[0055] Based on the classification information of the target song and the corrected classification information of other songs by the target singer, update the classification information corresponding to the target singer.

[0056] Thirdly, a computer device is provided, comprising a memory and a processor, the memory for storing computer instructions, and the processor for executing the computer instructions stored in the memory to cause the computer device to perform the methods provided in the first aspect and its possible implementations.

[0057] Fourthly, a computer-readable storage medium is provided, which stores computer program code, such that when the computer program code is executed by a computer device, the computer device performs the method provided in the first aspect and its possible implementations.

[0058] Fifthly, a computer program product is provided, comprising computer program code, wherein when the computer program code is executed by a computer device, the computer device executes the method provided by the first aspect and its possible implementations.

[0059] The beneficial effects of the technical solutions provided in this disclosure are:

[0060] In this embodiment, songs are subjected to multiple clustering processes to determine multiple initial classification information for each song. Each initial classification information comes from one clustering process. Given the potential for fluctuations in the random seed during clustering, the multiple initial classification information for a song may differ. In this case, the final classification information for the song can be determined based on these multiple initial classification information, i.e., corrected classification information. This can, to some extent, eliminate the influence of random seed fluctuations and improve the accuracy of song classification. Furthermore, determining the singer classification based on the song classification yields more accurate singer classification information. Attached Figure Description

[0061] Figure 1 This is a schematic diagram of the structure of a computer device provided in an embodiment of this disclosure;

[0062] Figure 2 This is a schematic flowchart of a method for classifying singers provided in an embodiment of this disclosure;

[0063] Figure 3 This is a schematic diagram of a process for initial clustering of songs provided in an embodiment of this disclosure;

[0064] Figure 4 This is a schematic diagram of a processing flow for determining the corrected classification information of a song based on initial classification information, provided by an embodiment of this disclosure;

[0065] Figure 5 This is a schematic diagram of a song classification process provided in an embodiment of this disclosure;

[0066] Figure 6 This is a schematic diagram of a device structure for classifying singers according to an embodiment of the present disclosure. Detailed Implementation

[0067] To make the objectives, technical solutions, and advantages of this disclosure clearer, the embodiments of this disclosure will be described in further detail below with reference to the accompanying drawings.

[0068] This disclosure provides a method for singer classification, and the execution entity of this method can be a server. The server can be a single server or a group of servers. If it is a single server, it can be responsible for all the processing in the following scheme. If it is a group of servers, different servers in the group can be responsible for different processing in the following scheme. The specific processing allocation can be arbitrarily set by technicians according to actual needs, and will not be elaborated here.

[0069] The server can be a backend server for an application, which may have audio playback functionality; the application may be a music player, etc. This embodiment uses the example of using a music player software's backend server for singer classification to provide a detailed explanation of the solution. Other cases are similar and will not be described in detail here.

[0070] Figure 1 This is a schematic diagram of a server structure provided in an embodiment of this disclosure. From a hardware perspective, the server structure can be as follows: Figure 1 As shown, it includes a processor 110, a memory 120, and a communication component 130.

[0071] The processor 110 can be a central processing unit (CPU) or a system on chip (SoC), etc. The processor 110 can be used to process various operation instructions, such as obtaining the attribute information of songs and calculating the similarity between songs.

[0072] The memory 120 may include various volatile or non-volatile memories, such as solid-state disks (SSDs) and dynamic random access memory (DRAM). The memory 120 can be used to store initial data, intermediate data, and result data used in the relevant processing, such as song attribute information and the similarity between songs.

[0073] The communication component 130 can be a wired network connector, an ultra-wideband (UWB) technology module, a wireless fidelity (WiFi) module, a Bluetooth module, a cellular network communication module, etc. The communication component 130 can be used to transmit data with other devices, such as other servers or terminals.

[0074] The following explains some terms used in the embodiments of this disclosure.

[0075] Word-to-vector (Word2Vec): A model used in natural language processing to generate word vectors.

[0076] Principal Component Analysis (PCA): A method for dimensionality reduction of high-dimensional vectors.

[0077] K-means clustering algorithm (K-Means): An unsupervised clustering algorithm that allows specifying the number of clusters, dividing vectors into a specified number of classes, with vectors in each class having high similarity.

[0078] People frequently use audio playback applications, which store a vast number of songs and artists on their backend servers. To facilitate song management, these servers often categorize artists. This categorization information enables various business functions, such as accurately recommending songs by artists that match a user's preferences, or directly displaying artist categorization information within the artist's profile.

[0079] This disclosure provides a method for classifying singers. The process of singer classification is described in detail below. Figure 2 As shown, it includes the following steps:

[0080] 201. Multiple clustering processes were performed on the attribute information of multiple songs to obtain multiple clustering results.

[0081] The song's attribute information represents specific characteristics of the song, such as song title, album name, instrument information, pitch information, rhythm information, etc. The random seeds used in multiple clustering processes can all be different. Each clustering result includes initial classification information for multiple songs. The initial classification information for each song can be a class number; for example, if the number of classes is 10, the initial classification information can be any value from 0 to 9. Generally, the initial classification information can simply indicate whether the types are the same or different, without specifying a particular type (e.g., rock, jazz, etc.). The number of classes input for the clustering process can be set by technical personnel according to actual needs.

[0082] 202. From multiple clustering results, obtain multiple initial classification information corresponding to each song, and form a classification information set corresponding to each song.

[0083] The classification information set is the set of initial classification information for each song after multiple clustering processes.

[0084] 203. Based on the classification information set corresponding to each song, determine the corrected classification information corresponding to each song.

[0085] Each song can have one or more correction classification information entries. The correction classification information can represent a specific type, such as rock, jazz, folk, blues, rap, love song, unison, solo, duet, high notes, mid-range, low notes, etc.

[0086] 204. Based on the corrected classification information corresponding to at least one song by the same singer, determine the classification information corresponding to the singer.

[0087] The classification information for a particular singer can be determined based on the corrected classification information of the songs corresponding to that singer. Users can select all songs by that singer, or, based on actual business needs, select at least one song by that singer relevant to those needs. Then, the singer's classification information is determined based on the corrected classification information of these songs obtained in step 203. Based on the singer's classification information, singers matching the user's preferences can be accurately recommended. Furthermore, the singer's classification information can also be used for retrieval. Users can input classification information on the terminal, which then sends a search request for that classification information to the server. The server finds singers with that classification information and displays the results on the terminal for further user operations, such as browsing singer information or playing the singer's songs.

[0088] There are several ways to determine the singer's category information. The following are some possible processing methods:

[0089] Method 1: Among the corrected classification information corresponding to at least one song by the same singer, the corrected classification information that appears most frequently is determined as the classification information corresponding to the singer.

[0090] For example, in the three corrected classification information of pop, rap and ballad for multiple songs by the same singer, the corresponding occurrence times are 8, 3 and 8 respectively. Pop and ballad have the same number of occurrences and the most occurrences. Therefore, pop and ballad can be identified as the classification information corresponding to this singer.

[0091] Method 2: In the corrected classification information corresponding to at least one song by the same singer, the corrected classification information that appears more than the number of times threshold is determined as the classification information corresponding to the singer.

[0092] For example, among the three corrected classification information of popular, rap and ballad for multiple songs by the same singer, the corresponding occurrence counts are 8, 3 and 8 respectively. According to actual business needs, the occurrence count threshold can be set to 5. From these three corrected classification information, the corrected classification information of popular and ballad with an occurrence count exceeding 5 can be selected as the classification information corresponding to the singer.

[0093] Method 3: In the corrected classification information corresponding to at least one song of the same singer, the corrected classification information that appears most frequently, as well as the corrected classification information of pop and ballad that appears more than 5 times, are jointly determined as the classification information corresponding to the singer.

[0094] For example, among the three corrected classification information of pop, rap and ballad for multiple songs by the same singer, the corresponding occurrence counts are 8, 3 and 9 respectively. According to actual business needs, the occurrence count threshold can be set to 5. The corrected classification information with more than 5 occurrences is pop and ballad, and the corrected classification information with the most occurrences is ballad. Therefore, ballad is selected as the classification information corresponding to this singer.

[0095] This approach prevents frequently occurring category information from being ignored, and also prevents the inability to select category information when the frequency of occurrence is very low (below the threshold), resulting in more accurate category information for singers.

[0096] Method 4: Input the corrected classification information corresponding to at least one song by the same singer into the singer type recognition model to obtain the classification information corresponding to the singer.

[0097] Among them, the singer type recognition model can be a machine learning model, such as a convolutional neural network or a perceptron.

[0098] The phrase "at least one song" means that some or all of the singer's songs are selected based on actual business needs.

[0099] In this embodiment of the disclosure, when performing the clustering process in step 201 above, song vectors can be generated first based on the song's attribute information, and then clustering can be performed based on the song vectors. The corresponding processing flow is as follows: Figure 3 As shown, it includes the following steps:

[0100] 301, retrieve the song's attribute information.

[0101] Obtain the song's attribute information, which is information that can represent a specific attribute of the song, such as at least one of the following: song name, album name, instrument information, pitch information, and rhythm information.

[0102] 302. Based on the attribute information of the song, generate a song vector.

[0103] Natural language processing (NLP) is performed on the attribute information of songs to determine the song vector for each song. For example, in this embodiment, a Word2Vec model can be selected to perform NLP on the song attribute information. The Word2Vec model can express each word in a given sentence or word sequence as a vector. For example, for each song stored in the database within a specified time period, at least one of the song title, album name, instrument information, pitch information, and rhythm information is listed as a sentence. This constructs the song's attribute information, which is then input into the Word2Vec model to output the song vector.

[0104] Alternatively, PCA (Programmable Array Analysis) can be used to reduce the dimensionality of song vectors, transforming the high-dimensional space into a lower-dimensional space and enabling vector space visualization. For example, PCA can be used to reduce song vectors to three dimensions. In this way, each song vector can be considered a coordinate point in three-dimensional space, allowing the song vectors to be represented in three-dimensional space and enabling visualization.

[0105] Optionally, the singer's attribute information can be obtained, and a singer vector can be generated based on the singer's attribute information.

[0106] 303. Multiple clustering processes are performed on the attribute information of multiple songs to obtain multiple clustering results.

[0107] The song vectors generated by the above steps are subjected to multiple clustering processes, with different random seeds used in each clustering process, resulting in multiple clustering results.

[0108] The above process uses song vectors as the feature information of the song. In addition, other forms of data can also be used to represent the feature information of the song, such as matrices, strings, etc.

[0109] In daily applications, clustering algorithms with random seeds are frequently used because clustering algorithms without random seeds are less efficient and have fewer application scenarios. However, different random seeds can lead to different clustering results. This disclosure provides a method for determining the corrected classification information of a song. The method involves changing the random seed in the clustering algorithm, performing multiple clustering operations, and recording the results of each clustering. Each clustering result serves as the initial classification information, and the corrected classification information of the song is determined based on this initial classification information. The process of determining the corrected classification information based on the initial classification information can be as follows: Figure 4 As shown, it includes the following steps:

[0110] 401. Combine multiple initial classification information from the classification information set corresponding to each song into a classification vector.

[0111] Each song undergoes multiple clustering processes, and the clustering result for each song in each process is recorded, which constitutes the initial classification information of the song. Optionally, a clustering algorithm can be used to determine the classification result of the song. For example, the K-Means algorithm can be selected. The classification result of this algorithm is influenced by a random seed. Ten clustering processes are performed, with the same number of target clusters selected in each process, and different random seeds for each process. The clustering results of each process are recorded, and the clustering results obtained from each process for each song are recorded as the initial classification information of the song. Each dimension of the above classification vector represents the clustering result obtained from each clustering process for each song. At this point, for any given song, it corresponds to multiple initial classification information, each of which comes from one clustering process. Some of these initial classification information may be the same, and some may be different. These multiple initial classification information can be arranged in the order of the clustering processes to form a sequence, which can be the classification vector corresponding to the song.

[0112] For example, five songs are labeled a1, a2, a3, a4, and a5. The song vectors of these five songs are clustered six times. The first clustering result is v11, v12, v13, v14, v15, which are the class numbers corresponding to a1, a2, a3, a4, and a5, respectively. The second clustering result is v21, v22, v23, v24, v25, ..., and the sixth clustering result is v61, v62, v63, v64, v65. Therefore, the classification vector corresponding to song a1 is (v11, v21, v31, v41, v51, v61), the classification vector corresponding to song a2 is (v12, v22, v32, v42, v52, v62), ..., and the classification vector corresponding to song a5 is (v15, v25, v35, v45, v55, v65).

[0113] Generally, for a song, most of the initial classification information is quite similar, or even has the same value. The elements of the corresponding classification vector also have this characteristic.

[0114] 402. Cluster the classification vectors of multiple songs to obtain clustering results, and determine the corrected classification information for each song based on the clustering results.

[0115] The clustering results obtained from the above operations are meaningless. To obtain meaningful corrected classification information, further processing of the clustering results is required. There are several ways to determine the corrected classification information for a song:

[0116] Clustering is performed on the classification vectors. The clustering algorithm can be arbitrarily chosen, and technicians can set the number of clusters input into the algorithm according to actual needs. The resulting clustering results can include secondary classification information corresponding to each classification vector, i.e., secondary classification information for each song. This secondary classification information can be a class number, used only for class differentiation and cannot be mapped to an actual class.

[0117] Furthermore, corrected classification information can be set for the songs corresponding to each category number. There are several possible processing methods; two are given below:

[0118] One approach is to manually add corrective classification information to the songs based on the clustering results.

[0119] For each cluster in the clustering results, the attribute information of at least one song corresponding to that cluster is displayed. In response to the output instruction for the classification information corresponding to that cluster, the input corrected classification information is obtained and determined as the corrected classification information for each song in that cluster. For example, if a cluster consists entirely of rap songs, the relevant personnel can set the corrected classification information for that cluster to rap.

[0120] Method 2: Corrective classification information can be added to the song automatically based on the clustering results using a computer.

[0121] For each cluster in the clustering results, the attribute information of at least one song corresponding to that cluster is input into the song type recognition model to obtain the type corresponding to each of the at least one song. The type that appears most frequently among the types corresponding to the at least one song is determined as the corrected classification information for each song in that cluster. The song type recognition model can be a machine learning model, such as a neural network or decision tree.

[0122] In addition, there are other possible methods, and multiple methods can be used in combination according to actual business needs.

[0123] Alternatively, the categories can be sorted in descending order based on the total number of songs in each category. The category with the smallest total number of songs in the last category may have weaker representativeness of the corrected classification information. Therefore, the corrected classification information of this category can be recorded as "other". Alternatively, the last few categories in the sort can be collectively referred to as "other".

[0124] In practice, the above Figure 2 The process can be performed periodically or triggered manually. Besides the initial artist categorization process, new songs may be added to the categorization, such as newly released songs, newly licensed songs, or songs from other platforms. When the attribute information of these songs is obtained, it may affect the categorization information of a particular artist. In this case, it is not necessary to re-categorize all songs. Figure 2 The processing can be done by processing only newly added songs to determine the corresponding singer's classification information. The corresponding processing flow is as follows: Figure 5 As shown, it includes the following steps:

[0125] 501, retrieves attribute information of the target song in addition to multiple songs.

[0126] When a new song is added to the artist category, the newly added target song can be a newly released song by the artist, a song for which the copyright has been newly acquired, or a song for which the platform does not have the copyright but the attribute information of the song can be obtained from other platforms. For example, for a newly released song by the artist, the attribute information of the aforementioned target song can be obtained.

[0127] 502, determine the similarity between the attribute information of the target song and the attribute information of multiple songs.

[0128] To determine the classification information of a newly added target song, we can first obtain its attribute information and then proceed with step 302. Based on the attribute information, we obtain the song vector of the target song, and then calculate the similarity between the song vector of the target song and the song vectors of each song in the database. This similarity is the similarity between the target song and the songs in the database. The similarity between song vectors can be derived from the cosine of the angle between the vectors.

[0129] 503. Among multiple songs, identify at least one song with the highest similarity to the attribute information of the target song, and determine the classification information of the target song based on the corrected classification information of the at least one song.

[0130] Among all the songs in the database, select the k songs with the highest similarity to the target song. Record the corrected classification information of these k songs, as well as the similarity between these k songs and the target song. Sum the similarity scores corresponding to the same corrected classification information to obtain the corrected classification information and the corresponding similarity score. Select the corrected classification information corresponding to the highest similarity score as the classification information of the target song.

[0131] In this way, when new songs are added, it is not necessary to re-process all songs. Figure 2 The processing method uses similarity to determine the classification information of songs, reducing the resources and time consumed.

[0132] Considering that the category information of the corresponding singers may be affected after new songs are added, the category information of the singers can be updated. The corresponding processing can be as follows:

[0133] Identify the target artist to whom the target song belongs. Based on the target song's category information and the corrected category information of the target artist's other songs, update the corresponding category information for the target artist. For example, if the artist's category information was previously rap, but newly added song categories are all R&B, and this category appears more frequently than rap, and actual business requirements dictate that the most frequently appearing category information should be used as the artist's category information, then the artist's category information is updated to R&B.

[0134] In this embodiment, songs are subjected to multiple clustering processes to determine multiple initial classification information for each song. Each initial classification information comes from one clustering process. Given the potential for fluctuations in the random seed during clustering, the multiple initial classification information for a song may differ. In this case, the final classification information for the song can be determined based on these multiple initial classification information, i.e., corrected classification information. This can, to some extent, eliminate the influence of random seed fluctuations and improve the accuracy of song classification. Furthermore, determining the singer classification based on the song classification yields more accurate singer classification information.

[0135] All of the above-mentioned optional technical solutions can be combined in any way to form optional embodiments of this disclosure, and will not be described in detail here.

[0136] This disclosure also provides an apparatus for classifying singers, which can be applied to the server in the above embodiments, such as... Figure 6 As shown, the device includes:

[0137] The initial song clustering module 610 is used to perform multiple clustering processes on the attribute information of multiple songs to obtain multiple clustering results, wherein each clustering result includes the initial classification information corresponding to the multiple songs respectively;

[0138] The song classification set module 620 is used to obtain multiple initial classification information for each song from multiple clustering results, and form a classification information set for each song respectively;

[0139] The song correction and classification module 630 is used to determine the correction and classification information corresponding to each song based on the classification information set corresponding to each song.

[0140] The singer classification module 640 is used to determine the classification information corresponding to the singer based on the corrected classification information corresponding to at least one song of the same singer.

[0141] In one possible implementation, the attribute information includes at least one of the following: song name, album name, instrument information, pitch information, and rhythm information.

[0142] In one possible implementation, the song correction and classification module 630 is used for:

[0143] For each song, multiple initial classification information sets are combined to form a classification vector.

[0144] The classification vectors of multiple songs are clustered to obtain clustering results. Based on the clustering results, the corrected classification information corresponding to each song is determined.

[0145] In one possible implementation, the song correction and classification module 630 is used for:

[0146] For each cluster in the clustering results, display the attribute information of at least one song corresponding to the cluster. In response to the classification information output instruction corresponding to the cluster, obtain the input corrected classification information and determine it as the corrected classification information for each song corresponding to the cluster.

[0147] In one possible implementation, the song correction and classification module 630 is used for:

[0148] For each cluster in the clustering results, the attribute information of at least one song corresponding to the class is input into the song type recognition model to obtain the type corresponding to the at least one song. The type that appears most frequently among the types corresponding to the at least one song is determined as the corrected classification information for each song corresponding to the class.

[0149] In one possible implementation, the singer classification module 640 is used for:

[0150] In the corrected classification information corresponding to at least one song by the same singer, the corrected classification information that appears most frequently is determined as the classification information corresponding to that singer; or,

[0151] In the corrected classification information corresponding to at least one song by the same singer, the corrected classification information that appears more than a certain number of times is determined as the classification information corresponding to the singer; or,

[0152] In the corrected classification information corresponding to at least one song by the same singer, the corrected classification information that appears most frequently, and the corrected classification information that appears more than a certain number of times, are jointly determined as the classification information corresponding to the singer; or,

[0153] The corrected classification information corresponding to at least one song by the same singer is input into the singer type recognition model to obtain the classification information corresponding to the singer.

[0154] In one possible implementation, the device further includes a target song classification module 650, used for:

[0155] Obtain attribute information of the target song other than the multiple songs mentioned above;

[0156] Determine the similarity between the attribute information of the target song and the attribute information of the multiple songs;

[0157] Among the multiple songs, at least one song with the highest similarity to the attribute information of the target song is identified, and the classification information of the target song is determined based on the corrected classification information of the at least one song.

[0158] In one possible implementation, the singer classification module 640 is further configured to:

[0159] Identify the target singer to whom the target song belongs;

[0160] Based on the classification information of the target song and the corrected classification information of other songs by the target singer, update the classification information corresponding to the target singer.

[0161] In this embodiment, songs are subjected to multiple clustering processes to determine multiple initial classification information for each song. Each initial classification information comes from one clustering process. Given the potential for fluctuations in the random seed during clustering, the multiple initial classification information for a song may differ. In this case, the final classification information for the song can be determined based on these multiple initial classification information, i.e., corrected classification information. This can, to some extent, eliminate the influence of random seed fluctuations and improve the accuracy of song classification. Furthermore, determining the singer classification based on the song classification yields more accurate singer classification information.

[0162] It should be noted that the singer classification device provided in the above embodiments is only an example of the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the singer classification device and the singer classification method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0163] This disclosure also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center that includes one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to perform a method of business processing, or instruct the computing device to perform a method of business processing.

[0164] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for classifying singers, characterized in that, The method includes: Multiple clustering processes are performed on the attribute information of multiple songs to obtain multiple clustering results, wherein each clustering result includes the initial classification information corresponding to the multiple songs respectively; In multiple clustering results, multiple initial classification information corresponding to each song is obtained, and these are used to form a set of classification information corresponding to each song. For each song, multiple initial classification information sets are combined to form a classification vector. Clustering is performed on the classification vectors of multiple songs to obtain clustering results. Based on the clustering results, the corrected classification information corresponding to each song is determined. Based on the corrected classification information corresponding to at least one song by the same singer, the classification information corresponding to the singer is determined.

2. The method according to claim 1, characterized in that, The attribute information includes at least one of the following: song name, album name, instrument information, pitch information, and rhythm information.

3. The method according to claim 1, characterized in that, The determination of the corrected classification information for each song based on the clustering results includes: For each cluster in the clustering results, display the attribute information of at least one song corresponding to the cluster. In response to the classification information output instruction corresponding to the cluster, obtain the input corrected classification information and determine it as the corrected classification information for each song corresponding to the cluster.

4. The method according to claim 1, characterized in that, The determination of the corrected classification information for each song based on the clustering results includes: For each cluster in the clustering results, the attribute information of at least one song corresponding to the class is input into the song type recognition model to obtain the type corresponding to the at least one song. The type that appears most frequently among the types corresponding to the at least one song is determined as the corrected classification information for each song corresponding to the class.

5. The method according to claim 1, characterized in that, The process of determining the classification information corresponding to the singer based on the corrected classification information corresponding to at least one song by the same singer includes: In the correction classification information corresponding to at least one song by the same singer, the correction classification information that appears most frequently and the correction classification information that appears more than once the frequency threshold are jointly determined as the classification information corresponding to the singer.

6. The method according to any one of claims 1-5, characterized in that, The method further includes: Obtain attribute information of the target song other than the multiple songs mentioned above; Determine the similarity between the attribute information of the target song and the attribute information of the multiple songs; Among the multiple songs, at least one song with the highest similarity to the attribute information of the target song is identified, and the classification information of the target song is determined based on the corrected classification information of the at least one song.

7. The method according to claim 6, characterized in that, The method further includes: Identify the target singer to whom the target song belongs; Based on the classification information of the target song and the corrected classification information of other songs by the target singer, update the classification information corresponding to the target singer.

8. A computer device, characterized in that, The computer device includes a memory and a processor, the memory being used to store computer instructions; The processor executes computer instructions stored in the memory to cause the computer device to perform the method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program code, which, when executed by a computer device, performs the method according to any one of claims 1 to 7.