Song Tag Sorting, Access Method and Its Device, Equipment, Medium, Product

By using the deep semantic labels extracted from the song list's own tag set and neural network model, combining song attributes and manual tags, the inductive labels of songs are determined, and the problems of large granularity and incomplete coverage in the existing technology are solved, and efficient and accurate song labeling and retrieval are achieved.

CN114048347BActive Publication Date: 2025-05-27GUANGZHOU GESHEN INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111333487.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-11
Publication Date
2025-05-27
Estimated Expiration
2041-11-11

AI Technical Summary

Technical Problem

In the existing song search technology, the song labels have large granularity and incomplete coverage, resulting in poor label classification results, high manual marking costs and difficult maintenance and updates.

Method used

By obtaining the own set of tags for multiple playlists to which the target song belongs, combining the song's own attributes and manual tag labels, a pre-trained neural network model is used to extract the deep semantic tags of the playlist description text, and count the tag hit frequency to determine the inductive tag.

Benefits of technology

It realizes efficient and accurate marking of song labels, reduces manual labeling costs, improves label coverage and recall, and simplifies the song organization and retrieval process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114048347B_ABST
    Figure CN114048347B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of song retrieval, and discloses a method and device, equipment, medium, and product for song label sorting and access. The sorting method includes: obtaining a target song in a song library; querying and determining a plurality of preset label subsets corresponding to the target song, and merging each label subset into a label set, wherein the first label subset is the union of the own label sets mapped by each playlist that collects the target song in a preset playlist library; counting the frequency of each label in the label set of the target song hitting the own label set corresponding to the playlist in the playlist library, and determining multiple labels whose frequencies meet the preset conditions as the inductive labels of the target song. The present application uses playlists as the source of label materials for target songs in the song library, and uses the own label sets corresponding to the playlists to label the target songs, which is convenient for preparing inductive labels for the songs collected by users, and facilitates users to accurately and efficiently query a large number of songs they have collected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of song retrieval, and particularly to a method for organizing and accessing song tags, and its corresponding device, computer equipment, computer-readable storage medium, and computer program product. Background Art

[0002] With the improvement of the quality of people's spiritual and cultural life, the total number of songs in personal song libraries is becoming increasingly large. To facilitate users to quickly access songs in their song libraries, the prior art provides users with various ways to achieve quick song search, including tagging using multiple preset attributes of songs, batch classification using playlists, indexing using other manual tags, etc. These methods can theoretically simplify the complexity of users' retrieval of target songs to a certain extent. However, these methods generally have problems such as large tag granularity and incomplete song coverage. Therefore, the tag classification effect is not good.

[0003] On the other hand, if the official platform providing the song access function is responsible for tagging songs and using authoritative tags to organize songs for users, it will involve the problem of a large number of personnel with a music background manually annotating song tags, as well as the corresponding maintenance and update problems. Therefore, the following specific problems will inevitably exist:

[0004] 1. High costs, requiring long-term accumulation of professional labor costs for annotation work;

[0005] 2. Since the number of songs in the existing song library reaches tens of millions, there will be problems such as sparse song tags, and the recall of tags is not optimistic either.

[0006] In summary, due to the problem of incomplete song coverage in the existing song tag data, it is difficult to apply the existing tag data in the direction of song organization. For the song organization requirement, a new approach needs to be explored. Summary of the Invention

[0007] The primary objective of this application is to solve at least one of the above problems and provide a method for organizing and accessing song tags, and its corresponding device, computer equipment, computer-readable storage medium, computer program product, so as to assist music creation.

[0008] To meet the various objectives of this application, the following technical solutions are adopted in this application:

[0009] A method for organizing song tags provided to meet one of the objectives of this application includes the following steps:

[0010] Obtain the target song in the song library;

[0011] Query to determine multiple preset tag subsets corresponding to the target song, and merge each tag subset into a tag set. Among them, the first tag subset is the union of the own tag sets mapped by each playlist that collects this target song in the preset playlist library;

[0012] Count the frequency of each tag in the tag set of the target song hitting the own tag set corresponding to the playlist in the playlist library, and determine multiple tags that meet the preset conditions as the inductive tags of the target song.

[0013] In a refined embodiment, in the step of querying to determine multiple preset tag subsets corresponding to the target song and merging each tag subset into a tag set:

[0014] The second tag subset of this tag set is the attribute tag set mapped by the attribute data carried by this target song, and / or,

[0015] The third tag subset of this tag set includes the manual marking tags carried by this target song, and this manual marking tag is used to represent the music genre to which the target song belongs.

[0016] In an extended embodiment, before the step of obtaining the target song in the song library, the following steps are included:

[0017] Use a pre-trained neural network model to extract multiple tags semantically related to the description text of each playlist in the playlist library, and store the extracted multiple tags as the own tag set corresponding to this playlist.

[0018] In a refined embodiment, the step of using a pre-trained neural network model to extract multiple tags semantically related to the description text of each playlist in the playlist library and storing the extracted multiple tags as the own tag set corresponding to this playlist includes the following steps:

[0019] Obtain the description text of each playlist in the playlist library;

[0020] Use the word vector model in the neural network model to extract the word vector representation of the description text;

[0021] Use the word encoding model in the neural network model to perform feature extraction on the word vector representation to obtain the corresponding text feature vector;

[0022] Use the classifier in the neural network model to classify the text feature vector to obtain multiple tags in the classification result, and store the multiple tags as the own tag set corresponding to this playlist.

[0023] In an extended embodiment, the method for sorting song tags in this application further includes the following steps of pre-training the neural network model:

[0024] Obtain a data set, which includes a plurality of training samples, and each training sample contains a description text corresponding to a playlist;

[0025] Perform rule matching on the description text of each training sample to obtain the label corresponding to the description text as the supervised label of the description text, so that the data set constitutes a training set;

[0026] In a single training task, use any training sample in the training set to train the neural network model, and obtain the classification label after classifying the description text of the training sample;

[0027] Calculate the loss value of the classification label according to the supervised label in the training sample, perform gradient update on the neural network model with the loss value, enable the next training sample in the training set to perform iterative training on the neural network model until the neural network model converges to complete the training of the current task, and obtain the intermediate state instance of the neural network model;

[0028] Use the intermediate state instance of the neural network model to classify and predict the description text of each training sample in the training set, use the classification structure to match the preset threshold to determine the classification label corresponding to each training sample, and update the supervised label of the corresponding training sample with the classification label to complete the update of the training set;

[0029] Enable the next training task, and use the updated training set to perform task iterative training on the neural network model until the intermediate state instance of the neural network model reaches the convergence state or its prediction accuracy reaches the preset threshold.

[0030] In a further embodiment, after the step of counting the frequency of each label in the label set of the target song hitting the corresponding own label set of the playlists in the playlist library, and determining multiple labels whose frequencies meet the preset conditions as the induction labels of the target song, the following steps are further included:

[0031] Respond to the user's song label sorting request, and obtain a song list composed of the songs collected by the user;

[0032] Obtain the induction labels held by each favorite song in the song list;

[0033] Cluster the favorite songs in the song list according to the induction labels, and store the mapping relationship data between each induction label and its subordinate favorite songs;

[0034] Push the induction list containing all induction labels to the user's terminal device for display.

[0035] In an extended embodiment, after the step of pushing the induction list containing all induction tags to the user's terminal device for display, the following steps are included:

[0036] In response to a song query request submitted by the user for any induction tag in the induction list, obtain an induction song list composed of the favorite songs corresponding to the induction tag according to the mapping relationship data;

[0037] Push the induction song list to the user's terminal device for display.

[0038] A song tag access method provided to meet one of the purposes of the present application includes the following steps:

[0039] In response to an operation instruction acting on a song sorting control, send a song tag sorting request to the server to obtain an induction list for encapsulating the induction tags of the favorite songs of the current user returned by the server after sorting the favorite songs;

[0040] Parse and display each induction tag in the induction list, and parse each induction tag into a control suitable for responding to a touch operation;

[0041] In response to an operation instruction acting on any of the induction tags, send a song query request to the server to obtain an induction song list composed of the favorite songs carrying the induction tag queried by the server according to the song query request;

[0042] Parse and display the favorite songs in the induction song list, and parse each favorite song into a form suitable for playing in response to a touch operation.

[0043] In a preferred embodiment, the induction tags in the induction list belong to the common tags of the own tag sets mapped to the song lists to which each of the subordinate songs belong.

[0044] A song tag sorting device provided to meet one of the purposes of the present application includes: a song library calling module, a tag query module, and a tag induction module. Among them, the song library calling module is used to obtain the target songs in the song library; the tag query module is used to query and determine multiple preset tag subsets corresponding to the target songs, and merge each tag subset into a tag set, where the first tag subset is the union of the own tag sets mapped to each song list that collects the target song in the preset song list library; the tag induction module is used to count the frequency of each tag in the tag set of the target song hitting the own tag set corresponding to the song list in the song list library, and determine multiple tags that meet the preset conditions as the induction tags of the target song.

[0045] A song tag access device provided to meet one of the purposes of the present application includes: a one-key sorting module, a tag listing module, a tag detailed search module, and a song display module. Among them, the one-key sorting module is used to send a song tag sorting request to the server in response to an operation instruction acting on a song sorting control, so as to obtain a summary list for encapsulating the summary tags of the user's favorite songs returned by the server after sorting the favorite songs of the current user; the tag listing module is used to parse and display each summary tag in the summary list and parse each summary tag into a control suitable for responding to a touch operation; the tag detailed search module is used to send a song query request to the server in response to an operation instruction acting on any of the summary tags, so as to obtain a summary song list composed of the favorite songs carrying the summary tag queried by the server according to the song query request; the song display module is used to parse and display the favorite songs in the summary song list and parse each favorite song into a form suitable for playing in response to a touch operation.

[0046] A computer device provided to meet one of the other purposes of the present application includes a central processing unit and a memory. The central processing unit is used to call and run a computer program stored in the memory to execute the steps of the song tag sorting and access method described in the present application.

[0047] A computer-readable storage medium provided to meet another purpose of the present application stores a computer program implemented according to the song tag sorting and access method in the form of computer-readable instructions. When the computer program is called and run by a computer, it executes the steps included in the method.

[0048] A computer program product provided to meet another purpose of the present application includes a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the method described in any embodiment of the present application are implemented.

[0049] Compared with the prior art, the advantages of the present application are as follows:

[0050] First, the present application utilizes multiple playlists to which a target song in a song library belongs to determine a label set composed of the own label sets corresponding to each playlist of the target song. Then, for each label in the label set, the frequency of hitting the own label set of the playlists in the playlist library is queried, so as to implement voting for each label in the label set corresponding to each target song. Then, by using the frequency levels of the labels in the label set, some labels that meet the preset conditions are determined as the induction labels of the corresponding target song. Since the voting unit is the own label set of the playlists in the playlist library, the higher the hit frequency of a label, the more suitable the label is for describing the corresponding target song according to the semantics of the description text of the playlists in which the target song is collected by the public. Therefore, it is more accurate to describe the target song with this label. Accordingly, efficient and accurate annotation of the songs in the song library can be achieved.

[0051] Secondly, in the present application, the labels for determining the target song are from the own label sets corresponding to the playlists. Generally, the description texts of the playlists are described in words such as scenes and moods. Therefore, the labels in their own label sets generally also carry corresponding scene and mood description words. The playlists have the characteristics of user-defined, and can describe the corresponding scenes, moods, etc. more precisely and comprehensively. Therefore, the corresponding labels often also carry these characteristics. The labels obtained accordingly have a finer granularity, are more in line with the natural thinking habits of users, and thus can more accurately reflect the classification and annotation effect, and can also overcome the deficiency of label sparsity and facilitate the improvement of label recall rate.

[0052] In addition, using the characteristics of the playlists to label the songs in the song library can reduce the necessity of manual annotation and effectively save the annotation cost of the songs. For the platform party that needs to process a large number of music files, it can not only greatly improve the sorting efficiency of the entire song library, but also save a huge amount of manual annotation cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The above and / or additional aspects and advantages of the present application will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:

[0054] Figure 1 is a schematic flowchart of a typical embodiment of the song label sorting method of the present application;

[0055] Figure 2 is a schematic flowchart of adding a pre-extracted own label set of the playlist in the embodiment of the present application;

[0056] Figure 3 is a schematic architecture diagram of the neural network model adopted in the embodiment of the present application;

[0057] Figure 4It is a flowchart showing the working process of the neural network model in the embodiments of the present application;

[0058] Figure 5 It is a flowchart showing the training process of the neural network model in the embodiments of the present application;

[0059] Figure 6 It is a flowchart showing the process of pushing the induction list in response to the song tag sorting request in the embodiments of the present application;

[0060] Figures 7(a), 7(b), and 7(c) are schematic diagrams of the graphical user interface of the terminal device in the embodiments of the present application, respectively showing the song sorting control interface, the induction label display interface, and the induction song list interface;

[0061] Figure 8 It is a flowchart showing the process of pushing the induction list and the induction song list in response to the song tag sorting request in the embodiments of the present application;

[0062] Figure 9 It is a flowchart showing a typical embodiment of the song tag access method of the present application;

[0063] Figure 10 It is a schematic block diagram of the song tag sorting device of the present application;

[0064] Figure 11 It is a schematic block diagram of the song tag access device of the present application;

[0065] Figure 12 It is a schematic diagram of the structure of a computer device adopted by the present application. Detailed implementation manners

[0066] The embodiments of the present application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be construed as a limitation to the present application.

[0067] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of this application means the presence of the stated features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more of the associated listed items.

[0068] Those skilled in the art can understand that, unless otherwise defined, all terms used herein (including technical terms and scientific terms) have the same meaning as the general understanding of those of ordinary skill in the art to which this application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless specifically defined as here.

[0069] Those skilled in the art can understand that the "client", "terminal", and "terminal device" used herein include both devices with wireless signal receivers that only have the ability to receive and no ability to transmit, and devices with receiving and transmitting hardware that can perform two-way communication on a two-way communication link. Such devices may include: cellular or other communication devices such as personal computers, tablet computers, etc., which have a single-line display or a multi-line display or a cellular or other communication device without a multi-line display; PCS (Personal Communications Service), which can combine voice, data processing, fax, and / or data communication capabilities; PDA (Personal Digital Assistant), which may include a radio frequency receiver, a pager, Internet / intranet access, a web browser, a notepad, a calendar, and / or a GPS (Global Positioning System) receiver; conventional laptop and / or palm computers or other devices, which are conventional laptop and / or palm computers or other devices with and / or including a radio frequency receiver. The "client", "terminal", and "terminal device" used herein can be portable, transportable, installed in a vehicle (air, sea, and / or land), or suitable for and / or configured to run locally, and / or run in a distributed form at any other location on the earth and / or in space. The "client", "terminal", and "terminal device" used herein can also be a communication terminal, an Internet access terminal, a music / video playback terminal, for example, it can be a PDA, a MID (Mobile Internet Device), and / or a mobile phone with music / video playback function, or it can also be a smart TV, a set-top box, etc.

[0070] The hardware referred to by names such as "server", "client", and "service node" in this application is essentially an electronic device with the equivalent capabilities of a personal computer, and is a hardware device with the necessary components disclosed by the von Neumann principle, including a central processing unit (including an arithmetic unit and a controller), a memory, an input device, and an output device. The computer program is stored in its memory, and the central processing unit loads the program stored in the external memory into the internal memory for execution, executes the instructions in the program, and interacts with the input / output devices to complete specific functions.

[0071] It should be noted that the concept of "server" in this application can similarly be extended to apply to server clusters. According to the network deployment principles understood by those skilled in the art, the servers should be logically divided. Physically, these servers can either be independent of each other but can be called through interfaces, or integrated into a physical computer or a set of computer clusters. Those skilled in the art should understand this flexibility and should not be restricted by this when implementing the network deployment method of this application.

[0072] One or several technical features of this application, unless expressly specified, can either be deployed on the server for implementation and accessed by the client remotely calling the online service interface provided by the server, or directly deployed and run on the client for implementation and access.

[0073] The neural network models cited or possibly cited in this application, unless expressly specified, can either be deployed on a remote server and remotely called on the client, or deployed on a client with sufficient device capabilities for direct calling. In some embodiments, when it runs on the client, its corresponding intelligence can be obtained through transfer learning to reduce the requirements for the client's hardware operation resources and avoid over-occupying the client's hardware operation resources.

[0074] All kinds of data involved in this application, unless expressly specified, can either be remotely stored on the server or stored on the local terminal device, as long as it is suitable for being called by the technical solution of this application.

[0075] Those skilled in the art should be aware that although the various methods of this application are described based on the same concept and thus show commonality with each other, unless otherwise specified, these methods can all be executed independently. Similarly, for the various embodiments disclosed in this application, they are all proposed based on the same inventive concept. Therefore, for the same expressed concepts, as well as concepts that are only appropriately transformed for convenience although the concept expressions are different, they should be equally understood.

[0076] For the various embodiments to be disclosed in this application, unless expressly pointed out that there is a mutually exclusive relationship between them, the relevant technical features involved in each embodiment can be cross-combined to flexibly construct new embodiments, as long as this combination does not deviate from the creative spirit of this application and can meet the requirements in the prior art or solve certain deficiencies in the prior art. Those skilled in the art should be aware of this flexibility.

[0077] A method for sorting song tags in this application can be programmed into a computer program product and deployed to run in a server. Thereby, the client can access the interface opened after the computer program product runs in the form of a web program or an application program, and achieve human-computer interaction with the process of the computer program product through the graphical user interface.

[0078] See also Figure 1 The song tag arrangement method of the present application, in its typical embodiment, includes the following steps:

[0079] Step S3100: Get the target song in the music library:

[0080] The backend server of the platform that provides online music download and playback Internet services maintains a music library consisting of a large number of song files. The data storage technology used in the music library can be flexibly selected and deployed, and can be distributed or centrally stored, without affecting the implementation of this application. The data format of the songs in the music library can cover all known or unknown formats, including but not limited to the following common audio formats: mp3, aac / mp4, ape / flac, wav, wma, amr, mid, in addition, considering that more and more songs are associated with video content and implemented as audio and video streaming media, therefore, the data format of the songs referred to in this application also includes but is not limited to the following common video formats: mp4 / m4v / 3gp / mpg, flv / f4v / swf, avi, gif, wmv, rmvb, mov, mts / m2t, webm / ogg / mkv. It should be understood that the songs in the music library of the present application refer to the content displayed after the corresponding multimedia files are played. The content displayed is music audible to human ears, usually a so-called song, including with or without lyrics. As for its storage form and storage format, it does not affect the embodiment of the creative spirit of the present application.

[0081] The massive amount of songs in the music library can be collected by different users on the platform so that different users can quickly access their own collected songs. The mapping relationship data between users and their collected songs is maintained by the background server. Based on this, users can access the songs they have legally obtained through the relevant mapping relationship data.

[0082] In this application, the songs in the music library can be labeled and each song can be marked to achieve classification and quick access to the songs. To this end, the target songs in the music library that need to be labeled need to be processed accordingly. In this regard, the target songs that need to be labeled can be called accordingly.

[0083] Step S3200: query and determine multiple preset tag subsets corresponding to the target song, and merge the tag subsets into a tag set, wherein the first tag subset is a collection of the own tag sets mapped to the various playlists in the preset playlist library that collect the target song:

[0084] In order to achieve the tagging arrangement of target songs, this application determines the tag set corresponding to each target song by means of the proprietary tag set carried by the playlist to which the target song belongs.

[0085] Generally speaking, the platform server supports encapsulating multiple songs in the form of a playlist. Each playlist generally uses a descriptive text as the title and encapsulates multiple songs selected from the music library. In terms of the background array organization, it is only necessary to establish the mapping relationship between the descriptive text and each corresponding song. Usually, the vocabulary used in the descriptive text of the playlist is relatively flexible, and various types of content such as scenes and moods can be described through various words to represent the common characteristics of the songs in the playlist, so that users can clearly understand the commonalities among the songs in it at a glance through the playlist title.

[0086] The presentation forms of playlists are also diverse. For example, playlists can be recommended by the platform or user-defined. Therefore, they can be presented as a recommended list on the platform, as an album list produced by a singer or a record company, or as a user-defined playlist list in different forms.

[0087] In order to describe the characteristics of each playlist, the platform can pre-tag each playlist so that each playlist has a proprietary tag set, which is composed of one or more tags. The tags in the proprietary tag set of the playlist can either be keywords extracted by the platform according to rules, defined by users themselves, or even obtained through classification mapping by extracting the deep semantic information of the descriptive text of the playlist as disclosed in the subsequent embodiments of this application.

[0088] Therefore, a vast number of playlists constitute a playlist library, which can be accessed and called by the background server. The playlist library can store the mapping relationship data between each playlist and its corresponding proprietary tag set, so that the corresponding proprietary tag set can be quickly obtained through a playlist.

[0089] To obtain the set of tags corresponding to each target song, in this embodiment, it can be achieved by querying the proprietary tag set of the playlist to which the target song belongs. It is not difficult to understand that the same song may be collected in different playlists. Therefore, the same target song may belong to multiple playlists, and each playlist has its own corresponding proprietary tag set. Therefore, for the same target song corresponding to multiple playlists, by querying the playlist library, multiple proprietary tag sets corresponding one by one to the multiple playlists to which the target song belongs can be obtained. The set of these proprietary tag sets constitutes the first tag subset of the tag set of the target song, that is, the first tag subset is the collection of multiple proprietary tag sets mapped by the corresponding target song. For the case where the same tag is included in multiple proprietary tag sets, it is naturally merged into the same tag. Thus, for a target song, its corresponding tag set can be obtained. In this embodiment, the tag set is mainly composed of the first tag subset therein.

[0090] In other embodiments implemented by making adaptations on the basis of this typical embodiment, the tag set corresponding to the target song itself may further include a second tag subset and / or a third tag subset.

[0091] The second tag subset mentioned above may be an attribute tag set composed of tags mapped by the attribute data carried by each target song itself, and can be obtained by those skilled in the art through rule matching or semantic mapping according to the attribute data carried by the song. For example: the release date of the song can be associated with the tag: classic old song; the language of the song can be associated with language tags such as Chinese, English, Cantonese, etc., and the song beat BPM (Beat Per Minute) can be associated with tags: fast song, slow song, etc. Those skilled in the art can set it flexibly accordingly to provide a classification reference based on the song's own attributes for the target song.

[0092] The third tag subset mentioned above may be the manually marked tags carried by each target song itself, which are used to represent the music genre to which the target song belongs. For example, genre tags include: ancient style, rock, folk, electronic music, DJ, etc. Generally speaking, such tags can be presented as professional classification tags with the support of professional personnel in the background, and are used to provide a professional classification reference for tagging the target song.

[0093] Both the second tag subset and the third tag subset can be flexibly implemented by those skilled in the art, and can subsequently form the tag set of the target song together with the inductive tags determined by this application for the target song, and be provided to the terminal device for calling and accessing.

[0094] Through the above process, it can be seen that the back-end server can obtain the corresponding tag set for each song in the music library through query. This tag set is not optimized and is the full set of tags extracted for each target song according to the above business logic. In theory, these tags describe the target song from different perspectives, so as to achieve tagging of the target song. This full set of tag sets can also be stored as mapping relationship data with the target song for subsequent calls.

[0095] Step S3300: Count the frequency of each tag in the tag set of the target song hitting the corresponding own tag set of the playlists in the playlist library, and determine multiple tags whose frequencies meet the preset conditions as the inductive tags of the target song.

[0096] To improve the refinement degree of the tagging description of the target song, this step is achieved by filtering and optimizing the tag set of each target song. In this exemplary embodiment, mainly the first tag subset in the tag set is filtered. Specifically, for each target song, count the frequency of each tag in its corresponding tag set hitting the corresponding own tag of the playlists in the playlist library, that is, use each own tag set to vote on each tag in the tag set of the target song. The frequency is the corresponding number of votes, so as to obtain a frequency sequence composed of the frequencies corresponding to each tag in this target song. Then, reverse sort the tags in the tag set of this target song according to the frequency so that the tags in the frequency sequence are arranged from high to low according to the frequency. Finally, select a predetermined number of N tags from the sorted frequency sequence as the final inductive tags corresponding to this target song, forming the inductive tag set of this song.

[0097] In the above example, the inductive tags selected from the frequency sequence are matched with a preset condition. The preset condition is set to preferably select the top N tags with the highest frequencies as the inductive tags of the target song, where N is a positive integer greater than 0. In other alternative exemplary embodiments, the preset condition can also be configured to preferably select tags with frequencies greater than a given preset threshold as inductive tags. And so on, indicating that the preset condition can be flexibly set and programmed by those skilled in the art, without affecting the implementation of this application.

[0098] It should be noted that in the embodiments where the tag set of the target song includes the second tag subset and the third tag subset, if the tag systems of the second tag subset and the third tag subset are independent of each other than the first tag subset, then when performing frequency statistics, neither the second tag subset nor the third tag subset needs to participate in the statistics, but can be directly output as inductive tags subsequently. On the contrary, if the first tag subset is compatible with the tags in the second tag subset and / or the third tag subset, the tags in the second tag subset and the third tag subset can be regarded as the tags in the first tag subset for frequency statistics, which does not affect the implementation of this application either.

[0099] According to the above process, for each target song in the song library, its corresponding inductive tag set can be obtained, that is, each target song obtains one or more corresponding inductive tags. Based on this, the mapping relationship data between each target song and its inductive tags can be associated and stored to quickly obtain the corresponding inductive tags of each target song in the song library subsequently. It is not difficult to understand that the same inductive tag may be mapped to multiple target songs in the song library, and the same target song generally will also be mapped to multiple inductive tags. Between the target song and the inductive tag, a many-to-many data organization form is constituted, and through the target song or through the inductive tag, mutual indexing can be achieved, providing a technical basis for finding inductive tags based on the target song or vice versa for finding songs based on inductive tags.

[0100] Through the above disclosure of this typical embodiment, it can be seen that the implementation of this application has many positive effects, including but not limited to the following aspects:

[0101] First of all, this application uses multiple playlists to which the target songs in the song library belong to determine the tag set composed of the own tag sets corresponding to each playlist of the target song. Then, for each tag in this tag set, query the frequency of the own tag set of the playlist hit in the playlist library, so as to realize the voting for each tag in the tag set corresponding to each target song. Then, use the frequency of each tag in the tag set to determine some tags that meet the preset conditions as the inductive tags of the corresponding target song. Since the voting unit is the own tag set of the playlist in the playlist library, the higher the hit frequency of a tag, the more suitable the tag is for describing its corresponding target song according to the semantics of the description text of the playlist in which the target song is collected by the public. Therefore, it is more accurate to describe the target song with this tag. Based on this, efficient and accurate annotation of the songs in the song library can be realized.

[0102] Secondly, in the present application, the tags used to determine the target song are from the proprietary tag set corresponding to the playlist. Generally, the description text of the playlist is described in words such as scenes and moods. Therefore, the tags in its proprietary tag set generally also carry corresponding scene and mood description words. The playlist has the characteristics of user-defined, and can describe the corresponding scenes, moods, etc. more precisely and comprehensively. Therefore, the corresponding tags often also carry these characteristics. The tags obtained accordingly have finer granularity and are more in line with the natural thinking habits of users. Therefore, they can more accurately reflect the classification and annotation effect, and can also overcome the deficiency of tag sparsity, which is convenient for improving the tag recall rate.

[0103] In addition, using the characteristics of the playlist to label the songs in the song library can reduce the necessity of manual annotation and effectively save the annotation cost of the songs. For the platform party that needs to process a large number of music files, it can not only greatly improve the sorting efficiency of the entire song library, but also save a huge amount of manual annotation cost.

[0104] Please refer to Figure 2 , in the extended embodiment, before the step S3100 of obtaining the target song in the song library, the following steps are included:

[0105] Step S2000: Use a pre-trained neural network model to extract multiple tags semantically related to the description text of each playlist in the playlist library, and store the extracted multiple tags as the corresponding proprietary tag set of the playlist:

[0106] In the present application, in order to obtain the proprietary tag set corresponding to the playlists in the playlist library, it is realized by classifying the deep semantic information extracted from the playlists using a pre-trained neural network model. After being classified by the neural network model, the classification probabilities corresponding to each tag in the preset tag system for the mapped deep semantic information are obtained, and multiple tags with classification probabilities exceeding the preset threshold are preferably selected, which can constitute the tags corresponding to the playlist. Using these tags to form the corresponding proprietary tag set of the playlist. Accordingly, each playlist in the playlist library can obtain its corresponding proprietary tag set. The extraction of the proprietary tag set can be realized by the server in the background.

[0107] The basic principle of the neural network model is to perform representation learning on the description text in the playlist, extract its deep semantic information, and then classify it. Therefore, it can be implemented by means of neural network models such as CNN and RNN. For example, models such as Bert, AlBert, and Electra are used to extract the deep semantic information of the description text, and Softmax is used for multi-classification processing based on the deep semantic information. Those skilled in the art can understand that as long as an appropriate neural network model is adopted according to the above principle and trained on this basis to enable it to learn the network architecture for obtaining the corresponding own label set of the playlist according to the description text of the playlist, it can be used in this application to extract the own label set for the playlist. After using this pre-trained neural network model to extract the own label set corresponding to each playlist, the playlist and the own label set can be constructed into mapping relationship data for storage for future invocation.

[0108] In a further embodiment, in step S2000, a pre-trained neural network model is used to extract multiple labels semantically related to the description text of each playlist in the playlist library, and the multiple extracted labels are stored as the corresponding own label set of the playlist. Please refer to Figure 3 , in this embodiment, a more specific neural network model architecture is adopted, in which AlBert is used to represent the description text as word vectors, and TextCNN is used to sort out the context of the word vectors. By utilizing the lightweight characteristics of the model, a more efficient label extraction effect is achieved. Accordingly, in this embodiment, as Figure 4 shown, step S2000 further includes the following steps:

[0109] Step S2100: Obtain the description text of each playlist in the playlist library:

[0110] Figure 3 In the model shown, for each playlist in the playlist library, the description file of the playlist is used as the input. Since the description text in the playlist is relatively concise, complex preprocessing operations do not need to be set, or appropriate text preprocessing can be performed, such as removing emoticons, punctuation marks, etc.

[0111] Step S2200: Use the word vector model in the neural network model to extract the word vector representation of the description text:

[0112] The AlBert model adopted by the neural network model in this embodiment is a lightweight model. AlBert is an open-source deep learning model framework by Google and is commonly used for text pre-training in the field of natural language processing NLP. In this embodiment, it is responsible for representing the description text as word vectors, encoding the description file into word vectors to initially obtain the deep semantic information of the playlist, and then providing it to TextCNN for context sorting.

[0113] Step S2300: Use the word encoding model in the neural network model to extract features from the word vector representation to obtain corresponding text feature vectors:

[0114] The word encoding model in the neural network model, namely the TextCNN model, encodes and combs the context information in the word vector representation corresponding to the description text of the playlist through one-dimensional convolution, so as to obtain the text feature vector corresponding to each playlist description text.

[0115] Step S2400: Use the classifier in the neural network model to classify the text feature vectors, obtain multiple labels in the classification result, and store the multiple labels as the corresponding self-owned label set of the playlist:

[0116] After the word encoding model obtains the text feature vector corresponding to the description text of the playlist, the text feature vector is fully connected and mapped to the classification space, and the multi-classifier calculates the classification probabilities corresponding to each label in the preset label classification system for the mapped text feature vector to obtain the classification result. At this point, according to the preset threshold corresponding to the classification probability, multiple labels with classification probabilities higher than the preset threshold can be screened out to form the corresponding self-owned label set of the playlist.

[0117] Here, the introduction of related embodiments that use a neural network model to construct the self-owned label set of the playlist enriches the technical materials required for tagging target songs in this application. Since the neural network model can highly intelligently and accurately map the description text of the playlist to multiple labels in the label system after pre-training, it is possible to save the labor annotation cost and can be fully automated for a large number of playlists, greatly improving the efficiency of obtaining the corresponding self-owned label set using the playlist.

[0118] Furthermore, since the playlist and the labels in its self-owned label set are mapped based on deep semantic information, and the vocabulary used in the playlist itself has the characteristic of flexible definition, the label data coverage is relatively comprehensive, which is convenient for extracting more refined labels, so that when tagging the target songs in the song library later, the corresponding inductive labels can more accurately describe the characteristics of the target songs.

[0119] Please refer to Figure 5 , in the extended embodiment, the song label sorting method of this application further includes the following steps of pre-training the neural network model used in the previous embodiment:

[0120] Step S1100: Obtain a data set, which includes multiple training samples, and each training sample contains the description text corresponding to a playlist:

[0121] For the platform side, playlists in its back-end database or playlists obtained externally can be used as data sets for preparing the training sets required for training the neural network model. Each playlist in the data set constitutes a corresponding training sample, and each training sample contains the description text of its corresponding playlist. Preferably, appropriate text preprocessing operations can be performed on the description file, including removing emoticons, punctuation marks, etc.

[0122] Step S1200: Perform rule matching on the description text of each training sample to obtain the label corresponding to the description text as the supervised label of the description text, so that the data set constitutes a training set:

[0123] In order to obtain the supervised label corresponding to each training sample, a corresponding label system can be pre-constructed. In one implementation, based on the principle of rule matching, it is determined whether there is text corresponding to the label in the preset label system in the description file of each training sample. If so, the corresponding label is marked as the supervised label corresponding to the description text. After all the training samples in the data set are associated with their corresponding supervised labels, a training set is formed, which can be used to perform supervised training on the neural network model.

[0124] It is not difficult to understand that the labels corresponding to the playlists obtained by rule matching are relatively simple. To improve efficiency, the rules applied may be relatively simple. Therefore, the number of labels corresponding to the playlists may not be comprehensive enough. Therefore, the supervised labels of the training samples in the training set can be continuously improved for iterative improvement later.

[0125] Step S1300: In a single training task, use any training sample in the training set to train the neural network model to obtain the classification label after classifying the description text of the training sample:

[0126] In this embodiment, the neural network model can be iteratively trained through multiple training tasks to continuously upgrade the intermediate state instances of the neural network model and finally obtain a mature and usable instance.

[0127] In a training task, by traversing each training sample in the training set, the description text of each training sample is used as the input of the neural network model. The neural network model extracts the text feature vectors corresponding to the deep semantic information of the neural network model for classification to obtain the corresponding classification results. Then, the supervision label corresponding to the training sample is used to supervise the classification results of the neural network model. In the classification results, one or more classification labels show relatively high classification probabilities. Among them, the classification labels with relatively high classification probabilities can be regarded as the classification labels mapped by the corresponding playlists, and the supervision label can be used to supervise them. The process of the neural network model extracting deep semantic information based on the description text can be referred to the disclosure of the previous embodiment and will not be elaborated here.

[0128] Step S1400: Calculate the loss value of the classification label according to the supervision label in the training sample, perform gradient update on the neural network model with this loss value, enable the next training sample in the training set to perform iterative training on the neural network model until the neural network model converges to complete the training of the current task, and obtain the intermediate state instance of the neural network model:

[0129] When using the supervision label to supervise the classification results of the neural network model, specifically, the loss value between the classification labels obtained from the training sample can be calculated according to the supervision label in the corresponding training sample, and then the neural network model is subjected to gradient update according to this loss value to correct its network weights so that the loss function of the neural network model continuously approaches convergence. Since classification is performed using a classifier, the cross-entropy loss function can be used to calculate the loss value here.

[0130] In each training task, it can be determined whether to continue to call the next training sample to continue training the neural network model by calculating whether the calculated loss value is close to 0, whether it reaches a preset range, or whether all training samples in the training set have been called. Accordingly, if any of the above judgment conditions is satisfied, the training of the current training task can be ended. Otherwise, return to step S1300, enable the next training sample in the training set, and continue to perform iterative training on the neural network model to make the neural network model continuously approach the convergence state through training.

[0131] When the neural network model is trained to convergence, the training of the current training task is completed, and thus the intermediate state instance of the neural network model after executing this training task is obtained. Theoretically, if for the sake of simplicity, this intermediate state instance can also be directly used in this application. Otherwise, the learning ability of the neural network model can also be further improved through the subsequent steps of this embodiment.

[0132] Step S1500: Use the intermediate state instance of the neural network model to perform classification prediction on the description texts of each training sample in the training set, determine the classification label corresponding to each training sample by using the classification structure to match a preset threshold, and update the supervision label of the corresponding training sample with the classification label to complete the update of the training set:

[0133] To improve the learning ability of the neural network model, this application trains it by adding training tasks. Accordingly, based on the intermediate state instance of the neural network model obtained after any training task, use this intermediate state instance to perform classification prediction on the playlist description texts in the training samples of the training set, so as to correct the supervision label of each training sample according to the classification result predicted for each training sample.

[0134] For the classification result predicted for each training sample, it contains the classification probabilities corresponding to each label in the preset label system mapped by the description text. This classification probability also represents the confidence of the model in the corresponding label. Therefore, by comparing the preset threshold corresponding to the classification probability with each classification probability, the label with a classification probability higher than the preset threshold can be preferably selected as one or more classification labels mapped to the corresponding playlist. Then, replace the supervision label of the corresponding training sample with these classification labels, and correct each training sample one by one in this way to complete the update of the training set. It is not difficult to understand that the supervision label corresponding to the training sample in the updated training set is relatively richer and more accurate, and its richness and accuracy also depend on the setting of the preset threshold. According to this principle, those skilled in the art can adjust it flexibly.

[0135] Step S1600: Enable the next training task, and use the updated training set to perform task iterative training on the neural network model until the intermediate state instance of the neural network model reaches a convergence state or its prediction accuracy reaches a preset threshold:

[0136] After updating the supervision label in the training set with the latest intermediate state instance, a new training set is obtained. Therefore, on this basis, a new training task can be started according to actual needs, and the latest training set can be used to continue the iterative training of a new training task for the neural network model. In this embodiment, it can be cycled back to and executed in step S1300.

[0137] The basis for determining whether to start the next training task to perform iterative training on the neural network model can be to use the intermediate state instance of the neural network model obtained from the previous training task to classify and predict the training samples in the preset test set, and then determine whether to restart the training by judging whether the corresponding prediction accuracy rate reaches a corresponding preset threshold. The test set can be a part of the data separated from the training set in advance, or independently prepared with reference to the training set. In this regard, those skilled in the art can implement it flexibly.

[0138] Specifically, use the intermediate state instance of the neural network model in the previous training task to predict each training sample in the test set one by one to obtain the corresponding classification labels, then compare the classification labels with the supervision labels of the corresponding training samples, and then count the prediction accuracy rate of each training sample being accurately predicted to have the same classification label as its supervision label. Compare whether this prediction accuracy rate exceeds the preset threshold. If it exceeds the preset threshold, new training tasks can no longer be enabled for iteration, otherwise new training tasks can be enabled for iteration.

[0139] In another alternative implementation method, those skilled in the art can decide whether to restart the task iterative training according to prior knowledge and other actual measurement methods. For example, after any training task is executed, observe whether the currently obtained intermediate state instance reaches the convergence state. When it reaches the convergence state, task iterative training can no longer be enabled for the neural network model.

[0140] The neural network model whose training is terminated can be put into use in this application. Using this neural network model, the corresponding classification labels can be predicted for each playlist in the playlist library, and these classification labels constitute the own label set of the corresponding playlist. It can be seen that the implementation of this embodiment enriches the technical effects of this application, at least manifested in the following aspects:

[0141] First of all, the neural network model of this embodiment extracts deep semantic information based on the description text of the playlist. Since the vocabulary in the playlist description file is rich and diverse, a more comprehensive label system with finer description granularity can be constructed. Therefore, the classification ability obtained after the neural network model is trained can determine finer labels based on the playlist to construct the own label set of the playlist. The fine and diverse descriptions are more convenient for retrieving target songs.

[0142] Secondly, during the training process of the neural network model in this embodiment, multiple task-based iterative trainings are performed using the same training set. The predicted classification labels of the training samples in the training set are used to update the corresponding supervision labels through the intermediate state instances obtained in each training to update the training set. Then, the intermediate state instances obtained in each training are used to predict the test set, and then it is analyzed whether the prediction accuracy reaches the standard. Accordingly, it is decided whether to start a new task iterative training. In the new training task, the updated training set is enabled. In this way, through cyclic iteration, only an appropriate amount rather than a large amount of training samples are needed to continuously improve the labeling ability of the neural network model, which not only saves training costs but also ensures the classification accuracy of the neural network model.

[0143] In addition, in this embodiment, the neural network model can map the corresponding labels based on the deep semantic information of the description text. The network architecture adopted is relatively simple and easy to implement. For the platform side, the cost is controllable and the efficiency is higher, which is suitable for in-depth data mining of the background data of the platform side to reflect the corresponding data value of the playlist.

[0144] Please refer to Figure 6 , in the further embodiment, after the step S3300 of statistically analyzing the frequency of each label in the label set of the target song hitting the corresponding own label set of the playlists in the playlist library, and determining multiple labels whose frequencies meet the preset conditions as the induction labels of the target song, each target song has its own corresponding induction label set. On this basis, the service of retrieving the songs collected by the terminal user by the induction label can be opened. To implement this function, in this embodiment, the song label sorting method further includes the following steps:

[0145] Step S4100, in response to the user's song label sorting request, obtain the song list composed of the songs collected by the user:

[0146] As shown in FIG. 7(a), the user can trigger a song label sorting request with one key in the graphical user interface displayed by the application installed on the terminal device. This request is sent to the server of the present application, and after receiving the request, the server responds to it.

[0147] The server responds to the above request and calls the corresponding personal song library data of the user. The personal song library data is composed of the songs collected by the user himself. The forms of the user's collection include but are not limited to the following ways: free purchase, paid purchase, adding to the favorites, adding to the playlist, etc. In this embodiment, for the convenience of understanding, the songs in the personal song library data can be understood as all the songs that can be played by the user associated with the user's personal account.

[0148] Accordingly, the server can regard the personal music library data corresponding to the user, that is, the set composed of all the songs associated with the user's personal account, as a song list, so as to sort out the corresponding induction labels for the user based on each favorite song in this song list.

[0149] Step S4200: Obtain the induction labels held by each favorite song in the song list:

[0150] As described above, in each of the foregoing embodiments of the present application, the operation of labeling induction labels has been performed on all the songs in the background music library. Therefore, for each song in the user's song list, there is theoretically one or more induction labels mapped to it. Thus, the server can retrieve the corresponding set of induction labels for each favorite song in the user's song list.

[0151] Step S4300: Cluster the favorite songs in the song list according to the induction labels, and store the mapping relationship data between each induction label and its subordinate favorite songs:

[0152] Based on the obtained mapping relationship data from the favorite songs in the song list to their set of induction labels, the server can use any clustering algorithm to cluster the favorite songs in the song list according to the induction labels. A relatively simple method is to count the favorite songs corresponding to each induction label and establish the mapping relationship data between each induction label and its associated favorite songs. Of course, more complex clustering algorithms can also be used to achieve this, such as the K-means clustering algorithm, the mean shift clustering algorithm, the density-based clustering algorithm, the expectation-maximization (EM) clustering algorithm using the Gaussian mixture model (GMM), the agglomerative hierarchical clustering algorithm, the graph community detection algorithm, etc., which can be selected by those skilled in the art according to the actual situation. Finally, all the favorite songs in the user's song list will form a set of induction labels, which can be regarded as an induction list. Each induction label in it has at least one favorite song mapped to it. Each induction label theoretically subordinates one or more of the favorite songs in the song list, and one or more of the favorite songs may carry multiple of the above-mentioned induction labels.

[0153] Step S4400: Push the induction list containing all the induction labels to the user's terminal device for display:

[0154] After determining the induction list corresponding to the user's song list, it can be pushed to the user's terminal device. The application program running on the terminal device parses it and displays it in the graphical user interface, as shown in Fig. 7(b). Thus, the user can achieve one-key sorting of the labels of all the songs he / she has collected.

[0155] In an alternative embodiment, before pushing the induction list to the user, the number of favorite songs under each induction label in the induction list may be counted first to obtain the total number of songs corresponding to each induction label, and then the total number of songs is associated with the corresponding induction label and included in the induction list, and pushed to the terminal device for display together.

[0156] Combined with the foregoing alternative embodiment, if the background includes the labels of the second label subset and the third label subset in the induction labels of the songs, these labels may also be included as induction labels as they are and form part of the induction list to be provided to the user terminal device for display.

[0157] This embodiment facilitates the user to implement one-key sorting of the labels of all favorite songs, provides a one-stop service for the terminal user to establish an indexing mechanism for sorting their favorite songs, can greatly improve the efficiency of the user sorting their favorite songs, enables the user to obtain more refined induction labels, and significantly enhances the user experience.

[0158] Please refer to Figure 8 , in an extended embodiment, after the step S4400 of pushing the induction list containing all induction labels to the terminal device of the user for display, the following steps are included:

[0159] Step S4500: Respond to the song query request submitted by the user for any induction label in the induction list, and obtain an induction song list composed of the favorite songs corresponding to the induction label according to the mapping relationship data:

[0160] The user can view the respective induction labels corresponding to the song list in the interface shown in FIG. 7(b). Among them, each induction label is displayed in the form of a control. Therefore, by touching any induction label, a song query request corresponding to the induction label can be sent to the server, and the request contains the specified information for the induction label to obtain the favorite songs carrying the induction label.

[0161] After the background server receives the song query request triggered by the user, in response to the request, according to the specified induction label, it queries the favorite songs carrying the corresponding induction label from the user's personal song library data and constructs an induction song list.

[0162] Step S4600: Push the induction song list to the terminal device of the user for display:

[0163] As shown in FIG. 7(c), after the server pushes the induction song list corresponding to the specified induction label to the terminal device where the user is located, the corresponding application program running on the terminal device parses the induction song list, and then outputs the induction song list in the graphical user interface, displays each favorite song in the list, and assigns corresponding play functions to each favorite song. By touching any favorite song through a touch operation, the user can execute corresponding functions on the favorite song, such as playing, downloading to the local, etc.

[0164] The human-computer interaction process based on induction labels further extended in this embodiment enables the user to not only access the songs he / she has collected based on induction labels, but also use the corresponding favorite songs more conveniently, reflecting the convenience of using the favorite songs after refined classification and demonstrating the improvement of the user experience.

[0165] A song label access method of the present application can be programmed as a computer program product and deployed to run in a client. Thereby, the client can access the interface opened after the computer program product runs in the form of a web program or an application program, and realize human-computer interaction with the process of the computer program product through a graphical user interface.

[0166] For a song label access method of the present application, please refer to Figure 9 , in its typical embodiment, includes the following steps:

[0167] Step S5100: Respond to an operation instruction acting on a song sorting control and send a song label sorting request to the server to obtain an induction list for encapsulating the induction labels of the favorite songs of the current user returned by the server after sorting the favorite songs of the current user:

[0168] Please refer back to FIG. 7(a). In the graphical user interface displayed by the application program installed on the terminal device, a song sorting control is displayed to enable the user to perform one-key tagging on his / her personal music library data. The application program is generally an application program that opens services on the background server, that is, the computer program product of this method.

[0169] After the user touches the song sorting control, the application constructs a song tag sorting request in response to the corresponding touch event. After this request is sent to the server, referring to the content disclosed in the foregoing embodiments of the present application, the corresponding server responds to this song tag sorting request and sorts the inductive tags of the favorite songs in the user's personal music library data for the user to obtain an inductive list, and sends this inductive list back to the current user. Since the storage list encapsulates multiple inductive tags obtained by clustering all the favorite songs corresponding to the user's personal music library data, therefore, the inductive list actually contains the index relationship between each inductive tag and all the favorite songs in the user's personal music library data.

[0170] Step S5200: Parse and display each inductive tag in the inductive list, and parse each inductive tag into a control suitable for responding to a touch operation:

[0171] In the application of the terminal device, after obtaining the inductive list, it parses the inductive list to obtain each inductive tag therein, and displays it as the interface effect shown in FIG. 7(b). Among them, each inductive tag is presented in the form of a control, so that after any control is touched, a song query request corresponding to the inductive tag represented by it can be triggered. Among them, if the inductive list carries the total number of songs corresponding to each inductive tag, it can also be displayed in the control.

[0172] Step S5300: Respond to the operation instruction acting on any of the inductive tags and send a song query request to the server to obtain an inductive song list composed of the favorite songs carrying the inductive tag queried by the server according to this song query request:

[0173] When the user touches the control corresponding to any inductive tag, for example, the control corresponding to "broken heart" in FIG. 7(b), it is equivalent to sending an operation instruction to the server to query the favorite songs corresponding to this inductive tag. Therefore, in response to this operation instruction, the application of the terminal device triggers a song query request, which contains the specified information of the corresponding inductive tag, and then sends this song query request to the server.

[0174] According to the various embodiments disclosed above, after the server responds to the song query request, according to the specified inductive tag in the request, it queries the inductive song list corresponding to this inductive tag from the user's personal music library data. This inductive song list contains multiple favorite songs, and each favorite song carries the specified inductive tag. Then, the server sends this inductive song list back to the terminal device.

[0175] Step S5400: Parse and display the favorite songs in the summarized song list, and parse each favorite song to be suitable for playing in response to a touch operation:

[0176] After the application program of the terminal device obtains the summarized song list pushed by the server, it parses the summarized song list, and then displays the summary information of each favorite song in the graphical user interface, as shown in FIG. 7(c). Among them, corresponding play, download, share controls, etc. can be associated with each favorite song, making it suitable for being played, downloaded, shared, etc. in response to a touch operation.

[0177] It can be seen that in the summarized list obtained by the terminal device touching the song sorting control, the summarized labels belong to the common label set of the self-owned label sets mapped by the song lists to which each of the subordinate favorite songs belongs. According to this business logic, the user can quickly and conveniently sort out the label system corresponding to his personal favorite songs, which is convenient for him to quickly find favorite songs and improve the song retrieval and access efficiency.

[0178] This embodiment improves the user experience at the terminal device. The user can sort out the data of his own personal music library with one key operation, establish a classification index, and the obtained summarized labels more precisely describe the characteristics of each favorite song, making it more convenient for the user to call the song data he owns.

[0179] Please refer to Figure 10 , a song label sorting device provided by the present application is functionally deployed to adapt to the song label sorting method of the present application, including: a music library calling module 3100, a label query module 3200, and a label summarization module 3300. Among them, the music library calling module 3100 is used to obtain the target song in the music library; the label query module 3200 is used to query and determine multiple preset label subsets corresponding to the target song, and merge each label subset into a label set, where the first label subset is the union of the self-owned label sets mapped by each playlist that collects the target song in the preset playlist library; the label summarization module 3300 is used to count the frequency of each label in the label set of the target song hitting the self-owned label set corresponding to the playlist in the playlist library, and determine multiple labels that meet the preset conditions as the summarized labels of the target song.

[0180] In a further embodiment, the label query module 3200 includes: the second label subset of the label set is the attribute label set mapped by the attribute data carried by the target song, and / or, the third label subset of the label set includes the manual marking label carried by the target song, and the manual marking label is used to represent the music genre to which the target song belongs.

[0181] In an extended embodiment, the song label sorting device of the present application further includes: a model training module, configured to extract multiple labels that are semantically related to the description text of each playlist in the playlist library by using a pre-trained neural network model, and store the extracted multiple labels as the corresponding own label set of the playlist.

[0182] In a further embodiment, the model training module includes: a playlist acquisition sub-module, configured to acquire the description text of each playlist in the playlist library; a vector encoding sub-module, configured to extract the word vector representation of the description text by using the word vector model in the neural network model; a feature extraction sub-module, configured to perform feature extraction on the word vector representation by using the word encoding model in the neural network model to obtain the corresponding text feature vector; a label determination sub-module, configured to classify the text feature vector by using the classifier in the neural network model to obtain multiple labels in the classification result, and store the multiple labels as the corresponding own label set of the playlist.

[0183] In an extended embodiment, the song label sorting device of the present application further includes: a data set calling sub-module, configured to acquire a data set, where the data set includes multiple training samples, and each training sample includes the description text corresponding to a playlist; an initial annotation sub-module, configured to perform rule matching on the description text of each training sample to obtain the label corresponding to the description text as the supervised label of the description text, so that the data set forms a training set; a training execution sub-module, configured to use any training sample in the training set to train the neural network model in a single training task to obtain the classification label after classifying the description text of the training sample; a sample iteration sub-module, configured to calculate the loss value of the classification label according to the supervised label in the training sample, perform gradient update on the neural network model with the loss value, enable the next training sample in the training set to perform iterative training on the neural network model until the neural network model converges to complete the training of the current task, and obtain the intermediate state instance of the neural network model; a prediction update sub-module, configured to perform classification prediction on the description text of each training sample in the training set by using the intermediate state instance of the neural network model, determine the classification label corresponding to each training sample by using the classification structure to match a preset threshold, and update the supervised label of the corresponding training sample with the classification label to complete the update of the training set; a training set iteration sub-module, configured to enable the next training task, and perform task iterative training on the neural network model by using the updated training set until the intermediate state instance of the neural network model reaches a convergence state or its prediction accuracy reaches a preset threshold.

[0184] In a refined embodiment, the tag induction module 3300 includes: an arrangement response module for responding to a user's song tag arrangement request and obtaining a song list composed of the songs collected by the user; a tag search module for obtaining the induction tags held by each collected song in the song list; a song clustering module for clustering the collected songs in the song list according to the induction tags and storing mapping relationship data between each induction tag and its subordinate collected songs; and an induction push module for pushing an induction list containing all the induction tags to the user's terminal device for display.

[0185] In an extended embodiment, the song tag arrangement device of the present application further includes: a tag song search module for responding to a song query request submitted by the user for any induction tag in the induction list and obtaining an induction song list composed of the collected songs corresponding to the induction tag according to the mapping relationship data; and a song push module for pushing the induction song list to the user's terminal device for display.

[0186] Please refer to Figure 11 , a song tag access device provided by the present application is functionally deployed according to the song tag access method of the present application, and includes: a one-key arrangement module 5100, a tag listing module 5200, a tag detailed search module 5300, and a song display module 5400. Among them, the one-key arrangement module 5100 is used to send a song tag arrangement request to the server in response to an operation instruction acting on a song arrangement control to obtain an induction list for encapsulating the induction tags of the collected songs of the current user returned by the server after arranging the collected songs; the tag listing module 5200 is used to parse and display each induction tag in the induction list and parse each induction tag into a control suitable for responding to a touch operation; the tag detailed search module 5300 is used to send a song query request to the server in response to an operation instruction acting on any of the induction tags to obtain an induction song list composed of the collected songs carrying the induction tag queried by the server according to the song query request; and the song display module 5400 is used to parse and display the collected songs in the induction song list and parse each collected song into a form suitable for playing in response to a touch operation.

[0187] In a preferred embodiment, the induction tags in the induction list belong to the common tags of the self-owned tag sets mapped to the song phases to which each of the subordinate songs belongs.

[0188] To solve the above technical problems, an embodiment of the present application further provides a computer device. As Figure 12As shown in the figure, it is a schematic diagram of the internal structure of a computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected through a system bus. Among them, the computer-readable storage medium of the computer device stores an operating system, a database, and computer-readable instructions. The database can store a control information sequence. When the computer-readable instructions are executed by the processor, the processor can implement a method for sorting and accessing song tags. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. The memory of the computer device can store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can execute the method for sorting and accessing song tags of the present application. The network interface of the computer device is used to connect and communicate with a terminal. Those skilled in the art can understand that Figure 12 The structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0189] In this embodiment, the processor is used to execute Figure 10 and Figure 11 the specific functions of each module and its sub-modules. The memory stores the program code and various types of data required to execute the above modules or sub-modules. The network interface is used for data transmission between the user terminal and the server. The memory in this embodiment stores the program code and data required to execute all modules / sub-modules in the song tag sorting and access device of the present application. The server can call the program code and data of the server to execute the functions of all sub-modules.

[0190] The present application also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the method for sorting and accessing song tags according to any embodiment of the present application.

[0191] The present application also provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by one or more processors, the steps of the method according to any embodiment of the present application are implemented.

[0192] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments of the present application can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a computer-readable storage medium such as a magnetic disk, an optical disc, a read-only memory (ROM), or a random access memory (RAM), etc.

[0193] In summary, the present application uses a playlist as the source of label materials for target songs in a song library, and uses the corresponding self-owned label set of the playlist to label the target songs, which is convenient for preparing and summarizing labels for the songs collected by users, and facilitates users to accurately and efficiently query a large number of songs they have collected.

[0194] Those skilled in the art of the present technology can understand that the steps, measures, and solutions in the various operations, methods, and processes discussed in the present application can be alternated, changed, combined, or deleted. Further, the other steps, measures, and solutions in the various operations, methods, and processes discussed in the present application can also be alternated, changed, rearranged, decomposed, combined, or deleted. Further, the steps, measures, and solutions in the prior art that are the same as those disclosed in the various operations, methods, and processes in the present application can also be alternated, changed, rearranged, decomposed, combined, or deleted.

[0195] The above are only some implementation manners of the present application. It should be noted that for those of ordinary skill in the art of the present technology, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A method for sorting song tags, characterized in that, it includes the following steps: Using a pre-trained neural network model to extract multiple tags semantically related to the description text of each playlist in the playlist library, and storing the extracted multiple tags as the corresponding own tag set of the playlist; Obtain the target song in the music library; Query and determine multiple preset tag subsets corresponding to the target song, and merge each tag subset into a tag set. Among them, the first tag subset is the union of the own tag sets mapped by each playlist that collects the target song in the preset playlist library; Count the frequency of each tag in the tag set of the target song hitting the corresponding own tag set of the playlist in the playlist library, and determine multiple tags whose frequencies meet the preset conditions as the inductive tags of the target song; It also includes the following steps of pre-training the neural network model: Obtain a data set, which includes multiple training samples, and each training sample contains a description text corresponding to a playlist; Perform rule matching on the description text of each training sample to obtain the tag corresponding to the description text as the supervised tag of the description text, so that the data set constitutes a training set; In a single training task, use any training sample in the training set to train the neural network model, and obtain the classification tag after classifying the description text of the training sample; Calculate the loss value of the classification tag according to the supervised tag in the training sample, perform gradient update on the neural network model with this loss value, enable the next training sample in the training set to perform iterative training on the neural network model until the neural network model converges and completes the training of the current task, and obtain the intermediate state instance of the neural network model; Use the intermediate state instance of the neural network model to classify and predict the description text of each training sample in the training set, and use the classification structure to match the preset threshold to determine the classification tag corresponding to each training sample, and update the supervised tag of the corresponding training sample with the classification tag to complete the update of the training set; Enable the next training task, and use the updated training set to perform task iterative training on the neural network model until the intermediate state instance of the neural network model reaches the convergence state or its prediction accuracy reaches the preset threshold.

2. The song tag sorting method according to claim 1, characterized in that, In the step of querying and determining multiple preset tag subsets corresponding to the target song and merging each tag subset into a tag set: The second tag subset of the tag set is the attribute tag set mapped by the attribute data carried by the target song, and / or, The third tag subset of the tag set includes the manually marked tags carried by the target song, and the manually marked tags are used to represent the music genre to which the target song belongs.

3. The song tag sorting method according to claim 1, characterized in that, The step of using a pre-trained neural network model to extract multiple tags semantically related to the description text of each playlist in the playlist library and storing the extracted multiple tags as the corresponding own tag set of the playlist includes the following steps: Obtain the description text of each playlist in the playlist library; Extract the word vector representation of the description text using the word vector model in the neural network model; Perform feature extraction on the word vector representation using the word encoding model in the neural network model to obtain the corresponding text feature vector; Classify the text feature vector using the classifier in the neural network model to obtain multiple labels in the classification result, and store the multiple labels as the proprietary label set corresponding to the playlist.

4. The song label sorting method according to any one of claims 1 to 3, wherein, after the step of counting the frequency of each label in the label set of the target song hitting the proprietary label set corresponding to the playlist in the playlist library, and determining multiple labels whose frequencies meet the preset conditions as the inductive labels of the target song, the following steps are further included: Respond to the user's song label sorting request, and obtain a song list composed of the songs collected by this user; Obtain the inductive labels held by each favorite song in the song list; Cluster the favorite songs in the song list according to the inductive labels, and store the mapping relationship data between each inductive label and its subordinate favorite songs; Push an induction list containing all inductive labels to the user's terminal device for display.

5. The song label sorting method according to claim 4, wherein, after the step of pushing the induction list containing all inductive labels to the user's terminal device for display, the following steps are included: Respond to the song query request submitted by the user for any inductive label in the induction list, and obtain an inductive song list composed of the favorite songs corresponding to this inductive label according to the mapping relationship data; Push the inductive song list to the user's terminal device for display.

6. A song label access method, wherein, includes the following steps: Respond to the operation instruction acting on the song sorting control and send a song label sorting request to the server to obtain an induction list for encapsulating the inductive labels of the favorite songs of the current user returned by the server after sorting the favorite songs; the inductive labels are obtained according to the song label sorting method according to any one of claims 1 to 5; Parse and display each inductive label in the induction list, and parse each inductive label into a control suitable for responding to touch operations; Respond to the operation instruction acting on any of the inductive labels and send a song query request to the server to obtain an inductive song list composed of the favorite songs carrying this inductive label queried by the server according to this song query request; Parse and display the favorite songs in the inductive song list, and parse each favorite song into being suitable for playing in response to touch operations.

7. The song label access method according to claim 6, wherein, the inductive labels in the induction list belong to the common labels of the proprietary label sets mapped to the playlists to which the subordinate songs belong.

8. A computer device, including a central processing unit and a memory, wherein, The central processing unit is configured to call and run a computer program stored in the memory to execute the steps of the method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that it stores, in the form of computer-readable instructions, a computer program implemented according to the method according to any one of claims 1 to 7, and when the computer program is called and run by a computer, it executes the steps included in the corresponding method.

10. A computer program product, comprising a computer program / instructions, characterized in that when the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Crawler-based music labelling method and system

    CN105718575A

  • Song recommendation method and device based on label theme model as well as storage medium

    CN108334601A

  • Music classification method based on adaptive CNN and semi-supervised self-training model

    CN112487237A