Playlist generation methods, media, devices and computing equipment
By extracting song information and determining style information from playlist screenshots, a recommended playlist is generated, solving the problem that users need to edit the playlist after importing it in existing technologies, thus improving the user experience.
Patent Information
- Application Number
- CN202310090321.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-12
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2043-01-12
AI Technical Summary
Existing online music platforms and software cannot provide personalized options when importing playlists based on screenshots of external playlists, resulting in cumbersome editing operations for users and a poor user experience.
By extracting song information from playlist screenshots, the style information of the songs is determined, and based on the number of songs and style distribution, a recommended playlist containing at least one style is generated, thus achieving automatic classification and recognition of playlists.
This reduces the amount of further categorization and processing required from the original playlist, improving operational efficiency and user satisfaction during playlist migration.
Smart Images

Figure CN116049479B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this disclosure relate to the field of Internet technology, and more specifically, the embodiments of this disclosure relate to a playlist generation method, medium, apparatus, and computing device. Background Technology
[0002] This section is intended to provide background or context for embodiments of this disclosure. The description herein is not intended to imply that it is prior art simply because it is included in this section.
[0003] Many existing online music platforms and software offer the function of importing external playlists. By inputting the link or screenshot of an external playlist, the system automatically identifies the songs in the playlist and imports them into the newly created playlist. This eliminates the need to repeatedly match and add songs one by one across different platforms or software in order to listen to the songs in the playlist, thus significantly reducing the user's usage costs.
[0004] Existing methods for importing playlists based on screenshots of external playlists can only import all tracks from the screenshot and generate a playlist, without providing personalized options. This often requires users to edit the generated playlist, making the process cumbersome and resulting in a poor user experience. Summary of the Invention
[0005] This disclosure provides a method, medium, apparatus, and computing device for generating playlists, in order to solve the problem in related technologies where playlists imported from external playlist screenshots require users to re-edit them, resulting in cumbersome operations.
[0006] In a first aspect of this disclosure, a playlist generation method is provided, comprising:
[0007] In response to the received playlist screenshot, extract the song information from the playlist screenshot;
[0008] Based on the song information, determine the corresponding style information of the song;
[0009] Based on the number of songs and the distribution of song styles in the playlist screenshot, a recommended playlist containing at least one style is generated.
[0010] In a second aspect of this disclosure, a computer-readable storage medium is provided, comprising:
[0011] The computer-readable storage medium stores computer-executable instructions that, when executed by a processor, are used to implement the playlist generation method as described in the first aspect of this disclosure.
[0012] In a third aspect of this disclosure, a playlist generation apparatus is provided, comprising:
[0013] The extraction module is used to extract song information from the received playlist screenshot in response to the screenshot.
[0014] The determination module is used to determine the style information corresponding to a song based on the song information;
[0015] The generation module is used to generate recommended playlists for songs based on style information.
[0016] In a fourth aspect of this disclosure, a computing device is provided, comprising: at least one processor;
[0017] and memory that is communicatively connected to at least one processor;
[0018] The memory stores instructions that can be executed by at least one processor to cause the computing device to perform the playlist generation method as described in the first aspect of this disclosure.
[0019] According to the playlist generation method, medium, apparatus, and computing device of this disclosure, when a playlist screenshot is received, song information is extracted from the screenshot. Then, based on the song information, the corresponding style information of the songs is determined. Based on the number of songs and the style distribution in the playlist screenshot, a recommended playlist containing at least one style is generated. Therefore, based on the received external playlist screenshot, different recommended playlists can be automatically generated according to characteristics such as song style, achieving automatic classification and recognition of playlist screenshots without requiring tedious editing by the user, simplifying operation and improving user experience. Attached Figure Description
[0020] The above and other objects, features, and advantages of this disclosure will become readily apparent from the following detailed description of exemplary embodiments, taken in conjunction with the accompanying drawings. Several embodiments of this disclosure are illustrated in the drawings by way of example and not limitation, in which:
[0021] Figure 1a An application scenario diagram illustrating an embodiment of the present disclosure is shown schematically;
[0022] Figure 1b schematically shown Figure 1a A schematic diagram of the structure of a playlist screenshot in the application scenario shown;
[0023] Figure 2 A flowchart illustrating a playlist generation method according to another embodiment of this disclosure is shown schematically;
[0024] Figure 3a A flowchart illustrating a playlist generation method according to yet another embodiment of this disclosure is shown schematically;
[0025] Figure 3bschematically shown Figure 3a The flowchart of the training method of the text detection model in the illustrated embodiment;
[0026] Figure 3c schematically shown Figure 3a A flowchart illustrating the training method of the text recognition model in the illustrated embodiment;
[0027] Figure 4a A flowchart illustrating a playlist generation method according to another embodiment of the present disclosure is shown schematically;
[0028] Figure 4b schematically shown Figure 4a The flowchart for determining the semantic vector of lyrics data in the illustrated embodiment;
[0029] Figure 4c schematically shown Figure 4a The flowchart shown in the embodiment illustrates the identification of song chord style and song style;
[0030] Figure 4d schematically shown Figure 4a The flowchart of the training method for the second classification network in the illustrated embodiment;
[0031] Figure 5 A schematic diagram of the structure of a storage medium according to another embodiment of the present disclosure is shown;
[0032] Figure 6 A schematic diagram of the structure of a playlist generation apparatus according to another embodiment of the present disclosure is shown.
[0033] Figure 7 A schematic diagram of the structure of a computing device according to another embodiment of the present disclosure is shown.
[0034] In the accompanying drawings, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed Implementation
[0035] The principles and spirit of this disclosure will now be described with reference to several exemplary embodiments. It should be understood that these embodiments are given merely to enable those skilled in the art to better understand and implement this disclosure, and are not intended to limit the scope of this disclosure in any way. Rather, these embodiments are provided to make this disclosure more thorough and complete, and to fully convey the scope of this disclosure to those skilled in the art.
[0036] Those skilled in the art will recognize that embodiments of this disclosure can be implemented as a system, apparatus, device, method, or computer program product. Therefore, this disclosure can be specifically implemented in the following forms: entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0037] According to embodiments of this disclosure, a playlist generation method, medium, apparatus, and computing device are proposed.
[0038] In this document, it should be understood that the terminology used is for convenience of understanding only and does not imply any limitation on its meaning. Furthermore, any number of elements in the accompanying drawings is for illustrative purposes only and not for limitation, and any naming is for distinction only and has no limiting meaning.
[0039] In addition, the data involved in this disclosure may be data authorized by the user or fully authorized by all parties. The collection, dissemination and use of the data shall comply with the requirements of relevant national laws and regulations. The implementation methods / executives of this disclosure may be combined with each other. Invention Overview
[0041] The inventors have discovered that many existing online music platforms and software offer the ability to automatically identify songs from external playlists by taking screenshots and importing them into existing or newly created playlists on the platform or software. This eliminates the need for users to repeatedly match and add songs one by one across different platforms or software in order to listen to songs from external playlists, significantly reducing user costs. However, existing methods for adding songs based on external playlist screenshots can only directly import all tracks from the screenshot, requiring users to further edit the songs, such as deleting copyrighted songs, incorrectly matched songs (e.g., version mismatch), or songs that are significantly different from other songs (e.g., songs with obvious stylistic differences). The songs ultimately added to the existing or newly created playlist are only a small portion of the original playlist screenshot, making the overall process cumbersome and unable to directly provide the specific songs the user needs, resulting in a poor user experience.
[0042] In this solution, by extracting the corresponding song information from the received playlist screenshot and generating one or more recommended playlists of different styles based on the distribution of song styles, it is easier for users to select and use recommended playlists. This reduces the amount of operation required for further classification of songs in the original playlist, improves the operational efficiency during the playlist migration process, and ultimately enhances user satisfaction.
[0043] After introducing the basic principles of this disclosure, various non-limiting embodiments of this disclosure will be described in detail below.
[0044] Application Scenarios Overview
[0045] First refer to Figure 1a As shown, during the playlist generation process, server 100 receives a screenshot of the playlist transmitted by client 110 (which can be a web client or an application client). Based on the playlist screenshot and the song-related data in database 120, multiple recommended playlists are generated, thereby completing the playlist generation process.
[0046] Secondly, refer to Figure 1b As shown, this is a structural diagram of a playlist screenshot. A playlist screenshot typically includes information such as playlist name (which can be omitted), song name, artist name, album name, duration (which can be omitted), and notes.
[0047] It should be noted that, Figure 1a The scenario shown uses only one server, client, and database as an example for illustration, but this disclosure is not limited to this; that is, the number of servers, clients, and databases can be arbitrary.
[0048] Exemplary methods
[0049] The following is combined Figure 1a and Figure 1b For application scenarios, refer to Figures 2 to 4d This document describes a method for generating playlists according to exemplary embodiments of the present disclosure. It should be noted that the above application scenarios are shown only to facilitate understanding of the spirit and principles of the present disclosure, and the embodiments of the present disclosure are not limited in any way. Rather, the embodiments of the present disclosure can be applied to any applicable scenario.
[0050] Figure 2 This is a flowchart illustrating a playlist generation method provided in one embodiment of this disclosure. Figure 2 As shown, the playlist generation method provided in this embodiment includes the following steps:
[0051] Step S201: In response to the received playlist screenshot, extract the song information from the playlist screenshot.
[0052] Specifically, a playlist screenshot can be taken from a smartphone app or from a computer app (such as a dedicated screenshot software or a non-dedicated screenshot software with screenshot functionality) and contain information about the songs in the playlist.
[0053] The song information includes information directly related to the song contained in the playlist screenshot, such as the song name, artist name, album name, etc. Based on the song information, the corresponding song data can be matched from the server's music library.
[0054] A playlist screenshot can contain only song information, or it can contain non-song information unrelated to the song information, such as the playlist name, background image, or icons unrelated to the song information. As long as the playlist screenshot contains song information, the server or processor can extract the song information from it and match the corresponding song data from the music library, without being affected by non-song information.
[0055] The song information in the playlist screenshot can be complete or incomplete. For example, some song information may be omitted due to excessive length. For instance, the album title may be "Top Ten Songs of the Year," but only "Top Ten Songs of the Year..." is actually displayed.
[0056] When song information is complete, the server or processor can directly match the corresponding song data based on the song information. If the song information is incomplete, the server or processor can also search for the closest song data in the music library based on the incomplete song information as the song data to match the song information. As in the example above, the music library has two albums with the names "Top Ten Songs of the Year" and "Top Ten Golden Hits of the Year" that can match the song information, but "Top Ten Golden Hits of the Year" does not contain a song with that song name. In this case, the server will determine that the song with that song name in "Top Ten Songs of the Year" is the matching song data (if multiple albums contain song data with the same performer and the same song name, the song data with the highest popularity or play count can be automatically selected as the matching song data based on the popularity or play count of these song data).
[0057] Step S202: Based on the song information, determine the style information corresponding to the song.
[0058] Specifically, after matching the corresponding song data, the server does not immediately return this song data to the user. Instead, it analyzes the song data (such as extracting its song content feature information through a neural network that specifically extracts song features or chord features, or extracting its style tag information through its configuration attributes, etc.) to obtain its style information, and based on the style information, it creates multiple recommended playlists corresponding to different styles.
[0059] Style information refers to the emotional characteristics of a song (such as cheerful, sad, or no specific emotion), song category (such as pop, classical, or blues), and other characteristic information. It can also include content information such as the main chords of the song (such as major triads or minor seventh chords). Through style information, we can determine the style characteristic tags corresponding to the song data. For example, songs with style characteristic tags such as pop, cheerful, and mainly using major triads are usually not suitable for being placed in the same recommended playlist as songs with style characteristic tags such as blues, sad, and mainly using minor seventh chords.
[0060] Therefore, by matching the song information with the corresponding song data, we can obtain the style information of the song, so as to better allocate the song to the corresponding recommended playlist.
[0061] Step S203: Based on the number of songs and the style distribution of the songs in the playlist screenshot, generate a recommended playlist containing at least one style.
[0062] Specifically, when generating a recommended playlist, it is necessary to consider both the style distribution of songs in the playlist screenshot and the number of songs. The song distribution mainly refers to the distribution of song style information (which can be obtained by clustering analysis of the style feature tags of each song data); the number of songs mainly refers to the number of songs that can be identified from the playlist screenshot.
[0063] When there are many songs and their styles are widely distributed, multiple recommended playlists may be generated to group songs with similar or identical style information into the same playlist as much as possible, while avoiding the inclusion of songs with significantly different style information in the same playlist. This improves user satisfaction when playing the generated playlists (generally, the user experience is better when the songs in the same playlist are similar in style).
[0064] When the number of songs is small, or the song styles are relatively concentrated, the number of recommended playlists may be small. For example, if the songs in the playlist screenshot have the same style information, such as coming from the same album, there may only be one recommended playlist. In this case, it is not necessary to generate too many recommended playlists to avoid excessive splitting and reduce the user experience (because users usually do not tend to have too few songs in a playlist, such as only two or three songs).
[0065] In one exemplary embodiment of this disclosure, different recommended playlists may contain songs from the same playlist screenshot. That is, songs from a playlist screenshot can be added to different recommended playlists simultaneously. For example, if the style information of a song is similar to that of songs in two recommended playlists, the song can be added to both playlists, rather than being limited to only one. This increases the number of songs in each recommended playlist while maximizing the consistency of style information within the playlists, thereby improving the user experience.
[0066] According to the playlist generation method of this disclosure, when a playlist screenshot is received, song information is extracted from the screenshot. Then, based on the song information, the corresponding style information of the songs is determined. Based on the number of songs and the style distribution in the playlist screenshot, a recommended playlist containing at least one style is generated. Therefore, based on the received external playlist screenshot, different recommended playlists can be automatically generated according to characteristics such as song style, achieving automatic classification and recognition of playlist screenshots without requiring tedious editing by the user, simplifying operation and improving user experience.
[0067] Figure 3a This is a flowchart illustrating a playlist generation method provided in one embodiment of this disclosure. Figure 3a As shown, the playlist generation method provided in this embodiment includes the following steps:
[0068] Step S301: Input the playlist screenshot into the pre-trained text detection model and output the text positions in the playlist screenshot.
[0069] Specifically, this embodiment will provide a detailed explanation of the steps for obtaining song information.
[0070] When a playlist screenshot is obtained, the server or processor will input the screenshot into a text detection model to locate the position of the text in the screenshot. This allows the text recognition model to identify the corresponding text later. (Because text recognition models can usually only recognize text, directly inputting the playlist screenshot into the text recognition model would severely reduce the recognition accuracy due to interference from images and non-text symbols contained in the screenshot. Therefore, pre-locating the text using a text detection model can significantly improve the recognition accuracy.)
[0071] The text detection model is used to locate each piece of text (or text box, such as the text box corresponding to the playlist name or the text box corresponding to the song name) in the playlist screenshot. After inputting the playlist screenshot, it can input the position information of the text or text box in the playlist screenshot (such as the diagonal coordinates of the text box), so that the text detection module can determine the content of each text box accordingly.
[0072] In one exemplary embodiment of this disclosure, before inputting the playlist screenshot into the text detection model, an adaptive binarization algorithm can be used to preprocess the playlist screenshot to enhance the features of the text portion and improve the detection performance of the text detection module. Any existing adaptive binarization algorithm can be used, and no limitation is made here.
[0073] In one exemplary embodiment of this disclosure, before inputting the playlist screenshot into the text detection module, a line segmentation algorithm can be used to split the text portion of the playlist screenshot, thereby improving the recognition accuracy of the file detection module. Because information such as song titles, artist names, and album names in the playlist screenshot are typically displayed in a single line, rather than on multiple lines, a line segmentation algorithm (any line segmentation algorithm is acceptable and not limited here) can be used to split the text in the playlist screenshot into several single lines of text. This reduces the size of the text region determined by the text detection module (achieved by reducing blank spaces in the detected text boxes, thereby reducing invalid or redundant information in the text boxes), and improves the accuracy of subsequent text recognition.
[0074] Furthermore, such as Figure 3b The diagram shown is a flowchart of the training method for a text detection model, which is trained in the following way:
[0075] Step S3011: Collect screenshots of the playlist as image samples.
[0076] The playlist screenshot contains actual text location information, including the position coordinates of the text box corresponding to the text and the size of the text box.
[0077] Specifically, the pre-trained samples are the collected playlist screenshots. Each playlist screenshot needs to be labeled with the actual text location information, that is, the position coordinates of the text or text box in the playlist screenshot (such as diagonal coordinates, center point coordinates, top left corner coordinates, etc.) and the size of the text box (which can be represented by length and width, and the unit can be millimeters, pixels, etc.).
[0078] Step S3012: Input the image sample into the text detection model and output the predicted text location information in the image sample.
[0079] Specifically, the text detection model can be implemented using existing convolutional neural network models for text detection, such as the YOLOv5 model. The text detection model will output the text position information of the predicted text boxes in the image samples. The format of the output predicted text position information is the same as the format of the actual text position information, including the corresponding position coordinates and text box size.
[0080] Step S3013: Based on the actual text location information and the predicted text location information, perform regression training on the text detection model.
[0081] Specifically, by predicting the difference between the text location information and the actual text location information, and with the optimization of the corresponding loss function as the objective, the parameters in the text detection model are corrected, thereby achieving regression training of the text detection model to obtain a text detection model that can be used for text location detection.
[0082] Step S302: Input the text position and the screenshot of the playlist into a pre-trained text recognition model, and output the text information in the screenshot of the playlist.
[0083] Specifically, after obtaining the position information (i.e., text position) of the text box where the text in the screenshot of the playlist is located through the text detection model, the text information in each text box can be sequentially recognized through the text recognition model, and then the text information in the entire screenshot of the playlist can be obtained.
[0084] Furthermore, as Figure 3c shown, it is a flowchart of the training method of the text recognition model, and this model is trained through the following method:
[0085] Step S3021: Generate pictures containing random text based on the text format in the screenshot of the playlist.
[0086] Specifically, the text format referred to here means the distribution method of the text boxes in the screenshot of the playlist, rather than the format of a single text such as font. For example, the text boxes in the same line contain the song name, album name, and duration, or the song name, album name, and singer name are distributed in the text boxes of different lines. Thus, the recognition accuracy of the trained text recognition model can be improved.
[0087] Step S3022: Add digital labels to the random text based on a preset text dictionary.
[0088] Specifically, in text recognition, the text recognition model matches the picture containing text with the text in the font library, and outputs the text in the font library with the highest matching degree with the text in the picture containing text as the recognition result. And the text recognition model usually does not directly output the recognized text itself, but outputs the digital label corresponding to the text (such as the digital 112 corresponds to the text "good"). Therefore, when training the text recognition model, a text dictionary for matching text and digital labels needs to be prepared first.
[0089] The text dictionary is generated based on the common text in the existing font library. By adding a digital label to each text in the text dictionary (such as adding the digital label 001 to "you"), the matching of text and digital labels is achieved.
[0090] Based on the corresponding relationship between the text and the digital label in the text dictionary, add the corresponding digital labels to the random text in the generated pictures respectively for training the text recognition model.
[0091] Step S3023: Input the pictures containing random text and the random text containing digital labels as training samples into the text recognition model for training, and obtain the trained text recognition model.
[0092] Specifically, text recognition models can employ convolutional recurrent neural networks that include connectionist temporal classification (CTC), which performs better in terms of text recognition accuracy and robustness compared to conventional convolutional neural networks.
[0093] The text recognition model is trained by inputting images containing random text and combinations of random text and number labels into the text recognition model.
[0094] Since the text in playlist screenshots typically only contains information such as song titles, artist names, and album names, without involving complex sentences or logical analysis, the vector model of the text output by the text recognition model can adopt a one-hot model. This model only outputs the results used to determine the text classification (i.e., the result only needs to determine the category to which the recognized text belongs, without considering the logical relationships between the texts; the category is the correspondence between the recognized text and the text in the dictionary). This improves the processing efficiency of the text recognition model.
[0095] Step S303: Based on the noise reduction model, perform noise reduction processing on the text information.
[0096] Specifically, after text recognition, the text information is obtained from all text boxes in the playlist screenshot. This text information often contains some distracting information with little relevance to the song information, such as ellipses in song titles and song numbers. To ensure the accuracy of the subsequently determined song information, this text information needs to be denoised to remove the distracting information.
[0097] At this point, a noise reduction algorithm can be used to remove interference information from the text. The specific noise reduction algorithm used can be based on regular expressions to write corresponding rules (such as keeping only the text corresponding to the song title, album name, and artist name according to the order of the text positions, and removing other text), and there are no restrictions here.
[0098] Step S304: Match the text information with the song information in the music library, and use the matching result as the song information in the playlist screenshot.
[0099] Specifically, the method for matching song information in a playlist screenshot can use the TF-IDF algorithm (term frequency–inverse document frequency) to match text information with song information in the music library. The song information with the highest matching degree with the text information is then identified as the song information that matches the text information.
[0100] In one exemplary embodiment of this disclosure, the text information obtained after noise reduction is represented in the form of a combination of song name, album name, and artist name (such as in the form of an array or vector). Therefore, each piece of text information used for matching contains at least one of the song name, album name, and artist name (usually at least the song name, while the album name and artist name can be omitted) to ensure the accuracy of the matching.
[0101] Step S305: Based on the song information, determine the style information corresponding to the song.
[0102] Step S306: Based on the number of songs and the style distribution of the songs in the playlist screenshot, generate a recommended playlist containing at least one style.
[0103] Specifically, steps S305 to S306 and Figure 2 Steps S202 to S203 in the illustrated embodiment are the same and will not be repeated here.
[0104] According to the playlist generation method of this disclosure, a screenshot of the playlist is input into a pre-trained text detection model, which outputs the text positions in the screenshot. Then, the text positions and the screenshot are input into a pre-trained text recognition model, which outputs the text information in the screenshot. Next, a noise reduction model is used to denoise the text information, and the text information is matched with song information in a music library. The matching result is used as the song information in the playlist screenshot. Finally, based on the song information, the corresponding style information is determined, and a recommended playlist is generated. This ensures the accuracy of song information recognition in the playlist screenshot, avoids mismatches between the generated recommended playlist and the songs in the screenshot, and effectively improves user satisfaction.
[0105] Figure 4a This is a flowchart illustrating a playlist generation method provided in one embodiment of this disclosure. Figure 4a As shown, the playlist generation method provided in this embodiment includes the following steps:
[0106] Step S401: In response to the received playlist screenshot, extract the song information from the playlist screenshot.
[0107] Specifically, step S401 and Figure 2 The steps of S201 in the illustrated embodiment are the same and will not be repeated here.
[0108] Step S402: Based on the song information, extract the audio data and lyrics data of the corresponding song from the music library.
[0109] Specifically, after determining the song information of all songs matched by the playlist screenshot, the server will directly extract the relevant data of these songs stored in the music library, namely audio data and lyrics data, to identify the style characteristics of the songs.
[0110] Step S403: Extract the Mel spectrum information from the audio data to obtain the audio vector corresponding to the audio data.
[0111] Specifically, Mel spectrum information can reflect the frequency distribution of audio data at different times. Based on this, feature vectors, i.e. audio vectors, of the audio data can be extracted through sampling and other processing methods. In order to determine the chord features of the song through the audio vectors, and classify the style of the song according to the chord features.
[0112] Step S404: Based on word segmentation tools and word vector conversion models, determine the semantic vector corresponding to the lyrics data.
[0113] Specifically, for lyrics data, it is necessary to extract their semantics in order to classify them according to semantic alignment, because songs of different styles usually have relatively fixed words, such as the word "classmates" usually belonging to the campus song category.
[0114] The semantics in lyrics data are represented in the form of semantic vectors. Extracting semantic vectors can be achieved using word segmentation tools and word vector conversion models.
[0115] In one exemplary embodiment of this disclosure, such as Figure 4b The diagram shown is a flowchart for determining the semantic vector of lyrics data. It includes the following steps:
[0116] Step S4041: Based on the word segmentation tool, perform word segmentation processing on the lyrics data.
[0117] Specifically, before converting lyrics data into semantic vectors, it is necessary to perform word segmentation, breaking the lyrics data into combinations of multiple words in order to extract semantics from the segmented words and convert them into semantic vectors.
[0118] The word segmentation tool used for word segmentation can be any existing word segmentation tool, such as word segmentation based on the jieba library or word segmentation based on the THULAC tool, etc., and there is no limitation here.
[0119] Step S4042: Input the segmented lyrics data into the word vector conversion model to obtain the word meaning vectors corresponding to the segmented lyrics data.
[0120] Specifically, after word segmentation is completed, the resulting words need to be converted into word meaning vectors, and semantic vectors are obtained based on the word meaning vectors.
[0121] Tools that convert split words into semantic vectors can be implemented using word vector conversion models. Specifically, the word2vec model can be used to ensure the processing efficiency of the word vector conversion process.
[0122] Step S4043: Perform a weighted average of all semantic vectors corresponding to the lyrics data to obtain the semantic vector corresponding to the lyrics data.
[0123] Specifically, because the length of lyrics varies from song to song, the number of words and semantic vectors derived from the lyric data also differs. Therefore, the style of a song cannot be directly evaluated based on semantic vectors (because different semantic vectors result in different amounts of data for evaluation, leading to poor accuracy and consistency of the results). Therefore, further integration processing of semantic vectors is required to ensure the accuracy and consistency of style features determined based on the features of the lyric data.
[0124] The integration of word meaning vectors can be achieved by averaging all word meaning vectors or by weighted averaging (e.g., setting different weights based on the frequency of occurrence of the same word) to obtain vectors that correspond one-to-one with the lyrics data, i.e., semantic vectors.
[0125] Step S405: Input the audio vector and semantic vector into the classification network respectively to identify the chord style and song style corresponding to the song.
[0126] Specifically, the style information of a song includes chord style and song style. Chord style refers to the chords that appear most frequently in the song and the types of chords. Song style refers to the genre or style to which the song belongs, such as jazz, classical, blues, rock, etc.
[0127] When identifying song style information using audio vectors and semantic vectors, it is necessary to identify the chord style and song style separately. Chord style, being solely related to audio features, can be identified based on audio vectors. Song style, however, involves both audio and lyrics (e.g., blues music uses relatively fixed chords, while folk music uses more universal lyrics and terminology). Therefore, audio vectors and semantic vectors can be combined to identify the song style.
[0128] In one exemplary embodiment of this disclosure, such as Figure 4c The diagram shows a flowchart for identifying song chord styles and song genres. If the classification network includes a first classification network for classifying chord categories and a second classification network for classifying song genres, the specific steps include:
[0129] Step S4051: Input the audio vector into the chord feature extraction network to obtain the chord features of the audio vector.
[0130] Specifically, since audio vectors contain not only chord features but also non-chord features, such as features corresponding to the main melody, it is necessary to first extract the chord features from the audio vectors when identifying chord styles.
[0131] The method for extracting chord features is a mature solution in the existing technology. Those skilled in the art can choose any chord feature extraction network (or extraction algorithm) to complete the chord feature extraction step, and there is no limitation here.
[0132] Step S4052: Input the chord features into the first classification network and output the chord style corresponding to the song.
[0133] Specifically, the first classification network is mainly used to classify chord features based on the frequency of different chords in the chord features, so that the chords with the highest frequency (such as four chords) are taken as the chord style corresponding to the chord features, that is, the chord style corresponding to the song.
[0134] Therefore, during training, the first classification network uses a large amount of chord feature data containing different chords (and chords appearing at different frequencies) as training samples, and marks the chords with the highest frequency in the chord feature data as their chord styles; then this training sample is input into the first classification network for training.
[0135] Based on the results of the first classification network output, the most frequent chords in the song can be identified, and thus its chord style can be determined.
[0136] Step S4053: Input the audio vector and semantic vector into the song style feature extraction network and output the song style features corresponding to the song.
[0137] Specifically, the song style feature extraction network is a neural network used for song style and genre identification. In the existing technology, the method for establishing a neural network for identifying song style based on the audio data of the song is relatively mature, and the method for establishing a song style feature extraction network based on audio vectors and semantic vectors is similar.
[0138] By collecting songs with a pre-determined style and pre-extracting their corresponding audio and semantic vectors as training samples, and then inputting the concatenated audio and semantic vectors into a neural network (such as a convolutional neural network or a recurrent convolutional neural network) for training, a neural network for recognizing song style features through audio and semantic vectors is obtained.
[0139] Since audio vectors and semantic vectors are feature information extracted from audio and lyrics data, they reduce data interference and improve recognition accuracy compared to directly using raw audio and lyrics data for training and judgment. Furthermore, compared to existing neural networks for song style and genre recognition that only use audio data, combining audio data with lyrics data (i.e., audio vectors and semantic vectors) can further improve the accuracy of song style recognition.
[0140] Step S4054: Perform a fusion process on the song's stylistic features and chord features.
[0141] Specifically, the song style features and chord features output by the song style feature extraction network are combined (e.g., by addition, offset addition, multiplication, or splicing) to obtain the fused features, so that the song style can be comprehensively determined based on the fused features.
[0142] Step S4055: Input the fused song style features and chord features into the second classification network and output the song style corresponding to the song.
[0143] Specifically, since the accuracy of existing conventional neural networks for identifying song styles and genres is usually limited (generally below 90%), even if audio vectors and semantic vectors are combined for identification, the accuracy of the resulting neural network cannot be guaranteed to be high enough. Therefore, based on the identified song style features, chord features can be combined and identified again through a second classification network to further improve the accuracy of identification.
[0144] Furthermore, such as Figure 4d The diagram shown is a flowchart of the training method for the second classification network. The second classification network is trained in the following way:
[0145] Step A1: Obtain audio sample data for training, and label the style and chords that appear more than a set number or more than a set percentage.
[0146] Specifically, the input data for the second classification network contains both chord features and song style features. Therefore, the training sample data used to train the second classification network needs to be pre-labeled with its style and chord features (i.e., the most frequent chords, chords with a frequency exceeding a set number, or chords with a frequency exceeding a set proportion).
[0147] Step 2A2: Input the audio sample data into the second classification network and output the predicted style from the audio sample data.
[0148] Specifically, by inputting sample data labeled with chord features and style into the second classification network, the network outputs the predicted style features, which are then compared with the labeled styles to determine the accuracy of the prediction results.
[0149] Step 3A3: Based on the labeled style and predicted style corresponding to the audio sample data, perform regression training on the second classification network.
[0150] Specifically, the process of regression training for the second classification network is similar to the training process for the song style feature extraction network and the chord feature extraction network, and will not be elaborated here.
[0151] In one exemplary embodiment of this disclosure, the neural networks used in steps S403 to S405 can also be integrated into a module, and the corresponding chord style and song style can be directly output by inputting the lyrics data and audio data extracted in step S402.
[0152] Step S406: Determine the number of playlists to be generated based on the number of songs in the playlist screenshot.
[0153] Specifically, after obtaining the chord progressions and style of a song, a recommended playlist can be created based on this information and the number of songs.
[0154] The appropriate number of playlists to generate varies depending on the number of songs in the playlist screenshot (e.g., if there are only 3 songs, there's no need to generate multiple playlists; however, if there are 20 songs, generating 2 to 5 playlists is acceptable). The specific correspondence between the number of songs and the number of playlists to be generated can be obtained through research or statistics on user habits regarding the number of songs; no specific limit is set here.
[0155] Step S407: Combine the chord style and song style in the style information and encode them into a style vector.
[0156] Specifically, based on the determined chord style and song style, they can be encoded into style vectors. For example, based on the most frequently occurring chord names in the chord style, corresponding codes can be selected (e.g., C chords can be encoded as 0001, Am chords as 0010), and all the most frequently occurring chords can be combined (the code length may vary depending on the number of chords; for example, if the most frequently occurring chords are C and Am, they can be encoded as 00010010). Song styles can also be encoded in a similar way (e.g., blues is encoded as 0000100000).
[0157] By concatenating the codes for chord style and song style, you can obtain a style vector (as in the previous example, concatenating C chord, Am chord, and blues results in 000100100000100000), which represents the song's corresponding style, chords, and other information.
[0158] Step S408: Determine the clustering distribution of style vectors to obtain the song clusters corresponding to the number of playlists to be generated.
[0159] Specifically, after determining the style vectors, cluster analysis can be used to cluster these style vectors, thereby clustering the songs in the playlist screenshot and dividing the songs into different playlists.
[0160] In one exemplary embodiment of this disclosure, the clustering analysis algorithm may select k-means clustering, which clusters style vectors into a corresponding number of categories based on a predetermined number of playlists, and assigns each category as a playlist to recommend to the user.
[0161] The number of songs in different clusters may vary, and different clusters may contain the same songs (e.g., a song may belong to both a rock playlist and a playlist with C chords). Therefore, the total number of songs in all categories may be greater than the number of songs in the playlist screenshot (because some songs may appear repeatedly in multiple categories).
[0162] Step S409: Generate a corresponding recommended playlist based on each song cluster, and use the style information of the song cluster as the feature information of the recommended playlist.
[0163] Specifically, for each cluster of songs, a corresponding recommended playlist can be generated. The information corresponding to the style vector of that cluster, or the tag information common to songs in that category (such as the same artist, the same year, the same label, etc., which cannot be expressed by style vectors), is used as the feature information of the recommended playlist. This allows users to intuitively understand the characteristics of different recommended playlists.
[0164] In one exemplary embodiment of this disclosure, the generated recommended playlist includes an option to retain it or not, so that users can choose to retain all or part of the recommended playlist as needed to improve the user experience.
[0165] In one exemplary embodiment of this disclosure, in addition to generating recommended playlists based on different clusters, a separate recommended playlist containing all the songs in the playlist screenshots is also generated to maximize the satisfaction of user needs (such as the user may need to create a playlist with mixed styles).
[0166] According to the playlist generation method of this disclosure, song information is extracted from a playlist screenshot, and the audio data and lyrics data of the corresponding songs in the music library are determined based on the song information. Then, the audio vectors corresponding to the audio data and the semantic vectors corresponding to the lyrics data are extracted and input into a classification network to obtain the chord style and song style corresponding to the songs. Then, based on the chord style, song style, and the number of songs in the playlist screenshot, multiple recommended playlists are generated based on the songs in the playlist screenshot. This allows the generated recommended playlists to be tailored to the characteristics of different songs in the playlist screenshot, reducing the tedious operation of manually splitting recommended playlists by the user, improving processing efficiency, and thus improving the user experience.
[0167] Exemplary media
[0168] After introducing the methods of exemplary embodiments of this disclosure, the following references are made. Figure 5 The storage medium of the exemplary embodiments of this disclosure will be described.
[0169] refer to Figure 5 As shown, a program product 50 for implementing the above-described method according to an embodiment of the present disclosure is described. This product may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto.
[0170] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0171] A readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying readable program code. This propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium.
[0172] Program code for performing the operations disclosed herein can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing devices can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN).
[0173] Exemplary device
[0174] Having introduced the medium of exemplary embodiments of this disclosure, the following references are made to... Figure 6 The playlist generation apparatus of the present disclosure is described in an exemplary embodiment, which is used to implement the playlist generation method in any of the above method embodiments. Its implementation principle and technical effect are similar to those of the corresponding methods described above, and will not be repeated here.
[0175] The playlist generation device 600 disclosed herein includes:
[0176] The extraction module 610 is used to extract song information from the received playlist screenshot in response to the received playlist screenshot.
[0177] The determination module 620 is used to determine the style information corresponding to the song based on the song information;
[0178] The generation module 630 is used to generate recommended playlists for songs based on style information.
[0179] In one exemplary embodiment of this disclosure, the extraction module 610 is specifically used to: input a playlist screenshot into a pre-trained text detection model and output the text position in the playlist screenshot; input the text position and the playlist screenshot into a pre-trained text recognition model and output the text information in the playlist screenshot; perform matching processing between the text information and the song information in the music library, and use the matching result as the song information of the song in the playlist screenshot.
[0180] In one exemplary embodiment of this disclosure, the extraction module 610 includes: training a text detection model by: acquiring a screenshot of a playlist as an image sample, the screenshot containing actual text location information, the actual text location information including the position coordinates of the text box corresponding to the text and the size of the text box; inputting the image sample into the text detection model, and outputting the predicted text location information in the image sample; and performing regression training on the text detection model based on the actual text location information and the predicted text location information.
[0181] In one exemplary embodiment of this disclosure, the extraction module 610 includes: training a text recognition model by generating an image containing random text based on the text format in the playlist screenshot; adding numeric labels to the random text based on a preset text dictionary; and inputting the image containing random text and the random text containing numeric labels as training samples into the text recognition model for training, thereby obtaining a trained text recognition model.
[0182] In one exemplary embodiment of this disclosure, the extraction module 610 is further configured to: perform noise reduction processing on the text information based on a noise reduction model before matching the text information with the song information in the music library and using the matching result as the song information of the song in the playlist screenshot.
[0183] In one exemplary embodiment of this disclosure, the determining module 620 is specifically used for: if the style information includes chord style and song style, extracting the audio data and lyrics data of the corresponding song from the music library based on the song information; extracting the Mel spectrum information from the audio data to obtain the audio vector corresponding to the audio data; determining the semantic vector corresponding to the lyrics data based on the word segmentation tool and the word vector conversion model; and inputting the audio vector and the semantic vector into the classification network respectively to identify the chord style and song style corresponding to the song.
[0184] In one exemplary embodiment of this disclosure, the determining module 620 is specifically used to: perform word segmentation processing on the lyrics data based on a word segmentation tool; input the segmented lyrics data into a word vector conversion model to obtain the word sense vectors corresponding to the segmented lyrics data; and perform weighted average processing on all word sense vectors corresponding to the lyrics data to obtain the semantic vectors corresponding to the lyrics data.
[0185] In one exemplary embodiment of this disclosure, the determining module 620 is specifically configured to: if the classification network includes a first classification network for classifying chord categories and a second classification network for classifying song styles, input the audio vector into the chord feature extraction network to obtain the chord features of the audio vector; input the chord features into the first classification network to output the chord style corresponding to the song; input the audio vector and the semantic vector into the song style feature extraction network to output the song style features corresponding to the song; perform a fusion processing on the song style features and the chord features; input the fused song style features and the chord features into the second classification network to output the song style corresponding to the song.
[0186] In one exemplary embodiment of this disclosure, the determining module 620 includes: training a second classification network by: acquiring audio sample data for training and labeling the styles and chords whose frequency of occurrence exceeds a set number or whose proportion of occurrence exceeds a set ratio corresponding to the audio sample data; inputting the audio sample data into the second classification network and outputting the predicted styles in the audio sample data; and performing regression training on the second classification network based on the labeled styles and predicted styles corresponding to the audio sample data.
[0187] In one exemplary embodiment of this disclosure, the generation module 630 is specifically used to: determine the number of playlists to be generated based on the number of songs in the playlist screenshot; combine the chord style and song style in the style information and encode them into a style vector; perform clustering processing on the style vector to obtain the song clusters corresponding to the number of playlists to be generated; generate a corresponding recommended playlist based on each song cluster, and use the style information corresponding to the song cluster as the feature information of the recommended playlist.
[0188] Exemplary computing device
[0189] Having described the methods, media, and apparatus of exemplary embodiments of this disclosure, the following references... Figure 7 A computing device according to an exemplary embodiment of the present disclosure will be described.
[0190] Figure 7 The computing device 70 shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0191] like Figure 7 As shown, the computing device 70 is presented in the form of a general-purpose computing device. The components of the computing device 70 may include, but are not limited to: at least one processing unit 701, at least one storage unit 702, and a bus 703 connecting different system components (including the processing unit 701 and the storage unit 702).
[0192] The 703 bus includes a data bus, a control bus, and an address bus.
[0193] Storage unit 702 may include readable media in the form of volatile memory, such as random access memory (RAM) 7021 and / or cache memory 7022, and may further include readable media in the form of non-volatile memory, such as read-only memory (ROM) 7023.
[0194] Storage unit 702 may also include a program / utility 7025 having a set (at least one) program module 7024, such program module 7024 including but not limited to: operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.
[0195] The computing device 70 can also communicate with one or more external devices 704 (e.g., keyboard, pointing device, etc.). This communication can be performed via the input / output (I / O) interface 705. Furthermore, the computing device 70 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via a network adapter 706. Figure 7 As shown, network adapter 706 communicates with other modules of computing device 70 via bus 703. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with computing device 70, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0196] It should be noted that although several units / modules or sub-units / modules of the supply chain strategy determination device and the object scoring model training device are mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.
[0197] Furthermore, although the operations of the methods disclosed herein are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.
[0198] While the spirit and principles of this disclosure have been described with reference to several specific embodiments, it should be understood that this disclosure is not limited to the disclosed specific embodiments, and the division of aspects does not imply that features in these aspects cannot be combined for benefit; such division is merely for convenience of expression. This disclosure is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.
Claims
1. A method of generating a playlist, characterized by, The method comprises: extracting song information of songs in the received playlist screenshot; determining style information corresponding to the songs based on audio data and lyric data of the songs, the style information comprising chord style and song style; combining the chord style and the song style in the style information to encode a style vector; determining a clustering distribution of the style vector to obtain a song cluster corresponding to the number of to-be-generated playlists; generating a corresponding recommended playlist for each song cluster based on the corresponding style information of the song cluster as feature information of the recommended playlist, and the recommended playlist comprising an option of whether to keep or not; determining the style information corresponding to the songs based on the audio data and the lyric data of the songs comprises: extracting audio data and lyric data of corresponding songs in a music library based on the song information; extracting mel-frequency spectrum information in the audio data to obtain an audio vector corresponding to the audio data; determining a semantic vector corresponding to the lyric data based on a word segmentation tool and a word vector conversion model; inputting the audio vector and the semantic vector into a classification network respectively to identify the chord style and the song style corresponding to the songs; the classification network comprises a first classification network for classifying chord categories and a second classification network for classifying song styles.
2. The playlist generating method according to claim 1, wherein The method comprises: inputting the playlist screenshot into a pre-trained text detection model to output text positions in the playlist screenshot; inputting the text positions and the playlist screenshot into a pre-trained text recognition model to output text information in the playlist screenshot; matching the text information with song information in a music library, and taking a matching result as song information of songs in the playlist screenshot.
3. The playlist generating method according to claim 2, wherein The text detection model is trained in the following manner: collecting playlist screenshot pictures as picture samples, the playlist screenshot pictures containing actual text position information, the actual text position information comprising position coordinates of a text box corresponding to the text and a size of the text box; inputting the picture samples into a text detection model to output predicted text position information in the picture samples; performing regression training on the text detection model based on the actual text position information and the predicted text position information.
4. The playlist generating method according to claim 2, characterized by, The text recognition model is trained in the following manner: generating pictures containing random texts based on a text format in a playlist screenshot; adding digital labels to the random texts based on a preset text dictionary; inputting the pictures containing random texts and the random texts containing digital labels as training samples into the text recognition model for training to obtain a trained text recognition model.
5. The playlist generating method according to claim 2, wherein Before the matching of the text information with song information in a music library and the taking of a matching result as song information of songs in the playlist screenshot, the method further comprises: performing noise reduction processing on the text information based on a noise reduction model.
6. The playlist generating method of claim 1, wherein, The word-based segmentation tool and the word vector conversion model are used to determine a semantic vector corresponding to the lyrics data, including: The lyrics data are processed by word-based segmentation based on the word-based segmentation tool; The lyrics data after the word-based segmentation are input into the word vector conversion model respectively to obtain word meaning vectors corresponding to the lyrics data after the word-based segmentation; All the word meaning vectors corresponding to the lyrics data are processed by weighted average to obtain a semantic vector corresponding to the lyrics data.
7. The playlist generation method of claim 1, wherein the audio vector and the semantic vector are input into a classification network respectively to identify chord styles and song styles corresponding to the songs, including: The audio vector is input into a chord feature extraction network to obtain chord features of the audio vector; The chord features are input into a first classification network to output chord styles corresponding to the songs; The audio vector and the semantic vector are input into a song style feature extraction network to output song style features corresponding to the songs; The song style features and the chord features are fused; The fused song style features and chord features are input into a second classification network to output song styles corresponding to the songs. The second classification network is trained in the following manner:
8. The playlist generating method according to claim 7, wherein Audio sample data used for training are obtained, and chords corresponding to the audio sample data and having a frequency of occurrence exceeding a set number or a proportion of occurrence exceeding a set proportion are labeled; The audio sample data are input into the second classification network to output predicted styles in the audio sample data; The second classification network is trained by regression based on the labeled styles corresponding to the audio sample data and the predicted styles. including:
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer execution instructions, and the computer execution instructions are executed by the processor to implement the playlist generation method of any one of claims 1 to 8. The device includes:
10. A playlist generating apparatus characterized by comprising: An extraction module configured to extract song information of songs in a received playlist screenshot; A determination module configured to determine style information corresponding to the songs based on audio data and lyrics data of the songs, the style information including chord styles and song styles; A generation module configured to determine a number of to-be-generated playlists based on a number of songs in the playlist screenshot, combine the chord styles and the song styles in the style information to encode as style vectors, determine clustering distribution of the style vectors to obtain song clusters corresponding to the to-be-generated playlists, generate a corresponding recommended playlist for each song cluster based on the song cluster, use corresponding style information of the song cluster as feature information of the recommended playlist, and generate the recommended playlist with an option of whether to keep or not to keep. The determining module is specifically configured to: extract audio data and lyric data of a corresponding song in a music library based on the song information; extract mel-frequency spectrum information in the audio data to obtain an audio vector corresponding to the audio data; determine a semantic vector corresponding to the lyric data based on a word segmentation tool and a word vector conversion model; and input the audio vector and the semantic vector into a classification network respectively to identify a chord style and a song style corresponding to the song.
11. The playlist generating apparatus according to claim 10, wherein The extracting module is specifically configured to: input the playlist screenshot into a pre-trained text detection model to output a text position in the playlist screenshot; input the text position and the playlist screenshot into a pre-trained text recognition model to output text information in the playlist screenshot; match the text information with song information in a music library, and use a matching result as the song information of a song in the playlist screenshot.
12. The playlist generating apparatus of claim 11, wherein, The extracting module includes: The text detection model is trained in the following manner: collect playlist screenshot pictures as picture samples, the playlist screenshot pictures containing actual text position information, the actual text position information including position coordinates of a text box corresponding to the text and a size of the text box; input the picture samples into the text detection model to output predicted text position information in the picture samples; perform regression training on the text detection model based on the actual text position information and the predicted text position information.
13. The playlist generating apparatus of claim 11, wherein, The extracting module includes: The text recognition model is trained in the following manner: generate a picture containing random text based on a text format in a playlist screenshot; add a digital label to the random text based on a preset text dictionary; input the picture containing random text and the random text containing a digital label as training samples into the text recognition model for training to obtain a trained text recognition model.
14. The playlist generating apparatus of claim 11, wherein The extracting module is further configured to: perform noise reduction processing on the text information based on a noise reduction model before matching the text information with song information in a music library and using a matching result as the song information of a song in the playlist screenshot.
15. The playlist generating apparatus of claim 11, wherein, The determining module is specifically configured to: perform word segmentation processing on the lyric data based on a word segmentation tool; input the lyric data after the word segmentation processing into the word vector conversion model respectively to obtain word meaning vectors corresponding to the lyric data after the word segmentation processing; perform weighted average processing on all word meaning vectors corresponding to the lyric data to obtain a semantic vector corresponding to the lyric data.
16. The playlist generating apparatus of claim 11, wherein The determining module is specifically configured to: input the audio vector into a chord feature extraction network to obtain a chord feature of the audio vector; input the chord feature into the first classification network to output a chord style corresponding to the song; input the audio vector and the semantic vector into a song style feature extraction network to output a song style feature corresponding to the song; perform fusion processing on the song style feature and the chord feature; and perform fusion processing on the song style feature and the chord feature. Input the song style features and chord features after the fusion processing into a second classification network, and output a song style corresponding to the song.
17. The playlist generating apparatus of claim 16, wherein The determination module comprises: The second classification network is trained in the following manner: Obtain audio sample data for training, and label the style corresponding to the audio sample data and the chords with an occurrence frequency exceeding a set number or an occurrence number ratio exceeding a set proportion; Input the audio sample data into the second classification network, and output a predicted style in the audio sample data; Based on the labeled style corresponding to the audio sample data and the predicted style, the second classification network is subjected to regression training.
18. A computing device, comprising: Comprise: At least one processor; And a memory connected in communication with the at least one processor; Wherein the memory has instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the computing device to perform the playlist generation method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Method and device for making song list, storage medium and processor
CN109325143A
Song list extraction method based on morphological method
CN113723401A
Audio classification method and apparatus, intelligent device, and storage medium
WO2019109787A1