Song tag expansion method and apparatus, device, medium, and product
By extracting keyword vectors from user access to songs and playlists within specific geographical regions and time periods, matching them with scene word vectors, and generating extended song tags, this technology solves the problem that existing song tags cannot meet personalized needs, and achieves efficient and accurate song recommendation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGZHOU GESHEN INFORMATION TECH CO LTD
- Filing Date
- 2021-12-16
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies that rely on song attributes for song tagging cannot meet users' personalized needs, and manual tagging is costly and inefficient, making it difficult to recommend suitable songs in specific scenarios.
By acquiring the songs and playlists accessed by users within a geographical region and time period, keyword vectors are extracted and matched with a scene corpus to generate a set of similar scene word vectors, which serve as extended tags for the songs.
It enables intelligent expansion of song tags, saves manpower costs, and recommends songs that suit users' playback needs in specific scenarios, thus improving the accuracy of personalized recommendations.
Smart Images

Figure CN114168789B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of music information retrieval technology, and in particular to a method for expanding song tags and the corresponding apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Technology
[0002] With the rapid development of technology and the economy, users are increasingly pursuing healthy and beneficial spiritual needs. Listening to music in various scenarios is gaining popularity, leading to a massive increase in the total number of songs in the libraries of online music platforms. To facilitate quick access to these libraries or to allow platforms to recommend songs similar to or the same as previously listened-to tracks, existing technologies typically tag songs in the library. These tags are generally based on language, genre, artist, and album attributes, allowing for quick matching of similar songs and recommendations. However, these tags, based solely on song attributes, are increasingly failing to meet users' personalized needs. Users now prefer to listen to songs that resonate with their specific circumstances.
[0003] On the other hand, if the official tagging of songs is handled by each online music platform, it would require a large number of professionals with music backgrounds to manually tag the songs. This would involve high costs and relatively low efficiency. Moreover, manually identifying applicable scenarios for tagging based on the melody, rhythm, lyrics, and real-life contexts of the songs would be quite difficult.
[0004] In summary, given the current problem that users' personalized preferences cannot be effectively met, the applicant has made corresponding explorations in this regard. Summary of the Invention
[0005] The primary objective of this application is to solve at least one of the above-mentioned problems by providing a song tag expansion method and corresponding apparatus, computer equipment, computer-readable storage medium, and computer program product.
[0006] To achieve the various objectives of this application, the following technical solution is adopted:
[0007] A song tag expansion method provided for one of the purposes of this application includes the following steps:
[0008] Get the songs accessed by online users within a specified geographical area and time period, and the playlists they belong to;
[0009] Extract keyword vectors corresponding to each keyword from the description text of the songs and playlists;
[0010] The keyword vectors are matched with the scene word vectors pre-stored in the scene corpus to obtain a set of scene word vectors that are similar to the keyword vectors.
[0011] Based on the set of scene word vectors, obtain the corresponding set of scene words, and mark one or more scene words in the set of scene words as tags for the song.
[0012] In a further embodiment, obtaining the songs accessed by online users within a specified geographical region and time period, and the playlists to which they belong, includes the following steps:
[0013] Get all songs accessed by all users online within this geographic region and time period;
[0014] The number of visits to all the songs is statistically analyzed to obtain multiple visit rankings. The number of visits can be any one or more of the following: click play count, search play count, and complete play count.
[0015] Filter out a limited number of target songs with the highest number of visits from the aforementioned access list;
[0016] Obtain the playlist to which the target song belongs.
[0017] In the extended embodiment, before the step of obtaining the songs accessed by online users within a set geographical region and time period and the playlists to which they belong, the following steps are included:
[0018] Respond to any online user's login event, obtain the online user's geographical location information, set the geographical region to which the user belongs based on the geographical location information, and set the time period based on the trigger time of the login event;
[0019] or:
[0020] Construct a geographical region table and a time period table, setting multiple geographical regions in the geographical region table and multiple time periods in the time period table.
[0021] In a further embodiment, keyword vectors corresponding to each keyword are extracted from the description text of the song and playlist, including the following steps:
[0022] Obtain the description text of the song and the description text of the playlist;
[0023] The description text of the songs and their playlists is segmented into words, and multiple keywords are extracted from the segmented words;
[0024] Each keyword is vectorized to obtain multiple corresponding keyword vectors.
[0025] In a preferred embodiment, the keyword vector is matched with pre-stored scene word vectors in a scene corpus to obtain a set of scene word vectors that are similar to the keyword vector, including the following steps:
[0026] For each keyword vector, perform similarity matching with the pre-stored scene word vectors in the scene corpus to obtain the similarity between the keyword vector and each scene word vector;
[0027] For each keyword vector, scene word vectors with similarity exceeding a preset threshold are identified as scene word vectors that are similar to that keyword vector;
[0028] Construct a set of scene word vectors by constructing scene word vectors that are similar to the vectors of each keyword.
[0029] In a further embodiment, obtaining the corresponding scene word set based on the scene word vector set, and marking one or more scene words in the scene word set as tags for the song, includes the following steps:
[0030] Search the scene corpus to determine the scene word corresponding to each scene word vector in the scene word vector set, and obtain the scene word set;
[0031] According to preset rules, multiple scene words in the scene word set are combined to obtain a combined tag set, wherein the combined tag set includes one or any number of scene words selected from the scene word set;
[0032] The song is annotated, and each scene word in the combined tag set is used as an extended tag for the song.
[0033] In a further embodiment, the scenario words are natural language words used to describe any one or more of natural phenomena, social activities, traffic phenomena, and geographical regions.
[0034] A song tag expansion device provided for one of the purposes of this application includes:
[0035] The region acquisition module is used to acquire songs accessed by online users within a specified geographical region and time period, as well as the playlists to which they belong.
[0036] The vector extraction module is used to extract and generate keyword vectors corresponding to each keyword from the description text of the song and the playlist.
[0037] The similarity matching module is used to perform similarity matching between the keyword vector and the scene word vectors pre-stored in the scene corpus to obtain a set of scene word vectors that are similar to the keyword vector;
[0038] The extended tag module is used to obtain the corresponding scene word set based on the scene word vector set, and to mark one or more scene words in the scene word set as tags for the song.
[0039] In a further embodiment, the region acquisition module includes:
[0040] The user song acquisition submodule is used to acquire all songs accessed by all users online within the specified geographical region and time period.
[0041] The access statistics submodule is used to count all the songs based on the access volume and obtain multiple access rankings. The access volume can be any one or any combination of click playback volume, search playback volume, and complete playback volume.
[0042] The access volume filtering submodule is used to filter out a limited number of target songs with the highest access volume in the access list;
[0043] The playlist acquisition submodule is used to acquire the playlist to which the target song belongs.
[0044] In an extended embodiment, prior to the user song acquisition submodule, the following steps are included:
[0045] The location area module is used to respond to login events of any online user, obtain the geographical location information of the online user, set the geographical area to which the user belongs based on the geographical location information, and set the time period based on the trigger time of the login event;
[0046] or:
[0047] The preset region module is used to construct a geographic region table and a time period table. Multiple geographic regions can be set in the geographic region table, and multiple time periods can be set in the time period table.
[0048] In a further embodiment, the vector extraction module includes:
[0049] The text acquisition submodule is used to acquire the description text of the song and the description text of the playlist;
[0050] The text segmentation submodule is used to segment the description text of the song and its playlist into words and extract multiple keywords from the segmented text.
[0051] The word vectorization submodule is used to vectorize each keyword to obtain multiple corresponding keyword vectors.
[0052] In a preferred embodiment, the similarity matching module includes:
[0053] The vector similarity matching submodule is used to perform similarity matching between each keyword vector and the pre-stored scene word vectors in the scene corpus to obtain the similarity between the keyword vector and each scene word vector;
[0054] The vector filtering submodule is used to identify scene word vectors with similarity exceeding a preset threshold as similar to the keyword vector for each keyword vector.
[0055] The vector set construction submodule is used to construct a scene word vector set by combining scene word vectors that are similar to the individual keyword vectors.
[0056] In a further embodiment, the expanded label module includes:
[0057] The scene word acquisition submodule is used to search the scene corpus, determine the scene word corresponding to each scene word vector in the scene word vector set, and obtain the scene word set.
[0058] The scene word combination submodule is used to combine multiple scene words in the scene word set according to preset rules to obtain a combination tag set, wherein the combination tag set includes one or any number of scene words selected from the scene word set;
[0059] An extended song tag submodule is used to tag the song, using each scene word in the combined tag set as an extended tag for the song.
[0060] A computer device provided for one of the purposes of this application includes a central processing unit and a memory, the central processing unit being configured to invoke and run a computer program stored in the memory to perform the steps of the song tag expansion method described in this application.
[0061] A computer-readable storage medium is provided for another purpose of this application, which stores, in the form of computer-readable instructions, a computer program implemented according to the described song tag expansion method, which, when invoked by a computer, performs the steps included in the method.
[0062] A computer program product provided for another purpose of this application includes a computer program / instructions that, when executed by a processor, implement the steps of the method described in any embodiment of this application.
[0063] Compared with existing technologies, the advantages of this application are as follows:
[0064] This application extracts keywords from songs accessed by users within a specific geographical area and time period, as well as from the descriptions of those songs and playlists, in real time. It then matches these keywords with semantically similar scene words and combines them to construct extended tags for the songs. This application saves significant manpower costs for intelligent song tagging. Based on the reality that similar scene events are highly probable within a target area and time period, it extracts and constructs extended tags from the descriptions of accessed songs and associated playlists revealed by real-time, on-site user behavior data. These extended tags contain scene characteristics specific to the time and place, resulting in recommended and retrieved songs highly suitable for users' playback needs in corresponding scenarios. Attached Figure Description
[0065] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0066] Figure 1 A flowchart illustrating a typical embodiment of the song tag expansion method of this application;
[0067] Figure 2 This is a schematic diagram illustrating the process of obtaining songs and their respective playlists in an embodiment of this application.
[0068] Figure 3 This is a schematic diagram illustrating the process of extracting and vectorizing keywords from the song and its playlist in this embodiment of the application.
[0069] Figure 4 This is a flowchart illustrating the similarity matching process in an embodiment of this application;
[0070] Figure 5 A schematic diagram illustrating the process of expanding the song tags based on matched scene words in this application embodiment;
[0071] Figure 6 A schematic diagram of the song tag expansion device of this application;
[0072] Figure 7 This is a schematic diagram of the structure of a computer device used in this application. Detailed Implementation
[0073] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.
[0074] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this application means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.
[0075] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.
[0076] Those skilled in the art will understand that the terms "client," "terminal," and "terminal device" as used herein include both devices that receive wireless signals, devices that only possess wireless signal receiver capabilities without transmission capabilities, and devices with receiving and transmitting hardware, devices that have receiving and transmitting hardware capable of bidirectional communication over a bidirectional communication link. Such devices may include: cellular or other communication devices such as personal computers or tablets, having single-line displays, multi-line displays, or cellular or other communication devices without multi-line displays; PCS (Personal Communications Service) that can combine voice, data processing, fax, and / or data communication capabilities; PDA (Personal Digital Assistant) that may include a radio frequency receiver, pager, internet / intranet access, web browser, notepad, calendar, and / or GPS (Global Positioning System) receiver; and conventional laptops and / or handheld computers or other devices that have and / or include radio frequency receivers. As used herein, "client," "terminal," and "terminal device" can be portable, transportable, installed in a means of transportation (air, sea, and / or land), or suitable and / or configured to operate locally and / or in a distributed manner, operating in any other location on Earth and / or in space. "Client," "terminal," and "terminal device" as used herein can also be a communication terminal, an internet access terminal, or a music / video playback terminal, such as a PDA, a MID (Mobile Internet Device), and / or a mobile phone with music / video playback capabilities, or a smart TV, set-top box, etc.
[0077] The hardware referred to by the names "server," "client," and "service node" in this application is essentially an electronic device with the equivalent capabilities of a personal computer. It is a hardware device with the necessary components revealed by the von Neumann architecture, such as a central processing unit (including an arithmetic logic unit and a control unit), memory, input devices, and output devices. The computer program is stored in its memory, and the central processing unit loads the program stored in the secondary storage into the main memory to run it, execute the instructions in the program, and interact with the input and output devices to complete specific functions.
[0078] It should be noted that the concept of "server" used in this application can also be extended to the case of server clusters. Based on the network deployment principles understood by those skilled in the art, the servers should be logically divided. Physically, these servers can be independent of each other but accessible through interfaces, or they can be integrated into a single physical computer or a computer cluster. Those skilled in the art should understand this flexibility and should not use it to constrain the implementation of the network deployment method in this application.
[0079] One or more of the technical features of this application, unless explicitly specified herein, can be deployed on a server and accessed by a client remotely calling the online service interface provided by the server, or can be directly deployed and run on a client for access.
[0080] Unless otherwise specified, the neural network models referenced or potentially referenced in this application may be deployed on a remote server and invoked remotely on the client, or deployed on a client with the capability to invoke directly. In some embodiments, when running on the client, the corresponding intelligence may be acquired through transfer learning in order to reduce the requirements on the client's hardware resources and avoid excessive consumption of the client's hardware resources.
[0081] Unless otherwise specified, all data involved in this application may be stored remotely on a server or on a local terminal device, as long as it is suitable for use by the technical solution of this application.
[0082] Those skilled in the art will understand that although the various methods in this application are described based on the same concept and thus present commonality among them, they can be performed independently unless otherwise specified. Similarly, the various embodiments disclosed in this application are all based on the same inventive concept; therefore, concepts expressed in the same way, as well as concepts that are appropriately changed for convenience but are expressed differently, should be understood equivalently.
[0083] Unless otherwise expressly stated, the various embodiments disclosed in this application can be combined in a cross-cutting manner to flexibly construct new embodiments, as long as such combination does not depart from the inventive spirit of this application and can meet the needs of the prior art or solve a certain deficiency in the prior art. Those skilled in the art should be aware of such modifications.
[0084] The song tag expansion method of this application can be programmed into a computer program product and deployed on a server to run. In this way, the client can access the interface opened by the computer program product after it runs in the form of a web program or application, and realize human-computer interaction with the process of the computer program product through a graphical user interface.
[0085] Please see Figure 1 The song tag expansion method of this application, in its typical embodiment, includes the following steps:
[0086] Step S1100: Obtain the songs accessed by online users within a set geographical area and time period, and the playlists to which they belong;
[0087] The defined geographic region can be either a real-time geographic region near the user's location or a pre-defined geographic region. The real-time geographic region is defined after the user logs in and authorizes the application installed on their terminal device to access the device's location information. This location information is then used as a base point, and the Manhattan distance (e.g., 5km) is used to calculate the corresponding boundary points using the Manhattan algorithm. The geographic region is then defined based on these boundary points and the base point. The pre-defined geographic region is a table element within a pre-constructed geographic region table. This table is divided by province, city, and district, with the corresponding geographic region at the district level being the table element.
[0088] The set time period is either a real-time time period set based on user login or a preset time period set in advance. The real-time time period is the time point corresponding to the user's login, and then the time period corresponding to one hour after that time point. The preset time period is a pre-built time period table, in which the table elements are time periods. The table is divided into 24 hours of a day according to the fact that each time period is one hour to obtain the various time periods.
[0089] The term "online user" refers to a user who performs the corresponding login operation in the graphical user interface displayed by the application installed on their terminal device, and remains logged in until they log out.
[0090] Furthermore, real-time tasks or scheduled tasks can be set accordingly. The real-time task involves the server accessing the user database to query the unique identifier of the online user within the real-time time period and the real-time geographical area, and then obtaining the song accessed corresponding to the unique identifier in the song database, as well as the playlist containing that song in the song database. The scheduled task involves similarly obtaining the songs accessed by online users within the preset geographical area and their corresponding playlists every day during a preset time period. In other words, by obtaining user behavior data of online users accessing the song database within a specific geographical area and a specific time period, the server can determine the songs accessed by these users based on their user behavior data, and further obtain the corresponding playlists for those songs.
[0091] Step S1200: Extract keyword vectors corresponding to each keyword from the description text of the song and playlist;
[0092] The description text of the song may include any one or more of the following information: song title, lyricist's name, composer's name, singer's name, lyrics, release date, album name, etc. Typically, the song title, singer's name, and lyrics are the main description text.
[0093] The playlists include user-created private playlists and / or official playlists created by the platform. The description text of each playlist mainly includes its introductory text and / or title text. The private playlists are generally composed of multiple songs of the same type, such as scene, language, song era, genre, and artist, which are freely selected by the user according to their subjective preferences. The official playlists are generally composed of multiple songs of the same type selected by the platform based on objective criteria such as the number of plays, comments, favorites, likes, and shares.
[0094] The current playlist contains the songs. In order to obtain the keywords of the songs and the playlists to which they belong, the main descriptive text of the songs and the introduction text and title text of the playlists are extracted. A text feature extraction model that has been pre-trained to convergence is used to preprocess the extracted text information of the playlists and songs by word segmentation to obtain their respective keywords. Then, the keywords are encoded to form keyword vectors for feature representation. The keyword vectors are feature vectors that represent the deep semantic information of each.
[0095] Step S1300: Perform similarity matching between the keyword vector and the scene word vectors pre-stored in the scene corpus to obtain a set of scene word vectors that are similar to the keyword vector;
[0096] The scene corpus is a database that pre-stores all scene words and their corresponding scene word vectors. These scene words are collected by relevant professionals based on real-world scenarios. A pre-trained text feature extraction model, having reached convergence, is invoked to extract deep semantic information features from all scene words, constructing corresponding scene word vectors. For ease of subsequent use, these scene word vectors are mapped and associated with their corresponding scene words and stored in the scene corpus. The text feature extraction model can be implemented using various pre-trained neural network models based on CNNs or RNNs, as available in existing technologies. It is important to note that the text feature extraction model described here is trained for a different task than the text feature extraction model used in the previous step.
[0097] Based on the scene corpus, the similarity between the keyword vector and each scene word vector in the scene corpus is calculated, and then scene word vectors with similarity greater than a preset threshold are selected and sorted according to their similarity to form a scene word vector set.
[0098] Step S1400: Obtain the corresponding scene word set based on the scene word vector set, and mark one or more scene words in the scene word set as tags of the song.
[0099] By establishing a one-to-one mapping relationship between scene word vectors and their corresponding scene words, the scene words corresponding to each scene word vector in the scene word vector set are obtained accordingly. Based on this, a scene word set is constructed. Then, the scene words in the scene word set are combined according to rules or arbitrarily, i.e., one or more are randomly selected for combination. Subsequently, each combination is used as a tag mapping associated with the song and stored to complete the expansion operation of the song tag set with scene class tags.
[0100] As disclosed in this exemplary embodiment, this application extracts keywords from songs accessed by users within a certain geographical area and time period, as well as from the description text of the playlists, in real time. It then matches these keywords with semantically highly similar scene words and further combines these scene words to construct extended tags for the songs. This application saves significant manpower costs for intelligent song tagging. Based on the reality that similar scene events are highly probable within a target area and time period, it extracts and constructs the extended tags using real-time user behavior data revealing the description information of accessed songs and their associated playlists. These extended tags contain scene characteristics specific to the time and place, resulting in recommended and retrieved songs highly suitable for users' playback needs in corresponding scenarios.
[0101] Please see Figure 2 In a further embodiment, obtaining the songs accessed by online users within a set geographical region and time period, and the playlists to which they belong, includes the following steps:
[0102] Step S1110: Obtain all songs accessed by all users online within the geographical region and time period;
[0103] In one embodiment, when performing the real-time task, based on the login operation of an online user, within a time period of one hour (which can be flexibly designed) after the current time point, the geographical area within a 5km (which can be flexibly adjusted) range of the user is calculated according to the real-time location of the user's authorized terminal device. Furthermore, the unique identifier code corresponding to all online users within the geographical area is obtained by accessing the user database, and then the access records of the song database are queried to obtain all the songs accessed by each online user from the song database.
[0104] In another embodiment, when each preset time period is reached, the timed task is triggered, and similarly, all songs accessed by all online users in each preset geographical region in the pre-built geographical region table can be obtained.
[0105] Step S1120: Statistically analyze all the songs based on the number of visits to obtain multiple visit rankings. The number of visits can be any one or more of the following: click play count, search play count, and complete play count.
[0106] Furthermore, the access records in the song database are queried to obtain the click play count, search play count, and complete play count for all songs. Based on this, each song is sorted in descending order of different play counts, thus obtaining three access lists: click play count list, search play count list, and complete play count list. The click play count is the number of online users who clicked and selected the song to play through the graphical user interface, including partial playback and complete playback. The search play count is the number of online users who searched and selected the song to play through the search entry provided in the graphical user interface, including partial playback and complete playback. The complete play count is the number of online users who played the song completely.
[0107] Step S1130: Filter out a limited number of target songs with the highest number of visits in the access list;
[0108] Based on the access rankings: click on the play count ranking, search for the play count ranking, and sort the play count in the complete play count ranking, and filter out the top 100 songs in each access ranking as target songs. The number of songs to be filtered based on this ranking can be flexibly set by those skilled in the art depending on the number of active online users in the actual scenario.
[0109] Step S1140: Obtain the playlist to which the target song belongs.
[0110] Query the playlists mapped to each target song in the song database, that is, the playlists that contain the target songs. The mapping association includes one target song associated with one or more playlists, thereby obtaining multiple playlists containing the target songs.
[0111] In this embodiment, by scientifically and feasiblely setting geographical regions and time periods, the songs accessed by users and their respective playlists can be obtained in real time. Furthermore, based on real-time data, the songs accessed by users are optimized. The selected songs not only have a certain degree of popularity that indicates the degree of user liking, but also have a certain representative significance in this geographical region at this time.
[0112] In the extended embodiment, before the step of obtaining the songs accessed by online users within a set geographical region and time period and the playlists to which they belong, the following steps are included:
[0113] Step S1000: Respond to the login event of any online user, obtain the geographical location information of the online user, set the geographical region to which the user belongs based on the geographical location information, and set the time period based on the trigger time of the login event;
[0114] The online user is defined as someone who performs the login operation through the graphical user interface displayed by the application installed on their terminal device. The user remains logged in until they log out, indicating they are in a "stay logged in" state. The application's server then responds to the login event. For privacy protection of user information and security protection of the user's terminal device's location information, authorization is required only with the user's consent. Therefore, after the login event is triggered, it checks whether permission has been granted to access the user's login input and the location information stored on the terminal device during location tracking. If the location information-related permissions are granted, it means the user has already granted the permissions and has not previously revoked them. Therefore, the application responds to the detection event without needing to re-grant the permissions. Otherwise, it means the user has not previously granted the permissions, or has previously granted them but subsequently revoked them. In this case, the application responds to the detection event and grants the permissions, displaying a pop-up window at the bottom of the application's graphical user interface requesting user information and location permission authorization. The user can then click the relevant buttons to grant or deny the permissions.
[0115] The time point corresponding to the server response to the login event triggered by the login operation of any online user is used as a reference to set the time point within one hour thereafter as the set time period. The specific duration can be flexibly set by those skilled in the art according to the actual business scenario.
[0116] Once the application obtains the corresponding permissions through user authorization, it calls the location information of the online user's terminal device to obtain the corresponding geographical location information. The geographical location is used as a base point to set an area within a 5km radius. This range can be flexibly set by those skilled in the art according to the actual business scenario, or the area corresponding to the district level of the province, city, or district to which the geographical location belongs can be set as the geographical area to which the online user belongs.
[0117] In another extended embodiment, unlike the previous embodiment, before the step of obtaining the songs accessed by online users within a set geographical region and time period and the playlists to which they belong, the following steps are included:
[0118] Step S1001: Construct a geographic region table and a time period table, setting multiple geographic regions in the geographic region table and multiple time periods in the time period table.
[0119] The geographical region table is divided according to province, city, and district. The district level corresponds to the geographical region data of the table, and the corresponding province and city are the header data of the table. For example, Guangdong Province / Guangzhou City is a header data in the table, and the corresponding Huangpu District, Tianhe District, Haizhu District, Yuexiu District, Baiyun District, etc. are the corresponding geographical region data in the table.
[0120] The time period table is divided into multiple time periods from 6:00 AM to 12:00 PM, with each time period limited to one hour. The data in the table represents each time period, such as 6:00 AM to 7:00 AM, 7:00 AM to 8:00 AM, etc.
[0121] In this embodiment, on the one hand, the geographical area and time period are objectively and feasiblely divided according to the actual situation of the user. On the other hand, the geographical area and time period are intuitively and effectively divided in advance, so as to set the corresponding real-time tasks and timed tasks according to these two aspects. Thus, the appropriate tasks can be selected and executed according to the actual situation.
[0122] Please see Figure 3 In a further embodiment, keyword vectors corresponding to each keyword are extracted from the description text corresponding to the song and playlist, including the following steps:
[0123] Step S1210: Obtain the description text of the song and the description text of the playlist;
[0124] For example, the song title, artist name, and lyrics text in the song description text, as well as the introduction text and title text in the playlist description text, can be obtained.
[0125] Step S1220: Segment the description text of the song and its playlist into words, and extract multiple keywords from the segmented text;
[0126] The text feature extraction model, pre-trained to convergence, is invoked to segment the description text of the song and the playlist. In one embodiment, the text feature extraction model is BERT. Based on two segmentation modules in BERT: BasicTokenizer and WordpieceTokenizer, the description text of the song and the playlist is input. First, BasicTokenizer is used to obtain a coarsely segmented list of tokens. Then, WordpieceTokenizer is used on each token to obtain the corresponding segmentation results as keywords. The BasicTokenizer is a preliminary segmenter. For a string to be segmented, the process is roughly as follows: Unicode, removal of empty characters, character replacement, control characters and strange characters such as whitespace, Chinese word segmentation, space segmentation, conversion of uppercase English letters to lowercase English letters, removal of diacritics, removal of punctuation for word segmentation, and space segmentation again. The WordpieceTokenizer is a further segmentation based on the BasicTokenizer result to obtain sub-words. For Chinese text, this segmenter module can be omitted during the word segmentation process.
[0127] In summary, the preferred text feature extraction model is a mature model such as BERT, Electra, or TextCNN, which can be flexibly adopted by those skilled in the art, as long as sufficient data samples are used for pre-training. Each data sample can be manually labeled with keywords from the description text corresponding to the playlist and songs. During training, these keywords are converted into corresponding encoded vectors, which are then processed by the text feature extraction model to obtain the corresponding text feature vectors. A classifier is then applied to these vectors for classification prediction. The cross-entropy loss of the model is obtained by using the difference between the supervision label corresponding to the data sample and the classification prediction result. When the loss value is not close to the preset threshold, the model weights are updated with gradients, and the next data is used for iterative training until the model's loss value reaches the preset threshold, confirming model convergence. The model can then be used in this application.
[0128] Step S1230: Vectorize each keyword to obtain multiple corresponding keyword vectors.
[0129] Furthermore, the text feature extraction model performs representation learning on each keyword to obtain the corresponding keyword vector.
[0130] In this embodiment, the text feature extraction model is used to quickly and easily extract keywords from songs and playlists, ensuring a certain level of accuracy while saving a lot of manpower, which is highly efficient.
[0131] Please see Figure 4 In a preferred embodiment, the keyword vector is matched with pre-stored scene word vectors in a scene corpus to obtain a set of scene word vectors that are similar to the keyword vector, including the following steps:
[0132] Step S1310: For each keyword vector, perform similarity matching with the scene word vectors pre-stored in the scene corpus to obtain the similarity between the keyword vector and each scene word vector;
[0133] The similarity between the keyword vector and the scene word vector is calculated by calling the data distance calculation formula, such as the Euclidean distance algorithm, cosine similarity algorithm, Jaccard algorithm, Pearson correlation coefficient algorithm, and other formulas familiar to those skilled in the art. Thus, a similarity sequence is obtained for each keyword vector, and the similarity value between the keyword vector and each scene word vector is stored in the similarity sequence.
[0134] Step S1320: For each keyword vector, determine the scene word vectors with similarity exceeding a preset threshold as scene word vectors that are similar to the keyword vector;
[0135] Based on the similarity sequence corresponding to the keyword vector, the similarity value of each element is compared with the preset threshold. When the former is greater than or equal to the latter, the scene word vector corresponding to the similarity value is determined. Thus, the scene word vector selected accordingly has the characteristic of being highly similar to the keyword vector in semantics. The preset threshold can be set by those skilled in the art based on relevant experience or experiments.
[0136] Step S1330: Construct a set of scene word vectors by forming scene word vectors that are similar to each keyword vector.
[0137] The scene word vectors that are highly semantically similar to each keyword vector are constructed into a scene word vector set, which is then encapsulated into a corresponding array. This makes it easy to quickly call the scene word vectors by traversing the array in subsequent steps.
[0138] In this embodiment, a similarity matching method is used to accurately match scene words from the scene corpus that are highly similar in semantics to the keywords in the song and the playlist to which it belongs. Compared with the simple XOR matching method of the keyword and the scene word, which requires that they be completely identical to match successfully and otherwise fail to match, it is evident that a richer and more valuable scene words can be matched.
[0139] Please see Figure 5 In a further embodiment, obtaining the corresponding scene word set based on the scene word vector set, and marking one or more scene words in the scene word set as tags for the song, includes the following steps:
[0140] Step S1410: Search the scene corpus, determine the scene word corresponding to each scene word vector in the scene word vector set, and obtain the scene word set;
[0141] Based on the mapping relationship between the scene word vectors in the scene word vector set and their corresponding scene words, the scene corpus is searched to obtain the corresponding scene words. Furthermore, each scene word is used as an array of data structures to construct a corresponding scene word set.
[0142] Step S1420: Combine multiple scene words in the scene word set according to preset rules to obtain a combined tag set, wherein the combined tag set includes one or any number of scene words selected from the scene word set;
[0143] The scenario terms are natural language words used to describe any one or more of the following: natural phenomena, social activities, traffic phenomena, and geographical regions. Natural phenomena include rain, cloudy days, sunny days, and snow. Social activities include prenatal education, sleep, going to work, and driving. Traffic phenomena include traffic jams, high-speed driving, and low-speed driving. Geographical regions include basketball courts, homes, gyms, and karaoke rooms.
[0144] The scene word set array contains multiple arrays, namely the natural phenomenon array, the social activity array, the traffic phenomenon array, and the geographical region array, with each array containing corresponding scene words.
[0145] The preset rule is to randomly select one or more arrays from the scene word set. If an array is selected, the scene words stored in the array are directly used to construct a combined tag set. If multiple arrays are selected, the scene words in each array are combined to construct a combined tag set. The specific implementation method can be selected by those skilled in the art as needed.
[0146] Step S1430: Tag the song and use each scene word in the combined tag set as an extended tag for the song.
[0147] The song tag set corresponding to the song is queried in the song tag library. The scene words corresponding to each combination tag in the combination tag set are added to the tag set as extended tags to complete the annotation of the song.
[0148] In this embodiment, the song and the scene words corresponding to its playlist are further used as extended tags to label the song, and the extended tags are enriched by randomly combining scene words, so that when the user is in a complex scene contained in the extended tags, the extended tags can quickly and accurately match the corresponding song for playback.
[0149] This application provides a song tag expansion device, functionally deployed to adapt to the song tag expansion method of this application, including: a region acquisition module 1100, a vector extraction module 1200, a similarity matching module 1300, and the expanded tag module 1400. The region acquisition module 1100 is used to acquire songs accessed by online users within a set geographical region and time period, and their respective playlists. The vector extraction module 1200 is used to extract keyword vectors corresponding to each keyword from the description text corresponding to the songs and playlists. The similarity matching module 1300 is used to perform similarity matching between the keyword vectors and pre-stored scene word vectors in a scene corpus to obtain a set of scene word vectors similar to the keyword vectors. The expanded tag module 1400 is used to obtain a set of scene words corresponding to the set of scene word vectors, and mark one or more scene words in the scene word set as tags for the songs.
[0150] In a further embodiment, the region acquisition module 1100 includes: a user song acquisition submodule, used to acquire all songs accessed by all users online within the geographical region and time period; an access volume statistics submodule, used to perform statistics on all songs based on access volume to obtain multiple access lists, wherein the access volume is any one or more of click playback volume, search playback volume, and complete playback volume; an access volume filtering submodule, used to filter out a limited number of target songs with the highest access volume in the access lists; and a playlist acquisition submodule, used to acquire the playlist to which the target songs belong.
[0151] In an extended embodiment, before the user song acquisition submodule, there is a: a location area module, used to respond to the login event of any online user, obtain the geographical location information of the online user, set the geographical area to which the user belongs based on the geographical location information, and set the time period based on the trigger time of the login event; or: a preset area module, used to construct a geographical area table and a time period table, set multiple geographical areas in the geographical area table, and set multiple time periods in the time period table.
[0152] In a further embodiment, the vector extraction module 1200 includes: a text acquisition submodule, used to acquire the description text of the song and the description text of the playlist; a text segmentation submodule, used to segment the description text of the song and the playlist into words and extract multiple keywords from the segmented words; and a word vectorization submodule, used to vectorize each keyword to obtain multiple corresponding keyword vectors.
[0153] In a preferred embodiment, the similarity matching module 1300 includes: a vector similarity matching submodule, used to perform similarity matching between each keyword vector and pre-stored scene word vectors in a scene corpus to obtain the similarity between the keyword vector and each scene word vector; a vector filtering submodule, used to determine scene word vectors with similarity exceeding a preset threshold as similar scene word vectors to the keyword vector for each keyword vector; and a vector set construction submodule, used to construct a scene word vector set from the scene word vectors that are similar to each keyword vector.
[0154] In a further embodiment, the extended tag module 1400 includes: a scene word acquisition submodule, used to search the scene corpus, determine the scene word corresponding to each scene word vector in the scene word vector set, and obtain a scene word set; a scene word combination submodule, used to combine multiple scene words in the scene word set according to preset rules to obtain a combination tag set, the combination tag set including one or any number of scene words selected from the scene word set; and an extended song tag submodule, used to tag the song, using each scene word in the combination tag set as an extended tag for the song.
[0155] To address the aforementioned technical problems, embodiments of this application also provide computer equipment. For example... Figure 7 The diagram shows the internal structure of a computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected via a system bus. The computer-readable storage medium stores an operating system, a database, and computer-readable instructions. The database may store a sequence of control information. When the computer-readable instructions are executed by the processor, they enable the processor to implement a song tag expansion method. The processor of the computer device provides computing and control capabilities, supporting the operation of the entire computer device. The memory of the computer device may store computer-readable instructions, which, when executed by the processor, enable the processor to execute the song tag expansion method of this application. The network interface of the computer device is used for communication with a terminal. Those skilled in the art will understand that… Figure 7The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0156] In this embodiment, the processor is used to execute... Figure 6 The system contains the specific functions of each module and its sub-modules. The memory stores the program code and various data required to execute these modules or sub-modules. The network interface is used for data transmission between the user terminal and the server. In this embodiment, the memory stores the program code and data required to execute all modules / sub-modules in the song tag expansion device of this application. The server can call the server's program code and data to execute the functions of all sub-modules.
[0157] This application also provides a storage medium storing computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to perform the steps of the song tag expansion method of any embodiment of this application.
[0158] This application also provides a computer program product, including a computer program / instructions that, when executed by one or more processors, implement the steps of the method described in any embodiment of this application.
[0159] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. This computer program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0160] In summary, this application obtains user-accessed songs and keywords from their playlists within a certain geographical area in real time or via a scheduled task. These keywords are then matched with scene words in a scene corpus to obtain similar scene words as extended tags for the songs. These extended tags are based on actual scenes and thus can represent scene characteristics. Subsequently, songs corresponding to certain scene characteristics can be obtained based on these extended tags and pushed to online users in the corresponding scene. This intelligent tagging method not only associates specific scenes to meet users' personalized and immediate needs but also saves a significant amount of manpower for tagging.
[0161] Those skilled in the art will understand that the steps, measures, and solutions in the various operations, methods, and processes discussed in this application can be alternated, modified, combined, or deleted. Furthermore, other steps, measures, and solutions in the various operations, methods, and processes discussed in this application can also be alternated, modified, rearranged, decomposed, combined, or deleted. Furthermore, steps, measures, and solutions in the prior art that are similar to those disclosed in this application can also be alternated, modified, rearranged, decomposed, combined, or deleted.
[0162] The above description is only a partial embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for expanding song tags, characterized in that, Includes the following steps: Respond to any online user's login event, obtain the online user's geographical location information, set the geographical region to which the user belongs based on the geographical location information, and set the time period based on the trigger time of the login event; Get the songs accessed by online users within a specified geographical area and time period, and the playlists they belong to; Extract keyword vectors corresponding to each keyword in the description text of the song, and extract keyword vectors corresponding to each keyword in the description text of the playlist. The keyword vectors represent the deep semantic information of the corresponding keywords. The keyword vectors are matched with pre-stored scene word vectors in a scene corpus to obtain a set of scene word vectors that are similar to the keyword vectors. This includes: for each keyword vector, matching it with pre-stored scene word vectors in the scene corpus to obtain the similarity between the keyword vector and each scene word vector, where the scene word vectors represent the deep semantic information of the corresponding scene words; for each keyword vector, scene word vectors with similarity exceeding a preset threshold are identified as scene word vectors that are similar to the keyword vector; and the scene word vectors that are similar to each keyword vector are constructed into a set of scene word vectors. Obtaining a corresponding set of scene words based on the set of scene word vectors, marking one or more scene words in the set of scene words as tags for the song, and recommending songs to the user based on the tags, includes the following steps: Search the scene corpus to determine the scene word corresponding to each scene word vector in the scene word vector set, and obtain the scene word set; By combining multiple scene words corresponding to various description types in the scene word set, a combined tag set is obtained. The combined tag set includes one or any number of scene words selected from the scene word set. The description types include natural phenomena, social activities, traffic phenomena, and any combination of geographical regions. The song is annotated, and each scene word in the combined tag set is used as an extended tag for the song; The process of obtaining songs accessed by online users within a specified geographical region and time period, along with their respective playlists, includes the following steps: Get all songs accessed by all users online within this geographic region and time period; The number of visits is used to count all the songs and obtain multiple visit rankings corresponding to different numbers of visits. The number of visits can be any of the following: click play count, search play count, and total play count. Select a limited number of songs with the highest number of visits from the aforementioned access list as the target songs to be obtained. The playlist to which the target song belongs is obtained as the target playlist.
2. The song tag expansion method according to claim 1, characterized in that, Before obtaining the songs accessed by online users within a specified geographical region and time period, and the playlists they belong to, the following steps are included: Construct a geographical region table and a time period table, setting multiple geographical regions in the geographical region table and multiple time periods in the time period table.
3. The song tag expansion method according to claim 1, characterized in that, Extracting keyword vectors corresponding to each keyword from the description text of the songs and playlists includes the following steps: Obtain the description text of the song and the description text of the playlist; The description text of the songs and their playlists is segmented into words, and multiple keywords are extracted from the segmented words; Each keyword is vectorized to obtain multiple corresponding keyword vectors.
4. The song tag expansion method according to any one of claims 1 to 3, characterized in that, The scenario terms are natural language terms used to describe any one or more of the following: natural phenomena, social activities, traffic phenomena, and geographical regions.
5. A computer device, comprising a central processing unit and a memory, characterized in that, The central processing unit is used to invoke and run a computer program stored in the memory to perform the steps of the method as described in any one of claims 1 to 4.
6. A computer-readable storage medium, characterized in that, It stores, in the form of computer-readable instructions, a computer program implemented according to any one of claims 1 to 4, which, when invoked by a computer, executes the steps included in the corresponding method.
7. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method described in any one of claims 1 to 4.