Music pushing method and device and electronic equipment
By acquiring multi-dimensional tags of candidate music and historical behavioral data of target users, and using a text generation model to generate personalized push text, the problems of low efficiency and severe homogenization in existing technologies are solved, thereby improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD
- Filing Date
- 2025-12-30
- Publication Date
- 2026-05-12
AI Technical Summary
In existing music recommendation methods, manually generated text content is inefficient and highly homogenized, while template-generated text content does not match the actual song information well, resulting in a poor user experience.
By acquiring multi-dimensional tags of candidate music and historical behavioral data of target users, a text generation model is used to generate personalized push text. Based on the target user's preference parameters, the target music is determined from the candidate music and pushed to the user.
It improved the user experience by making users more easily attracted through personalized push texts, with logical flow, and enhanced users' interest in music.
Smart Images

Figure CN122019828A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and technology, and more specifically, to a music streaming method, apparatus, and electronic device. Background Technology
[0002] With the relevant technology, music platforms typically generate music text content in two ways. One is that operations staff manually generate text content for each song and distribute it in batches to platform users; the other is that some matching rules are set in advance, each rule corresponds to a set of text templates, and when distributing, the user and song information are matched, and the placeholders in the text templates are replaced before distribution.
[0003] The aforementioned methods of manually generating text content are inefficient and result in highly homogenized content, leading to a poor user experience. Similarly, template-based text generation methods are prone to logical inconsistencies and stylistic mismatches when the template doesn't accurately match the actual song information, also resulting in a poor user experience. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a music push method, apparatus and electronic device, so that users are more easily attracted by the push text of the music, thereby generating interest in the music and improving the user experience.
[0005] In a first aspect, embodiments of the present invention provide a music push method, which includes: acquiring multiple candidate music tracks; each candidate music track having corresponding text content; the text content being generated by a text generation model based on relevant data of the corresponding candidate music track; the relevant data including multi-dimensional tags of the music track, the multi-dimensional tags including metadata of the music track, lyric features, and various user feedback data; determining a target music track from the multiple candidate music tracks based on the target user's preference parameters for the multiple candidate music tracks; the preference parameters being determined based on the target user's historical behavior data regarding the candidate music tracks; and pushing the target music track and the target text content corresponding to the target music track to the target user.
[0006] Secondly, embodiments of the present invention provide a music push device, the device comprising: a music acquisition module for acquiring multiple candidate music tracks; each candidate music track having corresponding text content; the text content being generated by a text generation model based on relevant data of the corresponding candidate music track; the relevant data including multi-dimensional tags of the music track, the multi-dimensional tags including metadata of the music track, lyric features, and various user feedback data; a target music determination module for determining a target music track from the multiple candidate music tracks based on the target user's preference parameters for the multiple candidate music tracks; the preference parameters being determined based on the target user's historical behavior data regarding the candidate music tracks; and a music push module for pushing the target music track and the target text content corresponding to the target music track to the target user.
[0007] Thirdly, embodiments of the present invention provide an electronic device, including a processor and a memory, wherein the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the music push method described above. Fourthly, embodiments of the present invention provide a machine-readable storage medium storing machine-executable instructions. When the machine-executable instructions are invoked and executed by a processor, the machine-executable instructions cause the processor to implement the aforementioned music push method.
[0008] The embodiments of the present invention bring the following beneficial effects: The aforementioned music push method, device, and electronic device acquire multiple candidate music tracks; each candidate music track has corresponding text content; the text content is generated using a text generation model based on relevant data of the corresponding candidate music track; the relevant data includes multi-dimensional tags for the music, which include metadata, lyric features, and various user feedback data; based on the target user's preference parameters for the multiple candidate music tracks, a target music track is determined from the multiple candidate music tracks; the preference parameters are determined based on the target user's historical behavior data regarding the candidate music tracks; and the target music track and its corresponding target text content are pushed to the target user. This method determines the target music track based on the target user's preference parameters for the candidate music tracks and sends the target music track and its corresponding text to the target user. This text is generated using artificial intelligence technology based on the multi-dimensional tags of the music track, ensuring it is relevant to the target music track and logically coherent, making it easier for users to be attracted by the text and thus generating interest in the target music track, thereby improving the user experience.
[0009] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.
[0010] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0011] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0012] Figure 1A flowchart of a music push method provided in an embodiment of the present invention; Figure 2 A system architecture diagram for implementing a music push method is provided in an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the interaction between a music platform system, data layer, module processing, AIGC module, user and feedback module, as provided in an embodiment of the present invention. Figure 4 This is a schematic diagram of the structure of a music streaming device provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0014] Music platforms typically generate push notifications in two ways: one is through manual generation by operations staff, which are then sent to users in batches; the other is by pre-setting matching rules, each corresponding to a set of text templates. During delivery, these templates are matched with user and song information, and placeholders in the text templates are replaced before delivery. Specifically, when generating text based on templates, push notifications can be generated through simple keyword replacement. The system pre-sets fixed sentence structures, such as templates like "[Singer Name]'s [Song Name] is so good, come listen!", which can then be tailored to specific song information.
[0015] The above method has the following disadvantages: 1. Severe content homogenization: Push notifications received by different users are highly similar, lacking personalized features. The texts generated by operations teams are often generic and fail to reflect individual user differences and preferences.
[0016] 2. Unstable template generation quality: Template-based text generation has obvious technical limitations: when the template does not match the actual song information well, problems such as logical confusion and style inconsistency are likely to occur; the text after template replacement lacks emotional warmth, the language is stiff, and the user experience is poor.
[0017] Based on this, the present invention provides a music push method, device and electronic device, which can be applied to scenarios of version code comparison.
[0018] See Figure 1 First, a music push method provided by an embodiment of the present invention will be introduced. This method includes the following steps: Step S102: Obtain multiple candidate music tracks; each candidate music track has corresponding text content; the text content is generated by a text generation model based on the relevant data of the corresponding candidate music track; the relevant data includes multi-dimensional tags of the music track, which include various elements such as the music track's metadata, lyrics features, and user feedback data.
[0019] The candidate music can be vocal or instrumental works, including both sung songs and purely instrumental pieces. The candidate music can be audio recorded by a recording device or computer-generated audio. Specific settings can be configured according to needs and are not limited here.
[0020] Music platforms typically include multiple tracks. Several popular tracks can be selected as candidate tracks. In practice, tracks with high play counts, high collection counts, or high comment counts can be chosen as candidates. Furthermore, pre-defined numerical criteria for play counts, collection counts, and comment counts can be used to select tracks that meet these criteria as candidates. For example, tracks with over 10,000 play counts, over 10,000 collection counts, and over 1,000 comment counts can be selected as candidates.
[0021] Candidate music can be updated at a certain frequency based on data such as play counts, favorites, and comments on music on a specific music platform. For example, the cumulative play counts, favorites, and comments for each song on the platform can be calculated monthly. Furthermore, music-related data can be standardized based on the release date to reduce the impact of release time.
[0022] After identifying candidate music tracks, corresponding text content, also known as "push notification copy," needs to be generated for each track. This can begin by acquiring data about the candidate music from the music platform, such as metadata like name, language, and lyrics, as well as user feedback data. When the candidate music is instrumental, it does not have lyrics. User feedback data can include user comments, likes, and favorites related to the music.
[0023] Multi-dimensional tags for candidate music can typically be generated based on the data described above. These tags usually include tags corresponding to multiple data points from the aforementioned dataset. Some tags can be directly derived from the data itself; for example, a tag representing the name of a candidate song can be represented by the characters corresponding to that name. Other tags can be obtained through analysis of the data. For instance, semantic analysis of lyrics can be performed, and the analysis results can be used as tags representing the lyrical features of candidate songs. The number of likes from user feedback data can also be used as tags representing the popularity of each region within the candidate song. Specific settings can be configured according to requirements and are not limited here.
[0024] After determining the multi-dimensional labels of candidate songs, these labels can be input into the text generation model. Text generation models are typically implemented using artificial intelligence techniques. In practical applications, models based on large language models (LLMs) can be used. A large language model (LLM) is a deep learning model trained on a large amount of text data, enabling it to generate natural language text or understand the meaning of language text. These models can provide in-depth knowledge and language production on various topics by training on massive datasets. Common large language models include ChatGPT, Wenxin Yiyan, Tongyi Qianwen, Newbing, and Bard. The specific large language model selected to implement the target model depends on the specific needs and is not limited here.
[0025] Step S104: Based on the target user's preference parameters for multiple candidate music, determine the target music from multiple candidate music; the preference parameters are determined based on the target user's historical behavior data regarding the candidate music.
[0026] The aforementioned target users can be one or more. A target user's behavior towards candidate music typically includes playing, saving, commenting, and sharing. Historical behavior data is used for these behaviors. Based on historical behavior data, parameters such as the number of times a target user played, the duration of playback, the number of times they saved, the number of times they commented, the number of words in their comments, and the number of times they shared the candidate music can be determined. Based on these parameters, the target user's preference parameters for the candidate music can be calculated. Generally speaking, as the values of these parameters increase, the higher the value of the target user's preference parameters for the candidate music, indicating a higher degree of liking for the candidate music.
[0027] It can update the target user's preference parameters for music on the music platform at a preset frequency, such as weekly. It can also determine the target user's preference parameters for candidate music before pushing music to them. Specific settings can be configured according to needs and are not limited here.
[0028] After determining the target user's preference parameters for each candidate music, the multiple candidate music tracks can be sorted, and the target user's most favorite candidate music can be selected as the target music, or multiple candidate music tracks that the target user likes can be selected as the target music. Specific settings can be configured according to needs and are not limited here.
[0029] Step S106: Push the target music and the target text content corresponding to the target music to the target user.
[0030] Once the target music is identified, the target music and its corresponding text content can be pushed to the target user. In practice, this can be achieved by using a long-lived link between the music platform's server and the target user's device to push the target music and text to the device's notification bar. Alternatively, it can be displayed in the graphical user interface (GUI) after the target user triggers a specified icon. No restrictions are imposed on this approach.
[0031] The aforementioned music recommendation method involves acquiring multiple candidate music tracks; each candidate music track has corresponding text content; the text content is generated using a text generation model based on relevant data of the corresponding candidate music; the relevant data includes multi-dimensional tags for the music, which include metadata, lyric features, and various user feedback data; based on the target user's preference parameters for the multiple candidate music tracks, the target music track is determined from the multiple candidate music tracks; the preference parameters are determined based on the target user's historical behavior data regarding the candidate music tracks; and the target music track and its corresponding target text content are pushed to the target user. This method determines the target music track based on the target user's preference parameters for the candidate music tracks and sends the target music track and its corresponding text to the target user. This text is generated using artificial intelligence technology based on the music track's multi-dimensional tags, ensuring it aligns with the target music track, is logically coherent, and makes it easier for users to be attracted by the text, thereby generating interest in the target music track and improving the user experience.
[0032] In one specific embodiment, the text content corresponding to the candidate music can be generated in the following manner.
[0033] First, we can obtain the data of candidate music stored on the servers of the music platform. Then, we need to generate multi-dimensional tags for the candidate music. We can construct multiple structured tags based on this data as multi-dimensional tags. We can also combine music comments to perform sentiment analysis, calculating the proportion of positive and negative comments, and thus determining the sentiment tags for the candidate music. For example, when there are more positive comments, the sentiment tags for the candidate music can be set to positive, inspirational, etc.; when there are more negative comments, the sentiment tags can be set to negative, sad, depressing, etc. Multi-dimensional tags can be represented as vectors; in this case, multi-dimensional tags can also be called multi-dimensional feature vectors, forming a complete music profile. For example, the tags for a song could be: pop style, Chinese language, emotional healing, inspirational. Furthermore, the multi-dimensional tags for the candidate music can be input into a text generation model, which then outputs the corresponding text content for the candidate music.
[0034] Text generation models typically include text reference information, usually input as prompts. This text reference information serves to indicate either text content features or text structure features, or both. Text content features can indicate one or more of the following parameters: keywords, sentiment expression, text length, and text style. Text structure features can indicate the position of tags for a specified category within the text content. For example, text reference information might include: a music title appearing at the end of the text, a metaphorical interview, a length of less than 30 words, and subtle expression.
[0035] The music on the music platform can be pre-categorized into multiple different music genres. Then, corresponding text reference information can be set for each music genre. For candidate music for which corresponding text content needs to be generated, the text reference information corresponding to the music genre to which the candidate music belongs can be obtained.
[0036] Furthermore, the multi-dimensional labels and text reference information corresponding to the candidate music can be input into the text generation model, so that the text generation model can generate the text content corresponding to the candidate music based on the multi-dimensional labels and text reference information corresponding to the candidate music.
[0037] The aforementioned text reference information is determined based on historical text content. Text reference information is typically not used when initially generating text, but it can also be generated manually. Once the target text content corresponding to the target music is sent to the target user, it can also be considered historical text content. Using the example of target text content being historical text content, we can explain how text reference information is determined.
[0038] After sending the target music and target text content to the target user, behavioral data about the target user's interaction with the target music can be obtained within a specified time period; this specified time period can be 24 hours or other durations, without restriction. The behavioral data typically records information related to actions such as playback, favorites, comments, and reposts. Based on this behavioral data, behavioral parameters of the target user's interaction with the target music are then determined. These parameters may include one or more of the following: number of plays, playback duration, favorite status, number of comments, and number of reposts. Finally, the text generation model is updated based on these behavioral parameters.
[0039] Typically, certain conditions are set in advance, which can be called "preset conditions." These preset conditions may include one or more of the following: the number of plays is greater than the first time; the playback duration is greater than a specified duration; the saved status is "saved"; the number of comments is greater than the second time; and the number of shares is greater than the third time. If the behavioral parameters meet the preset conditions, it indicates that the target text content has a certain appeal to users, and the text reference information can be updated based on the target text content.
[0040] If the behavioral parameters meet the specified conditions, the target text content can be analyzed and processed to determine its text content features and text structure features. When the text reference information differs for different music genres, the text reference information for the corresponding music genre can be obtained, and then updated based on the target text content features and text structure features.
[0041] In one specific embodiment, the target user's preference parameters can be determined in the following manner.
[0042] First, it's necessary to obtain the target user's historical behavior data regarding the candidate music. This historical behavior data typically includes: playback data, favorites status data, and comment data. Playback data usually records the number of times the target user played the candidate music and the playback duration. Favorites status data records whether the target user has favorited the candidate music. Comment data records the number of comments the target user made regarding the candidate music and the content of those comments. If the target user has not played any candidate music, playback data is not included in the historical behavior data; similarly, if the target user has not commented on any candidate music, comment data is not included.
[0043] Furthermore, based on playback data, playback parameters for candidate music by the target user can be determined. These parameters may include the number of plays, playback duration, or both. When no playback data exists, the playback parameter can be directly assigned a value of 0. Based on favorites data, favorites status parameters for candidate music by the target user can be determined. Favorites status parameters typically have only two values: one indicating that the candidate music is favorited (set to 1), and the other indicating that it is not favorited (set to 0). Based on comment data, comment parameters for candidate music by the target user can be determined. Comment parameters may include the number of comments, comment length, etc. Furthermore, based on playback parameters, favorites status parameters, and comment parameters, preference parameters for candidate music by the target user can be determined.
[0044] In one specific embodiment, the playback parameter is the number of plays, and the comment parameter is the number of comments. When calculating the target user's preference parameters, it is also necessary to consider the target user's behavioral habits. Some users are accustomed to commenting on songs multiple times, while others comment very infrequently. The maximum historical play count for the same song by the target user can be predetermined. Based on this maximum historical play count, the range of the target user's music play count can be determined as (0 to the maximum historical play count). Similarly, the maximum number of comments the target user can make on the same song can be determined, and the range of the target user's music comment count can be determined as (0 to the maximum number of comments).
[0045] Furthermore, based on the number of plays and the maximum historical number of plays, a standardized playback count is calculated. In practice, the ratio of the number of plays to the maximum historical number of plays can be used as the standardized playback count. Since users are unsure whether they will like a song before playing it, playing a song once does not reflect their level of liking. Therefore, the playback count can be subtracted by one, and the ratio of this subtracted playback count to the maximum historical number of plays can be used as the standardized playback count. Similarly, based on the number of comments and the maximum number of comments, a standardized comment count is calculated. Specifically, the ratio of the number of comments to the maximum number of comments can be used as the standardized comment count.
[0046] Finally, based on the standardized values of play counts, favorite status parameters, and comment counts, along with preset weight parameters, the target user's preference parameters for candidate music are calculated. This can be expressed using the following formula: Preference parameter = α × normalized value of play count + β × normalized value of collection rate + γ × normalized value of comment interaction.
[0047] Here, α, β and γ are weight parameters, which take values in the range (0,1) and add up to 1.
[0048] Music platforms typically have a large number of users, so the preference parameters of multiple users can be stored in the form of a matrix. An example of a user-preference matrix is shown below: Songs | Song A | Song B | Song C | Song D | Song E --------|-------|-------|-------|-------|------- User 1 | 0.92 | 0.75 | 0.00 | 0.31 | 0.88 User 2 | 0.15 | 0.00 | 0.95 | 0.72 | 0.24 User 3 | 0.68 | 0.89 | 0.45 | 0.00 | 0.62 In this matrix, rows represent users and columns represent music. The matrix element values are the user's preference parameters for the corresponding music (range 0-1), with larger values indicating a higher degree of user liking for the song. Only the top N songs by number of favorites are selected, not the entire music library.
[0049] The following embodiments provide a specific method for determining comments on target music from multiple candidate music based on the target user's preference parameters for multiple candidate music.
[0050] In practical applications, to prevent the same music from being pushed to users repeatedly within a short period, resulting in a poor user experience, a music push frequency control strategy can be adopted. This strategy ensures that the same music is not pushed to the same user multiple times within a preset time period. In practice, the target music needs to be determined from multiple candidate music options based on the target user's preference parameters and the push frequency control strategy.
[0051] To implement the push notification frequency control strategy, it's necessary to determine which music to be removed from a pool of candidate music based on the target user's historical music push data. This historical data includes information such as the music pushed to the target user and the push time. From this data, music to be pushed to the target user within a preset duration threshold can be determined. This duration threshold can be one week or one month, etc., and can be set according to needs. When these songs are included in the candidate music list, they can be identified as songs to be removed. That is, the time difference between the time the song to be removed was pushed to the target user and the current time is less than or equal to the duration threshold.
[0052] The process involves eliminating music from multiple candidate tracks and determining the target music from the remaining tracks based on the target user's preference parameters. In practice, the target music is typically selected from the remaining tracks based on the target user's preference for it.
[0053] The following embodiments also provide the application process of the music push method in a real-world scenario.
[0054] This method is achieved through, for example Figure 2 The system architecture shown is implemented as follows. This architecture comprises a combined architecture of a song tag construction module, a user behavior analysis module, an Artificial Intelligence Generated Content (AIGC) text generation module, an intelligent push decision module, and a feedback collection and optimization module. The interaction process between the music platform's system, data layer, module processing, AIGC module, and user and feedback modules is as follows: Figure 3 As shown.
[0055] The song tag building module is used to build a structured tag system based on song title, genre, language, and lyric keywords.
[0056] The user behavior analysis module is used to collect user behavior data on songs: play count, collection rate, number of comments, calculate user ratings of their liking for songs, and build a user-song preference matrix.
[0057] The AIGC text generation module generates multiple personalized push notification texts based on a large language model, combined with song tags and user preferences. Each text highlights different song features and emotional attributes, and AI can supplement basic song lyrics and music information, as well as popular comment data. This ensures the texts are logically coherent, stylistically consistent, and emotionally resonant. The intelligent push decision module sorts candidate songs daily based on user preferences; combines this with a frequency control strategy based on song dimensions to avoid over-pushing similar content; and selects the most suitable song and corresponding optimal text for the day. The feedback collection and optimization module is used to collect real-time performance data of push texts: click-through rate, effective play rate, and conversion rate; and to quantify and score text quality based on performance data, extract feature vectors of high-quality texts, and optimize subsequent generation strategies. Specifically, it can collect texts with performance data exceeding internal evaluation standards, extract their text content features, structural features, and user preference features, and improve and supplement the input for AIGC and push decisions.
[0058] For the above method embodiments, see Figure 4 The illustrated music streaming device includes: The music acquisition module 402 is used to acquire multiple candidate music tracks; each candidate music track has corresponding text content; the text content is generated by a text generation model based on the relevant data of the corresponding candidate music track; the relevant data includes multi-dimensional tags of the music track, which include the music track's metadata, lyric features, and various user feedback data. The target music determination module 404 is used to determine the target music from multiple candidate music based on the target user's preference parameters for multiple candidate music; the preference parameters are determined based on the target user's historical behavior data for candidate music. The music push module 406 is used to push target music and the target text content corresponding to the target music to the target user.
[0059] The aforementioned music recommendation device acquires multiple candidate music tracks; each candidate music track has corresponding text content; the text content is generated using a text generation model based on relevant data of the corresponding candidate music track; the relevant data includes multi-dimensional tags for the music, which include metadata, lyric features, and various user feedback data; based on the target user's preference parameters for the multiple candidate music tracks, the target music track is determined from the multiple candidate music tracks; the preference parameters are determined based on the target user's historical behavior data regarding the candidate music tracks; and the target music track and its corresponding target text content are pushed to the target user. This method determines the target music track based on the target user's preference parameters for the candidate music tracks and sends the target music track and its corresponding text to the target user. This text is generated using artificial intelligence technology based on the multi-dimensional tags of the music track, making it relevant to the target music track, logically coherent, and more easily attracting the user's attention, thereby generating interest in the target music track and improving the user experience.
[0060] The aforementioned device includes a text content determination module, used for: generating multi-dimensional labels corresponding to candidate music; inputting the multi-dimensional labels corresponding to candidate music into a text generation model; and outputting the text content corresponding to candidate music through the text generation model.
[0061] The aforementioned text generation model includes text reference information; the text reference information is used to indicate: text content features and / or text structure features; the text content features are used to indicate at least one of the following: keywords, sentiment expression, text length, and text style; the text structure features are used to indicate: the position of a label of a specified category in the text content; the text reference information is determined based on historical text content; the aforementioned device includes a text content determination module, used to: input the multi-dimensional labels and text reference information corresponding to the candidate music into the text generation model.
[0062] The aforementioned device further includes a preference parameter determination module, used to: acquire historical behavioral data of the target user for candidate music; the historical behavioral data includes: playback data, collection status data, and comment data; determine the playback parameters of the target user for candidate music based on the playback data; determine the collection status parameters of the target user for candidate music based on the collection status data; determine the comment parameters of the target user for candidate music based on the comment data; and determine the preference parameters of the target user for candidate music based on the playback parameters, collection status parameters, and comment parameters.
[0063] The playback parameters mentioned above include the number of plays; the comment parameters include the number of comments; the preference parameter determination module mentioned above is also used to: calculate the standardized value of the playback count based on the number of plays and the maximum historical number of plays by the target user for the same song; calculate the standardized value of the number of comments based on the number of comments and the maximum number of comments by the target user for the same song; and calculate the target user's preference parameters for candidate music based on the standardized value of the playback count, the collection status parameter, the standardized value of the number of comments, and the preset weight parameters.
[0064] The aforementioned target music determination module is also used for: determining music to be removed from multiple candidate music based on the historical music push data corresponding to the target user; the time difference between the time the music to be removed is pushed to the target user and the current time is less than or equal to a preset duration threshold; removing the music to be removed from multiple candidate music; and determining the target music from the removed candidate music based on the target user's preference parameters for multiple candidate music.
[0065] The aforementioned device further includes: a behavior data acquisition module for acquiring behavior data of the target user for the target music within a specified time period; a model update module for determining the behavior parameters of the target user for the target music based on the behavior data; and updating the text generation model based on the behavior parameters.
[0066] The aforementioned text generation model includes text reference information; the text reference information and labels generated based on relevant data of candidate music are input into the text generation model so that the text generation model generates text content corresponding to the candidate music; the model update module is also used to: update the text reference information based on the target text content if the behavior parameters meet preset conditions.
[0067] The aforementioned model update module is also used to: analyze and process the target text content to determine the text content features and text structure features of the target text content; the text content features are used to indicate at least one of the following: keywords, sentiment expression, text length, and text style; the text structure features are used to indicate: the position of the tags of the specified category in the target text content; and update the text reference information based on the text content features and text structure features of the target text content.
[0068] This embodiment also provides an electronic device, including a processor and a memory. The memory stores machine-executable instructions that can be executed by the processor. The processor executes the machine-executable instructions to implement the above-mentioned music push method, for example: Multiple candidate music tracks are acquired; each candidate music track has corresponding text content; the text content is generated using a text generation model based on the relevant data of the corresponding candidate music track; the relevant data includes multi-dimensional tags for the music track, which include metadata, lyric features, and various user feedback data; based on the target user's preference parameters for the multiple candidate music tracks, the target music track is determined from the multiple candidate music tracks; the preference parameters are determined based on the target user's historical behavior data regarding the candidate music tracks; the target music track and its corresponding target text content are pushed to the target user.
[0069] The above method determines the target music based on the target user's preference parameters for candidate music, and sends the target music and its corresponding text to the target user. This text is generated by artificial intelligence technology based on the multi-dimensional tags of the music, which is consistent with the target music and logically coherent, making it easier for users to be attracted by the text, thereby generating interest in the target music and improving the user experience.
[0070] Optionally, the above text content is determined by: generating multi-dimensional labels corresponding to candidate music; inputting the multi-dimensional labels corresponding to candidate music into the text generation model, and outputting the text content corresponding to candidate music through the text generation model.
[0071] Optionally, the above text generation model includes text reference information; the text reference information is used to prompt: text content features and / or text structure features; the text content features are used to indicate at least one of the following: keywords, sentiment expression, text length, and text style; the text structure features are used to indicate: the position of the label of the specified category in the text content; the text reference information is determined based on historical text content; the step of inputting the multi-dimensional labels corresponding to the candidate music into the text generation model includes: inputting the multi-dimensional labels corresponding to the candidate music and the text reference information into the text generation model.
[0072] Optionally, the above preference parameters are determined by the following: obtaining historical behavioral data of the target user for candidate music; historical behavioral data includes: playback data, collection status data, and comment data; determining the playback parameters of the target user for candidate music based on the playback data; determining the collection status parameters of the target user for candidate music based on the collection status data; determining the comment parameters of the target user for candidate music based on the comment data; and determining the preference parameters of the target user for candidate music based on the playback parameters, collection status parameters, and comment parameters.
[0073] Optionally, the playback parameters mentioned above include the number of plays; the comment parameters include the number of comments; the step of determining the target user's preference parameters for candidate music based on the playback parameters, favorite status parameters, and comment parameters includes: calculating a standardized value of the number of plays based on the number of plays and the target user's maximum historical number of plays for the same song; calculating a standardized value of the number of comments based on the number of comments and the target user's maximum number of comments for the same song; and calculating the target user's preference parameters for candidate music based on the standardized value of the number of plays, favorite status parameters, standardized value of the number of comments, and preset weight parameters.
[0074] Optionally, the above step of determining the target music from multiple candidate music based on the target user's preference parameters for multiple candidate music includes: determining the music to be removed from multiple candidate music based on the target user's historical music push data; the time difference between the time the music to be removed was pushed to the target user and the current time is less than or equal to a preset duration threshold; removing the music to be removed from multiple candidate music; and determining the target music from the removed candidate music based on the target user's preference parameters for multiple candidate music.
[0075] Optionally, after pushing the target music and the corresponding target text content to the target user, the method further includes: obtaining the target user's behavioral data on the target music within a specified time period; determining the target user's behavioral parameters on the target music based on the behavioral data; and updating the text generation model based on the behavioral parameters.
[0076] Optionally, the above text generation model includes text reference information; the text reference information and labels generated based on the relevant data of the candidate music are input into the text generation model so that the text generation model generates text content corresponding to the candidate music; the step of updating the text generation model based on behavioral parameters includes: if the behavioral parameters meet preset conditions, updating the text reference information based on the target text content.
[0077] Optionally, the above steps for updating text reference information based on target text content include: analyzing and processing the target text content to determine the text content features and text structure features of the target text content; the text content features are used to indicate at least one of the following: keywords, sentiment expression, text length, and text style; the text structure features are used to indicate: the position of a label of a specified category in the target text content; and updating the text reference information based on the text content features and text structure features of the target text content.
[0078] See Figure 5 As shown, the electronic device includes a processor 100 and a memory 101. The memory 101 stores machine-executable instructions that can be executed by the processor 100. The processor 100 executes the machine-executable instructions to implement the music push method described above.
[0079] Furthermore, Figure 5 The electronic device shown also includes a bus 102 and a communication interface 103, with the processor 100, the communication interface 103 and the memory 101 connected via the bus 102.
[0080] The memory 101 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 103 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc. The bus 102 may be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 5 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0081] Processor 100 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 100 or by instructions in software form. The processor 100 may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams of the invention in the embodiments of this invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method invented in conjunction with the embodiments of this invention can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software module can reside in a readily available storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 101, and the processor 100 reads the information from memory 101 and, in conjunction with its hardware, completes the steps of the method described in the foregoing embodiments.
[0082] This embodiment also provides a machine-readable storage medium storing machine-executable instructions. When the machine-executable instructions are called and executed by the processor, the machine-executable instructions cause the processor to implement the above-described music push method.
[0083] The music streaming method, apparatus, and electronic device provided in this invention include a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the preceding method embodiments, for example: Multiple candidate music tracks are acquired; each candidate music track has corresponding text content; the text content is generated using a text generation model based on the relevant data of the corresponding candidate music track; the relevant data includes multi-dimensional tags for the music track, which include metadata, lyric features, and various user feedback data; based on the target user's preference parameters for the multiple candidate music tracks, the target music track is determined from the multiple candidate music tracks; the preference parameters are determined based on the target user's historical behavior data regarding the candidate music tracks; the target music track and its corresponding target text content are pushed to the target user.
[0084] The above method determines the target music based on the target user's preference parameters for candidate music, and sends the target music and its corresponding text to the target user. This text is generated by artificial intelligence technology based on the multi-dimensional tags of the music, which is consistent with the target music and logically coherent, making it easier for users to be attracted by the text, thereby generating interest in the target music and improving the user experience.
[0085] Optionally, the above text content is determined by: generating multi-dimensional labels corresponding to candidate music; inputting the multi-dimensional labels corresponding to candidate music into the text generation model, and outputting the text content corresponding to candidate music through the text generation model.
[0086] Optionally, the above-mentioned text generation model includes text reference information; the text reference information is used to prompt: text content features and / or text structure features; the text content features are used to indicate at least one of the following: keywords, sentiment expression, text length, and text style; the text structure features are used to indicate: the position of the label of the specified category in the text content; the text reference information is determined based on historical text content; the step of inputting the multi-dimensional labels corresponding to the candidate music into the text generation model includes: inputting the multi-dimensional labels corresponding to the candidate music and the text reference information into the text generation model.
[0087] Optionally, the above preference parameters are determined by the following: obtaining historical behavioral data of the target user for candidate music; historical behavioral data includes: playback data, collection status data, and comment data; determining the playback parameters of the target user for candidate music based on the playback data; determining the collection status parameters of the target user for candidate music based on the collection status data; determining the comment parameters of the target user for candidate music based on the comment data; and determining the preference parameters of the target user for candidate music based on the playback parameters, collection status parameters, and comment parameters.
[0088] Optionally, the playback parameters mentioned above include the number of plays; the comment parameters include the number of comments; the step of determining the target user's preference parameters for candidate music based on the playback parameters, favorite status parameters, and comment parameters includes: calculating a standardized value of the number of plays based on the number of plays and the target user's maximum historical number of plays for the same song; calculating a standardized value of the number of comments based on the number of comments and the target user's maximum number of comments for the same song; and calculating the target user's preference parameters for candidate music based on the standardized value of the number of plays, favorite status parameters, standardized value of the number of comments, and preset weight parameters.
[0089] Optionally, the above step of determining the target music from multiple candidate music based on the target user's preference parameters for multiple candidate music includes: determining the music to be removed from multiple candidate music based on the target user's historical music push data; the time difference between the time the music to be removed was pushed to the target user and the current time is less than or equal to a preset duration threshold; removing the music to be removed from multiple candidate music; and determining the target music from the removed candidate music based on the target user's preference parameters for multiple candidate music.
[0090] Optionally, after pushing the target music and the corresponding target text content to the target user, the method further includes: obtaining the target user's behavioral data on the target music within a specified time period; determining the target user's behavioral parameters on the target music based on the behavioral data; and updating the text generation model based on the behavioral parameters.
[0091] Optionally, the above text generation model includes text reference information; the text reference information and labels generated based on the relevant data of the candidate music are input into the text generation model so that the text generation model generates text content corresponding to the candidate music; the step of updating the text generation model based on behavioral parameters includes: if the behavioral parameters meet preset conditions, updating the text reference information based on the target text content.
[0092] Optionally, the above steps for updating text reference information based on target text content include: analyzing and processing the target text content to determine the text content features and text structure features of the target text content; the text content features are used to indicate at least one of the following: keywords, sentiment expression, text length, and text style; the text structure features are used to indicate: the position of a label of a specified category in the target text content; and updating the text reference information based on the text content features and text structure features of the target text content.
[0093] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system and apparatus described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0094] Furthermore, in the description of the embodiments of the present invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in the present invention based on the specific circumstances.
[0095] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0096] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0097] Finally, it should be noted that the above embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A music push method, characterized in that, The method includes: Multiple candidate music tracks are obtained; each candidate music track has corresponding text content; the text content is generated by a text generation model based on the relevant data of the corresponding candidate music track; the relevant data includes multi-dimensional tags of the music track, which include various elements such as music metadata, lyric features, and user feedback data. The target music is determined from the multiple candidate music options based on the target user's preference parameters for the candidate music options; the preference parameters are determined based on the target user's historical behavior data regarding the candidate music options. The target music and its corresponding target text content are pushed to the target user.
2. The method according to claim 1, characterized in that, The text content is determined in the following way: Generate multi-dimensional tags for the candidate music; The multi-dimensional labels corresponding to the candidate music are input into the text generation model, and the text generation model outputs the text content corresponding to the candidate music.
3. The method according to claim 2, characterized in that, The text generation model includes text reference information; the text reference information is used to indicate: text content features and / or text structure features; the text content features are used to indicate at least one of the following: keywords, sentiment expression, text length, and text style; the text structure features are used to indicate: the position of a tag of a specified category in the text content; The text reference information is determined based on historical text content; The step of inputting the multi-dimensional labels corresponding to the candidate music into the text generation model includes: The multi-dimensional labels corresponding to the candidate music and the text reference information are input into the text generation model.
4. The method according to claim 1, characterized in that, The step of determining the target music from the multiple candidate music sources based on the target user's preference parameters includes: Based on the historical music push data corresponding to the target user, music to be removed is determined from the multiple candidate music; the time difference between the time when the music to be removed is pushed to the target user and the current time is less than or equal to a preset duration threshold. Remove the music to be removed from the plurality of candidate music; Based on the target user's preference parameters for the multiple candidate music tracks, the target music is determined from the eliminated candidate music tracks.
5. The method according to claim 1, characterized in that, After pushing the target music and its corresponding target text content to the target user, the method further includes: Obtain the target user's behavioral data regarding the target music within a specified time period; Based on the behavioral data, determine the behavioral parameters of the target user in relation to the target music; The text generation model is updated based on the behavioral parameters.
6. The method according to claim 5, characterized in that, The text generation model includes text reference information; the text reference information and tags generated based on the relevant data of the candidate music are input into the text generation model so that the text generation model generates text content corresponding to the candidate music. The step of updating the text generation model based on the behavioral parameters includes: If the behavior parameters meet the preset conditions, the text reference information is updated based on the target text content.
7. The method according to claim 6, characterized in that, The steps of updating text reference information based on the target text content include: The target text content is analyzed and processed to determine the text content features and text structure features of the target text content; The text reference information is updated based on the text content features and text structure features of the target text content.
8. A music streaming device, characterized in that, The device includes: The music acquisition module is used to acquire multiple candidate music tracks; each candidate music track has corresponding text content; the text content is generated by a text generation model based on the relevant data of the corresponding candidate music track; the relevant data includes multi-dimensional tags for the music track, which include various elements such as music metadata, lyric features, and user feedback data. The target music determination module is used to determine the target music from the multiple candidate music based on the target user's preference parameters for the multiple candidate music; the preference parameters are determined based on the target user's historical behavior data regarding the candidate music; The music push module is used to push the target music and the target text content corresponding to the target music to the target user.
9. An electronic device, characterized in that, The device includes a processor and a memory, the memory storing machine-executable instructions that can be executed by the processor, the processor executing the machine-executable instructions to implement the music push method according to any one of claims 1-7.
10. A machine-readable storage medium, characterized in that, The machine-readable storage medium stores machine-executable instructions, which, when invoked and executed by a processor, cause the processor to implement the music push method according to any one of claims 1-7.