A voice recommendation method, device, computer device, and storage medium
The method enhances voice recommendation accuracy by calculating preference weights based on playback and feedback behaviors with time decay and TF-IDF values, addressing the Matthew effect and cold-start issues in voice recommendation systems.
Patent Information
- Application Number
- CN202210179939.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-25
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-02-25
AI Technical Summary
The existing sound recommendation methods have serious Matthew effect, long-tail sound embedding effect is not ideal, and the sound cold start problem is serious. Traditional methods cannot effectively recommend new sounds.
By obtaining the historical behavior data and exposure data of the target object, combining the time attenuation coefficient and TF-IDF value, calculate the playback and positive feedback behavior weights of the anchor, determine the preferred anchor collection, and filter real-time negative feedback anchors to recall the voices of interest.
It improves the accuracy of sound recommendations and effectively solves the Matthew effect and sound cold start problems. The recommended anchors and sounds are more in line with user interests.
Smart Images

Figure CN114579797B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and specifically relates to a voice recommendation method, device, computer device, and storage medium. Background Art
[0002] With the continuous development of computer technology, online audio platforms have also been increasingly widely used, and users can listen to various voices through online audio platforms. Online audio platforms will recommend voices to users. Traditional voice recommendation methods include index-based recommendation, item2vec, and graph-based.
[0003] Among them, for the index-based recommendation method, it collects user label preferences through user behavior, and then recommends voices containing that label.
[0004] Disadvantages:
[0005] Considering that the granularity is too coarse, only the user's preference for labels is considered, without considering the user's preferences in finer dimensions.
[0006] For item2vec, it regards the user's behavior sequence as a sentence, the voices in the sequence as words, and then uses the word2vec model for training to obtain the embedding of the voice.
[0007] Disadvantages:
[0008] The Matthew effect is serious. Users have relatively few behaviors on long-tail voices, and the embedding effect of these long-tail voices is not ideal.
[0009] The voice cold start problem is serious. New voices have no user behaviors yet, so new voices will not appear in the user's behavior sequence, and thus the embedding of new voices cannot be obtained.
[0010] For graph-based, it constructs a graph based on the user's behavior on voices, and then starts from each node and performs random walks according to weights, so as to form new voice sequences. Finally, the word2vec model is continued to be used for these different voice sequences to obtain the embedding vector of the voice.
[0011] Disadvantages:
[0012] The random walk strategy will introduce relatively large noise.
[0013] The Matthew effect is serious. Users have relatively few behaviors on long-tail voices, and the embedding effect of these long-tail voices is not ideal.
[0014] The problem of voice cold start is serious. Since no user has generated behavior for the new voice yet, the new voice will not appear in the graph, and thus the embedding of the new voice cannot be obtained. Summary of the Invention
[0015] In view of the above problems, the present invention is proposed to provide a voice recommendation method, device, computer device, and storage medium that overcome the above problems or at least partially solve the above problems, improving the accuracy of voice recommendation.
[0016] In a first aspect, an embodiment of the present invention provides a voice recommendation method, which includes the following steps:
[0017] Obtain the historical behavior data of the target object playing voices within a set historical time period, the voice data exposed by the target object, the set of followed hosts, the set of unfollowed hosts, and the number of times the target object exposed to the hosts within the last N preset time periods in the historical time period; the historical behavior data includes play behavior data and positive feedback behavior data;
[0018] Based on the weight of the play behavior of the voices released by the host within each preset time period of the target object in the historical time period, the time decay coefficient, and the play behavior TF-IDF value, determine the play behavior weight of the target object for the corresponding host within the corresponding preset time period; based on the weight of the positive feedback behavior of the voices released by the host within each preset time period of the target object in the historical time period, the time decay coefficient, the positive feedback behavior TF-IDF value, and the number of positive feedback behaviors, determine the positive feedback behavior weight of the target object for the corresponding host within the corresponding preset time period; based on the play behavior weight of the target object for each host within each preset time period and the positive feedback behavior weight of the target object for each host within each preset time period, determine the preference weight of the target object for each host; based on the preference weight of the target object for each host, determine the set of preferred hosts of the target object;
[0019] Based on the preference weight of the target object for each host, the number of times the target object exposed to the hosts within the last N preset time periods, and the set of unfollowed hosts, determine the real-time negative feedback host set of the target object;
[0020] Filter the hosts in the real-time negative feedback host set from the set of preferred hosts of the target object, and filter the hosts in the real-time negative feedback host set from the set of followed hosts of the target object, and recall the voices released by the hosts according to the filtered set of preferred hosts and the set of followed hosts and recommend them to the target object.
[0021] Optionally, the weight of the playback behavior of the sound released by the host in each preset time period is determined based on the playback duration of the sound released by the corresponding host by the target object in the corresponding preset time period and the total exposure duration of the sound released by the corresponding host by the target object in the corresponding preset time period.
[0022] Optionally, the time decay coefficient in each preset time period is determined based on the ranking of the corresponding preset time period within the historical time period.
[0023] Optionally, the TF-IDF value of the playback behavior of the sound released by the host in each preset time period is determined based on the number of times the target object plays the sound released by the corresponding host in the corresponding preset time period, the number of times the target object plays the sounds released by all hosts in the corresponding preset time period, the number of times all users play the sounds released by all hosts in the corresponding preset time period, and the number of times all users play the sound released by the corresponding host in the corresponding preset time period.
[0024] Optionally, the TF-IDF value of the positive feedback behavior of the sound released by the host in each preset time period is determined based on the number of positive feedback behaviors of the target object on the sound released by the corresponding host in the corresponding preset time period, the number of positive feedback behaviors of the target object on the sounds released by all hosts in the corresponding preset time period, the number of positive feedback behaviors of all users on the sounds released by all hosts in the corresponding preset time period, and the number of positive feedback behaviors of all users on the sound released by the corresponding host in the corresponding preset time period.
[0025] Optionally, the types of the positive feedback behaviors include a like behavior and a favorite behavior.
[0026] Optionally, the method further includes:
[0027] If the target object has no playback behavior, obtain the information of the host corresponding to the sound played by the target object, assign the preference weight of the target object for the corresponding host as a preset weight value, and use the corresponding host as the preferred host of the target object;
[0028] If the target object has no playback behavior and has not been played with a sound, calculate the preference weights of each host for the user group with the same basic user profile as the target object, determine the basic preference weights of each host based on the preference weights of the user group for each host, and use the top M hosts with the highest basic preference weights as the preferred host set of the target object, and use the basic preference weights of the corresponding hosts as the preference weights of the target object for the corresponding hosts.
[0029] In a second aspect, an embodiment of the present invention provides a sound recommendation device, and the device includes:
[0030] An acquisition module, configured to acquire historical behavior data of a target object playing sounds within a set historical time period, sound data to which the target object is exposed, a set of followed live streamers, a set of unfollowed live streamers, and the number of times the target object exposes to live streamers within the last N preset time periods within the historical time period; the historical behavior data includes playing behavior data and positive feedback behavior data;
[0031] A preferred live streamer set determination module, configured to determine the playing behavior weight of the target object for a corresponding live streamer within a corresponding preset time period based on the weight of the playing behavior of the target object for the sounds released by the live streamer within each preset time period within the historical time period, a time decay coefficient, and the TF-IDF value of the playing behavior; determine the positive feedback behavior weight of the target object for a corresponding live streamer within a corresponding preset time period based on the weight of the positive feedback behavior of the target object for the sounds released by the live streamer within each preset time period within the historical time period, a time decay coefficient, the TF-IDF value of the positive feedback behavior, and the number of positive feedback behaviors; determine the preference weight of the target object for each live streamer based on the playing behavior weight of the target object for each live streamer within each preset time period and the positive feedback behavior weight of the target object for each live streamer within each preset time period; determine the preferred live streamer set of the target object based on the preference weight of the target object for each live streamer;
[0032] A real-time negative feedback live streamer set determination module, configured to determine the real-time negative feedback live streamer set of the target object based on the preference weight of the target object for each live streamer, the number of times the target object exposes to live streamers within the last N preset time periods, and the set of unfollowed live streamers;
[0033] A recommendation module, configured to filter the live streamers in the real-time negative feedback live streamer set from the preferred live streamer set of the target object, and filter the live streamers in the real-time negative feedback live streamer set from the followed live streamer set of the target object, and recall the sounds released by the live streamers and recommend them to the target object according to the filtered preferred live streamer set and followed live streamer set.
[0034] Optionally, the weight of the playing behavior of the target object for the sounds released by the live streamer within each preset time period is determined based on the playing duration of the target object for the sounds released by the corresponding live streamer within the corresponding preset time period and the total duration of the target object exposing to the sounds released by the corresponding live streamer within the corresponding preset time period.
[0035] Optionally, the time decay coefficient within each preset time period is determined based on the sorting of the corresponding preset time period within the historical time period.
[0036] Optionally, the TF-IDF value of the playback behavior of the voice released by the host in each preset time period is determined based on the number of times the target object plays the voice released by the corresponding host in the corresponding preset time period, the number of times the target object plays the voices released by all hosts in the corresponding preset time period, the number of times all users play the voices released by all hosts in the corresponding preset time period, and the number of times all users play the voice released by the corresponding host in the corresponding preset time period.
[0037] Optionally, the TF-IDF value of the positive feedback behavior of the voice released by the host in each preset time period is determined based on the number of positive feedback behaviors of the target object for the voice released by the corresponding host in the corresponding preset time period, the number of positive feedback behaviors of the target object for the voices released by all hosts in the corresponding preset time period, the number of positive feedback behaviors of all users for the voices released by all hosts in the corresponding preset time period, and the number of positive feedback behaviors of all users for the voice released by the corresponding host in the corresponding preset time period.
[0038] Optionally, the types of the positive feedback behaviors include a like behavior and a favorite behavior.
[0039] Optionally, the preferred host set determination module is further configured to:
[0040] If the target object has no playback behavior, obtain the information of the host corresponding to the voice played by the target object, assign the preference weight of the target object for the corresponding host as a preset weight value, and use the corresponding host as the preferred host of the target object;
[0041] If the target object has no playback behavior and no played voice, calculate the preference weights of each host for the user group with the same basic user profile as the target object, determine the basic preference weights of each host based on the preference weights of the user group for each host, and use the top M hosts with the highest basic preference weights as the preferred host set of the target object, and use the basic preference weights of the corresponding hosts as the preference weights of the target object for the corresponding hosts.
[0042] In a third aspect, an embodiment of the present invention provides a computer device, where the computer device includes:
[0043] One or more processors;
[0044] A memory for storing one or more programs;
[0045] When the one or more programs are executed by the one or more processors, the one or more processors implement the voice recommendation method as described in any item of the first aspect.
[0046] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium.
[0047] A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements the voice recommendation method described in any one of the first aspects.
[0048] In this embodiment, historical behavior data of the target object playing voices, voice data to which the target object is exposed, a set of followed anchors, a set of unfollowed anchors, and the number of times the target object exposes to an anchor within the last N preset time periods in the historical time period are obtained; the historical behavior data includes playing behavior data and positive feedback behavior data; based on the weight of the playing behavior of the target object on the voices released by the anchor within each preset time period in the historical time period, the time decay coefficient, and the TF-IDF value of the playing behavior, determine the playing behavior weight of the target object on the corresponding anchor within the corresponding preset time period; based on the weight of the positive feedback behavior of the target object on the voices released by the anchor within each preset time period in the historical time period, the time decay coefficient, the TF-IDF value of the positive feedback behavior, and the number of positive feedback behaviors, determine the positive feedback behavior weight of the target object on the corresponding anchor within the corresponding preset time period; based on the playing behavior weight of the target object on each anchor within each preset time period and the positive feedback behavior weight of the target object on each anchor within each preset time period, determine the preference weight of the target object on each anchor; determine the set of preferred anchors of the target object based on the preference weight of the target object on each anchor; determine the real-time negative feedback anchor set of the target object based on the preference weight of the target object on each anchor, the number of times the target object exposes to an anchor within the last N preset time periods, and the set of unfollowed anchors; filter the anchors in the real-time negative feedback anchor set from the set of preferred anchors of the target object, and filter the anchors in the real-time negative feedback anchor set from the set of followed anchors of the target object, and recall the voices released by the anchors and recommend them to the target object according to the filtered set of preferred anchors and the set of followed anchors. Collect playing behavior data and positive feedback behavior data and combine Newton's laws of thermodynamics to determine the preference portrait of the target object, determine the set of preferred anchors according to the preference portrait of the target object, and then filter the set of preferred anchors and the set of followed anchors. The filtered set of preferred anchors and the set of followed anchors can effectively hit the anchors that the target object is interested in, greatly improving the accuracy of voice recommendation. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0050] Figure 1Flowchart of a voice recommendation method provided in Embodiment 1 of the present invention;
[0051] Figure 2 Structural schematic diagram of a voice recommendation device provided in Embodiment 2 of the present invention;
[0052] Figure 3 Structural schematic diagram of a computer device provided in Embodiment 3 of the present invention. Detailed implementation manners
[0053] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of description, only parts related to the present invention are shown in the accompanying drawings rather than all structures.
[0054] With the continuous development of computer technology, online audio platforms have also been increasingly widely used, and users can listen to various voices through online audio platforms. Online audio platforms will recommend voices to users. Traditional voice recommendation methods include index-based recommendation, item2vec, and graph-based.
[0055] Among them, for the index-based recommendation method, it collects user label preferences through user behavior, and then recommends voices containing the label.
[0056] Disadvantages:
[0057] Considering that the granularity is too coarse, only the user's preference for labels is considered, without considering the user's preferences in finer dimensions.
[0058] For item2vec, it regards the user's behavior sequence as a sentence, the voices in the sequence as words, and then trains with a word2vec model to obtain the embedding of the voices.
[0059] Disadvantages:
[0060] The Matthew effect is serious. Users have relatively few behaviors on long-tail voices, and the embedding effects of these long-tail voices are not ideal;
[0061] The voice cold start problem is serious. For new voices, there are no user behaviors yet, so new voices will not appear in the user's behavior sequence, and thus the embedding of new voices cannot be obtained.
[0062] Based on a graph, it constructs a graph based on the user's behavior towards sounds, and then starts from each node and performs a random walk according to the weights, so as to form new sound sequences. Finally, the word2vec model is continued to be used for these different sound sequences to obtain the embedding vectors of the sounds.
[0063] Disadvantages:
[0064] The random walk strategy will introduce relatively large noise;
[0065] The Matthew effect is serious. Users have relatively few behaviors towards long-tail sounds, and the embedding effects of these long-tail sounds are not ideal;
[0066] The cold start problem of sounds is serious. New sounds have no user behaviors yet, so new sounds will not appear in the graph, and thus the embeddings of new sounds cannot be obtained.
[0067] To overcome the above problems or at least partially solve the above problems, the embodiments of the present application provide a sound recommendation method, which can improve the accuracy of sound recommendation.
[0068] The following is a detailed description through embodiments.
[0069] Embodiment 1
[0070] Figure 1 It is a flowchart of a sound recommendation method provided by Embodiment 1 of the present invention. This method can be executed by a sound recommendation device, which can be implemented by software and / or hardware and can be configured in a computer device, such as a server, a personal computer, etc. The sound recommendation method specifically includes the following steps:
[0071] Step 101: Obtain the historical behavior data of the target object playing sounds within a set historical time period, the sound data exposed by the target object, the set of followed hosts, the set of unfollowed hosts, and the number of times the target object exposes the hosts within the last N preset time periods in the historical time period.
[0072] The historical behavior data includes play behavior data and positive feedback behavior data.
[0073] During the process of a user listening to sounds on an online audio platform, various behavior data will be generated, such as play behavior data and positive feedback behavior data, and sounds will also be exposed. Among them, positive feedback behavior refers to behaviors that can express the user's positive emotions such as recognition and preference for sounds, such as liking and collecting.
[0074] The set historical time period and preset time period can be adjusted according to the actual situation. The set historical time period can be the most recent month, the most recent two months, the most recent year, or other time periods, and this embodiment does not limit this. Similarly, the preset time period can be one hour as a cycle, one day as a cycle, one week as a cycle, or other time cycles, and this embodiment does not limit this.
[0075] For example, obtain the historical behavior data of the user's played sounds, the sound data exposed by the user, the followed anchors, the unfollowed anchors, and the number of times the user exposes the anchor in a recent small time period, such as in the most recent 3 days, within a recent period of time, such as within the most recent 180 days.
[0076] Step 102: Determine the preference weights of the target object for each anchor based on the play behavior data, positive feedback behavior data, and the sound data exposed by the target object of the target object playing sounds within the historical time period, and determine the set of preferred anchors of the target object based on the preference weights of the target object for each anchor.
[0077] Specifically, step 102 includes:
[0078] Sub-step 1021: Determine the play behavior weight of the target object for the corresponding anchor within the corresponding preset time period based on the weight of the play behavior of the target object for the sounds released by the anchor within each preset time period within the historical time period, the time decay coefficient, and the play behavior TF-IDF value.
[0079] For example, the play behavior weight of user userA for anchor njA on the t-th day
[0080] W(userA, njA, t, type = play)
[0081] = f(play_time, total_time) * g(t) * tf-idf(userA, njA, t, type = play)
[0082] Wherein, f(play_time, total_time) is the weight of the play behavior of user userA for the sounds released by anchor njA on the t-th day, g(t) is the time decay coefficient on the t-th day, and tf-idf(userA, njA, t, type = play) is the play behavior TF-IDF value of user userA for the sounds released by anchor njA on the t-th day.
[0083] In one implementation, the weight of the playback behavior of the voice released by the host in each preset time period is determined based on the playback duration of the voice released by the corresponding host by the target object in the corresponding preset time period and the total exposure duration of the voice released by the corresponding host by the target object in the corresponding preset time period.
[0084] For example, the total playback duration of the voice (program) of host njA played by user userA on the t-th day is play_time, and the total exposure duration of the voice (program) of host njA exposed to user userA on the t-th day is total_time.
[0085] The weight of the playback behavior of the voice released by host njA by user userA on the t-th day
[0086]
[0087] In one implementation, the time decay coefficient in each preset time period is determined based on the ranking of the corresponding preset time period within the historical time period.
[0088] For example, the time decay coefficient on the t-th day
[0089]
[0090] In one implementation, the TF-IDF value of the playback behavior of the voice released by the host in each preset time period is determined based on the number of times the target object plays the voice released by the corresponding host in the corresponding preset time period, the number of times the target object plays the voices released by all hosts in the corresponding preset time period, the number of times all users play the voices released by all hosts in the corresponding preset time period, and the number of times all users play the voice released by the corresponding host in the corresponding preset time period.
[0091] For example, the TF-IDF value of the playback behavior of the voice released by host njA by user userA on the t-th day
[0092] tf-idf(userA, njA, t, type = playback)
[0093] = tf(userA, njA, t, type = playback) * idf(userA, njA, t, type = playback)
[0094] Where
[0095]
[0096]
[0097] C(userA, njA) represents the number of times user userA plays the voice (program) of anchor njA on the t-th day, and ∑∑C(user_i, nj_j) represents the total number of times all users play the voices (programs) of all anchors on the t-th day.
[0098] It should be noted that in the above-described embodiments, the calculation of the weight of the playback behavior of the target object on the voices released by the corresponding anchor in each preset time period, the calculation of the weight of the playback behavior of the target object on the voices released by the anchor in each preset time period, the calculation of the time decay coefficient in each preset time period, and the calculation of the TF-IDF value of the playback behavior of the target object on the voices released by the anchor in each preset time period are merely exemplary descriptions and do not limit this specification. In actual applications, the hyperparameters in the above calculations can be adjusted according to the specific business background.
[0099] Sub-step 1022: Determine the weight of the positive feedback behavior of the target object on the corresponding anchor in the corresponding preset time period based on the weight of the positive feedback behavior of the target object on the voices released by the anchor in each preset time period within the historical time period, the time decay coefficient, the TF-IDF value of the positive feedback behavior, and the number of positive feedback behaviors.
[0100] For example, user userA has a like behavior during the process of listening to voices (programs) on the online audio platform. The weight of user userA's positive feedback behavior on anchor njA on the t-th day (the positive feedback behavior is a like)
[0101] W(userA, njA, t, type = like) = 0.8 * g(t) * tf-idf(userA, njA, t, type = like) * log2(The number of voices liked by this anchor + 1)
[0102] Among them, 0.8 is the weight of the positive feedback behavior of the target object on the voices released by the anchor (the positive feedback behavior is a like), g(t) is the time decay coefficient on the t-th day, and tf-idf(userA, njA, t, type = like) is the TF-IDF value of the positive feedback behavior of user userA on the voices released by anchor njA on the t-th day (the positive feedback behavior is a like).
[0103] Another example, user userA has a collection behavior during the process of listening to voices (programs) on the online audio platform. The weight of user userA's positive feedback behavior on anchor njA on the t-th day (the positive feedback behavior is a collection)
[0104] W(userA, njA, t, type = collection) = 1.0 * g(t) * tf-idf(userA, njA, t, type = collection) * log2(The number of voices collected by this anchor + 1)
[0105] Among them, 1.0 is the weight of the positive feedback behavior of the target object towards the voice released by the host (the positive feedback behavior is collection), g(t) is the time decay coefficient on the t-th day, and tf-idf(userA, njA, t, type = collection) is the TF-IDF value of the positive feedback behavior of user userA towards the voice released by host njA on the t-th day (the positive feedback behavior is collection).
[0106] In one implementation, the time decay coefficient within each preset time period is determined based on the ranking of the corresponding preset time period within the historical time period.
[0107] For example, the time decay coefficient on the t-th day
[0108]
[0109] In one implementation, the TF-IDF value of the positive feedback behavior towards the voice released by the host within each preset time period is determined based on the number of positive feedback behaviors of the target object towards the voice released by the corresponding host within the corresponding preset time period, the number of positive feedback behaviors of the target object towards the voices released by all hosts within the corresponding preset time period, the number of positive feedback behaviors of all users towards the voices released by all hosts within the corresponding preset time period, and the number of positive feedback behaviors of all users towards the voice released by the corresponding host within the corresponding preset time period.
[0110] For example, the TF-IDF value of the positive feedback behavior of user userA towards the voice released by host njA on the t-th day (the positive feedback behavior is a like)
[0111] tf-idf(userA, njA, t, type = like)
[0112] = tf(userA, njA, t, type = like) * idf(userA, njA, t, type = like)
[0113] Among them,
[0114]
[0115]
[0116] C(userA, njA) represents the number of likes of user userA for the voice (program) of host njA on the t-th day, and ∑∑C(user_i, nj_j) represents the number of likes of all users for the voices (programs) of all hosts on the t-th day.
[0117] For another example, the TF-IDF value of the positive feedback behavior (the positive feedback behavior is collection) of user userA for the voice released by the host njA on the t-th day
[0118] tf-idf(userA, njA, t, type = collection)
[0119] = tf(userA, njA, t, type = collection) * idf(userA, njA, t, type = collection)
[0120] where
[0121]
[0122]
[0123] C(userA, njA) represents the number of times user userA has collected the voice (program) of host njA on the t-th day, and ∑∑C(user_i, nj_j) represents the total number of times all users have collected the voices (programs) of all hosts on the t-th day.
[0124] It should be noted that in the above-described embodiments, the weight assignment of the positive feedback behavior of the target object for the voice released by the corresponding host in each preset time period, the calculation of the time decay coefficient in each preset time period, and the calculation of the TF-IDF value of the positive feedback behavior of the target object for the voice released by the host in each preset time period are merely exemplary descriptions and do not limit this specification. In practical applications, the hyperparameters in the above assignments and calculations can be adjusted according to the specific business background.
[0125] Sub-step 1023: Determine the preference weight of the target object for each host based on the playback behavior weight of the target object for each host in each preset time period and the positive feedback behavior weight of the target object for each host in each preset time period.
[0126] For example, the preference weight of user userA for host njA
[0127]
[0128] Similarly, the preference weights of user userA for other hosts can be obtained.
[0129] It should be noted that in the above-described embodiments, the calculation of the preference weight of the target object for each host is merely an exemplary description and does not limit this specification. In practical applications, other formulas can be used for calculation.
[0130] Based on the preference weights of user userA for each anchor, the set of preferred anchors of user userA can be determined.
[0131] In one implementation, if the target object has no playback behavior, obtain the information of the anchor corresponding to the voice played to the target object, assign the preset weight value as the preference weight of the target object for the corresponding anchor, and use the corresponding anchor as the preferred anchor of the target object.
[0132] For example, if user userA has no playback behavior, obtain the portrait of the placement material of the user, obtain the anchor of the played voice, set the preference weight of user userA for the corresponding anchor to 1, and at the same time use the corresponding anchor as the preferred anchor of user userA.
[0133] In another implementation, if the target object has no playback behavior and has not been played with voice, calculate the preference weights of each anchor for the user group with the same basic user portrait as the target object, determine the basic preference weights of each anchor based on the preference weights of the user group for each anchor, and use the top M anchors with the highest basic preference weights as the set of preferred anchors of the target object, and use the basic preference weights of the corresponding anchors as the preference weights of the target object for the corresponding anchors.
[0134] For example, obtain the age group, gender, mobile phone brand, city where the user is located, and user type of user userA, find the user group with the same age group, gender, mobile phone brand, city where the user is located, and user type, calculate the preference weights of each anchor for the user group, and average the accumulated preference of the user group's anchor portrait to obtain the basic preference weights of each anchor. In the preference of the user group's anchor portrait, if the anchor portrait preference of anchor njA appears less than 100 times in the user group, then remove the anchor portrait preference of njA, and then select the top 20 anchors with the highest basic preference weights as the set of preferred anchors of user userA, and assign the preference weight value of user userA for the top 20 anchors as the basic preference weight of the corresponding anchor.
[0135] Step 103: Determine the real-time negative feedback anchor set of the target object based on the preference weights of the target object for each anchor, the exposure times of the anchors in the last N preset time periods, and the unfollowed anchor set.
[0136] If the preference weight of the user for the anchor is small, and the recent cumulative exposure times are many or the user has unfollowed the anchor, it can represent that the user is no longer interested in the anchor, so as to determine that the anchor is the real-time negative feedback anchor of the user.
[0137] For example, if the preference weight of user userA for anchor njA is less than 0.2, and the cumulative exposure times of anchor njA in the last 3 days exceed 20 times or user userA has unfollowed anchor njA, then njA is considered as the real-time negative feedback anchor of user userA.
[0138] Step 104: Filter the anchors in the real-time negative feedback anchor set from the set of preferred anchors of the target object, and filter the anchors in the real-time negative feedback anchor set from the set of followed anchors of the target object, and recall the voices released by the anchors according to the filtered set of preferred anchors and the set of followed anchors and recommend them to the target object.
[0139] After filtering the negative feedback anchors of the user, deduplicate the set of preferred anchors and the set of followed anchors. The deduplication logic can be: if a certain anchor is both in the set of preferred anchors and in the set of followed anchors, then remove this anchor from the set of preferred anchors.
[0140] The voice recommendation logic can be:
[0141] Strategy 1: For any anchor in the set of followed anchors of the user, recall the latest 10 voices (programs) released by this anchor, and also recall 10 voices (programs) released before.
[0142] Strategy 2: For any anchor in the set of preferred anchors of the user, first sort the voices (programs) released by this anchor according to the quality score, then remove the top 100 programs, and finally randomly select 20 voices (programs) as the recalled voices (programs).
[0143] It should be noted that in the above-described embodiments, regarding the deduplication logic and the voice recommendation logic, the above is only an exemplary description and does not limit this specification. In actual applications, it can be adjusted according to the specific business background.
[0144] In this embodiment, historical behavior data of the target object playing sounds, sound data to which the target object is exposed, a set of followed hosts, a set of unfollowed hosts, and the number of times the target object exposes to hosts in the last N preset time periods within the set historical time period are obtained; the historical behavior data includes playing behavior data and positive feedback behavior data; based on the weight of the playing behavior of the target object for the sounds released by the host, the time decay coefficient, and the TF-IDF value of the playing behavior in each preset time period within the historical time period, the playing behavior weight of the target object for the corresponding host in the corresponding preset time period is determined; based on the weight of the positive feedback behavior of the target object for the sounds released by the host, the time decay coefficient, the TF-IDF value of the positive feedback behavior, and the number of positive feedback behaviors in each preset time period within the historical time period, the positive feedback behavior weight of the target object for the corresponding host in the corresponding preset time period is determined; based on the playing behavior weight of the target object for each host in each preset time period and the positive feedback behavior weight of the target object for each host in each preset time period, the preference weight of the target object for each host is determined; based on the preference weight of the target object for each host, the set of preferred hosts of the target object is determined; based on the preference weight of the target object for each host, the number of times the target object exposes to the host in the last N preset time periods, and the set of unfollowed hosts, the set of real-time negative feedback hosts of the target object is determined; the hosts in the set of real-time negative feedback hosts in the set of preferred hosts of the target object are filtered, and the hosts in the set of real-time negative feedback hosts in the set of followed hosts of the target object are filtered, and the sounds released by the hosts are recalled and recommended to the target object according to the filtered set of preferred hosts and the set of followed hosts.
[0145] Collect playing behavior data and positive feedback behavior data and combine them with Newton's laws of thermodynamics to determine the preference portrait of the target object. According to the preference portrait of the target object, the set of preferred hosts is determined. Then, the set of preferred hosts and the set of followed hosts are filtered. The filtered set of preferred hosts and the set of followed hosts can effectively hit the hosts that the target object is interested in, greatly improving the accuracy of sound recommendation. The user behavior portrait, the material portrait, and the basic attributes are cascaded, which can effectively solve the problem of cold start of sounds.
[0146] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described action sequences, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.
[0147] Embodiment 2
[0148] Figure 2A structural schematic diagram of a voice recommendation device provided in the second embodiment of the present invention. The voice recommendation device may specifically include the following modules:
[0149] An acquisition module 201, configured to acquire historical behavior data of a target object playing voices within a set historical time period, voice data to which the target object is exposed, a set of followed anchors, a set of unfollowed anchors, and the number of times the target object exposes an anchor within the last N preset time periods within the historical time period; the historical behavior data includes play behavior data and positive feedback behavior data.
[0150] A preferred anchor set determination module 202, configured to determine the play behavior weight of the target object for a corresponding anchor within a corresponding preset time period based on the weight of the play behavior of the target object for the voices published by the anchor within each preset time period within the historical time period, a time decay coefficient, and the TF-IDF value of the play behavior; determine the positive feedback behavior weight of the target object for a corresponding anchor within a corresponding preset time period based on the weight of the positive feedback behavior of the target object for the voices published by the anchor within each preset time period within the historical time period, a time decay coefficient, the TF-IDF value of the positive feedback behavior, and the number of positive feedback behaviors; determine the preference weight of the target object for each anchor based on the play behavior weight of the target object for each anchor within each preset time period and the positive feedback behavior weight of the target object for each anchor within each preset time period; determine the preferred anchor set of the target object based on the preference weight of the target object for each anchor.
[0151] A real-time negative feedback anchor set determination module 203, configured to determine the real-time negative feedback anchor set of the target object based on the preference weight of the target object for each anchor, the number of times the target object exposes an anchor within the last N preset time periods, and the set of unfollowed anchors.
[0152] A recommendation module 204, configured to filter the anchors in the real-time negative feedback anchor set from the preferred anchor set of the target object, and filter the anchors in the real-time negative feedback anchor set from the followed anchor set of the target object, and recall the voices published by the anchors and recommend them to the target object according to the filtered preferred anchor set and followed anchor set.
[0153] In an implementation manner, the weight of the play behavior of the target object for the voices published by the anchor within each preset time period is determined based on the play duration of the target object for the voices published by the corresponding anchor within the corresponding preset time period and the total duration of the voices published by the corresponding anchor to which the target object is exposed within the corresponding preset time period.
[0154] In an implementation manner, the time decay coefficient within each preset time period is determined based on the sorting of the corresponding preset time period within the historical time period.
[0155] In one embodiment, the TF-IDF value of the playback behavior of the voice released by the host in each preset time period is determined based on the number of times the target object plays the voice released by the corresponding host in the corresponding preset time period, the number of times the target object plays the voices released by all hosts in the corresponding preset time period, the number of times all users play the voices released by all hosts in the corresponding preset time period, and the number of times all users play the voice released by the corresponding host in the corresponding preset time period.
[0156] In one embodiment, the TF-IDF value of the positive feedback behavior of the voice released by the host in each preset time period is determined based on the number of positive feedback behavior times of the target object for the voice released by the corresponding host in the corresponding preset time period, the number of positive feedback behavior times of the target object for the voices released by all hosts in the corresponding preset time period, the number of positive feedback behavior times of all users for the voices released by all hosts in the corresponding preset time period, and the number of positive feedback behavior times of all users for the voice released by the corresponding host in the corresponding preset time period.
[0157] In one embodiment, the types of positive feedback behaviors include a like behavior and a favorite behavior.
[0158] In one embodiment, the preferred host set determination module 202 is further configured to:
[0159] If the target object has no playback behavior, obtain the information of the host corresponding to the voice played by the target object, assign the preference weight of the target object for the corresponding host as a preset weight value, and use the corresponding host as the preferred host of the target object;
[0160] If the target object has no playback behavior and has not been played with a voice, calculate the preference weights of each host for the user group with the same basic user profile as the target object, determine the basic preference weights of each host based on the preference weights of the user group for each host, and use the top M hosts with the highest basic preference weights as the preferred host set of the target object, and use the basic preference weights of the corresponding hosts as the preference weights of the target object for the corresponding hosts.
[0161] The voice recommendation device provided by the embodiments of the present invention can execute the voice recommendation method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0162] Embodiment III
[0163] Figure 3 It is a schematic structural diagram of a computer device provided by Embodiment III of the present invention. Figure 3 A block diagram of an exemplary computer device 12 suitable for implementing the embodiments of the present invention is shown. Figure 3The computer device 12 shown is only an example and should not impose any limitation on the functions and the scope of use of the embodiments of the present invention.
[0164] As Figure 3 shown, the computer device 12 appears in the form of a general-purpose computing device. The components of the computer device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 that connects different system components (including the system memory 28 and the processing unit 16).
[0165] The bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus structures. By way of example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.
[0166] The computer device 12 typically includes a variety of computer system-readable media. These media can be any available media that can be accessed by the computer device 12, including volatile and non-volatile media, removable and non-removable media.
[0167] The system memory 28 may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system 34 can be used for reading and writing on a non-removable, non-volatile magnetic medium ( Figure 3 not shown, commonly referred to as a "hard disk drive"). Although Figure 3 not shown in the figure, a disk drive for reading and writing on a removable non-volatile disk (such as a "floppy disk"), and an optical disk drive for reading and writing on a removable non-volatile optical disk (such as a CD-ROM, DVD-ROM, or other optical medium) can be provided. In these cases, each drive can be connected to the bus 18 through one or more data medium interfaces. The memory 28 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the embodiments of the present invention.
[0168] A program / utilities 40 having a set (at least one) of program modules 42 can be stored, for example, in the memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment. The program modules 42 generally execute the functions and / or methods in the embodiments described in the present invention.
[0169] The computer device 12 can also communicate with one or more external devices 14 (such as a keyboard, a pointing device, a display 24, etc.), and can also communicate with one or more devices that enable a user to interact with the computer device 12, and / or communicate with any device that enables the computer device 12 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 22. Moreover, the computer device 12 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 20. As shown in the figure, the network adapter 20 communicates with other modules of the computer device 12 through the bus 18. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the computer device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0170] The processing unit 16 executes various functional applications and data processing by running the programs stored in the system memory 28, such as implementing the voice recommendation method provided by the embodiments of the present invention.
[0171] Embodiment 4
[0172] Embodiment 4 of the present invention also provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, it realizes each process of the above-mentioned voice recommendation method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0173] Among them, a computer-readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0174] Note that the above is only a preferred embodiment of the present invention and the applied technical principle. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A voice recommendation method, characterized in that, Including: Obtaining historical behavior data of the target object playing sounds within a set historical time period, sound data to which the target object is exposed, a set of followed streamers, a set of unfollowed streamers, and the number of times the target object exposes to streamers within the last N preset time periods within the historical time period; the historical behavior data includes playing behavior data and positive feedback behavior data; Determining the playing behavior weight of the target object for the corresponding streamer within the corresponding preset time period based on the weight of the playing behavior of the target object for the sounds released by the streamer within each preset time period within the historical time period, the time decay coefficient, and the TF-IDF value of the playing behavior; Determining the positive feedback behavior weight of the target object for the corresponding streamer within the corresponding preset time period based on the weight of the positive feedback behavior of the target object for the sounds released by the streamer within each preset time period within the historical time period, the time decay coefficient, the TF-IDF value of the positive feedback behavior, and the number of positive feedback behaviors; Determining the preference weight of the target object for each streamer based on the playing behavior weight of the target object for each streamer within each preset time period and the positive feedback behavior weight of the target object for each streamer within each preset time period; Determining the set of preferred streamers of the target object based on the preference weight of the target object for each streamer; Determining the real-time negative feedback streamer set of the target object based on the preference weight of the target object for each streamer, the number of times the target object exposes to streamers within the last N preset time periods, and the set of unfollowed streamers; Filtering the streamers in the real-time negative feedback streamer set from the set of preferred streamers of the target object and filtering the streamers in the real-time negative feedback streamer set from the set of followed streamers of the target object, and recalling the sounds released by the streamers according to the filtered set of preferred streamers and the set of followed streamers and recommending them to the target object; The weight of the playing behavior of the target object for the sounds released by the streamer within each preset time period is determined based on the playing duration of the target object for the sounds released by the corresponding streamer within the corresponding preset time period and the total duration of the target object exposing to the sounds released by the corresponding streamer within the corresponding preset time period; If the target object has no playing behavior, obtain the information of the streamer corresponding to the sound to which the target object is delivered, assign the preference weight of the target object for the corresponding streamer as a preset weight value, and use the corresponding streamer as the preferred streamer of the target object; If the target object has no playing behavior and has not been delivered any sounds, calculate the preference weights of each streamer for the user group with the same basic user profile as the target object, determine the basic preference weights of each streamer based on the preference weights of each streamer for the user group, and use the top M streamers with the highest basic preference weights as the set of preferred streamers of the target object, and use the basic preference weights of the corresponding streamers as the preference weights of the target object for the corresponding streamers.
2. The method according to claim 1, wherein: The time decay coefficient within each preset time period is determined based on the ranking of the corresponding preset time period within the historical time period.
3. The method according to claim 1, characterized in that: The TF-IDF value of the playback behavior of the voice released by the host in each preset time period is determined based on the number of times the target object plays the voice released by the corresponding host in the corresponding preset time period, the number of times the target object plays the voices released by all hosts in the corresponding preset time period, the number of times all users play the voices released by all hosts in the corresponding preset time period, and the number of times all users play the voice released by the corresponding host in the corresponding preset time period.
4. The method according to claim 1, wherein: The TF-IDF value of the positive feedback behavior of the voice released by the host in each preset time period is determined based on the number of positive feedback behavior times of the target object for the voice released by the corresponding host in the corresponding preset time period, the number of positive feedback behavior times of the target object for the voices released by all hosts in the corresponding preset time period, the number of positive feedback behavior times of all users for the voices released by all hosts in the corresponding preset time period, and the number of positive feedback behavior times of all users for the voice released by the corresponding host in the corresponding preset time period.
5. The method according to claim 1, characterized in that: The types of the positive feedback behavior include a like behavior and a favorite behavior.
6. A voice recommendation device, characterized in that, It includes: An obtaining module, configured to obtain the historical behavior data of the target object playing voices in a set historical time period, the voice data exposed by the target object, the set of followed hosts, the set of unfollowed hosts, and the number of times the target object exposes the hosts in the last N preset time periods within the historical time period; the historical behavior data includes playback behavior data and positive feedback behavior data. A preferred host set determining module, configured to determine the playback behavior weight of the target object for the corresponding host in the corresponding preset time period based on the weight of the playback behavior of the voice released by the host in each preset time period within the historical time period of the target object, the time decay coefficient, and the TF-IDF value of the playback behavior. Based on the weight of the positive feedback behavior of the voice released by the host in each preset time period within the historical time period of the target object, the time decay coefficient, the TF-IDF value of the positive feedback behavior, and the number of positive feedback behavior times, determine the positive feedback behavior weight of the target object for the corresponding host in the corresponding preset time period. Based on the playback behavior weight of the target object for each host in each preset time period and the positive feedback behavior weight of the target object for each host in each preset time period, determine the preference weight of the target object for each host. Based on the preference weight of the target object for each host, determine the preferred host set of the target object. A real-time negative feedback host set determining module, configured to determine the real-time negative feedback host set of the target object based on the preference weight of the target object for each host, the number of times the target object exposes the hosts in the last N preset time periods, and the set of unfollowed hosts. A recommendation module, configured to filter the hosts in the real-time negative feedback host set from the preferred host set of the target object, and filter the hosts in the real-time negative feedback host set from the set of followed hosts of the target object, and recall the voices released by the hosts and recommend them to the target object according to the filtered preferred host set and the set of followed hosts. The weight of the playback behavior of the voice released by the host in each preset time period is determined based on the playback duration of the voice released by the corresponding host by the target object in the corresponding preset time period and the total exposure duration of the voice released by the corresponding host by the target object in the corresponding preset time period; The preferred host set determination module is further configured to: If the target object has no playback behavior, obtain the information of the host corresponding to the voice played by the target object, assign the preference weight of the target object for the corresponding host as the preset weight value, and use the corresponding host as the preferred host of the target object; If the target object has no playback behavior and has not been played with a voice, calculate the preference weights of each host for the user group with the same basic user profile as the target object, determine the basic preference weights of each host based on the preference weights of each host for the user group, and use the top M hosts with the highest basic preference weights as the preferred host set of the target object, and use the basic preference weights of the corresponding hosts as the preference weights of the target object for the corresponding hosts.
7. A computer device, characterized in that, The computer device includes: One or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the voice recommendation method according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that: A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the voice recommendation method according to any one of claims 1-5 is implemented.
Citation Information
Patent Citations
Subscription anchor sorting method and apparatus, and terminal
CN107948752A