Feature Encoding Model Training Method and Apparatus, Media Object Recommendation Method and Apparatus
Through the feature encoding model training method, the music cold start problem is solved, and accurate recall and recommendation in the absence of user interaction data is achieved, which improves the recall accuracy of the recommendation system and user interest perception ability.
Patent Information
- Application Number
- CN202211308060.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-24
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-10-24
AI Technical Summary
It is difficult for existing recommendation systems to accurately recommend new music in music cold-start scenarios, and existing recall strategies such as content-based recall, collaborative filtering-based recall and characterization-based embedding are difficult to effectively solve, especially in the absence of user interaction data.
The feature encoding model training method is adopted, and the initial vectors of the target sample, positive sample, negative sample and comparison sample are obtained, feature encoding and mask processing are performed, similarity values are calculated, and the model is jointly trained to improve recall accuracy and solve the impact of the long-tail effect.
It realizes accurate recall of new media objects in cold start scenarios, improves recommendation accuracy and recall performance, can better perceive the characteristics of user interest groups and long-tail media objects, and improves recommendation accuracy.
Smart Images

Figure CN115600017B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of artificial intelligence technology, and particularly to a method and apparatus for training a feature encoding model, and a method and apparatus for recommending media objects. Background Art
[0002] The cold start problem is a very important problem in the recommendation system. For example, the cold start problem of music refers to the situation in a music platform where for new music, since there is no interaction with users, it is difficult for the music platform to recommend it to users who are interested in it. Therefore, those skilled in the art urgently need to solve the cold start problem of music. Summary of the Invention
[0003] According to one aspect of the embodiments of the present disclosure, there is provided a method for training a feature encoding model, including: constructing a feature encoding model to be trained, and obtaining initial vectors of a target sample, a positive sample, a negative sample, and a contrast sample; performing feature encoding on the initial vectors of the target sample, the positive sample, and the negative sample respectively according to the feature encoding model to obtain encoded vectors of each sample, and calculating correlation vectors between the encoded vectors of each sample and a preset user interest group matrix respectively; calculating a first similarity value between the correlation vector corresponding to the target sample and the correlation vector corresponding to the positive sample, and calculating a second similarity value between the correlation vector corresponding to the target sample and the correlation vector corresponding to the negative sample, and calculating a first loss value according to the first similarity value and the second similarity value; grouping the feature vectors corresponding to multiple features included in the initial vector of the contrast sample to obtain multiple feature vector groups corresponding to the contrast sample, and performing different feature masking processes on the multiple feature vector groups to obtain corresponding multiple masked vectors, and performing feature encoding on the multiple masked vectors respectively according to the feature encoding model to obtain multiple encoded vectors corresponding to the contrast sample; calculating a third similarity value between the multiple encoded vectors of the same contrast sample, and calculating a fourth similarity value between the encoded vectors of different contrast samples, and calculating a second loss value according to the third similarity value and the fourth similarity value; and jointly training the feature encoding model based on the first loss value and the second loss value.
[0004] According to one aspect of the embodiments of the present disclosure, a feature encoding model training device is provided, including: a construction module configured to construct a feature encoding model to be trained and obtain initial vectors of a target sample, a positive sample, a negative sample, and a contrast sample; a first extraction module configured to perform feature encoding on the initial vectors of the target sample, the positive sample, and the negative sample respectively according to the feature encoding model to obtain encoded vectors of each sample, and calculate association vectors between the encoded vectors of each sample and a preset user interest group matrix; a first calculation module configured to calculate a first similarity value between the association vector corresponding to the target sample and the association vector corresponding to the positive sample, and calculate a second similarity value between the association vector corresponding to the target sample and the association vector corresponding to the negative sample, and calculate a first loss value according to the first similarity value and the second similarity value; a second extraction module configured to group eigenvectors corresponding to multiple features included in the initial vector of the contrast sample to obtain multiple eigenvector groups corresponding to the contrast sample, and perform different feature masking processes on the multiple eigenvector groups to obtain corresponding multiple masked vectors, and perform feature encoding on the multiple masked vectors respectively according to the feature encoding model to obtain multiple encoded vectors corresponding to the contrast sample; a second calculation module configured to calculate a third similarity value between multiple encoded vectors of the same contrast sample, and calculate a fourth similarity value between encoded vectors of different contrast samples, and calculate a second loss value according to the third similarity value and the fourth similarity value; a joint training module configured to perform joint training on the feature encoding model based on the first loss value and the second loss value.
[0005] According to one aspect of the embodiments of the present disclosure, a media object recommendation method is provided, including: obtaining an initial vector of a media object to be recommended, and calculating an encoded vector of the media object to be recommended according to the initial vector by using a feature encoding model trained by using the feature encoding model training method as described above; calculating an association vector between the encoded vector and a preset user interest group matrix; querying in a vector pool for multiple target media objects whose association vectors are similar to the association vector of the media object to be recommended, where the vector pool is used to store association vectors of pre-collected media objects, and the association vectors of the pre-collected media objects are calculated based on the preset user interest group matrix after obtaining corresponding encoded vectors through the feature encoding model; sorting the multiple target media objects according to the approximation degree values between the association vectors of each target media object and the association vector of the media object to be recommended, and recommending the media object to be recommended to users associated with the target media objects based on the sorting result.
[0006] According to one aspect of the embodiments of the present disclosure, there is provided a media object recommendation device, including: an acquisition module configured to acquire an initial vector of a media object to be recommended, and calculate an encoded vector of the media object to be recommended according to the initial vector by using a feature encoding model trained by the feature encoding model training method as described above; a calculation module configured to calculate a correlation vector between the encoded vector and a preset user interest group matrix; a recall module configured to query, in a vector pool, a plurality of target media objects that are similar to the correlation vector of the media object to be recommended, where the vector pool is used to store the correlation vectors of pre-collected media objects, and the correlation vectors of the pre-collected media objects are calculated based on the preset user interest group matrix after obtaining the corresponding encoded vectors through the feature encoding model; and a recommendation module configured to sort the plurality of target media objects according to the similarity degree values between the correlation vectors of the respective target media objects and the correlation vector of the media object to be recommended, and recommend the media object to be recommended to the users associated with the target media objects based on the sorting result.
[0007] According to one aspect of the embodiments of the present disclosure, there is provided an electronic device, including: one or more processors; a storage device for storing one or more programs, which, when executed by the one or more processors, cause the electronic device to implement the feature encoding model training method or the media object recommendation method as described above.
[0008] According to one aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which computer-readable instructions are stored, which, when executed by a processor of a computer, cause the computer to execute the feature encoding model training method and the media object recommendation method as described above.
[0009] In the technical solution provided by the embodiments of the present disclosure, the trained feature encoding model is used to recall music in a music cold start scenario, so as to recommend new music to the users associated with the recalled music, in order to solve the problem of music cold start. Since the feature encoding model can not only perceive user interest groups based on the processing of target samples, positive samples, and negative samples during the training process, thereby improving the accuracy of the recall process, but also enhance the vector representation of the feature encoding model based on the processing of contrast samples, so that the trained feature encoding model can promote the accurate recall of music in the process of solving the music cold start problem, thereby realizing more accurate recommendation of music to relevant users. Description of the Drawings
[0010] Figure 1 is a schematic diagram of the implementation environment involved in this application;
[0011] Figure 2It is a flowchart of a feature encoding model training method shown in an exemplary embodiment;
[0012] Figure 3 It is a flowchart of the steps further included in the feature encoding model training method shown in another exemplary embodiment of the present application on the basis of the embodiment shown in Figure 2 ;
[0013] Figure 4 It is a schematic diagram of the process of calculating the correlation vectors between the encoding vectors of each sample and a preset user interest group matrix respectively;
[0014] Figure 5 It is Figure 2 An exemplary flowchart of step S220 in the embodiment shown;
[0015] Figure 6 It is a schematic diagram of the composition of the initial vector corresponding to an exemplary sample;
[0016] Figure 7 It is a schematic diagram of an exemplary feature correlation matrix;
[0017] Figure 8 It is Figure 2 An exemplary flowchart of step S240 in the embodiment shown;
[0018] Figure 9 It is a schematic diagram of the effects of using three data augmentation operators in an exemplary computer vision field;
[0019] Figure 10 It shows a schematic diagram of the overall process of obtaining multiple encoding vectors corresponding to comparison samples;
[0020] Figure 11 It is Figure 2 An exemplary flowchart of step S260 in the embodiment shown;
[0021] Figure 12 It is a flowchart of an exemplary media object recommendation method;
[0022] Figure 13 It is a block diagram of an exemplary feature encoding model training device;
[0023] Figure 14 It is a block diagram of an exemplary media object recommendation device;
[0024] Figure 15 It shows a schematic diagram of the structure of a computer system of an electronic device suitable for implementing the embodiments of the present disclosure. Detailed implementation manners
[0025] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0026] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0027] The flowcharts shown in the drawings are only exemplary illustrations and do not necessarily include all contents and operations / steps, nor do they have to be executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.
[0028] As used in the present disclosure, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0029] First of all, it should be noted that the cold start problem, also known as the Cold-Start Problem, is a very important problem in the recommendation system. The cold start problem can generally be divided into three categories, namely user cold start, item cold start, and system cold start. User cold start refers to how to recommend items to new users since there is no behavior data of them in the recommendation system. Item cold start refers to how to recommend a new item to users interested in it since there has been no interaction with users. System cold start refers to how to recommend interesting items to users for a newly developed platform since there is no user behavior and only some item information. It should be understood that the items mentioned in this application are in a broad sense, and the content on various platforms can be called items. For example, the goods on the commodity trading platform can be called items, the videos on the video platform can be called items, the music on the music platform can be called items, and the news on the news platform can also be called items. There should be no narrow understanding of items.
[0030] The type of cold start problem solved by the embodiments of this application is item cold start. For example, on a music platform, for new songs, since there is no interaction data generated with the users on the music platform, it is difficult for the music platform to recommend them to users who are interested in them.
[0031] In addition, it should be noted that a recommendation system usually includes a recall layer and a ranking layer. The recall layer, also known as Matching, is used to quickly reduce a large number of candidate sets to a smaller scale, and try to quickly filter out the items that users are interested in at this stage. The ranking layer, also known as Ranking, is used to obtain an accurate ranking result, that is, to accurately rank the items recalled by the recall layer according to rules. The recall layer and the ranking layer can be understood as software modules that make up the recommendation system.
[0032] Taking a music platform as an example, before recommending music to users, the recall layer is usually used to recall multiple candidate sets that users may be interested in from a large number of candidate musics. On the one hand, it can reduce unnecessary calculations. On the other hand, since the ranking layer is usually a single objective, the recommendation diversity and accuracy can be improved during the process of using multiple recalls. Therefore, the recall layer is an important module in the structure of the recommendation system, and the recall strategy adopted by the recall layer is crucial for the recommendation system.
[0033] In the prior art, the recall strategies adopted by the recall layer mainly include three categories, namely content-based recall, collaborative filtering-based recall, and representation embedding-based recall.
[0034] If still taking a music platform as an example, content-based recall mainly relies on the content tags of music to learn the correlation between musics. The content tags include information such as text and pictures. After establishing a content understanding model, the content vectors of existing musics are calculated offline and saved, and a vector similarity retrieval index is constructed. During online recall, the content vector of new music is calculated, and then the nearest neighbor index is queried to obtain the topK (that is, the existing musics whose content vector approximation degree ranks among the top K) recall set. However, implementing recall only from the level of content understanding will cause a large deviation in user recommendation in the case of music cold start.
[0035] Collaborative filtering-based recall constructs a similarity matrix between musics based on the interaction matrix of users and musics, and then stores it in the form of an inverted index. During online recall, triggered by the user's historical behavior, the inverted index is queried to obtain a set of similar musics, and finally the topK recall set is aggregated. However, this recall method depends on a large amount of user-music interaction data to mine co-occurrence information. For new music, since there is no interaction data with users, it is difficult to achieve accurate recall by this recall method, and thus it is also difficult to achieve accurate recommendation.
[0036] The recall based on representation embedding draws on the idea of Word2vec (a word embedding method used to calculate the distributed word vectors of each word in its given corpus environment). It regards the user's behavior as a sentence and believes that two adjacent musics within a sliding window are similar, so the word vectors of these two music vectors should also be approximate. According to this hypothesis, positive and negative samples are constructed, and a model is built for recall. Finally, a similarity retrieval index is constructed based on the music vectors learned by the model. During online recall, triggered by the user's historical behavior, the inverted similarity retrieval index is queried to obtain the topK recall set. However, this method captures the associations between musics from large-scale user music interaction data and is difficult to support the recall of new music.
[0037] As can be seen from the above, these existing recall strategies are all difficult to solve the cold start problem of music. In addition, there is usually a long-tail effect in music. The long-tail effect can be understood as being related to the popularity of music. Music with higher popularity can usually maintain its popularity for a longer time, while the popularity of music with lower popularity on the music platform often fades quickly. Based on this, when solving the cold start problem of music, the impact of the long-tail effect of music on the recall accuracy also needs to be considered.
[0038] To address the above issues comprehensively, the embodiments of this application respectively provide a feature encoding model training method and device, a media object recommendation method and device, an electronic device, and a computer-readable storage medium. These embodiments will be described in detail below. Additionally, it should be noted that the feature encoding model mentioned in the embodiments of this application acts on the recall layer of the recommendation system. The feature encoding model trained using the embodiments of this application can obtain an accurate topK recall set for new media objects in cold start scenarios, and then be used to achieve more accurate media object recommendations. The detailed operation process can be found in the descriptions of the following embodiments and will not be elaborated here. It should also be understood that the media objects mentioned in the embodiments of this application include, for example, music, videos, news, etc., and are not limited herein.
[0039] Please refer to Figure 1 , Figure 1 which is a schematic diagram of the implementation environment involved in this application. This implementation environment can be understood as a media object recommendation system, including a media terminal 110 and a media server 120. Data transmission occurs between the media terminal 110 and the media server 120 based on a pre-established wired or wireless communication connection.
[0040] The media terminal 110 is, for example, a device such as a smart phone, a tablet, a laptop computer, a computer, a vehicle-mounted terminal, etc. The media terminal 110 is usually oriented towards users and is used to provide a user interaction interface for interacting with users. For example, a media platform runs on the media terminal 110, such as a music platform, a video platform, a news consultation platform, etc. Users can operate these platforms on the media terminal 110, thereby performing relevant user interactions and generating user behavior data.
[0041] The media server 120 is, for example, a server and is used to provide data support for the operation of the media platform in the media terminal 110. For example, the media terminal 110 uploads user behavior data to the media server 120 for storage or processing. Or in a cold start scenario, the media terminal 110 transmits a new media object input by the user to the media server 120, so that the media server 120 recommends this new media object to users interested in it.
[0042] It should be noted that the functions of the media terminal 110 and the media server 120 can be determined according to actual application requirements, and the present implementation environment does not limit the specific functions of the media server 120. Additionally, it should be understood that the server mentioned in the present implementation environment can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or it can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. There is no limitation here either.
[0043] Please refer to Figure 2 , Figure 2 is a flowchart of a feature encoding model training method shown in an exemplary embodiment of the present application. This feature encoding training method can be applicable to Figure 1 the implementation environment shown in Figure 1 and is specifically executed by the media server 120 in the implementation environment shown in
[0044] to implement the recommendation of new media objects in a cold start scenario through the trained feature encoding model. Figure 1 It should also be understood that
[0045] the feature encoding model training method shown in Figure 2As shown, the exemplary feature encoding model method includes steps S210 - S260, which are introduced in detail as follows:
[0046] S210, construct a feature encoding model to be trained, and obtain the initial vectors of target samples, positive samples, negative samples, and comparison samples.
[0047] First, it should be noted that the feature encoding model is a machine learning model used to extract encoding vectors for input signals. By training the feature encoding model, a feature encoding model that meets the actual application requirements can be obtained. In the embodiments of the present application, that the feature encoding model meets the actual application requirements can be understood as that the encoding vectors extracted by using the trained feature encoding model can improve the recall accuracy in the media object recommendation process, so that finally the media object to be recommended can be more accurately recommended to users interested in it.
[0048] Exemplarily, the feature encoding model can adopt model structures such as Bidirectional Encoder Representation from Transformers (BERT), Multilayer Perception (MLP), Convolutional Neural Networks (CNN), Attention Network (AttNet), etc. The embodiments of the present application do not limit the specific model structure of the feature encoding model.
[0049] Target samples and positive samples refer to sample pairs extracted from a media object group where the recall from media object to media object has been achieved and the recall effect is good. The recall from media object to media object can be understood as recalling the subsequent media object similar to the previous media object. For example, the recall from media object A to media object B is to recall media object B similar to media object A. A good recall effect can be understood as that this recall has successfully achieved the recommendation of media objects. For example, media object A is recommended to the users of the recalled media object B, and these users have a high degree of interest in media object A. Based on this, media object A and media object B can be called a media object group, where media object A can be used as a target sample and media object B can be used as a positive sample.
[0050] As described above, positive samples can be understood as samples with good recall effects, while negative samples can be correspondingly understood as samples with poor recall effects. Comparison samples can be understood as samples used for recall effect comparison. For example, comparison samples can be randomly sampled samples, which have no direct relationship with the quality of the recall effect.
[0051] The initial vectors of the target sample, positive sample, negative sample, and comparison sample refer to the vectors obtained by performing feature embedding on the feature data of these samples. Feature embedding is also known as "embedding", which represents an object using a low-dimensional, dense, and continuous vector. In this embodiment, each sample is represented by the corresponding initial vector. It can be understood that the initial vectors of the target sample, positive sample, negative sample, and comparison sample each contain multiple features. For example, taking the media object as music, the multiple features may include category features, audio features, lyric features, etc. Each feature corresponds to its own feature vector, and the initial vectors of each sample are formed by combining the feature vectors of multiple features.
[0052] S220. According to the feature encoding model, perform feature encoding on the initial vectors of the target sample, positive sample, and negative sample respectively to obtain the encoded vectors of each sample, and calculate the correlation vectors between the encoded vectors of each sample and the preset user interest group matrix respectively.
[0053] In this embodiment, the initial vectors of the target sample, positive sample, and negative sample are respectively used as input signals and input into the feature encoding model to be trained, and the encoded vectors of each sample output by the feature encoding model can be obtained accordingly. It can be understood that the feature encoding model is used to extract deeper information from the input initial vectors. Therefore, the encoded vectors of each sample are deeper vector representations compared to the initial vectors of each sample.
[0054] The preset user interest group matrix is a matrix pre-collected from the user data of multiple user interest groups and extracted from these user data. Among them, a user interest group can be understood as a user group with unified interests. For example, taking the media object as music, a certain user group likes punk music, and another user group likes folk music.
[0055] In this embodiment, the correlation vectors between the encoded vectors of each sample and the preset user interest group matrix are calculated respectively, so as to reflect which user interest group each sample is liked or interested in through the correlation vectors, so that the first loss value can be calculated based on the correlation vectors later, and the feature encoding model can be trained according to the first loss value, which can make the finally trained feature encoding model perceive the user interest group during the encoding process, improve the interpretability of the model encoding, and later make the trained feature encoding model applicable to the accurate recall of media objects.
[0056] S230. Calculate the first similarity value between the correlation vector corresponding to the target sample and the correlation vector corresponding to the positive sample, and calculate the second similarity value between the correlation vector corresponding to the target sample and the correlation vector corresponding to the negative sample, and calculate the first loss value according to the first similarity value and the second similarity value.
[0057] To improve the recall performance, the distributions of the target samples and positive samples showing being liked or interested by each user interest group should be more consistent, while the distributions of the target samples and negative samples showing being liked or interested by each user interest group should emphasize the differences. Therefore, during the training process of the feature encoding model, the above situations of positive samples and negative samples should be considered simultaneously. Thus, in this embodiment, it is necessary to calculate the first similarity value between the relevance vector corresponding to the target sample and the relevance vector corresponding to the positive sample, and calculate the second similarity value between the relevance vector corresponding to the target sample and the relevance vector corresponding to the negative sample, and calculate the first loss value according to the first similarity value and the second similarity value.
[0058] S240, group the feature vectors corresponding to the multiple features included in the initial vector of the comparison sample to obtain multiple feature vector groups corresponding to the comparison sample, perform different feature masking processes on the multiple feature vector groups to obtain corresponding multiple masked vectors, and perform feature encoding on the multiple masked vectors respectively according to the feature encoding model to obtain multiple encoded vectors corresponding to the comparison sample.
[0059] As described above, the initial vector of the comparison sample is formed by combining the feature vectors corresponding to the multiple features of the comparison sample. In this embodiment, grouping the feature vectors corresponding to the multiple features included in the initial vector of the comparison sample may be to preset the number of feature vector groups, so as to perform uniform feature grouping according to the number of features included in the initial vector of the comparison sample to obtain corresponding multiple feature vector groups. Or, other grouping strategies may also be used to implement the grouping process of the feature vectors, which can be referred to the records in the subsequent embodiments. This embodiment will not elaborate here, and this embodiment also does not limit the specific grouping method.
[0060] The different feature masking processes on the multiple feature vector groups in this embodiment mean that different masking methods are used to mask or add noise to each feature vector group to obtain corresponding multiple masked vectors. Different masking methods include, for example, random masking, span masking, uniform noise, etc., which are not limited here. It can be understood that random masking means randomly sampling a certain proportion of vector element positions in the feature vector group for masking, span masking means randomly sampling a vector element position in the feature vector group and then continuously masking a certain proportion of vector element positions starting from this position, and uniform noise means uniformly adding noise data to the entire feature vector group.
[0061] Next, perform feature encoding on the multiple masked vectors obtained above according to the feature encoding model to be trained, whereby multiple encoded vectors corresponding to the comparison samples can be obtained. It can be seen that in this embodiment, first, the feature vectors corresponding to multiple features included in the initial vector of the comparison sample are grouped to obtain multiple feature vector groups corresponding to the comparison sample, then different feature masking processes are performed on each feature vector group to obtain the corresponding multiple masked vectors, and then feature masking is performed on the multiple masked vectors according to the feature encoding model to obtain multiple encoded vectors corresponding to the comparison sample. Through this process, different view representations of the same comparison sample can be obtained. It can also be understood that since each encoded vector is obtained by grouping the feature vectors in the initial vector of the same comparison sample, using different masking methods for masking, and performing feature encoding on the obtained masked vectors, the multiple encoded vectors obtained can be referred to as different view representations of the same comparison sample.
[0062] S250, calculate the third similarity value between the multiple encoded vectors of the same comparison sample, and calculate the fourth similarity value between the encoded vectors of different comparison samples, and calculate the second loss value according to the third similarity value and the fourth similarity value.
[0063] To improve better recall performance, the different view representations of the same comparison sample should be more consistent, while the view representations between different comparison samples should emphasize differences. Therefore, in the training process of the feature encoding model, such a situation should be considered. Therefore, this embodiment needs to calculate the third similarity value between the multiple encoded vectors of the same comparison sample, and calculate the fourth similarity value between the encoded vectors of different comparison samples, and calculate the second loss value according to the third similarity value and the fourth similarity value. Subsequently, training the feature encoding model based on the second loss value can enable the finally trained feature encoding model to increase its vector representation ability, and this ability can effectively solve the problem of the long-tail effect. For example, the feature encoding model trained in this embodiment can fully extract the representations of long-tail media objects during the feature encoding process.
[0064] It should be understood that still taking the media object as music as an example, the long-tail media object mentioned here refers to music with low popularity. In the music recommendation scenario, usually only 20% of the music on the music platform occupies 80% of the platform traffic. The proportion of long-tail music in the training samples is usually low, resulting in the feature encoding model obtained by ordinary training being able to accurately extract the encoded vectors of music with high popularity, but it is difficult to accurately extract the encoded vectors of long-tail music, thus affecting the accuracy of online recall. However, the feature encoding model trained based on the second loss value in this embodiment can effectively solve this problem.
[0065] S260. Jointly train the feature encoding model based on the first loss value and the second loss value.
[0066] As can be seen from the above description, in this embodiment, the feature encoding model is jointly trained based on the first loss value and the second loss value, so that the trained feature encoding model can not only perceive the user interest group during the feature encoding process, but also handle the inaccuracy problem caused by the long-tail effect. As a result, the trained feature encoding model can promote the accurate recall of media objects during the process of solving the cold start problem, thereby realizing more accurate recommendation of new media objects to relevant users.
[0067] To facilitate the understanding of the training process of the feature encoding model mentioned in the above embodiment, the following will describe this training process in detail in combination with the media object recommendation scenario. And as an example, the media object will also be described by taking music as an example, but it should not be understood that the media object is limited to music. As described in the foregoing embodiment, the media object includes but is not limited to media objects such as music, video, and news information.
[0068] As Figure 3 shown, exemplarily, before step S210 shown in Figure 2 shown, the training method of the feature encoding model further includes the following steps S310 - S330:
[0069] S310. Obtain the media object recommendation data log and the media object metadata. The media object recommendation data log contains the recall information from media object to media object.
[0070] S320. Calculate the recall success rate from media object to media object according to the media object recommendation data log, and extract multiple features from the media object metadata.
[0071] S330. Determine the reference media object and the recommended media object in the media object group with a recall success rate greater than the preset global probability as the target sample and the positive sample respectively, perform negative sampling on the determined positive sample to obtain corresponding multiple negative samples, and sample the media object recommendation data log to obtain comparison samples.
[0072] The above process reveals the acquisition processes of the target sample, positive sample, negative sample, and comparison sample. In step S310, the media object recommendation data log refers to the historical recommendation data of media objects recorded in the recommendation system. For example, it is the recommendation data log of the historical N days collected through the background log or database of the recommendation system, where N is an integer greater than 0. The recall information from media object to media object contained in the media object recommendation data log means that it can be learned from the media object recommendation data log that a certain media object is recalled via another media object. Media object metadata refers to the data used to describe certain characteristics of a media object. For example, taking music as an example, media metadata can refer to the data describing category characteristics, audio characteristics, and lyric characteristics.
[0073] In step S320, since the media object recommendation data log records the historical recommendation data of media objects, the corresponding recommendation feedback information can be obtained from it. According to the recommendation return information, the recall success rate from media object to media object can be determined. For example, taking music as an example, the red heart rate or the completion rate can be used as the recall success rate. The red heart rate refers to the probability that after recommending music A to the user of music B, the user behavior data records the probability that the user gives a red heart to music A. Here, the user giving a red heart to the music can be understood as the user liking the music. The completion rate refers to the probability that the user behavior data records the user playing music A completely. Since media object metadata refers to the data used to describe certain characteristics of a media object, multiple characteristics corresponding to the media object can be extracted from the media object metadata, such as the category characteristics, audio characteristics, lyric characteristics, etc. mentioned above.
[0074] In step S330, the preset global probability is a preset value. If the recall success rate of a group of media objects is greater than this preset value, it means that the recall effect of this group of media objects is good. Therefore, the reference media object and the recommended media object in this media object group are respectively determined as the target sample and the positive sample. It should be understood that in the exemplary scenario of recommending music A to the user of music B, music A is the recommended media object, that is, the positive sample, and music B is the reference media object, that is, the target sample. After determining the positive sample, negative sampling of the positive sample means sampling M negative samples for each positive sample, where M is an integer greater than 0. For example, negative samples can be collected from the music set that has the same users as the positive sample, or all the media objects that appear in the media object recommendation data log and the occurrence frequency of each media object can be counted, and then negative samples can be sampled from these media objects based on the occurrence frequency of the media objects. Therefore, media objects with higher occurrence frequencies are more likely to be sampled as negative samples, and it can be selected according to actual needs. Sampling the media object recommendation data log can be understood as uniform sampling or random sampling in the media object recommendation data log to obtain comparison samples.
[0075] As can be seen from the above, the target samples and positive samples obtained in this embodiment can reflect a good recall effect. The negative samples are not directly associated with the recall process. Therefore, the recall effect should be worse than that of the positive samples. Since the comparison samples are obtained by uniform sampling or random sampling from the media object recommendation data log, the comparison samples are relatively more independent than the positive samples and negative samples. Moreover, since the samples in this embodiment are obtained from the media object recommendation data log, these samples also contain collaborative filtering information.
[0076] The samples obtained in this embodiment can be collectively referred to as training samples for training the feature encoding model. That is, this embodiment can obtain the following quadruple sample set for training the feature encoding model
[0077]
[0078] where i a 、i + 、i - and i c respectively represent the media object identification codes of the target sample, positive sample, negative sample, and comparison sample, such as music ID. and respectively represent the positive sample set and negative sample set. represents the comparison sample set.
[0079] After obtaining the above quadruple sample set then, the feature encoding model can be trained based on each sample in the set. And as recorded in step S210, the initial vectors of each sample need to be obtained accordingly. For any sample x i , its initial vector can be represented as follows:
[0080]
[0081] The embedding representation of the feature, also known as the feature vector corresponding to each feature, K represents the maximum number of features of sample x i , and d represents the embedding dimension, that is, the dimension of the feature vector.
[0082] Accordingly, the following initial vector set of the training samples can be obtained
[0083]
[0084] Next, as recorded in step S220, a feature encoding model to be trained is constructed, and then, according to the feature encoding model to be trained, the initial vectors of the target sample, the positive sample, and the negative sample are respectively feature-encoded to obtain the encoded vectors of each sample, and the process of calculating the correlation vectors between the encoded vectors of each sample and the preset user interest group matrix can be expressed as follows Figure 4 as shown
[0085] See Figure 4 , represent the feature encoding model as f θ , then the process of feature encoding can be expressed as follows
[0086] z i = f θ (x i )
[0087] The initial vectors of the target sample, the positive sample, and the negative sample are calculated through the above formula to obtain the corresponding three encoded vectors, and θ represents the training parameter of the feature encoding model. Specifically, the initial vector of the target sample is feature-encoded to obtain the encoded vector The initial vector of the positive sample is feature-encoded to obtain the encoded vector The initial vector of the negative sample is feature-encoded to obtain the encoded vector The encoded vectors of each sample are the deep representations of their respective samples, and converge multiple features of their respective samples, with the dimension of d e .
[0088] The preset user interest group matrix can be expressed as where E represents the number of user interest groups, and the representation dimension of each user interest group is d e , which is the same as the dimension of the encoded features of each sample. Calculating the correlation vectors between the encoded vectors of each sample and the preset user interest group matrix can be expressed as the following formula
[0089]
[0090] Based on the above formula, the correlation vectors corresponding to each sample can be calculated, expressed as where τ represents the dimension coefficient, which is used to control the aggregation degree of the discrete distribution. Among the correlation vectors corresponding to each sample obtained, the distribution of the association degree between each sample and each user interest group can be known, for example Figure 4As shown, in the corresponding bar chart of relevance, the height of each bar reflects the degree of association between the sample and the corresponding user interest group. Therefore, based on the relevance vector between the encoded vector of each sample and the preset user interest group matrix, it can be well explained which user interest group the sample is favored by, so as to improve the interpretability of the trained feature encoding model. It should also be noted that in the cold start scenario, since the new media object has no user interaction data, it is impossible to be accurate to the user level during the recall process. However, the encoding features that can be accurate to the interest group level can be extracted based on the feature encoding model trained in this application, which proves the rationality of training the model based on the user interest group in this application and also meets the requirements of the recall layer.
[0091] Next, in step S230, the first loss value can be calculated by the following formula = s :
[0092]
[0093] where represents the training set composed of the target sample, positive sample and negative sample, and sim(·) represents the similarity between vectors, such as cosine similarity. represents the relevance vector corresponding to the target sample. represents the relevance vector corresponding to the positive sample. represents the relevance vector corresponding to the negative sample. In an exemplary embodiment, both the encoded vector and the preset user interest group matrix have been normalized by a linear layer. Therefore, the relevance vector is obtained by calculating the cosine similarity. However, since the range of the cosine similarity is relatively low, it is easy to lead to a relatively low upper limit when applying the gradient descent method during the training process. Therefore, an adjustment term e t is added to the above formula to stretch the value range, and t is understood as an adjustment coefficient.
[0094] Next, in step S240, the process of grouping the feature vectors corresponding to the multiple features included in the initial vector of the comparison sample to obtain multiple feature vector groups corresponding to the comparison sample can be seen in Figure 5 as shown, including steps S241 - S242 as follows:
[0095] S241, calculate the corresponding feature association matrix according to the initial vector of the comparison sample;
[0096] S242, sample seed features from the initial vector of the comparison sample, and group the feature vectors corresponding to the multiple features included in the initial vector based on the seed features and the feature association matrix to obtain multiple feature vector groups corresponding to the comparison sample.
[0097] First, it should be noted that for any sample x i The initial vector of which is composed of eigenvectors corresponding to multiple features. For example Figure 6 As shown, each rectangular box represents each feature of sample x i And each feature is associated with a corresponding eigenvector which is not shown in Figure 6 Therefore, the initial vector of the comparison sample also contains eigenvectors corresponding to multiple features respectively. Calculating the feature correlation matrix of the object according to the initial vector of the comparison sample in step S241 is also to calculate the correlation between multiple features in the initial vector of the comparison sample, so as to construct the feature correlation matrix.
[0098] In some exemplary embodiments, if it is assumed that the comparison sample is music, and the initial vector of the comparison sample contains 7 features: name, lyric, audio, language, region, category, and gender. By calculating the correlation between the eigenvectors corresponding to any two features and constructing an initial correlation matrix according to the correlation between the eigenvectors corresponding to any two features, a feature correlation matrix as shown in Figure 7 can be obtained. As shown in Figure 7 , these 7 features can be arranged in sequence to obtain a corresponding feature sequence. Using this feature sequence as both the vertical and horizontal feature arrangements of the feature correlation matrix, and filling the numerical values of the correlation between any two features in the vertical sequence and the horizontal sequence into the corresponding positions, a feature correlation matrix as shown in Figure 7 is obtained. Therefore, according to Figure 7 the correlation degree between the eigenvectors corresponding to any two features in the initial vector of the comparison sample can be clearly obtained.
[0099] Exemplarily, the correlation dCor ij between the eigenvectors corresponding to any two features can be calculated by the following formula:
[0100]
[0101] where i and S respectively represent any two features, that is, the i-th feature and the S-th feature in the initial vector of the comparison sample, e (.) represents the eigenvector corresponding to the feature, dVar(e (i) ) represents the distance variance of the eigenvector e (i) , and dCov(e (i) , e (j) ) represents the distance covariance skew of the eigenvectors corresponding to any two features.
[0102] In some other exemplary embodiments, considering that as the training progresses, the deep representations obtained by the feature encoding model for feature encoding should be dynamically updated, a self-guidance mechanism is introduced to dynamically adjust the correlation between feature vectors. In this scenario, the matrix constructed through the above process, such as Figure 7 shown, is not directly used as the feature correlation matrix, but as the initial correlation matrix. Subsequently, the initial correlation matrix needs to be dynamically updated according to the number of training steps, and the updated matrix is used as the feature correlation matrix.
[0103] Exemplarily, the dynamic adjustment of the correlation between feature vectors is achieved through the following formula:
[0104] C ij = αC ij + (1 - α)dCor ij
[0105] where C represents the feature correlation matrix, C ij represents the matrix element position corresponding to the correlation between the feature vectors corresponding to any two features in the feature correlation matrix, C ij is cumulative, dCor ij represents the correlation between the feature vectors corresponding to any two features in the current step, and α represents the update coefficient, which can be preset to 0.99, for example, to maintain a slow update process. The above formula uses the exponential moving weighted average method to update the feature correlation matrix. For example, it can be set to calculate and update every m steps, which is not limited here.
[0106] In step S242, sampling a seed feature from the initial vectors of the comparison samples means randomly sampling one feature from the multiple features included in the initial vectors as the seed feature. Below, taking the Figure 7 shown feature correlation matrix as an example, the process of grouping the feature vectors corresponding to the multiple features included in the initial vectors based on the seed feature and the feature correlation matrix will be introduced exemplarily.
[0107] For example, if two feature vector groups of the comparison samples are required, half of the features most relevant to the seed feature can be determined from the feature correlation matrix, and the feature vectors corresponding to these half of the features are divided into one feature vector group, and the feature vectors corresponding to the remaining half of the features are divided into another feature vector group accordingly. Referring to the Figure 7 shown feature correlation matrix, if it is assumed that the seed feature is name, the three features of lyric, language, and region have a higher correlation with the seed feature than the three features of audio, category, and gender with the seed feature. Therefore, the feature vectors corresponding to the three features of lyric, language, and region are divided into one feature vector group, and the feature vectors corresponding to the three features of audio, category, and gender are divided into another feature vector group.
[0108] Similarly, if it is necessary to obtain more than two feature vector groups, multiple features can be evenly divided into at least two groups in the same way according to the magnitude of the association with the seed feature shown in the feature association matrix, so as to obtain at least two feature vector groups accordingly. Still referring to Figure 7 the feature association matrix shown, the feature vectors corresponding to the two features of lyric and language can be divided into the first feature vector group, the feature vectors corresponding to the two features of region and audio can be divided into the second feature vector group, and the feature vectors corresponding to the two features of category and gender can be divided into the third feature vector group.
[0109] By analogy, the number of multiple feature vector groups can be expressed as n. A value can be obtained by subtracting 1 from the number of multiple features included in the initial vector. Multiplying this value by 1 / n can obtain the number of features corresponding to each feature vector group. Sort the other features except the seed feature among the multiple features according to the association with the seed feature from large to small or from small to large to obtain a feature sequence. Then, extract features matching the number of features from the feature sequence in turn to obtain multiple feature groups. Divide the feature vectors corresponding to the features included in each feature group into a feature vector group, so as to obtain multiple feature vector groups corresponding to the comparison sample.
[0110] Or in another exemplary embodiment, the number of multiple feature vector groups can still be expressed as n, and the number of multiple features included in the initial vector of the comparison sample can be expressed as m. After sampling the first seed feature, among the other features except the first seed feature in the multiple features included in the initial vector in the feature association matrix, select the first target feature whose association ranking with the first seed feature is topm / n. Divide the feature vector corresponding to the selected first target feature into a feature vector group, and use the features except the first target feature among the multiple features as the first candidate feature group; Next, resample the second seed feature from the first candidate feature group, and then according to the feature association matrix, select the second target feature whose association ranking with the second seed feature is topm / n from the other features except the second seed feature in the first candidate feature group. Divide the feature vector corresponding to the selected second target feature into a feature vector group, and at the same time use the features except the second target feature in the first candidate feature group as the second candidate feature group; Execute the process of dividing the target feature with the association ranking of topm / n from the candidate feature group obtained in the previous round in such a loop, and obtain a corresponding feature vector group and candidate feature group until the total number of obtained feature vector groups is n, which is regarded as the completion of the division of multiple feature vector groups.
[0111] Alternatively, based on the above embodiments, the feature vectors corresponding to the candidate feature groups obtained in the last round can also be divided into a feature vector group. When the total number of obtained feature vector groups reaches n, it is regarded as the completion of the division of multiple feature vector groups. It should be noted that the division method of the feature vector group as exemplified above can be selected according to actual application requirements, and this is not limited herein.
[0112] Next, the process of step S240 performing different feature masking processes on multiple feature vector groups to obtain corresponding multiple masked vectors may include steps S243 - S244 as Figure 8 shown:
[0113] S243, randomly sample multiple data augmentation operators from a preset set of data augmentation operators;
[0114] S244, perform feature masking processes on multiple feature vector groups respectively based on the multiple sampled data augmentation operators to obtain corresponding multiple masked vectors.
[0115] Considering that when using data augmentation operators to construct different video representations of contrast samples, randomly selecting features for masking or adding noise may make the training task too simple. For example, among the 7 features in the foregoing example, if only the language features are masked but the regional features are not masked, the feature encoding model can easily learn deeper-level representations from the regional features. Based on this, the embodiment of the present application proposes the foregoing association grouping mechanism, first obtains multiple feature vector groups of contrast samples, and then processes different feature vector groups with different feature masking methods respectively.
[0116] The preset set of data augmentation operators includes, for example, data augmentation operators such as random masking, span masking, and uniform noise, which are not limited herein. Different data augmentation operators represent different feature masking processing methods. For easy understanding, analogous to the field of computer vision, assume an exemplary sample is Figure 9 the picture shown. The set of data augmentation operators G contains three data augmentation operators: random masking, span masking, and uniform noise. Using these three data augmentation operators to perform random masking, span masking, and uniform noise feature masking processes on the feature vectors of the picture respectively, the Figure 9 processing effects shown can be obtained.
[0117] Correspondingly, the embodiment of the present application randomly samples multiple data augmentation operators from a preset set of data augmentation operators, and performs feature masking processes on multiple feature vector groups respectively based on the multiple sampled data augmentation operators. The corresponding multiple masked vectors reflect different masking effects.
[0118] After that, in step S240, the feature encoding model is further used to perform feature encoding on the multiple mask vectors respectively to obtain multiple encoded vectors corresponding to the comparison samples. It should be noted that the feature encoding model mentioned here shares network parameters with the feature encoding model mentioned in step S220.
[0119] It should also be noted that in some exemplary embodiments, when executing step S240, a mapping network can be further constructed on top of the feature encoding model to project the vectors output by the feature encoding model into a new vector space, thereby obtaining the corresponding encoded vectors. The specific calculation process is as follows:
[0120] h i = g φ (z i ), z i = f θ (x i )
[0121] where f θ (.) and g φ (.) represent the feature encoding model and the mapping network respectively, θ and φ represent the parameters of the feature encoding model and the mapping network respectively, x i represents the representation vector of any sample, and z i represents the result obtained by performing feature encoding processing on the representation vector x i by the feature encoding model.
[0122] Taking the example of obtaining two groups of feature vectors of the comparison sample, different feature encoding processes are performed on these two groups of feature vectors respectively to obtain the corresponding two mask vectors. Then, after performing feature encoding processing according to the feature encoding model, or the feature encoding model plus the mapping network, the corresponding two encoded vectors can be obtained. For example, they are represented as and
[0123] The overall process of obtaining multiple encoded vectors corresponding to the comparison sample as described above can be represented as Figure 10 shown in the process. As Figure 10 shown, exemplarily, the initial vector of the comparison sample is represented as The multiple features in the initial vector are grouped corresponding to the feature vectors to obtain two groups of feature vectors. Then, two data augmentation operators a' and a'' are sampled from the set G of data augmentation operators respectively for each group of feature vectors, and the corresponding groups of feature vectors are subjected to feature masking processing through the sampled data augmentation operators to obtain two mask vectors and Then, according to the feature encoding model f θFeature encode these two mask vectors respectively, and output corresponding vector results and Then, further process through the mapping network stacked on the upper layer of the feature encoder to output the corresponding encoded vectors and
[0124] Next, still taking the two encoded vectors corresponding to the comparison samples as examples, calculate the third similarity value between multiple encoded vectors of the same comparison sample in step S250, and calculate the fourth similarity value between the encoded vectors of different comparison samples. Calculate the second loss value L according to the third similarity value and the fourth similarity value c The process can be expressed by the following formula:
[0125]
[0126] Where represents the comparison sample set, i k and i c are two different comparison samples in this comparison sample set The two encoded vectors corresponding to i k are represented as and i c The two encoded vectors corresponding to are represented as and represents the third similarity between two encoded vectors of the same comparison sample, represents the fourth similarity between the encoded vectors of different comparison samples. Both the third similarity and the fourth similarity can be cosine similarity. τ1 represents the temperature coefficient, which is usually a preset value.
[0127] The training process of the feature encoding model provided in the embodiments of the present application can be regarded as a multi-task learning process. Specifically, the processes shown in steps S220 and S240 described above can be regarded as two different tasks, such as the task of improving recall accuracy and the task of solving the long-tail effect. Therefore, the process of jointly training the feature encoding model based on the first training loss value and the second training loss value in step S260 is also a multi-task learning process.
[0128] Exemplarily, the process of jointly training the feature encoding model based on the first training loss value and the second training loss value in step S260 may include steps S261-S263 as shown in Figure 11 :
[0129] S261, obtain the parameter regularization intensity value corresponding to the feature encoding model;
[0130] S262. Calculate the weighted sum between the parameter regularization strength value and the second loss value, and calculate the total value between this weighted sum and the first loss value.
[0131] S263. Use the total value as the total training loss to train the feature encoding model.
[0132] It should be noted that the parameter regularization strength value corresponding to the feature encoding model mentioned in step S261 refers to a preset model parameter of the feature encoding model. The model parameters of the feature encoding model will change continuously during the training process, but they cannot grow without limit. This preset model parameter is used to limit the change of the model parameters of the feature encoding model within a certain range during the training process. Exemplarily, the total training loss value L can be calculated by the following formula:
[0133]
[0134] where L s represents the first loss value, L c represents the second loss value, represents the parameter regularization strength value, and λ1 and λ2 are the weights corresponding to the parameter regularization strength value and the second loss value respectively, and can also be understood as preset hyperparameters.
[0135] Based on the above training process of the feature encoding model, it can be seen that the embodiments of the present application jointly train the feature encoding model based on the first loss value and the second loss value, which can make the trained feature encoding model not only be able to perceive the user interest group during the feature encoding process, but also be able to cope with the inaccuracy problem caused by the long-tail effect. As a result, the trained feature encoding model can promote the accurate recall of media objects in the process of solving the cold start problem, so as to realize more accurately recommending new media objects to relevant users.
[0136] Next, taking the application of the above-trained feature encoding model in the media object recommendation scenario as an example, the process of how the feature encoding model method proposed in the embodiments of the present application can solve the cold start problem will be described in detail.
[0137] Please refer to Figure 12 shown in Figure 12 which is a flowchart of a media object recommendation method shown in an exemplary embodiment of the present application. The method includes steps S1210 - 1240, which are introduced in detail as follows:
[0138] S1210. Obtain the initial vector of the media object to be recommended, and use the trained feature encoding model to calculate the encoded vector of the media object to be recommended according to this initial vector.
[0139] In the cold start scenario mentioned in this embodiment, the media object to be recommended refers to a new media object, such as newly uploaded music on a music platform. There is usually no user interaction data for the media object to be recommended in the cold start scenario.
[0140] The method for obtaining the initial vector of the media object to be recommended is the same as that for the initial vector of the training samples of the feature encoding model. For specific details, reference can be made to the description in the foregoing embodiment, and details will not be elaborated here.
[0141] The feature encoding model trained by using the feature encoding model training method mentioned in the foregoing embodiment is used to calculate the encoding vector of the media object to be recommended, so that the encoding vector not only contains information about the user interest group, but also contains information for solving the long-tail effect problem. Based on this encoding vector for recall, it can better explain why the recalled media object is similar to the media object to be recommended, and the media object recalled based on this encoding vector under the influence of the long-tail effect can avoid the recall bias caused by insufficient characterization extraction of the long-tail media object. Therefore, the accuracy of the media object recalled based on this encoding feature subsequently is higher, that is, the degree of interest of the user associated with the media object recalled based on this encoding feature in the media object to be recommended is also higher.
[0142] S1220, calculate the correlation vector between the encoding vector and the preset user interest group matrix.
[0143] The preset user interest group matrix mentioned in this embodiment is also the preset user interest group matrix mentioned in the feature encoding training process. The process of calculating the correlation vector between the encoding vector of the media object to be recommended and the preset user interest group matrix is the same as the correlation vector calculation process mentioned in the feature encoding training process. Details will not be elaborated here in this embodiment.
[0144] S1230, query multiple target media objects in the vector pool that are similar to the correlation vector of the media object to be recommended. The vector pool is used to store the correlation vectors of pre-collected media objects, and the correlation vectors of the pre-collected media objects are calculated based on the preset user interest group matrix after obtaining the corresponding encoding vectors through the feature encoding model.
[0145] Querying for multiple target media objects in the vector pool that are approximately similar to the relevance vector of the media object to be recommended reflects the online recall strategy. During online recall, after calculating the relevance vector between the encoded vector of the media object to be recommended and the preset user interest group matrix, a topK set is obtained through approximate nearest neighbor search in the vector pool. This topK set corresponds to multiple target media objects that are approximately similar to the relevance vector of the media object to be recommended. It can be understood that the multiple target media objects recalled in this embodiment are the media objects that are most similar to the media object to be recommended. Therefore, recommending the media object to be recommended to the users associated with the target media objects can ensure that these users are also interested in the media object to be recommended, thereby improving the accuracy of the recommendation.
[0146] S1240. Sort the multiple target media objects according to the degree of approximation between the relevance vector of each target media object and the relevance vector of the media object to be recommended, and recommend the media object to be recommended to the users associated with the target media objects based on the sorting result.
[0147] Since the online recall process calculates the degree of approximation between the relevance vector of each target media object and the relevance vector of the media object to be recommended, the multiple target media objects are sorted according to the order of these degrees of approximation from large to small or from small to large. This is also the function of the sorting model. Finally, the media object to be recommended is recommended to the users associated with the target media objects based on the sorting result, thereby achieving accurate media object recommendation to solve the corresponding cold start problem, such as the music cold start problem when the media object is music.
[0148] Please refer to Figure 13 , Figure 13 is a block diagram of a feature encoding model training device shown in an exemplary embodiment of the present application. The feature encoding model training device includes:
[0149] A construction module 1310, configured to construct a feature encoding model to be trained, and obtain the initial vectors of target samples, positive samples, negative samples, and contrast samples;
[0150] A first extraction module 1320, configured to perform feature encoding on the initial vectors of the target samples, positive samples, and negative samples respectively according to the feature encoding model to obtain the encoded vectors of each sample, and calculate the relevance vectors between the encoded vectors of each sample and the preset user interest group matrix;
[0151] A first calculation module 1330, configured to calculate a first similarity value between the relevance vector corresponding to the target sample and the relevance vector corresponding to the positive sample, and calculate a second similarity value between the relevance vector corresponding to the target sample and the relevance vector corresponding to the negative sample, and calculate a first loss value according to the first similarity value and the second similarity value;
[0152] The second extraction module 1340 is configured to group the feature vectors corresponding to multiple features included in the initial vector of the comparison sample to obtain multiple feature vector groups corresponding to the comparison sample, perform different feature masking processes on the multiple feature vector groups to obtain corresponding multiple masked vectors, and perform feature encoding on the multiple masked vectors respectively according to the feature encoding model to obtain multiple encoded vectors corresponding to the comparison sample;
[0153] The second calculation module 1350 is configured to calculate a third similarity value between multiple encoded vectors of the same comparison sample, and calculate a fourth similarity value between the encoded vectors of different comparison samples, and calculate a second loss value according to the third similarity value and the fourth similarity value;
[0154] The joint training module 1360 is configured to perform joint training on the feature encoding model based on the first loss value and the second loss value.
[0155] In another exemplary embodiment, the second extraction module 1340 includes:
[0156] A matrix calculation unit configured to calculate a corresponding feature correlation matrix according to the initial vector of the comparison sample;
[0157] A grouping unit configured to sample seed features from the initial vector of the comparison sample, and group the feature vectors corresponding to multiple features included in the initial vector based on the seed features and the feature correlation matrix to obtain multiple feature vector groups corresponding to the comparison sample.
[0158] In another exemplary embodiment, the matrix calculation unit includes:
[0159] A construction subunit configured to calculate the correlation between the feature vectors corresponding to any two features included in the initial vector of the comparison sample, and construct an initial correlation matrix according to the correlation between the feature vectors corresponding to any two features;
[0160] An update subunit configured to dynamically update each element in the initial correlation matrix according to the currently corresponding training step number to obtain a feature correlation matrix.
[0161] In another exemplary embodiment, the second extraction module 1340 further includes:
[0162] A sampling unit configured to randomly sample multiple data augmentation operators from a preset set of data augmentation operators;
[0163] A masking unit configured to perform feature masking processes on the multiple feature vector groups respectively based on the multiple sampled data augmentation operators to obtain corresponding multiple masked vectors.
[0164] In another exemplary embodiment, the joint training module 1360 includes:
[0165] An information acquisition unit configured to acquire the parameter regularization intensity value corresponding to the feature encoding model;
[0166] An information calculation unit configured to calculate the weight sum between the parameter regularization intensity value and the second loss value, and calculate the total value between the weight sum and the first loss value;
[0167] A training unit configured to train the feature encoding model using the total value as the total training loss value.
[0168] In another exemplary embodiment, the feature encoding model training apparatus further includes:
[0169] A first information extraction module configured to acquire the media object recommendation data log and the media object metadata, where the media object recommendation data log contains the recall information from the media object to the media object;
[0170] A second information extraction module configured to calculate the recall success rate from the media object to the media object according to the media object recommendation data log, and extract multiple features from the media object metadata;
[0171] A sample extraction module configured to respectively determine the reference media object and the recommended media object in the media object group with a recall success rate greater than the preset global probability as the target sample and the positive sample, perform negative sampling on the determined positive sample to obtain corresponding multiple negative samples, and sample the media object recommendation data log to obtain comparison samples.
[0172] It should be noted that the feature encoding model training apparatus provided in the above embodiment and the feature encoding model training method provided in the above embodiment belong to the same concept. The specific manners in which each module and unit perform operations have been described in detail in the method embodiment, and will not be elaborated here. In practical applications, the feature encoding model training apparatus provided in the above embodiment can, according to needs, allocate the above functions to different functional modules, that is, divide the internal structure of the apparatus into different functional modules to complete all or part of the functions described above, and this will not be limited here either.
[0173] Figure 14 is a block diagram of a media object recommendation apparatus shown in an exemplary embodiment of the present application. The media object recommendation apparatus includes:
[0174] An acquisition module 1410 configured to acquire the initial vector of the media object to be recommended, and calculate the encoded vector of the media object to be recommended according to the initial vector using the feature encoding model trained by the feature encoding model training method mentioned in the foregoing embodiment;
[0175] A computing module 1420, configured to compute a correlation vector between an encoded vector and a preset user interest group matrix;
[0176] A recall module 1430, configured to query, in a vector pool, a plurality of target media objects whose correlation vectors are approximate to the correlation vector of a media object to be recommended. The vector pool is used to store the correlation vectors of pre-collected media objects, and the correlation vectors of the pre-collected media objects are computed based on a preset user interest group matrix after obtaining corresponding encoded vectors through a feature encoding model;
[0177] A recommendation module 1440, configured to rank the plurality of target media objects according to the degree of approximation between the correlation vector of each target media object and the correlation vector of the media object to be recommended, and recommend the media object to be recommended to the users associated with the target media objects based on the ranking result.
[0178] It should also be noted that the media object recommendation apparatus provided in the above embodiments and the media object recommendation method provided in the above embodiments belong to the same concept. The specific manners in which each module and unit perform operations have been described in detail in the method embodiments, and will not be elaborated here. In practical applications, the media object recommendation apparatus provided in the above embodiments may, according to needs, allocate the above functions to different functional modules, that is, divide the internal structure of the apparatus into different functional modules to complete all or part of the functions described above. This is not limited herein either.
[0179] An embodiment of the present disclosure further provides an electronic device, including: one or more processors; a storage device for storing one or more programs, which, when executed by the one or more processors, cause the electronic device to implement the feature encoding model training method or the media object recommendation method provided in each of the above embodiments.
[0180] Figure 15 The structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present disclosure is shown. It should be noted that Figure 15 The computer system 1500 of the electronic device shown is only an example, and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0181] As Figure 15As shown, the computer system includes a Central Processing Unit (CPU) 1501, which can perform various appropriate actions and processes according to the program stored in the Read-Only Memory (ROM) 1502 or the program loaded from the storage section 1508 into the Random Access Memory (RAM) 1503, such as executing the method described in the above embodiments. In the RAM 1503, various programs and data required for system operation are also stored. The CPU 1501, ROM 1502, and RAM 1503 are connected to each other via a bus 1504. An Input / Output (I / O) interface 1505 is also connected to the bus 1504.
[0182] The following components are connected to the I / O interface 1505: an input section 1506 including a keyboard, a mouse, etc.; an output section 1507 including, for example, a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), etc. and a speaker, etc.; a storage section 1508 including a hard disk, etc.; and a communication section 1509 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1509 performs communication processing via a network such as the Internet. A drive 1510 is also connected to the I / O interface 1505 as needed. A removable medium 1511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1510 as needed so that a computer program read from it can be installed into the storage section 1508 as needed.
[0183] Specifically, according to the embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments of the present disclosure include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 1509, and / or installed from the removable medium 1511. When the computer program is executed by the Central Processing Unit (CPU) 1501, various functions defined in the system of the present disclosure are executed.
[0184] It should be noted that the computer-readable medium shown in the embodiments of the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure.
[0185] On the other hand, the present disclosure also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the feature encoding model training method or the media object recommendation method as described above. The computer-readable storage medium may be included in the electronic device described in the above embodiments, or may exist alone without being assembled into the electronic device.
[0186] The above content is only a preferred exemplary embodiment of the present disclosure and is not used to limit the implementation of the present disclosure. Those of ordinary skill in the art can easily make corresponding adaptations or modifications according to the main concepts and spirits of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope required by the claims.
Claims
1. A method for training a feature encoding model, characterized in that The method includes: Obtaining media object recommendation data logs and media object metadata, where the media object recommendation data logs contain recall information from media objects to media objects; Calculating the recall success rate from media objects to media objects according to the media object recommendation data logs, and extracting multiple features from the media object metadata; Determining the reference media object and the recommended media object in the media object group with a recall success rate greater than the preset global probability as the target sample and the positive sample respectively, performing negative sampling on the determined positive sample to obtain corresponding multiple negative samples, and sampling the media object recommendation data logs to obtain comparison samples; Constructing a feature encoding model to be trained, and obtaining the initial vectors of the target sample, the positive sample, the negative sample, and the comparison samples; Performing feature encoding on the initial vectors of the target sample, the positive sample, and the negative sample respectively according to the feature encoding model to obtain the encoded vectors of each sample, and calculating the correlation vectors between the encoded vectors of each sample and the preset user interest group matrix; the preset user interest group matrix is a matrix extracted from the user data of multiple user interest groups collected in advance; Calculating a first similarity value between the correlation vector corresponding to the target sample and the correlation vector corresponding to the positive sample, and calculating a second similarity value between the correlation vector corresponding to the target sample and the correlation vector corresponding to the negative sample, and calculating a first loss value according to the first similarity value and the second similarity value; Grouping the feature vectors corresponding to the multiple features included in the initial vector of the comparison sample to obtain multiple feature vector groups corresponding to the comparison sample, and performing different feature masking processes on the multiple feature vector groups to obtain corresponding multiple masked vectors, and performing feature encoding on the multiple masked vectors respectively according to the feature encoding model to obtain multiple encoded vectors corresponding to the comparison sample; Calculating a third similarity value between the multiple encoded vectors of the same comparison sample, and calculating a fourth similarity value between the encoded vectors of different comparison samples, and calculating a second loss value according to the third similarity value and the fourth similarity value; Jointly training the feature encoding model based on the first loss value and the second loss value.
2. The method according to claim 1, characterized in that The grouping of the feature vectors corresponding to the multiple features included in the initial vector of the comparison sample to obtain multiple feature vector groups corresponding to the comparison sample includes: Calculating a corresponding feature correlation matrix according to the initial vector of the comparison sample; Sampling seed features from the initial vector of the comparison sample, and grouping the feature vectors corresponding to the multiple features included in the initial vector based on the seed features and the feature correlation matrix to obtain multiple feature vector groups corresponding to the comparison sample.
3. The method according to claim 2, wherein The calculating a corresponding feature correlation matrix according to the initial vector of the comparison sample includes: Calculate the correlation between the feature vectors corresponding to any two features included in the initial vector of the comparison sample, and construct an initial correlation matrix according to the correlation between the feature vectors corresponding to any two features; Dynamically update each element in the initial correlation matrix according to the currently corresponding training step number to obtain the feature correlation matrix.
4. The method according to claim 1, wherein The performing different feature masking processes on the multiple feature vector groups to obtain corresponding multiple masked vectors includes: Randomly sample a plurality of data augmentation operators from a preset data augmentation operator set; Perform feature masking processes on the multiple feature vector groups respectively based on the sampled multiple data augmentation operators to obtain corresponding multiple masked vectors.
5. The method according to claim 1, wherein The jointly training the feature encoding model based on the first loss value and the second loss value includes: Obtain the parameter regularization strength value corresponding to the feature encoding model; Calculate the weight sum between the parameter regularization strength value and the second loss value, and calculate the total sum value between the weight sum and the first loss value; Use the total sum value as the total training loss value to train the feature encoding model.
6. A method for recommending media objects, characterized in that, The method includes: Obtain the initial vector of the media object to be recommended, and calculate the encoding vector of the media object to be recommended according to the initial vector by using the feature encoding model trained by the method described in any one of claims 1-5; Calculate the correlation vector between the encoding vector and a preset user interest group matrix; Query, in a vector pool, a plurality of target media objects that are approximate to the correlation vector of the media object to be recommended, where the vector pool is used to store the correlation vectors of pre-collected media objects, and the correlation vectors of the pre-collected media objects are calculated based on the preset user interest group matrix after obtaining the corresponding encoding vectors through the feature encoding model; Sort the plurality of target media objects according to the approximation degree value between the correlation vector of each target media object and the correlation vector of the media object to be recommended, and recommend the media object to be recommended to the users associated with the target media objects based on the sorting result.
7. A feature encoding model training device, characterized in that, The device includes: A first information extraction module configured to obtain a media object recommendation data log and media object metadata, where the media object recommendation data log contains recall information from media object to media object; A second information extraction module configured to calculate the recall success rate from media object to media object according to the media object recommendation data log, and extract a plurality of features from the media object metadata; A sample extraction module configured to respectively determine the reference media object and the recommended media object in the media object group with a recall success rate greater than a preset global probability as the target sample and the positive sample, perform negative sampling on the determined positive sample to obtain corresponding multiple negative samples, and sample the media object recommendation data log to obtain a comparison sample; A construction module configured to construct a feature encoding model to be trained, and obtain the initial vectors of the target sample, the positive sample, the negative sample, and the comparison sample; The first extraction module is configured to perform feature encoding on the initial vectors of the target sample, the positive sample, and the negative sample respectively according to the feature encoding model, obtain the encoded vectors of each sample, and calculate the correlation vectors between the encoded vectors of each sample and the preset user interest group matrix respectively; the preset user interest group matrix is a matrix collected in advance from the user data of multiple user interest groups and extracted from these user data. The first calculation module is configured to calculate the first similarity value between the correlation vector corresponding to the target sample and the correlation vector corresponding to the positive sample, and calculate the second similarity value between the correlation vector corresponding to the target sample and the correlation vector corresponding to the negative sample, and calculate the first loss value according to the first similarity value and the second similarity value. The second extraction module is configured to group the feature vectors corresponding to the multiple features included in the initial vector of the comparison sample to obtain multiple feature vector groups corresponding to the comparison sample, and perform different feature masking processes on the multiple feature vector groups to obtain corresponding multiple masked vectors, and perform feature encoding on the multiple masked vectors respectively according to the feature encoding model to obtain multiple encoded vectors corresponding to the comparison sample. The second calculation module is configured to calculate the third similarity value between the multiple encoded vectors of the same comparison sample, and calculate the fourth similarity value between the encoded vectors of different comparison samples, and calculate the second loss value according to the third similarity value and the fourth similarity value. The joint training module is configured to perform joint training on the feature encoding model based on the first loss value and the second loss value.
8. The device according to claim 7, characterized in that, The second extraction module includes: The matrix calculation unit is configured to calculate the corresponding feature correlation matrix according to the initial vector of the comparison sample. The grouping unit is configured to sample seed features from the initial vector of the comparison sample, and group the feature vectors corresponding to the multiple features included in the initial vector based on the seed features and the feature correlation matrix to obtain multiple feature vector groups corresponding to the comparison sample.
9. The device according to claim 8, wherein The matrix calculation unit includes: The construction subunit is configured to calculate the correlation between the feature vectors corresponding to any two features included in the initial vector of the comparison sample, and construct an initial correlation matrix according to the correlation between the feature vectors corresponding to any two features. The update subunit is configured to dynamically update each element in the initial correlation matrix according to the current corresponding training step number to obtain the feature correlation matrix.
10. The device according to claim 7, characterized in that, The second extraction module further includes: The sampling unit is configured to randomly sample multiple data augmentation operators from the preset data augmentation operator set. The masking unit is configured to perform feature masking processes on the multiple feature vector groups respectively based on the multiple sampled data augmentation operators to obtain corresponding multiple masked vectors.
11. The device according to claim 7, characterized in that, The joint training module includes: The information acquisition unit is configured to acquire the parameter regularization intensity value corresponding to the feature encoding model. An information calculation unit, configured to calculate a weighted sum between the parameter regularization strength value and the second loss value, and calculate a total value between the weighted sum and the first loss value; A training unit, configured to use the total value as the total training loss value to train the feature encoding model.
12. A media object recommendation device, characterized in that, The apparatus includes: An acquisition module, configured to acquire an initial vector of a media object to be recommended, and calculate an encoded vector of the media object to be recommended according to the initial vector by using a feature encoding model trained by the method according to any one of claims 1-5; A calculation module, configured to calculate a correlation vector between the encoded vector and a preset user interest group matrix; A recall module, configured to query, in a vector pool, a plurality of target media objects similar to the correlation vector of the media object to be recommended, where the vector pool is used to store correlation vectors of pre-collected media objects, and the correlation vectors of the pre-collected media objects are calculated based on the preset user interest group matrix after obtaining corresponding encoded vectors through the feature encoding model; A recommendation module, configured to sort the plurality of target media objects according to a similarity value between the correlation vector of each target media object and the correlation vector of the media object to be recommended, and recommend the media object to be recommended to a user associated with the target media object based on the sorting result.
13. An electronic device, characterized in that, including: a processor; and a memory, where computer-readable instructions are stored on the memory, and when the computer-readable instructions are executed by the processor, the method according to any one of claims 1-6 is implemented.
14. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the method according to any one of 1-6 is implemented.
Citation Information
Patent Citations
Information recommendation method and device, electronic equipment and storage medium
CN112395499A
Interest embedding vectors
US20180253496A1