A model training method, information recommendation method and device
By obtaining user historical operation information, using a self-supervised comparison algorithm to train the recommendation model, optimize the feature similarity of similar information and the feature differences of different types of information, the problem of high image similarity in the recommendation model is solved, and the accuracy and user experience of recommendation are improved.
Patent Information
- Application Number
- CN202210473098.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-29
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2042-04-29
AI Technical Summary
When the existing recommendation models recommend images to users, they are prone to highly similar situations between images, resulting in poor recommendation accuracy and inability to meet users' expectations.
By obtaining user historical operation information, determining the browsed and unbrowsed multimedia information, and using a self-supervised comparison algorithm to train the recommended model. The optimization goal is to minimize the feature deviation of similar information and maximize the feature differences between different types of information, including the joint optimization of feature sub-models and auxiliary sub-models.
It improves the accuracy of the recommendation model for user preferences, ensures that the recommended multimedia information is similar to the user's operation information, and is different from the unoperated information, improving the accuracy and user experience of recommendations.
Smart Images

Figure CN114860967B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and in particular to a model training method, an information recommendation method, and a device. Background Art
[0002] With the continuous development of electronic technology and network technology, in order to provide more convenience to users' daily lives, images are usually recommended to users based on image features and user preference features.
[0003] At present, when recommending images to users through existing recommendation models, there may be a situation where the images are highly similar, so that some of the recommended images do not meet the user's expectations, resulting in poor accuracy of the images recommended by the recommendation model, which in turn reduces the user experience.
[0004] Therefore, how to improve the accuracy of images recommended to users is an urgent problem to be solved. Summary of the Invention
[0005] This specification provides a model training method, information recommendation method, device, storage medium and electronic device to partially solve the above-mentioned problems existing in the prior art.
[0006] This manual adopts the following technical solutions:
[0007] This manual provides a model training method, including:
[0008] Obtaining the user's operation information on each recommended multimedia message within a set historical time period;
[0009] determining, from the recommended multimedia information, recommended multimedia information browsed by the user as first multimedia information, and determining recommended multimedia information not browsed by the user as second multimedia information, according to the operation information;
[0010] Inputting each of the first multimedia information and each of the second multimedia information into a recommendation model to be trained, determining a feature corresponding to each of the first multimedia information as a first feature, and a feature corresponding to each of the second multimedia information as a second feature;
[0011] For each first multimedia information, the recommendation model is trained with the optimization goals of minimizing the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and maximizing the deviation between the first feature corresponding to the first multimedia information and the second feature.
[0012] Optionally, based on the operation information, determining, from the recommended multimedia information, recommended multimedia information browsed by the user as each first multimedia information, and determining recommended multimedia information not browsed by the user as each second multimedia information, specifically includes:
[0013] According to the operation information, the recommended multimedia information is randomly transformed to obtain first transformed multimedia information and second transformed multimedia information;
[0014] Determining, from the first transformed multimedia information, recommended multimedia information browsed by the user as each first multimedia information, and determining recommended multimedia information not browsed by the user as each second multimedia information;
[0015] From the second transformed multimedia information, the recommended multimedia information browsed by the user is determined as each third multimedia information, and the recommended multimedia information not browsed by the user is determined as each fourth multimedia information.
[0016] Optionally, inputting each of the first multimedia information and each of the second multimedia information into a recommendation model to be trained, and determining a feature corresponding to each of the first multimedia information as a first feature, and a feature corresponding to each of the second multimedia information as a second feature, specifically includes:
[0017] Input the first multimedia information, the second multimedia information, the third multimedia information and the fourth multimedia information into the recommendation model to be trained, and determine the feature corresponding to each first multimedia information in the first transformed multimedia information as the first feature, the feature corresponding to each second multimedia information in the first transformed multimedia information as the second feature, the feature corresponding to each third multimedia information in the second transformed multimedia information as the third feature, and the feature corresponding to each fourth multimedia information in the second transformed multimedia information as the fourth feature.
[0018] Optionally, for each first multimedia information, the recommendation model is trained with minimizing the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and maximizing the deviation between the first feature corresponding to the first multimedia information and the second feature as optimization goals, specifically including:
[0019] For each first multimedia information, the recommendation model is trained with the optimization objectives of minimizing the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and maximizing the deviation between the first feature corresponding to the first multimedia information and the second feature; and for each third multimedia information, the recommendation model is trained with the optimization objectives of minimizing the deviation between the third feature corresponding to the third multimedia information and the third features corresponding to other third multimedia information, and maximizing the deviation between the third feature corresponding to the third multimedia information and the fourth feature.
[0020] Optionally, the recommendation model includes: a feature sub-model and an auxiliary sub-model;
[0021] Inputting each of the first multimedia information and each of the second multimedia information into a recommendation model to be trained, and determining a feature corresponding to each of the first multimedia information as a first feature, and a feature corresponding to each of the second multimedia information as a second feature, specifically including:
[0022] Inputting each of the first transformed multimedia information and each of the second transformed multimedia information into a feature sub-model to be trained, determining a feature corresponding to each of the first transformed multimedia information as a first model feature, and determining a feature corresponding to each of the second transformed multimedia information as a second model feature;
[0023] The first transformed multimedia information and the second transformed multimedia information are input into the auxiliary sub-model to be trained, and the features corresponding to each first transformed multimedia information are determined as the first auxiliary features, and the features corresponding to each second transformed multimedia information are determined as the second auxiliary features.
[0024] Optionally, for each first multimedia information, the recommendation model is trained with minimizing the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and maximizing the deviation between the first feature corresponding to the first multimedia information and the second feature as optimization goals, specifically including:
[0025] For each piece of recommended multimedia information, determining a similarity between the first model feature corresponding to the recommended multimedia information and the second auxiliary feature corresponding to the recommended multimedia information as a first similarity, and determining a difference between the first model feature corresponding to the recommended multimedia information and the second auxiliary feature corresponding to other pieces of recommended multimedia information as a first difference;
[0026] For each first multimedia information, the recommendation model is trained with the optimization goals of minimizing the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and maximizing the deviation between the first feature corresponding to the first multimedia information and the second feature, and maximizing the first similarity and the first difference as the optimization goals.
[0027] Optionally, for each first multimedia information, minimizing the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and maximizing the deviation between the first feature corresponding to the first multimedia information and the second feature as optimization goals, and maximizing the first similarity and the first difference as optimization goals, the recommendation model is trained, specifically including:
[0028] For each piece of recommended multimedia information, determining a similarity between the second model feature corresponding to the recommended multimedia information and the first auxiliary feature corresponding to the recommended multimedia information as a second similarity, and determining a difference between the second model feature corresponding to the recommended multimedia information and the first auxiliary feature corresponding to other pieces of recommended multimedia information as a second difference;
[0029] For each first multimedia information, the optimization goal is to minimize the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and to maximize the deviation between the first feature corresponding to the first multimedia information and the second feature, and the recommendation model is trained with the optimization goal of maximizing the first similarity, the first difference, the second similarity and the second difference.
[0030] Optionally, for each first multimedia information, minimizing the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and maximizing the deviation between the first feature corresponding to the first multimedia information and the second feature are optimization goals, and maximizing the first similarity, the first difference, the second similarity, and the second difference are optimization goals, the recommendation model is trained, specifically including:
[0031] For each first multimedia message, a deviation between a first feature corresponding to the first multimedia message and a first feature corresponding to another first multimedia message is used as a first deviation, and a deviation between the first feature corresponding to the first multimedia message and the second feature is used as a second deviation;
[0032] For each round of training, determining a weight corresponding to the first deviation and a weight corresponding to the second deviation according to the round of training, as the round weight, wherein a larger round corresponding to the training round has a larger round weight;
[0033] The recommendation model is trained according to the first deviation, the second deviation, the first similarity, the first difference, the second similarity, the second difference, and the round weight.
[0034] This specification provides an information recommendation method, including:
[0035] In response to an information acquisition request from a target user, acquiring candidate recommended multimedia information and a preference representation of the target user;
[0036] For each candidate recommended multimedia information, input the candidate recommended multimedia information and the target user's preference representation into a pre-trained recommendation model to predict the click-through rate corresponding to the candidate recommended multimedia information, wherein the recommendation model is trained using the above-mentioned model training method;
[0037] According to the click rates corresponding to the candidate recommended multimedia information, the recommended multimedia information to be recommended to the target user is determined as the target recommended multimedia information, and the target recommended multimedia information is recommended to the target user.
[0038] This specification provides a model training device, including:
[0039] An acquisition module is used to obtain the user's operation information on each recommended multimedia information within a set historical time period;
[0040] a determination module configured to determine, based on the operation information, from the recommended multimedia information, the recommended multimedia information browsed by the user as each first multimedia information, and to determine the recommended multimedia information not browsed by the user as each second multimedia information;
[0041] an input module, configured to input the first multimedia information and the second multimedia information into a recommendation model to be trained, and determine a feature corresponding to each first multimedia information as a first feature, and a feature corresponding to each second multimedia information as a second feature;
[0042] A training module is used to train the recommendation model for each first multimedia information with the optimization objectives of minimizing the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and maximizing the deviation between the first feature corresponding to the first multimedia information and the second feature.
[0043] This specification provides an information recommendation device, including:
[0044] A response module, configured to respond to a target user's information acquisition request and acquire candidate recommended multimedia information and a preference representation of the target user;
[0045] A prediction module is configured to input each candidate recommended multimedia information and the target user's preference representation into a pre-trained recommendation model to predict a click-through rate corresponding to the candidate recommended multimedia information, wherein the recommendation model is trained using the above-mentioned model training method;
[0046] The recommendation module is used to determine the recommended multimedia information to be recommended to the target user according to the click rate corresponding to each candidate recommended multimedia information, as the target recommended multimedia information, and recommend the target recommended multimedia information to the target user.
[0047] This specification provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned model training method and information recommendation method.
[0048] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the above-mentioned model training method and information recommendation method are implemented.
[0049] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects:
[0050] In the model training method provided in this specification, the user's operation information for each recommended multimedia information within a historical set time period is obtained. Secondly, based on the operation information, the recommended multimedia information browsed by the user is determined from each recommended multimedia information as each first multimedia information, and the recommended multimedia information not browsed by the user is determined as each second multimedia information. Then, each first multimedia information and each second multimedia information is input into the recommendation model to be trained, and the feature corresponding to each first multimedia information is determined as the first feature, and the feature corresponding to each second multimedia information is determined as the second feature. Finally, for each first multimedia information, the recommendation model is trained with the optimization goal of minimizing the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and maximizing the deviation between the first feature corresponding to the first multimedia information and the second feature.
[0051] As can be seen from the above method, this method can determine the similarity between the recommended multimedia information browsed by the user through the recommended multimedia information browsed by the user, and determine the difference between the recommended multimedia information browsed by the user and the recommended multimedia information not browsed by the user through the recommended multimedia information browsed by the user and the recommended multimedia information not browsed by the user. Finally, the recommendation model is trained with the optimization goal of minimizing the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and maximizing the deviation between the first feature corresponding to the first multimedia information and the second feature. This can ensure the correlation between the first features of the recommended multimedia information browsed by the user, and the difference between the first feature of the recommended multimedia information browsed by the user and the second feature of the recommended multimedia information not browsed by the user, thereby improving the accuracy of the multimedia information recommended to the user. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The exemplary embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation of this specification. In the drawings:
[0053] Figure 1 A flowchart of the model training method provided in the embodiments of this specification;
[0054] Figure 2 A schematic diagram of the model structure provided in the embodiments of this specification;
[0055] Figure 3 A flowchart of a method for providing information recommendation for the embodiments of this specification;
[0056] Figure 4 A schematic diagram of the structure of the device for model training provided in the embodiments of this specification;
[0057] Figure 5 A schematic diagram of the structure of a device recommended for providing information in accordance with the embodiments of this specification;
[0058] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this specification. DETAILED DESCRIPTION
[0059] To make the objectives, technical solutions, and advantages of this specification more clear, the following will clearly and completely describe the technical solutions of this specification in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this specification.
[0060] The technical solutions provided by the embodiments of this specification are described in detail below with reference to the accompanying drawings.
[0061] In the embodiment of this specification, when recommending information in response to the target user's information acquisition request, it is necessary to rely on a pre-trained recommendation model. Therefore, the following will first introduce the process of training the recommendation model. Figure 1 shown.
[0062] Figure 1 The flow chart of the model training method provided in the embodiment of this specification specifically includes the following steps:
[0063] S100: Acquire the user's operation information on each recommended multimedia information within a set historical time period.
[0064] In the embodiments of this specification, the execution entity for training the recommendation model can be a server or an electronic device such as a desktop computer. For the sake of ease of description, the following only uses the server as the execution entity to illustrate the training method of the recommendation model provided in this specification.
[0065] In real-world applications, users typically interact with recommended multimedia information based on a specific goal over a period of time. For example, if a user wants to buy a pet cat, they may click on images related to pet cats over a period of time. Based on this, the server can determine a set duration during which the user's goal remains unchanged. In other words, the recommended multimedia information clicked by the user during this period is of the same category.
[0066] In the embodiment of this specification, the server can obtain the user's operation information for each recommended multimedia information within a historical set time period. The user mentioned here can refer to any user who has historical operation information online in the past. In other words, the server can set a time limit and select the operation information within the set time limit from each recommended multimedia information of any user. For example, if the historical set time period is determined to be five minutes, the server can select five minutes of continuous operation information from each recommended multimedia information based on the user's historical operation information. The specific formula is as follows:
[0067] Seq i ={(It , C t )|t=0,1,2,3...T}
[0068] In the above formula, Seq i It can be used to represent the operation information of the i-th user for each recommended multimedia information within the historical set time. T can be used to represent the set time. t It can be used to represent the recommended multimedia information corresponding to time t. t Can be used to represent user browsing behavior, C t If it is 1, it means that the user browses the recommended multimedia information displayed to the user. t A value of 0 may indicate that the user has not browsed the recommended multimedia information displayed to the user. As can be seen from the above formula, the server selects continuous operation information of a set duration from the user's historical operation information, and arranges the operation information in reverse chronological order.
[0069] Based on this, the server can obtain the operation information of several users for each recommended multimedia information within the historical set time period, and use the obtained operation information of several users for each recommended multimedia information within the historical set time period as the user data set. The specific formula is as follows:
[0070] D={Seq i |i=0,1,2,3...N}
[0071] In the above formula, D can be used to represent the set of user operation records for each recommended multimedia message within a set period of time, which serves as the user dataset. During subsequent recommendation model training, the server can randomly select user operation records from the user dataset within the set period of time to train the recommendation model.
[0072] The multimedia information mentioned here includes images, videos, and other multimedia information. Recommended multimedia information can come from a variety of business scenarios. For example, search scenarios used to meet users' active search needs can include food search scenarios, travel search scenarios, and product search scenarios. Another example is recommendation scenarios used to meet users' passive personalized recommendation needs, such as food recommendation scenarios, travel recommendation scenarios, and product recommendation scenarios.
[0073] Operation information may include historical click information, historical order information, historical like information, etc. For example, it may include images of food items clicked by a user in a food search scenario, or images of products ordered by a user in a product recommendation scenario.
[0074] It should be noted that before the server obtains the user's operation information for each recommended multimedia message, it must first send an authorization request to the user. If the user does not agree to the authorization, the server cannot obtain the user's operation information for each recommended multimedia message. If the user agrees to the authorization, the server can delete the obtained user's operation information for each recommended multimedia message when the user cancels the service using this method.
[0075] S102: According to the operation information, determine the recommended multimedia information browsed by the user from the recommended multimedia information as each first multimedia information, and determine the recommended multimedia information not browsed by the user as each second multimedia information.
[0076] In the embodiment of this specification, the server can determine the recommended multimedia information browsed by the user from the recommended multimedia information as each first multimedia information, and determine the recommended multimedia information not browsed by the user as each second multimedia information based on the operation information.
[0077] In the embodiments of this specification, the recommendation model uses a self-supervised comparison algorithm. Positive samples in a self-supervised comparison algorithm are typically based on different transformations of the same image, while negative samples are typically other images. Based on this, the server can randomly transform each image in the user dataset to obtain the corresponding positive and negative samples.
[0078] In the embodiments of this specification, the server may randomly transform each recommended multimedia information based on the operation information to obtain each first transformed multimedia information and each second transformed multimedia information. The random transformation mentioned here includes random cropping, random color transformation, random Gaussian blurring, scaling to a fixed size, and other multimedia information transformation methods.
[0079] Then, the server may determine, from each first transformed multimedia information, the recommended multimedia information browsed by the user as each first multimedia information, and determine the recommended multimedia information not browsed by the user as each second multimedia information.
[0080] Similarly, the server can determine the recommended multimedia information browsed by the user from each second transformed multimedia information as each third multimedia information, and determine the recommended multimedia information not browsed by the user as each fourth multimedia information.
[0081] Among them, the self-supervised contrast algorithm used in the recommendation model applied by this method can be the Momentum Contrast Algorithm (MOCO), the Simple Framework for Contrastive Learning of Visual Representations (SimCLR), etc. This specification does not limit the specific form of the self-supervised contrast algorithm used by the recommendation model.
[0082] S104: Input each of the first multimedia information and each of the second multimedia information into the recommendation model to be trained, determine a feature corresponding to each of the first multimedia information as a first feature, and determine a feature corresponding to each of the second multimedia information as a second feature.
[0083] S106: For each first multimedia information, the recommendation model is trained with the optimization objectives of minimizing the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and maximizing the deviation between the first feature corresponding to the first multimedia information and the second feature.
[0084] In actual applications, when users browse recommended multimedia information, they usually have certain target needs to operate on the recommended multimedia information. Therefore, the server can assume that the categories of the recommended multimedia information operated by users within a period of time are consistent. In other words, the semantic correlation between the recommended multimedia information operated by users within a set time period is strong and has the same characteristics, while the semantic correlation between the recommended multimedia information operated by users within a set time period and the recommended multimedia information not operated by users within a set time period is weak and does not have the same characteristics. Based on this, the server can train the recommendation model by shortening the distance between the recommended multimedia information operated by users and increasing the distance between the recommended multimedia information operated by users and the recommended multimedia information not operated by users. In this way, in the process of actual application, the recommendation model can better characterize the characteristics of the recommended multimedia information.
[0085] In an embodiment of the present specification, the server may input each first multimedia information and each second multimedia information into a recommendation model to be trained, determine a feature corresponding to each first multimedia information as a first feature, and a feature corresponding to each second multimedia information as a second feature.
[0086] Then, for each first multimedia message, the server can train the recommendation model with the optimization goal of minimizing the deviation between the first feature corresponding to the first multimedia message and the first features corresponding to other first multimedia messages, and maximizing the deviation between the first feature corresponding to the first multimedia message and the second feature. The specific formula is as follows:
[0087]
[0088] In the above formula, f q1i It can be used to characterize the first feature corresponding to the i-th first multimedia information. It can be used to characterize the first feature corresponding to the j-th other first multimedia information. It can be used to represent the second feature corresponding to the j-th second multimedia information. N can be used to represent the number of first multimedia information. M can be used to represent the number of second multimedia information. The first average value of the distance between the first feature corresponding to each first multimedia message and the first features corresponding to other first multimedia messages may be used to characterize, for each first multimedia message, the first average value, wherein the smaller the first average value, the greater the similarity between the first feature corresponding to the first multimedia message and the first features corresponding to the other first multimedia messages. The second average value of the distance between the first feature corresponding to each first multimedia message and the second feature corresponding to the second multimedia message can be used to represent the difference between the first feature corresponding to the first multimedia message and the second feature corresponding to the second multimedia message. The larger the second average value, the greater the difference between the first feature corresponding to the first multimedia message and the second feature corresponding to the second multimedia message.
[0089] because, If the value of the training recommendation model is negative, the server cannot determine the value of the training recommendation model well. Based on this, the server sets the parameter m so that the difference between the first average value and the second average value is close to the parameter m. If the difference between the first average value and the second average value exceeds the parameter m, The result of is less than 0, so the formula result is determined to be 0, and the recommendation model training is determined to be completed. The parameter m can be manually set based on expert experience.
[0090] It can be seen from the above formula that this formula can make the features of recommended multimedia information of the same category determined by the recommendation model more similar, and make the features of recommended multimedia information of different categories determined by the recommendation model more different.
[0091] In practical applications, if the recommendation model can extract the features of the two transformed images obtained after different random transformations of the same image, then the image features of the two transformed images corresponding to the same image should be basically the same. Based on this, for each transformed image, the server can train the recommendation model by shortening the distance between the recommended multimedia information operated by the user in the transformed image, and increasing the distance between the recommended multimedia information operated by the user in the transformed image and the recommended multimedia information not operated by the user in the transformed image, thereby making the recommendation model more optimized.
[0092] In an embodiment of the present specification, the server can input each first multimedia information, each second multimedia information, each third multimedia information and each fourth multimedia information into the recommendation model to be trained, determine the feature corresponding to each first multimedia information in each first transformed multimedia information as the first feature, the feature corresponding to each second multimedia information in each first transformed multimedia information as the second feature, the feature corresponding to each third multimedia information in each second transformed multimedia information as the third feature, and the feature corresponding to each fourth multimedia information in each second transformed multimedia information as the fourth feature.
[0093] Similarly, the server can train the recommendation model for each third multimedia information by minimizing the deviation between the third feature corresponding to the third multimedia information and the third features corresponding to other third multimedia information, and by maximizing the deviation between the third feature corresponding to the third multimedia information and the fourth feature as the optimization goal. The specific formula is as follows:
[0094]
[0095] In the above formula, f q2i It can be used to characterize the third feature corresponding to the i-th third multimedia information. It can be used to characterize the third feature corresponding to the j-th other third multimedia information. It can be used to represent the fourth feature corresponding to the j-th fourth multimedia information. N can be used to represent the number of third multimedia information. M can be used to represent the number of fourth multimedia information. The third average value of the distance between the third feature corresponding to each third multimedia information and the third features corresponding to other third multimedia information can be used to characterize, for each third multimedia information, determining the third average value. The smaller the third average value, the greater the similarity between the third feature corresponding to the third multimedia information and the third features corresponding to the other third multimedia information. The fourth average value may be used to characterize, for each third multimedia message, a distance between a third feature corresponding to the third multimedia message and a fourth feature corresponding to the fourth multimedia message. A larger fourth average value indicates a greater difference between the third feature corresponding to the third multimedia message and the fourth feature corresponding to the fourth multimedia message.
[0096] Furthermore, the server may train the recommendation model for each first multimedia message with the optimization goal of minimizing the deviation between the first feature corresponding to the first multimedia message and the first features corresponding to other first multimedia messages, and maximizing the deviation between the first feature corresponding to the first multimedia message and the second feature, and for each third multimedia message with the optimization goal of minimizing the deviation between the third feature corresponding to the third multimedia message and the third features corresponding to other third multimedia messages, and maximizing the deviation between the third feature corresponding to the third multimedia message and the fourth feature. The specific formula is as follows:
[0097]
[0098] In the above formula, ∑ j∈+ L tri (f q1j , f q1 ) can be used to characterize to minimize the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and to maximize the deviation between the first feature corresponding to the first multimedia information and the second feature. j∈+ L tri (f q2j , f q2 ) can be used to characterize to minimize the deviation between the third feature corresponding to the third multimedia information and the third features corresponding to other third multimedia information, and to maximize the deviation between the third feature corresponding to the third multimedia information and the fourth feature.
[0099] In practical applications, after the same image undergoes different random transformations, the image features corresponding to the several transformed images obtained in models with different parameters should have similar relationships. Based on this, the server can train the recommendation model by reducing the distance between the image features corresponding to several transformed images of the same image and increasing the distance between the image features corresponding to different images, so that the image features determined by the recommendation model can better represent the characteristics of the image.
[0100] In the embodiment of this specification, the recommendation model includes: a feature sub-model and an auxiliary sub-model. The feature sub-model and the auxiliary sub-model have the same structure, but different model parameters.
[0101] In an embodiment of the present specification, the server can input each first transformed multimedia information and each second transformed multimedia information into a feature sub-model to be trained, determine the features corresponding to each first transformed multimedia information as the first model features, and determine the features corresponding to each second transformed multimedia information as the second model features.
[0102] Similarly, the server can input each first transformed multimedia information and each second transformed multimedia information into the auxiliary sub-model to be trained, determine the features corresponding to each first transformed multimedia information as the first auxiliary features, and determine the features corresponding to each second transformed multimedia information as the second auxiliary features.
[0103] Furthermore, the server can determine, for each recommended multimedia information, the similarity between the first model feature corresponding to the recommended multimedia information and the second auxiliary feature corresponding to the recommended multimedia information as the first similarity, and determine the difference between the first model feature corresponding to the recommended multimedia information and the second auxiliary features corresponding to other recommended multimedia information as the first difference. The recommendation model is trained with maximizing the first similarity and the first difference as the optimization goal. The specific formula is as follows:
[0104]
[0105] In the above formula, q1 can be used to represent the features corresponding to the first transformed multimedia information output by the feature sub-model. k2 can be used to represent the features corresponding to the second transformed multimedia information output by the auxiliary sub-model. τ can be used to represent the hyperparameter manually set based on expert experience. It can be used to represent, for each piece of recommended multimedia information, the similarity between the first model feature corresponding to the recommended multimedia information and the second auxiliary feature corresponding to the recommended multimedia information. It can be used to characterize, for each piece of recommended multimedia information, determining the sum of the differences between the first model feature corresponding to the recommended multimedia information and the second auxiliary features corresponding to other recommended multimedia information.
[0106] From the above formula, we can see that for each recommended multimedia information, The larger the value of L is, the higher the similarity between the first model feature corresponding to the recommended multimedia information and the second auxiliary feature corresponding to the recommended multimedia information is. nce The smaller (q1,k2). The smaller the value of L is, the higher the difference between the first model feature corresponding to the recommended multimedia information and the second auxiliary features corresponding to other recommended multimedia information is. nce The smaller (q1,k2).
[0107] Based on this, the server can train the recommendation model for each first multimedia information with the optimization goals of minimizing the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and maximizing the deviation between the first feature corresponding to the first multimedia information and the second feature, and with the optimization goals of maximizing the first similarity and the first difference.
[0108] Similarly, the server can determine, for each recommended multimedia information, the similarity between the second model feature corresponding to the recommended multimedia information and the first auxiliary feature corresponding to the recommended multimedia information as the second similarity, and determine the difference between the second model feature corresponding to the recommended multimedia information and the first auxiliary features corresponding to other recommended multimedia information as the second difference. The recommendation model is trained with maximizing the second similarity and the second difference as the optimization goal. The specific formula is as follows:
[0109]
[0110] In the above formula, q2 can be used to represent the features corresponding to the second transformed multimedia information output by the feature sub-model, and k1 can be used to represent the features corresponding to the first transformed multimedia information output by the auxiliary sub-model. It can be used to represent, for each piece of recommended multimedia information, the similarity between the second model feature corresponding to the recommended multimedia information and the first auxiliary feature corresponding to the recommended multimedia information. It can be used to characterize, for each piece of recommended multimedia information, determining the sum of the differences between the second model feature corresponding to the recommended multimedia information and the first auxiliary features corresponding to other recommended multimedia information.
[0111] From the above formula, we can see that for each recommended multimedia information, The larger the value of L is, the higher the similarity between the second model feature corresponding to the recommended multimedia information and the first auxiliary feature corresponding to the recommended multimedia information is. nce The smaller (q2,k1). The smaller the value of L is, the higher the difference between the second model feature corresponding to the recommended multimedia information and the first auxiliary feature corresponding to other recommended multimedia information is. nce The smaller (q1,k2).
[0112] Based on this, the server can, for each first multimedia information, take minimizing the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and maximizing the deviation between the first feature corresponding to the first multimedia information and the second feature as the optimization goals, and train the recommendation model with maximizing the first similarity, the first difference, the second similarity and the second difference as the optimization goals.
[0113] In actual applications, in the early stage of model training of the recommendation model, the features corresponding to the multimedia information determined by the recommendation model may not be accurate, which will lead to errors in the determination of each first multimedia information to minimize the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and to maximize the deviation between the first feature corresponding to the first multimedia information and the second feature, thereby reducing the efficiency of the model training of the recommendation model. Based on this, the server can reduce the weight of the similarity between the recommended multimedia information browsed by the user and the difference between the recommended multimedia information browsed by the user and the recommended multimedia information not browsed by the user in the model training process in the early stage, so that the early stage of model training focuses more on improving the accuracy of the features corresponding to the determined recommended multimedia information.
[0114] As the model training process continues to optimize, the recommendation model can determine relatively accurate features corresponding to recommended multimedia information. Then, in the later stages of model training, the server can increase the weighting of the similarities between recommended multimedia information viewed by users, as well as the differences between recommended multimedia information viewed by users and recommended multimedia information not viewed by users. This allows the server to focus more on improving the similarities between features corresponding to recommended multimedia information of the same category in the later stages of model training.
[0115] In an embodiment of this specification, the server may, for each first multimedia message, use the deviation between the first feature corresponding to the first multimedia message and the first features corresponding to other first multimedia messages as the first deviation, and the deviation between the first feature corresponding to the first multimedia message and the second feature as the second deviation. Secondly, for each round of training, a weight corresponding to the first deviation and a weight corresponding to the second deviation are determined based on the round number of the training round, as round weights, where a larger round number corresponds to the training round, and a larger round weight is assigned.
[0116] Finally, the server may train the recommendation model according to the first deviation, the second deviation, the first similarity, the first difference, the second similarity, the second difference, and the round weight.
[0117] L=L nce +β(t)*L tri
[0118]
[0119] In the above formula, L nce It can be used to represent the first similarity, the first difference, the second similarity and the second difference. tri It can be used to characterize the first deviation and the second deviation. β(t) can be used to characterize the round weight. α can be used to characterize the hyperparameter determined based on human experience. T can be used to characterize the total training rounds of the training model. t can be used to characterize the current training round of the training model. As can be seen from the above formula, the more current training rounds, the greater the round weight. That is to say, in the early stage of model training, more attention is paid to improving the accuracy of the features corresponding to the determined recommended multimedia information. In the later stage of model training, more attention is paid to improving the similarity between the features corresponding to the recommended multimedia information of the same category, as well as the difference between the recommended multimedia information browsed by the user and the recommended multimedia information not browsed by the user.
[0120] It can be seen from the above method that the recommendation model is used to bring the features corresponding to multimedia information of the same category closer, so as to avoid the situation where the recommendation model recommends highly similar multimedia information to the user in the process of recommending multimedia information based on user operations. The multimedia information recommended by the recommendation model is of the same category as the multimedia information operated by the user and has differences, thereby increasing the probability that the multimedia information meets the user's expectations.
[0121] In actual applications, users typically operate on recommended multimedia information based on a specific target requirement over a period of time. However, users may operate on other recommended multimedia information that does not meet their target requirement. Based on this, for each first multimedia information, the server can first determine several first multimedia information that are relatively close to the first multimedia information, then reduce the distance between the first multimedia information and the several first multimedia information that are relatively close to it. Furthermore, the server can determine several second multimedia information that are relatively far from the first multimedia information, and then increase the distance between the first multimedia information and the several second multimedia information that are relatively far from it.
[0122] In an embodiment of this specification, the server may determine the distance between the first feature corresponding to the first multimedia message and the first features corresponding to other first multimedia messages as the first distance, then sort the first distances to select a set number of first features corresponding to other first multimedia messages. Next, the server may determine the distance between the first feature corresponding to the first multimedia message and the second feature corresponding to the second multimedia message as the second distance, then sort the second distances to select a set number of second features corresponding to the second multimedia messages.
[0123] Similarly, the server may determine the distance between the third feature corresponding to the third multimedia message and the third features corresponding to other third multimedia messages as the third distance, and then sort the third distances to select a set number of third features corresponding to other third multimedia messages. Next, the server may determine the distance between the third feature corresponding to the third multimedia message and the fourth feature corresponding to the fourth multimedia message as the fourth distance, and then sort the fourth distances to select a set number of fourth multimedia messages corresponding to the fourth features.
[0124] Finally, for each first multimedia message, the recommendation model is trained with the optimization goals of minimizing the deviation between the first feature corresponding to the first multimedia message and the first features corresponding to a set number of other first multimedia messages that have been screened out, and maximizing the deviation between the first feature corresponding to the first multimedia message and the second feature corresponding to a set number of second multimedia messages that have been screened out. For each third multimedia message, the recommendation model is trained with the optimization goals of minimizing the deviation between the third feature corresponding to the third multimedia message and the third features corresponding to a set number of other third multimedia messages that have been screened out, and maximizing the deviation between the third feature corresponding to the third multimedia message and the fourth feature corresponding to a set number of fourth multimedia messages that have been screened out. After multiple rounds of iterative training, the deviation can be continuously reduced and converged within a numerical range, thereby completing the training process of the recommendation model.
[0125] Among them, the method for determining the training direction of the recommendation model can be a stochastic gradient descent method (Saccharomyces Genome Database, SGD).
[0126] It should be noted that in the recommendation model, the deviation can be continuously reduced after multiple rounds of iterative training. The feature sub-model is trained, and the auxiliary sub-model is updated with momentum based on the model parameter adjustment and set weights of the feature sub-model.
[0127] Figure 2 A schematic diagram of the model structure provided in the embodiments of this specification.
[0128] exist Figure 2 In the example, the server may randomly transform each recommended multimedia information in the user's historical operation information within a set time period to obtain each first transformed multimedia information and each second transformed multimedia information. Each first transformed multimedia information and each second transformed multimedia information are then input into a feature sub-model of the recommendation model to obtain a first model feature and a second model feature. Each first transformed multimedia information and each second transformed multimedia information are then input into an auxiliary sub-model of the recommendation model to obtain a first auxiliary feature and a second auxiliary feature.
[0129] Secondly, the server can determine, for each recommended multimedia information, the similarity between the first model feature corresponding to the recommended multimedia information and the second auxiliary feature corresponding to the recommended multimedia information as the first similarity, and determine the difference between the first model feature corresponding to the recommended multimedia information and the second auxiliary feature corresponding to other recommended multimedia information as the first difference.
[0130] For each recommended multimedia information, the similarity between the second model feature corresponding to the recommended multimedia information and the first auxiliary feature corresponding to the recommended multimedia information is determined as the second similarity, and the difference between the second model feature corresponding to the recommended multimedia information and the first auxiliary feature corresponding to other recommended multimedia information is determined as the second difference.
[0131] Then, the server may determine, based on each first model feature, a feature corresponding to each first multimedia message in each first transformed multimedia message as the first feature, and a feature corresponding to each second multimedia message in each first transformed multimedia message as the second feature;
[0132] And according to each second model feature, a feature corresponding to each first multimedia information in each second transformed multimedia information is determined as a third feature, and a feature corresponding to each second multimedia information in each second transformed multimedia information is determined as a fourth feature.
[0133] Finally, the server can, for each first multimedia information, minimize the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, minimize the deviation between the third feature corresponding to the first multimedia information and the third features corresponding to other first multimedia information, maximize the deviation between the first feature corresponding to the first multimedia information and the second feature, and maximize the deviation between the third feature corresponding to the first multimedia information and the fourth feature as optimization goals, and train the recommendation model with maximizing the first similarity, the first difference, the second similarity and the second difference as optimization goals.
[0134] It should be noted that the framework used by the above-mentioned recommendation model can be an Enc-Dec network structure such as a fully connected feedforward network (FFN), a fully connected neural network (FCNN), a multi-head self-attention mechanism network (MSA), or other forms of neural network structures. This specification does not limit the specific form of the framework used by the recommendation model.
[0135] As can be seen from the above process, this method can determine the similarity between the recommended multimedia information browsed by the user through the recommended multimedia information browsed by the user, and determine the difference between the recommended multimedia information browsed by the user and the recommended multimedia information not browsed by the user through the recommended multimedia information browsed by the user and the recommended multimedia information not browsed by the user. Finally, the recommendation model is trained with the optimization goals of minimizing the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and maximizing the deviation between the first feature corresponding to the first multimedia information and the second feature. This can ensure the correlation between the first features of the recommended multimedia information browsed by the user and the difference between the first feature of the recommended multimedia information browsed by the user and the second feature of the recommended multimedia information not browsed by the user, thereby improving the accuracy of the multimedia information recommended to the user. Furthermore, the server can also train the recommendation model with the optimization goals of maximizing the first similarity, the first difference, the second similarity, and the second difference to improve the accuracy of the features corresponding to the determined recommended multimedia information.
[0136] After the training of the recommendation model is completed, the recommendation model can be used to recommend information to users. The specific process is as follows: Figure 3 shown.
[0137] Figure 3 A flowchart of a method for recommending information provided in accordance with the embodiments of this specification.
[0138] S300: In response to an information acquisition request from a target user, acquiring candidate recommended multimedia information and a preference representation of the target user.
[0139] In an embodiment of this specification, a server may respond to a target user's information acquisition request and obtain candidate recommended multimedia information and the target user's preference representation. The target user herein refers to a user currently performing a service. The candidate recommended information may refer to filtered information within the service context or to all information within the service context. The target user's preference representation may be determined based on the user's historical order data and historical operation information.
[0140] S302: For each candidate recommended multimedia information, the candidate recommended multimedia information and the preference representation of the target user are input into a pre-trained recommendation model to predict the click rate corresponding to the candidate recommended multimedia information, where the recommendation model is trained using the above-mentioned model training method.
[0141] S304: Determine the recommended multimedia information to be recommended to the target user according to the click rate corresponding to each candidate recommended multimedia information, use the recommended multimedia information as the target multimedia information, and recommend the target recommended multimedia information to the target user.
[0142] In the embodiment of the present specification, the server may input each candidate recommended multimedia information and the target user's preference representation into a pre-trained recommendation model to predict the click rate corresponding to the candidate recommended multimedia information.
[0143] The method for determining the features corresponding to each candidate recommendation information is basically the same as the method mentioned in the above model training process, and will not be described in detail here.
[0144] Then, based on the click-through rates corresponding to the candidate recommended multimedia information, the recommended multimedia information to be recommended to the target user is determined as the target recommended multimedia information, and the target recommended multimedia information is recommended to the target user.
[0145] Specifically, the server can determine the click-through rate corresponding to each candidate recommendation information, sort the candidate recommendation information from large to small according to the click-through rate, and use the candidate recommendation information ranked before the set ranking position as the target recommendation information output at the current moment.
[0146] From the above content, it can be seen that the server can improve the recommendation effect of information recommendation to the target user by using the features corresponding to each candidate recommendation information determined by the recommendation model.
[0147] The above is a method for model training provided in one or more embodiments of this specification. Based on the same idea, this specification also provides a corresponding model training device, such as Figure 4 shown.
[0148] Figure 4 The schematic diagram of the structure of the model training device provided in the embodiment of this specification specifically includes:
[0149] The acquisition module 400 is used to obtain the user's operation information on each recommended multimedia information within a set historical time period;
[0150] a determination module 402 configured to determine, based on the operation information, from the recommended multimedia information, the recommended multimedia information browsed by the user as each first multimedia information, and to determine the recommended multimedia information not browsed by the user as each second multimedia information;
[0151] An input module 404 is configured to input each of the first multimedia information and each of the second multimedia information into a recommendation model to be trained, and determine a feature corresponding to each of the first multimedia information as a first feature, and a feature corresponding to each of the second multimedia information as a second feature;
[0152] The training module 406 is used to train the recommendation model for each first multimedia information with the optimization goals of minimizing the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and maximizing the deviation between the first feature corresponding to the first multimedia information and the second feature.
[0153] Optionally, the determination module 402 is specifically used to randomly transform the recommended multimedia information according to the operation information to obtain first transformed multimedia information and second transformed multimedia information, determine the recommended multimedia information browsed by the user from the first transformed multimedia information as first multimedia information, and determine the recommended multimedia information not browsed by the user as second multimedia information, determine the recommended multimedia information browsed by the user from the second transformed multimedia information as third multimedia information, and determine the recommended multimedia information not browsed by the user as fourth multimedia information.
[0154] Optionally, the input module 404 is specifically used to input the first multimedia information, the second multimedia information, the third multimedia information and the fourth multimedia information into the recommendation model to be trained, determine the feature corresponding to each first multimedia information in the first transformed multimedia information as the first feature, the feature corresponding to each second multimedia information in the first transformed multimedia information as the second feature, the feature corresponding to each third multimedia information in the second transformed multimedia information as the third feature, and the feature corresponding to each fourth multimedia information in the second transformed multimedia information as the fourth feature.
[0155] Optionally, the training module 406 is specifically used to train the recommendation model for each first multimedia information with the goal of minimizing the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and maximizing the deviation between the first feature corresponding to the first multimedia information and the second feature, and for each third multimedia information with the goal of minimizing the deviation between the third feature corresponding to the third multimedia information and the third features corresponding to other third multimedia information, and maximizing the deviation between the third feature corresponding to the third multimedia information and the fourth feature.
[0156] Optionally, the input module 404 is specifically configured to: the recommendation model includes: a feature sub-model and an auxiliary sub-model;
[0157] Input the first transformed multimedia information and the second transformed multimedia information into the feature sub-model to be trained, determine the feature corresponding to each first transformed multimedia information as the first model feature, and the feature corresponding to each second transformed multimedia information as the second model feature, input the first transformed multimedia information and the second transformed multimedia information into the auxiliary sub-model to be trained, determine the feature corresponding to each first transformed multimedia information as the first auxiliary feature, and the feature corresponding to each second transformed multimedia information as the second auxiliary feature.
[0158] Optionally, the training module 406 is specifically used to determine, for each recommended multimedia information, the similarity between the first model feature corresponding to the recommended multimedia information and the second auxiliary feature corresponding to the recommended multimedia information as the first similarity, and determine the difference between the first model feature corresponding to the recommended multimedia information and the second auxiliary feature corresponding to other recommended multimedia information as the first difference. For each first multimedia information, the optimization goal is to minimize the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and to maximize the deviation between the first feature corresponding to the first multimedia information and the second feature, and to train the recommendation model with maximizing the first similarity and the first difference as the optimization goal.
[0159] Optionally, the training module 406 is specifically used to determine, for each recommended multimedia information, the similarity between the second model feature corresponding to the recommended multimedia information and the first auxiliary feature corresponding to the recommended multimedia information as the second similarity, and determine the difference between the second model feature corresponding to the recommended multimedia information and the first auxiliary feature corresponding to other recommended multimedia information as the second difference. For each first multimedia information, the optimization goal is to minimize the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and to maximize the deviation between the first feature corresponding to the first multimedia information and the second feature. The recommendation model is trained with the optimization goal of maximizing the first similarity, the first difference, the second similarity and the second difference.
[0160] Optionally, the training module 406 is specifically used to, for each first multimedia information, use the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information as the first deviation, and the deviation between the first feature corresponding to the first multimedia information and the second feature as the second deviation. For each round of training, determine the weight corresponding to the first deviation and the weight corresponding to the second deviation according to the round of training, as round weights, wherein the larger the round corresponding to the round of training, the larger the round weight. The recommendation model is trained according to the first deviation, the second deviation, the first similarity, the first difference, the second similarity, the second difference and the round weight.
[0161] Figure 5 The schematic diagram of the structure of the device recommended for providing information in the embodiments of this specification specifically includes:
[0162] A response module 500 is configured to respond to a target user's information acquisition request and acquire candidate recommended multimedia information and a preference representation of the target user;
[0163] Prediction module 502 is configured to input each candidate recommended multimedia information and the target user's preference representation into a pre-trained recommendation model to predict the click-through rate corresponding to the candidate recommended multimedia information, wherein the recommendation model is trained using the above-mentioned model training method;
[0164] The recommendation module 504 is configured to determine the recommended multimedia information to be recommended to the target user according to the click-through rate of each candidate recommended multimedia information, as the target recommended multimedia information, and recommend the target recommended multimedia information to the target user.
[0165] This specification also provides a computer-readable storage medium, which stores a computer program that can be used to execute the above Figure 1 The model training method provided and the above Figure 3 The information provided is the recommended method.
[0166] This manual also provides Figure 6 The structural diagram of the electronic device shown in FIG. Figure 6 As mentioned above, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory and a non-volatile memory, and may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 The model training method and the above Figure 3The information provided recommends methods. Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. In other words, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0167] It should be noted that all actions of acquiring signals, information or data in this application are carried out in compliance with the relevant data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.
[0168] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD through their own programming, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly done using "logic compiler" software. This is similar to the software compiler used when developing programs. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.
[0169] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the means for implementing various functions included therein can also be considered as structures within the hardware component. Or even, the means for implementing various functions can be considered as both a software module implementing the method and a structure within the hardware component.
[0170] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0171] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0172] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0173] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0174] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0175] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0176] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0177] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0178] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0179] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0180] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Thus, this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0181] This specification may be described in the general context of computer-executable instructions, such as program modules, executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media, including storage devices.
[0182] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.
[0183] The foregoing is merely an example of the present invention and is not intended to limit the present invention. Various modifications and variations are possible within the scope of the present invention. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of the claims of the present invention.
Claims
1. A model training method, characterized in that: include: Obtaining the user's operation information on each recommended multimedia message within a set historical time period; determining, from the recommended multimedia information, recommended multimedia information browsed by the user as first multimedia information, and determining recommended multimedia information not browsed by the user as second multimedia information, according to the operation information; Inputting each of the first multimedia information and each of the second multimedia information into a recommendation model to be trained, determining a feature corresponding to each of the first multimedia information as a first feature, and a feature corresponding to each of the second multimedia information as a second feature; For each first multimedia information, the recommendation model is trained with the optimization goal of minimizing the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and maximizing the deviation between the first feature corresponding to the first multimedia information and the second feature; and based on the operation information, the recommended multimedia information browsed by the user is determined from the recommended multimedia information as each first multimedia information, and the recommended multimedia information not browsed by the user is determined as each second multimedia information, specifically including: According to the operation information, the recommended multimedia information is randomly transformed to obtain first transformed multimedia information and second transformed multimedia information; Determining, from the first transformed multimedia information, recommended multimedia information browsed by the user as each first multimedia information, and determining recommended multimedia information not browsed by the user as each second multimedia information; Determining, from the respective second transformed multimedia information, recommended multimedia information browsed by the user as respective third multimedia information, and determining recommended multimedia information not browsed by the user as respective fourth multimedia information; The recommendation model includes: a feature sub-model and an auxiliary sub-model; Inputting each of the first multimedia information and each of the second multimedia information into a recommendation model to be trained, and determining a feature corresponding to each of the first multimedia information as a first feature, and a feature corresponding to each of the second multimedia information as a second feature, specifically including: Inputting each of the first transformed multimedia information and each of the second transformed multimedia information into a feature sub-model to be trained, determining a feature corresponding to each of the first transformed multimedia information as a first model feature, and a feature corresponding to each of the second transformed multimedia information as a second model feature; inputting each of the first transformed multimedia information and each of the second transformed multimedia information into an auxiliary sub-model to be trained, determining a feature corresponding to each of the first transformed multimedia information as a first auxiliary feature, and a feature corresponding to each of the second transformed multimedia information as a second auxiliary feature; for each first multimedia information, training the recommendation model with the optimization goal of minimizing the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and maximizing the deviation between the first feature corresponding to the first multimedia information and the second feature. Specifically, the training includes: For each piece of recommended multimedia information, determining a similarity between the first model feature corresponding to the recommended multimedia information and the second auxiliary feature corresponding to the recommended multimedia information as a first similarity, and determining a difference between the first model feature corresponding to the recommended multimedia information and the second auxiliary feature corresponding to other pieces of recommended multimedia information as a first difference; For each first multimedia information, the recommendation model is trained with the optimization goals of minimizing the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and maximizing the deviation between the first feature corresponding to the first multimedia information and the second feature, and maximizing the first similarity and the first difference as the optimization goals.
2. The method according to claim 1, wherein Inputting each of the first multimedia information and each of the second multimedia information into a recommendation model to be trained, and determining a feature corresponding to each of the first multimedia information as a first feature, and a feature corresponding to each of the second multimedia information as a second feature, specifically including: Input the first multimedia information, the second multimedia information, the third multimedia information and the fourth multimedia information into the recommendation model to be trained, and determine the feature corresponding to each first multimedia information in the first transformed multimedia information as the first feature, the feature corresponding to each second multimedia information in the first transformed multimedia information as the second feature, the feature corresponding to each third multimedia information in the second transformed multimedia information as the third feature, and the feature corresponding to each fourth multimedia information in the second transformed multimedia information as the fourth feature.
3. The method according to claim 2, wherein For each first multimedia information, the recommendation model is trained with the optimization goal of minimizing the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and maximizing the deviation between the first feature corresponding to the first multimedia information and the second feature, specifically including: For each first multimedia information, the recommendation model is trained with the optimization objectives of minimizing the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and maximizing the deviation between the first feature corresponding to the first multimedia information and the second feature; and for each third multimedia information, the recommendation model is trained with the optimization objectives of minimizing the deviation between the third feature corresponding to the third multimedia information and the third features corresponding to other third multimedia information, and maximizing the deviation between the third feature corresponding to the third multimedia information and the fourth feature.
4. The method according to claim 1, wherein For each first multimedia information, minimizing the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and maximizing the deviation between the first feature corresponding to the first multimedia information and the second feature as optimization goals, and maximizing the first similarity and the first difference as optimization goals, the recommendation model is trained, specifically including: For each piece of recommended multimedia information, determining a similarity between the second model feature corresponding to the recommended multimedia information and the first auxiliary feature corresponding to the recommended multimedia information as a second similarity, and determining a difference between the second model feature corresponding to the recommended multimedia information and the first auxiliary feature corresponding to other pieces of recommended multimedia information as a second difference; For each first multimedia information, the optimization goal is to minimize the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and to maximize the deviation between the first feature corresponding to the first multimedia information and the second feature, and the recommendation model is trained with the optimization goal of maximizing the first similarity, the first difference, the second similarity and the second difference.
5. The method according to claim 4, wherein For each first multimedia information, minimizing the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and maximizing the deviation between the first feature corresponding to the first multimedia information and the second feature as optimization goals, and maximizing the first similarity, the first difference, the second similarity, and the second difference as optimization goals, the recommendation model is trained, specifically including: For each first multimedia message, a deviation between a first feature corresponding to the first multimedia message and a first feature corresponding to another first multimedia message is used as a first deviation, and a deviation between the first feature corresponding to the first multimedia message and the second feature is used as a second deviation; For each round of training, determining a weight corresponding to the first deviation and a weight corresponding to the second deviation according to the round of training, as the round weight, wherein a larger round corresponding to the training round has a larger round weight; The recommendation model is trained according to the first deviation, the second deviation, the first similarity, the first difference, the second similarity, the second difference, and the round weight.
6. A method for information recommendation, characterized in that: include: In response to an information acquisition request from a target user, acquiring candidate recommended multimedia information and a preference representation of the target user; For each candidate recommended multimedia information, input the candidate recommended multimedia information and the preference representation of the target user into a pre-trained recommendation model to predict the click-through rate corresponding to the candidate recommended multimedia information, wherein the recommendation model is trained by the method of any one of claims 1 to 5 above; According to the click rates corresponding to the candidate recommended multimedia information, the recommended multimedia information to be recommended to the target user is determined as the target recommended multimedia information, and the target recommended multimedia information is recommended to the target user.
7. A model training device, characterized in that: include: An acquisition module is used to obtain the user's operation information on each recommended multimedia information within a historical set time period; a determination module configured to determine, based on the operation information, from the recommended multimedia information, the recommended multimedia information browsed by the user as each first multimedia information, and to determine the recommended multimedia information not browsed by the user as each second multimedia information; an input module, configured to input the first multimedia information and the second multimedia information into a recommendation model to be trained, and determine a feature corresponding to each first multimedia information as a first feature, and a feature corresponding to each second multimedia information as a second feature; a training module, configured to train the recommendation model for each first multimedia information, with minimizing the deviation between a first feature corresponding to the first multimedia information and first features corresponding to other first multimedia information, and maximizing the deviation between the first feature corresponding to the first multimedia information and the second feature as optimization objectives; Determining, based on the operation information, from the recommended multimedia information, recommended multimedia information browsed by the user as first multimedia information, and determining recommended multimedia information not browsed by the user as second multimedia information, specifically includes: According to the operation information, the recommended multimedia information is randomly transformed to obtain first transformed multimedia information and second transformed multimedia information; Determining, from the first transformed multimedia information, recommended multimedia information browsed by the user as each first multimedia information, and determining recommended multimedia information not browsed by the user as each second multimedia information; Determining, from the respective second transformed multimedia information, recommended multimedia information browsed by the user as respective third multimedia information, and determining recommended multimedia information not browsed by the user as respective fourth multimedia information; The recommendation model includes: a feature sub-model and an auxiliary sub-model; Inputting each of the first multimedia information and each of the second multimedia information into a recommendation model to be trained, and determining a feature corresponding to each of the first multimedia information as a first feature, and a feature corresponding to each of the second multimedia information as a second feature, specifically including: Inputting each of the first transformed multimedia information and each of the second transformed multimedia information into a feature sub-model to be trained, determining a feature corresponding to each of the first transformed multimedia information as a first model feature, and a feature corresponding to each of the second transformed multimedia information as a second model feature; inputting each of the first transformed multimedia information and each of the second transformed multimedia information into an auxiliary sub-model to be trained, determining a feature corresponding to each of the first transformed multimedia information as a first auxiliary feature, and a feature corresponding to each of the second transformed multimedia information as a second auxiliary feature; for each first multimedia information, training the recommendation model with the optimization goal of minimizing the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and maximizing the deviation between the first feature corresponding to the first multimedia information and the second feature. Specifically, the training includes: For each piece of recommended multimedia information, determining a similarity between the first model feature corresponding to the recommended multimedia information and the second auxiliary feature corresponding to the recommended multimedia information as a first similarity, and determining a difference between the first model feature corresponding to the recommended multimedia information and the second auxiliary feature corresponding to other pieces of recommended multimedia information as a first difference; For each first multimedia information, the recommendation model is trained with the optimization goals of minimizing the deviation between the first feature corresponding to the first multimedia information and the first features corresponding to other first multimedia information, and maximizing the deviation between the first feature corresponding to the first multimedia information and the second feature, and maximizing the first similarity and the first difference as the optimization goals.
8. An information recommendation device, characterized in that: include: A response module, configured to respond to a target user's information acquisition request and acquire candidate recommended multimedia information and a preference representation of the target user; a prediction module for inputting, for each candidate recommended multimedia information, the candidate recommended multimedia information and the target user's preference representation into a pre-trained recommendation model to predict a click-through rate corresponding to the candidate recommended multimedia information, wherein the recommendation model is trained using the method of any one of claims 1 to 5; The recommendation module is used to determine the recommended multimedia information to be recommended to the target user according to the click rate corresponding to each candidate recommended multimedia information, as the target recommended multimedia information, and recommend the target recommended multimedia information to the target user.
9. A computer-readable storage medium, characterized in that The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Click rate prediction method, prediction model training method and device and equipment
CN110363346A
Model training method and resource recommendation method
CN113469298A
Model training method and device, image depth prediction method and device, equipment and medium
CN113743517A