Recommended Model Training Method and Device

By training a fusion recommendation model and updating the model parameters using the sorting evaluation indicators, the problem of insufficiently accurate recommendation results in the existing technology is solved, and a more accurate recommendation effect is achieved.

CN113536105BActive Publication Date: 2025-06-13TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202011223031.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-05
Publication Date
2025-06-13
Estimated Expiration
2040-11-05

AI Technical Summary

Technical Problem

When establishing artificial intelligence recommendation models, the existing technology usually establishes separate models for each sub-business, resulting in the incomplete recommendation results.

Method used

By obtaining the sub-label set of training samples and trained sub-target models, a fusion recommendation model is trained, and the model parameters are updated using the sorting evaluation indicators to improve the accuracy of the recommended results.

Benefits of technology

More accurate recommendation results are achieved, and the accuracy and effectiveness of recommendations are improved by integrating the output of multiple sub-business models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113536105B_ABST
    Figure CN113536105B_ABST
Patent Text Reader

Abstract

The present application relates to a method and apparatus for training a recommendation model. The method includes: obtaining training samples and obtaining sub-label sets respectively corresponding to at least two trained sub-goal models; inputting the training samples into the trained sub-goal models to obtain sub-recommendation degree sets; inputting each sub-recommendation degree set into an initial fusion recommendation model to obtain a fusion recommendation degree set, and sorting each historical recommendation target based on the fusion recommendation degree to obtain a historical recommendation target sequence; obtaining sub-label sequences respectively corresponding to the trained sub-goal models based on the historical recommendation target sequence; determining target sorting evaluation information based on the sorting evaluation information corresponding to each sub-label sequence; updating the initial fusion recommendation model based on the target sorting evaluation information, and when the training is completed, obtaining a target fusion recommendation model, where the target fusion recommendation model is used to recommend information to be recommended. Using this method can improve the accuracy of the target fusion recommendation model when making recommendations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technologies, and particularly to a method and apparatus for training a recommendation model, a computer device, and a storage medium. Background Art

[0002] With the development of artificial intelligence technologies, recommendation technologies based on artificial intelligence have emerged, such as video recommendation, product recommendation, news recommendation, advertisement recommendation, and so on. Currently, when establishing an artificial intelligence recommendation model, models are usually established separately for each sub-business. For example, a video click-through rate recommendation model usually performs video recommendation based on the click-through rate and does not pay attention to other features of the video. Finally, the outputs of each sub-business model are fused to obtain a fused result, and recommendations are made according to the fused result. Currently, when performing fusion, weights are usually set for the outputs of each sub-business model and weighted fusion is performed to obtain a fused recommendation result. However, simply performing weighted fusion will result in the problem that the fused recommendation result is not accurate enough. Summary of the Invention

[0003] Based on this, there is provided a method and apparatus for training a recommendation model, a computer device, and a storage medium that can improve the accuracy of recommendation results.

[0004] A method for training a recommendation model, the method includes:

[0005] Obtain training samples, where the training samples include various historical recommendation targets, and obtain sub-label sets respectively corresponding to at least two trained sub-target models, and each sub-label set includes sub-labels corresponding to various historical recommendation targets;

[0006] Input the training samples into the trained sub-target models to obtain sub-recommendation degree sets output by each trained sub-target model, and each sub-recommendation degree set includes sub-recommendation degrees corresponding to various historical recommendation targets;

[0007] Input each sub-recommendation degree set into an initial fusion recommendation model to obtain a fusion recommendation degree set, where the fusion recommendation degree set includes fusion recommendation degrees corresponding to various historical recommendation targets, and sort various historical recommendation targets based on the fusion recommendation degrees to obtain a historical recommendation target sequence;

[0008] Sort the sub-labels corresponding to various historical recommendation targets in each sub-label set according to the order of the historical recommendation target sequence to obtain sub-label sequences respectively corresponding to each trained sub-target model;

[0009] Determine sorting evaluation information corresponding to each sub-label sequence based on a sorting evaluation index, and determine target sorting evaluation information based on the sorting evaluation information corresponding to each sub-label sequence;

[0010] Update the initial fusion recommendation model based on the target sorting evaluation information. When the training is completed, obtain the target fusion recommendation model, which is used to recommend the information to be recommended.

[0011] In one embodiment, the training sample includes the historical user identifier and each historical recommendation target corresponding to the historical user identifier;

[0012] Inputting the training sample into the trained sub-target model to obtain the sub-recommendation degree sets output by each trained sub-target model includes:

[0013] Obtain the historical attribute features corresponding to the historical user identifier and the historical recommendation target features corresponding to each historical recommendation target;

[0014] Input the historical attribute features and the historical recommendation target features into the trained sub-target model to obtain the sub-recommendation degree sets corresponding to the historical user identifier output by each trained sub-target model.

[0015] A recommendation method, the method includes:

[0016] Obtain the user identifier and obtain the attribute features based on the user identifier;

[0017] Obtain each target to be recommended and the corresponding target attribute features, input the attribute features and the target attribute features into at least two trained sub-target models to obtain the sub-target recommendation degree sets output by each trained sub-target model, and the sub-target recommendation degree sets include the sub-target recommendation degrees corresponding to each target to be recommended;

[0018] Input each sub-recommendation degree set into the target fusion recommendation model to obtain the fusion recommendation degrees corresponding to each target to be recommended. The target fusion recommendation model is trained using the training sample and the sub-label sets corresponding to at least two trained sub-target models respectively. The training sample includes each historical recommendation target, and the sub-label sets include the sub-labels corresponding to each historical recommendation target;

[0019] Sort each target to be recommended based on the fusion recommendation degree to obtain the sequence of targets to be recommended;

[0020] Select a preset number of targets to be recommended from the sequence of targets to be recommended and recommend the preset number of targets to be recommended to the user identifier.

[0021] A recommendation model training device, the device includes:

[0022] A sample acquisition module, configured to acquire a training sample, the training sample includes each historical recommendation target, and acquire the sub-label sets corresponding to at least two trained sub-target models respectively. The sub-label sets include the sub-labels corresponding to each historical recommendation target;

[0023] A sub - recommendation degree obtaining module, configured to input training samples into a trained sub - target model, and obtain a set of sub - recommendation degrees output by each trained sub - target model. The set of sub - recommendation degrees includes sub - recommendation degrees corresponding to each historical recommendation target;

[0024] A target sequence obtaining module, configured to input each set of sub - recommendation degrees into an initial fusion recommendation model, obtain a set of fusion recommendation degrees. The set of fusion recommendation degrees includes fusion recommendation degrees corresponding to each historical recommendation target, and sort each historical recommendation target based on the fusion recommendation degrees to obtain a historical recommendation target sequence;

[0025] A sub - sequence obtaining module, configured to sort sub - labels corresponding to each historical recommendation target in a sub - label set based on the order of the historical recommendation target sequence, and obtain sub - label sequences respectively corresponding to each trained sub - target model;

[0026] An evaluation module, configured to determine sorting evaluation information corresponding to each sub - label sequence based on a sorting evaluation index, and determine target sorting evaluation information based on the sorting evaluation information corresponding to each sub - label sequence;

[0027] An update module, configured to update the initial fusion recommendation model based on the target sorting evaluation information. When the training is completed, a target fusion recommendation model is obtained. The target fusion recommendation model is used to recommend information to be recommended.

[0028] A recommendation device, the device includes:

[0029] A feature acquisition module, configured to acquire a user identifier and acquire attribute features based on the user identifier;

[0030] A feature input module, configured to acquire each target to be recommended and corresponding target attribute features, and input the attribute features and the target attribute features into at least two trained sub - target models to obtain a set of sub - target recommendation degrees output by each trained sub - target model. The set of sub - target recommendation degrees includes sub - target recommendation degrees corresponding to each target to be recommended;

[0031] A fusion module, configured to input each set of sub - recommendation degrees into a target fusion recommendation model to obtain fusion recommendation degrees corresponding to each target to be recommended. The target fusion recommendation model is trained using training samples and sub - label sets respectively corresponding to at least two trained sub - target models. The training samples include each historical recommendation target, and the sub - label sets include sub - labels corresponding to each historical recommendation target;

[0032] A sorting module, configured to sort each target to be recommended based on the fusion recommendation degrees to obtain a target - to - be - recommended sequence;

[0033] A recommendation module for selecting a preset number of recommended targets from a sequence of targets to be recommended and recommending the preset number of recommended targets to a user identifier.

[0034] A computer device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the following steps are implemented:

[0035] Obtain training samples, where the training samples include various historical recommended targets, and obtain sub-label sets corresponding to at least two trained sub-target models respectively. The sub-label sets include sub-labels corresponding to various historical recommended targets;

[0036] Input the training samples into the trained sub-target models to obtain sub-recommendation degree sets output by the trained sub-target models respectively. The sub-recommendation degree sets include sub-recommendation degrees corresponding to various historical recommended targets;

[0037] Input the sub-recommendation degree sets into an initial fusion recommendation model to obtain a fusion recommendation degree set. The fusion recommendation degree set includes fusion recommendation degrees corresponding to various historical recommended targets. Sort the various historical recommended targets based on the fusion recommendation degrees to obtain a historical recommended target sequence;

[0038] Sort the sub-labels corresponding to various historical recommended targets in the sub-label sets based on the order of the historical recommended target sequence to obtain sub-label sequences corresponding to the trained sub-target models respectively;

[0039] Determine sorting evaluation information corresponding to the sub-label sequences based on a sorting evaluation index, and determine target sorting evaluation information based on the sorting evaluation information corresponding to the sub-label sequences;

[0040] Update the initial fusion recommendation model based on the target sorting evaluation information. When the training is completed, obtain a target fusion recommendation model, and the target fusion recommendation model is used to recommend information to be recommended.

[0041] A computer device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the following steps are implemented:

[0042] Obtain a user identifier and obtain attribute features based on the user identifier;

[0043] Obtain various targets to be recommended and corresponding target attribute features, and input the attribute features and the target attribute features into at least two trained sub-target models to obtain sub-target recommendation degree sets output by the trained sub-target models respectively. The sub-target recommendation degree sets include sub-target recommendation degrees corresponding to various targets to be recommended;

[0044] Input each sub - recommendation degree set into the target fusion recommendation model to obtain the fusion recommendation degrees corresponding to each target to be recommended. The target fusion recommendation model is trained using training samples and sub - label sets respectively corresponding to at least two trained sub - target models. The training samples include each historical recommended target, and the sub - label sets include sub - labels corresponding to each historical recommended target;

[0045] Sort each target to be recommended based on the fusion recommendation degrees to obtain a sequence of targets to be recommended;

[0046] Select a preset number of targets to be recommended from the sequence of targets to be recommended, and recommend the preset number of targets to be recommended to the user identifier.

[0047] A computer - readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0048] Obtain training samples, where the training samples include each historical recommended target, and obtain sub - label sets respectively corresponding to at least two trained sub - target models. The sub - label sets include sub - labels corresponding to each historical recommended target;

[0049] Input the training samples into the trained sub - target models to obtain sub - recommendation degree sets output by each trained sub - target model. The sub - recommendation degree sets include sub - recommendation degrees corresponding to each historical recommended target;

[0050] Input each sub - recommendation degree set into the initial fusion recommendation model to obtain a fusion recommendation degree set. The fusion recommendation degree set includes fusion recommendation degrees corresponding to each historical recommended target. Sort each historical recommended target based on the fusion recommendation degrees to obtain a sequence of historical recommended targets;

[0051] Sort the sub - labels corresponding to each historical recommended target in the sub - label sets based on the order of the sequence of historical recommended targets to obtain sub - label sequences respectively corresponding to each trained sub - target model;

[0052] Determine sorting evaluation information corresponding to each sub - label sequence based on sorting evaluation metrics, and determine target sorting evaluation information based on the sorting evaluation information corresponding to each sub - label sequence;

[0053] Update the initial fusion recommendation model based on the target sorting evaluation information. When the training is completed, obtain the target fusion recommendation model, which is used to recommend information to be recommended.

[0054] A computer - readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0055] Obtain a user identifier, and obtain attribute features based on the user identifier;

[0056] Obtain each target to be recommended and its corresponding target attribute features, input the attribute features and the target attribute features into at least two trained sub-target models, and obtain a set of sub-target recommendation degrees output by each trained sub-target model. The set of sub-target recommendation degrees includes the sub-target recommendation degrees corresponding to each target to be recommended;

[0057] Input each set of sub-recommendation degrees into the target fusion recommendation model to obtain the fusion recommendation degrees corresponding to each target to be recommended. The target fusion recommendation model is trained using training samples and sub-label sets corresponding to at least two trained sub-target models respectively. The training samples include each historical recommended target, and the sub-label sets include the sub-labels corresponding to each historical recommended target;

[0058] Sort each target to be recommended based on the fusion recommendation degrees to obtain a sequence of targets to be recommended;

[0059] Select a preset number of targets to be recommended from the sequence of targets to be recommended, and recommend the preset number of targets to be recommended to the user identifier.

[0060] For the above-mentioned recommendation model training method, device, computer device, and storage medium, sort each historical recommended target through the obtained set of fusion recommendation degrees to obtain a sequence of historical recommended targets, and then sort the sub-labels corresponding to each historical recommended target in the sub-label set based on the order of the sequence of historical recommended targets to obtain sub-label sequences corresponding to each trained sub-target model respectively. Determine the target sorting evaluation information according to the sorting evaluation information corresponding to each sub-label sequence, and then update the initial fusion recommendation model according to the target sorting evaluation information. When the training is completed, obtain the target fusion recommendation model, and the target fusion recommendation model is used to recommend the information to be recommended. That is, determine each sub-label sequence through the sequence of historical recommended targets, and determine the target sorting evaluation information according to the sorting evaluation information corresponding to each sub-label sequence, so as to make the obtained target sorting evaluation information more accurate, and then use the target sorting evaluation information to update the initial fusion recommendation model, so that the trained target fusion recommendation model can be more accurate when performing target fusion recommendation. Brief Description of the Drawings

[0061] Figure 1 It is an application environment diagram of the recommendation model training method in an embodiment;

[0062] Figure 2 It is a flowchart of the recommendation model training method in an embodiment;

[0063] Figure 3 It is a flowchart of calculating the first sorting evaluation information in an embodiment;

[0064] Figure 4Schematic diagram of the process for calculating the second sorting evaluation information in an embodiment;

[0065] Figure 5 Schematic diagram of the process for determining the number of positive order pairs in an embodiment;

[0066] Figure 6 Schematic diagram for determining the number of positive order pairs in a specific embodiment;

[0067] Figure 7 Schematic diagram of the process for obtaining the second target sorting evaluation information in an embodiment;

[0068] Figure 8 Schematic diagram of the process for obtaining the third target sorting evaluation information in an embodiment;

[0069] Figure 9 Schematic diagram of the process for obtaining the target fusion recommendation model in an embodiment;

[0070] Figure 10 Schematic diagram of the process for obtaining partial derivatives in an embodiment;

[0071] Figure 11 Schematic diagram of the process for obtaining the initial fusion recommendation model in an embodiment;

[0072] Figure 12 Schematic diagram of the process for the recommendation method in an embodiment;

[0073] Figure 13 Schematic diagram of the process for the recommendation model training method in a specific embodiment;

[0074] Figure 14 Structural block diagram of the recommendation model training device in an embodiment;

[0075] Figure 15 Structural block diagram of the recommendation device in an embodiment;

[0076] Figure 16 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0077] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0078] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making.

[0079] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0080] The solution provided in the embodiments of this application relates to technologies such as machine learning in artificial intelligence and will be specifically described through the following embodiments:

[0081] The recommended model training method provided in this application can be applied, for example, to Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The terminal 102 sends a model training instruction to the server 104, and the server 104 obtains training samples according to the model training instruction. The training samples include each historical recommendation target, and obtains at least two sub-label sets corresponding to the trained sub-target models, and the sub-label set includes the sub-labels corresponding to each historical recommendation target; the server 104 inputs the training samples into the trained sub-target models to obtain the sub-recommendation degree sets output by each trained sub-target model, and the sub-recommendation degree sets include the sub-recommendation degrees corresponding to each historical recommendation target; the server 104 inputs each sub-recommendation degree set into the initial fusion recommendation model to obtain a fusion recommendation degree set, and the fusion recommendation degree set includes the sub-recommendation degrees corresponding to each historical recommendation target. The fusion recommendation degree of each historical recommendation target is sorted based on the fusion recommendation degree to obtain a historical recommendation target sequence; the sub-tags corresponding to each historical recommendation target in the sub-tag set are sorted based on the order of the historical recommendation target sequence to obtain sub-tag sequences corresponding to each trained sub-target model; the server 104 determines the sorting evaluation information corresponding to each sub-tag sequence based on the sorting evaluation index, and determines the target sorting evaluation information based on the sorting evaluation information corresponding to each sub-tag sequence; the server 104 updates the initial fusion recommendation model based on the target sorting evaluation information, and when the training is completed, the target fusion recommendation model is obtained, and the target fusion recommendation model is used to recommend the recommended information. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers and portable wearable devices, and the server 104 can be implemented by an independent server or a server cluster composed of multiple servers.

[0082] In one embodiment, Figure 2 As shown in FIG, a recommendation model training method is provided, and the method is applied to Figure 1 The server in the example is used for explanation. It can be understood that the method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0083] Step 202, obtaining training samples, the training samples include each historical recommendation target, and obtaining sub-label sets corresponding to at least two trained sub-target models, the sub-label sets include sub-labels corresponding to each historical recommendation target.

[0084] Among them, the target refers to the object that needs to be recommended to the user through the network, and different recommendation application scenarios have different targets. The target may include at least one of advertisements, videos, commodities, social objects, texts, music, and pictures. The historical recommendation target refers to the target that has been recommended to the user at a historical time. For example, in the video recommendation scenario, the historical recommendation target may be each historical recommended video. In the commodity recommendation scenario, the historical recommendation target may be each historical recommended commodity. In the social object recommendation scenario, the historical recommendation target may be each historical recommended social user. In the music recommendation scenario, the historical recommendation target may be each historical recommended music. In the picture recommendation scenario, the historical recommendation target may be each historical recommended picture.

[0085] The sub-goal refers to different business goals corresponding to the target. For example, in the video recommendation scenario, the sub-goals corresponding to the video may be different business goals such as the completion rate, attention rate, like rate, and video stay duration. For another example, in the commodity recommendation scenario, the sub-goals may be different business goals such as the commodity click-through rate, commodity collection rate, and commodity purchase rate. For example, in the social object recommendation scenario, the sub-goals may be the proportion of common social objects, the age similarity of social objects, and so on. The sub-goal model is trained using artificial intelligence algorithms with training samples and corresponding sub-label sets. The artificial intelligence algorithm may be a linear regression algorithm, a deep neural network algorithm, a decision tree algorithm, a random forest algorithm, a support vector machine algorithm, and so on. Different sub-goal models may use different artificial intelligence algorithms or the same artificial intelligence algorithm. The sub-label sets corresponding to different sub-goal models are different. The sub-label refers to the true result of the sub-goal corresponding to the historical recommendation target. For example, in the video recommendation scenario, the sub-label corresponding to the video stay duration prediction model may be the video stay duration. In the commodity recommendation scenario, the sub-labels corresponding to the commodity collection model include the collection label and the non-collection label. The training sample is a sample used for model training and can be used to train the sub-goal model or the target fusion recommendation model.

[0086] Specifically, the server can obtain training samples from the database. The training samples include each historical recommendation target. The server can also obtain training samples from a third party, where the third party refers to the service provider storing the training samples. The server can also collect training samples from the Internet. At the same time, the server obtains the sub-label sets corresponding to at least two trained sub-goal models respectively, and the sub-label sets include the sub-labels corresponding to each historical recommendation target.

[0087] Step 204, input the training samples into the trained sub-goal models to obtain a set of sub-recommendation degrees output by each trained sub-goal model. The set of sub-recommendation degrees includes the sub-recommendation degrees corresponding to each historical recommendation target.

[0088] Among them, the sub-recommendation degree refers to the recommendation degree output by the trained sub-goal model, which is used to characterize the recommendability of the corresponding training samples under this business goal, that is, the output result of the sub-goal model is used to describe the tendency degree on the corresponding business goal. For example, in the video recommendation scenario, the higher the output result of the favorite model, the more likely the user is to favorite, that is, the more likely the user is to pay attention.

[0089] Specifically, the server pre-trains to obtain the trained sub-goal model, deploys the trained sub-goal model in the server. When a training sample is obtained, the training sample is input into the trained sub-goal model to obtain a set of sub-recommendation degrees output by each trained sub-goal model. The set of sub-recommendation degrees includes the sub-recommendation degrees corresponding to each historical recommendation goal.

[0090] In one embodiment, the server can send the training sample to the server where the trained sub-goal model is deployed, that is, the trained sub-goal model can be deployed in other servers, such as cloud servers and third-party servers. The server obtains the set of sub-recommendation degrees returned by the server where the trained sub-goal model is deployed. The set of sub-recommendation degrees includes the sub-recommendation degrees corresponding to each historical recommendation goal in the training sample.

[0091] Step 206, input each set of sub-recommendation degrees into the initial fusion recommendation model to obtain a set of fusion recommendation degrees. The set of fusion recommendation degrees includes the fusion recommendation degrees corresponding to each historical recommendation goal. Sort each historical recommendation goal based on the fusion recommendation degree to obtain a historical recommendation goal sequence.

[0092] Among them, the initial fusion recommendation model refers to the fusion recommendation model with initialized model parameters. The fusion recommendation degree is used to characterize the fused recommendation degree corresponding to the historical recommendation goal. This fusion recommendation degree is obtained by fusing the sub-recommendation degrees output by the historical recommendation goal in different sub-goal models. For example, in the commodity recommendation scenario, the recommendation degree output by the commodity click model and the recommendation output by the commodity favorite model are fused to obtain the fused commodity recommendation degree. The historical recommendation goal sequence refers to the sequence obtained by sorting the historical recommendation goals according to the fusion recommendation degree.

[0093] Specifically, the server inputs the sub-recommendation degrees corresponding to the same historical recommendation goal in each set of sub-recommendation degrees into the initial fusion recommendation model at the same time to obtain the fusion recommendation degree of this historical recommendation goal in the set of fusion recommendation degrees output. The set of fusion recommendation degrees includes the fusion recommendation degrees corresponding to each historical recommendation goal. Then the server sorts each historical recommendation goal according to the size of the fusion recommendation degree to obtain a historical recommendation goal sequence. Among them, each historical recommendation goal can be sorted from large to small according to the fusion recommendation degree, or each historical recommendation goal can be sorted from small to large according to the fusion recommendation degree.

[0094] Step 208: Sort the sub - labels corresponding to each historical recommendation target in the sub - label set according to the order of the historical recommendation target sequence, to obtain the sub - label sequences corresponding to each trained sub - target model respectively.

[0095] Among them, the sub - label sequence refers to the sequence obtained by sorting the sub - labels of the same sub - target model according to the order of the historical recommendation target sequence.

[0096] Specifically, the server sorts the sub - labels in the sub - label sets corresponding to different trained sub - target models according to the order of the historical recommendation target sequence. For example, the sub - label corresponding to the first historical recommendation target in the historical recommendation target sequence is the first sub - label in the sub - label sequence. Different trained sub - target models have different sub - label sets, thus obtaining the corresponding sub - label sequences.

[0097] Step 210: Determine the sorting evaluation information corresponding to each sub - label sequence based on the sorting evaluation metric, and determine the target sorting evaluation information based on the sorting evaluation information corresponding to each sub - label sequence.

[0098] Among them, the sorting evaluation metric is used to evaluate the sorting accuracy corresponding to each sub - label sequence. This sorting evaluation metric can include AUC (Area Under Curve, the area enclosed by the ROC curve and the coordinate axes), the proportion of positive pairs, NDCG (Normalized Discounted Cumulative Gain), or a - NDCG, etc. The sorting evaluation information refers to the sorting evaluation metric values corresponding to each sub - label sequence. The target sorting evaluation information refers to the sorting evaluation metric value used to evaluate the sorting accuracy corresponding to the historical recommendation target sequence.

[0099] Specifically, the server can obtain the label data types corresponding to each sub - label sequence. The label data type is used to characterize the data type corresponding to the label. Different sub - label sequences can have different label data types, or the same label data type. Then, based on the label data type, obtain the corresponding sorting evaluation metric. Different label data types have different sorting evaluation metrics. And calculate the sorting evaluation information corresponding to each sub - label sequence according to the sorting evaluation metric. Finally, the server calculates the target sorting evaluation information based on the sorting evaluation information corresponding to each sub - label sequence and the pre - set weights. Different sub - targets can be set with different weights.

[0100] Step 212: Update the initial fusion recommendation model based on the target sorting evaluation information. When the training is completed, obtain the target fusion recommendation model, and the target fusion recommendation model is used to recommend the information to be recommended.

[0101] Among them, the information to be recommended refers to the information that needs to be recommended to the user, including at least one of advertisements, videos, products, social objects, texts, music, and pictures.

[0102] Specifically, the server updates the model parameters in the initial fusion recommendation model based on the target sorting evaluation information. Among them, algorithms such as the gradient descent algorithm, Adagrad (Adaptive Gradient) algorithm, Adadelta (an improvement of the Adagrad algorithm), RMSprop (an improvement of the Adagrad algorithm), and Adam (Adaptive Moment Estimation) algorithm can be used as optimizers to update the model parameters in the initial fusion recommendation model. When the model parameters converge, the training is completed, and the target fusion recommendation model is obtained. The model parameter convergence condition can be that the model parameters no longer change, or the model parameters decrease, or the number of training times reaches the maximum number of iterations, etc. The obtained target fusion recommendation model can be used to recommend various information to be recommended. For example, when the information to be recommended is multiple videos to be recommended, the multiple videos to be recommended are input into the target fusion recommendation model, and an output video sequence is obtained. The video can be recommended to the user according to the output video sequence.

[0103] In the above recommendation model training method, each historical recommendation target is sorted through the obtained fusion recommendation degree set to obtain a historical recommendation target sequence. Then, the sub-labels corresponding to each historical recommendation target in the sub-label set are sorted based on the order of the historical recommendation target sequence to obtain the sub-label sequences corresponding to each trained sub-target model. The target sorting evaluation information is determined according to the sorting evaluation information corresponding to each sub-label sequence. Then, the initial fusion recommendation model is updated according to the target sorting evaluation information. When the training is completed, the target fusion recommendation model is obtained. The target fusion recommendation model is used to recommend the information to be recommended. That is, each sub-label sequence is determined through the historical recommendation target sequence, and the target sorting evaluation information is determined according to the sorting evaluation information corresponding to each sub-label sequence, so that the obtained target sorting evaluation information can be more accurate. Then, the initial fusion recommendation model is updated using the target sorting evaluation information, so that the obtained target fusion recommendation model can be more accurate when performing target fusion recommendation.

[0104] In one embodiment, the training sample includes a historical user identifier and each historical recommendation target corresponding to the historical user identifier;

[0105] Step 204, input the training sample into the trained sub-target model to obtain a sub-recommendation degree set output by each trained sub-target model, including:

[0106] Obtain the historical attribute features corresponding to the historical user identifier and the historical recommendation target features corresponding to each historical recommendation target; input the historical attribute features and the historical recommendation target features into the trained sub-target models to obtain the set of sub-recommendation degrees corresponding to the historical user identifier output by each trained sub-target model.

[0107] Among them, the historical user identifier is used to uniquely identify the historical user and can be a number, a string, a name, etc. The historical attribute features are used to characterize the attribute features of the historical user and can include basic attribute features and behavioral attribute features. Among them, the basic attribute features can be age attribute features, gender attribute features, etc. The historical users in different application scenarios have different behavioral attribute features. For example, in the video recommendation scenario, the behavioral attribute features of users can include viewing behavior features, like behavior features, collection behavior features, etc. The historical recommendation target features are used to characterize the attribute features of the historical recommendation target itself. The historical recommendation targets in different application scenarios have different historical recommendation target features. For example, in the video recommendation scenario, the historical video features can include video duration features, video view count features, video like count features, etc.

[0108] Specifically, the server can search for the historical attributes and the historical recommendation target attributes corresponding to each historical recommendation target in the database according to the historical user identifier, extract the historical attribute features based on the found historical attributes and extract the historical recommendation target features corresponding to the historical recommendation target attributes to obtain the historical attribute features corresponding to the historical user identifier and the historical recommendation target features corresponding to each historical recommendation target. Then the server inputs the historical attribute features and each historical recommendation target feature into the trained sub-target models respectively to obtain the set of sub-recommendation degrees corresponding to the historical user identifier output by each trained sub-target model. For example, the server obtains a sample pair of the user and the historical recommendation target according to the historical attribute features and each historical recommendation target feature, and this sample pair includes the historical attribute features and a historical recommendation target feature. Input each sample pair into each trained sub-target model for calculation to obtain the set of sub-recommendation degrees corresponding to each sample pair output by each trained sub-target model.

[0109] In one embodiment, the training samples can include multiple historical user identifiers and each historical recommendation target corresponding to each historical user identifier. Input the historical attribute features and the historical recommendation target features corresponding to each historical user identifier into the trained sub-target models to obtain the set of sub-recommendation degrees corresponding to each historical user identifier output by each trained sub-target model.

[0110] In the above embodiments, by inputting the historical attribute features and each historical recommendation target feature into the trained sub-goal model, a set of sub-recommendation degrees corresponding to the historical user identifiers output by each trained sub-goal model is obtained. That is, by using the historical attribute features and historical recommendation target features as the input of the model, the obtained sub-recommendation degrees can be made more accurate.

[0111] In one embodiment, determining the sorting evaluation information corresponding to each sub-label sequence based on a sorting evaluation index includes:

[0112] Obtaining the label data type corresponding to each sub-label sequence, determining the sorting evaluation index corresponding to each sub-label sequence based on the label data type, and calculating the sorting evaluation information corresponding to each sub-label sequence based on the sorting evaluation index corresponding to each sub-label sequence.

[0113] Among them, the label data type is used to characterize the data type corresponding to the label.

[0114] Specifically, the server obtains the label data type corresponding to each sub-label sequence. The label data type corresponding to the sub-label can be pre-set and stored in the server. Different sub-goal models can have different label data types corresponding to their sub-labels. Among them, the label corresponding to the classification model can be set to a discrete data type. For example, the label corresponding to the click model can be set to a discrete data type. The label corresponding to the linear model can be set to a continuous data type. For example, the label corresponding to the video stay duration model can be set to a continuous data type. The server determines the sorting evaluation index corresponding to each sub-label sequence based on the label data type. Among them, the server can determine that the sorting evaluation index corresponding to the sub-label sequence according to the discrete data type can be AUC, or other indexes. The server can determine that the sorting evaluation index corresponding to the sub-label sequence according to the discrete data type can be the proportion of positive-order pairs. Then the server calculates the sorting evaluation information corresponding to each sub-label sequence based on the sorting evaluation index corresponding to each sub-label sequence.

[0115] In a specific embodiment, define the sorting index of the k-th sub-label sequence as P k , where k is a positive integer. When the sub-label sequence is of discrete data type, P k is calculated using AUC. When the sub-label sequence is of continuous data type, P k is calculated using the proportion of positive-order pairs. Define the k-th sub-label sequence corresponding to the i-th historical user identifier as , where i is a positive integer. Define the fusion score sequence corresponding to the i-th historical user identifier as W refers to the model parameters of the unfinalized fusion recommendation model. Then, calculating the sorting evaluation information of the k-th sub-label sequence corresponding to the i-th historical user identifier is specifically shown in formula (1):

[0116]

[0117] where p ki (W) represents the sorting evaluation information of the k-th sub-tag sequence corresponding to the i-th historical user identifier. It represents that according to the fusion score sequence corresponding to the i-th historical user identifier it is determined that the k-th sub-tag sequence corresponding to the i-th historical user identifier is Then, the sorting index of the k-th sub-tag sequence is used as P k to calculate the sorting evaluation information of the k-th sub-tag sequence corresponding to the i-th historical user identifier.

[0118] In the above embodiment, by determining the sorting evaluation index corresponding to each sub-tag sequence based on the tag data type, and then calculating the sorting evaluation information corresponding to each sub-tag sequence based on the sorting evaluation index corresponding to each sub-tag sequence, the obtained sorting evaluation information corresponding to each sub-tag sequence is more accurate.

[0119] In one embodiment, the tag data type includes a discrete data type; as Figure 3 shown, determining the sorting evaluation index corresponding to each sub-tag sequence based on the tag data type, and calculating the sorting evaluation information corresponding to each sub-tag sequence based on the sorting evaluation index corresponding to each sub-tag sequence includes:

[0120] Step 302, when the tag data type corresponding to the first sub-tag sequence is a discrete data type, determine the number of first-category sub-tags and the number of second-category sub-tags in the first sub-tag sequence.

[0121] Among them, the first sub-tag sequence refers to the sub-tag sequence of the discrete data type. The number of first-category sub-tags refers to the tags representing the first category in the tags of the binary classification model. For example, the click tag in the click model. The number of first-category sub-tags refers to the number of first-category sub-tags in the first sub-tag sequence. The second-category sub-tag refers to the tag representing the second category in the tags of the binary classification model, for example, the non-click tag in the click model. The number of second-category sub-tags refers to the number of second-category sub-tags in the second sub-tag sequence.

[0122] Specifically, when the server determines that the tag data type corresponding to the first sub-tag sequence is a discrete data type, it counts the number of first-category tags and second-category tags in the first sub-tag sequence to obtain the number of first-category sub-tags and the number of second-category sub-tags.

[0123] Step 304: Determine the historical recommended target position identifiers corresponding to each first-category sub-tag from the historical recommended target sequence, and calculate the sum of the identifiers of the historical recommended target position identifiers corresponding to each first-category sub-tag.

[0124] Among them, the historical recommended target position identifier is used to uniquely identify the position of the historical recommended target corresponding to the first-category sub-tag in the historical recommended target sequence, which can be a number, a code, etc. The sum of the identifiers is obtained by summing up each historical recommended target position identifier.

[0125] Specifically, the server determines the historical recommended target position identifiers corresponding to each first-category sub-tag from the historical recommended target sequence, and then adds up the historical recommended target position identifiers corresponding to each first-category sub-tag to obtain the sum of the identifiers.

[0126] For example, in the video recommendation application scenario, the historical recommended sequence is (Video 3, Video 2, Video 5, Video 1, Video 4). The corresponding sub-tag sequence is (1, 0, 0, 1, 1), where 1 represents the first-category sub-tag and 0 represents the second-category sub-tag. The historical recommended targets corresponding to the first-category sub-tags are Video 3, Video 1, and Video 4. The position of Video 3 in the sequence is the first, so the position identifier is 1. The position of Video 1 in the sequence is the 4th, so the position identifier is 4. The position of Video 4 in the sequence is the 4th, so the position identifier is 5. Adding the position identifier 1, the position identifier 4, and the position identifier 5, the sum of the identifiers is 10.

[0127] In one embodiment, the server can also sort the fusion recommendation degrees corresponding to each historical recommended target from small to large to obtain a fusion recommendation degree sequence, use the fusion recommendation degree sequence to determine the fusion recommendation degree position identifiers corresponding to each first-category sub-tag, and calculate the sum of the identifiers of the fusion recommendation degree position identifiers corresponding to each first-category sub-tag.

[0128] Step 306: Calculate the first sorting evaluation information corresponding to the first sub-tag sequence based on the number of first-category tags, the number of second-category tags, and the sum of the identifiers.

[0129] Among them, the first sorting evaluation information refers to the sorting evaluation information corresponding to the first sub-tag sequence.

[0130] Specifically, the server calculates the first sorting evaluation information corresponding to the first sub-tag sequence using the number of first-category tags, the number of second-category tags, and the sum of the identifiers, and specifically can be calculated using the following formula (2).

[0131]

[0132] Among them, ACU1 refers to the first sorted evaluation information. M refers to the number of first category labels, and N refers to the number of second category labels. refers to the identification sum. y(i) represents the sub-label corresponding to the i-th historical recommended target, and i is a positive integer. y(i) ∈ pos represents the first category label. rank(x i ) represents the position identification of the i-th historical recommended target. x i represents the i-th training sample, that is, by calculating the product of the number of first category labels and the number of first category labels plus 1, then calculating the ratio of the product to the preset value 2, and calculating the difference between the identification sum and the ratio, so as to use the ratio of the difference to the product of the number of first category labels and the number of second category labels as the first sorted evaluation information.

[0133] In the above embodiment, by calculating the first sorted evaluation information corresponding to the first sub-label sequence through the number of first category labels, the number of second category labels, and the identification sum, the first sorted evaluation information can be quickly calculated, improving the efficiency of obtaining the first sorted evaluation information.

[0134] In one embodiment, the label data type includes a continuous data type; as Figure 4 shown, determining the sorting evaluation indicators corresponding to each sub-label sequence based on the label data type, and calculating the sorting evaluation information corresponding to each sub-label sequence based on the sorting evaluation indicators corresponding to each sub-label sequence, including:

[0135] Step 402, when the label data type corresponding to the second sub-label sequence is a continuous data type, calculate the number of positive pairs and the total number of sequence pairs in the second sub-label sequence.

[0136] Step 406, calculate the ratio of the number of positive pairs to the total number of sequence pairs to obtain the second sorted evaluation information corresponding to the second sub-label sequence.

[0137] Among them, a positive pair means that in a sequence, if the number in front of the sorting is greater than the number behind the sorting, then these two numbers are called a positive pair. The second sub-label sequence refers to the sub-label sequence corresponding to the continuous data type. The number of positive pairs refers to the number of positive pairs included in the second sub-label sequence. The total number of sequence pairs refers to the total number of sequence pairs included in the second sub-label sequence. The second sorted evaluation information refers to the sorting evaluation information corresponding to the second sub-label sequence.

[0138] Specifically, when the server determines that the tag data type corresponding to the second sub-tag sequence is a continuous data type, the server calculates the number of positive pairs in the second sub-tag sequence. That is, the server can sequentially traverse each sub-tag in the second sub-tag sequence and compare it with the sub-tags after sorting to obtain the number of positive pairs. Then the server can obtain the total number of sequence pairs by calculating the combination number in the second sub-tag sequence. At this time, the server calculates the ratio of the number of positive pairs to the total number of sequence pairs and uses this ratio as the second sorting evaluation information corresponding to the second sub-tag sequence.

[0139] In one embodiment, as Figure 5 shown, step 402, calculating the number of positive pairs in the second sub-tag sequence, includes the steps:

[0140] Step 502, dividing the second sub-tag sequence to obtain a second left sub-tag sequence and a second right sub-tag sequence.

[0141] Step 502, calculating the first number of positive pairs of the second left sub-tag sequence and calculating the second number of positive pairs of the second right sub-tag sequence.

[0142] Step 502, calculating the number of interactive positive pairs between the second left sub-tag sequence and the second right sub-tag sequence, and determining the number of positive pairs based on the first number of positive pairs, the second number of positive pairs, and the number of interactive positive pairs.

[0143] Among them, the second left sub-tag sequence refers to the sub-tag sequence of the first part after division. The second right sub-tag sequence refers to the sub-tag sequence of the second part after division. The first number of positive pairs refers to the number of positive pairs in the second left sub-tag sequence. The second number of positive pairs refers to the number of positive pairs in the second right sub-tag sequence. The number of interactive positive pairs refers to the number of positive pairs between the second left sub-tag sequence and the second right sub-tag sequence.

[0144] Specifically, the server uses a recursive algorithm to calculate the number of positive pairs in the second sub-tag sequence. That is, the second sub-tag sequence is divided to obtain a second left sub-tag sequence and a second right sub-tag sequence, and then the second left sub-tag sequence and the second right sub-tag sequence are respectively divided until there is only one sub-tag in the second left sub-tag sequence and the second right sub-tag sequence. At this time, the server counts the first number of positive pairs of the second left sub-tag sequence, calculates the second number of positive pairs of the second right sub-tag sequence, and then calculates the number of interactive positive pairs between the second left sub-tag sequence and the second right sub-tag sequence. Then the server calculates the sum of the first number of positive pairs, the second number of positive pairs, and the number of interactive positive pairs to obtain the number of positive pairs.

[0145] In a specific embodiment, as Figure 6As shown in the figure, in a video recommendation scenario, the obtained sub-label sequence of video dwell time is (4, 6, 5, 7, 8, 1, 2, 3), and z represents the number of positive order pairs. The sub-label sequence is recursively partitioned until there is only one sub-label in the second sub-label left sequence and the second sub-label right sequence. Among them, the number of positive order pairs in the second sub-label left sequence (4) is 0. The number of positive order pairs in the second sub-label right sequence (6) is 0. Then, the interactive positive order pair number between (4) and (6) is calculated by merge sorting to be 0. Until all the partitioned sequences are calculated, the number of positive order pairs of the sub-label sequence of video dwell time (4, 6, 5, 7, 8, 1, 2, 3) is 16.

[0146] In the above embodiment, when the label data type corresponding to the second sub-label sequence is a continuous data type, calculate the number of positive order pairs and the total number of sequence pairs in the second sub-label sequence. Calculate the ratio of the number of positive order pairs to the total number of sequence pairs to obtain the second sorting evaluation information corresponding to the second sub-label sequence, so that the obtained second sorting evaluation information is more accurate.

[0147] In one embodiment, step 210, determining the target sorting evaluation information based on the sorting evaluation information corresponding to each sub-label sequence includes:

[0148] Obtain the preset weights corresponding to each trained sub-target model, and perform weighted calculation on the sorting evaluation information of each sub-label sequence based on the preset weights corresponding to each trained sub-target model to obtain the first target sorting evaluation information.

[0149] Among them, the preset weight is the weight occupied by the trained sub-target model set in advance. Different trained sub-target models can be set with different weights. The first target sorting evaluation information refers to the target sorting evaluation information obtained by weighting the sorting evaluation information of each sub-label sequence.

[0150] Specifically, the server pre-sets the preset weights corresponding to each trained sub-target model. When in use, the server can obtain the preset weights corresponding to each trained sub-target model from the memory. The server can obtain the preset weights corresponding to each trained sub-target model input by the user through the terminal. Then the server uses the preset weights corresponding to each trained sub-target model to weight and sum the sorting evaluation information of each sub-label sequence to obtain the first target sorting evaluation information. In a specific embodiment, the following formula (3) can be used to calculate the first target sorting evaluation information.

[0151]

[0152] Among them, W represents the model parameters of the unfinalized fusion recommendation model, and P(W) represents the first target ranking evaluation information of the fusion recommendation model established using the model parameters W. For example, when the model parameters W are the initialized model parameters, the first target ranking evaluation information of the initial fusion recommendation model can be calculated using formula (2). K represents the total number of trained sub-goal models. p i (W) refers to the ranking evaluation information of the sub-label sequence corresponding to the i-th trained sub-goal model. θ i represents the preset weight corresponding to the i-th trained sub-goal model. represents calculating the sum by weighting the ranking evaluation information of each sub-label sequence based on the preset weights corresponding to each trained sub-goal model.

[0153] In the above embodiment, by obtaining the preset weights corresponding to each trained sub-goal model and performing weighted calculation on the ranking evaluation information of each sub-label sequence based on the preset weights corresponding to each trained sub-goal model, the first target ranking evaluation information is obtained, further making the obtained first target ranking evaluation information more accurate.

[0154] In one embodiment, the training samples include each historical user identifier and each historical recommendation target corresponding to each historical user identifier;

[0155] Such as Figure 7 shown, before obtaining the preset weights corresponding to each trained sub-goal model, it further includes:

[0156] Step 702, obtaining the ranking evaluation information corresponding to each sub-label sequence of each historical user identifier, and obtaining the total number of historical users.

[0157] Step 704, performing an average calculation based on the ranking evaluation information corresponding to each sub-label sequence of each historical user identifier and the total number of historical users to determine the average ranking evaluation information corresponding to each sub-label sequence.

[0158] Among them, the total number of historical users refers to the total number of historical user identifiers. The average ranking evaluation information refers to the averaged ranking evaluation information corresponding to the sub-label sequence.

[0159] Specifically, different historical users have different historical recommendation goals. The training samples include various historical user identifiers and the respective historical recommendation goals corresponding to each historical user identifier. Obtain the sub-label sets corresponding to each historical user identifier for at least two trained sub-goal models. Then, input the training samples including various historical user identifiers and the respective historical recommendation goals corresponding to each into the trained sub-goal models to obtain the sub-recommendation degree sets corresponding to each historical user identifier output by each trained sub-goal model. Input the sub-recommendation degree sets corresponding to each historical user identifier into the initial fusion recommendation model to obtain the fusion recommendation degree sets corresponding to each historical user identifier, and sort the historical recommendation goals corresponding to each historical user identifier based on the fusion recommendation degree sets to obtain the historical recommendation goal sequences corresponding to each historical user identifier. Sort the sub-labels corresponding to the respective historical recommendation goals in the sub-label sets according to the historical recommendation goal sequences corresponding to each historical user identifier to obtain the sub-label sequences corresponding to each historical user identifier for each trained sub-goal model. Evaluate each sub-label sequence corresponding to each historical user identifier based on the sorting evaluation index to obtain the sorting evaluation information corresponding to each sub-label sequence of each historical user identifier. At this time, the server counts the total number of historical user identifiers to obtain the total number of historical users. Calculate the sum of the sorting evaluation information corresponding to the same sub-label sequence for each historical user identifier, and then calculate the ratio of the sum of the sorting evaluation information corresponding to the same sub-label sequence to the total number of historical users to obtain the average sorting evaluation information corresponding to each sub-label sequence. The same sub-label sequence refers to the sub-label sequence corresponding to the same trained sub-goal model.

[0160] In a specific embodiment, the average sorting evaluation information can be calculated using the following formula (4)

[0161]

[0162] where U represents the total number of historical users, p k (W) represents the average sorting evaluation information corresponding to the k-th sub-label sequence. p ki (W) represents the sorting evaluation information corresponding to the k-th sub-label sequence of the i-th historical user identifier. represents calculating the sum of the sorting evaluation information corresponding to the k-th sub-label sequence for each historical user identifier. represents the ratio of the sum of the sorting evaluation information corresponding to the k-th sub-label sequence of each historical user identifier to the total number of historical users

[0163] Perform weighted calculation based on the preset weights corresponding to each trained sub-goal model and the sorting evaluation information of each sub-label sequence to obtain the target sorting evaluation information, including:

[0164] Step 706: Perform weighted calculation based on the preset weights corresponding to each trained sub-goal model and the average ranking evaluation information corresponding to each sub-label sequence to obtain the second target ranking evaluation information.

[0165] Among them, the second target ranking evaluation information refers to the target ranking evaluation information obtained by weighting the average ranking evaluation information of each sub-label sequence.

[0166] Specifically, the server performs weighted summation calculation on the average ranking evaluation information corresponding to each sub-label sequence according to the preset weights corresponding to each trained sub-goal model set, to obtain the second target ranking evaluation information.

[0167] In the above embodiment, by calculating the average ranking evaluation information and then using the average ranking evaluation information for weighted calculation to obtain the second target ranking evaluation information, it can further make the obtained second target ranking evaluation information more accurate.

[0168] In one embodiment, the training samples include each historical user identifier and each historical recommended target corresponding to each historical user identifier;

[0169] Such as Figure 8 shown, before obtaining the preset weights corresponding to each trained sub-goal model, it further includes:

[0170] Step 802: Obtain the ranking evaluation information corresponding to each sub-label sequence of each historical user identifier, and obtain the number of historical recommended targets corresponding to each historical user identifier.

[0171] Step 804: Perform weighted calculation on the ranking evaluation information corresponding to each sub-label sequence of each historical user identifier based on the number of historical recommended targets corresponding to each historical user identifier, to obtain the weighted ranking evaluation information corresponding to each sub-label sequence.

[0172] Among them, the weighted ranking evaluation information refers to the ranking evaluation information corresponding to each sub-label sequence obtained by weighting using the number of historical recommended targets corresponding to each historical user.

[0173] Specifically, the server can pre-calculate the sorting evaluation information corresponding to each sub-tag sequence of each historical user identifier and save the sorting evaluation information corresponding to each sub-tag sequence of each historical user identifier. Then, when needed, it can be directly obtained. The server can also obtain from a third party the sorting evaluation information corresponding to each sub-tag sequence of each historical user identifier. This third party is a service provider for providing the sorting evaluation information corresponding to each sub-tag sequence of the historical user identifier. The server obtains the historical recommendation targets corresponding to each historical user identifier and then counts the number of historical recommendation targets corresponding to each historical user identifier. The number of historical recommendation targets corresponding to different historical user identifiers is different. For example, there may be a historical user identifier corresponding to 2 historical recommendation targets. The server can also directly obtain from a third party the number of historical recommendation targets corresponding to each historical user identifier, and this third party can also be used to provide the number of historical recommendation targets. Then, the server respectively performs weighted calculation on the sorting evaluation information corresponding to each sub-tag sequence of each historical user identifier using the corresponding number of historical recommendation targets to obtain the weighted sorting evaluation information corresponding to each sub-tag sequence.

[0174] Step 806, calculate the total number of historical recommendation targets based on the number of historical recommendation targets corresponding to each historical user, calculate the ratio of the weighted sorting evaluation information corresponding to each sub-tag sequence to the total number of historical recommendation targets, and obtain the specific sorting evaluation information corresponding to each sub-tag sequence.

[0175] Among them, the specific sorting evaluation information refers to the sorting evaluation information obtained after weighted averaging the sorting evaluation information corresponding to each sub-tag sequence using the number of historical recommendation targets.

[0176] Specifically, the server calculates the sum of the number of historical recommendation targets corresponding to each historical user to obtain the total number of historical recommendation targets. Then, it adds up the weighted sorting evaluation information corresponding to the same sub-tag sequence of each historical user to obtain the total sum of the weighted sorting evaluation information corresponding to each sub-tag sequence. Then, the server calculates the ratio of the total sum of the weighted sorting evaluation information to the total number of historical recommendation targets to obtain the specific sorting evaluation information corresponding to each sub-tag sequence.

[0177] In a specific embodiment, the specific sorting evaluation information corresponding to each sub-tag sequence can be calculated using the following formula (5).

[0178]

[0179] Among them, p k (W) represents the specific sorting evaluation information corresponding to the Kth sub-tag sequence. m represents the number of historical recommendation targets. m i represents the number of historical recommendation targets corresponding to the ith sub-tag sequence. pki (W) represents the sorting evaluation information corresponding to the k-th sub-tag sequence of the i-th historical user identifier. It represents calculating the weighted sum of the sorting evaluation information corresponding to the k-th sub-tag sequence of each historical user identifier. It represents the total number of historical recommendation targets. That is, by calculating the ratio of the weighted sum of the sorting evaluation information corresponding to the k-th sub-tag sequence of each historical user identifier to the total number of historical recommendation targets, the specific sorting evaluation information corresponding to the k-th sub-tag sequence is obtained.

[0180] Based on the preset weights corresponding to each trained sub-goal model and the sorting evaluation information of each sub-tag sequence, weighted calculation is performed to obtain the target sorting evaluation information, including:

[0181] Step 808, based on the preset weights corresponding to each trained sub-goal model and the specific sorting evaluation information corresponding to each sub-tag sequence, weighted calculation is performed to obtain the third target sorting evaluation information.

[0182] Among them, the third target sorting evaluation information refers to the sorting evaluation information obtained by performing weighted calculation on the specific sorting evaluation information corresponding to each sub-tag sequence.

[0183] Specifically, the server uses the preset weights corresponding to each trained sub-goal model and the specific sorting evaluation information of the corresponding each sub-tag sequence to perform weighted calculation to obtain the third target sorting evaluation information.

[0184] In the above embodiment, by calculating the specific sorting evaluation information of each sub-tag sequence, and then using the specific sorting evaluation information for weighted calculation to obtain the third target sorting evaluation information, it avoids the situation where the sorting evaluation information is not accurate enough when the number of historical recommendation targets corresponding to the historical user identifier is extremely small. For example, when there are only 2 historical recommendation targets for the historical user identifier, the AUC or the proportion of positive pairs will show extreme situations such as being equal to 1 or 0. By calculating the specific sorting evaluation information, the error influence brought by extreme situations is avoided, and the accuracy of the third target sorting evaluation information is further improved.

[0185] In one embodiment, before calculating the target sorting evaluation information, the historical user identifier can be preprocessed. For example, filter out the historical user identifiers whose corresponding historical recommendation targets are less than the preset number, so as to make the obtained target sorting evaluation information more accurate.

[0186] In one embodiment, as Figure 9 shown, step 212, update the initial fusion recommendation model based on the target sorting evaluation information. When the training is completed, the target fusion recommendation model is obtained, including:

[0187] Step 902, when the initial fusion recommendation model meets the preset conditions, simulate and calculate the simulated gradient of the initial model parameters in the initial fusion recommendation model based on the target ranking evaluation information.

[0188] Among them, the preset conditions refer to the evaluation index conditions that the results output by the preset fusion recommendation model meet. The preset conditions can include multiple ones, which can be respectively R 1 (W(,R 2 (W),...,R Q (W). Among them, Q identifies the number of preset conditions. R represents the preset conditions. In a specific embodiment, for example, in the video recommendation scenario, the preset conditions may include that the proportion of videos with a video duration exceeding the preset duration among the videos recommended to the user should exceed the preset duration proportion threshold. The preset conditions may include that the proportion of new videos among the videos recommended to the user should exceed the preset new video proportion threshold. The preset conditions can be specifically set according to the application scenario. The simulated gradient refers to the gradient of the model parameters in the initial fusion recommendation model obtained by simulating the standard gradient descent.

[0189] Specifically, the server determines whether the initial fusion recommendation model meets the preset conditions, that is, the server obtains the historical recommendation target sequence through the initial fusion recommendation model, and determines whether the historical recommendation target sequence meets the preset conditions. When it meets the preset conditions, based on the target ranking evaluation information, calculate the partial derivative of each initial model parameter in the initial fusion recommendation model through simulation to obtain the simulated gradient of the initial model parameters in the initial fusion recommendation model.

[0190] Step 904, update the initial model parameters in the initial fusion recommendation model based on the simulated gradient and the preset learning rate to obtain an updated fusion recommendation model.

[0191] Among them, the preset learning rate refers to the learning rate for training the preset fusion recommendation model.

[0192] Specifically, the server uses the simulated gradient and the preset learning rate to update the initial model parameters in the initial fusion recommendation model to obtain an updated fusion recommendation model. Among them, the initial model parameters in the initial fusion recommendation model can be updated using the following formula (6).

[0193]

[0194] Among them, W1 represents the model parameters before update, and W2 represents the model parameters after update. λ refers to the preset learning rate, is the simulated gradient. t represents the total number of model parameters, represents the partial derivative of the model parameters. is the partial derivative of the first model parameter, w 1Refers to the first model parameter. When using formula (6), by calculating the difference between the product of each initial model parameter in the initial fusion recommendation model and the preset learning rate and the corresponding partial derivative, the updated model parameters in the initial fusion recommendation model are obtained.

[0195] Step 906: When the updated fusion recommendation model meets the preset training completion condition, obtain the target fusion recommendation model.

[0196] Specifically, when the updated fusion recommendation model does not meet the preset training completion condition, use the updated fusion recommendation model as the initial fusion recommendation model and return to step 204 to continue execution, that is, return to the step of inputting each sub-recommendation degree into the initial fusion recommendation model to obtain a set of fusion recommendation degrees. The set of fusion recommendation degrees includes the fusion recommendation degrees corresponding to each historical recommendation target. Sort each historical recommendation target based on the fusion recommendation degrees to obtain the historical recommendation target sequence. Continue to execute until the updated fusion recommendation model meets the preset training completion condition. When the updated fusion recommendation model meets the preset training completion condition, use the updated fusion recommendation model at this time as the target fusion recommendation model.

[0197] In the above embodiment, when the initial fusion recommendation model meets the preset conditions, the simulated gradient of the initial model parameters in the initial fusion recommendation model is simulated based on the target ranking evaluation information, and then the initial fusion recommendation model is updated using the simulated gradient, so as to obtain the target fusion recommendation model, improving the accuracy of the obtained target fusion recommendation model.

[0198] In one embodiment, step 902: Simulating and calculating the simulated gradient of the initial model parameters in the initial fusion recommendation model based on the target ranking evaluation information includes the steps of:

[0199] Calculating the partial derivative of the initial model parameters in the initial fusion recommendation model based on the target ranking evaluation information; determining the simulated gradient based on the partial derivative of the initial model parameters in the initial fusion recommendation model.

[0200] Specifically, the server can use the function derivative formula to calculate the partial derivative of each initial model parameter in the initial fusion recommendation model through the target ranking evaluation information, and then combine the partial derivatives of each initial model parameter in the initial fusion recommendation model to obtain the simulated gradient. Among them, the partial derivative of the initial model parameter can be calculated using formula (6) or formula (7) shown below.

[0201]

[0202] Among them, f′(x) refers to the partial derivative of the initial model parameters. x refers to the initial model parameters. f refers to the calculated target ranking evaluation information, and Δx refers to the change amount of the initial model parameters, which is a very small value. Among them, if the partial derivative of the initial model parameters is calculated using formula (6), then by calculating the ranking evaluation information of the difference between the initial model parameters and the change amount, and calculating the difference between the ranking evaluation information of this difference and the ranking evaluation information of the initial model parameters, and then by calculating the ratio of this difference to the change amount, the partial derivative of the initial model parameters is obtained. If the partial derivative of the initial model parameters is calculated using formula (7), then by calculating the ranking evaluation information of the difference between the initial model parameters and the change amount, calculating the ranking evaluation information of the sum value between the initial model parameters and the change amount, further calculating the difference between the ranking evaluation information of this difference and the ranking evaluation information of the sum value, and finally calculating the ratio between this difference and the change amount, the partial derivative of the initial model parameters is obtained.

[0203] In one embodiment, as Figure 10 shown, calculating the partial derivative of the initial model parameters in the initial fusion recommendation model based on the target ranking evaluation information includes:

[0204] Step 1002, obtain a preset first parameter micro-variable, adjust the initial model parameters of the initial fusion recommendation model based on the preset first parameter micro-variable to obtain first adjusted model parameters, and determine a first adjusted fusion recommendation model based on the first adjusted model parameters.

[0205] Among them, the preset first parameter micro-variable is a tiny change amount of the pre-set model parameters, and this preset first parameter micro-variable is used to calculate the simulated gradient. The first adjusted model parameters refer to the model parameters adjusted using the preset first parameter micro-variable. Among them, each initial model parameter of the initial fusion recommendation model can be adjusted using the preset first parameter micro-variable in sequence.

[0206] Specifically, the server obtains the preset first parameter micro-variable, and using the first parameter micro-variable to adjust the initial model parameters of the initial fusion recommendation model can be to increase the initial model parameters by the preset first parameter micro-variable or to decrease the initial model parameters by the preset first parameter micro-variable to obtain the first adjusted model parameters, and determine the first adjusted fusion recommendation model based on the first adjusted model parameters. Among them, one of the model parameters in this first adjusted fusion recommendation model is the parameter adjusted using the first parameter micro-variable, and the other model parameters are the same as the initial model parameters of the initial fusion recommendation model.

[0207] Step 1004, determine first adjusted ranking evaluation information based on the first adjusted fusion recommendation model and the training samples.

[0208] Among them, the first adjusted sorting evaluation information is used to characterize the accuracy of the first historical recommendation target sequence obtained by using the first adjusted fusion recommendation model.

[0209] Specifically, the server inputs the respective sub-recommendation degrees corresponding to the training samples into the first adjusted fusion recommendation model to obtain a first fusion recommendation degree set, sorts the respective historical recommendation targets based on the first fusion recommendation segment set to obtain a first historical recommendation target sequence, and sorts the sub-labels corresponding to the respective historical recommendation targets in the sub-label set based on the order of the first historical recommendation target sequence to obtain a first sub-label sequence corresponding to each trained sub-target model. The server determines the sorting evaluation information corresponding to each first sub-label sequence based on the sorting evaluation index, and determines the first adjusted sorting evaluation information based on the sorting evaluation information corresponding to each first sub-label sequence.

[0210] Step 1006: Calculate the difference between the first adjusted sorting evaluation information and the target sorting evaluation information, and calculate the ratio of the difference in sorting evaluation information to the preset first parameter micro-variable to obtain the partial derivative corresponding to the first adjusted model parameter.

[0211] Specifically, the server calculates the difference between the first adjusted sorting evaluation information and the target sorting evaluation information, and then calculates the ratio of the difference in sorting evaluation information to the preset first parameter micro-variable to obtain the partial derivative corresponding to the first adjusted model parameter.

[0212] In a specific embodiment, the partial derivative of the initial model parameter in the initial fusion recommendation model can also be calculated using the following formula (8).

[0213]

[0214] Among them, represents the partial derivative of the l-th initial model parameter. t represents the total number of initial model parameters. l is selected from 1 to t. Δw represents the preset first parameter micro-variable. P([w 1 ,w 2 ,...,w l-1 ,w l +Δw,w l+1 ,...,w t ) represents the adjusted sorting evaluation information obtained when the l-th initial model parameter is increased by the preset first parameter micro-variable. P([w 1 ,w 2 ...,w l-1 ,w l -Δw,w l+1 ,...,w t) represents the adjusted ranking evaluation information obtained when the l-th initial model parameter is reduced by a preset first parameter micro-variable. The partial derivatives of each initial model parameter are calculated using formula (8), and the simulated gradient obtained is shown in the following formula (9):

[0215]

[0216] where represents the partial derivative of the first initial model parameter. represents the partial derivative of the second initial model parameter. represents the partial derivative of the last initial model parameter.

[0217] In the above embodiment, by using the preset first parameter micro-variable to adjust the initial model parameters, the adjusted ranking evaluation information is calculated, and then the simulated gradient is calculated through the adjusted ranking evaluation information, so that any application scenario that needs to use the fusion recommendation model for target recommendation can be applicable, expanding the application scenario.

[0218] In one embodiment, as Figure 11 shown in step 212, based on the target ranking evaluation information, the initial fusion recommendation model is updated, and when the training is completed, the target fusion recommendation model is obtained, including:

[0219] Step 1102, when the initial fusion recommendation model does not meet the preset conditions, calculate the specific evaluation index information corresponding to the preset conditions based on the historical recommended target sequence.

[0220] Among them, the specific evaluation index information is used to characterize the actual value of the evaluation index condition that the result output by the initial fusion recommendation model does not meet.

[0221] Specifically, the server determines whether the initial fusion recommendation model meets the preset conditions, that is, the server obtains the historical recommended target sequence through the initial fusion recommendation model, and determines whether the historical recommended target sequence meets the preset conditions. When the historical recommended target sequence does not meet the preset conditions, it means that the initial fusion recommendation model does not meet the preset conditions. At this time, the server calculates the specific evaluation index information corresponding to the preset conditions according to the historical recommended target sequence.

[0222] Step 1104, obtain the preset second parameter micro-variable, adjust the initial model parameters of the initial fusion recommendation model based on the preset second parameter micro-variable to obtain the second adjusted model parameters, and determine the second adjusted fusion recommendation model based on the second adjusted model parameters.

[0223] Among them, the preset second parameter micro-variable refers to the tiny change amount of the preset model parameters. The preset second parameter micro-variable can be the same as or different from the preset first parameter micro-variable. The preset second parameter micro-variable is also used to calculate the simulated gradient. The second adjusted model parameter refers to the model parameter obtained after adjustment using the preset second parameter micro-variable. The second adjusted fusion recommendation model is the fusion recommendation model obtained using the second adjusted model parameter.

[0224] Specifically, the server obtains the preset second parameter micro-variable, which can be pre-set in the server or obtained through the terminal. The server uses the preset second parameter micro-variable to adjust the initial model parameters of the initial fusion recommendation model, which can be to increase the initial model parameters by the preset second parameter micro-variable or to decrease the initial model parameters by the preset second parameter micro-variable, to obtain the second adjusted model parameter, and determines the second adjusted fusion recommendation model based on the second adjusted model parameter.

[0225] Step 1106, determine the target historical recommendation target sequence based on the second adjusted fusion recommendation model and the training sample.

[0226] Among them, the target historical recommendation target sequence is the historical recommendation target sequence obtained using the second adjusted fusion recommendation model.

[0227] Specifically, the server inputs the respective sub-recommendation degrees corresponding to the training sample into the second adjusted fusion recommendation model to obtain the second fusion recommendation degree set, and sorts the respective historical recommendation targets based on the second fusion recommendation segment set to obtain the target historical recommendation target sequence.

[0228] Step 1108, calculate the target specific evaluation index information corresponding to the preset condition based on the target historical recommendation target sequence.

[0229] Step 1110, calculate the specific evaluation information difference between the target specific evaluation index information and the specific evaluation index information, and calculate the ratio of the specific evaluation information difference to the preset second parameter micro-variable to obtain the partial derivative corresponding to the second adjusted model parameter.

[0230] Among them, the target specific evaluation index information is the specific evaluation index information corresponding to the index historical recommendation target sequence. The specific evaluation information difference refers to the information difference between the target specific evaluation index information and the specific evaluation index information.

[0231] Specifically, the server calculates the target specific evaluation index information corresponding to the preset condition according to the target historical recommendation target sequence, and then uses the function derivative formula to calculate the partial derivative corresponding to the second adjusted model parameter.

[0232] For example, when the target historical recommendation target sequence is the target historical recommendation video sequence, the proportion of new videos in the target historical recommendation video sequence can be calculated to obtain the target specific evaluation index information. Then, calculate the proportion difference between the proportion of new videos in the target historical recommendation video sequence and the proportion of new videos in the historical recommendation video sequence, and calculate the ratio of the proportion difference to the preset second parameter micro-variable to obtain the partial derivative corresponding to the second adjustment model parameter.

[0233] Step 1112, determine the target simulation gradient corresponding to the initial fusion recommendation model based on the partial derivative corresponding to the second adjustment model parameter.

[0234] Step 1114, update the initial model parameters in the initial fusion recommendation model based on the target simulation gradient and the preset target learning rate to obtain the target updated fusion recommendation model.

[0235] Among them, the target simulation gradient refers to the simulation gradient calculated using the specific evaluation index information. The preset target learning rate refers to the learning rate set in advance.

[0236] Specifically, the server combines the partial derivatives corresponding to each second adjustment model parameter to obtain the target simulation gradient corresponding to the initial fusion recommendation model. Then, calculate the parameter update amount using the target simulation gradient and the preset target learning rate, and use the parameter update amount to update the initial model parameters to obtain the target updated fusion recommendation model.

[0237] Step 1116, when the target updated fusion recommendation model meets the preset conditions, use the target updated fusion recommendation model as the initial fusion recommendation model.

[0238] Specifically, when the server continues to determine whether the target updated fusion recommendation model meets the preset conditions, when it meets the preset conditions, use the target updated fusion recommendation model as the initial fusion recommendation model. When it does not meet the preset conditions, use the target updated fusion recommendation model as the initial fusion recommendation model and return to step 1102 to continue iterating until the target updated fusion recommendation model meets the preset conditions.

[0239] In a specific embodiment, the target simulation gradient can also be calculated using the following formula (10).

[0240]

[0241] Among them, represents the partial derivative of the l-th initial model parameter, R q refers to the specific evaluation index information of the first preset condition q that is not satisfied. R q ([w 1 , w 2 ,..., w l-1 , wl +Δw, w l+1 ,..., w t ) refers to the specific evaluation index information calculated when the l-th initial model parameter is increased by a preset second parameter micro-variable. R q ([w 1 , w 2 ..., w l-1 , w l -Δw, w l+1 ,..., w t ) refers to the specific evaluation index information calculated when the l-th initial model parameter is decreased by a preset second parameter micro-variable. By using formula (10), the difference between the specific evaluation index information calculated when each initial model parameter is increased by the preset second parameter micro-variable and the specific evaluation index information calculated when it is decreased by the preset second parameter micro-variable is calculated in sequence. Then, the ratio of the difference to twice the preset second parameter micro-variable is calculated to obtain the partial derivative of each initial model parameter, and further obtain the target simulation gradient.

[0242] In the above embodiment, when the initial fusion recommendation model does not meet the preset conditions, the specific evaluation index information is used to update the parameters in the initial fusion recommendation model until the initial fusion recommendation model meets the preset conditions, so that the result output by the trained target fusion recommendation model meets the preset conditions, thereby enabling the target fusion recommendation model to meet the different requirements of different application scenarios and improving the applicability of the target fusion recommendation model.

[0243] In one embodiment, as Figure 12 shown, a recommendation method is provided. Taking the example that this method is applied to the Figure 1 server as an example for illustration, it can be understood that this method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0244] Step 1202, obtain a user identifier, and obtain attribute features based on the user identifier.

[0245] Among them, the user identifier is the unique identifier of the user to be recommended. The attribute features are used to characterize the features of the attributes, and may include basic attribute features and behavioral attribute features. Among them, the basic attribute features may be age attribute features, gender attribute features, etc. Users in different application scenarios have different behavioral attribute features. For example, in the video recommendation scenario, the behavioral attribute features may include video click features, video favorite features, viewing duration features, and so on.

[0246] Specifically, the server can obtain the user identifier of the target to be recommended, which can be the user identifier of the target to be recommended obtained from the terminal or the user identifier set in the server obtained in advance. By using this user identifier to obtain the attribute features, the server can search for the attributes corresponding to the user identifier in the database and then extract the attribute features corresponding to the attributes.

[0247] Step 1204: Obtain each target to be recommended and its corresponding target attribute features, input the attribute features and the target attribute features into at least two pre-trained sub-target models, and obtain a set of sub-target recommendation degrees output by each pre-trained sub-target model. The set of sub-target recommendation degrees includes the sub-target recommendation degrees corresponding to each target to be recommended.

[0248] Among them, the target to be recommended refers to the target that can be recommended and is saved in the server. For example, in the short video recommendation scenario, the target to be recommended can be the short video to be recommended. The target attribute feature refers to the attribute feature of the target to be recommended. In different scenarios, different targets to be recommended can have different target attribute features, which can be set according to requirements.

[0249] Specifically, at least two pre-trained sub-target models are pre-deployed in the server. When recommendation is needed, the server obtains each target to be recommended and its corresponding target attribute features, inputs the attribute features and the target attribute features into at least two pre-trained sub-target models, and obtains a set of sub-target recommendation degrees output by each pre-trained sub-target model. The set of sub-target recommendation degrees includes the sub-target recommendation degrees corresponding to each target to be recommended.

[0250] Step 1206: Input each set of sub-recommendation degrees into the target fusion recommendation model to obtain the fusion recommendation degrees corresponding to each target to be recommended. The target fusion recommendation model is trained using the training samples and the sub-label sets respectively corresponding to at least two pre-trained sub-target models. The training samples include each historical recommended target, and the sub-label sets include the sub-labels corresponding to each historical recommended target.

[0251] Specifically, a target fusion recommendation model trained using any one of the above model training methods is pre-deployed in the server. At this time, the server inputs each set of sub-recommendation degrees into the target fusion recommendation model to obtain the fusion recommendation degrees corresponding to each target to be recommended. The target fusion recommendation model is trained using the training samples and the sub-label sets respectively corresponding to at least two pre-trained sub-target models. The training samples include each historical recommended target, and the sub-label sets include the sub-labels corresponding to each historical recommended target.

[0252] Step 1208: Sort each target to be recommended based on the fusion recommendation degrees to obtain a sequence of targets to be recommended.

[0253] Step 1210: Select a preset number of target candidates from the target candidate sequence to be recommended, and recommend the preset number of target candidates to the user identifier.

[0254] Specifically, the server sorts all the target candidates from largest to smallest according to the size of the fusion recommendation degree to obtain a target candidate sequence, and then sequentially selects a preset number of target candidates from the target candidate sequence, and recommends the preset number of target candidates to the terminal corresponding to the user identifier.

[0255] In the above embodiment, by using the target fusion recommendation model to fuse each sub-recommendation degree, the obtained fusion recommendation degree is more accurate, and then a preset number of target candidates are selected according to the fusion recommendation degree and recommended to the user identifier, improving the accuracy of the recommendation.

[0256] In a specific embodiment, as Figure 13 shown, a method for training a recommendation model is provided, including the following steps:

[0257] Step 1302: Obtain training samples, where the training samples include each historical recommendation target, and obtain sub-label sets respectively corresponding to at least two trained sub-target models, and the sub-label sets include sub-labels corresponding to each historical recommendation target.

[0258] Step 1304: Input the training samples into the trained sub-target models to obtain a set of sub-recommendation degrees output by each trained sub-target model, and the set of sub-recommendation degrees includes sub-recommendation degrees corresponding to each historical recommendation target.

[0259] Step 1306: Input each set of sub-recommendation degrees into the initial fusion recommendation model to obtain a set of fusion recommendation degrees, and the set of fusion recommendation degrees includes fusion recommendation degrees corresponding to each historical recommendation target. Sort each historical recommendation target based on the fusion recommendation degree to obtain a historical recommendation target sequence.

[0260] Step 1308: Sort the sub-labels corresponding to each historical recommendation target in the sub-label sets based on the order of the historical recommendation target sequence to obtain sub-label sequences respectively corresponding to each trained sub-target model;

[0261] Step 1310a: When the label data type corresponding to the first sub-label sequence is a discrete data type, determine the number of first-category sub-labels and the number of second-category sub-labels from the first sub-label sequence, determine the historical recommendation target position identifiers corresponding to each first-category sub-label from the historical recommendation target sequence, calculate the sum of the identifiers of the historical recommendation target position identifiers corresponding to each first-category sub-label, and calculate the first sorting evaluation information corresponding to the first sub-label sequence based on the number of first-category labels, the number of second-category labels, and the sum of the identifiers.

[0262] Step 1310b, when the tag data type corresponding to the second sub-tag sequence is a continuous data type, calculate the number of positive pairs and the total number of sequence pairs in the second sub-tag sequence, and calculate the ratio of the number of positive pairs to the total number of sequence pairs to obtain the second sorting evaluation information corresponding to the second sub-tag sequence.

[0263] Step 1312, obtain the sorting evaluation information corresponding to each sub-tag sequence of each historical user identifier, and obtain the total number of historical users. Based on the sorting evaluation information corresponding to each sub-tag sequence of each historical user identifier and the total number of historical users, perform an average calculation to determine the average sorting evaluation information corresponding to each sub-tag sequence. Based on the preset weights corresponding to each trained sub-goal model and the average sorting evaluation information corresponding to each sub-tag sequence, perform a weighted calculation to obtain the second target sorting evaluation information.

[0264] Step 1314, when the initial fusion recommendation model meets the preset conditions, obtain the preset first parameter micro-variable, adjust the initial model parameters of the initial fusion recommendation model based on the preset first parameter micro-variable to obtain the first adjusted model parameters, and determine the first adjusted fusion recommendation model based on the first adjusted model parameters; determine the first adjusted sorting evaluation information based on the first adjusted fusion recommendation model and the training samples; calculate the sorting evaluation information difference between the first adjusted sorting evaluation information and the target sorting evaluation information, and calculate the ratio of the sorting evaluation information difference to the preset first parameter micro-variable to obtain the partial derivative corresponding to the first adjusted model parameters, and determine the simulated gradient based on the partial derivative of the initial model parameters in the initial fusion recommendation model.

[0265] Step 1316, update the initial model parameters in the initial fusion recommendation model based on the simulated gradient and the preset learning rate to obtain the updated fusion recommendation model. When the updated fusion recommendation model reaches the preset training completion condition, obtain the target fusion recommendation model.

[0266] In a specific embodiment, a recommendation model training method is provided, which specifically includes:

[0267] Obtain a specific sample x, which includes a sample pair composed of an attribute feature and a historical recommendation target feature. For the specific sample x, the outputs of each trained sub-goal model are f i (x), i = 1, 2... n, where n represents the total number of trained sub-goal models.

[0268] A large number of sample data are obtained, which altogether include U user identifiers. The jth specific sample of the ith user identifier is x ij。 The ith user identifier corresponds to m i specific samples, and these specific samples include the attribute feature and m i historical recommendation target features. The sub-tag set of the Kth trained sub-goal model is yijk The sub - recommendation degree calculated using the K - th trained sub - goal model is s ijk Randomly initialize the model parameters W of the fusion recommendation model to obtain the initial fusion recommendation model. Use the initial fusion recommendation model to obtain a specific sample as x ij The corresponding fusion recommendation degree is g ij = G(s ij1 , s ij2 ... s ijn , W), where G represents the fusion recommendation model, and the calculated fusion recommendation degree sequence corresponding to the i - th user identifier is According to the fusion recommendation degree sequence Obtain the sub - label sequence of the K - th trained sub - goal model as Fusion recommendation degree sequence And the sub - label sequence Have the same length of m i At this time, use the sorting evaluation index to calculate the sorting evaluation information p ki (W) corresponding to the sub - label sequence of each trained sub - goal model. Then calculate the average sorting evaluation information p k (W) of all user identifiers. Obtain the weights of each trained sub - goal model as θ 1 ... θ n . Weight the average sorting evaluation information p k (W) through the weights of each trained sub - goal model to obtain p(W).

[0269] Then calculate whether the initial fusion recommendation model under the initial model parameters W meets each preset condition. When all preset conditions are met, use formula (8) to calculate the partial derivative of each initial model parameter in the initial fusion recommendation model, obtain the simulated gradient according to the partial derivative of each initial model parameter using formula (9), and then use formula (6) to update the initial model parameters in the initial fusion recommendation model to obtain the updated fusion recommendation model, and then continuously perform cyclic iteration. When the training completion condition is reached, obtain the target fusion recommendation model. When the preset conditions are not met, obtain the first specific evaluation index that is not met, then use formula (10) to calculate the partial derivative of each initial model parameter, obtain the simulated gradient according to the partial derivative of each initial model parameter, use the simulated gradient to update the initial model parameters, obtain the updated fusion recommendation model, and continuously perform cyclic iteration until the updated fusion recommendation model meets all the preset conditions. At this time, continue to execute the steps after meeting all the preset conditions to obtain the target fusion recommendation model.

[0270] This application also provides an application scenario that applies the above - mentioned fusion model training method.

[0271] Specifically, the application of the fusion model training method in this application scenario is as follows:

[0272] In the video recommendation application scenario, the server obtains training samples, where the training samples include various historical recommended videos, and obtains the sub-label sets corresponding to the completion rate video recommendation model, the like rate video recommendation model, and the follow rate video recommendation model respectively. The sub-label sets include the sub-labels corresponding to each historical recommended video. For example, the sub-label corresponding to the completion rate video recommendation model can be the user's video viewing duration. The sub-label corresponding to the like rate video recommendation model can be the label corresponding to whether the user likes it. The sub-label corresponding to the follow rate video recommendation model can be the label corresponding to whether the user follows it. Input the training samples into the completion rate video recommendation model, the like rate video recommendation model, and the follow rate video recommendation model simultaneously to obtain the output recommendation score set, where the recommendation score set includes the recommendation scores corresponding to each historical recommended video. Input the recommendation scores corresponding to each historical recommended video into the initial fusion recommendation model to obtain the fusion score set, where the fusion score set includes the fusion scores corresponding to each historical recommended video. Sort each historical recommended video based on the fusion scores to obtain the historical recommended video sequence. Sort the sub-labels corresponding to each historical recommended video in the sub-label set according to the order of the historical recommended video sequence to obtain the sub-label sequences corresponding to the completion rate video recommendation model, the like rate video recommendation model, and the follow rate video recommendation model respectively. Determine the sorting evaluation information corresponding to each sub-label sequence based on the sorting evaluation index, and determine the target sorting evaluation information based on the sorting evaluation information corresponding to each sub-label sequence; update the initial fusion recommendation model based on the target sorting evaluation information. When the training is completed, obtain the target fusion recommendation model. Deploy the target fusion recommendation model to the server for video recommendation.

[0273] This application also provides another application scenario, which applies the above-mentioned fusion model training method. Specifically, the application of the fusion model training method in this application scenario is as follows:

[0274] In the advertising recommendation application scenario, the server obtains training samples, where the training samples include various historical recommended ads, and obtains the sub-label sets corresponding to the ad click-through rate recommendation model, the ad viewing duration recommendation model, and the ad video playback rate recommendation model respectively. The sub-label sets include the sub-labels corresponding to each historical recommended ad. For example, the sub-label corresponding to the ad click-through rate can be the label corresponding to whether the user clicks on the ad. The sub-label corresponding to the ad viewing duration recommendation model can be the label corresponding to the duration for which the user views the ad. The sub-label corresponding to the ad video playback rate recommendation model can be the label corresponding to whether the user plays the ad video. The training samples are input into the ad click-through rate recommendation model, the ad viewing duration recommendation model, and the ad video playback rate recommendation model simultaneously to obtain the output ad recommendation score set, where the ad recommendation score set includes the ad recommendation scores corresponding to each historical recommended ad. The ad recommendation scores corresponding to each historical recommended ad are input into the initial fusion recommendation model to obtain the ad fusion score set, where the ad fusion score set includes the ad fusion scores corresponding to each historical recommended ad. Based on the ad fusion scores, each historical recommended ad is sorted to obtain the historical recommended ad sequence. Based on the order of the historical recommended ad sequence, the sub-labels corresponding to each historical recommended ad in the sub-label set are sorted to obtain the sub-label sequences corresponding to the ad click-through rate recommendation model, the ad viewing duration recommendation model, and the ad video playback rate recommendation model respectively. Based on the sorting evaluation index, the sorting evaluation information corresponding to each sub-label sequence is determined, and based on the sorting evaluation information corresponding to each sub-label sequence, the target sorting evaluation information is determined; based on the target sorting evaluation information, the initial fusion recommendation model is updated. When the training is completed, the target fusion recommendation model is obtained. The target fusion recommendation model is deployed to the server for advertising recommendation.

[0275] It should be understood that although each step in the flowchart in Figures 2 - 5 and Figures 7 - 13 is shown in sequence according to the arrow indication, these steps do not necessarily need to be executed in the order indicated by the arrow. Unless otherwise clearly stated in this article, there is no strict order limit for the execution of these steps, and these steps can be executed in other orders. Moreover, Figures 2 - 5 and Figures 7 - 13 at least a part of the steps in

[0276] In one embodiment, as shown in Figure 14As shown, a recommendation model training device is provided. This device can be a software module, a hardware module, or a combination of both, and becomes part of a computer device. Specifically, the device includes: a sample acquisition module 1402, a sub - recommendation degree obtaining module 1404, a target sequence obtaining module 1406, a subsequence obtaining module 1408, an evaluation module 1410, and an update module 1412, where:

[0277] The sample acquisition module 1402 is used to acquire training samples. The training samples include various historical recommendation targets, and acquire sub - label sets respectively corresponding to at least two pre - trained sub - target models. The sub - label sets include sub - labels corresponding to various historical recommendation targets;

[0278] The sub - recommendation degree obtaining module 1404 is used to input the training samples into the pre - trained sub - target models, and obtain a set of sub - recommendation degrees output by each pre - trained sub - target model. The set of sub - recommendation degrees includes sub - recommendation degrees corresponding to various historical recommendation targets;

[0279] The target sequence obtaining module 1406 is used to input each set of sub - recommendation degrees into an initial fusion recommendation model, obtain a set of fusion recommendation degrees. The set of fusion recommendation degrees includes fusion recommendation degrees corresponding to various historical recommendation targets, and sort various historical recommendation targets based on the fusion recommendation degrees to obtain a historical recommendation target sequence;

[0280] The subsequence obtaining module 1408 is used to sort the sub - labels corresponding to various historical recommendation targets in the sub - label sets based on the order of the historical recommendation target sequence, and obtain sub - label sequences respectively corresponding to each pre - trained sub - target model;

[0281] The evaluation module 1410 is used to determine sorting evaluation information corresponding to each sub - label sequence based on a sorting evaluation index, and determine target sorting evaluation information based on the sorting evaluation information corresponding to each sub - label sequence;

[0282] The update module 1412 is used to update the initial fusion recommendation model based on the target sorting evaluation information. When the training is completed, a target fusion recommendation model is obtained. The target fusion recommendation model is used to recommend information to be recommended.

[0283] In one embodiment, the training samples include historical user identifiers and various historical recommendation targets corresponding to the historical user identifiers; the sub - recommendation degree obtaining module 1404 is further used to acquire historical attribute features corresponding to the historical user identifiers and historical recommendation target features corresponding to various historical recommendation targets; input the historical attribute features and historical recommendation target features into the pre - trained sub - target models, and obtain a set of sub - recommendation degrees corresponding to the historical user identifiers output by each pre - trained sub - target model.

[0284] In one embodiment, the evaluation module 1410 includes:

[0285] A type acquisition unit is configured to acquire the tag data types corresponding to each sub-tag sequence, determine the sorting evaluation metrics corresponding to each sub-tag sequence based on the tag data types, and calculate the sorting evaluation information corresponding to each sub-tag sequence based on the sorting evaluation metrics corresponding to each sub-tag sequence.

[0286] In one embodiment, the tag data type includes a discrete data type;

[0287] The type acquisition unit is further configured to, when the tag data type corresponding to the first sub-tag sequence is a discrete data type, determine the number of first-category sub-tags and the number of second-category sub-tags in the first sub-tag sequence; determine the historical recommended target position identifiers corresponding to each first-category sub-tag from the historical recommended target sequence, and calculate the sum of the identifiers of the historical recommended target position identifiers corresponding to each first-category sub-tag; calculate the first sorting evaluation information corresponding to the first sub-tag sequence based on the number of first-category tags, the number of second-category tags, and the sum of the identifiers.

[0288] In one embodiment, the tag data type includes a continuous data type; the type acquisition unit is further configured to, when the tag data type corresponding to the second sub-tag sequence is a continuous data type, calculate the number of positive pairs and the total number of sequence pairs in the second sub-tag sequence; calculate the ratio of the number of positive pairs to the total number of sequence pairs to obtain the second sorting evaluation information corresponding to the second sub-tag sequence.

[0289] In one embodiment, the type acquisition unit is further configured to divide the second sub-tag sequence to obtain a second left sub-tag sequence and a second right sub-tag sequence; calculate the number of first positive pairs in the second left sub-tag sequence and calculate the number of second positive pairs in the second right sub-tag sequence; calculate the number of interactive positive pairs between the second left sub-tag sequence and the second right sub-tag sequence, and determine the number of positive pairs based on the number of first positive pairs, the number of second positive pairs, and the number of interactive positive pairs.

[0290] In one embodiment, the evaluation module 1410 includes:

[0291] A first target information obtaining unit is configured to obtain the preset weights corresponding to each trained sub-target model, and perform weighted calculation on the sorting evaluation information of each sub-tag sequence based on the preset weights corresponding to each trained sub-target model to obtain the first target sorting evaluation information.

[0292] In one embodiment, the training samples include each historical user identifier and each historical recommended target corresponding to each historical user identifier; the evaluation module 1410 further includes:

[0293] An average information determination unit is configured to obtain the sorting evaluation information corresponding to each sub-tag sequence of each historical user identifier, and obtain the total number of historical users; perform an average calculation based on the sorting evaluation information corresponding to each sub-tag sequence of each historical user identifier and the total number of historical users to determine the average sorting evaluation information corresponding to each sub-tag sequence.

[0294] The first target information obtaining unit is further configured to perform a weighted calculation based on the preset weights corresponding to each trained sub-target model and the average sorting evaluation information corresponding to each sub-tag sequence to obtain the second target sorting evaluation information.

[0295] In one embodiment, the training samples include each historical user identifier and each historical recommendation target corresponding to each historical user identifier; the evaluation module 1410 further includes:

[0296] A specific information obtaining unit is configured to obtain the sorting evaluation information corresponding to each sub-tag sequence of each historical user identifier, and obtain the number of historical recommendation targets corresponding to each historical user identifier; perform a weighted calculation on the sorting evaluation information corresponding to each sub-tag sequence of each historical user identifier based on the number of historical recommendation targets corresponding to each historical user identifier to obtain the weighted sorting evaluation information corresponding to each sub-tag sequence; calculate the total number of historical recommendation targets based on the number of historical recommendation targets corresponding to each historical user, and calculate the ratio of the weighted sorting evaluation information corresponding to each sub-tag sequence to the total number of historical recommendation targets to obtain the specific sorting evaluation information corresponding to each sub-tag sequence.

[0297] The first target information obtaining unit is further configured to perform a weighted calculation based on the preset weights corresponding to each trained sub-target model and the specific sorting evaluation information corresponding to each sub-tag sequence to obtain the third target sorting evaluation information.

[0298] In one embodiment, the update module 1412 includes:

[0299] A gradient calculation unit is configured to, when the initial fusion recommendation model meets the preset conditions, simulate and calculate the simulated gradient of the initial model parameters in the initial fusion recommendation model based on the target sorting evaluation information;

[0300] A parameter update unit is configured to update the initial model parameters in the initial fusion recommendation model based on the simulated gradient and the preset learning rate to obtain an updated fusion recommendation model;

[0301] A model obtaining unit is configured to, when the updated fusion recommendation model reaches the preset training completion condition, obtain the target fusion recommendation model.

[0302] In one embodiment, the gradient calculation unit includes:

[0303] A partial derivative calculation sub-unit, configured to calculate partial derivatives of initial model parameters in an initial fusion recommendation model based on target sorting evaluation information;

[0304] A simulated gradient determination sub-unit, configured to determine a simulated gradient based on partial derivatives of initial model parameters in the initial fusion recommendation model.

[0305] In one embodiment, the partial derivative calculation sub-unit is further configured to: obtain a preset first parameter micro-variable, adjust the initial model parameters of the initial fusion recommendation model based on the preset first parameter micro-variable to obtain first adjusted model parameters, and determine a first adjusted fusion recommendation model based on the first adjusted model parameters; determine first adjusted sorting evaluation information based on the first adjusted fusion recommendation model and training samples; calculate a sorting evaluation information difference between the first adjusted sorting evaluation information and the target sorting evaluation information, and calculate a ratio of the sorting evaluation information difference to the preset first parameter micro-variable to obtain partial derivatives corresponding to the first adjusted model parameters.

[0306] In one embodiment, the update module 1412 is further configured to, when the initial fusion recommendation model does not meet a preset condition, calculate specific evaluation index information corresponding to the preset condition based on a historical recommendation target sequence; obtain a preset second parameter micro-variable, adjust the initial model parameters of the initial fusion recommendation model based on the preset second parameter micro-variable to obtain second adjusted model parameters, and determine a second adjusted fusion recommendation model based on the second adjusted model parameters; determine a target historical recommendation target sequence based on the second adjusted fusion recommendation model and training samples; calculate target specific evaluation index information corresponding to the preset condition based on the target historical recommendation target sequence; calculate a specific evaluation information difference between the target specific evaluation index information and the specific evaluation index information, and calculate a ratio of the specific evaluation information difference to the preset second parameter micro-variable to obtain partial derivatives corresponding to the second adjusted model parameters; determine a target simulated gradient corresponding to the initial fusion recommendation model based on the partial derivatives corresponding to the second adjusted model parameters; update the initial model parameters in the initial fusion recommendation model based on the target simulated gradient and a preset target learning rate to obtain a target updated fusion recommendation model; when the target updated fusion recommendation model meets the preset condition, use the target updated fusion recommendation model as the initial fusion recommendation model.

[0307] In one embodiment, as Figure 15 shown, a recommendation device is provided. The device can be a software module, a hardware module, or a combination of both to form a part of a computer device. The device specifically includes: a feature acquisition module 1502, a feature input module 1504, a fusion module 1506, a sorting module 1508, and a recommendation module 1510, where:

[0308] The feature acquisition module 1502 is configured to obtain a user identifier and obtain attribute features based on the user identifier;

[0309] A feature input module 1504 is configured to obtain each target to be recommended and corresponding target attribute features, and input the attribute features and the target attribute features into at least two trained sub-target models, so as to obtain a set of sub-target recommendation degrees output by each trained sub-target model. The set of sub-target recommendation degrees includes the sub-target recommendation degrees corresponding to each target to be recommended;

[0310] A fusion module 1506 is configured to input each set of sub-recommendation degrees into a target fusion recommendation model to obtain the fusion recommendation degrees corresponding to each target to be recommended. The target fusion recommendation model is trained using training samples and sub-label sets respectively corresponding to at least two trained sub-target models. The training samples include each historical recommended target, and the sub-label sets include the sub-labels corresponding to each historical recommended target;

[0311] A sorting module 1508 is configured to sort each target to be recommended based on the fusion recommendation degrees to obtain a sequence of targets to be recommended;

[0312] A recommendation module 1510 is configured to select a preset number of targets to be recommended from the sequence of targets to be recommended, and recommend the preset number of targets to be recommended to a user identifier.

[0313] For the specific definitions of the recommendation model training device and the recommendation device, reference may be made to the definitions of the recommendation model training method and the recommendation method in the foregoing text, which will not be elaborated herein. Each module in the foregoing recommendation model training device and recommendation device can be implemented in whole or in part by software, hardware, and their combination. The foregoing modules can be embedded in or independent of a processor in a computer device in the form of hardware, or stored in a memory in the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the foregoing modules.

[0314] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as Figure 16 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Wherein, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store sample data or data of targets to be recommended. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a recommendation model training method or a recommendation method.

[0315] Those skilled in the art can understand, Figure 16The structure shown is only a block diagram of some of the structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0316] In one embodiment, a computer device is also provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0317] In one embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0318] In one embodiment, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device implements the steps in the above method embodiments.

[0319] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application may include at least one of non-volatile and volatile memories. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0320] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0321] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A recommendation model training method, It is characterized in that The method comprises: Acquire training samples, the training samples include each historical recommendation target, and acquire sub-label sets corresponding to at least two trained sub-target models, the sub-label sets include sub-labels corresponding to each historical recommendation target; Inputting the training samples into the trained sub-goal models to obtain sub-recommendation degree sets output by each trained sub-goal model, wherein the sub-recommendation degree sets include sub-recommendation degrees corresponding to each of the historical recommendation goals; Inputting each sub-recommendation degree set into the initial fusion recommendation model to obtain a fusion recommendation degree set, wherein the fusion recommendation degree set includes the fusion recommendation degrees corresponding to each of the historical recommendation targets, and sorting each of the historical recommendation targets based on the fusion recommendation degrees to obtain a historical recommendation target sequence; Sort the sub-labels corresponding to each historically recommended target in the sub-label set based on the order of the historically recommended target sequence to obtain sub-label sequences corresponding to each trained sub-target model; Determining the ranking evaluation information corresponding to each sub-label sequence based on the ranking evaluation index, and determining the target ranking evaluation information based on the ranking evaluation information corresponding to each sub-label sequence, including: performing weighted calculation based on the preset weights corresponding to each trained sub-target model and the average ranking evaluation information corresponding to each sub-label sequence to obtain the second target ranking evaluation information, the average ranking evaluation information is obtained by obtaining the ranking evaluation information corresponding to each sub-label sequence of each historical user identifier, and obtaining the total number of historical users, and based on the ranking evaluation information corresponding to each sub-label sequence of each historical user identifier and the total number of historical users. The training sample includes each historical user identifier and each historical recommendation target corresponding to each historical user identifier; The initial fusion recommendation model is updated based on the target ranking evaluation information. When the training is completed, a target fusion recommendation model is obtained, and the target fusion recommendation model is used to recommend the information to be recommended.

2. The method according to claim 1, It is characterized in that The determining of the ranking evaluation information corresponding to each sub-label sequence based on the ranking evaluation index includes: The tag data type corresponding to each subtag sequence is obtained, the ranking evaluation index corresponding to each subtag sequence is determined based on the tag data type, and the ranking evaluation information corresponding to each subtag sequence is calculated based on the ranking evaluation index corresponding to each subtag sequence.

3. The method according to claim 2, It is characterized in that The tag data type includes a discrete data type; The determining, based on the tag data type, the ranking evaluation index corresponding to each sub-tag sequence, and calculating, based on the ranking evaluation index corresponding to each sub-tag sequence, the ranking evaluation information corresponding to each sub-tag sequence includes: When the tag data type corresponding to the first sub-tag sequence is a discrete data type, determining the number of first category sub-tags and the number of second category sub-tags from the first sub-tag sequence; Determine the historical recommended target position identifiers corresponding to each first-category sub-tag from the historical recommended target sequence, and calculate the sum of the identifiers of the historical recommended target position identifiers corresponding to each first-category sub-tag; Calculate the first sorting evaluation information corresponding to the first sub-tag sequence based on the number of first-category sub-tags, the number of second-category sub-tags, and the sum of the identifiers.

4. The method according to claim 2, wherein, the tag data type includes a continuous data type; the determining the sorting evaluation indicators corresponding to each sub-tag sequence based on the tag data type, and calculating the sorting evaluation information corresponding to each sub-tag sequence based on the sorting evaluation indicators corresponding to each sub-tag sequence includes: when the tag data type corresponding to the second sub-tag sequence is a continuous data type, calculate the number of positive pairs and the total number of sequence pairs in the second sub-tag sequence; Calculate the ratio of the number of positive pairs to the total number of sequence pairs to obtain the second sorting evaluation information corresponding to the second sub-tag sequence.

5. The method according to claim 4, wherein, the calculating the number of positive pairs in the second sub-tag sequence includes: Divide the second sub-tag sequence to obtain a second sub-tag left sequence and a second sub-tag right sequence; Calculate the number of first positive pairs in the second sub-tag left sequence and calculate the number of second positive pairs in the second sub-tag right sequence; Calculate the number of interactive positive pairs between the second sub-tag left sequence and the second sub-tag right sequence, and determine the number of positive pairs based on the number of first positive pairs, the number of second positive pairs, and the number of interactive positive pairs.

6. The method according to claim 1, wherein, the determining the target sorting evaluation information based on the sorting evaluation information corresponding to each sub-tag sequence includes: Obtain the preset weights corresponding to each trained sub-target model, and perform weighted calculation on the sorting evaluation information of each sub-tag sequence based on the preset weights corresponding to each trained sub-target model to obtain the first target sorting evaluation information.

7. The method according to claim 6, wherein, the training samples include each historical user identifier and each historical recommended target corresponding to each historical user identifier; Before obtaining the preset weights corresponding to each trained sub-target model, it further includes: Obtain the sorting evaluation information corresponding to each sub-tag sequence of each historical user identifier, and obtain the number of historical recommended targets corresponding to each historical user identifier; Perform weighted calculation on the sorting evaluation information corresponding to each sub-tag sequence of each historical user identifier based on the number of historical recommended targets corresponding to each historical user identifier to obtain the weighted sorting evaluation information corresponding to each sub-tag sequence; Calculate the total number of historical recommended targets based on the number of historical recommended targets corresponding to each historical user, and calculate the ratio of the weighted sorting evaluation information corresponding to each sub-tag sequence to the total number of historical recommended targets to obtain the specific sorting evaluation information corresponding to each sub-tag sequence; Performing weighted calculation based on the preset weights corresponding to the respective trained sub-goal models and the sorting evaluation information of the respective sub-label sequences to obtain target sorting evaluation information, including: Performing weighted calculation based on the preset weights corresponding to the respective trained sub-goal models and the specific sorting evaluation information corresponding to the respective sub-label sequences to obtain third target sorting evaluation information.

8. The method according to claim 1, wherein, updating the initial fusion recommendation model based on the target sorting evaluation information, and when the training is completed, obtaining a target fusion recommendation model, including: when the initial fusion recommendation model meets the preset conditions, calculating the simulated gradient of the initial model parameters in the initial fusion recommendation model based on the target sorting evaluation information; updating the initial model parameters in the initial fusion recommendation model based on the simulated gradient and the preset learning rate to obtain an updated fusion recommendation model; when the updated fusion recommendation model reaches the preset training completion condition, obtaining the target fusion recommendation model.

9. The method according to claim 8, wherein, calculating the simulated gradient of the initial model parameters in the initial fusion recommendation model based on the target sorting evaluation information, including: calculating the partial derivative of the initial model parameters in the initial fusion recommendation model based on the target sorting evaluation information; determining the simulated gradient based on the partial derivative of the initial model parameters in the initial fusion recommendation model.

10. The method according to claim 9, wherein, calculating the partial derivative of the initial model parameters in the initial fusion recommendation model based on the target sorting evaluation information, including: obtaining a preset first parameter micro-variable, adjusting the initial model parameters of the initial fusion recommendation model based on the preset first parameter micro-variable to obtain first adjusted model parameters, and determining a first adjusted fusion recommendation model based on the first adjusted model parameters; determining first adjusted sorting evaluation information based on the first adjusted fusion recommendation model and the training samples; calculating the difference between the first adjusted sorting evaluation information and the target sorting evaluation information, and calculating the ratio of the difference in sorting evaluation information to the preset first parameter micro-variable to obtain the partial derivative corresponding to the first adjusted model parameters.

11. The method according to claim 1, wherein, updating the initial fusion recommendation model based on the target sorting evaluation information, and when the training is completed, obtaining a target fusion recommendation model, including: when the initial fusion recommendation model does not meet the preset conditions, calculating specific evaluation index information corresponding to the preset conditions based on the historical recommendation target sequence; obtaining a preset second parameter micro-variable, adjusting the initial model parameters of the initial fusion recommendation model based on the preset second parameter micro-variable to obtain second adjusted model parameters, and determining a second adjusted fusion recommendation model based on the second adjusted model parameters; determining a target historical recommendation target sequence based on the second adjusted fusion recommendation model and the training samples; calculating target specific evaluation index information corresponding to the preset conditions based on the target historical recommendation target sequence; Calculate the specific evaluation information difference between the target specific evaluation index information and the specific evaluation index information, and calculate the ratio of the specific evaluation information difference to the preset second parameter micro-variable to obtain the partial derivative corresponding to the second adjustment model parameter; Determine the target simulation gradient corresponding to the initial fusion recommendation model based on the partial derivative corresponding to the second adjustment model parameter; Update the initial model parameters in the initial fusion recommendation model based on the target simulation gradient and the preset target learning rate to obtain a target updated fusion recommendation model; When the target updated fusion recommendation model meets the preset conditions, use the target updated fusion recommendation model as the initial fusion recommendation model.

12. The method according to claim 1, wherein, the training sample includes a historical user identifier and each historical recommended target corresponding to the historical user identifier; The step of inputting the training sample into the trained sub-target models to obtain a set of sub-recommendation degrees output by each trained sub-target model includes: Obtain the historical attribute features corresponding to the historical user identifier and the historical recommended target features corresponding to each historical recommended target; Input the historical attribute features and the historical recommended target features into the trained sub-target models to obtain the set of sub-recommendation degrees corresponding to the historical user identifier output by each trained sub-target model.

13. A recommendation method, wherein, the method includes: Obtain a user identifier, and obtain attribute features based on the user identifier; Obtain each target to be recommended and the corresponding target attribute features, and input the attribute features and the target attribute features into at least two trained sub-target models to obtain a set of sub-target recommendation degrees output by each trained sub-target model, and the set of sub-target recommendation degrees includes the sub-target recommendation degrees corresponding to each target to be recommended; Input each set of sub-recommendation degrees into a target fusion recommendation model to obtain the fusion recommendation degrees corresponding to each target to be recommended, where the target fusion recommendation model is trained using a training sample and the sub-label sets respectively corresponding to the at least two trained sub-target models, the training sample includes each historical recommended target, and the sub-label sets include the sub-labels corresponding to each historical recommended target; the target fusion recommendation model is trained using the method described in any one of claims 1 to 12 in the above model training method; Rank each target to be recommended based on the fusion recommendation degree to obtain a sequence of targets to be recommended; Select a preset number of targets to be recommended from the sequence of targets to be recommended, and recommend the preset number of targets to be recommended to the user identifier.

14. A recommendation model training device, wherein, the device includes: A sample acquisition module, configured to acquire a training sample, where the training sample includes each historical recommended target, and acquire the sub-label sets respectively corresponding to at least two trained sub-target models, and the sub-label sets include the sub-labels corresponding to each historical recommended target; A sub-recommendation degree obtaining module, used for inputting the training sample into the trained sub-goal model to obtain a sub-recommendation degree set output by each trained sub-goal model, wherein the sub-recommendation degree set includes a sub-recommendation degree corresponding to each historical recommendation goal; A target sequence obtaining module, used for inputting each sub-recommendation degree set into the initial fusion recommendation model to obtain a fusion recommendation degree set, wherein the fusion recommendation degree set includes the fusion recommendation degrees corresponding to each historical recommendation target, and sorting each historical recommendation target based on the fusion recommendation degree to obtain a historical recommendation target sequence; A subsequence obtaining module, used to sort the sublabels corresponding to each historically recommended target in the sublabel set based on the order of the historically recommended target sequence, to obtain sublabel sequences corresponding to each trained subtarget model; An evaluation module, used to determine the ranking evaluation information corresponding to each sub-label sequence based on the ranking evaluation index, and determine the target ranking evaluation information based on the ranking evaluation information corresponding to each sub-label sequence, including: performing weighted calculation based on the preset weights corresponding to each trained sub-target model and the average ranking evaluation information corresponding to each sub-label sequence to obtain the second target ranking evaluation information, the average ranking evaluation information is obtained by obtaining the ranking evaluation information corresponding to each sub-label sequence of each historical user identifier and obtaining the total number of historical users, and based on the ranking evaluation information corresponding to each sub-label sequence of each historical user identifier and the total number of historical users. The training sample includes each historical user identifier and each historical recommendation target corresponding to each historical user identifier; An updating module is used to update the initial fusion recommendation model based on the target ranking evaluation information. When the training is completed, a target fusion recommendation model is obtained, and the target fusion recommendation model is used to recommend the information to be recommended.

15. The device according to claim 14, It is characterized in that The evaluation module comprises: A type acquisition unit is used to obtain the label data type corresponding to each sub-label sequence, determine the sorting evaluation index corresponding to each sub-label sequence based on the label data type, and calculate the sorting evaluation information corresponding to each sub-label sequence based on the sorting evaluation index corresponding to each sub-label sequence.

16. The device according to claim 15, It is characterized in that The tag data type includes a discrete data type; The type acquisition unit is also used to determine the number of first category subtags and the number of second category subtags from the first subtag sequence when the tag data type corresponding to the first subtag sequence is a discrete data type; determine the historical recommendation target position identifiers corresponding to each first category subtag from the historical recommendation target sequence, and calculate the identifier sum of the historical recommendation target position identifiers corresponding to each first category subtag; and calculate the first ranking evaluation information corresponding to the first subtag sequence based on the number of first category subtags, the number of second category subtags and the identifier sum.

17. The device according to claim 15, It is characterized in that The label data type includes a continuous data type; The type acquisition unit is further configured to calculate the number of positive pairs and the total number of sequence pairs in the second sub-label sequence when the label data type corresponding to the second sub-label sequence is a continuous data type; calculate a ratio of the number of positive pairs to the total number of sequence pairs to obtain second sorting evaluation information corresponding to the second sub-label sequence.

18. The apparatus according to claim 17, wherein, the type acquisition unit is further configured to divide the second sub-label sequence to obtain a second sub-label left sequence and a second sub-label right sequence; calculate a first number of positive pairs of the second sub-label left sequence and calculate a second number of positive pairs of the second sub-label right sequence; calculate an interactive number of positive pairs between the second sub-label left sequence and the second sub-label right sequence, and determine the number of positive pairs based on the first number of positive pairs, the second number of positive pairs, and the interactive number of positive pairs.

19. The apparatus according to claim 14, wherein, the evaluation module includes: a first target information obtaining unit, configured to obtain preset weights corresponding to the respective trained sub-target models, and perform weighted calculation on the sorting evaluation information of the respective sub-label sequences based on the preset weights corresponding to the respective trained sub-target models to obtain first target sorting evaluation information.

20. The apparatus according to claim 19, wherein, the training samples include respective historical user identifiers and respective historical recommended targets corresponding to each historical user identifier; the evaluation module further includes: a specific information obtaining unit, configured to obtain sorting evaluation information corresponding to respective sub-label sequences of each historical user identifier, and obtain the number of historical recommended targets corresponding to each historical user identifier; perform weighted calculation on the sorting evaluation information corresponding to the respective sub-label sequences of each historical user identifier based on the number of historical recommended targets corresponding to each historical user identifier to obtain weighted sorting evaluation information corresponding to the respective sub-label sequences; calculate a total number of historical recommended targets based on the number of historical recommended targets corresponding to each historical user, and calculate a ratio of the weighted sorting evaluation information corresponding to the respective sub-label sequences to the total number of historical recommended targets to obtain specific sorting evaluation information corresponding to the respective sub-label sequences; the first target information obtaining unit is further configured to perform weighted calculation based on the preset weights corresponding to the respective trained sub-target models and the specific sorting evaluation information corresponding to the respective sub-label sequences to obtain third target sorting evaluation information.

21. The apparatus according to claim 14, wherein, the update module includes: a gradient calculation unit, configured to, when the initial fusion recommendation model meets a preset condition, simulate and calculate a simulated gradient of initial model parameters in the initial fusion recommendation model based on the target sorting evaluation information; a parameter update unit, configured to update the initial model parameters in the initial fusion recommendation model based on the simulated gradient and a preset learning rate to obtain an updated fusion recommendation model; A model obtaining unit, configured to obtain the target fusion recommendation model when the updated fusion recommendation model meets a preset training completion condition.

22. The apparatus according to claim 21, wherein, the gradient calculation unit includes: a partial derivative calculation subunit, configured to calculate the partial derivatives of the initial model parameters in the initial fusion recommendation model based on the target ranking evaluation information; a simulated gradient determination subunit, configured to determine the simulated gradient based on the partial derivatives of the initial model parameters in the initial fusion recommendation model.

23. The apparatus according to claim 22, wherein, the partial derivative calculation subunit is configured to obtain a preset first parameter micro-variable, adjust the initial model parameters of the initial fusion recommendation model based on the preset first parameter micro-variable to obtain first adjusted model parameters, and determine a first adjusted fusion recommendation model based on the first adjusted model parameters; determine first adjusted ranking evaluation information based on the first adjusted fusion recommendation model and the training samples; calculate the ranking evaluation information difference between the first adjusted ranking evaluation information and the target ranking evaluation information, and calculate the ratio of the ranking evaluation information difference to the preset first parameter micro-variable to obtain the partial derivatives corresponding to the first adjusted model parameters.

24. The apparatus according to claim 14, wherein, the updating module is further configured to calculate specific evaluation index information corresponding to the preset condition based on the historical recommendation target sequence when the initial fusion recommendation model does not meet the preset condition; obtain a preset second parameter micro-variable, adjust the initial model parameters of the initial fusion recommendation model based on the preset second parameter micro-variable to obtain second adjusted model parameters, and determine a second adjusted fusion recommendation model based on the second adjusted model parameters; determine a target historical recommendation target sequence based on the second adjusted fusion recommendation model and the training samples; calculate target specific evaluation index information corresponding to the preset condition based on the target historical recommendation target sequence; calculate the specific evaluation information difference between the target specific evaluation index information and the specific evaluation index information, and calculate the ratio of the specific evaluation information difference to the preset second parameter micro-variable to obtain the partial derivatives corresponding to the second adjusted model parameters; determine the target simulated gradient corresponding to the initial fusion recommendation model based on the partial derivatives corresponding to the second adjusted model parameters; update the initial model parameters in the initial fusion recommendation model based on the target simulated gradient and a preset target learning rate to obtain a target updated fusion recommendation model; when the target updated fusion recommendation model meets the preset condition, use the target updated fusion recommendation model as the initial fusion recommendation model.

25. The apparatus according to claim 14, wherein, the training samples include historical user identifiers and respective historical recommendation targets corresponding to the historical user identifiers. The sub - recommendation degree obtaining module is further configured to obtain the historical attribute features corresponding to the historical user identifier and the historical recommendation target features corresponding to each historical recommendation target; input the historical attribute features and the historical recommendation target features into the trained sub - target model, and obtain the set of sub - recommendation degrees corresponding to the historical user identifier output by each trained sub - target model.

26. A recommendation device, characterized in that, the device includes: a feature obtaining module, configured to obtain a user identifier and obtain attribute features based on the user identifier; a feature input module, configured to obtain each target to be recommended and the corresponding target attribute features, input the attribute features and the target attribute features into at least two trained sub - target models, and obtain a set of sub - target recommendation degrees output by each trained sub - target model, where the set of sub - target recommendation degrees includes the sub - target recommendation degrees corresponding to each target to be recommended; a fusion module, configured to input each set of sub - recommendation degrees into a target fusion recommendation model to obtain the fusion recommendation degrees corresponding to each target to be recommended, where the target fusion recommendation model is trained using training samples and the sub - label sets respectively corresponding to the at least two trained sub - target models, the training samples include each historical recommendation target, and the sub - label sets include the sub - labels corresponding to each historical recommendation target; the target fusion recommendation model is trained using the method described in any one of claims 1 to 12 in the above model training method; a sorting module, configured to sort each target to be recommended based on the fusion recommendation degree to obtain a sequence of targets to be recommended; a recommendation module, configured to select a preset number of targets to be recommended from the sequence of targets to be recommended and recommend the preset number of targets to be recommended to the user identifier.

27. A computer device, including a memory and a processor, where the memory stores a computer program, characterized in that, when the processor executes the computer program, the steps of the method described in any one of claims 1 to 13 are implemented.

28. A computer - readable storage medium, storing a computer program, characterized in that, when the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 13 are implemented.

29. A computer program product, including a computer program, characterized in that, when the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 13 are implemented.

Citation Information

Patent Citations

  • Object recommendation method and device, equipment and medium

    CN111143543A