Recommendation model training method, object recommendation method, recommendation system and computing equipment
Through the adversarial pre-training and meta-learning methods, the differences in user data distribution are eliminated, the meta-learning task set for multiple recommended scenarios is constructed, and the adaptive recommendation model is trained, which solves the problems of sparse and distribution differences of user data in the target recommendation scenario, and achieves high-precision personalized recommendation and improvement of training efficiency.
Patent Information
- Application Number
- CN202510583718.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-19
AI Technical Summary
The existing recommendation model faces the problems of sparse user data and distribution differences in the target recommendation scenario, which leads to insufficient generalization capabilities of the model and it is difficult to achieve high-precision personalized recommendations.
Through adversarial pre-training and meta-learning methods, the domain classifier is used to eliminate the differences in user data distribution, and a meta-learning task set for multiple recommended scenarios is constructed, and the adaptive recommendation model is trained in combination with the adaptive task set for the target recommended scenario.
The accuracy and adaptability of the recommended model is improved under low data density, the balance between sparse data and high precision is achieved, and the training efficiency and real-time performance of personalized recommendations is improved.
Smart Images

Figure CN120509465A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the technical field of machine learning, and in particular to training a recommendation model, an object recommendation method, a recommendation system, and a computing device. Background Art
[0002] With the development of internet technology, recommendation models are widely used in e-commerce platforms, social networks, content communities, and other fields. Currently, recommendation models rely on a pre-training and fine-tuning training paradigm or rule-based heuristic strategies to complete targeted training for target recommendation scenarios.
[0003] However, due to the influx of target users in target recommendation scenarios, the target users' user data for the recommended objects is highly sparse, resulting in significant distribution differences between the target users' user data and the reference users' user data (for example, in a content community, target users prefer content from top creators, while reference users prefer niche, vertical content). Traditional training paradigms and heuristic strategies are difficult to support the targeted training of recommendation models. Recommendation models trained based on user data from large-scale reference users produce negative transfer effects in target recommendation scenarios, resulting in insufficient model generalization and accuracy of the trained recommendation models. Balancing the conflict between the low user data density of target users and the high precision required for object recommendation in target recommendation scenarios is an urgent problem that needs to be solved. Summary of the Invention
[0004] In view of this, embodiments of this specification provide a method for training a recommendation model. One or more embodiments of this specification also relate to an object recommendation method, a recommendation system, a computing device, a computer-readable storage medium, and a computer program product to address technical deficiencies in the prior art.
[0005] According to a first aspect of an embodiment of this specification, a method for training a recommendation model is provided, comprising:
[0006] Obtaining user data of sample users for sample objects, wherein the sample users include reference sample users and target sample users;
[0007] Based on the encoding features of the user data output by the initial recommendation model, the initial recommendation model is adversarially pre-trained through a domain classifier of user type to obtain a pre-trained recommendation model, wherein the domain classifier determines whether the user data corresponds to a reference sample user or a target sample user based on the encoding features of the user data;
[0008] Divide user data into meta-learning task sets for multiple recommendation scenarios, and train the pre-trained recommendation model based on the meta-learning task sets to obtain a meta-parameter recommendation model;
[0009] Based on the meta-learning task set of the target recommendation scenario, an adaptation task set of the target recommendation scenario is constructed, and based on the adaptation task set, the meta-parameter recommendation model is trained to obtain the adaptation recommendation model of the target recommendation scenario.
[0010] According to a second aspect of the embodiments of this specification, a method for recommending an object is provided, including:
[0011] Obtaining user information of a target user, object information of multiple candidate recommendation objects, and an adaptation recommendation model for a target recommendation scenario, wherein the adaptation recommendation model for the target recommendation scenario is trained according to the recommendation model training method described above;
[0012] Input user information and object information of multiple candidate recommendation objects into the adaptation recommendation model to obtain the target recommendation result;
[0013] Based on the target recommendation result, determining the target recommendation object from multiple candidate recommendation objects;
[0014] Recommend the target object to the target user.
[0015] According to a third aspect of an embodiment of this specification, there is provided a recommendation system, comprising a data processing module, a meta-learning module, an adaptation module, and a recommendation generation module;
[0016] A data processing module is configured to obtain user data of sample users for sample objects, wherein the sample users include reference sample users and target sample users;
[0017] a meta-learning module configured to perform adversarial pre-training on the initial recommendation model based on the encoding features of the user data output by the initial recommendation model through a domain classifier of user type to obtain a pre-trained recommendation model, wherein the domain classifier determines whether the user data corresponds to a reference sample user or a target sample user based on the encoding features of the user data;
[0018] The data processing module is further configured to divide the user data into a set of meta-learning tasks for multiple recommendation scenarios;
[0019] The meta-learning module is further configured to train the pre-trained recommendation model based on the meta-learning task set to obtain a meta-parameter recommendation model;
[0020] The adaptation module is configured to construct an adaptation task set for the target recommendation scenario based on a meta-learning task set of the target recommendation scenario, and train a meta-parameter recommendation model based on the adaptation task set to obtain an adaptation recommendation model for the target recommendation scenario;
[0021] The recommendation generation module is configured to obtain user information of the target user, object information of multiple candidate recommendation objects, and an adaptive recommendation model for the target recommendation scenario, input the user information and the object information of multiple candidate recommendation objects into the adaptive recommendation model, obtain the target recommendation result, determine the target recommendation object from multiple candidate recommendation objects based on the target recommendation result, and recommend the target recommendation object to the target user.
[0022] According to a fourth aspect of the embodiments of this specification, a computing device is provided, including:
[0023] memory and processor;
[0024] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the training method of the recommendation model or the object recommendation method are implemented.
[0025] According to a fifth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned recommendation model training method or object recommendation method.
[0026] According to a sixth aspect of the embodiments of this specification, a computer program product is provided, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned recommendation model training method or object recommendation method.
[0027] In one embodiment of the present specification, user data of sample users for sample objects are obtained, wherein the sample users include reference sample users and target sample users; based on the encoding features of the user data output by the initial recommendation model, the initial recommendation model is adversarially pre-trained through a domain classifier of the user type to obtain a pre-trained recommendation model, wherein the domain classifier determines whether the user data corresponds to a reference sample user or a target sample user based on the encoding features of the user data; the user data is divided into meta-learning task sets for multiple recommendation scenarios, and based on the meta-learning task sets, the pre-trained recommendation model is trained to obtain a meta-parameter recommendation model; based on the meta-learning task set of the target recommendation scenario, an adaptation task set of the target recommendation scenario is constructed, and based on the adaptation task set, the meta-parameter recommendation model is trained to obtain an adaptation recommendation model for the target recommendation scenario.
[0028] The user data of sample users for sample objects is obtained, where the sample users include reference sample users and target sample users. Based on the encoding features of the user data output by the initial recommendation model, the initial recommendation model is adversarially pre-trained through the user type domain classifier to obtain a pre-trained recommendation model. This allows the pre-trained recommendation model to effectively eliminate the data distribution differences between the reference sample users and the target sample users, enhancing the model's generalization ability for cross-domain user features. This adversarial training mechanism forces the encoding features to align with the feature distributions of different user types in the latent space, thereby suppressing the model's learning of user type sensitivity and significantly reducing recommendation bias caused by differences in sample user groups. At the same time, the pre-training process optimizes the feature encoder through reverse gradient propagation of the domain classifier, making the generated user data encoding features discriminative regardless of user type, providing a distributionally robust feature representation foundation for the subsequent meta-learning stage. User data is divided into meta-learning task sets for multiple recommendation scenarios, and the pre-trained recommendation model is trained based on the meta-learning task sets to obtain a meta-parameter recommendation model. In this way, the model can learn universal parameters that adapt to distribution differences during the meta-learning training phase, avoiding the negative transfer effect caused by using only the user data of reference sample users for meta-learning. By utilizing the common knowledge transfer between meta-learning tasks of multiple recommendation scenarios, the model can quickly learn the common preferences of users in multiple recommendation scenarios represented by user data, which not only alleviates the data sparsity problem, but also enhances the adaptability and flexibility of the model in different recommendation scenarios. Based on the meta-learning task set of the target recommendation scenario, an adaptation task set of the target recommendation scenario is constructed, and based on the adaptation task set, the meta-parameter recommendation model is trained to obtain the adaptation recommendation model of the target recommendation scenario, so that the adaptation recommendation model of the target recommendation scenario generates recommendations in the target recommendation scenario, which not only inherits the generalization basis of the meta-parameter model, but also integrates the specific preferences of the target user. When the user data of the target user is low in data density, the object recommendation accuracy in the target recommendation scenario is still guaranteed, and a balance is achieved between sparse data and high-precision requirements. At the same time, the adaptation recommendation model can be quickly trained based on a small amount of target user data to adapt to the recommendation generation requirements of the target recommendation scenario, thereby improving the training efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 This is a flowchart of a training method for a recommendation model provided by one embodiment of this specification;
[0030] Figure 2 This is a flowchart of an object recommendation method provided by one embodiment of this specification;
[0031] Figure 3 This is a flowchart of a processing process of an object recommendation method applied to a content community provided by one embodiment of this specification;
[0032] Figure 4 This is a front-end schematic diagram of a content recommendation method applied to a content community provided by an embodiment of this specification;
[0033] Figure 5 This is a schematic diagram of the structure of a recommendation system provided by an embodiment of this specification;
[0034] Figure 6 This is a structural block diagram of a computing device provided by one embodiment of this specification. DETAILED DESCRIPTION
[0035] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0036] The terms used in one or more embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present invention. The singular forms "a", "the" and "the" used in one or more embodiments of the present invention are also intended to include plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present invention refers to and includes any or all possible combinations of one or more associated listed items.
[0037] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of the present invention, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present invention, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0038] In addition, it should be noted that the data involved in one or more embodiments of the present invention are information and data authorized by the user or fully authorized by all parties, and the statistics, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0039] First, the terms involved in one or more embodiments of this specification are explained.
[0040] Cold start: In recommendation systems, it is difficult to accurately recommend new users or new items due to a lack of sufficient historical interaction data. The cold start problem can be divided into two cases: user cold start and item cold start.
[0041] Meta-learning: A machine learning method that aims to extract common knowledge from the learning process of multiple related tasks, enabling the model to quickly adapt to new but similar tasks. Meta-learning typically consists of two stages: meta-learning and meta-testing.
[0042] Adaptive training: The process of further adjusting model parameters based on a meta-learning-trained model to better adapt it to a specific task or scenario. Adaptive training typically uses a small amount of data from the target task or scenario for fine-tuning.
[0043] Negative transfer effect: In transfer learning, when there are significant differences between the source domain and the target domain, transferring knowledge from the source domain to the target domain will lead to a decrease in model performance.
[0044] Model-Agnostic Meta-Learning (MAML): A meta-learning method that optimizes the initial model parameters so that the model can achieve good performance on new tasks after a small number of gradient updates. It is applicable to various types of models and tasks.
[0045] Reptile: A simplified meta-learning algorithm that iteratively updates the initial model parameters to improve the model's generalization ability on new tasks. It is particularly suitable for small-sample learning problems.
[0046] Reinforcement learning: A type of machine learning method that optimizes cumulative rewards by measuring the actions performed by the model in the environment, calculating the reward function value based on the reward function, and feeding it back to the model for learning.
[0047] Contextual Bandit: A form of reinforcement learning in which the agent chooses an action at a time and immediately receives a reward function value for that action, while taking into account contextual information to optimize long-term returns.
[0048] Support Set: In meta-learning, a support set is a set of labeled data used to learn task-specific parameters. It is often used in the in-task learning phase of meta-learning training. The support set simulates a small amount of user data during the cold start phase and is used to quickly update the model parameters.
[0049] Query Set: In meta-learning, a query set is a set of data used to evaluate model performance. It is often used in conjunction with a support set to validate the model's performance on a specific task. The query set is used to validate the model's training results and simulate the diversity of user preferences.
[0050] Long-tail data: refers to data with low frequency of occurrence. The number of individual long-tail data is low, but the overall long-tail data accounts for a considerable proportion in the data set.
[0051] Twin Tower Model: A recommendation model that independently models user features and object features, and finally calculates the matching score between the two through some method (such as dot product or cosine similarity).
[0052] Sequence models: Models used to process sequential data, such as time series data or text series data, are often used in recommendation systems based on user behavior sequences. For example, the Transformer architecture can be used to capture the migration of user browsing interests from text and image content to short videos.
[0053] Multimodal fusion models: These models can integrate information from different modalities (such as text, images, and audio) to provide richer and more accurate recommendation results. For example, they can combine sentiment analysis of user text reviews with product visual features through a cross-attention mechanism to predict user preferences.
[0054] Adversarial training: A training technique that enhances the robustness and generalization ability of a model by introducing adversarial examples.
[0055] Adversarial loss: An objective function that measures the difference between the model output and the true output. It is an important component of training a generative adversarial network.
[0056] Binary Cross Entropy Loss (BCE) is a commonly used loss function used to measure the difference between the probability distribution of the model output and the actual label in a binary classification problem.
[0057] Domain classifier: In adversarial training, a classifier is used to distinguish whether a sample is from the source domain or the target domain, which helps to identify cross-domain differences.
[0058] Cosine similarity: measures the cosine value of the angle between the directions of two non-zero eigenvectors. As an indicator of similarity between the two, it is widely used in text mining, recommendation systems and other fields.
[0059] KL Divergence (Kullback-Leibler Divergence, abbreviated as KL Divergence): A statistical distance measurement method used to compare the differences between two probability distributions. It is often used as one of the indicators to measure the quality of the generative model.
[0060] Maximal Marginal Relevance (MMR): A document ranking algorithm that aims to balance the relevance and diversity of documents and is widely used in information retrieval and recommendation systems.
[0061] Policy gradient methods: A class of algorithms in reinforcement learning that directly optimize a reward function, i.e., the probability distribution over actions taken in a given state, to maximize the expected return.
[0062] Proximal Policy Optimization (PPO): A policy gradient method that avoids performance crashes caused by large policy updates by limiting the extent of each update.
[0063] Q-learning optimization strategy: A model-free reinforcement learning algorithm that finds a target policy by learning an action-value function (i.e., Q-function), which is the expected reward function value that can be obtained by taking an action in a given state.
[0064] With the development of internet technology, recommendation models are widely used in scenarios such as e-commerce platforms, social networks, and content communities. Currently, a core challenge facing recommendation models is the cold start problem, particularly for target users. Due to a lack of sufficient user data on target users, traditional recommendation models struggle to accurately capture their preferences, resulting in poor recommendation performance.
[0065] Currently, training recommendation models based on pre-training and fine-tuning methods or rule-based heuristic strategies has the following limitations: 1. Sparse data problem: The target users have very little user data, which makes it difficult to support the training and optimization of traditional models. 2. Distribution difference problem: The user data distribution of target users and reference users may differ significantly, resulting in insufficient model generalization capabilities. 3. Insufficient real-time performance: Recommendations in the cold start phase require a fast response, but traditional methods are often computationally expensive and time-consuming. 4. Insufficient personalization: The user preferences of target users are often more diverse, and traditional methods find it difficult to quickly adapt to their personalized needs.
[0066] Currently, meta-learning technology has become a potential solution to the cold start problem due to its excellent performance in few-shot learning. Meta-learning can quickly adapt to the target recommendation scenario by learning shared knowledge from multiple recommendation scenarios, and is particularly suitable for scenarios with sparse data. However, the application of existing meta-learning methods in recommendation models still faces the following challenges: 1. Rationality of task construction: How to construct a meta-learning task set that is close to the target user distribution from user data to improve the generalization ability of the model. 2. Adaptability of distribution differences: How to solve the distribution difference problem between the user data of the target user and the reference user to avoid the negative transfer effect of the model. 3. Computational efficiency and real-time performance: How to reduce computing costs and achieve fast cold start while ensuring recommendation results, thereby improving model training efficiency.
[0067] To address the above issues, this specification provides a training method for a recommendation model. This specification also involves an object recommendation method, a recommendation system, a computing device, a computer-readable storage medium, and a computer program product, which are described in detail one by one in the following embodiments.
[0068] See also Figure 1 , Figure 1 A flowchart of a training method for a recommendation model provided by an embodiment of this specification is shown, including the following specific steps:
[0069] Step 102: Obtain user data of sample users for sample objects, wherein the sample users include reference sample users and target sample users.
[0070] The embodiments of this specification are applicable to applications or system platforms with model training capabilities, such as application servers or cloud computing platforms dedicated to model training. The trained recommendation models adapted to the target scenarios are used in areas such as product recommendations on e-commerce platforms, account recommendations on social networks, or content recommendations on content communities.
[0071] Sample users are the set of users involved in the recommendation model training process, and are used to provide user data to support model training. Sample users include reference sample users with rich user data and target sample users with sparse user data. Reference sample users are sample users with rich user data, and they serve as the benchmark user group for the source of knowledge transfer in recommendation model training. The user data of reference sample users has high density and distribution stability. Target sample users are sample users with sparse user data, and they are the user group that needs to perform scenario preference enhancement in recommendation model training. Their feature data exhibits low density and distribution bias. For example, in a content community application, reference sample users are users who have been registered for more than a certain period of time, have a large amount of content browsing history, and have complete and rich user information. Target sample users are newly registered users with a small amount of content browsing history and incomplete user information.
[0072] Sample objects are physical or virtual objects that can be recommended to users during the recommendation model training process. They are used to provide object data to support model training. Sample objects include, but are not limited to: products, services, advertisements, user accounts, and content (multimedia content, graphic content, etc.). For example, in a content community application, sample objects include video content, graphic content, and plain text content.
[0073] The user data of sample users for sample objects represents a multidimensional feature set representing the relationship between the sample users and the sample objects. It can include at least one of user information (user static data), object information (object static data), and user behavior data (dynamic interaction data) of sample users for the sample objects. For example, in a content community application, user data includes: user information (age, gender, region, preference tags, etc.); object information (multimodal data such as text and images, object identifiers, object categories, etc.); and user behavior data (click behavior, browsing behavior, play completion behavior, and interactive behavior (comments, likes, favorites, etc.).
[0074] To obtain user data of sample users for sample objects, one optional method is to collect user data of users for objects within a preset time period as the user data of sample users for sample objects. Another optional method is to obtain user data of sample users for sample objects from a sample database. Another optional method is to analyze user logs to obtain user data of sample users for sample objects. There is no limitation here.
[0075] Optionally, obtaining user data of the sample user for the sample object includes the following specific steps: obtaining initial user data of the sample user for the sample object; and performing data preprocessing on the initial user data to obtain user data of the sample user for the sample object, wherein the preprocessing includes at least one of data cleaning, data deduplication, data labeling, and feature extraction. Feature extraction may include extracting multimodal features (e.g., text descriptions of products and image features).
[0076] For example, a content community application attracts a large number of new users in a short period of time due to successful promotion and marketing. These new users do not have rich user data and have content preferences that are highly different from those of existing users. Meta-learning training is needed to complete a cold start of knowledge transfer across user groups and capture the real-time dynamic preferences of new users. User information of some existing users is collected from the user resource database, including registration duration, device fingerprint, region, user interest tags, etc., as well as user information of new users, including registration channel, initial device characteristics, region, etc. Content information of recommended content is obtained from the content database, including multimodal data, content identifiers, content categories, etc. User behavior data of these existing users on the recommended content is obtained from the user behavior database, including click behavior, browsing behavior, completion behavior, and interaction behavior. Furthermore, user behavior data of new users on the recommended content is obtained, including click behavior, browsing behavior, completion behavior, and interaction behavior.
[0077] In step 102, user data of sample users for sample objects is obtained, wherein the sample users include reference sample users and target sample users, providing a data basis for the subsequent construction of a meta-learning task set for multiple recommendation scenarios.
[0078] Step 104: Based on the encoding features of the user data output by the initial recommendation model, the initial recommendation model is adversarially pre-trained through a domain classifier of user type to obtain a pre-trained recommendation model, wherein the domain classifier determines whether the user data corresponds to a reference sample user or a target sample user based on the encoding features of the user data.
[0079] The initial recommendation model is the basic model architecture configured before adversarial pre-training. It can be a dual-tower model, a sequence model, or a multimodal fusion model adapted for recommendation scenarios. The initial recommendation model typically includes a user feature encoder, an object feature encoder, and a matching prediction module, which maps user data into low-dimensional latent space encoding features and generates user-object matching predictions.
[0080] The encoding features of user data are the vector representations obtained after the initial recommendation model extracts features from the user data.
[0081] Adversarial pre-training is a pre-training process in which a domain classifier competes with the initial recommendation model in an adversarial game. During the back-propagation process of adversarial pre-training, the domain classifier optimizes its own classification weights through supervised learning to accurately distinguish user types; while the initial recommendation model updates the encoder parameters through gradient reversal operations, making the generated user data encoding features domain invariant, that is, minimizing the difference in feature distribution between the reference sample users and the target sample users. Specifically, the objective function of adversarial pre-training consists of the recommendation loss and the domain confusion loss: L total =L recommend_ -λL Adversarial Among them, L total is the target loss function, L recommend is the recommendation loss, L Adversarial is the adversarial loss, and λ is the weight. Optionally, during training, L is adjusted by gradient reversal. Adversarial The gradient of the target loss function L is reversed, forcing the initial recommendation model to optimize the target loss function L total When ,the recommendation accuracy and domain indistinguishability are improved simultaneously.
[0082] The pre-trained recommendation model is a basic model architecture configured before meta-learning training. It maintains the same architecture as the initial recommendation model, but with adjustments to the parameters of the user feature encoder. The pre-trained recommendation model is a recommendation model with generalized encoding capabilities across user groups, obtained through adversarial pre-training.
[0083] The domain classifier is a binary classification network module built based on static user features. Based on the encoded features of user data, the domain classifier determines whether the user data corresponds to a reference sample user or a target sample user. For example, in a content community application, input features include: user registration duration (e.g., number of days), initial device fingerprint (e.g., device model), registration channel (e.g., social media jump, organic traffic), number of interactions in the previous 24 hours (e.g., number of content clicks, number of favorites), etc. For example, the domain classifier uses a two-layer fully connected network with an input dimension of 8-dimensional feature vectors and an output of a probability value in the interval [0, 1], indicating the likelihood that the user belongs to the target sample user (new user).
[0084] For example, the initial recommendation model uses a dual-tower model. The domain classifier is constructed as a binary classification network with a gradient reversal layer. During forward propagation, the encoded features of the user data output by two fully connected layers are used to predict domain prediction probabilities. During backward propagation, the gradient reversal layer multiplies the encoder gradient by -λ = -0.3, forcing the encoder to generate encoded features of user data that are unrelated to user type. After multiple rounds of iterative adversarial training, the domain classifier's accuracy continues to decline, indicating that the encoded features successfully confuse user type discrimination.
[0085] In step 104, based on the encoded features of the user data output by the initial recommendation model, the initial recommendation model is adversarially pre-trained using a domain classifier for user type to obtain a pre-trained recommendation model. This pre-trained recommendation model effectively eliminates the data distribution differences between the reference sample users and the target sample users, enhancing the model's generalization ability for cross-domain user features. This adversarial training mechanism forces the encoded features to align with the feature distributions of different user types in the latent space, thereby suppressing the model's learning of user type sensitivity and significantly reducing recommendation bias caused by differences in sample user groups. Simultaneously, the pre-training process optimizes the feature encoder through reverse gradient propagation of the domain classifier, ensuring that the generated user data encoded features have user type-independent discriminability, providing a robust feature representation foundation for the subsequent meta-learning stage.
[0086] Step 106: Divide the user data into meta-learning task sets for multiple recommendation scenarios, and train the pre-trained recommendation model based on the meta-learning task sets to obtain a meta-parameter recommendation model.
[0087] The recommendation scenario is an application scenario for object recommendations corresponding to the preference characteristics of user data. For example, 1. Age: Teenagers may be more inclined to entertainment content, while adults may be more interested in resources related to career development or family. 2. Region: Product recommendations from nearby stores, local activities or news, etc. 3. Interest tags: Recommendations for electronic product enthusiasts, recommendations for sports enthusiasts. 4. New user guidance: Provide new users with initial personalized recommendations to help them quickly discover content or products of interest and promote user retention. 5. Active user incentives: Provide more in-depth and diverse recommendations for active users to encourage them to explore more unknown areas. 6. Lost user recall: For users who were once active but are now inactive, push specially customized content or offers based on their past behavior patterns to try to reactivate these users.
[0088] The meta-learning task set of multiple recommendation scenarios is the meta-learning basic unit after the recommendation scenarios are divided based on the user data preference characteristics, and is a set of training data at multiple task levels. The meta-learning task set of multiple recommendation scenarios is used to guide the recommendation model to learn common knowledge across scenarios, so that the meta-parameter recommendation model has the ability to quickly adapt to the migration of the target recommendation scenario. Optionally, the meta-learning task set of any recommendation scenario includes the meta-learning support set and meta-learning query set of the recommendation scenario. For example, in a content community application, the meta-learning task set of the local life guide recommendation scenario includes: a meta-learning support set of the user's geo-fence trigger records and point of interest (POI) collection density, and a meta-learning query set for offline store check-in behavior prediction.
[0089] The meta-parameter recommendation model is a scenario-generalizing recommendation model obtained through cross-scenario meta-learning. It has transfer learning capabilities that allow for rapid adaptation to target user data from a small number of target recommendation scenarios. The meta-model parameter of the pre-trained recommendation model is θ. After meta-learning training, the meta-model parameter θ gradually converges, resulting in the meta-parameter recommendation model.
[0090] An optional method for dividing user data into meta-learning task sets for multiple recommendation scenarios is to divide user data into meta-learning task sets for multiple recommendation scenarios based on the preference characteristics of the user data.
[0091] User data preference features represent a user's tendency to select an object. These can be explicit or implicit, including but not limited to: user interest tags in user information, user behavior patterns in user behavior data, and applicable user tags in object information. For example, in a content community application, user behavior data may include a user's completion of a particular piece of content, indicating a high user preference for that content.
[0092] Based on the meta-learning task set of multiple recommendation scenarios, the pre-trained recommendation model is trained to obtain a meta-parameter recommendation model. One optional method is: based on the meta-learning task set of multiple recommendation scenarios, the pre-trained recommendation model is supervised trained to obtain a meta-parameter recommendation model. Another optional method is: based on the meta-learning task set of multiple recommendation scenarios, the pre-trained recommendation model is reinforced learned to obtain a meta-parameter recommendation model. There is no limitation here.
[0093] For example, the preference characteristics of user data include: user interest tags: beauty ingredient research, exploration of niche travel destinations, and home renovation inspiration; user behavior patterns: high-density collection behavior (collecting >20 articles per day), content secondary editing behavior, and continuous participation in brand topics.
[0094] Based on the preference characteristics of user data, a meta-learning task set of recommendation scenarios corresponding to five preference characteristics is constructed:
[0095] 1. Meta-learning task set for the in-depth analysis and recommendation scenario of cosmetic ingredients: Meta-learning support set: clustering of "ingredient comparison analysis" content collected by users in the past 7 days; meta-learning query set: predicting the user's continued attention coefficient for professional skincare content.
[0096] 2. Meta-learning task set for the scenario of recommending travel guides for unpopular destinations: Meta-learning support set: geographic semantic clustering of users who searched for "non-popular attractions" three times in a row; meta-learning query set: verifying the visit depth of the outdoor equipment content channel.
[0097] 3. Meta-learning task set for dynamic recommendation scenarios for home space renovation: Meta-learning support set: user image annotation behavior sequence for "space optimization plan" content; meta-learning query set: predicting the content conversion funnel of home co-branded products.
[0098] 4. Meta-learning task set for highly active user content supply scenarios: Meta-learning support set: cross-category browsing path graph of users who collect more than 20 articles per day; meta-learning query set: optimizing the long-tail content penetration rate of the recommendation system.
[0099] 5. Meta-learning task set for the brand topic immersive recommendation scenario: Meta-learning support set: User generated content (UGC) interaction heat of users participating in the topic of "sunscreen technology"; Meta-learning query set: Predicting the acceptance threshold of seeding content for new sunscreen products.
[0100] The model-independent meta-learning (MAML) algorithm is used to supervise the pre-trained recommendation model. The parameters θ of the pre-trained recommendation model are set. The model architecture is a dual-tower model, in which the user tower and object tower extract user feature vectors and object multimodal feature vectors respectively, and calculate the matching score through cosine similarity:
[0101] First, we train on the meta-learning task set of the beauty ingredient deep analysis recommendation scenario, and calculate the temporary update parameters of the model through the support set: Where L_{Support} is the loss function of the support set, and α is the learning rate. Meta-learning training for the training set of the meta-learning task set is completed, and training for the remaining four meta-learning task sets is continued until the meta-parameter θ converges, thus obtaining the meta-parameter recommendation model.
[0102] In step 106, the user data is divided into meta-learning task sets for multiple recommendation scenarios, and the pre-trained recommendation model is trained based on the meta-learning task sets to obtain a meta-parameter recommendation model, so that the model learns universal parameters that adapt to distribution differences during the meta-learning training phase, avoiding the negative transfer effect caused by using only the user data of reference sample users for meta-learning. By utilizing the common knowledge transfer between meta-learning tasks of multiple recommendation scenarios, the model can quickly learn the common preferences of users in multiple recommendation scenarios represented by user data, which not only alleviates the data sparsity problem, but also enhances the adaptability and flexibility of the model in different recommendation scenarios.
[0103] Step 108: Based on the meta-learning task set of the target recommendation scenario, an adaptation task set of the target recommendation scenario is constructed, and based on the adaptation task set, the meta-parameter recommendation model is trained to obtain an adaptation recommendation model of the target recommendation scenario.
[0104] From the meta-learning task set for the target recommendation scenario, we extract target user data for target sample objects, and construct an adaptation task set for the target recommendation scenario. Target sample users are the specific user group for whom personalized recommendations are required in the target recommendation scenario. Target sample objects are the set of candidate objects that can be recommended in the target recommendation scenario. For example, for a travel guide to unpopular attractions, the target sample objects are outdoor adventure graphic content newly added in the past week.
[0105] The target user data of target sample users for target sample objects represents a multi-dimensional feature set representing the relationship between target sample users and target sample objects in a target recommendation scenario. This data may include at least one of user information (user static data), object information (object static data), and user behavior data of target sample users for target sample objects (dynamic interaction data). For example, the interaction data of new users for unpopular tourist attractions may include refined indicators such as average stay time ≥ 120 seconds and repeat visit rate ≥ 30%.
[0106] The adaptation task set for the target recommendation scenario is a fine-tuned dataset extracted from the meta-learning task set and tailored to the target recommendation scenario. It is a single-task-level training data set. The adaptation task set for the target recommendation scenario is used to guide the meta-parameter recommendation model to learn the personalized knowledge of the target recommendation scenario, training the adaptive recommendation model to have the ability to capture real-time dynamic preferences and provide personalized recommendations for the scenario. Optionally, the adaptation task set includes the adaptation support set and adaptation query set for the target recommendation scenario.
[0107] The adaptive recommendation model for the target recommendation scenario is a scenario-customized recommendation model that is fine-tuned based on the meta-parameter recommendation model through the adaptation task set of the target recommendation scenario. The adaptive recommendation model retains the general representation ability obtained by the meta-parameter recommendation model from multi-scenario meta-learning. The adaptive recommendation model uses a small amount of user data unique to the target scenario and adjusts the model parameters through adaptive training to capture the distribution characteristics of the target user group.
[0108] Based on the meta-learning task set of the target recommendation scenario, an adaptation task set of the target recommendation scenario is constructed. One optional method is: from the meta-learning task set of the target recommendation scenario, the target user data of the target sample user for the target sample object is extracted to construct the adaptation task set of the target recommendation scenario.
[0109] Furthermore, from the meta-learning task set of the target recommendation scenario, the target user data of the target sample user for the target sample object is extracted to construct the adaptation task set of the target recommendation scenario. One optional method is: from the meta-learning task set of the target recommendation scenario, the target user data of the target sample user for the target sample object is randomly extracted to construct the adaptation task set of the target recommendation scenario. Another optional method is: based on the time characteristics of the user data, the target user data of the target sample user for the target sample object is extracted from the meta-learning task set of the target recommendation scenario to construct the adaptation task set of the target recommendation scenario. Another optional method is: based on the user behavior pattern of the user data, the target user data of the target sample user for the target sample object is extracted from the meta-learning task set of the target recommendation scenario to construct the adaptation task set of the target recommendation scenario. Another optional method is: based on the user interest tags of the user data, the target user data of the target sample user for the target sample object is extracted from the meta-learning task set of the target recommendation scenario to construct the adaptation task set of the target recommendation scenario. There is no limitation here.
[0110] Based on the adaptation task set of the target recommendation scenario, the meta-parameter recommendation model is trained to obtain the adaptation recommendation model of the target recommendation scenario. One optional method is: based on the adaptation task set of the target recommendation scenario, the meta-parameter recommendation model is supervised trained to obtain the adaptation recommendation model of the target recommendation scenario. Another optional method is: based on the adaptation task set of the target recommendation scenario, the meta-parameter recommendation model is reinforced learning trained to obtain the adaptation recommendation model of the target recommendation scenario. There is no limitation here.
[0111] For example, in the meta-learning task set of the target recommendation scenario "brand topic immersive recommendation scenario", the interaction records of users participating in the latest user-generated content on the topic of "sunscreen technology" are extracted to capture short-term interest fluctuations.
[0112] The meta-parameter recommendation model is trained using the Contextual Bandit reinforcement learning algorithm:
[0113] When initializing the adaptive recommendation model, the meta-parameter θ is inherited. In the adaptive support set of the target recommendation scenario "brand topic immersive recommendation scenario", the personalized parameter θ′ is calculated based on the first 10 behavior records of new users through second-order optimization: Where L_{Support} is the loss function for the support set, and α is the learning rate. A reward function, R = 0.7*CTR + 0.2 min(T / 60, 1) + 0.1*CVR, is designed, where T is the number of seconds spent on a page and CVR is the conversion rate of favorites to likes. Based on this reward function, reinforcement learning training of the meta-parameter recommendation model is completed, resulting in an adapted recommendation model for the target recommendation scenario, "immersive brand topic recommendation scenario."
[0114] In step 108, based on the meta-learning task set of the target recommendation scenario, an adaptation task set of the target recommendation scenario is constructed, and based on the adaptation task set, the meta-parameter recommendation model is trained to obtain the adaptation recommendation model of the target recommendation scenario, so that the adaptation recommendation model of the target recommendation scenario generates recommendations in the target recommendation scenario, inheriting the generalization basis of the meta-parameter model and integrating the specific preferences of the target user. When the user data of the target user is at low data density, the object recommendation accuracy in the target recommendation scenario is still guaranteed, thereby achieving a balance between sparse data and high precision requirements.
[0115] In the embodiments of this specification, the accuracy of object recommendations in the target recommendation scenario is guaranteed while the user data of the target user is at low data density, thereby achieving a balance between sparse data and high-precision requirements. At the same time, an adaptive recommendation model can be quickly trained based on a small amount of target user data to adapt to the recommendation generation requirements of the target recommendation scenario, thereby improving training efficiency.
[0116] In an optional embodiment of the present specification, step 102 includes the following specific steps: obtaining reference user data of the reference sample user for the sample object, and target user data of the target sample user for the sample object; sampling the reference user data to obtain sampled reference user data; and constructing user data of the sample user for the sample object based on the sampled reference user data and the target user data.
[0117] The reference user data of the reference sample user for the sample object represents a multidimensional feature set of the association relationship between the reference sample user and the sample object, and may include at least one of user information (user static data), object information (object static data), and user behavior data (dynamic interaction data) of the reference sample user for the sample object.
[0118] The target user data of the target sample user for the sample object is a multidimensional feature set that characterizes the association relationship between the target sample user and the sample object, and may include at least one of user information (user static data), object information (object static data), and user behavior data (dynamic interaction data) of the target sample user for the sample object.
[0119] Compared to target user data, reference user data is a high-density data set with a complete behavioral chain, encompassing explicit preference characteristics and implicit behavioral patterns throughout the user's lifecycle. Target user data exhibits a significant data distribution shift compared to reference user data. For example, in e-commerce scenarios, data from established users includes a complete three-year search-click-add-to-cart-order chain, duration spent on product detail pages, and cross-category browsing paths. However, data from new users includes fewer than five clicks within 30 days, concentrated on products within a specific price range, significantly different from the consumption habits of established users.
[0120] The sampled reference user data is a multi-dimensional feature set obtained through downsampling processing, and is closer to the target user distribution than the reference user data.
[0121] To obtain reference user data of reference sample users for sample objects, one optional method is to collect user data of users for objects within a preset time period as reference user data of reference sample users for sample objects. Another optional method is to obtain reference user data of reference sample users for sample objects from a sample database. Another optional method is to analyze user logs to obtain reference user data of reference sample users for sample objects. There is no limitation here.
[0122] To obtain the target user data of the target sample user for the sample object, one optional method is to collect the user data of users for the object within a preset time period as the target user data of the target sample user for the sample object. Another optional method is to obtain the target user data of the target sample user for the sample object from the sample database. Another optional method is to analyze the user logs to obtain the target user data of the target sample user for the sample object. There is no limitation here.
[0123] The reference user data is sampled to obtain sampled reference user data. One optional method is to sample the reference user data based on the time characteristics of the user data to obtain sampled reference user data. Another optional method is to sample the reference user data based on the user distribution of the user data to obtain sampled reference user data. Another optional method is to sample the reference user data based on the data distribution of the user data to obtain sampled reference user data. The above methods are not limited here.
[0124] For example, behavioral data from old users for the past three years are collected, including complete interactive links for commodities such as "high-end skin care products" and "digital products". Behavioral data from new users for the past 30 days are collected, which only includes a small number of click records for commodities such as "affordable clothing" and "daily necessities". The behavioral data of old users for the past 30 days are retained, and early low-relevance records are filtered out. The sampling ratio of "high-end commodities" among old users is reduced, and data related to the preferences of target users (such as affordable commodities) is increased. The sampled old user data (such as 100,000) and the target user data (such as 50,000) are merged to form a sample data set with a total size of 150,000. The merged data is standardized to ensure the coding consistency of features such as "price range" and "category ID".
[0125] In the examples of this specification, downsampling and feature selection are used to reduce the distribution differences between reference user data and target user data, thus avoiding the negative transfer effect in meta-learning. By controlling the size of the reference user data, the computational overhead of the meta-learning phase is reduced while retaining key behavioral patterns. Common features in the reference user data are used to fill in missing information about the target user, improving the model's ability to model sparse data.
[0126] In an optional embodiment of the present specification, the user data includes user behavior data of the sample user for the sample object, user information of the sample user, and object information of the sample object; sampling the reference user data to obtain the sampled reference user data includes at least one of the following:
[0127] Based on the time distribution of user behavior data, sample the reference user data to obtain sampled reference user data;
[0128] Based on the similarity between the user information of the reference sample user and the user information of the target sample user, sampling the reference user data to obtain sampled reference user data;
[0129] Based on the data sparsity of the reference user data, the reference user data is sampled to obtain sampled reference user data.
[0130] User behavior data is a collection of records of dynamic interactions between sample users and sample objects, for example, explicit feedback: clicks, favorites, purchases, ratings; implicit feedback: length of stay, page scrolling depth, number of repeat visits; behavior sequence: interaction object IDs and their contextual information (such as search keywords, recommended locations) sorted by time.
[0131] The user information of sample users includes static attributes and dynamic portrait features, such as basic attributes: age, gender, region, and registration channel; dynamic portraits: recent interest tags (generated through behavioral clustering), device fingerprints, and network environment; social relationships: follow lists, number of fans, and group affiliation.
[0132] The object information of the sample object is a multi-dimensional description of the recommendable entity, for example, static features: product category, content tag, publisher information; dynamic features: real-time popularity (such as 24-hour click volume), contextual association (such as matching purchased products); multimodal data: text description, image features, video frame embedding vector.
[0133] For example, user data includes: user behavior data for a certain product "browsed the details page for 120 seconds → added to the shopping cart → placed an order the next day"; user information "25 years old, female, Beijing, iOS user"; object information "dress, price range 300-500 yuan", and dress pictures.
[0134] The temporal distribution of user behavior data describes the statistical characteristics of user-object interactions over time, including the temporal order, frequency, and time span of these interactions. For example, new users' interactions tend to occur within the initial registration period (e.g., within 48 hours), while existing users' behavior may exhibit long-term, stable, and cyclical patterns. These differences in temporal distribution make it difficult for models to capture the short-term interests and preferences of target users.
[0135] The similarity between the user information of reference sample users and target sample users indicates the degree of match between the two user groups in terms of static attributes and dynamic features. This is quantified by calculating the cosine similarity or Euclidean distance between the feature vectors. For example, if 80% of the target users are female between the ages of 18 and 25, the higher the proportion of females in the same age group among the reference users, the higher the similarity. This metric is used to select reference user data that aligns with the target user profile and reduce distribution bias.
[0136] The data sparsity of reference user data is the density of the reference user's behavior records, typically measured by the number of interactions or interaction partners per unit time. For example, high-sparsity users with fewer than 10 monthly interactions have fragmented behavior patterns, while low-sparsity users with more than 50 monthly interactions have complete behavior chains. Using sparsity-stratified sampling, we can simulate the low data density characteristics of target users.
[0137] For example, in a content community application, the target user is a newly registered user, whose behavior data is concentrated within 48 hours of registration. The reference user data is sampled using a time-decay weight, with a 90% probability of retaining recent behavior (e.g., within 24 hours) and a 30% probability of retaining behavior from three days ago, ensuring that the time distribution is close to the target user.
[0138] For example, in a content community application, 60% of the target user group is female, aged between 18 and 25, and uses specific electronic devices. The reference user data is filtered out to identify a subset where the female user accounts for 60%. From this subset, users with the same age and device distribution as the target user group are randomly sampled to ensure user profile alignment.
[0139] For example, in a content community application, some low-activity users (average monthly interactions <10 times) are included in the reference user base, and their behavior sparsity is similar to that of the target users. These users are sampled with a 2x weight, while highly active users (average monthly interactions >50 times) are sampled with a 0.5x weight to balance the data distribution.
[0140] In the embodiments of this specification, through time window sampling, attribute similarity matching and data sparsity sampling, a training task that is closer to the distribution of new users is constructed, which solves the problem of large differences in the distribution of old user data and new user data, makes the training task more targeted, and improves the model generalization ability.
[0141] In an optional embodiment of the present specification, the user data includes user behavior data of sample users for sample objects and user information of the sample users, and the preference features of the user data include user interest tags in the user information and / or user behavior patterns in the user behavior data; in step 106, the user data is divided into a plurality of meta-learning task sets for recommendation scenarios, including the following specific steps: dividing the user data into a plurality of meta-learning task sets for recommendation scenarios according to the user interest tags; and / or dividing the user data into a plurality of meta-learning task sets for recommendation scenarios according to the user behavior patterns.
[0142] User interest tags in user information can be explicit interest categories selected by users, such as "beauty", "travel", "technology", etc., or they can be implicit interest vectors generated through behavioral clustering or semantic analysis, such as BERT embedding clustering based on user favorite content to generate potential interest clusters such as "ingredients" and "outdoor adventure".
[0143] User behavior patterns in user behavior data can be time-ordered sequences of interactions, such as "search → browse details → add to favorites," or they can be the frequency of interactions within a unit of time, such as >20 add to favorites per day. For example, we can use LSTM or Transformer to encode behavior sequences and cluster them to generate behavior pattern clusters. We can define pattern rules (such as "visiting the same category for three consecutive days"), filter user behavior data that meets these rules, and identify user behavior patterns.
[0144] For example, in a content community application, the user's explicit label "home renovation" is mapped to the "space optimization solution recommendation scenario", and its meta-learning support set is the user's collection of "small apartment layout" content, and the query set is to predict the user's click rate on smart home products; the implicit label "ingredients" is generated through clustering and mapped to the "professional skin care analysis scenario", the support set is the interactive sequence of ingredient comparison content, and the query set is to predict the user's conversion rate for high-concentration essences.
[0145] For example, in a content community application, the target user's behavior pattern of "short-term, high-frequency browsing" (browsing ≥ 20 articles in a single session but with a dwell time of < 30 seconds) is detected and mapped to a "shallow interest capture scenario." The meta-learning support set consists of user behavior sequences with similar patterns among reference users (with 30% random jump noise added to simulate the target user).
[0146] In the embodiments of this specification, user data is divided into meta-learning task sets for multiple recommendation scenarios based on user interest tags in user information and / or user behavior patterns in user behavior data, thereby improving the diversity of the constructed meta-learning task sets, improving the effectiveness of the multi-task training mechanism to integrate common knowledge, and further improving the generalization ability of the model.
[0147] In an optional embodiment of the present specification, before step 106, the following specific steps are further included: sampling the user data according to a preset ratio to obtain sampled user data, wherein the preset ratio is set according to the quantity distribution of the sample objects.
[0148] The preset ratio is set based on the distribution of sample objects to ensure that the representativeness and diversity of different sample objects are fully considered during the sampling process. For example, in a product recommendation scenario, different sampling ratios are used for a large number of mainstream products and a small number of long-tail products to balance their proportions in the training data. A lower sampling rate is used for mainstream products, while a higher sampling rate is used for long-tail products. This ensures that the model's recommendation effect on popular products is enhanced while enhancing its ability to discover long-tail products.
[0149] For example, in content community applications, for popular content types such as videos and graphics, a lower sampling rate can be set because they have a large amount of user interaction data; while for relatively niche content types such as audio and live broadcasts, the sampling rate should be appropriately increased to ensure that these long-tail contents can be fully represented in the meta-learning task set. In addition, when constructing the query set, a certain proportion of long-tail content is specially introduced, that is, content with less user interaction but may have unique value. User data is sampled according to the preset ratio to obtain sampled user data, where the preset ratio is set according to the number distribution of sample objects.
[0150] In the embodiments of this specification, by sampling user data according to a preset ratio based on the distribution of the number of sample objects, not only can the diversity and representativeness of the meta-learning task set be effectively improved, but the different needs of mainstream and long-tail sample objects can also be taken into account during the training process, thereby improving the performance of the recommendation model in handling sparse data and cold start problems, and achieving more accurate and personalized recommendation services. At the same time, this method can also help alleviate the bias problem caused by data imbalance, so that the ultimately trained recommendation model has stronger generalization ability and higher recommendation accuracy.
[0151] In an optional embodiment of the present specification, the user data includes user behavior data of sample users for sample objects, user information of sample users, and object information of sample objects; in step 106, the pre-trained recommendation model is trained based on the meta-learning task set to obtain a meta-parameter recommendation model, including the following specific steps: extracting the meta-learning task set of the current recommendation scenario from the meta-learning task set of multiple recommendation scenarios; extracting the current group of user data from the meta-learning task set of the current recommendation scenario; inputting the user information in the current group of user data and the object information in the current group of user data into the adaptation recommendation model to obtain a predicted recommendation result; calculating the recommendation loss value based on the difference between the predicted recommendation result and the user behavior data in the current group of user data; adjusting the model parameters of the pre-trained recommendation model based on the recommendation loss value, returning to execute the step of extracting the current group of user data from the meta-learning task set, until all groups of user data are extracted, returning to execute the step of extracting the meta-learning task set of the current recommendation scenario from the meta-learning task set of multiple recommendation scenarios, until all groups of user data are extracted, and obtaining the meta-parameter recommendation model.
[0152] The current recommendation scenario is a specific recommendation application scenario processed sequentially during the meta-learning training process, and its meta-learning task set includes a subset of user data associated with that scenario. For example, in a content community application, the current recommendation scenario can be specific to the "beauty ingredient in-depth analysis recommendation scenario," and its meta-learning task set includes the user interaction sequence and associated features of the skincare ingredient analysis content.
[0153] The current set of user data is a batch of training data units extracted from the meta-learning task set for the current recommendation scenario. It contains combined samples of the support set and the query set. For example, in the "Unpopular Travel Guide Recommendation Scenario," the current set of user data might include user information such as registration channels and device characteristics for 50 users, as well as user click and favorite data on outdoor adventure content, along with the content's multimodal features and location tags.
[0154] The predicted recommendation result is the probabilistic output of the pre-trained recommendation model for user-object matching, specifically representing the user's preference score for the object. For example, in a dual-tower model architecture, the user tower outputs the user feature vector u, and the object tower outputs the object feature vector v. The predicted recommendation result is the cosine similarity score sim(u, v), which is converted to a click probability p∈[0,1] using a Sigmoid function.
[0155] The difference between the predicted recommendation results and the user behavior data in the current group of users is quantified using a preset loss function to quantify the degree of deviation between the model's prediction and the actual interaction behavior. For example, when the user behavior data contains a binary label y∈{0,1} (click / not click), the difference is calculated using the binary cross-entropy loss function: L = -[y·log(p)+(1-y)·log(1-p)], where p is the model's predicted click probability.
[0156] The recommendation loss is the loss function calculated based on the current set of user data and is used to guide model parameter updates. For example, in the "Dynamic Home Renovation Recommendation Scenario," the average BCE loss calculated for a batch of 100 users was 0.32, reflecting the degree of deviation between the model's current predictions and the user's actual collection behavior.
[0157] Based on the recommendation loss value, adjust the model parameters of the pre-trained recommendation model. One optional method is to determine the gradient of the pre-trained recommendation model based on the recommendation loss value, and adjust the model parameters of the pre-trained recommendation model according to the gradient of the pre-trained recommendation model. For example, in the third iteration, the gradient calculation shows that the embedding layer parameters of the user tower need to be adjusted in the direction of reducing the long-tail content prediction error. The parameters are updated accordingly: Among them, L_{Support} is the loss function of the support set and α is the learning rate.
[0158] For example, in the meta-learning training process of the recommendation model for content community applications, the meta-learning task sets of five recommendation scenarios are processed in sequence. For the "brand topic immersive recommendation scenario", the pre-trained recommendation model first loads the support set of the scenario (interaction records of users participating in user-generated content on the topic of sunscreen) and obtains the temporary parameter θ' through three gradient updates. The meta-loss is then calculated on the query set (exposure data of new sunscreen product content) and the meta-parameter θ is updated through backpropagation. After iterating each scenario 200 times, it switches to the next scenario until the meta-loss of all scenarios converges to below the threshold of 0.25, and the meta-parameter recommendation model is output.
[0159] In the embodiments of this specification, an iterative multi-task training mechanism is used to alternate parameter optimization across multiple recommendation scenarios, allowing the model to gradually learn common representation capabilities across scenarios. The construction of a meta-learning task set for multiple recommendation scenarios simulates the potential behavioral distribution of target users, while the batch training strategy effectively balances computational efficiency and generalization performance. While ensuring that the model quickly adapts to the target recommendation scenario, the difference-driven loss calculation effectively mitigates the negative transfer effect, making it suitable for target recommendation scenarios where user behavior is sparse and the distribution shift is significant.
[0160] In an optional embodiment of the present specification, the recommendation loss value is calculated based on the predicted recommendation results and the user behavior data in the current group of user data, including the following specific steps: based on the predicted recommendation results and the user behavior data in the current group of user data, the recommendation loss value is calculated according to the preset weight, wherein the preset weight is set based on the similarity between the current recommendation scenario and the target sample user.
[0161] The preset weight is a weighting coefficient dynamically assigned based on the degree of match between the recommended scenario and the target user group's behavioral distribution. Its numerical range is mapped to the interval [0, 1] through normalization. Specifically, the preset weight is calculated using the following formula: wi = softmax(similarity i), where similarity i represents the distribution similarity score between the i-th recommended scenario and the target sample user. For example, in content community applications, high-weight tasks (such as wi>0.3) are prioritized for meta-parameter updates to strengthen knowledge transfer related to target user preferences.
[0162] The similarity between the current recommendation scenario and the target sample user represents the degree of match between the two in terms of user behavior patterns and interest distribution, and is quantified using at least one of the following methods:
[0163] 1. Cosine Similarity: Calculates the cosine angle between the behavior feature vector of the meta-learning support set of users in the recommendation scenario and the behavior feature vector of the target user. For example, in the "cosmetic ingredient recommendation scenario," the frequency distribution of the user's favorite ingredient analysis content is extracted as a feature vector, and the similarity score calculated with the target user's behavior vector from the previous 24 hours is 0.78.
[0164] 2. KL Divergence: Measures the difference between the recommended scenario user behavior distribution P and the target user distribution Q. The calculation formula is: For example, in the "unpopular travel guide recommendation scenario", the KL divergence value of the target user's click rate distribution Q for outdoor equipment content and the scenario support set distribution P is 1.25, so the similarity score is si = -D KL .
[0165] For example, in the meta-learning process of the content community application, the similarity of each recommendation scenario is calculated for the target user group (users newly registered within 48 hours):
[0166] Scenario 1 (Beauty ingredient analysis): Calculate the cosine similarity score of 0.85 by comparing the user behavior vector (ingredient content click-through rate, collection time) with the target user behavior.
[0167] Scenario 2 (home renovation plan): KL divergence-based distribution difference score -1.32 (similarity score 1.32)
[0168] Scenario 3 (Brand Topic Immersion): Calculate the similarity score to 0.65
[0169] After softmax normalization, the preset weight distribution is w1 = 0.42, w2 = 0.37, w3 = 0.21. In the cross-task optimization stage, the meta-parameter update formula is adjusted to: Where L_{Query}^i is the loss value for the query set in the i-th scenario, and β is the meta-learning rate. This mechanism enables highly similar cosmetic ingredient analysis scenarios to dominate parameter updates, strengthening the model's ability to capture the target user's skincare preferences.
[0170] In the embodiments of this specification, a dynamic task reweighting strategy is used to adaptively adjust the training weights based on the distribution similarity between the recommended scenarios and the target users, prioritize knowledge transfer for highly relevant scenarios, improve the accuracy of the model in capturing cold-start user preferences, balance the model's rapid adaptation to new user data and its generalization capabilities for diverse tasks, optimize recommendation performance, and improve model robustness.
[0171] In an optional embodiment of the present specification, the user data includes user behavior data of sample users for sample objects, user information of sample users, and object information of sample objects; step 108 includes the following specific steps: extracting the current group user data from the adaptation task set of the target recommendation scenario; inputting the user information in the current group user data and the object information in the current group user data into the pre-trained recommendation model to obtain a predicted recommendation result; calculating the reward function value based on the similarity between the predicted recommendation result and the user behavior data in the current group user data; adjusting the model parameters of the meta-parameter recommendation model based on the reward function value, and returning to execute the step of extracting the current group user data from the adaptation task set of the target recommendation scenario, until the reward function value reaches a threshold, thereby obtaining an adapted recommendation model for the target recommendation scenario.
[0172] The current set of user data is a batch of training data units extracted from the adaptation task set of the current recommendation scenario, and includes combined samples of the support set and the query set. For example, in the "cosmetic ingredient recommendation scenario," the current set of user data might include user information such as registration channels and device characteristics for 100 users, as well as user click and favorite behavior data on cosmetic ingredient recommendations, along with the associated multimodal features and location tags of the content.
[0173] The predicted recommendation result is the probabilistic output of the meta-parameter recommendation model for user-object matching, specifically representing the user's preference score for the object. For example, in a dual-tower model architecture, the user tower outputs the user feature vector u, and the object tower outputs the object feature vector v. The predicted recommendation result is the cosine similarity score sim(u, v), which is converted to a click probability p∈[0,1] using the Sigmoid function.
[0174] The similarity between the predicted recommendation results and the user behavior data in the current group of user data represents the degree of consistency between the model's predicted preferences and actual user interaction behaviors, and is quantified by a preset reward function. For example, a multi-dimensional reward mechanism is adopted, including immediate click feedback and long-term interest indicators: 1. Instant feedback indicators: click-through rate (CTR), standardized value of dwell time (T / 60 seconds), and interaction conversion rate (CVR). 2. Long-term interest indicators: secondary visit rate, cross-session content consistency (the similarity of the topic distribution of the content browsed by users in adjacent sessions is calculated using cosine similarity). The reward function is defined as: R = 0.5*CTR + 0.2·min(T / 60, 1) + 0.15*CVR + 0.15*secondary visit rate, where CTR is the click-through rate, T is the number of seconds of dwell time, and CVR is the collection / like conversion rate.
[0175] The reward function value is a weighted score calculated based on multi-dimensional indicators, reflecting the degree to which the recommendation result matches the target user's preferences. For example, in the "Beauty Ingredient Recommendation Scenario", a new user clicks on 3 articles (CTR = 0.3), stays for an average of 45 seconds (T / 60 = 0.75), saves 1 article (CVR = 0.1), and has a secondary visit rate of 0.4. The reward value is calculated as: R = 0.5*0.3+0.2*0.75+0.15*0.1+0.15*0.4=0.15+0.15+0.015+0.06=0.375
[0176] Based on the reward function value, the model parameters of the meta-parameter recommendation model are adjusted. One optional method is to use the policy gradient method to adjust the model parameters of the meta-parameter recommendation model based on the reward function value. The policy gradient method is a proximal policy optimization algorithm or a Q-learning optimization strategy. For example, the proximal policy optimization algorithm is used to avoid performance fluctuations by limiting the policy update step size. The specific steps include: calculating the probability ratio of the current strategy to the old strategy Constructing an alternative objective reward function: in, is the advantage function estimate, and ∈=0.2 is the clipping threshold. For example, when the parameters are updated, if the advantage value of a recommended action is And the probability ratio r t (θ)=1.3, min(1.3*0.5, 1.2*0.5)=0.6, ensuring that the update amplitude is controllable.
[0177] For example, during the adaptation training of a recommendation model for a content community application, 100 sets of new user data were extracted from the adaptation task set for the "unpopular travel guide recommendation scenario." Each set contained user device characteristics, registration channels, and the top 10 behavior records. The meta-parameter recommendation model outputs a predicted click probability distribution based on user characteristics and multimodal features of outdoor adventure content (image and text semantic vectors, geolocation embedding). User feedback on the recommendation results is collected in real time, the reward value is calculated, and the model parameters are updated. For example, a user clicks on the "Uninhabited Island Adventure Guide" content (CTR + 0.1), stays for 80 seconds (T / 60 = 1.33 → take 1), collects the content (CVR + 0.1), and visits the topic channel a second time (second visit rate + 0.2). The single-step reward value R = 0.5*0.1+0.2*1+0.15*0.1+0.15*0.2=0.05+0.2+0.015+0.03=0.295. The PPO algorithm is used to accumulate multi-step rewards. When the average reward value exceeds the threshold of 0.35 for three consecutive batches, the training is terminated and an adapted recommendation model for the "unpopular travel guide recommendation scenario" is output.
[0178] In the embodiments of this specification, a dynamic parameter adjustment mechanism driven by reinforcement learning is used to achieve real-time capture of user preferences in target scenarios and rapid model adaptation. The model parameters are dynamically adjusted based on the target reward function, and the real-time recommendation effect in the cold start phase is optimized through the exploration-utilization mechanism to quickly adapt to the preference drift of cold start users.
[0179] In an optional embodiment of the present specification, step 104 includes the following specific steps: inputting the user data of the sample user for the sample object into the initial recommendation model, and outputting the encoding features of the user data; judging, based on the encoding features of the user data, by a domain classifier, whether the user data corresponds to a reference sample user or a target sample user, and obtaining a predicted classification result; calculating an adversarial loss value based on the predicted classification result; adjusting the parameters of the initial recommendation model based on the adversarial loss value, and returning to execute the step of inputting the user data of the sample user for the sample object into the initial recommendation model, and obtaining a pre-trained recommendation model when the adversarial loss value reaches a preset threshold.
[0180] The predicted classification result is the binary probability distribution output by the domain classifier, which is converted into a discrete label using a threshold (e.g., 0.5). For example, for a newly registered user who has been registered for three days, uses a specific device, registered via an ad link, and clicked on five pieces of content in the previous 24 hours, the domain classifier outputs a probability value of 0.82, making them a "target sample user."
[0181] Adversarial loss is a binary cross-entropy loss calculated based on the predicted classification results and the true user source labels. It represents the error in the domain classifier's judgment of user type. The specific formula is: L_{Adversarial}=BCE(D(user_features,domain_labels), where D is the domain classifier and BCE is the binary cross-entropy loss.
[0182] The preset threshold is the stopping condition for adversarial training. For example, training is terminated when the moving average of the adversarial loss value for three consecutive training rounds is lower than 0.15, or when the number of training rounds reaches 100.
[0183] Based on the adversarial loss value, the parameters of the initial recommendation model are adjusted. One optional method is: based on the adversarial loss value, the gradient of the initial recommendation model is determined; the gradient of the initial recommendation model is reversed to obtain a reversed gradient; and the model parameters of the initial recommendation model are adjusted according to the reversed gradient.
[0184] Reversed gradient is the adjustment amount after multiplying the gradient by the reverse coefficient (usually -1), which is used for adversarial training. The specific implementation is: λ is the weight of adversarial training, is the gradient of the initial recommendation model.
[0185] For example, in the “new user interest cold start recommendation scenario”, the domain classifier receives the following feature inputs:
[0186] User A: Registered 2 days ago, device a, registered via a short video ad, clicked on 8 beauty content articles in the first 24 hours. User B: Registered 180 days ago, device b, registered via organic traffic, and clicked on an average of 20 home furnishing content articles per day in the past 7 days.
[0187] The domain classifier outputs the predicted classification result pA=0.91 for user A and the true label yA=1; and outputs pB=0.05 for user B and the true label yB=0.
[0188] For example, in a content community scenario, gradient reversal forces the feature extraction layer of the recommendation model to generate encoded features of user data that are difficult for the domain classifier to distinguish the source. Calculate the batch adversarial loss value L_{Adversarial} = 0.12. Perform gradient reversal based on the adversarial loss value and calculate the gradient of the initial recommendation model. Reverse the gradient of the initial recommendation model to obtain the reverse gradient According to the inverted gradient, the parameters of the initial recommendation model are adjusted so that user A's beauty content click behavior and user B's home preferences are mapped to the same feature space, generating an input vector that is invariant across domains.
[0189] In the examples of this specification, by introducing an adversarial training mechanism combining domain classifiers and gradient reversal, we significantly improve the recommendation performance for cold-start users in content communities. The gradient reversal operation forces the initial recommendation model to learn domain-invariant feature representations across user groups, effectively mitigating the negative transfer effect caused by differences in data distribution between target and reference users.
[0190] In an optional embodiment of the present specification, after step 108, the following specific steps are also included: obtaining user data of the target user for the recommended object, wherein the target user and the recommended object are matched through an adapted recommendation model; and fine-tuning the adapted recommendation model of the target recommendation scenario based on the user data.
[0191] Target users are cold-start users who require real-time recommendations in target recommendation scenarios. Their user data is low-density and dynamically evolving. For example, in a content community application, target users are newly registered users within 48 hours with fewer than 10 behavior records, and their initial preferences are not yet stable.
[0192] Recommended content is a collection of candidate content that can be distributed in real time within the target recommendation scenario. For example, in the "sunscreen technology topic recommendation scenario," recommended content includes beauty ingredient analysis content, sunscreen product review videos, and related topic discussion threads published in the past 24 hours.
[0193] The target user's user data for recommended content includes incremental behavioral data generated by interactions with recommended content, including but not limited to real-time clickthroughs, session duration, cross-content jump paths, and secondary visit frequency. For example, a new user's behavioral data after their first recommendation might include in-depth browsing of three sunscreen articles (staying on each article for >90 seconds), saving one ingredient comparison analysis article, and quickly skipping two advertisements.
[0194] For example, in the "sunscreen technology topic recommendation scenario" of a content community application, real-time user interaction data, including click behavior, dwell time, and cross-session jump paths, is collected through tracking technology. For example, user U1 clicks on two sunscreen ingredient analysis articles within 5 minutes of the initial recommendation, dwelling on each article for 120 seconds and 95 seconds respectively, and collects one of them. A 5-minute sliding window is set, and user behavior data within the window is aggregated to generate incremental training batches. For example, each batch contains the device fingerprints, behavior sequences, and multimodal content features of 32 users. Model parameters are updated based on the gradient of the reward function calculated in real time. The reward function is defined as: R = 0.6*CTR + 0.25*min(60 / T, 1) + 0.15*CVR, where CTR is the click-through rate, T is the dwell time in seconds, and CVR is the collection conversion rate. The learning rate is decayed according to user activity. The learning rate is set to 0.01 and drops to 0.001 after 48 hours to balance rapid adaptation and stability. Inject 10% random perturbations into user device features (such as replacing the suffix characters of the device model) to enhance the robustness of the adaptive recommendation model to sparse noise.
[0195] In the embodiments of this specification, through the online fine-tuning mechanism, the adaptive recommendation model can capture the user's dynamic behavior patterns in real time, quickly adapt to the evolution trend of personalized preferences, and while ensuring the recommendation accuracy, greatly reduce the computational overhead of model training, and achieve efficient real-time response in the cold start phase.
[0196] See also Figure 2 , Figure 2 A flowchart of an object recommendation method provided by an embodiment of this specification is shown, including the following specific steps:
[0197] Step 202: Obtain user information of the target user, object information of multiple candidate recommendation objects, and an adaptation recommendation model for the target recommendation scenario, wherein the adaptation recommendation model for the target recommendation scenario is trained according to the above-mentioned recommendation model training method.
[0198] The embodiments of this specification are applied to an application or system platform for object recommendation, such as an e-commerce platform, a social media application, or a content community application.
[0199] The target user's user information is a multi-dimensional feature set that represents the user's identity and behavioral preferences, including static attributes (such as age, gender, and registration channel) and dynamic behavioral data (such as recent click history and search keywords). For example, in a content community application, user information includes registration time, device model, initial interest tags, and keyword distribution of content browsed in the previous 24 hours.
[0200] The object information of multiple candidate recommendation objects is a multi-dimensional description of the recommended entity, including content features (such as text semantic vectors and image visual features) and metadata (such as release time and popularity value). For example, in the "sunscreen technology topic recommendation scenario", the object information of the candidate recommendation object includes the text summary of the sunscreen content, the multimodal features of the component analysis diagram, and the topic relevance score.
[0201] The adaptive recommendation model for target recommendation scenarios is a scenario-customized model trained through a meta-learning framework. Its model parameters incorporate common knowledge across scenarios and the unique characteristics of the target scenario. For example, the model uses a dual-tower architecture: the user tower encodes user information to generate a preference vector, while the object tower encodes object information to generate a content feature vector. Matching scores are calculated based on similarity.
[0202] To obtain the user information of the target user, one optional method is to extract real-time interaction records from the user behavior log database and generate structured feature vectors through feature engineering; another optional method is to call the application programming interface (API) of the user portrait system to obtain dynamic portrait data; another optional method is to use point-of-sale technology to collect user device features and session behavior sequences in real time, which is not limited here.
[0203] To obtain the object information of multiple candidate recommendation objects, one optional method is to query the candidate content list under the target scenario from the content management system and extract its multimodal features; another optional method is to generate semantic embedding vectors of text and images through a content understanding model; another optional method is to obtain the dynamic weight value of the content in combination with a real-time heat calculation module, which is not limited here.
[0204] To obtain the adapted recommendation model for the target recommendation scenario, one optional method is to load the pre-trained adapted recommendation model parameters from the model repository; another optional method is to call the model service interface in real time through a distributed computing cluster; another optional method is to load the model instance from the local cache based on the scenario identifier, which is not limited here.
[0205] For example, in a content community scenario, the target user is a newly registered user, whose user information includes device type, registration channel, and initial interest tags. Candidate recommendation objects are collections of content related to specific technical topics, with object information including text summaries, visual features, and interaction popularity. The adaptive recommendation model is trained using a meta-learning framework and deployed in a cloud-based inference service.
[0206] Obtain the user information of the target user, the object information of multiple candidate recommendation objects, and the adaptive recommendation model of the target recommendation scenario. By dynamically obtaining the user's real-time behavior data and multimodal object features, it effectively solves the data sparsity problem in the cold start phase. The adaptive recommendation model integrates common knowledge across scenarios, significantly improving the generalization ability of the target scenario, while reducing the computational overhead of model training and ensuring the real-time response efficiency of the recommendation system.
[0207] Step 204: Input the user information and the object information of multiple candidate recommendation objects into the adaptation recommendation model to obtain the target recommendation result.
[0208] The target recommendation result is a set of matching scores between the candidate recommendation objects and the target user. For example, the output vector is [0.85, 0.72, 0.63, …], which represents the probability of each sunscreen content matching the preference of user U1.
[0209] User information and object information of multiple candidate recommendation objects are input into the adaptation recommendation model to obtain the target recommendation result. An optional method is: the user feature encoding module maps the user information into a low-dimensional dense vector; the object feature encoding module converts the multimodal object information into a feature representation in a unified semantic space; the matching degree calculation module generates a user-object association score through a similarity metric (such as cosine similarity); and after normalization, the score is mapped to a predicted probability in the interval [0,1].
[0210] For example, the user feature vector and the feature vector of a certain technical analysis content are calculated to have a similarity of 0.92, and the predicted click probability after normalization is 0.88, ranking first in the candidate list.
[0211] User information and object information of multiple candidate recommendation objects are input into the adaptation recommendation model to obtain the target recommendation results. Based on the parallel processing of the dual-tower architecture, efficient matching of user preferences and object features is achieved, thereby improving the accuracy and interpretability of the recommendation results.
[0212] Step 206: Based on the target recommendation result, determine a target recommendation object from multiple candidate recommendation objects.
[0213] The target recommendation object is a content collection selected based on matching ranking and diversity constraints.
[0214] Based on the target recommendation results, the target recommendation object is determined from multiple candidate recommendations. One option is to sort the candidates in descending order of match score, introduce a maximum marginal relevance algorithm to balance relevance and diversity, and select the top K highly rated content with diverse topic distributions. For example, in a technical topic recommendation scenario, analysis, evaluation, and practice content are each selected to account for 30% to avoid content homogeneity.
[0215] For example, the target recommendation results include 3 technical analysis contents, 2 product evaluation contents and 1 user practice report, covering multi-dimensional information needs.
[0216] Based on the target recommendation results, the target recommendation object is determined from multiple candidate recommendation objects. According to the target recommendation results generated in the previous step, the recommendation objects suitable for the target user are screened out, which not only enhances the accuracy of the recommendation decision, but also effectively improves the user experience and satisfaction.
[0217] Step 208: Recommend the target recommendation object to the target user.
[0218] To recommend the target recommendation object to the target user, one optional way is to push the sorting results to the user terminal page through the content distribution interface; another optional way is to adaptively render the recommendation results based on the user terminal type; another optional way is to dynamically optimize the display strategy based on the real-time feedback mechanism.
[0219] For example, the top 6 items of content are displayed in the personalized recommendation flow of the target user according to the matching degree, and are marked as "generated based on your interests". At the same time, the user's real-time interaction behavior is recorded for model iterative optimization.
[0220] In the embodiments of this specification, by employing a meta-learning-based recommendation model training method, this object recommendation method can provide highly personalized recommendations for specific target recommendation scenarios. By employing a meta-parameter recommendation model and further adapting training to specific target recommendation scenarios, it can reduce computing resources and time costs while ensuring recommendation accuracy.
[0221] In an optional embodiment of the present specification, step 206 includes the following specific steps: based on the target recommendation result, determining at least one target recommendation object from multiple candidate recommendation objects, and sorting the at least one target recommendation object to obtain a recommendation object list;
[0222] Step 208 includes the following specific steps: based on the recommendation object list, recommending the target recommendation object to the target user.
[0223] The recommended object list is a set of ranked content generated based on the target recommendation results, which contains an ordered arrangement of candidate objects screened by matching scores and diversity constraints. Optionally, the recommended object list is constructed in the following ways: Candidate screening: Preliminary screening of highly relevant objects from the candidate pool based on a matching threshold (such as click probability ≥ 0.7), such as screening out graphic and text content of potential interest to users in the content community. Diversity weighting: Use the maximum marginal relevance (MMR) algorithm to balance the content topic distribution and matching degree, and suppress the aggregation of homogeneous content. For example, in an e-commerce scenario, ensure that the recommendation list includes products from different categories such as clothing, home furnishings, and electronics. Dynamic sorting: Dynamically adjust the sorting weights based on real-time popularity, user context (such as device type, network environment) and feedback data. For example, mobile users prioritize short video content, and PC users prioritize long graphic and text analysis.
[0224] Sort at least one target recommendation object to obtain a list of recommended objects. One optional method is: multi-dimensional weighted sorting: comprehensive matching score, real-time popularity value and content novelty index, generate a weighted total score and sort. For example, the ranking score is defined as: S = 0.6*matching+0.3*popularity+0.1*novelty. Another optional method is: dynamic feedback adjustment: dynamically adjust the ranking weight based on the user's real-time interactive behavior (such as clicks, length of stay). For example, if the user's click rate for a certain type of content exceeds 50%, the ranking weight coefficient of this type of content is increased. Another optional method is: context-aware sorting: adaptively adjust the display order according to the user terminal type, network environment and current session context. For example, mobile users give priority to displaying graphic content, and PC users give priority to displaying long video content. This is not limited here.
[0225] For example, in a content community application, the candidate content library contains 1,000 graphic and text contents, and the adaptation recommendation model outputs the matching score of each content. The top-50 content is selected in descending order of score, and the MMR algorithm is applied to calculate the subject difference between the content, and duplicate content with the same theme is eliminated (for example, only the one with the highest matching degree is retained for 3 articles belonging to the same "sunscreen ingredient analysis"). Finally, 20 highly relevant and thematically diverse content are retained. The ranking is dynamically adjusted based on real-time interactive data, and the ranking of content with a collection increase of more than 15% in the past hour is increased by 5 places to form the final recommendation list. For example, a certain "new sunscreen technology actual test" content jumped from the original 8th place to the 3rd place due to a 20% increase in collections in the past hour. The optimized list covers three types of subject content: ingredient analysis, product evaluation, and user practice.
[0226] In the examples of this specification, a multi-dimensional weighted ranking and dynamic feedback adjustment mechanism are used to optimize the generation of recommendation lists. By evaluating and ranking, and dynamically adjusting ranking weights based on user interactions, this ensures that recommendations meet users' personalized needs while also being timely and diverse. This significantly improves the relevance of recommendation results and user experience, enhancing user engagement and satisfaction.
[0227] In an optional embodiment of the present specification, there are at least three target recommendation objects, and the at least three target recommendation objects include a first target recommendation object, a second target recommendation object, and a third target recommendation object; before step 208, the following specific steps are also included: sorting the at least three target recommendation objects from high to low to obtain a high-order first target recommendation object and at least two low-order second target recommendation objects; based on the similarity between the at least two second target recommendation objects and the first target recommendation object, filtering out the second target recommendation objects with high similarity.
[0228] The first target recommendation object is the candidate with the highest matching score. For example, in a content community application, a piece of content called "In-depth Analysis of Sunscreen Ingredients" is listed as the first target recommendation object because it highly matches the user's interest tag (scoring 0.92). The second target recommendation object is the candidate with a lower matching score than the first target recommendation object. For example, two pieces of content called "Sunscreen Product Reviews" and "Outdoor Sunscreen Practice Guide" with scores of 0.88 and 0.85, respectively, are listed as the second target recommendation objects.
[0229] Based on the similarity between at least two second target recommendation objects and the first target recommendation object, the second target recommendation objects with high similarity are screened out, that is, the MMR algorithm is applied to maximize the list diversity while retaining high-scoring objects. The specific formula is: Among them, Sim1 is the user-object matching degree, Sim2 is the subject similarity between objects, and λ=0.6 is the balance coefficient.
[0230] For example, in the "Sunscreen Technology Topic Recommendation Scenario" of the content community application, the first target recommendation object is the "Analysis of New Sunscreen Ingredients" note with a score of 0.92. The second target recommendation objects include the "Sunscreen Product Comparison and Evaluation" note with a score of 0.88 (topic similarity 0.75) and the "Outdoor Sunscreen Practical Skills" note with a score of 0.85 (topic similarity 0.35). Calculating the topic similarity between the two second target recommendation objects and the first target recommendation object, it was found that the similarity of the "Sunscreen Product Comparison and Evaluation" note exceeded the threshold of 0.7, so it was eliminated. The second highest-scoring "Sunscreen Technology Development Review" note (score 0.83, topic similarity 0.25) was selected from the candidate pool and added to the recommendation list to ensure that the content covers four topics: analysis, evaluation, practice, and review, to avoid homogenization of recommendation results.
[0231] In the embodiments of this specification, the maximum marginalization algorithm is used to balance the matching degree and the topic difference, ensuring that the final recommendation list covers multiple topic areas and avoiding the aggregation of homogeneous content, thereby improving the efficiency and quality of users' acquisition of information and further optimizing the performance of the recommendation system and the user interaction experience.
[0232] The following combined Figure 3 , taking the application of the object recommendation method provided in this specification in the content community as an example, the object recommendation method is further explained. Figure 3 A flowchart of a process for an object recommendation method applied to a content community provided by an embodiment of this specification is shown, including the following specific steps:
[0233] Step 302: Obtain old user data of old sample users for sample content and new user data of new sample users for sample content, wherein the user data includes user behavior data of sample users for sample content and user information of sample users.
[0234] For example, the data of old sample users may include their browsing, liking, and collection behaviors under the "Photography" category, as well as information such as user age, gender, and registration time; the data of new sample users (such as newly registered users) include their first interaction records with the "Fashion" category content, as well as characteristics such as device model and geographic location.
[0235] Step 304: Based on the time distribution of the user behavior data, the old user data is sampled to obtain the sampled old user data; based on the similarity between the user information of the old sample users and the user information of the new sample users, the old user data is sampled to obtain the sampled old user data; based on the data sparsity of the old user data, the old user data is sampled to obtain the sampled old user data.
[0236] For example, based on the time distribution of user behavior, the interaction data of old users in the "evening active period" (such as 20:00-24:00) is sampled first; by calculating the Jaccard similarity of the interest tags of old users and new users (such as both contain "photography" and "outfit" tags), old users with similarity ≥ 0.7 are screened; for old users with sparse data (such as only browsing 10 pieces of content), the interpolation method is used to supplement their behavior sequence.
[0237] Step 306: Based on the sampled old user data and new user data, construct user data of the sample user for the sample content.
[0238] For example, the collection behavior data of old users on "photography tutorial" content is combined with the browsing time data of new users on "dressing tutorial" to construct a cross-scene comparison sample set for analyzing the patterns of user interest migration.
[0239] Step 308: Sampling the user data according to a preset ratio to obtain sampled user data, wherein the preset ratio is set according to the quantity distribution of the sample content.
[0240] For example, if the "photography" category accounts for 60% of the sample content and the "travel" category accounts for 20%, the sampling amount of the "photography" category data will be reduced to 50% of the original data, and the sampling amount of the "travel" category will be increased to 150% to balance the scene coverage.
[0241] Step 310: Input the user data of the sample user for the sample object into the initial recommendation model, output the encoding features of the user data, determine whether the user data corresponds to the reference sample user or the target sample user based on the encoding features of the user data through the domain classifier, obtain a predicted classification result, calculate the adversarial loss value based on the predicted classification result, adjust the parameters of the initial recommendation model based on the adversarial loss value, return to the step of inputting the user data of the sample user for the sample object into the initial recommendation model, and obtain a pre-trained recommendation model when the adversarial loss value reaches a preset threshold.
[0242] For example, the initial recommendation model uses a Transformer encoder to process user behavior sequences, mapping the click duration of the user's browsed "Photography Tips" graphic content and "Styling Tutorials" video into a 256-dimensional feature vector. The domain classifier, a two-layer fully connected network structure, receives this feature vector and outputs a predicted classification result, indicating whether the user belongs to an old sample user (with probability 0.9) or a new sample user (with probability 0.1). In this case, the adversarial loss function is defined as a binary cross-entropy loss. When the actual user is a new user who has registered for 7 days, the loss value is calculated as -(1log(0.1)+0log(0.9))=2.3. Adversarial training is implemented through a gradient reversal layer. During backpropagation, the domain classifier gradient is multiplied by a weight of -0.5. This allows the recommendation model to simultaneously optimize two objectives when updating parameters: retaining sufficient user features for the recommendation task while confusing the domain classifier's ability to distinguish between new and old users. When the moving average of the adversarial loss value for three consecutive epochs drops below 0.65, the pre-training process is terminated. At this time, the classification index of the domain classifier drops from the initial 0.88 to 0.62, indicating that the recommendation model has learned a universal representation across new and old user groups.
[0243] Step 312: Divide the user data into multiple meta-learning task sets for recommendation scenarios based on user interest tags, and divide the user data into multiple meta-learning task sets for recommendation scenarios based on user behavior patterns, wherein any meta-learning task set includes a support set and a query set.
[0244] For example, based on user interest tags, the task sets are divided into scenario sets such as "mother and baby care" and "technology and digital"; based on user behavior patterns (such as "high frequency of likes and low collections" and "only browsing without interaction"), the task sets are divided into scenario sets such as "light browsers" and "deep interactors".
[0245] Step 314: Extract the meta-learning task set of the current recommendation scenario from the meta-learning task set of multiple recommendation scenarios, extract the current group user data from the support set of the meta-learning task set of the current recommendation scenario, input the user information of the sample user in the current group user data into the domain classifier, obtain the predicted classification result of whether the sample user is an old sample user or a new sample user, input the user information in the current group user data and the content information in the current group user data into the pre-trained recommendation model, obtain the predicted recommendation result, calculate the recommendation loss value according to the preset weight based on the predicted recommendation result and the user behavior data in the current group user data, and calculate the recommendation loss value based on the predicted recommendation result. Classify the results, calculate the adversarial loss value, determine the gradient of the initial recommendation model based on the adversarial loss value, reverse the gradient of the initial recommendation model to obtain the reversed gradient, adjust the model parameters of the domain classifier according to the reversed gradient, adjust the model parameters of the pre-trained recommendation model based on the recommendation loss value, return to the step of extracting the current group of user data from the meta-learning task set, until the verification is completed with the query set, return to the step of extracting the meta-learning task set of the current recommendation scenario from the meta-learning task set of multiple recommendation scenarios, until the meta-learning task sets of each recommendation scenario are extracted, and the meta-parameter recommendation model is obtained.
[0246] For example, the domain classifier predicts whether the user is an old user or a new user based on the user's gender and age. The pre-trained recommendation model inputs the user's interest tags and content tags and outputs "predicted click probability 0.8"; the recommendation loss value is calculated based on the cross-entropy loss between the predicted click rate and the actual click behavior, and the adversarial loss value is adjusted by backpropagating the domain prediction accuracy to adjust the domain classifier parameters.
[0247] Step 316: Extract new user data of new sample users for new sample content from the meta-learning task set of the new recommendation scenario, and construct an adaptation task set of the new recommendation scenario, wherein any adaptation task set includes a support set and a query set.
[0248] For example, the reward function value is calculated based on the weighted calculation of the user's stay time and click-through rate on the recommended content (such as stay time * 0.6 + click-through rate * 0.4). When the reward value exceeds the 0.7 threshold for three consecutive iterations, the model parameters are frozen and the scene adaptation is completed.
[0249] Step 318: Extract the current group user data from the support set of the adaptation task set of the new recommendation scenario, input the user information and content information in the current group user data into the pre-trained recommendation model to obtain the predicted recommendation result, calculate the reward function value based on the similarity between the predicted recommendation result and the user behavior data in the current group user data, adjust the model parameters of the meta-parameter recommendation model based on the reward function value, and return to execute the step of extracting the current group user data from the adaptation task set of the new recommendation scenario, until the adaptation recommendation model of the new recommendation scenario is obtained after verification is completed using the query set.
[0250] For example, the reward function value is calculated based on the weighted calculation of the user's stay time and click-through rate on the recommended content (such as stay time * 0.6 + click-through rate * 0.4). When the reward value exceeds the 0.7 threshold for three consecutive iterations, the model parameters are frozen and the scene adaptation is completed.
[0251] Step 320: Obtain user information of the new user, content information of multiple candidate recommended contents, and an adapted recommendation model for the new recommendation scenario.
[0252] For example, user information includes "18-24 years old", "registered for 3 days", and "device is device b"; candidate recommendation content information includes "title keywords", "tag weight", and "release time"; the adapted recommendation model is a meta-parameter model that has completed the "novice scenario" iteration.
[0253] Step 322: Input the user information and the content information of multiple candidate recommended contents into the adaptation recommendation model to obtain a new recommendation result.
[0254] For example, by inputting user interest tags (such as "photography" and "student") and content tags (such as "outfit" and "college life"), the model outputs a matching score between the content and the user's interests (such as 0.85) and generates a "recommendation priority" ranking.
[0255] Step 324: Based on the recommendation results, determine at least one target recommended content from multiple candidate recommended content, sort the at least one target recommended content from high to low, obtain a high-order first target recommended content and at least two low-order second target recommended content, and based on the similarity between the at least two second target recommended content and the first target recommended content, screen out the second target recommended content with high similarity.
[0256] For example, after sorting, the first recommendation is "Take photos of retro texture in minutes", and the second recommendation is "Beauty and storage tips for college dormitories". Because the label similarity between the latter and the first content exceeds the 0.6 threshold, the content is filtered out.
[0257] Step 326: Sort the at least three target recommended contents to obtain a recommended content list, and recommend the target recommended content to the new user based on the recommended content list.
[0258] For example, the recommendation list is arranged from high to low according to the "matching score" and displayed in the form of a double-column card. When the user scrolls down, more low-similarity content is dynamically loaded.
[0259] Figure 4 A front-end schematic diagram of a content recommendation method applied to a content community provided by an embodiment of this specification is shown. Figure 4 As shown:
[0260] The homepage of the content community application includes two vertical columns of recommended content cards. Each card contains a content cover image, a short title, a user avatar, and the name of the publishing account.
[0261] In the embodiments of this specification, time window sampling, attribute similarity matching and data sparsity sampling techniques are used to construct a training task set close to the distribution of new users from old user data, effectively solving the problem of data distribution differences between new and old users. The initial parameters with rapid adaptability are learned through a meta-learning framework, and adversarial training of domain classifiers and pre-trained recommendation models is combined to enable the model to efficiently adapt to the sparse data distribution of new users under few sample conditions. A dynamic task reweighting mechanism is introduced to adjust the task weights in real time based on similarity metrics such as cosine similarity, balancing the model's rapid adaptation to new user data and its generalization ability for diversified tasks. By designing the support set and query set differently, scenarios such as interest matching, sparse data and long-tail preferences are simulated, thereby enhancing the model's adaptability to unpopular content and long-tail users. Combining multimodal feature extraction and reinforcement learning technology, dynamic sorting of recommendation results and screening of similar content are achieved, significantly improving the robustness and real-time performance of the recommendation system. Compared with the traditional pre-training-fine-tuning method, this solution can still maintain high-precision recommendations in data-sparse scenarios, reducing the training cost in the cold start phase. At the same time, it supports joint modeling of new and old users and continuous adaptation of dynamic data distribution, effectively solving the problem of insufficient prediction accuracy in the recommendation system due to changes in user behavior distribution.
[0262] Corresponding to the above method embodiment, this specification also provides a recommendation system embodiment, Figure 5 FIG1 shows a schematic diagram of a structure of a recommendation system provided by an embodiment of this specification. Figure 5 As shown, the recommendation system includes a data processing module 502, a meta-learning module 504, an adaptation module 506 and a recommendation generation module 508;
[0263] The data processing module 502 is configured to obtain user data of sample users for sample objects, wherein the sample users include reference sample users and target sample users;
[0264] The meta-learning module 504 is configured to perform adversarial pre-training on the initial recommendation model based on the encoding features of the user data output by the initial recommendation model using a domain classifier of user type to obtain a pre-trained recommendation model, wherein the domain classifier determines whether the user data corresponds to a reference sample user or a target sample user based on the encoding features of the user data;
[0265] The data processing module 502 is further configured to divide the user data into a plurality of meta-learning task sets for recommendation scenarios;
[0266] The meta-learning module 504 is further configured to train the pre-trained recommendation model based on the meta-learning task set to obtain a meta-parameter recommendation model;
[0267] The adaptation module 506 is configured to construct an adaptation task set for the target recommendation scenario based on the meta-learning task set of the target recommendation scenario, and train the meta-parameter recommendation model based on the adaptation task set to obtain an adaptation recommendation model for the target recommendation scenario;
[0268] The recommendation generation module 508 is configured to obtain user information of the target user, object information of multiple candidate recommendation objects, and an adaptive recommendation model for the target recommendation scenario, input the user information and the object information of multiple candidate recommendation objects into the adaptive recommendation model, obtain a target recommendation result, determine a target recommendation object from multiple candidate recommendation objects based on the target recommendation result, and recommend the target recommendation object to the target user.
[0269] In an optional embodiment of the present specification, the data processing module 502 is further configured to: obtain reference user data of the reference sample user for the sample object, and target user data of the target sample user for the sample object; sample the reference user data to obtain sampled reference user data; and construct user data of the sample user for the sample object based on the sampled reference user data and the target user data.
[0270] In an optional embodiment of the present specification, the data processing module 502 is further configured to: sample the user data according to a preset ratio to obtain sampled user data, wherein the preset ratio is set according to the quantity distribution of the sample objects.
[0271] In an optional embodiment of the present specification, the user data includes user behavior data of sample users for sample objects, user information of sample users, and object information of sample objects; the meta-learning module 504 is further configured to: extract the meta-learning task set of the current recommendation scenario from the meta-learning task set of multiple recommendation scenarios; extract the current group user data from the meta-learning task set of the current recommendation scenario; input the user information in the current group user data and the object information in the current group user data into the pre-trained recommendation model to obtain a predicted recommendation result; calculate the recommendation loss value based on the difference between the predicted recommendation result and the user behavior data in the current group user data; adjust the model parameters of the pre-trained recommendation model based on the recommendation loss value, return to execute the step of extracting the current group user data from the meta-learning task set of the current recommendation scenario, until all groups of user data are extracted, return to execute the step of extracting the meta-learning task set of the current recommendation scenario from the meta-learning task set of multiple recommendation scenarios, until all meta-learning task sets of recommendation scenarios are extracted, and obtain a meta-parameter recommendation model.
[0272] In an optional embodiment of the present specification, the meta-learning module 504 is further configured to: calculate the recommendation loss value according to the preset weight based on the predicted recommendation results and the user behavior data in the current group user data, wherein the preset weight is set based on the similarity between the current recommendation scenario and the target sample user.
[0273] In an optional embodiment of the present specification, the user data includes user behavior data of sample users for sample objects, user information of sample users and object information of sample objects; the adaptation module 506 is further configured to: extract the current group user data from the adaptation task set of the target recommendation scenario; input the user information in the current group user data and the object information in the current group user data into the pre-trained recommendation model to obtain a predicted recommendation result; calculate the reward function value based on the similarity between the predicted recommendation result and the user behavior data in the current group user data; adjust the model parameters of the meta-parameter recommendation model based on the reward function value, and return to execute the step of extracting the current group user data from the adaptation task set of the target recommendation scenario, until the reward function value reaches a threshold, thereby obtaining an adapted recommendation model for the target recommendation scenario.
[0274] In an optional embodiment of the present specification, the meta-learning module 504 is further configured to: input the user data of the sample user for the sample object into the initial recommendation model, and output the encoding features of the user data; determine, based on the encoding features of the user data, by a domain classifier, whether the user data corresponds to a reference sample user or a target sample user, and obtain a predicted classification result; calculate the adversarial loss value based on the predicted classification result; adjust the parameters of the initial recommendation model based on the adversarial loss value, and return to the step of inputting the user data of the sample user for the sample object into the initial recommendation model, and obtain a pre-trained recommendation model when the adversarial loss value reaches a preset threshold.
[0275] The recommendation system provided in the embodiments of this specification ensures recommendation accuracy under low data density, achieves a balance between sparse data and high precision, and improves training efficiency. By adopting a meta-learning-based recommendation model training method, the object recommendation method can perform highly personalized recommendations for specific target recommendation scenarios.
[0276] The above is a schematic diagram of a recommendation system according to this embodiment. It should be noted that the technical solution of this recommendation system shares the same concept as the technical solutions for the recommendation model training method and the object recommendation method described above. For details not described in detail in the technical solution of the recommendation system, please refer to the description of the technical solutions for the recommendation model training method and the object recommendation method described above.
[0277] Figure 6 6 shows a block diagram of a computing device according to an embodiment of the present disclosure. Components of the computing device 600 include, but are not limited to, a memory 610 and a processor 620. The processor 620 is connected to the memory 610 via a bus 630, and a database 650 is used to store data.
[0278] The computing device 600 also includes an access device 640 that enables the computing device 600 to communicate via one or more networks 660. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 640 may include one or more of any type of network interface (e.g., a network interface card (NIC)) of wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC).
[0279] In one embodiment of the present specification, the above components of the computing device 600 and Figure 6 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 6 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art may add or replace other components as needed.
[0280] Computing device 600 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, content book computer, netbook, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 600 can also be a mobile or stationary server.
[0281] The processor 620 is configured to execute the following computer program / instructions, which, when executed by the processor, implement the steps of the above-mentioned recommendation model training method or object recommendation method.
[0282] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of this computing device is based on the same concept as the technical solutions of the recommendation model training method and the object recommendation method described above. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solutions of the recommendation model training method or the object recommendation method described above.
[0283] An embodiment of the present specification further provides a computer-readable storage medium storing a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned recommendation model training method or object recommendation method.
[0284] The above is a schematic diagram of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium is based on the same concept as the technical solutions of the recommendation model training method and the object recommendation method described above. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solutions of the recommendation model training method or the object recommendation method described above.
[0285] An embodiment of the present specification further provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned recommendation model training method or object recommendation method.
[0286] The above is a schematic diagram of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product shares the same concept as the technical solutions of the recommendation model training method and the object recommendation method described above. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solutions of the recommendation model training method or the object recommendation method described above.
[0287] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0288] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content of the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.
[0289] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.
[0290] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0291] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to specific embodiments. Obviously, many modifications and variations are possible based on the content of the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A training method for a recommendation model, characterized in that: include: Acquire user data of sample users for sample objects, wherein the sample users include reference sample users and target sample users; Based on the encoding features of the user data output by the initial recommendation model, the initial recommendation model is adversarially pre-trained using a domain classifier of user type to obtain a pre-trained recommendation model, wherein the domain classifier determines, based on the encoding features of the user data, whether the user data corresponds to a reference sample user or a target sample user; Dividing the user data into a plurality of meta-learning task sets for recommendation scenarios, and training the pre-trained recommendation model based on the meta-learning task sets to obtain a meta-parameter recommendation model; Based on the meta-learning task set of the target recommendation scenario, an adaptation task set of the target recommendation scenario is constructed, and based on the adaptation task set, the meta-parameter recommendation model is trained to obtain an adaptation recommendation model of the target recommendation scenario.
2. The method according to claim 1, characterized in that The obtaining of user data of the sample user for the sample object includes: Obtain reference user data of the reference sample user for the sample object, and target user data of the target sample user for the sample object; Sampling the reference user data to obtain sampled reference user data; Based on the sampled reference user data and the target user data, user data of the sample user for the sample object is constructed.
3. The method according to claim 1, characterized in that Before dividing the user data into a plurality of meta-learning task sets for recommendation scenarios, the method further includes: The user data is sampled according to a preset ratio to obtain the sampled user data, wherein the preset ratio is set according to the quantity distribution of the sample objects.
4. The method according to claim 1, wherein The user data includes user behavior data of the sample user with respect to the sample object, user information of the sample user, and object information of the sample object; The step of training the pre-trained recommendation model based on the meta-learning task set to obtain a meta-parameter recommendation model includes: Extracting a meta-learning task set for a current recommendation scenario from the meta-learning task sets for the multiple recommendation scenarios; Extracting user data of the current group from the meta-learning task set of the current recommendation scenario; Inputting user information in the current group of user data and object information in the current group of user data into the pre-trained recommendation model to obtain a predicted recommendation result; Calculating a recommendation loss value based on a difference between the predicted recommendation result and the user behavior data in the current group of user data; Based on the recommendation loss value, adjust the model parameters of the pre-trained recommendation model, return to execute the step of extracting the current group of user data from the meta-learning task set of the current recommendation scenario, until all groups of user data are extracted, return to execute the step of extracting the meta-learning task set of the current recommendation scenario from the meta-learning task set of the multiple recommendation scenarios, until all meta-learning task sets of the recommendation scenarios are extracted, and obtain a meta-parameter recommendation model.
5. The method according to claim 4, characterized in that The calculating the recommendation loss value based on the difference between the predicted recommendation result and the user behavior data in the current group of user data includes: Based on the difference between the predicted recommendation result and the user behavior data in the current group of user data, the recommendation loss value is calculated according to the preset weight, wherein the preset weight is set based on the similarity between the current recommendation scenario and the target sample user.
6. The method according to claim 1, characterized in that The user data includes user behavior data of the sample user with respect to the sample object, user information of the sample user, and object information of the sample object; The step of training the meta-parameter recommendation model based on the adaptation task set to obtain an adaptation recommendation model for the target recommendation scenario includes: Extracting user data of the current group from the adaptation task set of the target recommendation scenario; Inputting user information in the current group of user data and object information in the current group of user data into the pre-trained recommendation model to obtain a predicted recommendation result; Calculating a reward function value based on the similarity between the predicted recommendation result and the user behavior data in the current group of user data; Based on the reward function value, the model parameters of the meta-parameter recommendation model are adjusted, and the step of extracting the current group of user data from the adaptation task set of the target recommendation scenario is returned to be executed until the reward function value reaches a threshold, thereby obtaining the adaptation recommendation model of the target recommendation scenario.
7. The method according to any one of claims 1 to 6, characterized in that The method of performing adversarial pre-training on the initial recommendation model based on the encoded features of the user data output by the initial recommendation model by using a domain classifier of user type to obtain a pre-trained recommendation model includes: Inputting the user data of the sample user for the sample object into an initial recommendation model, and outputting the encoding features of the user data; Determining, by the domain classifier based on the encoding features of the user data, whether the user data corresponds to a reference sample user or a target sample user, and obtaining a predicted classification result; Calculating an adversarial loss value based on the predicted classification result; Based on the adversarial loss value, the parameters of the initial recommendation model are adjusted, and the step of inputting the user data of the sample user for the sample object into the initial recommendation model is returned to execute, and when the adversarial loss value reaches a preset threshold, a pre-trained recommendation model is obtained.
8. An object recommendation method, characterized in that: include: Obtaining user information of a target user, object information of multiple candidate recommendation objects, and an adaptation recommendation model for a target recommendation scenario, wherein the adaptation recommendation model for the target recommendation scenario is obtained by training using the recommendation model training method according to any one of claims 1 to 7; Inputting the user information and the object information of the multiple candidate recommendation objects into the adaptation recommendation model to obtain a target recommendation result; Based on the target recommendation result, determining a target recommendation object from the multiple candidate recommendation objects; Recommend the target recommendation object to the target user.
9. The method according to claim 8, characterized in that The determining a target recommendation object from the plurality of candidate recommendation objects based on the target recommendation result includes: Based on the target recommendation result, determining at least one target recommendation object from the multiple candidate recommendation objects, and sorting the at least one target recommendation object to obtain a recommendation object list; The recommending the target recommendation object to the target user includes: Based on the recommendation object list, the target recommendation object is recommended to the target user.
10. A recommendation system, characterized in that Includes data processing module, meta-learning module, adaptation module and recommendation generation module; The data processing module is configured to obtain user data of sample users for sample objects, wherein the sample users include reference sample users and target sample users; The meta-learning module is configured to perform adversarial pre-training on the initial recommendation model based on the encoding features of the user data output by the initial recommendation model through a domain classifier of user type to obtain a pre-trained recommendation model, wherein the domain classifier determines whether the user data corresponds to a reference sample user or a target sample user based on the encoding features of the user data; The data processing module is further configured to divide the user data into a plurality of meta-learning task sets for recommendation scenarios; The meta-learning module is further configured to train the pre-trained recommendation model based on the meta-learning task set to obtain a meta-parameter recommendation model; The adaptation module is configured to construct an adaptation task set for the target recommendation scenario based on the meta-learning task set of the target recommendation scenario, and train the meta-parameter recommendation model based on the adaptation task set to obtain an adaptation recommendation model for the target recommendation scenario; The recommendation generation module is configured to obtain user information of the target user, object information of multiple candidate recommendation objects, and an adaptation recommendation model of the target recommendation scenario, input the user information and the object information of the multiple candidate recommendation objects into the adaptation recommendation model, obtain a target recommendation result, determine a target recommendation object from the multiple candidate recommendation objects based on the target recommendation result, and recommend the target recommendation object to the target user.
11. A computing device, characterized in that include: memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 9 are implemented.
12. A computer-readable storage medium, characterized in that It stores a computer program / instruction, which implements the steps of the method according to any one of claims 1 to 9 when executed by a processor.
13. A computer program product, characterized in that The method comprises a computer program / instruction which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 9.
Citation Information
Cited By
Natural travel destination identification method and system based on outdoor trajectory data
CN120744424A
Meteorological decision intelligent processing method and device based on intelligent agent, and electronic equipment
CN121029830A
Recommendation model training method and device, recommendation method and device, equipment and distributed system
CN121188284A