Training Method, Device, Equipment and Medium for Content Recommendation Control and Recall Model
By downsampling the training sample data set and meta-learning optimization of the hyperparameter group, the problem of inaccurate hyperparameter configuration in the existing recommendation system is solved, and efficient automation and accuracy improvement of the recall model is achieved.
Patent Information
- Application Number
- CN202011477384.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-15
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2040-12-15
AI Technical Summary
The existing recommendation system relies on manual adjustment in hyperparameter configuration, resulting in low accuracy of recall models and consuming human resources, which cannot effectively improve the accuracy of multimedia content recommendations.
By downsampling the training sample dataset, the sub-sample dataset is obtained, and the meta-learning method is used to optimize the hyperparameter group, migrate to large-scale data, and select hyperparameters automatically to improve the accuracy of the recall model.
It improves the targeted recommendation capability of the recommendation system, improves the recall rate, and achieves the efficiency of hyperparameter selection and the accuracy of the recall model.
Smart Images

Figure CN114638372B_ABST
Abstract
Description
Background Art
[0002] Recommendation systems are widely used in e-commerce, search, advertising and other fields to recommend personalized items to users. For example, in the advertising scenario, a personalized recommendation system can push advertisements to users based on their characteristics and "preferences". If the user finally generates a click-through conversion behavior, it can be considered that the advertisement push is successful; otherwise, the push fails. All of these can be achieved based on the recall model in the recommendation system.
[0003] In actual application scenarios, the data volume of users and items is very large and has the characteristic of inconsistent distribution over time. It is necessary to manually adjust the hyperparameters of the recall model, such as the dimension of attribute features, the number of neurons in the neural network, the learning rate of model training, etc. This method mainly involves manual configuration by operators based on experience when training the model. However, using this method is likely to lead to inaccurate hyperparameter configuration and consume human resources, ultimately resulting in an inaccurate recall model of the recommendation system and a low accuracy of multimedia content recommendation. Summary of the Invention
[0004] The embodiments of the present application provide a method, device, equipment and medium for content recommendation control and training of a recall model, so as to improve the directional recommendation ability of the recommendation system and increase the recall rate.
[0005] A content recommendation control method provided by an embodiment of the present application includes:
[0006] Performing downsampling on a training sample data set to obtain a plurality of sub-sample data sets, where the training sample data set is used to train a first recall model of a target recommendation system;
[0007] Based on each sub-sample data set respectively, performing hyperparameter optimization on the first recall model to obtain an intermediate hyperparameter group corresponding to each sub-sample data set;
[0008] Respectively predicting the accuracy of each alternative hyperparameter group on the first recall model according to the accuracy of the intermediate hyperparameter group corresponding to each sub-sample data set on the first recall model, where each alternative hyperparameter group is initialized based on the training sample data set;
[0009] According to the accuracy of each alternative hyperparameter group on the first recall model, selecting a hyperparameter group from each alternative hyperparameter group to perform hyperparameter configuration and hyperparameter adjustment on the first recall model to obtain a target recall model;
[0010] Performing multimedia content recommendation to a target object based on the target recall model.
[0011] A method for training a recall model provided by an embodiment of the present application includes:
[0012] Downsample the training sample dataset to obtain multiple sub - sample datasets, where the training sample dataset is used to train the first recall model of the target recommendation system. Among them, the training samples in the training sample dataset include sample object features, sample multimedia content features, and an operation label of the sample object and the sample multimedia content. The operation label indicates whether the sample object has a target operation behavior on the sample multimedia content when the sample multimedia content is recommended to the sample object;
[0013] Based on each sub - sample dataset respectively, perform hyperparameter optimization on the first recall model to obtain intermediate hyperparameter groups corresponding to each sub - sample dataset;
[0014] Based on the accuracy of each intermediate hyperparameter group corresponding to each sub - sample dataset on the first recall model, predict the accuracy of each alternative hyperparameter group on the first recall model. Each alternative hyperparameter group is initialized based on the training sample dataset;
[0015] According to the accuracy of each alternative hyperparameter group on the first recall model, select a hyperparameter group from each alternative hyperparameter group to perform hyperparameter configuration and hyperparameter adjustment on the first recall model to obtain a target recall model, where the target recall model is used to recommend multimedia content to a target object.
[0016] A content recommendation control device provided by an embodiment of the present application includes:
[0017] A first data acquisition unit, configured to downsample the training sample dataset to obtain multiple sub - sample datasets, where the training sample dataset is used to train the first recall model of the target recommendation system;
[0018] A first hyperparameter optimization unit, configured to perform hyperparameter optimization on the first recall model based on each sub - sample dataset respectively to obtain intermediate hyperparameter groups corresponding to each sub - sample dataset;
[0019] A first prediction unit, configured to predict the accuracy of each alternative hyperparameter group on the first recall model based on the accuracy of each intermediate hyperparameter group corresponding to each sub - sample dataset on the first recall model. Each alternative hyperparameter group is initialized based on the training sample dataset;
[0020] A first configuration unit, configured to select a hyperparameter group from each alternative hyperparameter group according to the accuracy of each alternative hyperparameter group on the first recall model to perform hyperparameter configuration and hyperparameter adjustment on the first recall model to obtain a target recall model;
[0021] A recommendation unit for recommending multimedia content to a target object based on the target recall model.
[0022] Optionally, the first data acquisition unit is specifically configured to:
[0023] Screen training samples that meet the following anti-interference conditions from the training sample dataset as training samples in the sub-sample dataset: the size of the adversarial perturbation added to the training sample is less than a preset threshold, and when reclassifying the training sample with the added adversarial perturbation based on the first recall model, the training sample with an incorrect classification result; or
[0024] Screen training samples that meet the following quantity conditions from the training sample dataset as training samples in the sub-sample dataset: add a certain amount of adversarial perturbation to each training sample, and when reclassifying the training sample with the added adversarial perturbation based on the first recall model, the number of training samples with an incorrect classification result reaches the sum of the training samples in each sub-sample dataset.
[0025] Optionally, the first prediction unit is specifically configured to:
[0026] Extract the meta-features of each sub-sample dataset according to the training samples in each sub-sample dataset respectively;
[0027] Input the meta-features corresponding to each sub-sample dataset and the intermediate hyperparameter group into the first accuracy prediction model respectively, and perform multiple rounds of iterative training on the first accuracy prediction model based on the difference between the first predicted accuracy output by the first accuracy prediction model and the corresponding true accuracy, so as to obtain a second accuracy prediction model, where the true accuracy refers to the accuracy of the intermediate hyperparameter group on the first recall model;
[0028] Input the meta-features corresponding to the training sample dataset and the alternative hyperparameter groups into the second accuracy prediction model respectively, and predict the second predicted accuracy corresponding to each alternative hyperparameter group, where the second predicted accuracy refers to the accuracy of the predicted alternative hyperparameter group on the first recall model.
[0029] Optionally, each sub-sample dataset includes two major categories: a training subset and a validation subset; the first hyperparameter optimization unit is specifically configured to:
[0030] For any sub-sample dataset, initialize the hyperparameters corresponding to the first recall model;
[0031] Configure the first recall model based on the obtained hyperparameters, and train the first recall model based on the training subset in the sub-sample dataset to obtain a second recall model;
[0032] Input the validation subset in the sub-sample dataset into the second recall model, obtain the accuracy corresponding to the current hyperparameters based on the output result of the second recall model, and use the current hyperparameters as a set of intermediate hyperparameter groups corresponding to the sub-sample dataset. Among them, the accuracy obtained based on the output result of the second recall model is the true accuracy corresponding to the intermediate hyperparameter group;
[0033] Adjust the current hyperparameters through Bayesian optimization to obtain a new set of hyperparameters, return to configure the first recall model based on the obtained hyperparameters, and train the first recall model based on the training subset in the sub-sample dataset to obtain the second recall model;
[0034] Optionally, the training samples in the training sample dataset include sample object features, sample multimedia content features, and operation labels of the sample object and the sample multimedia content. The operation label indicates whether the sample object has a target operation behavior on the sample multimedia content when the sample multimedia content is recommended to the sample object;
[0035] The first prediction unit is specifically used for:
[0036] For any sub-sample dataset, input each training sample in the sub-sample dataset into the second recall model, and based on the second recall model, extract the expected score corresponding to each training sample. The expected score indicates the score of the sample object having a target operation behavior on the sample multimedia content when the sample multimedia content is recommended to the sample object;
[0037] Obtain the feature mean of the expected scores corresponding to each training sample, and determine the distribution feature of the operation label according to the sample object features, sample multimedia content features, and operation labels included in each training sample;
[0038] Combine the feature mean and the distribution feature as the meta-feature of the sub-sample dataset.
[0039] Optionally, the first prediction unit is specifically used to perform the following process in each round of iterative training:
[0040] Select a certain number of target sub-sample datasets from the unselected sub-sample datasets;
[0041] Input the meta-features and intermediate hyperparameter groups corresponding to each target sub-sample dataset into the first accuracy prediction model respectively, and predict the first prediction accuracy corresponding to each target sub-sample dataset;
[0042] Adjust the network parameters of the first accuracy prediction model according to the difference between the first predicted accuracy and the corresponding true accuracy.
[0043] Optionally, the first predicted accuracy includes a meta-training predicted accuracy and a meta-validation predicted accuracy; specifically, the first prediction unit is configured to:
[0044] Use a part of the target sub-sample data set selected from the target sub-sample data set as the meta-training set, and use another part of the target sub-sample data set as the meta-validation set;
[0045] Input the meta-features and intermediate hyperparameter groups corresponding to each meta-training set into the first accuracy prediction model respectively, and predict the meta-training predicted accuracy corresponding to each meta-training set;
[0046] Update the network parameters of the first accuracy prediction model based on the difference between the meta-training predicted accuracy and the corresponding true accuracy;
[0047] Input the meta-features and hyperparameter groups corresponding to each meta-validation set into the first accuracy prediction model after parameter update respectively, and predict the meta-learning predicted accuracy corresponding to each meta-validation set;
[0048] Update the network parameters of the first accuracy prediction model again based on the difference between the meta-learning predicted accuracy and the corresponding true accuracy.
[0049] Optionally, the meta-feature of the training sample data set is obtained by accumulating the meta-features of each sub-sample data set.
[0050] A training device for a recall model provided by an embodiment of the present application includes:
[0051] A second data acquisition unit, configured to downsample a training sample data set to obtain a plurality of sub-sample data sets, where the training sample data set is used to train a first recall model of a target recommendation system. The training samples in the training sample data set include sample object features, sample multimedia content features, and an operation label of the sample object and the sample multimedia content, where the operation label indicates whether the sample object has a target operation behavior on the sample multimedia content when the sample multimedia content is recommended to the sample object;
[0052] A second hyperparameter optimization unit, configured to optimize the hyperparameters of the first recall model based on each sub-sample data set respectively, and obtain intermediate hyperparameter groups corresponding to each sub-sample data set respectively;
[0053] A second prediction unit, configured to predict the accuracy of each alternative hyperparameter group on the first recall model according to the accuracy of the corresponding intermediate hyperparameter group of each sub-sample data set on the first recall model, where each alternative hyperparameter group is obtained by initializing based on the training sample data set;
[0054] A second configuration unit, configured to select a hyperparameter group from each alternative hyperparameter group according to the accuracy of each alternative hyperparameter group on the first recall model, configure the hyperparameters of the first recall model and perform hyperparameter adjustment to obtain a target recall model, where the target recall model is used to recommend multimedia content to a target object.
[0055] An electronic device provided in an embodiment of the present application includes a processor and a memory. Among them, the memory stores program code, and when the program code is executed by the processor, the processor is caused to execute the steps of any of the above content recommendation control methods or the steps of any of the above recall model training methods.
[0056] An embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of any of the above content recommendation control methods or the steps of any of the above recall model training methods.
[0057] An embodiment of the present application provides a computer-readable storage medium, which includes program code. When the program product runs on an electronic device, the program code is used to cause the electronic device to execute the steps of any of the above content recommendation control methods or the steps of any of the above recall model training methods.
[0058] The beneficial effects of the present application are as follows:
[0059] An embodiment of the present application provides a method, device, equipment, and medium for content recommendation control and recall model training. Since in the embodiment of the present application, a meta-learning method is used to transfer the hyperparameter group on the sub-sample data set to large-scale data, the efficiency of hyperparameter selection in the recommendation system is improved. Compared with manual hyperparameter selection, a large improvement in recommendation accuracy can be achieved. On this basis, a set of hyperparameters applied to the training sample data set is optimized based on the above method, and then the hyperparameters of the first recall model are configured and hyperparameter adjustment is performed using the hyperparameter group to obtain a target recall model; in this case, the obtained target recall model is more accurate, so when using the target recall model to recommend multimedia content, the directional recommendation ability of the recommendation system can be effectively improved, and the recall rate can be increased.
[0060] Other features and advantages of the present application will be described in the following specification, and in part will be obvious from the specification, or can be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the written specification, claims, and drawings. Description of the Drawings
[0061] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The schematic embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0062] Figure 1 is an optional schematic diagram of an application scenario in an embodiment of the present application;
[0063] Figure 2 is a schematic flowchart of a content recommendation control method in an embodiment of the present application;
[0064] Figure 3 is a schematic structural diagram of a recall model of an object recommendation system in an embodiment of the present application;
[0065] Figure 4 is a schematic diagram of an adversarial sampler in an embodiment of the present application;
[0066] Figure 5 is a schematic diagram of hyperparameter optimization in an embodiment of the present application;
[0067] Figure 6 is a schematic diagram of an automatic hyperparameter optimization system in an embodiment of the present application;
[0068] Figure 7 is a schematic diagram of a model performance value predictor in an embodiment of the present application;
[0069] Figure 8 is a schematic flowchart of a training method of a recall model in an embodiment of the present application;
[0070] Figure 9 is a schematic flowchart of a content recommendation method in an embodiment of the present application;
[0071] Figure 10 is a complete flowchart of a hyperparameter optimization method in an embodiment of the present application;
[0072] Figure 11 is a schematic structural diagram of the composition of a content recommendation control device in an embodiment of the present application;
[0073] Figure 12 is a schematic structural diagram of the composition of a training device of a recall model in an embodiment of the present application;
[0074] Figure 13 It is a schematic diagram of a hardware composition structure of an electronic device applying an embodiment of the present application. Detailed implementation manners
[0075] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the technical solutions of the present application, rather than all of them. Based on the embodiments described in this application document, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the technical solutions of the present application.
[0076] The following introduces some concepts involved in the embodiments of the present application.
[0077] Recommendation system: Personalized recommendation recommends information and products that a user is interested in according to the user's interest characteristics and purchase behavior. With the continuous expansion of the scale of e-commerce, the number and variety of products have increased rapidly, and customers need to spend a lot of time to find the products they want to buy. This process of browsing a large amount of irrelevant information and products will undoubtedly cause continuous loss of consumers overwhelmed by the problem of information overload. To solve these problems, the personalized recommendation system came into being. The personalized recommendation system is an advanced business intelligence platform based on massive data mining to help e-commerce websites provide completely personalized decision-making support and information services for their customers' shopping. The target recommendation system in the embodiments of the present application refers to a recommendation system for multimedia content, which is used to recommend multimedia content, such as advertisements, videos, etc., to users.
[0078] Recall Model and Hyperparameter Sets: The recall model refers to the machine learning model in the target recommendation system, which is used to screen out a part of multimedia content that the target object prefers more from a large amount of multimedia content and recommend it to the user. In a machine learning model, hyperparameters are parameters whose values are set before the start of the learning process, rather than parameters obtained through training, such as the model learning rate, etc. Usually, it is necessary to optimize the hyperparameters and set a set of optimal hyperparameters for the machine learning model to improve the learning performance and effect. The intermediate hyperparameter set and the alternative hyperparameter set in the embodiments of this application are both for the hyperparameters of the recall model. Among them, the intermediate hyperparameter set is for the subsample dataset, while the alternative hyperparameter set is for the training sample dataset. In addition, the first recall model, the second recall model, and the target recall model in the embodiments of this application all belong to the recall models of the target recommendation system. Among them, the second recall model is further trained based on the first recall model, and the target recall model is configured with the hyperparameter set obtained based on the hyperparameter optimization method in the embodiments of this application.
[0079] Adversarial Attack: Methods that use the drawbacks of deep learning to disrupt the recognition system can generally be referred to as adversarial attacks, that is, making special changes to the recognition object, which are not visible to the naked eye of a person, but will cause the recognition model to malfunction. Since the input form of machine learning algorithms is a numerical vector, the attacker will design a targeted numerical vector, that is, adversarial data, so that the machine learning model makes misjudgments. Different from other attacks, adversarial attacks mainly occur when constructing adversarial data, and then the adversarial data is input into the machine learning model like normal data and a deceived recognition result is obtained. In the embodiments of this application, mainly based on adversarial attacks, the training sample dataset is downsampled to obtain multiple subsample datasets.
[0080] Meta-Learning: The mapping relationship between the state characteristics and quality parameters of the neural network at each stage of the machine learning framework can be mined through supervised learning, and the performance of the neural network can be optimized according to the characteristics of the new learning task. The core idea of meta-learning is to learn the initial parameters of the neural network from a large number of training tasks, and these initial parameters can enable the new machine learning task to quickly converge to a better solution under the condition of small samples. In the embodiments of this application, by applying meta-learning-based hyperparameter optimization to the deep learning recall model under large-scale data, the targeted recommendation ability of the advertising system is improved, and the recall rate is increased.
[0081] The embodiments of this application relate to artificial intelligence (AI) and machine learning technologies, and are designed based on computer vision technology and machine learning (ML) in artificial intelligence.
[0082] Artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems for perceiving the environment, acquiring knowledge, and using knowledge to achieve the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence.
[0083] Artificial intelligence also involves researching the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technologies mainly include several major directions such as computer vision technology, natural language processing technology, and machine learning / deep learning. With the research and progress of artificial intelligence technology, artificial intelligence has been studied and applied in multiple fields, such as common smart homes, intelligent customer service, virtual assistants, smart speakers, intelligent marketing, driverless, autonomous driving, robots, intelligent healthcare, etc. It is believed that with the development of technology, artificial intelligence will be applied in more fields and play an increasingly important role.
[0084] Machine learning is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance.
[0085] Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, and inductive learning. In the embodiments of this application, when performing multimedia content recommendation, a recall model based on machine learning or deep learning is adopted, and multimedia content is recommended to the target object based on this model. Based on the method in the embodiments of this application, hyperparameter automatic optimization can be achieved, the efficiency of hyperparameter selection in the recommendation system is improved, and the recommendation accuracy of the recommendation system is greatly enhanced.
[0086] The method for training a recall model proposed in the embodiments of the present application can be divided into two parts, including a training part and an application part; among them, the training part involves the technical field of machine learning. In the training part, an accuracy prediction model is trained through machine learning. Specifically, the meta-features and intermediate hyperparameter groups corresponding to each subsample dataset obtained by downsampling the training sample dataset are used as the input parameters of the first accuracy prediction model, and the true accuracy corresponding to each intermediate hyperparameter group is used as a label to train the first accuracy prediction model, and a trained second accuracy prediction model is obtained. Then, based on the trained second accuracy prediction model, the accuracy of each candidate hyperparameter group corresponding to the training sample dataset on the target recommendation system recall model is predicted. Finally, a group of hyperparameter groups is selected based on the predicted accuracy, and the first recall model is configured with hyperparameters and hyperparameter adjustment is performed based on this group of hyperparameter groups to obtain the target recall model. The application part is used to use the target recall model trained in the training part to predict the preferences of the target object for each multimedia content to be recommended for multimedia content recommendation, such as advertising recommendation, etc.
[0087] The following briefly introduces the design concept of the embodiments of the present application:
[0088] Recommendation systems are widely used in fields such as e-commerce, search, and advertising. Through the attributes and historical records of users, as well as the attribute characteristics of items, personalized items are recommended to users. For example, in the advertising scenario, a personalized recommendation system can push advertisements to users based on the characteristics and "preferences" of the users. If the user finally generates a click-through conversion behavior, it can be considered that the advertisement push is successful, otherwise the push fails. In the training process of the recommendation system, usually, whether the user purchases an item or clicks on an advertisement is used as a binary classification prediction task, and a prediction model is constructed to complete this binary classification prediction task. Using a large amount of user and item data in the industrial environment, the recall model of the recommendation system can be trained to predict the behavior of users such as purchasing or clicking on corresponding items.
[0089] When optimizing the hyperparameters of the recall model, in addition to the method of manual configuration by operators according to experience listed in the background art, in the related art, there are already methods aimed at optimizing hyperparameters on the entire dataset. When the data scale is relatively large, it takes a long time and the efficiency is very low; in addition, it is not possible to achieve rapid optimization of hyperparameters under large-scale data, ultimately resulting in a low accuracy of the recommendation system.
[0090] In view of this, embodiments of the present application propose a method, apparatus, device, and medium for training a content recommendation control and recall model. In the embodiments of the present application, a meta-learning method is used to transfer hyperparameter groups on a subsample dataset to large-scale data, improving the efficiency of hyperparameter selection in the recommendation system. Compared with manually selecting hyperparameters, a significant improvement in recommendation accuracy can be achieved. On this basis, a set of hyperparameters applied to the training sample dataset is optimized based on the above method, and then the hyperparameter group is used to configure and adjust the hyperparameters of the first recall model to obtain a target recall model. In this case, the obtained target recall model is more accurate. Therefore, when using the target recall model for multimedia content recommendation, the directional recommendation ability of the recommendation system can be effectively improved, and the recall rate can be increased.
[0091] The preferred embodiments of the present application will be described below with reference to the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only for the purpose of illustrating and explaining the present application, and are not used to limit the present application. And without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.
[0092] As Figure 1 shown, it is a schematic diagram of the application scenario of the embodiments of the present application. The application scenario diagram includes two terminal devices 110 and a server 120. The terminal devices 110 and the server 120 can communicate through a communication network.
[0093] In an alternative embodiment, the communication network is a wired network or a wireless network. The terminal devices 110 and the server 120 can be directly or indirectly connected through wired or wireless communication methods, and the present application does not limit this here.
[0094] In the embodiments of the present application, the terminal device 110 is an electronic device used by a user. The electronic device can be a personal computer, mobile phone, tablet computer, notebook, e-reader, smart home, etc., which are computer devices with certain computing capabilities and running instant messaging software and websites or social software and websites. Each terminal device 110 communicates with the server 120 through a wireless network. The server 120 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0095] Among them, browsers or social applications can be installed in each terminal device. For example, when users use short-video social applications, advertisements are often pushed, which can be implemented based on the recall model in the recommendation system. The applications involved in the embodiments of the present application can be software, or web pages, applets, etc. The server is the application server corresponding to the software, web page, applet, etc., and the specific type of the client is not limited.
[0096] Among them, the recall model or accuracy prediction model can be deployed on the server 120 for training. A large number of training samples obtained from the target recommendation system can be stored in the server 120, including sample object features, sample multimedia content features, and operation labels of the sample object and the sample multimedia content, etc., for training the recall model or accuracy prediction model. Optionally, after the recall model or accuracy prediction model is trained based on the training method in the embodiments of the present application, the trained recall model or accuracy prediction model can be directly deployed on the terminal device 110 or on the server 120. In the embodiments of the present application, the recall model is used to recommend multimedia content to users, while the accuracy prediction model is used to predict the optimal hyperparameters on the full dataset.
[0097] In the embodiments of the present application, when the recall model is deployed on the terminal device 110, the multimedia content set to be recommended can be obtained from the target recommendation system, and the multimedia content features of each multimedia content to be recommended in the multimedia content set and the object features of the target object can be extracted. Based on these features, the preference degree of the target object for each multimedia content can be predicted, that is, the expected score of the target object for performing the target operation behavior on each multimedia content to be recommended. Then, at least one target multimedia content to be recommended can be screened out and recommended to the target object. When the recall model or accuracy prediction model is deployed on the server 120, the terminal device 110 can obtain the multimedia content set to be recommended and upload it to the server. The server extracts the multimedia content features of each multimedia content to be recommended in the multimedia content set and the object features of the target object, and based on these features, the preference degree of the target object for each multimedia content can be predicted, that is, the expected score of the target object for performing the target operation behavior on each multimedia content to be recommended. Then, the server 120 can return the predicted expected scores of each multimedia content to the terminal device 110, and the terminal device 110 can screen out at least one target multimedia content to be recommended and recommend it to the target object, or the server 120 can directly return at least one target multimedia content to be recommended screened out to the terminal device 110, etc. However, generally, the recall model or accuracy prediction model is directly deployed on the server 120, and no specific limitation is made here.
[0098] It should be noted that the training samples used in different scenarios are different. Taking the scenario of video recommendation as an example, the multimedia content is a video, and the sample multimedia content features in the training samples can refer to features such as the title and content of the video; in the scenario of advertising recommendation, the multimedia content is an advertisement, and the sample multimedia content features in the training samples mainly refer to features such as advertising slogans and advertising content.
[0099] Taking advertising recommendation as an example, in the embodiments of the present application, advertising triggering is an important link in the online advertising placement system, and its function is advertising recall: that is, based on the user's context scenario, retrieve a candidate advertisement set from the advertisement library, and then provide it for subsequent modules to calculate and optimize the exposure advertisement set. The main triggering strategies of the current advertising placement system mainly include the following two:
[0100] First, label triggering: population targeting based on a label system. Advertisers select the population attribute labels mined by the system to determine the target population. In contrast, the advertising placement system recalls the advertisements of all advertisers whose purchased labels are hit through the label set carried by the current user as the candidate advertisement set.
[0101] Second, intelligent triggering: With the help of the first-party or platform historical placement second-party effect data accumulated by advertisers, finely calculate the matching degree between <scenario, user> and the advertising target of the advertiser, and then determine the target population for placement. It gets rid of the manual selection, experiments and comparison of advertisers, and the labor cost of label and label combination. From the perspective of the advertising placement system, represent the current user as a high-dimensional vector, and recall a set of advertisements with higher similarity in the advertisement library as advertisement candidates through vector retrieval.
[0102] Among them, the background model of intelligent directional triggering is a deep learning model. However, in deep learning, the setting of parameters often has a certain impact on the final effect. For example, the size of the representation vector of the input features, the number of network layers, the number of nodes in each layer, the selection of activation functions, the use of sum pooling and average pooling for multi-valued input features, and the network structure design, etc. If the artificial trial-and-error method given in the related technology is used to find the optimal hyperparameters, this tuning method is very test the experience of the tuner, and it is time-consuming and laborious. For example, the setting of the size of different feature representation vectors on the user side is generally set in a human heuristic way. For features with smaller values such as age, use a smaller embedding dimension, and for features with larger values such as user interests, use a larger embedding dimension. In the embodiments of the present application, by studying the application of meta-learning-based hyperparameter optimization in the deep learning recall model under large-scale data, the directional recommendation ability of the advertising system can be effectively improved, and the recall rate can be increased.
[0103] In a possible application scenario, the training samples in this application can be stored using cloud storage technology. Cloud storage is a new concept extended and developed from the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as the storage system) refers to a storage system that combines a large number of different types of storage devices (storage devices are also called storage nodes) in the network through functions such as cluster applications, grid technology, and distributed file systems, and works together through application software or application interfaces to jointly provide data storage and service access functions to the outside world.
[0104] In a possible application scenario, to facilitate reducing communication latency, servers 120 can be deployed in each region, or for load balancing, different servers 120 can respectively serve the regions corresponding to each terminal device 110. Multiple servers 120 can achieve data sharing through blockchain. Multiple servers 120 are equivalent to a data sharing system composed of multiple servers 120. For example, the terminal device 110 is located at location a and is communicatively connected to server 120, and the terminal device 110 is located at location b and is communicatively connected to another server 120. Among them, blockchain is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity (anti-counterfeiting) of the information and generate the next block. Blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer.
[0105] For each server 120 in the data sharing system, it has a node identifier corresponding to that server 120. Each server 120 in the data sharing system can store the node identifiers of other servers 120 in the data sharing system, so as to broadcast the generated block to other servers 120 in the data sharing system according to the node identifiers of other servers 120 later. A node identifier list as shown in the following table can be maintained in each server 120, and the server 120 name and node identifier are correspondingly stored in this node identifier list. Among them, the node identifier can be an IP (Internet Protocol) address and any other information that can be used to identify the node. Only the IP address is used as an example in Table 1 for illustration.
[0106] Table 1
[0107] Server Name Node Identifier Node 1 119.115.151.174 Node 2 118.116.189.145 … … Node N 119.124.789.258
[0108] The following mainly takes the scenario of advertising recommendation as an example to introduce in detail the content recommendation control method and the training method of the recall model in the embodiments of the present application. It should be noted that the above application scenarios are only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not limited in this regard.
[0109] Refer to Figure 2 As shown, it is a flowchart of the implementation of a content recommendation control method provided by an embodiment of the present application. The specific implementation process of this method is as follows:
[0110] S21: Downsample the training sample data set to obtain multiple sub-sample data sets. The training sample data set is used to train the first recall model of the target recommendation system;
[0111] Optionally, the training samples in each sub-sample data set are training samples with anti-interference ability lower than the target condition sampled from the training sample data set based on adversarial perturbations. Among them, there can be many kinds of target conditions. In the embodiments of the present application, the following is taken as an example for illustration:
[0112] Target condition 1: Judging according to the magnitude of the adversarial perturbation, screening the training samples that meet the following anti-interference conditions from the training sample data set as the training samples in the sub-sample data set: the magnitude of the adversarial perturbation added to the training sample is less than the preset threshold, and when reclassifying the training sample with the added adversarial perturbation based on the first recall model, the training sample with an incorrect classification result.
[0113] For example, set an upper and lower limit of the adversarial perturbation, add adversarial perturbations of different magnitudes within this interval to the training samples, select the training samples with incorrect classification from them, and then sort them according to the magnitude of the added adversarial perturbation, and screen out the training samples with smaller adversarial perturbations, such as less than the preset threshold, and add them to the sub-sample data set.
[0114] For example, first add a large adversarial perturbation to each training sample in the training sample data set to obtain the adversarial samples corresponding to each training sample; respectively obtain the classification results of each adversarial sample based on the first recall model, and use the adversarial samples with classification results inconsistent with the labels of the training samples as the target adversarial samples; then continuously reduce the magnitude of the adversarial perturbation corresponding to each target adversarial sample, and determine the distance between each target adversarial sample and the target classification surface; the target classification surface in the embodiments of the present application mainly represents the category of user clicking on an advertisement. Then, according to the distance between each target adversarial sample and the target classification surface, screen the training samples in the training sample data set to obtain multiple sub-sample data sets, where each sub-sample data set contains at least one training sample corresponding to a target adversarial sample.
[0115] In the above embodiments, the smaller the perturbation noise of the misclassified samples, the closer the samples are to the target classification surface. Therefore, based on this, training samples with relatively low anti-interference ability can be screened out, and an accurate and diverse sub-dataset can be collected.
[0116] Target condition two: Determine according to the number of training samples in the sub-sample dataset. From the training sample dataset, screen out training samples that meet the following quantity conditions as the training samples in the sub-sample dataset: When a certain adversarial perturbation is added to each training sample and the first recall model is used to re-classify the training samples after adding the adversarial perturbation, the number of training samples with incorrect classification results reaches the sum of the training samples in each sub-sample dataset.
[0117] For example, when the number of sub-sample datasets is fixed and the number of training samples in the sub-sample datasets can also be fixed. First, set an upper limit for the adversarial perturbation. At the beginning, add a relatively large adversarial perturbation to each training sample, select the training samples with incorrect classification among them, and add these samples to the sub-sample dataset; then continuously reduce the magnitude of the adversarial perturbation added to the training samples, and further screen out the training samples with incorrect classification and add them to the sub-sample dataset until a certain number of sub-sample datasets are obtained, so as to collect an accurate and diverse sub-dataset.
[0118] In addition, these two target conditions can also be combined, that is, given the size and number of sub-sample datasets, as well as the upper and lower limits of the adversarial perturbation, an accurate and diverse sub-dataset can be collected from the entire data sample set through the technology of adversarial perturbation, etc., which will not be specifically limited here.
[0119] S22: Based on each sub-sample dataset respectively, perform hyperparameter optimization on the recall model of the target recommendation system to obtain the intermediate hyperparameter groups corresponding to each sub-sample dataset;
[0120] S23: Predict the accuracy of each alternative hyperparameter group on the first recall model respectively according to the accuracy of the intermediate hyperparameter group corresponding to each sub-sample dataset on the recall model. Each alternative hyperparameter group is initialized based on the training sample dataset.
[0121] In the embodiments of the present application, each training sample in the training sample dataset (also called the full dataset) includes sample object features, sample multimedia content features, and an operation label for the sample object and the sample multimedia content. The operation label indicates whether the sample object has a target operation behavior on the sample multimedia content when the sample multimedia content is recommended to the sample object. Among them, the target operation behavior can be operations such as clicking, browsing, and liking.
[0122] In the application scenario of advertising recommendation, taking the target operation behavior as the click behavior as an example, the training samples include the features of users, the features of advertisements, and the label indicating whether the user clicks on the advertisement. For example, clicking on the advertisement is 1, and otherwise it is 0. The data structure of each training sample is as follows:
[0123]
[0124] Based on the above training sample data, the recall model in the target recommendation system can be trained, and the hyperparameters of the recall model in the target recommendation system can be optimized to obtain the intermediate hyperparameter groups corresponding to each sub-sample data set. Among them, the input data of the recall model: user ID u, the set of advertisements S clicked by the user u , and the output results include: user feature matrix P usercount*k , advertisement feature matrix Q itemcount*k (usercount and itemcount represent the total number of users and the total number of advertisements respectively), the expected score of the recall model for a certain user-advertisement pair (u, i) Here The larger it is, the greater the probability that the user clicks on the advertisement.
[0125] Refer to Figure 3 shown, which is a schematic structural diagram of a recall model in an embodiment of the present application. This model is a two-tower model. As Figure 3 shown, the model is divided into two main parts. One part is used to extract features from the advertisement data to obtain the advertisement feature matrix Q, and the other part is used to extract features from the user data to obtain the user feature matrix P. By learning the user features and advertisement features through a machine learning model, from calculate the degree of preference of this user for all advertisements to determine whether the user will click on the advertisement.
[0126] Among them, the accuracy is also called the model performance value. In the embodiments of the present application, the model performance value can refer to the recall rate y, or other evaluation indicators, such as the F value, ROC (receiver operating characteristic curve), etc. In the following, the recall rate is mainly used as an example for illustration.
[0127] Among them, the accuracy of the intermediate hyperparameter group corresponding to each sub-sample data set on the recall model indicates that when the recall model is configured based on different intermediate hyperparameter groups, the training samples in the sub-sample data set are input into the recall model. After training the recall model, the model performance value y of the recall model obtained based on the validation set. Furthermore, based on the accuracy of the intermediate hyperparameter group corresponding to each sub-sample data set on the recall model, the accuracy of each alternative hyperparameter group on the first recall model is predicted.
[0128] In the embodiments of the present application, the intermediate hyperparameters corresponding to each sub-sample data set are obtained by optimizing the hyperparameters based on the sub-sample data set, while each alternative hyperparameter group corresponding to the training sample data set is obtained by random initialization.
[0129] S24: According to the accuracy of each alternative hyperparameter group on the first recall model, select a hyperparameter group from each alternative hyperparameter group to configure the hyperparameters of the recall model and perform hyperparameter adjustment to obtain a target recall model;
[0130] S25: Based on the target recall model, perform multimedia content recommendation to the target object.
[0131] In the embodiments of the present application, in the field of recommendation systems, for the problem of hyperparameter optimization of models in recommendation systems, large-scale data is sampled based on adversarial attacks, effective sub-sample data sets are selected, and the hyperparameter groups on the sub-sample data sets are migrated to large-scale data using the method of meta-learning, which improves the efficiency of hyperparameter selection in recommendation systems. Compared with manual hyperparameter selection, it can achieve a large improvement in recommendation accuracy.
[0132] Optionally, when predicting the accuracy of each alternative hyperparameter group on the first recall model according to the accuracy of the intermediate hyperparameter group corresponding to each sub-sample data set, it can be implemented based on a deep learning model, which is called a hyperparameter learner (also called a model performance value predictor or an accuracy prediction model) in the embodiments of the present application.
[0133] Generally speaking, the hyperparameter optimization system based on meta-learning under large-scale data involved in the embodiments of the present application can be divided into two parts: an adversarial sampler and a hyperparameter learner. The adversarial sampler is responsible for collecting accurate and diverse sub-sample data sets from the full data sample set through the technology of adversarial perturbation; the hyperparameter learner receives the sub-sample data sets collected by the adversarial sampler, performs hyperparameter optimization on each sub-sample data set, combines the meta-features of the sub-sample data sets, and trains the hyperparameter learner. Furthermore, in the test stage, the hyperparameter learner can directly predict the hyperparameter group that performs well on the full data set. Based on this hyperparameter group, the hyperparameters of the recall model are configured and hyperparameter adjustment is performed to obtain the target recall model.
[0134] The following provides a detailed introduction to the adversarial sampler and hyperparameter learner given in the embodiments of the present application respectively:
[0135] Referring to Figure 4 As shown, it is a schematic diagram of an adversarial sampler in the embodiments of the present application. Generally speaking, the adversarial sampler continuously adds perturbation noise to sample points and observes whether the samples are misclassified. The smaller the perturbation noise of the misclassified samples, the closer the sample is to the classification surface, and in the embodiments of the present application, such samples are more inclined to be sampled. Specifically, fixing the size of the sampled sub-sample dataset, for the adversarial attack sampling strategy, the optimization goal is to find the minimum perturbation for each sample individual such that the model makes a wrong judgment on the sample after adding the perturbation, and the adversarial sample is still within a reasonable range.
[0136] Based on the above basic idea, the following optimization equation is defined in the embodiments of the present application:
[0137]
[0138] Among them, x represents the sample individual, ∈ represents the perturbation, λ1 and λ2 are hyperparameters, the L2 distance is used as a measure of the perturbation, W represents the weight value of the features of each sample learned by the adversarial sampler, and in order to ensure that ∈ is within a reasonable range, the tanh function is used to limit the perturbation to a reasonable interval, and w is used instead of ∈ for optimization, so that x + ∈ ∈ [min(x), max(x)] d , where d represents the dimension of the sample features. At the same time, g is defined as a constraint function for attacking the sample. By constraining g, it is ensured that adding perturbation can misclassify the sample, so as to select more sensitive sample points for sampling. For the classification task, g means that the sample x is misclassified after adding the perturbation, and Z(x) i represents the probability that x is classified into the i-th class, t represents the target attack class, and generally the second largest Z(x) is specified i as the attack target class, and c is a constant representing the control of the attack intensity; for the regression task, g represents the distance between the predicted regression value of the sample and the true value, where z(x) represents the predicted regression value for the sample x, and by setting a positive threshold σ to measure whether the change in the predicted value before and after the attack exceeds the specified range.
[0139] In the above embodiment, the idea of adversarial attack is introduced into the sampling strategy, and a carefully designed attack on the network is carried out using the gradient descent method, so as to find the sample points around the decision boundary of the network. On the one hand, these samples play a more important role in the decision-making of the model, and on the other hand, they are also the sample points where the model is weakest in its judgment. Compared with the sample points far from the decision boundary, only a smaller perturbation needs to be added to make the sample cross the decision boundary and produce a misclassification result.
[0140] After downsampling the training sample dataset based on the above method to obtain multiple sub-sample datasets, further hyperparameter optimization is performed on the sub-sample datasets. Additionally, after sampling, it is necessary to calculate the meta-features of different sub-sample datasets, such as the feature distributions of users and advertisements, the historical behaviors of users, etc., and design a meta-learning module to mine the relationship between the meta-features of the sub-sample datasets and the hyperparameter groups, so as to realize the accuracy prediction of the hyperparameter groups on the sub-sample datasets to the hyperparameter groups on the entire dataset, greatly improving the hyperparameter optimization efficiency on the entire dataset.
[0141] Specifically, in the embodiment of the present application, after sampling to obtain the sub-sample dataset, the embodiment of the present application proposes to optimize the hyperparameters on the sub-sample dataset to obtain the corresponding hyperparameter group θ and its model performance value y. Among them, each sub-sample dataset includes two major categories: a training subset and a validation subset. For any sub-sample dataset, the hyperparameter optimization training process is as follows:
[0142] S1: Initialize the hyperparameters corresponding to the first recall model;
[0143] S2: Configure the first recall model based on the obtained hyperparameters, and train the first recall model based on the training subset in the sub-sample dataset to obtain the second recall model;
[0144] For example, set a group of initial hyperparameters as θ, and use the training set in the data to learn the recall model of the recommendation system The specific model is as Figure 3 shown. Input the training samples used as the training set in the sub-sample dataset into the Figure 3 shown recall model respectively, and the feature matrices of users and advertisements can be obtained. Based on these data, perform iterative training on the Figure 3 shown recall model to obtain the trained recall model, that is, the second recall model. Next, the second recall model can be verified based on the training samples in the validation set to obtain the corresponding accuracy, and step S3 is executed.
[0145] S3: Input the validation subset in the sub-sample dataset into the second recall model, obtain the accuracy corresponding to the current hyperparameters based on the output result of the second recall model, and use the current hyperparameters and the corresponding accuracy as a set of intermediate hyperparameter groups and true accuracies corresponding to the sub-sample dataset respectively;
[0146] In S2, after training the recall model based on the training set, in S3, the corresponding recall rate y (or other evaluation metrics, such as F-value, ROC, etc.) is obtained on the validation set as the true accuracy corresponding to the current set of hyperparameters. Among them, in the first iteration process, the "current hyperparameters" are those initialized in S1, and in subsequent iteration processes, the "current hyperparameters" refer to those re-obtained after optimization in S4.
[0147] S4: Adjust the current hyperparameters through Bayesian optimization to re-obtain a set of hyperparameters and return to S2.
[0148] That is, select the next set of hyperparameters and repeat S2 to S4. Finally, for any subsample dataset, the input data is the initialized hyperparameters θ, and the output data is all the optimized hyperparameter sets θ and the corresponding model performance values y for each hyperparameter set.
[0149] Among them, after hyperparameter optimization on the subsample dataset, it is also necessary to further extract meta-features from the subsample dataset.
[0150] In the embodiments of the present application, in order to make full use of the data characteristics of the subsample dataset and achieve diverse and personalized hyperparameter prediction, meta-features M are extracted for each subsample dataset, including user feature distribution, advertisement feature distribution, and the interaction between users and advertisements, that is, whether the user clicks on the advertisement. For any subsample dataset, the specific feature extraction method is as Figure 5 shown, including the following steps:
[0151] S1: Input each training sample in the subsample dataset into the second recall model, and based on the second recall model, extract the expected score corresponding to each training sample, where the expected score represents the score of the target operation behavior of the sample object on the sample multimedia content when the sample multimedia content is recommended to the sample object;
[0152] S2: Obtain the feature mean of the expected scores corresponding to each training sample, and determine the distribution feature of the operation label according to the sample object features, sample multimedia content features, and operation labels included in each training sample;
[0153] Assume that the input training sample data is x and the label is y. Based on the already trained recall model First, extract the features of each piece of data and calculate the feature mean (m: the number of samples); then calculate the distribution feature of y under the condition of x according to the label y as where Φ y is the label matrix, Φ xis the feature matrix of the input data, I is the identity matrix, and λ is a hyperparameter to avoid overfitting.
[0154] S3: Combine the feature mean and the distribution feature as the meta-feature of the subsample dataset.
[0155] Finally, the meta-feature of the subsample dataset is obtained as M = [μ x , μ y|x .
[0156] Refer to Figure 6 shown, which is a schematic diagram of an automatic hyperparameter optimization system in an embodiment of the present application. In the embodiment of the present application, the scale of the full dataset is very large. Here, downsampling is performed through adversarial perturbations to obtain N subsample datasets. For each subsample dataset, its meta-feature M is extracted, and then the corresponding hyperparameter group θ and the model performance value y are obtained by using the hyperparameter optimization method, that is, the intermediate hyperparameter group and the accuracy corresponding to the subsample dataset in the embodiment of the present application. Furthermore, a model performance value predictor based on meta-learning is constructed, whose input is the meta-feature M and the hyperparameter group θ of the subsample dataset, and the performance value is predicted through the model performance value predictor and compared with the true model performance value y. This model performance value predictor is trained using the meta-learning algorithm to avoid overfitting when the training data {M, θ, y} is small.
[0157] In the embodiment of the present application, after obtaining the meta-feature M of the subsample dataset, the hyperparameter group θ of each subsample dataset, and the corresponding model performance value y, these data are used to train a model performance value predictor, as Figure 7 shown, aiming to predict the model performance value for any hyperparameter group given the meta-feature. The specific process is as follows:
[0158] Respectively input the meta-feature and the intermediate hyperparameter group corresponding to each subsample dataset into the first accuracy prediction model, and perform multiple rounds of iterative training on the first accuracy prediction model based on the difference between the first predicted accuracy output by the first accuracy prediction model and the corresponding true accuracy, so as to obtain the second accuracy prediction model, where the true accuracy refers to the accuracy of the intermediate hyperparameter group on the first recall model; furthermore, input the meta-feature corresponding to the training sample dataset and the alternative hyperparameter groups into the second accuracy prediction model respectively, and predict the second predicted accuracy corresponding to each alternative hyperparameter group, where the second predicted accuracy refers to the accuracy of the predicted alternative hyperparameter group on the first recall model.
[0159] Specifically, when the model performance value predictor based on meta-learning predicts the accuracy of each alternative hyperparameter group corresponding to the training sample dataset on the recall model, it is first necessary to iteratively train the model performance value predictor based on the meta-features and intermediate hyperparameter groups corresponding to the sub-sample dataset. Among them, each round of iterative training performs the following process:
[0160] Select a certain number of target sub-sample datasets from the unselected sub-sample datasets; input the meta-features and intermediate hyperparameter groups corresponding to each target sub-sample dataset into the first accuracy prediction model respectively, and predict the first prediction accuracy corresponding to each target sub-sample dataset; adjust the network parameters of the first accuracy prediction model according to the difference between the first prediction accuracy and the corresponding true accuracy.
[0161] Among them, the first prediction accuracy in the embodiments of the present application includes meta-training prediction accuracy and meta-validation prediction accuracy; the specific training process of the model performance predictor is as follows:
[0162] S1: Randomly select a sub-sample dataset, use a part of the target sub-sample datasets in the selected target sub-sample dataset as the meta-training set, and use another part of the target sub-sample datasets as the meta-validation set;
[0163] Among them, step S1 means that the data input into the first model performance predictor includes sub-sample dataset meta-feature M, hyperparameter group θ, model performance value y, learning rate α, β; specifically, the model performance value y is used as a label for adjusting model parameters as a reference.
[0164] First, initialize the parameters Ψ of the model performance value predictor (the first accuracy prediction model), randomly select k sub-sample datasets, and divide the training data into the meta-training set D s and the meta-test set D q . Then perform the following steps.
[0165] S2: Input the meta-features and intermediate hyperparameter groups corresponding to each meta-training set into the first accuracy prediction model respectively, and predict the meta-training prediction accuracy corresponding to each meta-training set;
[0166] S3: Update the network parameters of the first accuracy prediction model based on the difference between the meta-training prediction accuracy and the corresponding true accuracy;
[0167] Among them, steps S2 to S3 mean that for each sample in the meta-training set D s calculate the gradient Then use stochastic gradient descent to update the parameters where L(D sIt refers to a loss function constructed based on the difference between the meta-training prediction accuracies corresponding to each meta-training set obtained through prediction and the true accuracies.
[0168] S4: Input the meta-features and hyperparameter groups corresponding to each meta-validation set into the first accuracy prediction model with updated parameters, and predict the meta-learning prediction accuracies corresponding to each meta-validation set.
[0169] S5: Based on the difference between the meta-learning prediction accuracy and the corresponding true accuracy, update the network parameters of the first accuracy prediction model again.
[0170] Among them, steps S4 to S5 indicate that for the meta-training set D s For each sample, use stochastic gradient descent to update the parameters Among them, L(Ψ ′ , D q ) represents the loss function constructed based on the difference between the meta-learning prediction accuracies corresponding to each meta-validation set predicted by the updated first accuracy prediction model and the corresponding true accuracies.
[0171] Through multiple loop iterations, repeat the above S1 to S5. The final output data is the trained model parameters Ψ, and the trained second accuracy prediction model, that is, the trained model performance predictor, is obtained.
[0172] Considering that the process of training a deep model is complex and time-consuming, and the cost of training a deep model and testing its performance value for any hyperparameter group is very high, which is unacceptable in practical applications. The advantage of the model performance predictor proposed in the embodiments of the present application is that it does not require retraining the model, and can predict the performance value of any hyperparameter group according to the existing data features and past hyperparameter groups, which greatly improves the efficiency of hyperparameter optimization.
[0173] Optionally, the model performance predictor can be composed of a multi-layer perceptron. The difficulty in its design and training lies in that the data volume of the existing dataset meta-features and hyperparameter groups is small, usually dozens of groups of data, and the model is prone to overfitting. To solve this problem, the embodiments of the present application propose to use the framework of meta-learning to train the model performance predictor. Specifically, a small amount of training data is divided into a meta-training set and a meta-test set. For the model parameters Ψ, first use the gradient descent method on the meta-training set to obtain Ψ ′ , then obtain the gradient value of the original model parameters Ψ on the meta-test set, and finally update Ψ. This two-layer meta-learning optimization method can avoid overfitting of the model, improve the generalization ability and prediction accuracy of the model. The loss function in the optimization process is the absolute value of the difference between the predicted value and the true value.
[0174] After training the second accuracy prediction model, the hyperparameter group on the full dataset can be predicted based on the second accuracy prediction model. That is, after completing the training of the model performance predictor in the meta-learning framework, the predictor can be directly used to optimize the hyperparameter group on the full dataset. The prediction process is as follows:
[0175] First, extract the meta-features of the full dataset as input conditions, and use the quasi-Newton method L-BFGS to optimize the hyperparameter group to obtain the optimal hyperparameters. L-BFGS is a method for finding the minimum point of a model. By sampling multiple groups of initial input hyperparameters θ and applying the trained model performance predictor Ψ in 5, the predicted output y = Ψ(θ) corresponding to each group of input hyperparameters θ is obtained. Further, the optimization direction is calculated by estimating the first derivative and second derivative of Ψ(θ), and the minimum point of the model is estimated. This method is more efficient than obtaining the model performance value by training on a deep recommendation model. At the same time, it can also use historical learning data to mine the relationship between datasets, find better hyperparameters, and improve the recommendation accuracy of the model.
[0176] Among them, the meta-features of the training sample dataset are obtained by accumulating the meta-features of each sub-sample dataset. When no training samples are added to the training sample dataset, directly accumulate the meta-features of each sub-sample dataset, and the obtained result can be used as the meta-features of the training sample dataset; when training samples are added to the training sample dataset, it is also necessary to extract the meta-features of the sub-sample dataset composed of the added training samples. The specific implementation process is the same as the extraction method of the meta-features of the sub-sample dataset listed above. Then, accumulate the newly extracted meta-features with the meta-features of each sub-sample dataset to obtain the meta-features of the training sample dataset.
[0177] Compared with the existing hyperparameter optimization technology on the full dataset in this application embodiment, the training process is faster and the running efficiency is higher; compared with the hyperparameter optimization technology based on downsampling (Fabolas), this application embodiment proposes a downsampling method based on adversarial perturbation, which can sample more representative sub-sample datasets, and realizes transfer learning of hyperparameters through the meta-features between sub-sample datasets. The hyperparameters obtained by the system have better performance on the full dataset and can perform more accurate advertising targeting recommendations; compared with the hyperparameter transfer technology on multiple datasets, this application embodiment proposes a transfer learning method based on the meta-learning framework, which can directly predict the hyperparameter group on the full dataset, select hyperparameters more efficiently, and can train an accurate recommendation system.
[0178] Refer to Figure 8 As shown, it is a flowchart of the implementation of a content recommendation control method provided by an embodiment of this application. The specific implementation process of this method is as follows:
[0179] S81: Downsample the training sample dataset to obtain multiple sub-sample datasets. The training sample dataset is used to train the first recall model of the target recommendation system. Among them, the training samples in the training sample dataset include sample object features, sample multimedia content features, and operation labels of the sample object and the sample multimedia content. The operation label indicates whether the sample object has a target operation behavior on the sample multimedia content when the sample multimedia content is recommended to the sample object;
[0180] S82: Based on each sub-sample dataset respectively, optimize the hyperparameters of the recall model of the target recommendation system to obtain intermediate hyperparameter groups corresponding to each sub-sample dataset;
[0181] S83: Predict the accuracy of each alternative hyperparameter group on the first recall model respectively according to the accuracy of the intermediate hyperparameter group corresponding to each sub-sample dataset on the recall model. Each alternative hyperparameter group is initialized based on the training sample dataset;
[0182] S84: According to the accuracy of each alternative hyperparameter group on the first recall model, select a hyperparameter group from each alternative hyperparameter group to configure and adjust the hyperparameters of the recall model to obtain the target recall model. The target recall model is used to recommend multimedia content to the target object.
[0183] Optionally, when recommending multimedia content to the target object based on the target recall model, the specific implementation process can refer to Figure 9 . As Figure 9 shown, it is the implementation flowchart of a content recommendation control method provided by an embodiment of the present application. The specific implementation process of this method is as follows:
[0184] S91: Obtain the set of multimedia content to be recommended from the target recommendation system, and obtain the multimedia content features of each multimedia content to be recommended in the multimedia content set and the object features of the target object;
[0185] S92: Input the object features of the target object and the multimedia content features of each multimedia content to be recommended into the target recall model respectively, and predict the expected scores of the target object's target operation behavior on each multimedia content to be recommended based on the target recall model;
[0186] S93: Based on the expected scores corresponding to each multimedia content to be recommended, screen at least one target multimedia content to be recommended from the multimedia content set, and recommend the target multimedia content to be recommended to the target object.
[0187] In summary, the embodiment of the present application designs an automatic hyperparameter optimization system based on meta-learning for large-scale data. By using the method of adversarial sampling to downsample the full dataset, and then using the meta-learning framework to achieve the transfer and prediction of hyperparameters, the efficiency of hyperparameter optimization on large-scale data is improved, and a more accurate recommendation effect is achieved.
[0188] In addition, in the embodiment of the present application, the two parts of the adversarial sampler and the hyperparameter learner are emphasized during the description process. When using adversarial perturbation for downsampling in the adversarial sampler part, it can also be replaced by other sampling methods, such as the cluster-based sampling method, but its performance is inferior to that of adversarial sampling. In the hyperparameter learner part, what is proposed in the embodiment of the present application is to use the meta-features and hyperparameter groups of the sub-dataset as inputs and the model performance value as the output, and perform model training under the meta-learning framework to finally predict the hyperparameter group with better performance on the full dataset / new dataset. This part may be replaced by other mapping methods from hyperparameter groups to model performance values, such as the Bayesian estimation method.
[0189] Refer to Figure 10 As shown, it is a complete method flow chart for hyperparameter optimization in the embodiment of the present application, specifically including the following steps:
[0190] S1001: Obtain the training sample dataset based on the target recommendation system, and downsample the training sample dataset to obtain multiple sub-sample datasets;
[0191] S1002: Select a previously unselected sub-sample dataset, and initialize the hyperparameters corresponding to the first recall model for this sub-sample dataset;
[0192] S1003: Configure the first recall model based on the obtained hyperparameters, and train the first recall model based on the training subset in the sub-sample dataset to obtain the second recall model;
[0193] S1004: Input the validation subset in the sub-sample dataset into the second recall model, obtain the accuracy corresponding to the current hyperparameters based on the output result of the second recall model, and use the current hyperparameters and the corresponding accuracy as a set of intermediate hyperparameter groups and true accuracy corresponding to the sub-sample dataset respectively;
[0194] S1005: Adjust the current hyperparameters through Bayesian optimization to obtain a new set of hyperparameters, and return to S1003;
[0195] S1006: Determine whether the hyperparameters of the first recall model have been optimized based on all sub-sample datasets. If so, execute S1007; otherwise, return to S1002;
[0196] S1007: Extract the meta-features of each sub-sample dataset based on the training samples in each sub-sample dataset respectively;
[0197] S1008: Select a certain number of target sub-sample datasets from the unselected sub-sample datasets;
[0198] S1009: Use part of the target sub-sample datasets in the selected target sub-sample datasets as the meta-training set, and use the other part of the target sub-sample datasets as the meta-validation set;
[0199] S1010: Adjust the network parameters of the first accuracy prediction model according to the difference between the first predicted accuracy and the corresponding true accuracy;
[0200] S1011: Input the meta-features and intermediate hyperparameter groups corresponding to each meta-training set into the first accuracy prediction model respectively, and predict the meta-training prediction accuracies corresponding to each meta-training set;
[0201] S1012: Update the network parameters of the first accuracy prediction model based on the difference between the meta-training prediction accuracy and the corresponding true accuracy;
[0202] S1013: Input the meta-features and hyperparameter groups corresponding to each meta-validation set into the first accuracy prediction model with updated parameters respectively, and predict the meta-learning prediction accuracies corresponding to each meta-validation set;
[0203] S1014: Update the network parameters of the first accuracy prediction model again based on the difference between the meta-learning prediction accuracy and the corresponding true accuracy;
[0204] S1015: Perform multiple rounds of iterative training on the first accuracy prediction model to obtain the second accuracy prediction model;
[0205] S1016: Input the meta-features corresponding to the training sample dataset and the alternative hyperparameter groups into the second accuracy prediction model respectively, and predict the second predicted accuracies corresponding to each alternative hyperparameter group, that is, the accuracies of the predicted alternative hyperparameter groups on the first recall model;
[0206] S1017: Select hyperparameter groups from each alternative hyperparameter group to configure and adjust the hyperparameters of the first recall model according to the accuracies of each alternative hyperparameter group on the first recall model, and obtain the target recall model.
[0207] Based on the same inventive concept, an embodiment of the present application also provides a content recommendation control device. As Figure 11 shown, it is a schematic structural diagram of a content recommendation control device 1100, which may include:
[0208] The first data acquisition unit 1101 is used to downsample the training sample dataset to obtain multiple sub-sample datasets, and the training sample dataset is used to train the first recall model of the target recommendation system;
[0209] The first hyperparameter optimization unit 1102 is used to respectively optimize the hyperparameters of the recall model of the target recommendation system based on each sub-sample dataset, and obtain the intermediate hyperparameter groups corresponding to each sub-sample dataset;
[0210] The first prediction unit 1103 is used to respectively predict the accuracy of each alternative hyperparameter group on the first recall model according to the accuracy of the intermediate hyperparameter groups corresponding to each sub-sample dataset on the first recall model, and each alternative hyperparameter group is initialized based on the training sample dataset;
[0211] The first configuration unit 1104 is used to select a hyperparameter group from each alternative hyperparameter group according to the accuracy of each alternative hyperparameter group on the first recall model, configure the hyperparameters of the first recall model and adjust the hyperparameters to obtain the target recall model;
[0212] The recommendation unit 1105 is used to recommend multimedia content to the target object based on the target recall model.
[0213] Optionally, the training samples in each sub-sample dataset are training samples with anti-interference ability lower than the target condition sampled from the training sample dataset based on adversarial perturbations.
[0214] Optionally, the first data acquisition unit 1101 is specifically used for:
[0215] Screen the training samples that meet the following anti-interference conditions from the training sample dataset as the training samples in the sub-sample dataset: the size of the adversarial perturbation added to the training sample is less than the preset threshold, and when reclassifying the training sample with the added adversarial perturbation based on the first recall model, the training sample with an incorrect classification result; or
[0216] Screen the training samples that meet the following quantity conditions from the training sample dataset as the training samples in the sub-sample dataset: add a certain adversarial perturbation to each training sample, and when reclassifying the training sample with the added adversarial perturbation based on the first recall model, the number of training samples with an incorrect classification result reaches the total number of training samples in each sub-sample dataset.
[0217] Optionally, the first prediction unit 1103 is specifically used for:
[0218] Extract the meta-features of each sub-sample dataset respectively according to the training samples in each sub-sample dataset;
[0219] Respectively input the meta-features corresponding to each sub-sample data set and the intermediate hyperparameter group into the first accuracy prediction model, and perform multiple rounds of iterative training on the first accuracy prediction model based on the difference between the first predicted accuracy output by the first accuracy prediction model and the corresponding true accuracy, so as to obtain the second accuracy prediction model, where the true accuracy refers to the accuracy of the intermediate hyperparameter group on the first recall model;
[0220] Input the meta-features corresponding to the training sample data set and the alternative hyperparameter groups into the second accuracy prediction model respectively, and predict the second predicted accuracy corresponding to each alternative hyperparameter group, where the second predicted accuracy refers to the accuracy of the predicted alternative hyperparameter group on the first recall model.
[0221] Optionally, each sub-sample data set includes two major categories: a training subset and a validation subset; the first prediction unit 1103 is specifically configured to:
[0222] For any sub-sample data set, initialize the hyperparameters corresponding to the first recall model;
[0223] Configure the first recall model based on the obtained hyperparameters, and train the first recall model based on the training subset in the sub-sample data set to obtain the second recall model;
[0224] Input the validation subset in the sub-sample data set into the second recall model, obtain the accuracy corresponding to the current hyperparameters based on the output result of the second recall model, and use the current hyperparameters as a set of intermediate hyperparameter groups corresponding to the sub-sample data set. Among them, the accuracy obtained based on the output result of the second recall model is the true accuracy corresponding to the intermediate hyperparameter group;
[0225] Adjust the current hyperparameters through Bayesian optimization to obtain a new set of hyperparameters, and return to the step of configuring the first recall model based on the obtained hyperparameters and training the first recall model based on the training subset in the sub-sample data set to obtain the second recall model.
[0226] Optionally, the training samples in the training sample data set include sample object features, sample multimedia content features, and operation labels of the sample object and the sample multimedia content. The operation label indicates whether the sample object has a target operation behavior on the sample multimedia content when the sample multimedia content is recommended to the sample object;
[0227] The first prediction unit 1103 is specifically configured to:
[0228] For any sub - sample data set, input each training sample in the sub - sample data set into the second recall model. Based on the second recall model, extract the expected score corresponding to each training sample. The expected score represents the score of the target operation behavior of the sample object for the sample multimedia content when the sample multimedia content is recommended to the sample object.
[0229] Obtain the feature mean value of the expected scores corresponding to each training sample, and determine the distribution feature of the operation label according to the sample object feature, the sample multimedia content feature and the operation label included in each training sample.
[0230] Combine the feature mean value and the distribution feature as the meta - feature of the sub - sample data set.
[0231] Optionally, the first prediction unit 1103 is specifically configured to perform the following process in each round of iterative training:
[0232] Select a certain number of target sub - sample data sets from the unselected sub - sample data sets.
[0233] Input the meta - features and the intermediate hyper - parameter groups corresponding to each target sub - sample data set into the first accuracy prediction model respectively, and predict the first prediction accuracy corresponding to each target sub - sample data set.
[0234] Adjust the network parameters of the first accuracy prediction model according to the difference between the first prediction accuracy and the corresponding true accuracy.
[0235] Optionally, the first prediction accuracy includes the meta - training prediction accuracy and the meta - validation prediction accuracy. The first prediction unit 1103 is specifically configured to:
[0236] Use some of the target sub - sample data sets in the selected target sub - sample data set as the meta - training set, and use the other part of the target sub - sample data set as the meta - validation set.
[0237] Input the meta - features and the intermediate hyper - parameter groups corresponding to each meta - training set into the first accuracy prediction model respectively, and predict the meta - training prediction accuracy corresponding to each meta - training set.
[0238] Update the network parameters of the first accuracy prediction model based on the difference between the meta - training prediction accuracy and the corresponding true accuracy.
[0239] Input the meta - features and the hyper - parameter groups corresponding to each meta - validation set into the first accuracy prediction model with updated parameters respectively, and predict the meta - learning prediction accuracy corresponding to each meta - validation set.
[0240] Update the network parameters of the first accuracy prediction model again based on the difference between the meta - learning prediction accuracy and the corresponding true accuracy.
[0241] Optionally, the meta-features of the training sample dataset are obtained by accumulating the meta-features of each sub-sample dataset.
[0242] Based on the same inventive concept, an embodiment of the present application further provides a training device for a recall model. As Figure 12 shown, it is a schematic structural diagram of a training device 1200 for a recall model, which may include:
[0243] A second data acquisition unit 1201, configured to perform downsampling on the training sample dataset to obtain a plurality of sub-sample datasets. The training sample dataset is used to train a first recall model of the target recommendation system. Among them, the training samples in the training sample dataset include sample object features, sample multimedia content features, and operation labels of the sample object and the sample multimedia content. The operation label indicates whether the sample object has a target operation behavior on the sample multimedia content when the sample multimedia content is recommended to the sample object;
[0244] A second hyperparameter optimization unit 1202, configured to respectively optimize the hyperparameters of the recall model of the target recommendation system based on each sub-sample dataset to obtain intermediate hyperparameter groups corresponding to each sub-sample dataset;
[0245] A second prediction unit 1203, configured to respectively predict the accuracy of each alternative hyperparameter group on the first recall model according to the accuracy of the intermediate hyperparameter group corresponding to each sub-sample dataset on the first recall model. Each alternative hyperparameter group is initialized based on the training sample dataset;
[0246] A second configuration unit 1204, configured to select a hyperparameter group from each alternative hyperparameter group according to the accuracy of each alternative hyperparameter group on the first recall model, perform hyperparameter configuration on the first recall model and perform hyperparameter adjustment to obtain a target recall model. The target recall model is used to recommend multimedia content to a target object.
[0247] For the convenience of description, the above parts are divided into various modules (or units) according to functions and described separately. Of course, when implementing the present application, the functions of each module (or unit) can be implemented in the same or multiple software or hardware.
[0248] After introducing the content recommendation control method and device of the exemplary embodiment of the present application, next, an electronic device according to another exemplary embodiment of the present application is introduced.
[0249] Those skilled in the art can understand that various aspects of the present application can be implemented as a system, a method, or a program product. Therefore, various aspects of the present application can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuits", "modules", or "systems" here.
[0250] Based on the same inventive concept as the above method embodiment, an electronic device is also provided in an embodiment of the present application. This electronic device can be used for content recommendation control. In one embodiment, this electronic device can be a server, such as Figure 1 the server 120 shown. In this embodiment, the structure of the electronic device can be as Figure 13 shown, including a memory 1301, a communication module 1303, and one or more processors 1302.
[0251] The memory 1301 is used to store the computer program executed by the processor 1302. The memory 1301 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system and programs required to run the instant messaging function, etc.; the data storage area can store various instant messaging information and operation instruction sets, etc.
[0252] The memory 1301 can be a volatile memory, such as a random-access memory (RAM); the memory 1301 can also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or the memory 1301 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 1301 can be a combination of the above memories.
[0253] The processor 1302 can include one or more central processing units (CPUs) or be a digital processing unit, etc. The processor 1302 is used to implement the above content recommendation control method when calling the computer program stored in the memory 1301.
[0254] The communication module 1303 is used to communicate with the terminal device and other servers.
[0255] In the embodiments of the present application, the specific connection medium between the above-mentioned memory 1301, communication module 1303 and processor 1302 is not limited. In the embodiments of the present disclosure, Figure 13 it is shown that the memory 1301 and the processor 1302 are connected through a bus 1304. The bus 1304 is represented by a thick line in Figure 13 . The connection manners between other components are only for illustrative purposes and are not to be taken as limiting. The bus 1304 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 13 only one thick line is used to represent it in, but it does not mean that there is only one bus or one type of bus.
[0256] The memory 1301 stores a computer storage medium, and the computer storage medium stores computer-executable instructions, which are used to implement the content recommendation control method of the embodiments of the present application. The processor 1302 is used to execute the above-mentioned content recommendation control method, as Figure 2 shown. The processor 1302 can also be used to execute the above-mentioned training method of the recall model, as Figure 8 shown.
[0257] The program product of the embodiments of the present application can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0258] The program product of the embodiments of the present application can adopt a portable compact disk read-only memory (CD-ROM) and include program code, and can run on a computing device. However, the program product of the present application is not limited to this. In this document, the readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with a command execution system, apparatus, or device.
[0259] A readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium can also be any readable medium other than a readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with a command execution system, apparatus, or device. The program code contained on the readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.
[0260] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including those of the above method embodiments. The foregoing storage medium includes: various media that can store program code, such as a removable storage device, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc. Alternatively, if the above integrated unit in the embodiments of the present application is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the related art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media that can store program code, such as a removable storage device, a ROM, a RAM, a magnetic disk, or an optical disc.
[0261] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these changes and modifications of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.
Claims
1. A content recommendation control method, characterized in that, The method includes: Downsampling a training sample dataset to obtain multiple subsample datasets, where the training sample dataset is used to train a first recall model of a target recommendation system; wherein, the training samples in the training sample dataset include sample object features, sample multimedia content features, and operation labels of the sample object and the sample multimedia content; Based on each subsample dataset respectively, hyperparameter optimization is performed on the first recall model to obtain intermediate hyperparameter groups respectively corresponding to each subsample dataset; Based on the accuracy of each intermediate hyperparameter group corresponding to each subsample dataset on the first recall model and a hyperparameter learner respectively, the accuracy of each alternative hyperparameter group on the first recall model is predicted, and each alternative hyperparameter group is initialized based on the training sample dataset; the hyperparameter learner is: trained with the accuracy of each intermediate hyperparameter group corresponding to each subsample dataset on the first recall model as training labels and with each intermediate hyperparameter group corresponding to each subsample dataset and meta-features as inputs; Based on the accuracy of each alternative hyperparameter group on the first recall model, a hyperparameter group is selected from each alternative hyperparameter group to perform hyperparameter configuration and hyperparameter adjustment on the first recall model to obtain a target recall model; Based on the target recall model, multimedia content is recommended to a target object.
2. The method according to claim 1, wherein The training samples in each subsample dataset are training samples sampled from the training sample dataset based on adversarial perturbations and having anti-interference ability lower than a target condition.
3. The method according to claim 2, wherein The training samples sampled from the training sample dataset and having anti-interference ability lower than a target condition specifically include: From the training sample dataset, training samples that meet the following anti-interference condition are selected as the training samples in the subsample dataset: the size of the adversarial perturbation added to the training sample is less than a preset threshold, and when the first recall model re-classifies the training sample with the added adversarial perturbation, the training sample with a wrong classification result; or From the training sample dataset, training samples that meet the following quantity condition are selected as the training samples in the subsample dataset: a certain adversarial perturbation is added to each training sample, and when the first recall model re-classifies the training sample with the added adversarial perturbation, the number of training samples with a wrong classification result reaches the sum of the training samples of each subsample dataset.
4. The method according to claim 1, wherein The hyperparameter learner to be trained is a first accuracy prediction model, and the trained hyperparameter learner is a second accuracy prediction model; the step of respectively predicting the accuracy of each alternative hyperparameter group on the first recall model based on the accuracy of each intermediate hyperparameter group corresponding to each subsample dataset on the first recall model and the hyperparameter learner specifically includes: Based on the training samples in each subsample dataset respectively, the meta-features of each subsample dataset are extracted; Respectively input the meta-features corresponding to each sub-sample data set and the intermediate hyperparameter group into the first accuracy prediction model, and perform multiple rounds of iterative training on the first accuracy prediction model based on the difference between the first predicted accuracy output by the first accuracy prediction model and the corresponding true accuracy, so as to obtain a second accuracy prediction model, where the true accuracy refers to the accuracy of the intermediate hyperparameter group on the first recall model; Input the meta-features corresponding to the training sample data set and the alternative hyperparameter groups into the second accuracy prediction model respectively, and predict the second predicted accuracy corresponding to each alternative hyperparameter group, where the second predicted accuracy refers to the accuracy of the predicted alternative hyperparameter group on the first recall model.
5. The method according to claim 4, wherein Each sub-sample data set includes two major categories: a training subset and a validation subset; when respectively optimizing the hyperparameters of the first recall model based on each sub-sample data set to obtain the intermediate hyperparameter groups corresponding to each sub-sample data set, for any one sub-sample data set, it specifically includes: Initialize the hyperparameters corresponding to the first recall model; Configure the first recall model based on the obtained hyperparameters, and train the first recall model based on the training subset in the sub-sample data set to obtain a second recall model; Input the validation subset in the sub-sample data set into the second recall model, obtain the accuracy corresponding to the current hyperparameters based on the output result of the second recall model, and use the current hyperparameters as a set of intermediate hyperparameter groups corresponding to the sub-sample data set, where the accuracy obtained based on the output result of the second recall model is the true accuracy corresponding to the intermediate hyperparameter group; Adjust the current hyperparameters through Bayesian optimization to obtain a new set of hyperparameters, and return to the step of configuring the first recall model based on the obtained hyperparameters and training the first recall model based on the training subset in the sub-sample data set to obtain a second recall model.
6. The method according to claim 5, wherein The operation label indicates whether the sample object has a target operation behavior on the sample multimedia content when the sample multimedia content is recommended to the sample object; When respectively extracting the meta-features of each sub-sample data set according to the training samples in each sub-sample data set, for any one sub-sample data set, it specifically includes: Input each training sample in the sub-sample data set into the second recall model, and extract the expected score corresponding to each training sample based on the second recall model, where the expected score represents the score of the sample object having a target operation behavior on the sample multimedia content when the sample multimedia content is recommended to the sample object; Obtain the feature mean of the expected scores corresponding to each training sample, and determine the distribution feature of the operation label according to the sample object feature, sample multimedia content feature, and operation label included in each training sample; Combine the feature mean and the distribution feature as the meta-feature of the sub-sample data set.
7. The method according to claim 4, characterized in that Each round of iterative training performs the following process: Select a certain number of target sub-sample data sets from the unselected sub-sample data sets; Input the meta-features and intermediate hyperparameter groups corresponding to each target sub-sample data set into the first accuracy prediction model respectively, and predict the first prediction accuracies corresponding to each target sub-sample data set; Adjust the network parameters of the first accuracy prediction model according to the difference between the first prediction accuracy and the corresponding true accuracy.
8. The method according to claim 7, wherein The first prediction accuracy includes meta-training prediction accuracy and meta-validation prediction accuracy; Input the meta-features and intermediate hyperparameter groups corresponding to each target sub-sample data set into the first accuracy prediction model respectively, and predict the first prediction accuracies corresponding to each target sub-sample data set; Adjust the network parameters of the first accuracy prediction model according to the difference between the first prediction accuracy and the corresponding true accuracy, specifically including: Use a part of the target sub-sample data sets in the selected target sub-sample data set as the meta-training set, and use another part of the target sub-sample data sets as the meta-validation set; Input the meta-features and intermediate hyperparameter groups corresponding to each meta-training set into the first accuracy prediction model respectively, and predict the meta-training prediction accuracies corresponding to each meta-training set; Update the network parameters of the first accuracy prediction model based on the difference between the meta-training prediction accuracy and the corresponding true accuracy; Input the meta-features and hyperparameter groups corresponding to each meta-validation set into the first accuracy prediction model with updated parameters respectively, and predict the meta-learning prediction accuracies corresponding to each meta-validation set; Update the network parameters of the first accuracy prediction model again based on the difference between the meta-learning prediction accuracy and the corresponding true accuracy.
9. The method according to any one of claims 4 to 8, characterized in that, The meta-features of the training sample data set are obtained by accumulating the meta-features of each sub-sample data set.
10. A training method for a recall model, characterized in that, The method includes: Perform downsampling on the training sample data set to obtain multiple sub-sample data sets. The training sample data set is used to train the first recall model of the target recommendation system. Among them, the training samples in the training sample data set include sample object features, sample multimedia content features, and operation labels of the sample object and the sample multimedia content. The operation label indicates whether the sample object has a target operation behavior on the sample multimedia content when the sample multimedia content is recommended to the sample object; Based on each sub-sample data set respectively, optimize the hyperparameters of the first recall model to obtain intermediate hyperparameter groups corresponding to each sub-sample data set respectively; Predict the accuracies of each alternative hyperparameter group on the first recall model respectively according to the accuracies of the intermediate hyperparameter groups corresponding to each sub-sample data set on the first recall model and the hyperparameter learner. Each alternative hyperparameter group is initialized based on the training sample data set; The hyperparameter learner is: trained with the accuracies of the intermediate hyperparameter groups corresponding to each sub-sample data set on the first recall model as training labels and the intermediate hyperparameter groups and meta-features corresponding to each sub-sample data set as inputs; Select a hyperparameter group from each alternative hyperparameter group according to the accuracy of each alternative hyperparameter group on the first recall model, configure the hyperparameters of the first recall model and adjust the hyperparameters to obtain a target recall model, where the target recall model is used to recommend multimedia content to a target object.
11. A content recommendation control device, characterized in that, Including: A first data acquisition unit, configured to downsample a training sample data set to obtain a plurality of sub-sample data sets, where the training sample data set is used to train a first recall model of a target recommendation system; wherein, the training samples in the training sample data set include sample object features, sample multimedia content features, and operation labels of the sample object and the sample multimedia content; A first hyperparameter optimization unit, configured to respectively optimize the hyperparameters of the first recall model based on each sub-sample data set to obtain intermediate hyperparameter groups respectively corresponding to each sub-sample data set; A first prediction unit, configured to respectively predict the accuracy of each alternative hyperparameter group on the first recall model according to the accuracy of the intermediate hyperparameter group corresponding to each sub-sample data set on the first recall model and a hyperparameter learner, where each alternative hyperparameter group is initialized based on the training sample data set; the hyperparameter learner is: trained with the accuracy of the intermediate hyperparameter group corresponding to each sub-sample data set on the first recall model as training labels and the intermediate hyperparameter group corresponding to each sub-sample data set and meta-features as inputs; A first configuration unit, configured to select a hyperparameter group from each alternative hyperparameter group according to the accuracy of each alternative hyperparameter group on the first recall model, configure the hyperparameters of the first recall model and adjust the hyperparameters to obtain a target recall model; A recommendation unit, configured to recommend multimedia content to a target object based on the target recall model.
12. The device according to claim 11, wherein The training samples in each sub-sample data set are training samples sampled from the training sample data set based on adversarial perturbations and having anti-interference ability lower than a target condition.
13. A training device for a recall model, characterized in that, Including: A second data acquisition unit, configured to downsample a training sample data set to obtain a plurality of sub-sample data sets, where the training sample data set is used to train a first recall model of a target recommendation system, where the training samples in the training sample data set include sample object features, sample multimedia content features, and operation labels of the sample object and the sample multimedia content, and the operation label indicates whether the sample object has a target operation behavior on the sample multimedia content when the sample multimedia content is recommended to the sample object; A second hyperparameter optimization unit, configured to respectively optimize the hyperparameters of the first recall model based on each sub-sample data set to obtain intermediate hyperparameter groups respectively corresponding to each sub-sample data set; A second prediction unit, configured to predict the accuracy of each alternative hyperparameter group on the first recall model respectively according to the accuracy of the intermediate hyperparameter group corresponding to each sub-sample data set on the first recall model and a hyperparameter learner, where each alternative hyperparameter group is obtained by initializing based on the training sample data set; the hyperparameter learner is: trained with the accuracy of the intermediate hyperparameter group corresponding to each sub-sample data set on the first recall model as the training label, and with the intermediate hyperparameter group corresponding to each sub-sample data set and meta-features as the input; A second configuration unit, configured to select a hyperparameter group from each alternative hyperparameter group according to the accuracy of each alternative hyperparameter group on the first recall model, perform hyperparameter configuration and hyperparameter adjustment on the first recall model, and obtain a target recall model, where the target recall model is used to recommend multimedia content to a target object.
14. An electronic device, characterized in that, It includes a processor and a memory, where the memory stores program code, and when the program code is executed by the processor, the processor is caused to execute the steps of any one of claims 1 to 9 or the steps of the method according to claim 10.
15. A computer-readable storage medium, characterized in that, It includes program code, and when the program code runs on an electronic device, the program code is used to cause the electronic device to execute the steps of any one of claims 1 to 9 or the steps of the method according to claim 10.
Citation Information
Patent Citations
Fast hyperparameter search for machine-learning program
US20190122141A1