Training method and device of recall model and storage medium thereof
Through the sorting learning method, the recall model is optimized, and the matrix is constructed using the pre-trained fine-scheduled model and recall model, and the loss is calculated and the parameters are updated. The problem of inconsistency between the recall model and the fine-scheduled model is solved, and the number of calculations is reduced and the consistency of the model is improved.
Patent Information
- Application Number
- CN202410009625.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-03
- Publication Date
- 2025-07-04
AI Technical Summary
The existing recall model is inconsistent with the fine sorting model, with large calculations, cannot meet the time constraints of the advertising system, and poor sorting effect.
By obtaining the training sample set, the object features are extracted, the information features and cross features are promoted, and the pre-trained fine-scheduled model and recall model are used to predict, the first and second matrices are constructed, the loss is calculated and the parameters of the recall model iteratively is updated until the preset conditions are met.
It improves the consistency between the recall model and the fine-scheduling model, reduces the number of calculations, improves the model training speed, and enhances the accuracy of the recall model.
Smart Images

Figure CN120258164A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and particularly to a training method, device and computer storage medium of a recall model. Background Art
[0002] Since the recall model needs to calculate all items, if all ads are scored online in real time, the time constraints of the ad system cannot be met. Using the two-tower model for recall is the current mainstream recall solution. The two-tower model completely decouples the object tower and the promoted object tower, and completes the embedding calculation of all recommended items offline. Only the user's embedding calculation is performed online. Due to the limitations of the recall model, the sorting effect often has a large gap from the fine-rank model.
[0003] The existing recall model does not use a sorting learning scheme for optimization, and the recall model is inconsistent with the fine-rank model. In addition, the number of inference calculations of the pairwise scheme in the existing recall model is O(n 2 ), and the computational complexity is too large. Summary of the Invention
[0004] In view of this, the embodiments of the present application provide a training method, device and storage medium of a recall model, which can improve the performance of the recall model, improve the consistency between the recall model and the fine-rank model, and reduce the number of inference calculations.
[0005] According to a first aspect of the present application, there is provided a method for training a recall model, the method comprising: obtaining a training sample set, the training sample set including positive samples and negative samples, the positive samples including promoted information implicitly feedback by an object, and the negative samples including randomly sampled promoted information; for each sample in the training sample set, extracting object features, promoted information features, and cross features of the object and the promoted information therefrom; inputting the object features, promoted information features, and cross features of each sample into a pre-trained fine-ranking model to obtain a first prediction result for each sample, the first prediction result including a fine-ranking score indicating the recommendation degree of the sample; inputting the object features and promoted information features of each sample into the recall model to obtain a second prediction result for each sample, the second prediction result including a recall score indicating the recommendation degree of the sample; constructing a first matrix based on the first prediction results of the samples in the training sample set, the set of elements of the first matrix corresponding one-to-one to all first prediction result ordered pairs formed by taking any two first prediction results from the first prediction results of the samples in the training sample set, and each element in the first matrix representing the magnitude relationship between the fine-ranking scores of the two first prediction results in the corresponding first prediction result ordered pair; constructing a second matrix based on the second prediction results of the samples in the training sample set, the set of elements of the second matrix corresponding one-to-one to all second prediction result ordered pairs formed by taking any two second prediction results from the second prediction results of the samples in the training sample set, and each element in the second matrix representing the magnitude relationship between the recall scores of the two second prediction results in the corresponding second prediction result ordered pair; calculating a first loss based on the difference between the corresponding elements of the first matrix and the second matrix; and iteratively updating the parameters of the recall model at least based on the first loss until a preset condition is satisfied.
[0006] In some embodiments, for each sample in the training sample set, extracting object features, promoted information features, and cross features of the object and the promoted information therefrom includes: for each sample in the training sample set, extracting object features therefrom, the object features including object discrete features and object continuous features; for each sample in the training sample set, extracting promoted information features therefrom, the promoted information features including promoted information discrete features and promoted information continuous features; and for each sample in the training sample set, generating cross features based on the object features and the promoted information features.
[0007] In some embodiments, inputting the object feature and the promotion information feature of each sample into a recall model to obtain the second prediction result for each sample includes: inputting the object feature and the promotion information feature of each sample into a two-tower model, where the two-tower model includes at least an object tower and a promotion information tower, and the promotion information tower includes a first promotion information tower for inputting positive samples of promotion information features and a second promotion information tower for inputting negative samples of promotion information features. The parameters of the first promotion information tower and the second promotion information tower are the same and are synchronized during iterative updates, and output the second prediction result for each promotion information in the training samples.
[0008] In some embodiments, inputting the object feature, the promotion information feature, and the cross feature of each sample into a pre-trained fine-ranking model to obtain the first prediction result for each sample includes: concatenating the object feature, the promotion information feature, and the cross feature of the object and the promotion information of each sample, inputting the concatenated result into the pre-trained fine-ranking model, and outputting the first prediction result for each promotion information in the training samples and an embedding vector of M*(m + n + k) dimensions, where m is the number of object features, n is the number of promotion information features, k is the number of cross features of the object and the promotion information, and M is a positive integer.
[0009] In some embodiments, inputting the object feature and the promotion information feature of each sample into a recall model to obtain the second prediction result for each sample includes: concatenating the object features to obtain an embedding vector of m*M dimensions, where m is the number of object features and M is a positive integer; concatenating the promotion information features to obtain an embedding vector of n*M dimensions, where n is the number of promotion information features; and inputting the embedding vector of m*M dimensions and the embedding vector of n*M dimensions into the recall model to obtain the second prediction result for each promotion information in the training samples.
[0010] In some embodiments, constructing a first matrix based on the first prediction results of each sample in the training sample set includes: for each ordered pair of first prediction results: comparing the fine-ranking scores y i and y j in the ordered pair of first prediction results, where i = 1, 2,..., M, j = 1, 2,..., M, and M represents the number of samples in the training set; in response to y i being greater than y j , setting the element y ij corresponding to the ordered pair of first prediction results in the first matrix to 1; in response to y i being less than or equal to y j , setting the element y ij of the first matrix to 0; and based on the element y ijConstruct the first matrix.
[0011] In some embodiments, constructing the second matrix based on the second prediction results of each sample in the training sample set includes: for each pair of second prediction results: comparing the recall scores p i and p j where i = 1, 2, …, M, j = 1, 2, …, M, and M represents the number of samples in the training set; in response to p i being greater than p j , setting the element p ij corresponding to the pair of second prediction results in the second matrix to 1; in response to p i being less than or equal to p j , setting the element p ij of the second matrix to 0; and constructing the second matrix based on the elements p ij of the second matrix.
[0012] In some embodiments, the first loss includes Bayesian personalized ranking loss.
[0013] In some embodiments, iteratively updating the parameters of the recall model based at least on the first loss includes: calculating a second loss based on the second matrix and the positive and negative sample labels of each sample in the training sample set; calculating a weighted sum of the first loss and the second loss to determine the target loss of the recall model; and iteratively updating the parameters of the recall model based on the target loss.
[0014] In some embodiments, the second loss includes cross-entropy loss.
[0015] In some embodiments, the generalized information of the object implicit feedback includes at least one of the following: the promoted information clicked by the object, the promoted information browsed by the object, and the promoted information purchased by the object.
[0016] According to a second aspect of the present application, there is provided a training device for a recall model, including an acquisition module configured to acquire a training sample set, where the training sample set includes positive samples and negative samples, the positive samples include promotion information implicitly fed back by an object, and the negative samples include randomly sampled promotion information; an extraction module configured to extract, for each sample in the training sample set, object features, promotion information features, and cross features between the object and the promotion information; a fine-ranking module configured to input the object features, promotion information features, and cross features of each sample into a pre-trained fine-ranking model to obtain a first prediction result for each sample, where the first prediction result includes a fine-ranking score indicating the recommendation degree of the sample; a recall module configured to input the object features and promotion information features of each sample into the recall model to obtain a second prediction result for each sample, where the second prediction result includes a recall score indicating the recommendation degree of the sample; a first matrix construction module configured to construct a first matrix based on the first prediction results of the samples in the training sample set, where the set of elements of the first matrix corresponds one-to-one to all ordered pairs of the first prediction results formed by taking any two first prediction results from the first prediction results of the samples in the training sample set, and each element in the first matrix represents the magnitude relationship between the fine-ranking scores of the two first prediction results in the corresponding ordered pair of the first prediction results; a second matrix construction module configured to construct a second matrix based on the second prediction results of the samples in the training sample set, where the set of elements of the second matrix corresponds one-to-one to all ordered pairs of the second prediction results formed by taking any two second prediction results from the second prediction results of the samples in the training sample set, and each element in the second matrix represents the magnitude relationship between the recall scores of the two second prediction results in the corresponding ordered pair of the second prediction results; a loss calculation module configured to calculate a first loss based on the difference between the corresponding elements of the first matrix and the second matrix; an iterative update module configured to iteratively update the parameters of the recall model at least based on the first loss until a preset condition is met.
[0017] According to a third aspect of the present application, there is provided a computing device, including: a memory and a processor, where a computer program is stored in the memory, and when the computer program is executed by the processor, it causes the processor to execute the training method of the recall model according to some embodiments of the present application.
[0018] According to a fourth aspect of the present application, there is provided a computer-readable storage medium, on which computer-readable instructions are stored, and when the computer-readable instructions are executed, they implement the training method of the recall model according to some embodiments of the present application.
[0019] According to a fifth aspect of the present application, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the training method of the recall model according to some embodiments of the present application.
[0020] The training method, device, and storage medium of the recall model provided by the embodiments of the present application at least include the following beneficial effects: The present application uses ranking learning to optimize the recall model, improving the consistency between the recall model and the fine-ranking model. By using the first matrix to identify the ranking between any two prediction results in the fine-ranking model prediction results, and using the second matrix to identify the ranking between any two prediction results in the recall model prediction results, an exponential reduction in the number of calculations is achieved. In addition, since the training method of this recall model does not require sampling, all information in the queue is retained.
[0021] According to the embodiments described below, these and other aspects of the present application will be clear and will be elucidated with reference to the embodiments described below. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In the following description of exemplary embodiments with reference to the accompanying drawings, more details, features, and advantages of the present application are disclosed. In the drawings:
[0023] Figure 1 is a schematic diagram of an exemplary implementation environment provided by an exemplary embodiment of the present application;
[0024] Figure 2 is a flowchart of a training method of a recall model provided by an exemplary embodiment of the present application;
[0025] Figure 3 is a schematic diagram of the construction of a fine-ranking model provided by an exemplary embodiment of the present application;
[0026] Figure 4 is a schematic diagram of the construction of a recall model provided by an exemplary embodiment of the present application;
[0027] Figure 5 is a schematic diagram of the construction of a fine-ranking model in the related art;
[0028] Figure 6 is a block diagram of a training device of a recall model provided by an exemplary embodiment of the present application; and
[0029] Figure 7 is an example block diagram of a computing device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] The technical solutions in the present application will be clearly and completely described below with reference to the accompanying drawings in the present application. The described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.
[0031] Artificial intelligence technology is a comprehensive discipline that involves a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, intelligent transportation, and automatic control.
[0032] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and active learning.
[0033] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, common ones include smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, intelligent healthcare, intelligent customer service, vehicle networking, autonomous driving, and intelligent transportation. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0034] Before introducing the embodiments of the present application in detail, some related concepts are first explained.
[0035] 1. Learning To Rank (LTR): Usually abbreviated as LTR, learning to rank is the application of machine learning in information retrieval systems, and its goal is to build a ranking model for ranking lists. Typical applications of learning to rank include search lists, recommendation lists, and advertising lists, etc. The purpose of list ranking is to rank multiple entries. LTR is usually divided into three categories: Pointwise, Pairwise, and Listwise, and the three have different input and output spaces respectively. Based on different assumption conditions, different loss functions are used.
[0036] 2. Bayesian Personalized Ranking Loss (BPR): BPR loss is often used in conjunction with Pairwise in Learning to Rank (LTR) to evaluate the ranking quality of two different advertisements. The BPR loss function was initially proposed by Steffen Rendle et al. in the paper "BPR: Bayesian Personalized Ranking from implicit feedback". It is a loss function for learning user personalized preferences in a recommendation system. In a recommendation system, users' historical behavior data usually exists in an implicit form, such as users' browsing, purchasing, or clicking behavior. Compared with explicit feedback data (such as users' ratings), implicit feedback data is sparser and more difficult to interpret. Therefore, the recommendation system needs to develop models and algorithms suitable for implicit feedback data to recommend items. The BPR loss function was proposed to solve the recommendation problem under implicit feedback data. Its basic idea is: given a user and two items, the model needs to rank the item that the user prefers more before the item that the user prefers less, so as to learn the user's personalized preferences. Compared with other recommendation algorithms, BPR has better performance and scalability, and can handle large-scale implicit feedback data. Therefore, the BPR loss function has been widely used in both academia and industry.
[0037] In the existing technical solutions for Learning to Rank, the prediction effect of the fine-ranking model is better than that of the recall model; and the results recalled by the recall model ultimately need to be provided to the fine-ranking model for scoring. This application uses the method of Learning to Rank to improve the recall model. The object that the recall model learns is the fine-ranking model, which improves the consistency between the recall model and the fine-ranking model, and avoids the recall model recalling too many advertisements that the fine-ranking model considers inappropriate, wasting system resources.
[0038] In Learning to Rank, Pairwise is a frequently selected ranking learning solution. However, the advertisement queue waiting to be ranked often has a large number, and since Pairwise requires pairwise combination, the number of model training times is of the order of O(n 2 ). When n is very large, O(n 2 ) will cause the training time to increase explosively. This application proposes a fast calculation solution, which reduces the number of computational inference times of the order of O(n 2 ) to the order of O(n), greatly reducing the computational amount of Pairwise BPR loss.
[0039] Figure 1 Schematically shows a schematic diagram of an exemplary scenario 100 provided according to an exemplary embodiment of the present application.
[0040] As Figure 1As shown in the figure, scenario 100 includes a computing device 101. The recall model / rank refinement model provided by the embodiments of the present disclosure can be deployed on the computing device 101. The recall model / rank refinement model at least includes a deep learning model trained using training samples. The computing device 101 can include, but is not limited to, mobile phones, computers, intelligent voice interaction devices, intelligent household appliances, vehicle-mounted terminals, aircraft, etc. The embodiments of the present disclosure can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, etc.
[0041] Exemplarily, user 102 can use the computing device 101 to perform a recall service on promotional information. For example, user 102 can input instructions through the user interface provided by the computing device 101, such as through relevant physical or virtual buttons, through text, voice, or gesture instructions, etc., so as to initiate a recall service deployed on the computing device 101 and / or the server 103.
[0042] Scenario 100 may also include a server 103. Optionally, the training method of the recall model provided by the embodiments of the present disclosure can also be deployed on the server 103. Or, optionally, the training method of the recall model provided by the embodiments of the present disclosure can also be deployed on a combination of the computing device 101 and the server 103. The present disclosure does not make specific limitations in this regard. For example, user 102 can access the server 103 via the computing device 101 through the network 105 to obtain the services provided by the server 103.
[0043] The server 103 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. In addition, it should be understood that the server 103 is only shown as an example. In fact, other devices or combinations of devices with computing power and storage capacity can also be used alternatively or additionally to provide corresponding services.
[0044] Optionally, the computing device 101 and / or the server 103 can be linked to the database 104 via the network 105, so as to obtain relevant data of training samples from the database 104, etc. Exemplarily, the database 104 can be an independent data storage device or a cluster of devices, or can also be a backend data storage device or a cluster of devices related to other online services. The method for completing the knowledge graph proposed in this article can be applied to application programs such as vector graphics editing software, vector drawing applications, image processing software, etc. The recall model / ranking model is a deep learning model that runs in the above application programs in the form of a plug-in. When a user uses the above application programs such as vector graphics editing software, vector drawing applications, image processing software, etc., the deep learning model can be invoked via the user interface of the application program. For example, the user inputs demand description information in the demand input box of the above application program, and invokes the plug-in of the text generation model in the application program based on the demand description information; and obtains the promotion information generated by the text generation model plug-in based on the demand description information.
[0045] In addition, in the present disclosure, the network 105 can be a wired network connected via, for example, cables, optical fibers, etc., or can also be a wireless network such as 2G, 3G, 4G, 5G, Wi-Fi, Bluetooth, ZigBee, Li-Fi, etc., or can also be the internal connection lines of one or several devices, etc.
[0046] Figure 2 The flowchart of the training method 200 of the recall model provided by an embodiment of the present application is shown.
[0047] In step S210, a training sample set is obtained. The training sample set includes positive samples and negative samples. The positive samples include the promotion information implicitly feedback by the object, and the negative samples include the promotion information randomly sampled. Specifically, the promotion information clicked by the user is a positive sample, and the promotion information randomly sampled from the entire library is a negative sample. Here, the promotion information refers to the object for which the user performs recall, including advertisements, music, and any other objects for recall. The promotion information implicitly feedback by the object includes at least one of the following: the promotion information clicked by the object, the promotion information browsed by the object, and the promotion information purchased by the object. In another example, for example, the promotion information rejected by the user can also be used as a negative sample.
[0048] In step S220, for each sample in the training sample set, object features, promotion information features, and cross-features of the object and the promotion information are extracted therefrom. In one embodiment, object features are extracted from the training samples, and the object features include object discrete features and object continuous features. The object discrete features may include, for example, user identity ID, etc. The object continuous features may include, for example, object portrait features (such as age, gender, etc.); object behavior features, etc. The object behavior features may be, for example, the article reading records, advertisement click records, advertisement conversion records, etc. of the object extracted within the past 14 days. In one embodiment, in step S220, object features for the recall model are extracted respectively for the recall model, and object features for the re-rank model are extracted for the re-rank model. In one embodiment, promotion information features are extracted from the training samples, and the promotion information features include promotion information discrete features and promotion information continuous features. Here, taking the promotion information as an advertisement as an example, the advertisement discrete features include the advertisement ID. The advertisement continuous features may include features such as advertisement category, advertiser, etc. In one embodiment, cross-features of the object and the promotion information are constructed for the re-rank model. Here, the cross-features of the object and the promotion information refer to the features that have restrictions in both the object feature dimension and the cross-feature dimension of the object and the promotion information. For example, in the scenario where a user purchases shoes, the feature that is restricted in both the dimensions of the user being female and the shoe brand being A is the cross-feature of the object and the promotion information.
[0049] In step S230, the object features, promotion information features, and cross-features of each sample are input into the pre-trained re-rank model to obtain a first prediction result for each sample, and the first prediction result includes a re-rank score indicating the recommendation degree of the sample.
[0050] In one embodiment, n advertisements to be sorted for a user are constructed into a batch, where n is a positive integer greater than or equal to 2. In one embodiment, the fine-ranking model is pre-trained; the object features, promotion information features, and cross-features of the object and promotion information are input into the pre-trained fine-ranking model. There are no special requirements for the fine-ranking model here, and any trained fine-ranking model can be used. A brief introduction to the fine-ranking model is given below. The trained fine-ranking model is used to predict the n advertisements in this batch to obtain the fine-ranking of the n advertisements. For example, first, the object features, the promotion information features, and the cross-features of the object and promotion information extracted in step S220 are passed through the embedding layer. For discrete-type features, the discrete-type digital fields are mapped into M-dimensional vectors, where M is a positive integer. In one example, M can be, for example, 64. As understood by those skilled in the art, M is an example and can vary according to the actual situation of using fine-ranking. For continuous-type features, such as the consumption amount in the object features, it needs to be discretized by binning, and then the discretized features are mapped into M-dimensional embedding vectors.
[0051] In one embodiment, the object features, the promotion information features, and the cross-features of the object and promotion information are input into the fine-ranking model to obtain a first prediction result for each promotion information in the training sample. Figure 3 FIG. schematically shows a schematic diagram of a fine-ranking model 300 according to an embodiment of the present application. Specifically, the object features, the promotion information features, and the cross-features of the object and promotion information are connected and input into the fine-ranking model, and a first prediction result for each promotion information in the training sample and an M*(m + n + k)-dimensional embedding vector are output, where m is the number of object features, n is the number of promotion information features, k is the number of cross-features of the object and promotion information, and M is a positive integer. For example, when N is 64, the object features, the promotion information features, and the embedding vectors of the user and promotion information types are connected, and forward calculation is performed to obtain the scores y1, y2,..., y of the n advertisements n , and at the same time, a 64*(m + n + k)-dimensional embedding vector is obtained.
[0052] In step S240, the object feature and the promotion information feature of each sample are input into the recall model to obtain a second prediction result for each sample, where the second prediction result includes a recall score indicating the recommendation degree of the sample. First, the object feature, the promotion information feature, and the cross feature of the object and the promotion information extracted in step S220 are passed through the embedding layer of the recall model. For discrete type features, the discrete type digital fields are mapped into M-dimensional vectors, where M is a positive integer. In one example, N can be, for example, 64. As understood by those skilled in the art, N is an example and can vary according to the actual situation of using the refined ranking. For continuous type features, such as the consumption amount in the object feature, it needs to be discretized by binning, and then the discretized feature is mapped into an M-dimensional embedding vector. The embedding vectors of the object features are concatenated to obtain an m*M-dimensional vector, where m is the number of object features; the embedding vectors of the advertisement features are concatenated to obtain an n*M-dimensional vector, where n is the number of advertisement features. The above m*M-dimensional vector and n*M-dimensional vector are input into the corresponding recall network, and through forward calculation, the prediction result logits = [p1, p2, … p n .
[0053] Figure 4 FIG. shows a schematic diagram of a recall model 400 provided by an embodiment of the present application. The recall model here adopts a two-tower model. Since the recall model needs to calculate all advertisements, scoring all advertisements online in real time cannot meet the time constraints of the advertisement system. Adopting a two-tower model including an object tower and a promotion information tower can completely decouple the user and the promotion information. Thus, the embedding calculation of the advertisement is completed offline, and only the user embedding calculation is performed online in real time.
[0054] In one embodiment, inputting the object feature and the promotion information feature into the recall model to obtain a second prediction result for each promotion information in the training sample includes: inputting the object feature and the promotion information feature into a two-tower model, where the two-tower model at least includes an object tower and a promotion information tower, and the promotion information tower includes a first promotion information tower for inputting positive samples of the promotion information feature and a second promotion information tower for inputting negative samples of the promotion information feature. The parameters of the first promotion information tower and the second promotion information tower are the same and are synchronized during iterative update; outputting a second prediction result for each promotion information in the training sample.
[0055] In step S250, a first matrix is constructed based on the first prediction results of each sample in the training sample set. The set of elements of the first matrix corresponds one-to-one with all the first prediction result ordered pairs formed by taking any two first prediction results from the first prediction results of each sample in the training sample set. Each element in the first matrix represents the magnitude relationship between the re-ranking scores of the two first prediction results in the corresponding first prediction result ordered pair. A prediction difference matrix, that is, the first matrix, is constructed as follows: First, any two first prediction results are taken from the first prediction results of each sample in the training sample set to form a first prediction result ordered pair. The set of elements of the first matrix corresponds one-to-one with all the first prediction result ordered pairs. Each element in the first matrix represents the magnitude relationship between the re-ranking scores of the two first prediction results in the corresponding first prediction result ordered pair. In one example, the first prediction result is denoted as y, then the element of the first matrix is y - tf.transpose(y). In one example, the first matrix is shown in Table 1. If the re-ranking of the i-th promoted information is higher than the re-ranking of the j-th promoted information, then the element y of the first matrix ij = 1; if the re-ranking of the i-th promoted information is equal to or lower than the re-ranking of the j-th promoted information, then the element y of the first matrix ij = 0.
[0056] 0 ....... <![CDATA[y ij > ....... 0 ....... <![CDATA[y ji > ....... 0
[0057] Table 1 Schematic diagram of an example of the first matrix
[0058] In step S260, a second matrix is constructed based on the second prediction results of each sample in the training sample set. The set of elements of the second matrix corresponds one-to-one with all the second prediction result ordered pairs formed by taking any two second prediction results from the second prediction results of each sample in the training sample set. Each element in the second matrix represents the magnitude relationship between the recall scores of the two second prediction results in the corresponding second prediction result ordered pair. A prediction difference matrix, that is, the second matrix, is constructed based on the second prediction results. The construction method of the first matrix is as follows: First, any two second prediction results are taken from the second prediction results of each sample in the training sample set to form a second prediction result ordered pair. The set of elements of the second matrix corresponds one-to-one with all the second prediction result ordered pairs. Each element in the second matrix represents the magnitude relationship between the recall scores of the two second prediction results in the corresponding second prediction result ordered pair. In one example, the second prediction result is denoted as p, then the element of the second matrix is p - tf.transpose(p). In one example, the second matrix is shown in Table 2. Similar to the first matrix, if the recall ranking of the i-th promoted information is higher than the recall ranking of the j-th promoted information, then the element p of the second matrix ij= 1. If the recall ranking of the i-th promoted information is equal to or lower than that of the j-th promoted information, then the element p of the second matrix ij = 0.
[0059] 0 ....... <![CDATA[p ij > ....... 0 ....... <![CDATA[p ji > ....... 0
[0060] Schematic diagram of an example of the second matrix in Table 2
[0061] In step S270, calculate the first loss based on the difference between the corresponding elements of the first matrix and the second matrix.
[0062] In one embodiment, calculate the first loss based on the difference between the element y of the first matrix ij and the element p of the second matrix ij ; calculate the second loss based on the difference between the second prediction result and the sample label; calculate the target loss according to the weighted sum of the first loss and the second loss.
[0063] Here, the first loss adopts the Bayesian Personalized Ranking (BPR) loss. Specifically, the calculation method of the BPR loss is as follows: where p ij and y ij are calculated as described above. N is the number of promoted information in the batch to be sorted for the user. The second loss adopts the cross-entropy loss. where N is the number of promoted information in the batch to be sorted for the user, y i is the sample label, and p i is the prediction result. The target loss Loss = β * LossMain + (1 - β) * LossBPR is calculated according to the weighted sum of the first loss and the second loss. β is a hyperparameter, which is set according to empirical values and is used to balance the first loss and the second loss.
[0064] In step S280, at least based on the first loss, iterate and update the parameters of the recall model until a preset condition is met. Specifically, stop the iteration when the target loss meets a predetermined condition. The predetermined condition can be that the target loss is less than a predetermined threshold, or the number of training iterations reaches a predetermined number.
[0065] The training method, device, and storage medium of the recall model provided by the embodiments of the present application at least include the following beneficial effects: The present application uses ranking learning to optimize the recall model, improving the consistency between the recall model and the fine-ranking model. By using the first matrix to identify the ranking between any two prediction results in the fine-ranking model prediction results, and using the second matrix to identify the ranking between any two prediction results in the recall model prediction results, an exponential reduction in the number of calculations is achieved. In addition, since the training method of this recall model does not require sampling, all information in the queue is retained. The training method 200 of the recall model significantly improves the consistency between the recall model and the fine-ranking model. Specifically, the recall and fine-ranking consistency recall@10_100 is increased by 20%, where recall@10_100 represents the ratio of the number of elements in the intersection of the top 100 promoted information retrieved and the top 10 promoted information in the fine-ranking to 100.
[0066] In addition, the number of inferences of the training method 200 of the recall model is reduced from O(n 2 ) to O(n), and the model training speed is increased by 6 times. Figure 5 Shows a schematic diagram of the construction of the fine-ranking model 500 in the related art. In Figure 5 it, the fine-ranking model 200 adopts the original pairwise pairing method Pairwise. For example, Figure 5 in it, the first promoted information and the second promoted information form a pair (user, ad1), (user, ad2); the second promoted information and the third promoted information form a pair (user, ad2), (user, ad3), and so on.
[0067] Figure 6The block diagram of a training device 600 for a recall model provided by an embodiment of the present application is shown. The training device 600 for the recall model includes: an acquisition module 610 configured to acquire a training sample set, the training sample set including positive samples and negative samples, the positive samples including promotion information implicitly feedback by an object, and the negative samples including randomly sampled promotion information; an extraction module 620 configured to extract, for each sample in the training sample set, object features, promotion information features, and cross features of the object and the promotion information; a fine-ranking module 630 configured to input the object features, promotion information features, and cross features of each sample into a pre-trained fine-ranking model to obtain a first prediction result for each sample, the first prediction result including a fine-ranking score indicating the recommendation degree of the sample; a recall module 640 configured to input the object features and promotion information features of each sample into the recall model to obtain a second prediction result for each sample, the second prediction result including a recall score indicating the recommendation degree of the sample; a first matrix construction module 650 configured to construct a first matrix based on the first prediction results of the respective samples in the training sample set, the set of elements of the first matrix corresponding one-to-one to all first prediction result ordered pairs formed by taking any two first prediction results from the first prediction results of the respective samples in the training sample set, and each element in the first matrix representing the magnitude relationship between the fine-ranking scores of the two first prediction results in the corresponding first prediction result ordered pair; a second matrix construction module 660 configured to construct a second matrix based on the second prediction results of the respective samples in the training sample set, the set of elements of the second matrix corresponding one-to-one to all second prediction result ordered pairs formed by taking any two second prediction results from the second prediction results of the respective samples in the training sample set, and each element in the second matrix representing the magnitude relationship between the recall scores of the two second prediction results in the corresponding second prediction result ordered pair; a loss calculation module 670 configured to calculate a first loss based on the difference between the corresponding elements of the first matrix and the second matrix; and an iterative update module 680 configured to iteratively update the parameters of the recall model at least based on the first loss until the target loss meets a preset condition.
[0068] It should be understood that the training device 600 for the recall model can be implemented in a software, hardware, or software-hardware combination manner. Multiple different modules in the device can be implemented in the same software or hardware structure, or one module can be implemented by multiple different software or hardware structures.
[0069] In addition, the training device 600 for the recall model can be used to implement the training method 200 for the recall model described above. The relevant details have been described in detail above. For the sake of brevity, they will not be repeated here. Additionally, these devices can have the same features and advantages as those described in the corresponding methods.
[0070] Figure 7FIG. illustrates an example system 700 that includes an example computing device 710 representative of one or more systems and / or devices that can implement the various methods described herein. The computing device 710 can be, for example, a server of a service provider, a device associated with the server, a system-on-chip, and / or any other suitable computing device or computing system. The training apparatus 600 of the recall model described above with reference to Figure 6 can take the form of the computing device 710. Alternatively, Figure 6 the training apparatus 600 of the recall model described above can be implemented as a computer program in the form of an application 716.
[0071] As illustrated, the example computing device 710 includes a processing system 711, one or more computer-readable media 712, and one or more I / O interfaces 713 that are communicatively coupled to each other. Although not shown, the computing device 710 may also include a system bus or other data and command transfer system that couples the various components to each other. The system bus can include any one or combination of different bus structures, such as a memory bus or memory controller, a peripheral bus, a universal serial bus, and / or a processor or local bus utilizing any one of a variety of bus architectures. Also contemplated are various other examples, such as control and data lines.
[0072] The processing system 711 represents the functionality to perform one or more operations using hardware. Accordingly, the processing system 711 is illustrated as including hardware elements 714 that can be configured as processors, functional blocks, etc. This can include being implemented in hardware as an application-specific integrated circuit or other logic device formed using one or more semiconductors. The hardware elements 714 are not limited by the materials from which they are formed or the processing mechanisms employed therein. For example, a processor can be composed of (multiple) semiconductors and / or transistors (e.g., an electronic integrated circuit (IC)). In such a context, the executable instructions of the processor can be electronically executable instructions.
[0073] The computer-readable media 712 is illustrated as including a memory / storage device 715. The memory / storage device 715 represents the memory / storage capacity associated with one or more computer-readable media. The memory / storage device 715 can include volatile media (such as random access memory (RAM)) and / or non-volatile media (such as read-only memory (ROM), flash memory, optical discs, magnetic disks, etc.). The memory / storage device 715 can include fixed media (e.g., RAM, ROM, fixed hard disk drive, etc.) and removable media (e.g., flash memory, removable hard disk drive, optical disc, etc.). The computer-readable media 712 can be configured in a variety of other ways as described further below.
[0074] One or more I / O interfaces 713 represent the functionality that allows a user to input commands and information into the computing device 710 using various input devices and optionally also allows information to be presented to the user and / or other components or devices using various output devices. Examples of input devices include keyboards, cursor control devices (e.g., mice), microphones (e.g., for voice input), scanners, touch capabilities (e.g., capacitive or other sensors configured to detect physical touch), cameras (e.g., that can detect motion not involving touch as a gesture using visible or invisible wavelengths such as infrared frequencies), and so on. Examples of output devices include display devices, speakers, printers, network cards, haptic response devices, and the like. Thus, the computing device 710 can be configured in various ways as further described below to support user interaction.
[0075] The computing device 710 also includes an application 716. The application 716 can be, for example, a software instance of the training device 600 for the Figure 6 recall model described, and implements the techniques described herein in combination with other elements in the computing device 710.
[0076] This application provides a computer program product or a computer program that includes computer instructions stored in a computer-readable storage medium. The processor of the computing device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions such that the computing device performs the method for presenting visual data provided in the above various alternative implementations.
[0077] Various techniques may be described herein in the general context of software, hardware elements, or program modules. Generally, these modules include routines, programs, objects, elements, components, data structures, etc. that perform particular tasks or implement particular abstract data types. The terms "module", "function", and "component" as used herein generally represent software, firmware, hardware, or a combination thereof. The techniques described herein are platform-independent, meaning that these techniques can be implemented on various computing platforms having various processors.
[0078] Implementations of the described modules and techniques can be stored on or transmitted across some form of computer-readable medium. Computer-readable media can include various media accessible by the computing device 710. By way of example and not limitation, computer-readable media can include "computer-readable storage media" and "computer-readable signal media".
[0079] Contrary to mere signal transmission, carrier, or the signal itself, a "computer-readable storage medium" refers to a medium and / or device that can persistently store information, and / or a tangible storage device. Thus, a computer-readable storage medium refers to a non-signal-bearing medium. Computer-readable storage media include hardware such as volatile and non-volatile, removable and non-removable media and / or storage devices implemented in a method or technology suitable for storing information such as computer-readable instructions, data structures, program modules, logic elements / circuits, or other data. Examples of computer-readable storage media can include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD), or other optical storage devices, hard disks, magnetic cassettes, magnetic tape, magnetic disk storage devices, or other magnetic storage devices, or other storage devices, tangible media, or articles of manufacture suitable for storing the desired information and accessible by a computer.
[0080] A "computer-readable signal medium" refers to a signal-bearing medium configured to send instructions to a computing device 710, such as via a network. Signal media typically can embody computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave, data signal, or other transmission mechanism. Signal media also include any information delivery medium. The term "modulated data signal" refers to a signal in which one or more of the characteristics are set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media include wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, RF, infrared, and other wireless media.
[0081] As previously, hardware elements 714 and computer-readable media 712 represent instructions, modules, programmable device logic, and / or fixed device logic implemented in hardware, which in some embodiments can be used to implement at least some aspects of the techniques described herein. Hardware elements can include integrated circuits or systems-on-a-chip, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs), and components of other hardware devices implemented in silicon or other hardware. In this context, a hardware element can serve as a processing device that executes program tasks defined by the instructions, modules, and / or logic embodied by the hardware element, and as a hardware device for storing instructions for execution, e.g., the previously described computer-readable storage medium.
[0082] The foregoing combinations can also be used to implement the various techniques and modules herein. Accordingly, software, hardware, or program modules and other program modules can be implemented as one or more instructions and / or logic on a computer-readable storage medium of a certain form and / or embodied by one or more hardware elements 714. The computing device 710 can be configured to implement specific instructions and / or functions corresponding to the software and / or hardware modules. Thus, for example, by using the computer-readable storage medium of the processing system and / or the hardware element 714, the module can be implemented at least in part in hardware as a module executable by the computing device 710 as software. The instructions and / or functions can be executable / operable by one or more articles of manufacture (e.g., one or more computing devices 710 and / or processing systems 711) to implement the techniques, modules, and examples herein.
[0083] In various embodiments, the computing device 710 can assume a variety of different configurations. For example, the computing device 710 can be implemented as a computer-like device including a personal computer, a desktop computer, a multi-screen computer, a laptop computer, a netbook, etc. The computing device 710 can also be implemented as a mobile device-like device including mobile devices such as mobile phones, portable music players, portable game devices, tablet computers, multi-screen computers, etc. The computing device 710 can also be implemented as a television-like device, which includes a device having or connected to a generally larger screen in a leisure viewing environment. These devices include televisions, set-top boxes, game consoles, etc.
[0084] The techniques described herein can be supported by these various configurations of the computing device 710 and are not limited to the specific examples of the techniques described herein. The functionality can also be implemented in whole or in part using a distributed system, such as on the "cloud" 720 via a platform 722 as described below.
[0085] The cloud 720 includes and / or represents a platform 722 for resources 724. The platform 722 underlies the hardware (e.g., servers) and software resources of the cloud 720. The resources 724 can include applications and / or data that can be used when performing computer processing on a server remote from the computing device 710. The resources 724 can also include services provided via the Internet and / or via a subscriber network such as a cellular or Wi-Fi network.
[0086] The platform 722 can abstract resources and functionality to connect the computing device 710 with other computing devices. The platform 722 can also be used to abstract the hierarchy of resources to provide a corresponding level of hierarchy for the demand for resources 724 encountered via the platform 722. Thus, in an interconnected device embodiment, the implementation of the functions described herein can be distributed throughout the system 700. For example, the functionality can be implemented in part on the computing device 710 and via the platform 722 that abstracts the functionality of the cloud 720.
[0087] It should be understood that, for clarity, embodiments of the present application have been described with reference to different functional units. However, it will be apparent that, without departing from the present application, the functionality of each functional unit can be implemented in a single unit, in multiple units, or as part of other functional units. For example, functionality described as being performed by a single unit can be performed by multiple different units. Thus, the reference to a particular functional unit is only considered as a reference to an appropriate unit for providing the described functionality, rather than indicating a strict logical or physical structure or organization. Accordingly, the present application can be implemented in a single unit, or can be physically and functionally distributed among different units and circuits.
[0088] In embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of that module or unit.
[0089] Although the present application has been described in connection with some embodiments, it is not intended to be limited to the specific forms set forth herein. Instead, the scope of the present application is only limited by the appended claims. Additionally, although individual features may be included in different claims, these may possibly be advantageously combined, and including in different claims does not imply that a combination of features is not feasible and / or advantageous. The order of features in the claims does not imply that the features must work in any particular order. Further, in the claims, the word "comprising" does not exclude other elements, and the terms "a" or "an" do not exclude a plurality. The reference numerals in the claims are provided only as illustrative examples and should not be construed as limiting the scope of the claims in any way.
[0090] It can be understood that, in the specific implementation of the present application, it involves data related to entities such as promotion information and object implicit feedback. When the above embodiments of the present application are applied to specific products or technologies, obtaining user permission or consent is required, and the collection, use, and processing of relevant data are required to comply with relevant laws, regulations, and standards of relevant countries and regions.
Claims
1. A training method for a recall model, characterized in that, The method includes: Obtaining a training sample set, where the training sample set includes positive samples and negative samples, the positive samples include generalized information of object implicit feedback, and the negative samples include randomly sampled generalized information; For each sample in the training sample set, extracting object features, generalized information features, and cross features of the object and the generalized information therefrom; Inputting the object features, generalized information features, and cross features of each sample into a pre-trained fine-ranking model to obtain a first prediction result for each sample, where the first prediction result includes a fine-ranking score indicating the recommendation degree of the sample; Inputting the object features and generalized information features of each sample into a recall model to obtain a second prediction result for each sample, where the second prediction result includes a recall score indicating the recommendation degree of the sample; Constructing a first matrix based on the first prediction results of each sample in the training sample set, where the set of elements of the first matrix corresponds one-to-one with all the first prediction result ordered pairs formed by taking any two first prediction results from the first prediction results of each sample in the training sample set, and each element in the first matrix represents the magnitude relationship between the fine-ranking scores of the two first prediction results in the corresponding first prediction result ordered pair; Constructing a second matrix based on the second prediction results of each sample in the training sample set, where the set of elements of the second matrix corresponds one-to-one with all the second prediction result ordered pairs formed by taking any two second prediction results from the second prediction results of each sample in the training sample set, and each element in the second matrix represents the magnitude relationship between the recall scores of the two second prediction results in the corresponding second prediction result ordered pair; Calculating a first loss based on the difference between the corresponding elements of the first matrix and the second matrix; Iteratively updating the parameters of the recall model at least based on the first loss until a preset condition is satisfied.
2. The method according to claim 1, wherein For each sample in the training sample set, extracting object features, generalized information features, and cross features of the object and the generalized information therefrom includes: For each sample in the training sample set, extracting object features, where the object features include object discrete features and object continuous features; For each sample in the training sample set, extracting generalized information features, where the generalized information features include generalized information discrete features and generalized information continuous features; For each sample in the training sample set, generating cross features based on the object features and the generalized information features.
3. The method according to claim 1, characterized in that, Inputting the object features and generalized information features of each sample into the recall model to obtain a second prediction result for each sample includes: Inputting the object features and the generalized information features of each sample into a two-tower model, where the two-tower model includes at least an object tower and a generalized information tower, the generalized information tower includes a first generalized information tower for inputting positive samples of generalized information features and a second generalized information tower for inputting negative samples of generalized information features, and the parameters of the first generalized information tower and the second generalized information tower are the same and are synchronized during iterative update, and outputting a second prediction result for each generalized information in the training sample.
4. The method according to claim 1, characterized in that, Inputting the object features, generalized information features, and cross features of each sample into a pre-trained fine-ranking model to obtain a first prediction result for each sample includes: Concatenate the object features, the promotion information features, and the cross - features of the object and the promotion information of each sample, and input them into a pre - trained fine - ranking model. Output the first prediction result for each promotion information in the training sample and an embedding vector of M*(m + n + k) dimensions, where m is the number of object features, n is the number of promotion information features, k is the number of cross - features of the object and the promotion information, and M is a positive integer.
5. The method according to claim 1, wherein Input the object features and promotion information features of each sample into a recall model to obtain the second prediction result for each sample, including: Concatenate the object features to obtain an embedding vector of m*M dimensions, where m is the number of object features and M is a positive integer; Concatenate the promotion information features to obtain an embedding vector of n*M dimensions, where n is the number of promotion information features; Input the m*M - dimensional embedding vector and the n*M - dimensional embedding vector into the recall model to obtain the second prediction result for each promotion information in the training sample.
6. The method according to claim 1, wherein Construct a first matrix based on the first prediction results of each sample in the training sample set, including: For each pair of first prediction results: Compare the fine-ranking scores y of the two first prediction results in the first prediction result ordered pair i and y j to determine their magnitudes, where i = 1, 2, …, M, j = 1, 2, …, M, and M represents the number of samples in the training set; In response to y i Greater than y j , set the element y corresponding to the ordered pair of the first prediction result in the first matrix ij to 1; In response to y i Less than or equal to y j , set the element y of the first matrix ij to 0; and Element y of the first matrix ij Construct the first matrix.
7. The method according to claim 1, wherein Construct a second matrix based on the second prediction results of each sample in the training sample set, including: For each pair of second prediction results: Compare the recall scores p i and p j in the two second prediction results of the second prediction result ordered pair, where i = 1, 2, …, M, j = 1, 2, …, M, and M represents the number of samples in the training set; In response to p i Greater than p j , set the element p corresponding to the second prediction result ordered pair in the second matrix ij to 1; In response to p i Less than or equal to p j , set the element p of the second matrix ij to 0; and Based on the element p of the second matrix ij Construct the second matrix.
8. The method according to claim 1, wherein The first loss includes Bayesian personalized ranking loss.
9. The method according to claim 1, characterized in that, The at least iterative update of the parameters of the recall model based on the first loss includes: Calculate a second loss based on the second matrix and the positive and negative sample labels of each sample in the training sample set; Calculate the weighted sum of the first loss and the second loss to determine the target loss of the recall model; Iteratively update the parameters of the recall model based on the target loss.
10. The method according to claim 9, characterized in that The second loss includes cross - entropy loss.
11. The method according to claim 1, wherein The promotion information with implicit object feedback includes at least one of the following: the promotion information clicked by the object, the promotion information browsed by the object, and the promotion information purchased by the object.
12. A training device for a recall model, characterized in that, The device includes: An acquisition module configured to acquire a training sample set, the training sample set including positive samples and negative samples, the positive samples including promotion information with implicit object feedback, and the negative samples including randomly sampled promotion information; An extraction module configured to extract, for each sample in the training sample set, object features, promotion information features, and cross - features of the object and the promotion information; A fine - ranking module configured to input the object features, promotion information features, and cross - features of each sample into a pre - trained fine - ranking model to obtain the first prediction result for each sample, and the first prediction result includes a fine - ranking score indicating the recommendation degree of the sample; A recall module configured to input the object features and promotion information features of each sample into a recall model to obtain the second prediction result for each sample, and the second prediction result includes a recall score indicating the recommendation degree of the sample. A first matrix construction module, configured to construct a first matrix based on the first prediction results of each sample in a training sample set, where the set of elements of the first matrix corresponds one-to-one to all first prediction result ordered pairs formed by taking any two first prediction results from the first prediction results of each sample in the training sample set, and each element in the first matrix represents the magnitude relationship between the two first prediction results in the corresponding first prediction result ordered pair; A second matrix construction module, configured to construct a second matrix based on the second prediction results of each sample in a training sample set, where the set of elements of the second matrix corresponds one-to-one to all second prediction result ordered pairs formed by taking any two second prediction results from the second prediction results of each sample in the training sample set, and each element in the second matrix represents the magnitude relationship between the two second prediction results in the corresponding second prediction result ordered pair; A loss calculation module, configured to calculate a first loss based on the difference between the corresponding elements of the first matrix and the second matrix; An iterative update module, configured to iteratively update the parameters of the recall model at least based on the first loss until a preset condition is satisfied.
13. A computing device, characterized in that, Comprising: A memory, configured to store computer-executable instructions; A processor, configured to execute the method according to any one of claims 1-11 when the computer-executable instructions are executed by the processor.
14. A computer-readable storage medium, characterized in that, It stores computer-executable instructions, and the computer-executable instructions, when executed, implement the method according to any one of claims 1-11.
15. A computer program product, characterized in that, Comprising a computer program, and the computer program, when executed, implements the steps of the method according to any one of claims 1-11.