Method for training rough ranking scoring model for e-commerce commodities and related products
By introducing a multi-objective scoring module and a coarse ranking distillation module into the coarse ranking model of the e-commerce platform and calculating the minimum and maximum distillation loss, the problems of poor coarse ranking effect and complex loss function in the existing technology are solved, and more efficient knowledge transfer and personalized recommendations are achieved.
Patent Information
- Application Number
- CN202410425132.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-09
- Publication Date
- 2025-10-14
AI Technical Summary
Existing distillation learning methods cannot effectively learn the order of coarse and fine ranking in the product recommendation system of e-commerce platforms, and the loss function calculation is complex, which affects the recommendation effect. It is especially difficult to control the knowledge transfer effect in multi-objective business tasks.
A multi-objective scoring module and a rough ranking distillation module are used to calculate the minimum and maximum distillation loss, optimize the rough ranking model, combine the multi-objective scoring results, improve the consistency of the rough and fine ranking, reduce the computational complexity of the loss function, and control the degree of knowledge transfer.
The scoring accuracy and ranking effect of the coarse ranking model are improved, and the knowledge of the fine ranking model can be easily transferred to the coarse ranking model under multi-objective business tasks, thereby enhancing the personalized ranking ability of the recommendation system.
Smart Images

Figure CN120782508A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application generally relates to the technical field of recommendation system. More specifically, the present application relates to a method, device and computer readable storage medium for training a coarse ranking model for e-commerce goods. BACKGROUND
[0002] Under the trend of rapid development of global e-commerce, e-commerce platforms have penetrated into people's lives and affected users' shopping habits and met users' demand for personalized shopping. With the development and expansion of e-commerce platforms, the number of goods that e-commerce platforms can provide for users to purchase is increasing dramatically. In order to find and recommend goods that meet users' personalized shopping needs from a large number of goods, the recommendation system in e-commerce platforms is crucial. At present, with the help of mature artificial intelligence technology, the recommendation system has also become relatively mature, which is mainly divided into a recall stage and a ranking stage according to the recommendation process. A large number of goods are first recalled by the recall layer to obtain a batch of goods that users may be interested in, thereby narrowing down the recommended goods set, and then the batch of goods are finely ranked in the ranking stage and the goods that users are most interested in are displayed to users.
[0003] Since the goods returned by the recall generally have more than a thousand goods, directly entering a large number of goods into the ranking stage for fine ranking will consume huge computing resources, so the ranking stage is divided into a coarse ranking stage and a fine ranking stage. Among them, the personalized model in the coarse ranking stage is responsible for quickly ranking the recalled goods under the requirement of less time consumption, and then inputting the top n goods in the sequence to the fine ranking stage, while the fine ranking stage consumes more time to rank and push the goods more personalized and fine to users. Therefore, the personalized ranking effect of coarse ranking is generally worse than that of fine ranking. The existing method is to let the coarse ranking learn the ranking results of the fine ranking, and this learning process is called cascade learning, and this process also uses the idea of distillation learning or transfer learning to transfer (also known as distill) the knowledge of the fine ranking to the coarse ranking. However, the existing distillation learning idea cannot learn the sequence of the whole scoring and ranking well, thereby affecting the learning effect of ranking, and the loss function calculation is relatively complex. In addition, the existing method takes the distillation loss as an auxiliary loss, which cannot well adjust the influence of the knowledge transferred by the fine ranking model (i.e. teacher model) on the final output score of the coarse ranking model (i.e. student model) when multiple target business tasks are involved.
[0004] Therefore, it is urgent to provide a scheme for training a coarse ranking model for e-commerce goods, so as to efficiently calculate the loss function, reduce the calculation complexity of the loss function, enable the coarse ranking model to learn the knowledge of the fine ranking model better, and also enable the target ability learned by the fine ranking to be easily transferred to the coarse ranking model without affecting the existing coarse ranking knowledge under multiple target business tasks. SUMMARY
[0005] To at least solve one or more technical problems as mentioned above, the present application proposes, in multiple aspects, a scheme for training a coarse ranking scoring model for e-commerce commodities.
[0006] In a first aspect, the present application provides a method for training a coarse ranking scoring model for e-commerce commodities, wherein the coarse ranking scoring model comprises a multi-objective scoring module and a coarse ranking distillation module, and the method comprises: generating a training sample according to a list of recommended commodities of an e-commerce platform returned by a user request; performing, based on the training sample, a scoring operation on ranking of commodities under a target task using the multi-objective scoring module to obtain a multi-objective scoring result; performing, based on the training sample, distillation learning on scoring knowledge of fine ranking of commodities for a fine ranking model using the coarse ranking distillation module, and calculating a minimax distillation loss; and optimizing a final coarse ranking result of coarse ranking of commodities of the coarse ranking scoring model according to the multi-objective scoring result and the minimax distillation loss to train the coarse ranking scoring model for e-commerce commodities.
[0007] In a second aspect, the present application provides a device for training a coarse ranking scoring model for e-commerce commodities, comprising: a processor; and a memory having computer instructions for training a coarse ranking scoring model for e-commerce commodities stored thereon, which, when executed by the processor, cause the implementation of the embodiments in the foregoing first aspect.
[0008] In a third aspect, the present application provides a non-transitory computer-readable storage medium having computer program instructions for training a coarse ranking scoring model for e-commerce commodities stored thereon, which, when executed by one or more processors, cause the implementation of the embodiments in the foregoing first aspect.
[0009] Through the scheme for training a coarse ranking scoring model for e-commerce commodities as provided above, the embodiments of the present application improve the consistency of coarse ranking and fine ranking scoring by simply calculating a minimax distillation loss in distillation learning of scoring knowledge of fine ranking of commodities for a fine ranking model in a coarse ranking distillation module, so as to efficiently calculate a loss function, reduce the calculation complexity of the loss function, and make the coarse ranking model better fit the fine ranking result, thereby improving the scoring precision of the coarse ranking model. Further, the embodiments of the present application distill fine ranking knowledge into an independent model target tower (i.e., the coarse ranking distillation module), and optimize the final coarse ranking result of coarse ranking of commodities of the coarse ranking scoring model by combining the multi-objective scoring result. Based on this, even under multi-objective business tasks, the target ability of fine ranking learning can be easily migrated to the coarse ranking model without affecting the existing coarse ranking knowledge, so that the knowledge degree of the migrated fine ranking model is controllable. BRIEF DESCRIPTION OF DRAWINGS
[0010] The above and other objects, features and advantages of the present application will become more apparent from the following detailed description read in conjunction with the accompanying drawings, in which:
[0011] Figure 1 is an exemplary flow chart illustrating a method of training a coarse ranking model for e-commerce goods according to an embodiment of the present application;
[0012] Figure 2 is an exemplary flow chart illustrating generation of training samples according to an embodiment of the present application;
[0013] Figure 3 is an exemplary flow chart illustrating generation of an overall training sample according to an embodiment of the present application;
[0014] Figure 4 is an exemplary flow chart illustrating calculation of minimax distillation loss according to an embodiment of the present application;
[0015] Figure 5 is an exemplary schematic diagram illustrating calculation of minimax distillation loss according to an embodiment of the present application;
[0016] Figure 6 is an exemplary schematic diagram illustrating an overall coarse ranking model according to an embodiment of the present application;
[0017] Figure 7 is an exemplary flow chart illustrating a method of coarse ranking of e-commerce goods according to an embodiment of the present application;
[0018] Figure 8 is an exemplary structural block diagram of a device for training a coarse ranking model for e-commerce goods according to an embodiment of the present application. DETAILED DESCRIPTION
[0019] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0020] It should be understood that the terms "comprises" and "comprising" used in the specification and claims of the application indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0021] It should also be understood that the terms used in the specification and the claims of the application are for the purpose of describing particular embodiments and are not intended to be limiting of the application. As used in the specification and the claims, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "and / or," as used in the specification and the claims, indicate and are used to mean one and / or any combination of the associated listed items.
[0022] As used in the specification and claims, the term "if" can be interpreted as meaning "when" or "once" or "in response to a determination" or "in response to detecting" depending on the context. Similarly, the phrase "if it is determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once it is determined" or "in response to a determination" or "once [the described condition or event] is detected" or "in response to detecting [the described condition or event]" depending on the context.
[0023] As described in the above background section, the ranking stage includes a coarse ranking stage and a fine ranking stage. The personalized model of the coarse ranking stage is responsible for quickly ranking the recalled commodities under less time-consuming requirements, and then inputting the top n commodities in the sequence to the fine ranking stage, while the fine ranking stage takes more time to sort the commodities more personalized and detailed and push them to the user. Therefore, the personalized ranking effect of the coarse ranking is generally worse than that of the fine ranking. It can be understood that learning better ranking methods by the model is collectively referred to as Learning To Rank ("LTR"), and common methods of LTR include Point-Wise, Pair-Wise and List-Wise. The cascade learning described above can also be divided into Point-Wise, Pair-Wise and List-Wise, but there are not many research methods for List-Wise cascade learning, and the method for calculating List-Wise cascade learning is also relatively complex.
[0024] In deep learning, the idea of distillation learning is to let the student model output the ranking score or the embedding vector of the intermediate output of the model to approximate the ranking score or the embedding vector of the intermediate output of the teacher model as much as possible. In cascade learning, the coarse ranking model corresponds to the student model, and the fine ranking model corresponds to the teacher model. The loss function used in this learning process is generally Point-Wise loss. However, in the Point-Wise learning method, the coarse ranking model does not learn the relationship between the order of multiple commodities in the fine ranking sequence. In the field of recommendation algorithms, there have been attempts to use Pair-Wise Loss to let the coarse ranking model learn the order of the fine ranking sequence. For example, when a fine ranking request returns a list, a commodity is randomly sampled from the top n commodities in the sequence head, and then a commodity is randomly sampled from the sequence tail, and the two commodities are combined into a pair of samples to let the model learn to pull apart the distance between the two samples. Therefore, this method can only learn the relationship between the two samples, and cannot learn the relationship between all the commodities in the fine ranking list.
[0025] By extending the above problem to the List-Wise Loss learning method, this learning method often uses, for example, RankNet, LambdaMART, etc., but this List-Wise method always needs to first convert the ranking into a pair of samples through sampling, and then compare the scores of each pair to calculate the loss. This method needs to compare the fine ranking scores of any two commodities in a ranking, making the calculation too complex.
[0026] In addition, in most cascade learning algorithms, the distillation learning method uses the loss used in distillation as an auxiliary loss to train the final output of the model. This method cannot well control the degree of influence of the knowledge distilled from the teacher model to the student model on the final output of the student model, especially in the case of multiple business objectives (such as click rate, add car rate, and conversion rate) in the recommendation field. In this case, if the fine ranking model (i.e. the teacher model) is a conversion rate model, and the coarse ranking model (i.e. the student model) is a click rate model, using distillation learning to let the coarse ranking model learn the fine ranking sequence of the fine ranking model will cause the optimization goal of the coarse ranking to change from click rate improvement to both click rate improvement and conversion rate improvement, and it is difficult to control whether the final coarse ranking should be more inclined to optimize the click rate effect or the conversion rate effect.
[0027] Based on this, the application proposes a scheme for training a coarse ranking scoring model for e-commerce goods. By calculating a simple minimax distillation loss, the consistency of coarse ranking and fine ranking scoring is improved, so that the coarse ranking model can better fit the fine ranking result and improve the scoring accuracy of the coarse ranking model. By distilling the fine ranking knowledge into an independent model target tower and combining the multi-target scoring result to optimize the coarse ranking scoring model, the degree of knowledge transfer of the fine ranking model is controllable.
[0028] The specific embodiments of the application will be described in detail below with reference to the accompanying drawings.
[0029] Figure 1 is an exemplary flow block diagram illustrating a method 100 for training a coarse ranking scoring model for e-commerce goods according to an embodiment of the application. In one implementation scenario, the coarse ranking scoring model can include a multi-target scoring module and a coarse ranking distillation module. The aforementioned multi-target scoring module can include a multilayer perceptron (MLP), a click rate tower module, and a conversion rate tower module. In some embodiments, the MLP, the click rate tower module, and the conversion rate tower module in the multi-target scoring module have loaded trained parameter weights. That is, the multi-target scoring module has been trained. When training the coarse ranking scoring model, only the coarse ranking distillation module needs to be optimized to make the degree of knowledge transfer of the fine ranking model controllable.
[0030] As shown in Figure 1 , at step S101, training samples are generated from the e-commerce platform recommended goods list returned according to user requests. In one embodiment, first, it is determined whether there is a target task under the goods in the e-commerce platform recommended goods list returned according to the user request, and then the training data set is constructed according to the determination result of whether there is a target task under the goods in the e-commerce platform recommended goods list returned according to the user request to generate training samples. In some implementation scenarios, the aforementioned target task can include but is not limited to click operation and conversion operation, and the aforementioned conversion operation can include but is not limited to adding a car operation and / or placing an order operation.
[0031] As an example, assume that trace_id is used to identify the request ID of the recommended goods list result once the user browses the e-commerce platform APP, and the jth request is named G j . Then, let the sample index in the jth request G j be i, and the sample in the jth request G j with a click operation, an add-to-car operation, or a place order operation be named P ji . That is, when generating training samples, first, it is determined whether there is a sample P ji, and further constructing training data set according to the judgment result to generate training sample. In an embodiment, first sample data and second sample data can be constructed according to the judgment result of whether the goods under the target task exist in the goods list recommended by the e-commerce platform returned according to the user request, and further constructing training data set based on the first sample data and the second sample data to generate training sample. That is, judging whether the sample P ji comes from the goods list recommended by the e-commerce platform returned according to the user request, and then putting the first sample data and the second sample data of all request groups into the training data set listwise as training sample.
[0032] In an implementation scenario, in response to the existence of the goods under the target task in the goods list recommended by the e-commerce platform returned according to the user request, the goods ranking score and the goods features of the goods except the goods under the target task are aggregated into a plurality of first list features, and the first list features are added to the goods under the target task to form the first sample data. In response to the non-existence of the goods under the target task in the goods list recommended by the e-commerce platform returned according to the user request, the goods ranking score and the goods features of the goods except the goods with the largest ranking score are aggregated into a plurality of second list features, and the second list features are added to the goods with the largest ranking score to form the second sample data.
[0033] Specifically, for the existence of the sample P ji in the goods list returned according to the user request, the goods ranking score and the goods features of other samples except the sample P ji are aggregated into a plurality of first list features list1, and the first list features list1 are added to the sample P ji to form the first sample data. For the non-existence of the sample P ji in the goods list returned according to the user request, that is, the existence of the sample exposed only without the click operation / add-to-cart operation / order operation, which is recorded as N ji . In this scenario, the sample S ji with the largest ranking score is determined from the sample N j . The goods ranking score and the goods features of other samples except the sample S j are aggregated into a plurality of second list features list2, and the second list features list2 are added to the sample S j to form the second sample data. Further, the first list features list1 and the second list features list2 are put into the training data set listwise to generate the training sample.
[0034] In some embodiments, assuming that the label of the exposure sample is 0, the label of the click sample is 1, the label of the add-car sample is 2, and the label of the order sample is 3, the above operation of generating the training set can also be understood as requesting G j The current sample s k All label values outside are less than s k The sample of label [s1, s2, ..., s m ] product ID, online ranking score are aggregated into a product list [goods1, goods2, ..., goods m ] and the refined score list [score1, score2, ..., score m ] is added to sample s as a new sequence feature k Then, the list features corresponding to each request are synthesized into training samples. j If there is no sample with label value greater than 0, then the product list sequence features and the refined ranking list sequence features are added to the sample S j and currently requesting G j Only return the exposure sample S with the highest ranking score j Based on this, by using the maximum fine-ranked sample in the request as the cascade positive sample for training, the utilization rate of the exposure sample and the cascade learning effect are improved.
[0035] Based on the training samples generated above, at step S102, based on the training samples, a multi-objective scoring module is used to perform a scoring operation for sorting the products under the target task to obtain a multi-objective scoring result. Specifically, in one embodiment, a multi-layer perceptron is used to extract features from the training samples, and a scoring operation for sorting the products under the click operation and the conversion operation is performed respectively through the click rate tower module and the conversion rate tower module to obtain the click rate scoring result and the conversion rate scoring result respectively. That is, the training samples are input into the multi-objective scoring module, and the features are first extracted through the MLP in the multi-objective scoring module, and then the corresponding click rate scoring result and the conversion rate scoring result are output respectively through the click rate tower module and the conversion rate tower module. Among them, the click rate is related to the aforementioned click operation, and the conversion rate is related to the aforementioned add-to-cart operation or order operation. As can be seen from the foregoing, the MLP, click rate tower module and conversion rate tower module in the multi-objective scoring module have been loaded with the trained parameter weights, that is, the baseline model knowledge is loaded using the incremental learning method.
[0036] Then, at step S103, based on the training sample, the scoring knowledge of the fine ranking model for fine ranking of the commodity is distilled using the coarse ranking distillation module, and a min-max distillation loss is calculated. In an implementation scenario, the training sample is first subjected to feature extraction via the aforementioned MLP, and then the scoring knowledge of the fine ranking model for fine ranking of the commodity is distilled via the coarse ranking distillation module, and a min-max distillation loss is calculated. Specifically, in one embodiment, a fine ranking score of the fine ranking model for fine ranking of the commodity under the training sample is obtained, a coarse ranking score of the coarse ranking distillation module for fine ranking of the commodity under the training sample is output, and the coarse ranking score is normalized to obtain a normalized coarse ranking score, so as to calculate the min-max distillation loss according to the fine ranking score and the normalized coarse ranking score. That is, the fine ranking score and the coarse ranking score are obtained by the fine ranking model and the coarse ranking distillation module respectively, and then the min-max distillation loss is calculated by the normalized coarse ranking score and the fine ranking score.
[0037] More specifically, in one embodiment, the maximum fine ranking score and the minimum fine ranking score are determined according to the fine ranking score, and the commodity index values corresponding to the maximum fine ranking score and the minimum fine ranking score are determined, the normalized coarse ranking score of the maximum fine ranking commodity and the normalized coarse ranking score of the minimum fine ranking commodity are determined based on the commodity index values corresponding to the maximum fine ranking score and the minimum fine ranking score respectively, and then the min-max distillation loss is calculated based on the normalized coarse ranking score of the maximum fine ranking commodity and the normalized coarse ranking score of the minimum fine ranking commodity. In one implementation scenario, the min-max distillation loss is calculated based on the negative logarithm of the normalized coarse ranking score of the maximum fine ranking commodity and the normalized coarse ranking score of the minimum fine ranking commodity. That is, the corresponding commodities corresponding to the minimum fine ranking score and the maximum fine ranking score are found, the normalized coarse ranking scores of the corresponding commodities corresponding to the minimum fine ranking score and the maximum fine ranking score are found, and then the min-max distillation loss is calculated by the negative logarithm of the normalized coarse ranking score of the commodity corresponding to the maximum fine ranking score and the normalized coarse ranking score of the commodity corresponding to the minimum fine ranking score.
[0038] In one exemplary scenario, it is assumed that the batch size of a training sample is B, the exposure commodity ID sequence feature of the i-th sample in a batch is g i , and the exposure commodity ID sequence corresponds to a fine ranking score sequence z i . Wherein, the feature length of g i and z i is N, indicating that there are N exposure commodities and corresponding N fine ranking scores. It is assumed that the predicted coarse ranking score of the exposure commodity list of the i-th sample by the coarse ranking model is y i , and the length of y i is N. In this scenario, let the index of the maximum fine ranking score in the fine ranking of the i-th sample (that is, the index of the most exposure commodity with the highest fine ranking score in the fine ranking) be k = argmax(z i), the index of the exposure commodity with the minimum fine ranking score in the fine ranking is j = argmin(z i ). Thus, the aforementioned minimax distillation loss can be expressed by the following formula:
[0039]
[0040] wherein Loss represents the minimax distillation loss, softmax(y i ) j represents the normalized coarse ranking score of the commodity j with the minimum fine ranking score, softmax(y i ) k represents the normalized coarse ranking score of the commodity k with the maximum fine ranking score, the Softmax normalized coarse ranking score represents the probability that the fine ranking model ranks the corresponding commodity at the top of the recommendation list in the predicted commodity list, and -log represents the negative logarithm.
[0041] It can be understood that the minimax distillation loss of the embodiments of the present application is equivalent to maximizing the coarse ranking prediction score of the commodity corresponding to the maximum fine ranking score in the ranking score returned by the first fine ranking request or minimizing the negative logarithm form of the coarse ranking prediction score of the commodity with the maximum fine ranking score, so as to minimize the coarse ranking prediction score of the commodity with the minimum fine ranking score. Based on the minimax distillation loss, the gap between the coarse ranking prediction score of the commodity with the maximum fine ranking score, the coarse ranking prediction score of the commodity with the minimum fine ranking score, and the coarse ranking prediction score of other commodities can be widened. By using the scheme of the present application, the calculation complexity of the minimax distillation loss of a sample is O(N), and N is the length of the scoring sequence of a List-Wise sample.
[0042] Compared with existing methods (such as LambdaMART, RankNet, etc.), the minimax distillation loss of the embodiments of the present application does not need to sample the sample, so that the calculation is simpler. In addition, the minimax distillation loss of the embodiments of the present application is only related to the fine ranking score, and is independent of the size of the fine ranking score. This avoids the problem that in a recommendation system, due to the serious imbalance of sample categories, the model scoring approaches 0, the distance of the fine ranking scores of different commodities and the loss also approach 0, causing the model to fail to learn the knowledge of the scoring sequence of different commodities in the same list, improves the consistency of the coarse ranking and fine ranking scoring, so that the coarse ranking model can better fit the fine ranking result, and improves the scoring accuracy of the coarse ranking model.
[0043] Furthermore, at step S104, the final rough ranking result of the rough ranking model for the rough ranking of products is optimized based on the multi-objective scoring results and the minimum-maximum distillation loss, thereby training the rough ranking scoring model for e-commerce products. In one implementation scenario, weight coefficients are set for the click-through rate scoring results, the conversion rate scoring results, and the predicted rough ranking results output by the rough ranking distillation module adjusted based on the minimum-maximum distillation loss, and then the corresponding weight coefficients are adjusted to optimize the final rough ranking result of the rough ranking model for the rough ranking of products, thereby training the rough ranking scoring model for e-commerce products.
[0044] For example, suppose the click-through rate score result is recorded as o1, the conversion rate score result is recorded as o2, and the predicted rough ranking result output by the rough ranking distillation module based on the minimum-maximum distillation loss adjustment is recorded as o3, and the corresponding weight coefficients are recorded as w1, w2, and w3 respectively. In an implementation scenario, the final rough ranking result of the rough ranking of products by the rough ranking scoring model can be obtained based on w1*o1+w2*o2+w3*o3. During the training process, only the above-mentioned minimum-maximum distillation loss can be used to fine-tune the rough ranking distillation module to distill the knowledge of the fine ranking into the rough ranking distillation module, while the click-through rate tower module and the conversion rate tower module remain unchanged. Furthermore, by adjusting the aforementioned weight coefficients w1, w2, and w3, the influence of the output of the rough ranking distillation module on the final score of the rough ranking scoring model is determined, thereby adjusting the influence of the knowledge of the fine ranking model on the rough ranking output effect.
[0045] In some embodiments, assuming that the refined ranking model is a conversion rate model, the final score of the coarse ranking scoring model trained based on the embodiments of the present application can better improve the conversion rate. Assuming that the refined ranking model is a click-through rate model, the final score of the coarse ranking scoring model trained based on the embodiments of the present application can better improve the click-through rate effect. Thus, without affecting the existing coarse ranking knowledge, the target capabilities learned from refined ranking can be easily transferred to the coarse ranking model, thereby improving the accuracy of the coarse ranking scoring model.
[0046] Combined with the above description, it can be seen that the embodiment of the present application calculates a simple minimum-maximum distillation loss in the distillation learning of the scoring knowledge of the fine ranking model for the fine ranking of goods in the coarse ranking distillation module, and can calculate the loss quickly without traversing each pair of Pair-Wise samples, thereby reducing the computational complexity. The minimum-maximum distillation loss calculated by the embodiment of the present application can avoid the problem that the coarse ranking model cannot learn too much knowledge due to the serious imbalance of samples in the recommendation system, resulting in the fine ranking model score or the fine ranking score being very small, and avoids the problem that when the fine ranking score is too small, the score gap between different goods in the same List-Wise sample also becomes very small, which is not conducive to the coarse ranking model learning the order and gap of the scores of different goods, so that the coarse ranking model can better fit the fine ranking ranking result and improve the accuracy of the coarse ranking scoring model.
[0047] Further, by using the incremental learning and multi-target modeling (e.g., the multi-target scoring module described above), and introducing a coarse ranking distillation module and corresponding weight coefficients dedicated to distillation learning, the embodiment of the present application can easily superimpose the knowledge of the fine ranking model into the coarse ranking model without affecting the existing knowledge of the coarse ranking model, by only using the above minimum-maximum distillation loss to fine-tune the coarse ranking distillation module, so that the degree of knowledge transfer of the fine ranking model is controllable. In addition, the embodiment of the present application also adds fine ranking maximum commodity samples in the training samples to improve the utilization rate of exposure samples and the cascading learning effect.
[0048] Figure 2 is an exemplary flow chart showing the generation of training samples according to the embodiment of the present application. It should be understood that, Figure 2 is one embodiment of step S101 in the above Figure 1 , so the description made in the above Figure 1 also applies to Figure 2 .
[0049] As shown in Figure 2 , at step S201, the list of recommended commodities of the e-commerce platform according to the user request is returned. That is, the original sample is obtained. Then, at step S202, it is determined whether there is a commodity under the target task in the list of recommended commodities of the e-commerce platform returned according to the user request. As described above, the target task may, for example, be a click operation, a car operation, and an order operation, etc. Based on whether there is a commodity under the target task in the above-mentioned list of commodities, a training data set can be constructed to generate training samples.
[0050] Specifically, when there is a commodity under the target task in the list of recommended commodities of the e-commerce platform returned according to the user request, at step S203, the fine ranking score and the commodity features of the commodities other than the commodity under the target task are aggregated into a plurality of first list features, which are added to the commodity under the target task to form a first sample data. For example, taking the target task as a car as an example, when there is a car commodity in the list of commodities, the fine ranking score and the commodity features of the samples of, for example, clicks other than the car can be aggregated into a plurality of list features, which are added to the car commodity to form a first sample data.
[0051] When the list of products recommended by the e-commerce platform returned according to the user's request contains products under the target task, at step S204, the multiple second list features formed by aggregating the fine ranking scores and product features of products other than the product with the highest fine ranking score are added to the product with the highest fine ranking score to form the second sample data. It can be understood that when there are no products under the target task (such as click / add to cart / place an order) in the product list, there are only exposed products. In this scenario, the multiple second list features formed by aggregating the fine ranking scores and features of samples other than the sample with the highest fine ranking score are added to the sample with the highest fine ranking score to form the second sample data.
[0052] Furthermore, at step S205, the first sample data and the second sample data are added to the training data set to generate training samples at step S206. Based on this, by also using the maximum precise ranking score in the request as the cascade positive sample training set, the utilization rate of the exposure samples and the cascade learning effect can be improved.
[0053] Figure 3 FIG. 1 is an exemplary flow chart showing the overall process of generating training samples according to an embodiment of the present application. Figure 3 As shown in , at step S301, the operation of generating training samples is started. In the process of generating training samples, first at step S302, the product list is returned based on the user request, that is, the original sample is obtained. Then, at step S303, the aforementioned original samples are aggregated according to the request ID trace_id feature to obtain n groups ("group"). Among them, trace_id is recorded as the request ID for identifying the recommended product list result when the user browses the e-commerce platform APP. At step S304, each of the aforementioned groups is traversed, and the j-th request in the current group is named G j .
[0054] Based on the G in the current group j At step S305, determine the G j Is there a click / add to cart / order sample? j When there is a click / add cart / order sample, at step S306, traverse the G j The positive sample of each click / add cart / order is recorded as P ji , i represents the jth request G j In step S307, the sample index is removed from the current positive sample P. ji The refined ranking scores and product features of other exposure samples are aggregated into multiple first list features list1, and the multiple first list features list1 are added to the positive sample P ji , to form the first sample data in the context of this application.
[0055] G in the current group j In the absence of click / add-to-cart / order samples, i.e. only exposure samples are present, at step S308, G in the current group is searched j The exposure sample with the largest inner ranking score is recorded as S j , and at step S309, the fine ranking score and item features of other samples except sample S j are aggregated into a plurality of second list features list2, and the second list features list2 are added to sample S j to form second sample data.
[0056] After the first and second sample data are formed, at step S310, samples S j and samples P ji of all groups are added to a final sample list (i.e. a training data set) to obtain final listwise samples at step S311 to generate training samples. Finally, at step S312, the operation of generating training samples is ended.
[0057] In the embodiments of the present application, based on the generated training samples, a multi-target scoring module can be used to perform scoring operations on the ranking of items under target tasks to obtain multi-target scoring results, and a coarse ranking distillation module can be used to distill the scoring knowledge of the fine ranking model for the fine ranking of items, and calculate a minimax distillation loss, and then optimize the final coarse ranking result of the coarse ranking of items by the coarse ranking scoring model. Next, the calculation of the minimax distillation loss will be described first in combination with Figure 4 .
[0058] Figure 4 is an exemplary flow chart illustrating the calculation of the minimax distillation loss according to the embodiments of the present application. It should be understood that Figure 4 is one embodiment of step S103 in Figure 1 , and thus the description made in Figure 1 also applies to Figure 4 .
[0059] As shown in Figure 4 , at step S401, the coarse ranking score for the fine ranking of items under the training samples is output by the coarse ranking distillation module. In one implementation scenario, the training samples are first subjected to feature extraction by the MLP in the multi-target scoring module, and then the coarse ranking score for the fine ranking of items is output by the coarse ranking distillation module. At step S402, the coarse ranking scores of the samples are normalized to obtain corresponding normalized coarse ranking scores. In some embodiments, the normalized coarse ranking scores can be output by, for example, softmax, such as softmax(y i ), where y idenotes the coarse ranking score.
[0060] Further, at step S403, the fine ranking model is used to obtain the fine ranking scores of the training samples. Based on the obtained fine ranking scores, at step S404, the maximum fine ranking score and the minimum fine ranking score are determined, and at step S405, the respective corresponding item index values of the maximum fine ranking score and the minimum fine ranking score are determined. As an example, the corresponding item index value of the maximum fine ranking score is k = argmax(z i ), and the corresponding item index value of the minimum fine ranking score is j = argmin(z i ), where z i denotes the fine ranking score.
[0061] Then, at step S405, based on the respective corresponding item index values of the maximum fine ranking score and the minimum fine ranking score, the normalized coarse ranking score of the maximum fine ranking score item and the normalized coarse ranking score of the minimum fine ranking score item are determined, which are denoted as softmax(y i ) k and softmax(y i ) j respectively. Finally, at step S406, the minimum-maximum distillation loss is calculated based on the negative logarithm of the normalized coarse ranking score of the maximum fine ranking score item and the normalized coarse ranking score of the minimum fine ranking score item. Specifically, the minimum-maximum distillation loss of the present embodiment can be calculated according to the above formula (1). The calculation of the aforementioned minimum-maximum distillation loss will be described in detail again below in conjunction with a specific example in Figure 5 .
[0062] Figure 5 is an exemplary schematic diagram illustrating the calculation of the minimum-maximum distillation loss according to the present embodiment. As shown in Figure 5 , it is assumed that the coarse ranking model is used to perform coarse ranking scoring on items in the same order, and the coarse ranking scores of the items are obtained. For example, for item 1, item 2 and item 3, their respective corresponding coarse ranking scores are 0.001, 0.5 and 0.06. Then, the respective coarse ranking scores are normalized to obtain the corresponding normalized coarse ranking scores. For example, as shown in the figure, the respective corresponding normalized coarse ranking scores of item 1, item 2 and item 3 are 0.269, 0.444 and 0.286.
[0063] Assuming that the same order of goods is scored by using the fine ranking model, for the goods item1, item2 and item3, their respective fine ranking scores are 0.1, 0.001 and 0.4. According to the fine ranking scores, the maximum fine ranking score 0.4 and the minimum fine ranking score 0.001 can be determined, which correspond to the goods index values of 3 and 2 respectively. Then, based on the respective goods index values corresponding to the maximum fine ranking score and the minimum fine ranking score, the normalized coarse ranking score of the maximum fine ranking score goods and the normalized coarse ranking score of the minimum fine ranking score goods are determined, for example, the normalized coarse ranking score corresponding to the maximum fine ranking score goods item3 is determined to be 0.286 and the normalized coarse ranking score of the minimum fine ranking score goods item2 is determined to be 0.444. Substituting them into the above formula (1), i.e. Loss = 0.444 - log(0.286).
[0064] The minimum maximum distillation loss calculation of the embodiments of the present application is simple and efficient, and the gap between the coarse ranking prediction score of the goods with the maximum fine ranking score, the coarse ranking prediction score of the goods with the minimum fine ranking score and the coarse ranking prediction scores of other goods is widened. Therefore, the problem that the fine ranking model scoring or fine ranking score is very small due to the serious imbalance of samples in the recommendation system, and the coarse ranking model cannot learn too much knowledge can be avoided, and the problem that when the fine ranking score is too small, the scoring gap of different goods in the same List-Wise sample is also very small, which is not conducive to the coarse ranking model to learn the knowledge of the order and gap of different goods scoring can be avoided, so that the coarse ranking model can better fit the fine ranking ranking result, and the consistency of coarse ranking and fine ranking scoring is improved.
[0065] In the embodiments of the present application, the fine ranking knowledge is also distilled to an independent model target tower (i.e. the coarse ranking distillation module in the embodiments of the present application), and the coarse ranking scoring model is optimized in combination with the multi-target scoring results (such as click rate and conversion rate), so that the degree of knowledge of the migrated fine ranking model is controllable, for example Figure 6 as shown.
[0066] Figure 6 is an exemplary schematic diagram showing the overall coarse ranking scoring model 600 according to the embodiments of the present application. As Figure 6As shown in , the rough ranking scoring model 600 may include a multi-objective scoring module (which includes an MLP 601, a click-through rate tower module 602, and a conversion rate tower module 603) and a rough ranking distillation module 604. In an implementation scenario, by inputting the above-mentioned training sample 605 into the rough ranking scoring model, the training sample 605 is first subjected to feature extraction via the MLP 601, and the scoring operation of ranking the products under the click operation and the conversion operation is performed respectively via the click-through rate tower module 602 and the conversion rate tower module 603 to obtain the click-through rate scoring result and the conversion rate scoring result accordingly. At the same time, after the aforementioned MLP 601 extracts the features of the training sample 605, it also outputs the predicted rough ranking result (i.e., the rough ranking score) for the precise ranking score of the products under the training sample through the rough ranking distillation module 604.
[0067] Furthermore, the aforementioned click-through rate score results, conversion rate score results, and predicted rough ranking results are combined to obtain a final rough ranking result 606 of the rough ranking of products by the rough ranking scoring model. Specifically, weight coefficients w1, w2, and w3 can be set for the aforementioned click-through rate score results, conversion rate score results, and predicted rough ranking results, and then the final rough ranking result 606 is obtained based on the weighted summation. For example, assuming that the click-through rate score result is recorded as o1, the conversion rate score result is recorded as o2, and the predicted rough ranking result is recorded as o3, the final rough ranking result of the rough ranking of products by the rough ranking scoring model can be obtained based on w1*o1+w2*o2+w3*o3.
[0068] It's understandable that the multi-objective scoring module already has pre-trained parameter weights loaded into the MLP, CTR, and conversion rate tower modules, using incremental learning to load baseline model knowledge. During training, simply adjusting the rough distillation module based on the aforementioned maximum and minimum distillation losses, or adjusting the corresponding weight coefficients w1, w2, and w3 to determine their impact on the final rough ranking results, can easily incorporate knowledge from the refined ranking model into the rough ranking model without significantly impacting the existing rough ranking knowledge, thereby improving the accuracy of the rough ranking scoring model.
[0069] Figure 7 FIG. 7 is an exemplary flow chart showing a method 700 for roughly ranking and scoring e-commerce products according to an embodiment of the present application. Figure 7 As shown in , at step S701, the list of products recommended by the e-commerce platform returned by the user request is obtained. At step S702, based on the product list, the trained coarse ranking scoring model is used to perform coarse ranking scoring to obtain a coarse ranking scoring result. That is, based on the product list returned by the user request, the coarse ranking scoring model trained above is used to perform coarse ranking scoring to obtain a coarse ranking scoring result. For more details about the coarse ranking scoring model training, please refer to the above Figures 1-6 The description of , this application will not be repeated here.
[0070] Figure 8 is an exemplary structural block diagram of the device 800 for training a coarse ranking model for e-commerce commodities according to an embodiment of the present application. As shown in Figure 8 the device 800 of the present application can include a processor 801 and a memory 802, wherein the processor 801 and the memory 802 communicate with each other through a bus. The memory 802 stores program instructions for training a coarse ranking model for e-commerce commodities, which when executed by the processor 801, cause the implementation of the method steps described above in conjunction with the drawings: generating a training sample from a list of recommended commodities of an e-commerce platform returned according to a user request; based on the training sample, performing a scoring operation for ranking of commodities under a target task using the multi-objective scoring module to obtain a multi-objective scoring result; based on the training sample, using the coarse ranking distillation module to perform distillation learning on scoring knowledge of a fine ranking model for fine ranking of commodities, and calculating a minimax distillation loss; and optimizing a final coarse ranking result of the coarse ranking model for coarse ranking of commodities according to the multi-objective scoring result and the minimax distillation loss to train the coarse ranking model for e-commerce commodities.
[0071] In some embodiments, the above-mentioned memory 802 can also store program instructions for coarse ranking of e-commerce commodities, which when executed by the processor 801, cause the implementation of the method steps described above in conjunction with the drawings: obtaining a list of recommended commodities of an e-commerce platform returned according to a user request; and based on the list of commodities, using the trained coarse ranking model to perform coarse ranking to obtain a coarse ranking result.
[0072] According to the above description in conjunction with the drawings, those skilled in the art can also understand that the embodiments of the present application can also be implemented by software programs. Therefore, the present application also provides a non-transitory computer readable storage medium. The computer readable storage medium has stored thereon computer readable instructions for training a coarse ranking model for e-commerce commodities or coarse ranking of e-commerce commodities, which when executed by one or more processors, implement the embodiments of the present application described above in conjunction with the drawings. Figure 1 The described method for training a coarse ranking model for e-commerce commodities or the method for coarse ranking of e-commerce commodities. Figure 7 The described method for coarse ranking of e-commerce commodities.
[0073] Those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary universal hardware platform, and of course can also be implemented by hardware, through the above description of the embodiments. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0074] It should be noted that although the operations of the method of the present application are described in a specific order in the accompanying drawings, this does not require or imply that the operations must be performed in that specific order, or that all of the shown operations must be performed to achieve the desired result. On the contrary, the steps depicted in the flowchart can change the order of execution. Additionally or alternatively, some steps can be omitted, combined into one step, and / or divided into multiple steps.
[0075] The foregoing can be better understood in light of the following clauses:
[0076] Clause A1, a method for training a coarse ranking scoring model for e-commerce goods, wherein the coarse ranking scoring model comprises a multi-objective scoring module and a coarse ranking distillation module, and the method comprises: generating training samples according to a list of recommended goods of an e-commerce platform returned in response to a user request; performing scoring operations for ranking goods under a target task using the multi-objective scoring module based on the training samples to obtain multi-objective scoring results; performing distillation learning for scoring knowledge of a fine ranking model for fine ranking of goods using the coarse ranking distillation module based on the training samples, and calculating a minimax distillation loss; and optimizing a final coarse ranking result of the coarse ranking scoring model for coarse ranking of goods according to the multi-objective scoring results and the minimax distillation loss to train the coarse ranking scoring model for e-commerce goods.
[0077] Clause A2, the method of clause A1, wherein generating training samples according to a list of recommended goods of an e-commerce platform returned in response to a user request comprises: determining whether there are goods under the target task in the list of recommended goods of the e-commerce platform returned in response to the user request; and constructing a training data set according to the determination result of whether there are goods under the target task in the list of recommended goods of the e-commerce platform returned in response to the user request to generate the training samples.
[0078] Clause A3, the method of clause A2, wherein constructing the training data set according to the determination of whether the target task exists in the list of recommended commodities of the e-commerce platform returned according to the user request to generate the training sample comprises: constructing first sample data and second sample data according to the determination of whether the target task exists in the list of recommended commodities of the e-commerce platform returned according to the user request; and constructing the training data set based on the first sample data and the second sample data to generate the training sample.
[0079] Clause A4, the method of clause A3, wherein constructing the first sample data and the second sample data according to the determination of whether the target task exists in the list of recommended commodities of the e-commerce platform returned according to the user request comprises: in response to the target task existing in the list of recommended commodities of the e-commerce platform returned according to the user request, aggregating the commodity fine ranking score and the commodity feature of the commodities other than the commodity under the target task into a plurality of first list features, and adding the first list features to the commodity under the target task to form the first sample data; and in response to the target task not existing in the list of recommended commodities of the e-commerce platform returned according to the user request, aggregating the commodity fine ranking score and the commodity feature of the commodities other than the commodity with the largest fine ranking score into a plurality of second list features, and adding the second list features to the commodity with the largest fine ranking score to form the second sample data.
[0080] Clause A5, the method of any one of clauses A1-A4, wherein the target task at least includes a click operation and a conversion operation, and the conversion operation at least includes a car adding operation and / or an order placing operation.
[0081] Clause A6, the method of clause A5, wherein the multi-target scoring module comprises a multi-layer perception machine, a click rate tower module, and a conversion rate tower module, and based on the training sample, performing a scoring operation on the ordering of the commodity under the target task using the multi-target scoring module to obtain a multi-target scoring result comprises: using the multi-layer perception machine to perform feature extraction on the training sample, and performing a scoring operation on the ordering of the commodity under the click operation and the conversion operation via the click rate tower module and the conversion rate tower module, respectively, to correspondingly obtain a click rate scoring result and a conversion rate scoring result.
[0082] Clause A7, the method of clause A6, wherein distilling learning of scoring knowledge of a fine ranking model for fine ranking of items using the coarse ranking distillation module and calculating a min-max distillation loss comprises: obtaining fine ranking scores of the fine ranking model scoring for fine ranking of items under the training sample; outputting coarse ranking scores of the coarse ranking distillation module scoring for fine ranking of items under the training sample and normalizing the coarse ranking scores to obtain normalized coarse ranking scores; and calculating the min-max distillation loss according to the fine ranking scores and the normalized coarse ranking scores.
[0083] Clause A8, the method of clause A7, wherein calculating the min-max distillation loss according to the fine ranking scores and the normalized coarse ranking scores comprises: determining a maximum fine ranking score and a minimum fine ranking score according to the fine ranking scores and determining item index values corresponding to the maximum fine ranking score and the minimum fine ranking score respectively; determining a normalized coarse ranking score of a maximum fine ranking item and a normalized coarse ranking score of a minimum fine ranking item based on the item index values corresponding to the maximum fine ranking score and the minimum fine ranking score respectively; and calculating the min-max distillation loss based on the normalized coarse ranking score of the maximum fine ranking item and the normalized coarse ranking score of the minimum fine ranking item.
[0084] Clause A9, the method of clause A8, wherein calculating the min-max distillation loss based on the normalized coarse ranking score of the maximum fine ranking item and the normalized coarse ranking score of the minimum fine ranking item comprises: calculating the min-max distillation loss based on a negative logarithm of the normalized coarse ranking score of the maximum fine ranking item and the normalized coarse ranking score of the minimum fine ranking item.
[0085] Clause A10, the method of clause A6, wherein optimizing a final coarse ranking result of coarse ranking of items by the coarse ranking scoring model according to the multi-objective scoring results and the min-max distillation loss to train the coarse ranking scoring model for e-commerce items comprises: setting weight coefficients to the click rate scoring results, the conversion rate scoring results, and the predicted coarse ranking result output by the coarse ranking distillation module adjusted based on the min-max distillation loss respectively; and adjusting the respective weight coefficients to optimize the final coarse ranking result of coarse ranking of items by the coarse ranking scoring model to train the coarse ranking scoring model for e-commerce items.
[0086] Clause A11, a device for training a coarse ranking scoring model for e-commerce items, comprising: a processor; and a memory having stored thereon computer instructions for training a coarse ranking scoring model for e-commerce items, which, when executed by the processor, cause the device to implement the method according to clauses A1-A10.
[0087] Clause A12, a method for coarse ranking e-commerce items, comprising: obtaining a list of recommended items returned by an e-commerce platform in response to a user request; and performing coarse ranking on the list of items using a trained coarse ranking model according to any one of clauses A1-A10 to obtain a coarse ranking result.
[0088] Clause A13, an apparatus for coarse ranking e-commerce items, comprising: a processor; and a memory having computer instructions for coarse ranking e-commerce items stored thereon, which when executed by the processor, cause the implementation of the method according to clause A12.
[0089] Clause A14, a non-transitory computer-readable storage medium having stored thereon computer program instructions for training a coarse ranking model for e-commerce items, which when executed by one or more processors, cause the implementation of the method according to any one of clauses A1-A10; or having stored thereon computer program instructions for coarse ranking e-commerce items, which when executed by one or more processors, cause the implementation of the method according to clause A12.
[0090] It should be understood that when the terms "first", "second", "third", and "fourth" and the like are used in the specification and claims of this application, these terms are used only to distinguish different objects, and are not used to describe a particular order. The terms "include" and "contain" used in the specification and claims of this application indicate the presence of the described features, whole, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, whole, steps, operations, elements, components and / or sets thereof.
[0091] It should also be understood that the terms used in this application are only for the purpose of describing specific embodiments, and are not intended to limit the application. As used in the specification and claims of this application, the singular forms "a", "an" and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should be further understood that the term "and / or" used in the specification and claims of this application means any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.
[0092] Although the embodiments of the present application are as described above, the above description is only for the purpose of understanding the embodiments of the present application, and is not intended to limit the scope and application of the present application. Any person skilled in the art of the technology described in the present application can make any modification and change in the form and details of the implementation without departing from the spirit and scope of the present application, but the patent protection scope of the present application shall be subject to the scope defined by the appended claims.
[0093] In addition, the collection and acquisition of various data in this application comply with relevant legal provisions and are authorized by the data provider. Any organization or individual that needs to obtain external data should obtain authorization and ensure data security in accordance with the law, and shall not illegally collect, use, process, transmit, sell, provide or disclose unauthorized or unprotected data.
Claims
1. A method for training a coarse-rank scoring model for e-commerce products, wherein the coarse-rank scoring model includes a multi-objective scoring module and a coarse-rank distillation module, and the method comprises: Generate training samples based on the product list recommended by the e-commerce platform returned by the user request; Based on the training samples, using the multi-objective scoring module to perform a scoring operation on the order of the products under the target task to obtain a multi-objective scoring result; Based on the training samples, using the rough ranking distillation module to perform distillation learning on the scoring knowledge of the fine ranking model for the fine ranking of products, and calculating the minimum and maximum distillation loss; and The final rough ranking result of the rough ranking of products by the rough ranking scoring model is optimized according to the multi-objective scoring result and the minimum maximum distillation loss, so as to train the rough ranking scoring model for e-commerce products.
2. The method according to claim 1, wherein generating training samples based on the product list recommended by the e-commerce platform returned by the user request comprises: Determine whether the product under the target task exists in the list of products recommended by the e-commerce platform returned according to the user request; as well as A training data set is constructed based on a judgment result of whether the product under the target task exists in the list of products recommended by the e-commerce platform returned by the user request to generate the training samples.
3. The method according to claim 2, wherein constructing a training dataset based on a determination result of whether the product under the target task exists in the list of products recommended by the e-commerce platform returned by the user request to generate the training samples comprises: Constructing first sample data and second sample data based on a determination result of whether the product under the target task exists in the list of products recommended by the e-commerce platform returned by the user request; as well as The training data set is constructed based on the first sample data and the second sample data to generate the training samples.
4. The method according to claim 3, wherein constructing the first sample data and the second sample data based on a determination result of whether the product under the target task exists in the list of products recommended by the e-commerce platform returned by the user request comprises: In response to the presence of the target task item in the list of items recommended by the e-commerce platform returned in response to the user request, aggregating the refined ranking scores and item features of items other than the target task item into a plurality of first list features, and adding the first list features to the target task item to form the first sample data; as well as In response to the fact that the product under the target task does not exist in the list of products recommended by the e-commerce platform returned according to the user request, the product ranking scores and product features except for the product with the largest ranking score are aggregated into multiple second list features, and the second list features are added to the product with the largest ranking score to form the second sample data.
5. The method according to claim 1, wherein the target task includes at least a click operation and a conversion operation, and the conversion operation includes at least an add-to-cart operation and / or an order operation, the multi-objective scoring module includes a multi-layer perceptron, a click-through rate tower module, and a conversion rate tower module, and based on the training sample, using the multi-objective scoring module to perform a scoring operation to sort products under the target task to obtain a multi-objective scoring result comprises: The multi-layer perceptron is used to extract features from the training samples, and the click-through rate tower module and the conversion rate tower module are used to perform scoring operations to sort the products under the click operation and the conversion operation, respectively, so as to obtain corresponding click-through rate scoring results and conversion rate scoring results.
6. The method according to claim 5, wherein using the rough ranking distillation module to perform distillation learning on the scoring knowledge of the refined ranking model for the refined ranking of products, and calculating the minimum and maximum distillation loss comprises: Obtaining a ranking score for the product ranking by the ranking model under the training sample; Outputting a rough ranking score for the refined ranking of products under the training sample using the rough ranking distillation module, and normalizing the rough ranking score to obtain a normalized rough ranking score; and The minimum-maximum distillation loss is calculated according to the refined discharge fraction and the normalized rough discharge fraction.
7. The method according to claim 6, wherein calculating the minimum-maximum distillation loss according to the fine separation score and the normalized rough separation score comprises: Determining a maximum and a minimum precise ranking score according to the precise ranking scores, and determining commodity index values corresponding to the maximum and the minimum precise ranking scores; Determine a normalized coarse ranking score of the product with the maximum fine ranking score and a normalized coarse ranking score of the product with the minimum fine ranking score based on the product index values corresponding to the maximum fine ranking score and the minimum fine ranking score respectively; as well as The minimum-maximum distillation loss is calculated based on the normalized rough fraction of the largest polished commodity and the normalized rough fraction of the smallest polished commodity.
8. The method according to claim 5, wherein optimizing the final rough ranking result of the rough ranking scoring model for rough ranking of products based on the multi-objective scoring result and the minimum-maximum distillation loss to train the rough ranking scoring model for e-commerce products comprises: Setting weight coefficients for the click-through rate scoring result, the conversion rate scoring result, and the predicted rough sorting result output by the rough sorting distillation module based on the minimum-maximum distillation loss adjustment, respectively; and The corresponding weight coefficients are adjusted to optimize the final rough ranking result of the rough ranking of products by the rough ranking scoring model, so as to train the rough ranking scoring model for e-commerce products.
9. A device for training a coarse ranking and scoring model for e-commerce products, comprising: processor; as well as A memory storing computer instructions for training a coarse ranking and scoring model for e-commerce products, wherein when the computer instructions are executed by a processor, the device implements the method according to any one of claims 1-8.
10. A non-transitory computer-readable storage medium storing computer program instructions for training a coarse ranking scoring model for e-commerce products, wherein the computer program instructions, when executed by one or more processors, implement the method according to any one of claims 1-8.