Method and apparatus for predicting click-through rate based on product representation

By extracting unique and common features in the multimodal information of the product and processing it in combination with deep neural networks, the problem of inaccurate click-through rate prediction in the existing technology is solved, and the accuracy of product representation and the benefits of e-commerce platforms are improved.

CN113641889BActive Publication Date: 2025-06-27ALIBABA GROUP HOLDING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010392970.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-05-11
Publication Date
2025-06-27
Estimated Expiration
2040-05-11

AI Technical Summary

Technical Problem

The existing multimodal product representation cannot effectively reflect the product characteristics, resulting in inaccurate click-through rate prediction.

Method used

By obtaining the user behavior of the target user and the multimodal information of the target product, the first model and the second model respectively extract unique and common features, and processing them in combination with the deep neural network, the probability of the target user clicking on the target product is obtained.

Benefits of technology

It improves the accuracy and efficiency of product representation, enhances the accuracy of click-through rate prediction, and thus improves the revenue of e-commerce platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113641889B_ABST
    Figure CN113641889B_ABST
Patent Text Reader

Abstract

The present invention discloses a click-through rate prediction method and device based on commodity representation, which relates to the technical field of data processing. The main purpose of the present invention is to perform more accurate commodity representation on multimodal commodities to improve the accuracy of click-through rate prediction. The main technical solution of the present invention is as follows: obtaining a set of multimodal information corresponding to the user behavior of a target user and a target commodity respectively; using a first model to determine the unique features corresponding to each group of multimodal information, where the unique features include features existing in one modality and have dynamic weights for different modalities; using a second model to determine the common features corresponding to each group of multimodal information, where the common features include features co-existing in multiple modalities; determining the user representation of the target user and the commodity representation of the target commodity respectively according to the unique features and the common features; and using a deep neural network to process the user representation and the commodity representation to obtain the probability that the target user clicks on the target commodity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular, to a click-through rate prediction method and device based on product representation. Background Art

[0002] Existing large-scale e-commerce portals provide billions of products for hundreds of millions of users through mobile applications and PC websites. To obtain a better user experience and business results, more accurate and effective representation of products helps to recommend products that users are interested in and like to them, thereby increasing the click-through rate of the recommended products. The more accurate the prediction of the click-through rate of products, the more effectively the revenue of e-commerce and platforms can be increased.

[0003] Currently, products in e-commerce usually have multiple heterogeneous modality representations, such as modalities of product names, images, titles, and statistical features. For product representation, features are extracted from different modalities according to preset modality weights, regardless of the type of product. However, an ideal product representation should be able to dynamically weigh different modalities according to different types of products, thereby emphasizing more useful modality signals. It can be seen that the existing multi-modal product representation cannot effectively reflect product characteristics, resulting in inaccurate click-through rate prediction based on this product representation. Summary of the Invention

[0004] In view of the above problems, the present invention proposes a click-through rate prediction method and device based on product representation, and the main purpose is to perform more accurate product representation on multi-modal products to improve the accuracy of click-through rate prediction.

[0005] To achieve the above object, the present invention mainly provides the following technical solutions:

[0006] On the one hand, the present invention provides a click-through rate prediction method based on product representation, specifically including:

[0007] Obtain a set of multi-modal information corresponding to the user behavior of the target user and the target product respectively;

[0008] Use a first model to determine the unique features corresponding to each set of multi-modal information, where the unique features include features existing in one modality and have dynamic weights for different modalities;

[0009] Use a second model to determine the common features corresponding to each set of multi-modal information, where the common features include features that coexist in multiple modalities;

[0010] Determine the user representation of the target user and the product representation of the target product according to the unique features and the common features respectively;

[0011] Process the user representation and the product representation using a deep neural network to obtain the probability that the target user clicks on the target product.

[0012] On the other hand, the present invention provides a click-through rate prediction device based on product representation, specifically including:

[0013] An acquisition unit for acquiring a set of multimodal information corresponding to the user behavior of the target user and the target product respectively;

[0014] A first determination unit for using a first model to determine the unique features corresponding to each group of multimodal information acquired by the acquisition unit, where the unique features include features existing in one modality and have dynamic weights for different modalities;

[0015] A second determination unit for using a second model to determine the common features corresponding to each group of multimodal information acquired by the acquisition unit, where the common features include features that coexist in multiple modalities;

[0016] A representation unit for respectively determining the user representation of the target user and the product representation of the target product according to the unique features determined by the first determination unit and the common features determined by the second determination unit;

[0017] A prediction unit for processing the user representation and the product representation obtained by the representation unit using a deep neural network to obtain the probability that the target user clicks on the target product.

[0018] On the other hand, the present invention provides a processor, and the processor is used to run a program. When the program runs, it executes the above-mentioned click-through rate prediction method based on product representation.

[0019] By means of the above technical solution, a click-through rate prediction method and device based on product representation provided by the present invention predict the probability that a product is clicked by optimizing the representation of the product. For the optimization of the product representation, the embodiments of the present invention propose to further split the multimodal information of the product, select the unique features and common features between different modalities from the multimodal information, and use these features to re-represent the product. Moreover, when determining the unique features, dynamic weights will be assigned to the corresponding unique features according to different modalities of the product to highlight the modal features that the product is more concerned by users. And by extracting the common features, the processing of redundant information in different modalities can be reduced, and the representation efficiency and accuracy of the product can be improved. Therefore, based on this optimized product representation, the click-through rate of the preset product can be increased, thereby increasing the revenue of the e-commerce and the platform.

[0020] The above description is only an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are given below. Description of the Drawings

[0021] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0022] Figure 1 Shows a flowchart of a click-through rate prediction method based on commodity representation proposed by an embodiment of the present invention;

[0023] Figure 2 Shows a flowchart of obtaining unique features of multimodal commodities in an embodiment of the present invention;

[0024] Figure 3 Shows a flowchart of obtaining common features of multimodal commodities in an embodiment of the present invention;

[0025] Figure 4 Shows a flowchart of the user representation process for a target user in an embodiment of the present invention;

[0026] Figure 5 Shows a block diagram of a click-through rate prediction device based on commodity representation proposed by an embodiment of the present invention;

[0027] Figure 6 Shows a block diagram of another click-through rate prediction device based on commodity representation proposed by an embodiment of the present invention. Detailed Embodiments

[0028] The exemplary embodiments of the present invention will be described in more detail below with reference to the drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.

[0029] In a click-through rate prediction method based on product representation proposed in an embodiment of the present invention, a new product representation method is proposed. Specifically for products provided in e-commerce, they are usually represented by multiple heterogeneous modalities. Existing product representations only extract features from different modalities of products according to fixed weights for representation. However, for different types of products, their corresponding modality weights should be different, making the existing product representation methods unable to effectively adapt to complex types of products. Therefore, it is also impossible to accurately predict the click-through rate of products using existing product representations. For this reason, the following embodiments of the present invention propose a click-through rate prediction method based on product representation to improve the representation accuracy of complex types of products and thus improve the accuracy of click-through rate prediction. The specific steps of the embodiments of the present invention are as Figure 1 shown and include:

[0030] Step 101, obtain a set of multi-modal information corresponding to the user behavior of the target user and the target product respectively.

[0031] Among them, the content of click-through rate prediction in the embodiments of the present invention is to estimate the probability that the target user clicks on the target product. The target product and the target user can be preset or extracted from the database according to a preset policy. In this step, the target product and the target user can be regarded as being given in advance. The user behavior of the target user refers to the behavior performed by the target user within a specified historical time period, and the number of behaviors is not limited. This behavior includes operating behaviors on a certain product, such as clicking to view the product, purchasing the product, adding the product to the shopping cart, etc. The information included in this user behavior at least includes the type of operating behavior and the time when the behavior is executed, etc.

[0032] For multi-modal information, for the target product, it can be understood as a set of feature information representing the product from different dimensions. For e-commerce products, common modalities include: product name (ID), product picture, product title, statistical information, etc. And there are also multiple features in each modality. For example, for the product name modality, it can include the name of the store where the product is located, the name of the brand to which the product belongs, the name number of the product, etc. It can be seen that each modality can be represented as a set of multi-dimensional vectors. For the user behavior of the target user, it means that a certain user behavior of the target user corresponds to a set of multi-modal information, and this set of multi-modal information can be regarded as the multi-modal information of the product corresponding to this user behavior.

[0033] Step 102, use the first model to determine the unique features corresponding to each group of multi-modal information.

[0034] Among them, the unique feature includes the unique feature existing in one modality, and this unique feature has a dynamic weight for different modalities. The unique feature can be a feature unique to a certain modality, or a feature unique to a certain type of modality, or a feature that is most distinguishable from other modalities in a certain modality.

[0035] The first model in this step can learn the dynamic weights corresponding to different modalities of the commodity based on the multi-modal fusion network, and output a unique feature representation that can represent the commodity based on the dynamic weights and the modality unique features in each modality of the commodity. That is to say, when processing a set of multi-modal information corresponding to a certain commodity, the first model can identify the modality unique features existing in different modalities, and obtain a unique feature representation for the commodity according to the learned dynamic weights corresponding to different modalities.

[0036] In this step, the first model determines the unique feature representations corresponding to each commodity for the target commodity and the behavior commodities corresponding to the user behavior of the target user one by one. The unique features of the commodity can reflect the complementarity of different modalities of the commodity.

[0037] Step 103: Use the second model to determine the common features corresponding to each group of multi-modal information.

[0038] Among them, the common feature includes the features that coexist in multiple modalities, and what the common features of the commodity reflect is the redundancy of different modalities of the commodity.

[0039] The second model in this step learns the modality common features existing in different modalities of the commodity based on the multi-modal adversarial network, and represents these modality common features jointly, and outputs the common feature representation corresponding to the commodity.

[0040] It can be seen that the second model in this step has the opposite processing function to the first model in the above step, that is, it outputs the common feature representation and the unique feature representation corresponding to the multi-modal information. It should be noted that the first model and the second model are two independent models. Therefore, there is no sequential relationship in the processing logic between step 103 and step 102.

[0041] Step 104: Determine the user representation of the target user and the commodity representation of the target commodity according to the unique features and the common features respectively.

[0042] Among them, the commodity representation of the target commodity is obtained by extracting and combining the unique features and the common features of the multi-modal information of the commodity.

[0043] The user representation of the target user is based on the multimodal information corresponding to each user behavior of the target user. The multimodal information is used to represent the user behavior as a combination of unique features and common features, and then the user representation of the target user is comprehensively obtained according to the representations of each user behavior.

[0044] Step 105: Use a deep neural network to process the user representation and the product representation to obtain the probability that the target user clicks on the target product.

[0045] Since the user representation of the target user is obtained based on multiple user behaviors of the target user, it can represent the preference bias of the target user for products. Using this user representation and the product representation of the target product, the probability that the target user clicks on the target product can be predicted through a deep neural network. Among them, using a deep neural network for click-through rate prediction is a commonly used prediction method in the art. In the embodiments of the present invention, when using a deep neural network for click-through rate prediction, the main difference lies in the different representation methods of the target user and the target product output. Therefore, the specific processing process of the user representation and the product representation in the deep neural network will not be described further.

[0046] Through the description of the above embodiments, a click-through rate prediction method based on product representation provided by the present invention mainly provides a new representation method for the target user and the target product. By splitting the multimodal information of the product into unique features and common features and using these features for product representation, while the user representation of the target user is represented as a combination of unique features and common features based on the products corresponding to the existing user behaviors of the target user. Since the unique features and common features can reflect the complementarity and redundancy of multiple modalities, this representation method can obtain a more discriminative product representation, thereby making the user representation of the target user and the product representation of the target product more accurate and enabling more accurate prediction of the click-through rate of the target product.

[0047] Further, the embodiments of the present invention will specifically describe the implementation of each step of the above Figure 1 described method one by one.

[0048] First, for the above step 101, that is, the specific implementation of obtaining a set of multimodal information corresponding to the user behavior of the target user and the target product is divided into: obtaining multimodal information for the user behavior of the target user, and obtaining multimodal information for the target product.

[0049] To obtain multimodal information for the user behavior of the target user, it is necessary to first determine the corresponding behavior product according to the user behavior. For example, if the user behavior of the target user is to click and view a product, then the multimodal information corresponding to the product is the multimodal information corresponding to the user behavior, that is, the user behavior is converted into the corresponding behavior product.

[0050] After that, according to the preset modality, the feature information of each modality corresponding to the behavioral commodity and the target commodity is obtained. The preset modality for e-commerce commodities can be divided into modalities such as name, picture, title, and statistical information. According to these modalities, the corresponding feature information of different modalities is extracted from the commodity information. Generally, the collected feature information is a high-dimensional vector, which is not conducive to subsequent processing. Therefore, it is necessary to map these feature information to low-dimensional vectors based on a preset strategy to obtain a set of multi-modal information corresponding to the behavioral commodity and the target commodity respectively. Among them, the preset strategy is a corresponding dimensionality reduction mapping strategy set according to different modalities. In this regard, this embodiment is not particularly limited.

[0051] After the dimensionality reduction mapping process, each modality can be regarded as a sequence of vectors, and each vector in it represents a feature of that modality. Thus, each commodity can be represented as a combination of sequences of vectors corresponding to multiple modalities.

[0052] For Figure 1 For step 102 shown in the figure, use the first model to determine the unique features corresponding to each group of multi-modal information, that is, extract the unique features from the multi-modal information for each of the target commodity and each user behavior of the target user. The specific process of determining the unique features of a certain commodity is as Figure 2 shown, including:

[0053] Step 201: Process the multi-modal information using the uniqueness projection layer to obtain the modality-unique features in each modality.

[0054] Among them, the uniqueness projection layer is used to find the modality-unique features in each modality. The modality-unique feature is a feature that only exists in one modality. The uniqueness projection layer is a non-linear projection layer. And because the uniqueness of commodities in different modalities is different, the uniqueness projection layer is determined according to different modalities, that is, to determine the modality-unique features in different modalities, the uniqueness projection layer corresponding to that modality needs to be applied.

[0055] Step 202: Obtain the category of the commodity corresponding to the multi-modal information.

[0056] Among them, the category of the commodity can be determined by obtaining the specified category field in the commodity information, or by analyzing the commodity information through a learning network for commodity categories.

[0057] Step 203: Determine the modality weights corresponding to multiple modalities of the commodity according to the category and the modality-unique features.

[0058] This step can be to learn the modal weights corresponding to multiple modalities of different commodities through a multi-modal attention fusion network. Among them, the attention of different modalities can be represented by the tanh function. By learning the modal uniqueness features of existing commodities and the corresponding commodity categories, the modal weights of different modalities in different categories of commodities are determined.

[0059] Step 204: Determine the unique features of the corresponding commodity by using the modal weights and the modal unique features.

[0060] Since the representation of the unique features of a commodity is based on all the modal unique features, in this embodiment, one way to represent the unique features can be to perform a weighted sum of each modal unique feature according to the modal weights to represent the unique features of the commodity. In addition to this method, it is also possible to screen the modal weights, and then perform a weighted sum of the modal unique features with modal weights to represent the unique features of the commodity. The specific method of specifically using the modal weights and the modal unique features to represent the unique features of the commodity is not specifically limited in this step.

[0061] Furthermore, in order to improve the accuracy of the representation of the unique features of the commodity, the embodiment of the present invention can be realized by improving the determination of the modal unique features. For this purpose, an auxiliary discriminator can be set up to filter the modal unique features obtained in step 201. This auxiliary discriminator is essentially a trained multi-class classifier, and its classes correspond to different modalities. That is to say, the greater the probability that a modal unique feature is recognized by this auxiliary discriminator, the stronger the uniqueness of this modal unique feature.

[0062] For Figure 1 the step 103 shown, use the second model to determine the common features corresponding to each group of multi-modal information, that is, extract the common features from the multi-modal information for each user behavior of the target commodity and the target user respectively. The specific process of determining the common features for a certain commodity is as Figure 3 shown, including:

[0063] Step 301: Process the multi-modal information by using the commonality projection layer to obtain the modal common features existing in each modality.

[0064] Among them, this commonality projection layer corresponds to the above-mentioned uniqueness projection layer, and both are used to determine a type of modal feature from a certain modal information. The difference is that the uniqueness projection layer needs to select the corresponding projection layer according to different modalities, while the commonality projection layer is applicable to different modalities.

[0065] Step 302: Use the first discriminator to determine the candidate common features that cannot be recognized from the modal common features.

[0066] According to the content of step 103, the second model is constructed based on a multi-modal adversarial network. For a conventional adversarial network, it consists of a generator and a discriminator. The generator is used to generate data by machine with the aim of "fooling" the discriminator, while the discriminator is used to determine whether the data generated by the input generator is real data or fabricated fake data. However, this kind of adversarial network is used to seek an effective common subspace across different modalities. However, this cross-modal adversarial approach focuses more on one-to-one confrontation, while the products in e-commerce involve multiple heterogeneous modalities, that is, it is difficult to achieve strict one-to-one confrontation for the features in the modalities. In addition, different modal features contain different degrees of shared latent features. Therefore, in order to perform adversarial training on the products in e-commerce, the embodiment of the present invention proposes a multi-modal adversarial network with a dual discriminator to realize the expansion of one-to-one confrontation to a multi-modal scenario. Among them, the first discriminator is used to identify the features that may come from the common subspace of the modalities, that is, the candidate common features. The first discriminator can be a multi-class discriminator, and its classification corresponds to different modalities. That is, the input of the first discriminator is the modal common features obtained in step 301, and the modal common features that the discriminator cannot identify or the recognition probability is lower than a certain threshold are output as candidate common features. In addition, the first discriminator is also used to input the candidate common features it outputs into the second discriminator to confuse the second discriminator. And the second discriminator is used to drive the knowledge transfer between modalities in order to learn the common latent subspace across multiple modalities. The description of the second discriminator will be described in the next step.

[0067] Step 303: Use the second discriminator to perform combined recognition on the candidate common features and the modal common features to determine the most unrecognizable candidate common feature in each modality.

[0068] The second discriminator is also a multi-class discriminator, and the difference between it and the first discriminator lies in the different discrimination strategies used. The purpose of the second discriminator is to reduce the dispersion between different modalities. Its input is the combination of the candidate common features and the modal common features, so as to confuse the candidate common features. It can be understood that the role of the second discriminator is to judge whether the candidate common features in different modalities are the same feature, and find a most unrecognizable candidate common feature from each modality, that is, the recognition probability of this candidate common feature is the lowest in this modality.

[0069] Step 304: Combine the most unrecognizable candidate common features in each modality to obtain the common features corresponding to the multi-modal information.

[0070] Among them, the combination of candidate common features is obtained by weighted summation according to the weights of each modality. In this way, the common features of a certain commodity are composed of the most common features within different modalities selected from its multi-modal information according to the corresponding modality weights. The common features are not single or multiple features in a certain modality, but are jointly composed of features in multiple different modalities.

[0071] For Figure 1 In step 104 shown, determining the user representation of the target user and the commodity representation of the target commodity according to the unique features and common features can be specifically divided into two steps. One is the commodity representation of the target commodity, and the other is the user representation of the target user.

[0072] Among them, the commodity representation is obtained by using the multi-modal information of the target commodity to obtain the corresponding unique features and common features, and then splicing and combining these features to obtain the commodity representation. For the user representation of the target user, it is obtained based on the user behavior corresponding to the target user. Each user behavior has a corresponding behavior commodity. Therefore, the representation of each user behavior is equivalent to the commodity representation of the behavior commodity. Therefore, the user representation of the target user is obtained by synthesizing the representations of multiple user behaviors. However, considering the temporal characteristics of the user behavior sequence in e-commerce, a GRU network can be used to model the dependence relationship between user behaviors. Among them, the input of the GRU network is the commodities sorted by time, that is, the behavior commodities corresponding to the user behaviors sorted by time. However, the hidden state output by the conventional GRU network can only capture the dependence relationship between user behaviors, and cannot obtain important commodities in a group of behavior sequences to represent different interests of the target user, so it cannot better represent the target user. For this reason, the embodiment of the present invention optimizes and improves the conventional GRU network, takes the behavior attributes of user behaviors as input to reflect the importance of different behaviors to the target user, and uses the attention network to configure corresponding weights for the hidden state output by the network and the commodity representation of the target commodity. Finally, the weighted summation of the hidden states corresponding to different user behaviors is used to represent the user representation of the target user relative to the target commodity. The specific steps of this process are as Figure 4 shown, including:

[0073] Step 401: Combine the unique features and common features of the behavior commodities corresponding to the user behaviors of the target user to obtain the behavior commodity representation.

[0074] This behavior commodity representation is the same as the commodity representation of the above-mentioned target commodity. The extraction of unique features and common features can refer to the above Figure 2 And Figure 3 shown content. It will not be elaborated here.

[0075] Step 402: Determine the behavior state of the target user based on the behavior attributes of the user behavior and the behavior commodity representation.

[0076] In this step, the optimized GRU network model is used to process the behavior commodity representation and the corresponding behavior attributes obtained in the previous step as the input of the model. Among them, the behavior attributes include information such as the type and time of the user behavior. The behavior type includes click, purchase, etc., which is used to measure the importance of the user behavior and assign a higher weight to the user behavior. In the e-commerce field, generally, the importance of the purchase behavior is higher than that of the click behavior because the purchase behavior represents that the user has a very high interest in the commodity and has been converted into a transaction amount. And the behavior time can sort the user behaviors according to time. Since the hidden state output by the model also depends on the hidden state of the previous user behavior, in addition to the behavior attributes and the behavior commodity representation, the input of the model also needs to input the hidden state of the previous user behavior. This hidden state represents the behavior state of the target user when performing this user behavior. This behavior state is used to represent the degree of preference of the target user for performing the user behavior on the behavior commodity. The measurement of this degree of preference is obtained by comparing all the user behaviors of the target user.

[0077] Step 403: Use the behavior state to match the relevance between the behavior commodity and the target commodity, and determine the weight corresponding to the behavior commodity.

[0078] Specifically, this step can be implemented through an attention network model. The input of the model is the behavior states corresponding to each user behavior obtained in the previous step and the commodity representation of the target commodity, and the output is the weights corresponding to each user behavior. The processing logic of the model is to focus on the relevance between the behavior commodity corresponding to the user behavior and the target commodity. The higher the relevance of the user behavior, the greater the weight assigned to it. In this way, when representing the target user, the user behavior related to the target commodity can be more prominent, so as to improve the accuracy of predicting the click-through rate of the target commodity in the future.

[0079] Step 404: Perform a weighted sum of all the user behaviors of the target user according to the weights to obtain the user representation of the target user.

[0080] Specifically, the user representation of the target user is the result obtained by performing a weighted sum of the behavior states corresponding to all user behaviors combined with the assigned weights.

[0081] Through the above Figure 4 description, it can be seen that in the embodiment of the present invention, when determining the user representation of the target user, first, Figure 2 、 Figure 3The network model in it processes the multi-modal information of behavioral commodities into unique features and common features to obtain the representation of behavioral commodities, and then uses the optimized GRU network model to transform the representation of behavioral commodities into the behavioral state of the target user. At the same time, according to the obtained behavioral state and the commodity representation of the target commodity, the attention network model is used to allocate corresponding weights to determine the behavior of the target user related to the target commodity. Finally, the user representation of the target user relative to the target commodity is obtained by weighted summation of the behavioral states.

[0082] Based on the above detailed description of Figure 1 each step, it can be seen that the embodiment of the present invention mainly proposes a new way of representing commodities for the multi-modal information of commodities, taking into account the complementarity and redundancy between various modal information, representing commodities with a combination of unique features and common features, so as to obtain a more discriminative commodity representation. At the same time, through the optimization and improvement of the GRU network, the user representation of the target user can also consider more the behavioral states corresponding to the behavioral commodities related to the target commodity, making the correlation between the user representation and the commodity representation stronger, thereby improving the accuracy of predicting the probability that the target user clicks on the target commodity.

[0083] Furthermore, it should be noted that the method provided by the embodiment of the present invention is not only applicable to e-commerce fields such as Taobao, Tmall, and Kaola, but also applicable to travel scenarios such as Fliggy, such as destination hotels, air tickets, travel habits, etc.; it can also be applicable to local life scenarios such as Ele.me, such as food delivery preferences, drug purchases, etc.; it can also be applicable to logistics scenarios such as Cainiao, such as price, speed, etc.; and so on.

[0084] Further, as an implementation of the above Figure 1-4 shown method, the embodiment of the present invention provides a click-through rate prediction device based on commodity representation. The main purpose of this device is to perform more accurate commodity representation on multi-modal commodities to improve the accuracy of click-through rate prediction. For ease of reading, the details of the foregoing method embodiments will not be described one by one in the device embodiments, but it should be clear that the device in this embodiment can correspondingly implement all the contents of the foregoing method embodiments. The device is as Figure 5 shown and specifically includes:

[0085] An acquisition unit 51, configured to acquire a set of multi-modal information corresponding to the user behavior of the target user and the target commodity respectively;

[0086] A first determination unit 52, configured to use a first model to determine the unique features corresponding to each group of multi-modal information acquired by the acquisition unit 51. The unique features include features existing in one modality and have dynamic weights for different modalities;

[0087] A second determination unit 53, configured to determine, by using a second model, common features corresponding to each group of multimodal information acquired by the acquisition unit 51, where the common features include features that coexist in multiple modalities;

[0088] A representation unit 54, configured to determine a user representation of the target user and a product representation of the target product respectively according to the unique features determined by the first determination unit 52 and the common features determined by the second determination unit 53;

[0089] A prediction unit 55, configured to process the user representation and the product representation obtained by the representation unit 54 by using a deep neural network to obtain a probability that the target user clicks on the target product.

[0090] Further, as Figure 6 shown, the acquisition unit 51 includes:

[0091] A determination module 511, configured to determine a behavior product corresponding to the user behavior;

[0092] An acquisition module 512, configured to acquire, according to a preset modality, feature information of each modality corresponding to the behavior product determined by the determination module 511 and the target product;

[0093] A generation module 513, configured to map the feature information of the acquisition module 512 to a low-dimensional vector based on a preset policy to obtain a group of multimodal information corresponding to each of the behavior product and the target product.

[0094] Further, as Figure 6 shown, the first determination unit 52 includes:

[0095] A mapping module 521, configured to process the multimodal information by using a uniqueness projection layer to obtain modality-unique features in each modality;

[0096] An acquisition module 522, configured to acquire a category of a product corresponding to the multimodal information;

[0097] A determination module 523, configured to determine modality weights corresponding to multiple modalities of the product according to the category obtained by the acquisition module 522 and the modality-unique features obtained by the mapping module 521;

[0098] A generation module 524, configured to determine unique features corresponding to the product by using the modality weights obtained by the determination module 523 and the modality-unique features.

[0099] Further, the generation module 524 is specifically configured to perform weighted summation on the modality-unique features according to the modality weights to obtain the unique features of the product.

[0100] Further, asFigure 6 As shown, the first determination unit 52 further includes:

[0101] A filtering module 525, configured to filter the modality-unique features obtained by the mapping module 521 by using an auxiliary discriminator, so as to obtain the modality-unique features recognizable in each modality.

[0102] Furthermore, as Figure 6 shown, the second determination unit 53 includes:

[0103] A mapping module 531, configured to process the multimodal information by using a commonality projection layer to obtain the modality-common features possessed in each modality;

[0104] A first recognition module 532, configured to determine the candidate common features that cannot be recognized from the modality-common features obtained by the mapping module 531 by using a first discriminator;

[0105] A second recognition module 533, configured to perform combined recognition on the candidate common features obtained by the first recognition module 532 and the modality-common features by using a second discriminator, so as to determine one of the candidate common features that is the most unrecognizable in each modality;

[0106] A generation module 534, configured to combine the candidate common features that are the most unrecognizable in each modality obtained by the second recognition module 533 to obtain the common features corresponding to the multimodal information.

[0107] Furthermore, as Figure 6 shown, the representation unit 54 is specifically configured to combine the unique features corresponding to the target commodity with the common features to obtain the commodity representation; and determine the user representation of the target user by using the behavioral attributes of the user behavior, where the behavioral attributes include the type and time information of the user behavior.

[0108] Furthermore, as Figure 6 shown, the representation unit 54 includes:

[0109] A combination module 541, configured to combine the unique features and the common features of the behavioral commodity corresponding to the user behavior of the target user to obtain the behavioral commodity representation;

[0110] A determination module 542, configured to determine the behavioral state of the target user according to the behavioral attributes of the user behavior and the behavioral commodity representation obtained by the combination module 541, where the behavioral state is used to represent the preference degree of the target user for performing the user behavior on the behavioral commodity;

[0111] A matching module 543, configured to match the relevance between the behavioral commodity and the target commodity by using the behavioral state determined by the determination module 542, and determine the weight corresponding to the behavioral commodity.

[0112] A representation module 544 is configured to perform a weighted sum of all user behaviors of the target user according to the weights obtained by the matching module 543, so as to obtain a user representation of the target user.

[0113] In addition, an embodiment of the present invention further provides a processor, which is configured to run a program, where when the program runs, it executes the click-through rate prediction method based on item representation provided in any one of the above embodiments.

[0114] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0115] It can be understood that the relevant features in the above methods and apparatuses can be referred to each other. In addition, the "first", "second", etc. in the above embodiments are used to distinguish the embodiments, and do not represent the advantages and disadvantages of the embodiments.

[0116] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described in detail here.

[0117] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings herein. The structure required to construct such a system will be apparent from the above description. In addition, the present invention is not directed to any particular programming language. It should be understood that the content of the present invention described herein can be implemented using various programming languages, and the description of a particular language above is for the purpose of disclosing the best mode of the present invention.

[0118] In addition, the memory may include non-permanent memory in a computer-readable medium, random access memory (RAM), and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.

[0119] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0120] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0121] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0122] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0123] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0124] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.

[0125] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0126] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.

[0127] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0128] The above are only embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A click-through rate prediction method based on commodity representation, the method comprising: Obtaining a set of multimodal information respectively corresponding to the user behavior of a target user and a target commodity; Using a first model to determine unique features corresponding to each group of multimodal information, the unique features including features existing in one modality and having dynamic weights for different modalities, including: Processing the multimodal information by using a uniqueness projection layer to obtain modality-unique features in each modality, wherein the uniqueness projection layer is a non-linear projection layer, and there is a corresponding uniqueness projection layer in different modalities; Obtaining the category of the commodity corresponding to the multimodal information; Determining modality weights corresponding to multiple modalities of the commodity according to the category and the modality-unique features; Determining unique features corresponding to the commodity by using the modality weights and the modality-unique features; Using a second model to determine common features corresponding to each group of multimodal information, the common features including features co-existing in multiple modalities, including: Processing the multimodal information by using a commonality projection layer to obtain modality-common features in each modality; Using a first discriminator to determine candidate common features with an identification probability lower than a certain threshold from the modality-common features, wherein the first discriminator is a multi-class discriminator, and its classification corresponds to different modalities; Using a second discriminator to perform combined identification on the candidate common features and the modality-common features to determine one of the candidate common features with the lowest identified probability in each modality, the second discriminator being a multi-class discriminator; Combining one of the candidate common features with the lowest identified probability in each modality to obtain the common features corresponding to the multimodal information; Respectively determining a user representation of the target user and a commodity representation of the target commodity according to the unique features and the common features, the user representation being used to represent the preference of the target user for commodities; Processing the user representation and the commodity representation by using a deep neural network to obtain the probability that the target user clicks on the target commodity.

2. The method according to claim 1, wherein The obtaining a set of multimodal information respectively corresponding to the user behavior of a target user and a target commodity includes: Determining corresponding behavior commodities according to the user behavior; Obtaining feature information of each modality corresponding to the behavior commodities and the target commodity according to a preset modality; Mapping the feature information into low-dimensional vectors based on a preset strategy to obtain a set of multimodal information respectively corresponding to the behavior commodities and the target commodity.

3. The method according to claim 1, wherein Determining unique features corresponding to the commodity by using the modality weights and the modality-unique features includes: Performing weighted summation on the modality-unique features according to the modality weights to obtain the unique features of the commodity.

4. The method according to claim 1, wherein The method further includes: Filtering the modality-unique features by using an auxiliary discriminator to obtain modality-unique features recognizable in each modality.

5. The method according to claim 1, characterized in that Respectively determining a user representation of the target user and a commodity representation of the target commodity according to the unique features and the common features includes: Combining the unique features and the common features corresponding to the target commodity to obtain the commodity representation; Determine the user representation of the target user by using the behavioral attributes of the user behavior, where the behavioral attributes include the type and time information of the user behavior.

6. The method according to claim 5, wherein Determining the user representation of the target user by using the behavioral attributes of the user behavior includes: Combining the unique features and common features of the behavioral items corresponding to the user behavior of the target user to obtain the representation of the behavioral items; Determine the behavioral state of the target user according to the behavioral attributes of the user behavior and the representation of the behavioral items, where the behavioral state is used to represent the degree of preference of the target user for performing the user behavior on the behavioral items; Match the relevance between the behavioral item and the target item by using the behavioral state to determine the weight corresponding to the behavioral item; Perform weighted summation on all user behaviors of the target user according to the weight to obtain the user representation of the target user.

7. A click-through rate prediction device based on item representation, the device includes: An acquisition unit, configured to acquire a set of multimodal information corresponding to the user behavior of the target user and the target item respectively; A first determination unit, configured to determine the unique features corresponding to each set of multimodal information acquired by the acquisition unit by using a first model, where the unique features include features existing in one modality and have dynamic weights for different modalities. The first determination unit includes: A mapping module, configured to process the multimodal information by using a uniqueness projection layer to obtain the modality-unique features in each modality, where the uniqueness projection layer is a non-linear projection layer, and there is a corresponding uniqueness projection layer in different modalities; An acquisition module, configured to acquire the category of the item corresponding to the multimodal information; A determination module, configured to determine the modality weights corresponding to multiple modalities of the item according to the category obtained by the acquisition module and the modality-unique features obtained by the mapping module; A generation module, configured to determine the unique features corresponding to the item by using the modality weights obtained by the determination module and the modality-unique features; A second determination unit, configured to determine the common features corresponding to each set of multimodal information acquired by the acquisition unit by using a second model, where the common features include features that coexist in multiple modalities. The second determination unit includes: A mapping module, configured to process the multimodal information by using a commonality projection layer to obtain the modality-common features in each modality; A first recognition module, configured to determine candidate common features with an identification probability lower than a certain threshold from the modality-common features obtained by the mapping module by using a first discriminator, where the first discriminator is a multi-class discriminator, and its classification corresponds to different modalities; A second recognition module, configured to perform combined recognition on the candidate common features obtained by the first recognition module and the modality-common features by using a second discriminator to determine one of the candidate common features with the lowest recognized probability in each modality, where the second discriminator is a multi-class discriminator; A generation module, configured to combine the candidate common features with the lowest recognized probability in each modality obtained by the second recognition module to obtain the common features corresponding to the multimodal information; A representation unit, configured to respectively determine a user representation of the target user and a product representation of the target product according to the unique features determined by the first determination unit and the common features determined by the second determination unit, where the user representation is used to represent the preference of the target user for products; A prediction unit, configured to process the user representation and the product representation obtained by the representation unit by using a deep neural network, so as to obtain the probability that the target user clicks on the target product.

8. A processor, characterized in that, The processor is configured to run a program, where when the program runs, it executes the method for predicting click-through rate based on product representation according to any one of claims 1-6.

Citation Information

Patent Citations

  • Image-text cross-modal feature unentanglement method based on depth mutual information constraint

    CN110807122A

  • Click rate prediction method, recommendation method, click rate prediction model, click rate prediction device and equipment

    CN111046294A