A method and device for optimizing a recommendation model

By extending the gain network on the backbone network of the recommendation model, learning user preferences and preference gains, and optimizing recommendation results, the problem of poor recommendation results of existing recommendation models is solved, and higher recommendation accuracy is achieved.

CN119669581BActive Publication Date: 2025-05-06浙江飞猪网络技术有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510199093.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-06
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

When the existing recommendation model recommends similar objects to users, there is a problem of poor recommendation results and it is difficult to accurately meet user needs.

Method used

An optimization method for recommendation model is proposed, by extending the gain network based on the backbone network, learning user preferences and preference gains, and combining user query history and information differences between recommended objects, optimize the recommendation results.

Benefits of technology

By learning user preferences and preference gains, the recommendation model can more accurately recommend similar objects that meet the needs to users and improve recommendation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119669581B_ABST
    Figure CN119669581B_ABST
Patent Text Reader

Abstract

A method for optimizing a recommendation model, wherein the recommendation model is used to recommend similar objects corresponding to objects queried by the user to a user; the recommendation model comprises a backbone network and at least one gain network; wherein the backbone network is used to learn the user's preference for the similar objects based on the user's historical operation behavior data for the similar objects; the gain network is used to learn the preference gain corresponding to the user's preference based on the information difference contained in the historical objects queried by the user and the recommended similar objects corresponding to the historical objects; the preference gain represents the degree of preference of the user's preference; the training loss of the recommendation model comprises the user preference loss of the user's preference learned by the backbone network; and the preference gain loss of the preference gain learned by the gain network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of artificial intelligence technology, and in particular to a method and device for optimizing a recommendation model. Background Art

[0002] In some service platforms with recommendation functions, in order to attract users, a recommendation model is usually deployed in the service platform. When users query the objects provided by the platform to users on the service platform, the recommendation model is used to recommend to users some similar objects corresponding to the objects queried by the users.

[0003] For example, when users search for products on e-commerce platforms, their product needs may be somewhat flexible. Products similar to the one they searched for may also meet their specific needs. Therefore, while users are searching for products on e-commerce platforms, recommendation models deployed on the platforms can be used to recommend similar products to them, potentially increasing their conversion rate for these products.

[0004] However, the recommendation models deployed on the service platform usually make personalized recommendations based on user preferences learned from the user's historical operation data on the recommended objects. In actual applications, there may be problems with poor recommendation effects.

[0005] Therefore, how to optimize the existing recommendation model so that it can more accurately recommend objects that can meet user needs has become an urgent problem that needs to be solved in the industry. Summary of the Invention

[0006] This specification provides a method for optimizing a recommendation model, wherein the recommendation model is used to recommend similar objects corresponding to an object queried by the user to a user; the recommendation model includes a backbone network and at least one gain network; wherein the backbone network is used to learn the user's preference for the similar objects; the gain network is used to learn a preference gain corresponding to the user's preference based on the information difference between the historical objects queried by the user and the recommended similar objects corresponding to the historical objects; the preference gain represents the degree of preference of the user's preference; the method includes:

[0007] Obtaining a training sample set for training the recommendation model; wherein the training sample set includes a plurality of historical operation behavior data samples; the historical operation behavior data samples include attribute information of historical objects queried by the user, attribute information of recommended similar objects corresponding to the historical objects, and a label indicating whether the user has performed an operation on the similar objects;

[0008] Inputting the historical operation behavior data samples in the training sample set into the backbone network for supervised learning and training, and calculating the user preference loss of the user preference learned by the backbone network based on the output result of the backbone network; and calculating the information difference between the attribute information of the historical object and the attribute information of the similar object contained in the historical operation behavior data sample, and further inputting the calculated information difference into the gain network for supervised learning and training, and calculating the preference gain loss of the preference gain learned by the gain network based on the output result of the gain network;

[0009] The user preference loss and the preference gain loss are used as the training losses of the recommendation model, and the model parameters contained in the recommendation model are adjusted to complete the learning and training of the recommendation model; wherein the trained recommendation model is used to recommend similar objects corresponding to the objects queried by the user to the user based on the user preferences learned by the backbone network and the preference gains learned by the gain network.

[0010] This specification also proposes an object recommendation method based on a recommendation model, wherein the recommendation model is used to recommend similar objects corresponding to an object queried by the user to a user; the recommendation model includes a backbone network and at least one gain network; wherein the backbone network is used to learn the user's preference for the similar objects; the gain network is used to learn a preference gain corresponding to the user's preference based on the information difference between the historical objects queried by the user and the recommended similar objects corresponding to the historical objects; the preference gain represents the degree of preference of the user's preference; the method includes:

[0011] Obtaining an object queried by a user, and constructing a recommendation data sample based on the object queried by the user and similar objects corresponding to the object to be recommended; wherein the recommendation data sample includes attribute information of the object queried by the user and attribute information of the similar objects to be recommended;

[0012] Inputting the recommended data sample into the backbone network for learning and calculation, and obtaining the user preference of the user for the similar object output by the backbone network; and calculating the information difference between the attribute information of the object queried by the user contained in the recommended data sample and the attribute information of the similar object to be recommended, further inputting the calculated information difference into the gain network for learning and calculation, and obtaining the preference gain corresponding to the user preference output by the gain network;

[0013] Based on the user preference learned by the backbone network and the preference gain learned by the gain network, it is determined whether to recommend the similar object to the user.

[0014] This specification also proposes a recommendation model based on a neural network, wherein the recommendation model is used to recommend similar objects corresponding to the object queried by the user to the user, including:

[0015] A backbone network, configured to learn user preferences for the similar objects based on historical operation behavior data of the users for the similar objects;

[0016] at least one gain network, configured to learn a preference gain corresponding to a user preference based on information differences between a historical object queried by the user and a recommended similar object corresponding to the historical object; the preference gain indicating a degree of preference for the user preference;

[0017] The training loss of the recommendation model includes the user preference loss of the user preference learned by the backbone network; and the preference gain loss of the preference gain learned by the gain network.

[0018] In the above embodiment, in the scenario where similar objects corresponding to the objects queried by the user are recommended to the user based on the recommendation model, at least one gain network is expanded on the basis of the backbone network of the recommendation model, and a preference gain that can reflect the user's preference for similar objects is learned in the gain network based on the information difference contained in the historical objects queried by the user and the recommended similar objects corresponding to the historical objects. This allows the recommendation model to recommend similar objects that are more likely to be operated by the user based on the user preferences learned by the backbone network and the preference gains learned by the gain network, thereby improving the recommendation accuracy of the recommendation model. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of this specification, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0020] Figure 1 This is a schematic diagram of the architecture of an Internet service system shown in an embodiment of this specification;

[0021] Figure 2 This is a flow chart of a method for optimizing a recommendation model shown in an embodiment of this specification;

[0022] Figure 3 This is a network architecture diagram of a recommended model shown in an embodiment of this specification;

[0023] Figure 4is a network architecture diagram of another recommended model shown in an embodiment of this specification;

[0024] Figure 5 This is a network architecture diagram of a backbone network of a recommended model shown in an embodiment of this specification;

[0025] Figure 6 is a network architecture diagram of another recommended model shown in an embodiment of this specification;

[0026] Figure 7 This is a network architecture diagram of a gain network of a recommended model shown in an embodiment of this specification;

[0027] Figure 8 is a network architecture diagram of a gain network of another recommended model shown in an embodiment of this specification;

[0028] Figure 9 is a network architecture diagram of another recommended model shown in an embodiment of this specification;

[0029] Figure 10 is a network architecture diagram of a gain network of another recommended model shown in an embodiment of this specification;

[0030] Figure 11 is a flowchart of an object recommendation method based on a recommendation model shown in an embodiment of this specification;

[0031] Figure 12 is a schematic structural diagram of an electronic device shown in an embodiment of this specification;

[0032] Figure 13 is a block diagram of an optimization device for a recommendation model shown in an embodiment of this specification;

[0033] Figure 14 This is a block diagram of an object recommendation device based on a recommendation model shown in one embodiment of this specification. DETAILED DESCRIPTION

[0034] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The implementations described in the following exemplary embodiments are not intended to represent all implementations consistent with one or more embodiments of this specification. Rather, they are merely examples of apparatuses and methods consistent with certain aspects of one or more embodiments of this specification, as detailed in the appended claims.

[0035] It should be noted that in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this specification. In some other embodiments, the method may include more or fewer steps than those described in this specification. In addition, a single step described in this specification may be broken down into multiple steps for description in other embodiments, and multiple steps described in this specification may be combined into a single step for description in other embodiments.

[0036] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0037] This specification aims to propose a technical solution for optimizing a recommendation model in a scenario where similar objects corresponding to objects queried by the user are recommended to the user based on the recommendation model, so that the recommendation model can, based on the learned user preferences for the above-mentioned similar objects, further learn a preference gain that can reflect the user's preference for the above-mentioned similar objects based on the information differences contained in the historical objects queried by the user and the recommended similar objects corresponding to the historical objects, thereby improving the recommendation accuracy of the recommendation model.

[0038] Figure 1 This is a schematic diagram of the architecture of an Internet service system provided by an exemplary embodiment. Figure 1 As shown, the system may include a server 11 , a network 12 , and several electronic devices, such as a PC (Personal Computer) 13 , a mobile phone 14 , and the like.

[0039] The server 11 may be a physical server including an independent host, or a virtual server hosted by a host cluster. During operation, the server 11 may run a server-side program of an application to implement the relevant functions of the application. For example, when the server 11 runs a program for an Internet service, it may be implemented as a corresponding Internet service platform.

[0040] PC 13 and mobile phone 14 are only some types of electronic devices that can be used by users. In fact, users can obviously also use electronic devices such as the following types: tablet devices, laptops, PDAs (Personal Digital Assistants), wearable devices (such as smart glasses, smart watches, etc.), etc., and one or more embodiments of this specification do not limit this. During operation, the electronic device can run the client-side program of a certain application to implement the relevant functions of the application. For example, when the electronic device runs the program of a certain Internet service, it can be implemented as the client of the Internet service. Among them, the application program of the client of the above-mentioned Internet service can be started and run on the electronic device. The client-side program can be a native application installed on the electronic device, or the client-side program can be a small program, a quick application or other similar forms.

[0041] Of course, when using web technologies such as HTML5 or similar, relevant functions can be implemented through pages displayed by a browser. The browser here can be an independent browser application or a browser module embedded in certain applications.

[0042] Regarding the network 12 for interaction between electronic devices such as PC 13 and mobile phone 14 and server 11, communication can be achieved using a wired or wireless network based on the communication methods supported by the corresponding electronic devices, and this specification does not limit this. For example, if PC 13 supports both wired and wireless communication, then communication can be achieved using a wired or wireless network as needed, while mobile phone 14 generally only supports wireless communication and thus can achieve communication using a wireless network.

[0043] It should be noted that the above-mentioned Internet services may include any services implemented on the Internet; for example, the above-mentioned Internet services may specifically include travel transaction services that can provide users with travel products (such as air tickets), e-commerce services such as fresh food transaction services, or other types of services, etc., and this manual does not limit this.

[0044] The technical solution of this specification is described in detail below with reference to the accompanying drawings.

[0045] See Figure 2 , Figure 2 This is a flowchart of a method for optimizing a recommendation model shown in this specification, which includes the following execution process:

[0046] Step 202: Obtain a training sample set for training the recommendation model; wherein the training sample set includes a number of historical operation behavior data samples; the historical operation behavior data samples include attribute information of historical objects queried by the user, attribute information of recommended similar objects corresponding to the historical objects, and a label indicating whether the user has performed an operation on the similar objects;

[0047] The recommendation model described above can be used to recommend similar objects to the object queried by the user.

[0048] In one embodiment shown, the above-mentioned object can be a product, and the above-mentioned recommendation model can recommend similar products corresponding to the product queried by the user to the user, thereby maximizing the product conversion rate of the e-commerce platform.

[0049] For example, in one example, the aforementioned products may specifically include air tickets sold to users by a travel transaction service platform. Users can search for relevant air tickets by entering the required ODD (Origin-Destination-Departure Time) information in the query interface provided by the travel transaction service platform. In this scenario, to improve user conversion rates for air tickets, the aforementioned similar products may specifically include air tickets with similar itineraries to the user's searched air ticket. For example, such air tickets with similar itineraries may specifically include air tickets on nearby routes, with the departure and / or destination being less than a threshold distance from the user's searched air ticket.

[0050] In practical applications, the above recommendation model can specifically be a deep learning model that adopts a deep learning network architecture.

[0051] Please refer to FIG3 , which is a network architecture diagram of a recommendation model shown in an embodiment of this specification.

[0052] As shown in FIG3 , the above recommendation model may specifically include an input layer, a network layer, and an output layer.

[0053] The network layer is the core layer in the recommendation model, and specifically may include a backbone network and at least one gain network extended on the basis of the backbone network.

[0054] The backbone network can be used to learn user preferences for similar objects recommended by the recommendation model. In practical applications, the backbone network can specifically learn user preferences for similar objects based on historical user action data on recommended similar objects corresponding to historical objects queried by the user.

[0055] The gain network can be specifically used to learn preference gains corresponding to the user preferences. It should be noted that the preference gains can specifically represent the degree of preference of the user preferences. In practical applications, the gain network can specifically learn preference gains corresponding to the user preferences based on the information differences between historical objects retrieved by the user and recommended similar objects corresponding to the historical objects.

[0056] It should be noted that the specific data formats of the user preferences learned by the backbone network and the preference gains learned by the gain network are not particularly limited in this specification.

[0057] For example, in one embodiment shown, the user preference learned by the backbone network can be specifically a user preference score for representing the user preference; correspondingly, the preference gain learned by the gain network can also be specifically a preference gain score for representing the preference gain. Of course, in actual applications, if the backbone network and the gain network are both multidimensional networks that support multidimensional data processing, the user preference learned by the backbone network can be specifically a multidimensional user preference score vector composed of user preference scores corresponding to each dimension. Correspondingly, the preference gain learned by the gain network can also be specifically a multidimensional preference gain score vector composed of preference gain scores corresponding to each dimension.

[0058] The above input layer is specifically used to obtain input data for the recommendation model.

[0059] For example, Figure 3 As shown, the input layer may include at least two inputs, one of which is used to input historical operation behavior data of users' preferences for similar objects into the backbone network, and the other is used to input information differences of preference gains corresponding to the above user preferences into the gain network.

[0060] The above output layer is specifically used to output the recommendation results of the recommendation model.

[0061] It should be emphasized that in practical applications, in different recommendation scenarios, the network architecture of the above recommendation model can be flexibly expanded in terms of the number of gain networks based on the specific type of object to be recommended, on the basis of the network architecture shown in FIG3 .

[0062] Please refer to FIG4 , which is a network architecture diagram of another recommendation model shown in an embodiment of this specification.

[0063] For example, consider products as the recommended items. Product attributes typically include price information and other contextual information related to the product. The information differences used by the aforementioned gain network when learning preference gains can specifically include price and contextual information differences between historical products searched by the user and recommended similar products.

[0064] For example, if the product is an air ticket that the user has searched for, the similar products may be air tickets with similar itineraries to the air ticket that the user has searched for. The context information may include itinerary information other than price information, such as the distance and duration of the trip.

[0065] Correspondingly, the preference gain learned by the gain network may specifically include the price preference gain learned based on the price difference and the context information gain learned based on the context information difference.

[0066] As shown in Figure 4, Figure 3 Based on the network architecture shown, a first gain network and a second gain network are expanded.

[0067] The first gain network may be specifically configured to learn a price preference gain corresponding to the user preference based on price differences between historical products queried by the user and recommended similar products corresponding to the historical products.

[0068] The price preference gain may be used to specifically indicate the user's sensitivity to the price difference between the recommended similar products and the product the user has searched for.

[0069] It should be noted that, in actual applications, the above-mentioned price difference can be either an absolute price difference or a relative price difference. The so-called absolute price difference specifically refers to the specific value of the price difference between the recommended similar products and the products queried by the user. The relative price difference specifically refers to the ratio of the price difference between the recommended similar products and the products queried by the user to the total price of the products queried by the user. In actual applications, the absolute price difference can be used to indicate the price reduction amount of the recommended similar products compared to the products queried by the user; and the relative price difference can be used to indicate the price reduction ratio or price reduction range of the recommended similar products compared to the products queried by the user.

[0070] For example, in practical applications, by learning the user's price preference gain, the recommendation model can recommend similar products with lower prices or larger price reductions to users as much as possible when recommending similar products to users.

[0071] The second gain network can be specifically used to learn the context information gain corresponding to the user preference based on the difference in context information contained in the historical products queried by the user and the recommended similar products corresponding to the historical products.

[0072] The context information gain may be used to indicate the difference in context information between the recommended similar products and the product queried by the user, and the information gain on whether the user will operate the recommended similar products.

[0073] For example, in actual applications, based on the learned price preference gain, by further learning the context information gain corresponding to the above user preference, it is possible to recommend to the user lower prices or similar products with larger price reductions as much as possible while ensuring that the context information difference between the similar products recommended to the user and the products queried by the user is within the acceptable range for the user, rather than recommending similar products with the lower prices the better.

[0074] In this specification, the above-mentioned backbone network and gain network can specifically be a network architecture based on deep learning.

[0075] In one embodiment shown, both the backbone network and the gain network can be neural networks based on the attention mechanism.

[0076] See Figure 5 , Figure 5 This is a network architecture diagram of a backbone network of a recommendation model shown in one embodiment of this specification.

[0077] like Figure 5 As shown in FIG, the backbone network may include an embedding layer, an attention pooling layer, and an attention layer.

[0078] The above-mentioned Embedding layer can be used to encode the input data into a vector, and further input the encoded vector into the attention convergence layer.

[0079] The attention aggregation layer can aggregate multiple input data into a vector representation of a fixed length based on the information aggregation method of the attention mechanism, and further input the vector into the attention layer.

[0080] The attention layer is used to learn features related to the network's learning objectives from the vector output by the attention aggregation layer based on the attention mechanism. For example, based on the attention mechanism, the importance weights of different parts of the input vector can be calculated, and the vectors can be weightedly aggregated according to these weights, allowing the model to "focus" on the most relevant parts of the input data.

[0081] In one embodiment shown, the context information gain learned by the second gain network is only the information gain generated by the context information difference on the user preference and cannot directly reflect the user's preference. The price preference gain learned by the first gain network directly reflects the user's price preference for the product. Therefore, in actual applications, the price preference gain learned by the first gain network can also be used as one input of the backbone network and input into the backbone network.

[0082] In this case, the backbone network may specifically learn the user's preference for similar objects based on the historical operation behavior data and the price preference gain learned by the first gain network.

[0083] For example, see Figure 6 , Figure 6 This is a network architecture diagram of another recommendation model shown in an embodiment of this specification.

[0084] like Figure 6 As shown, the backbone network adopts Figure 5 Taking the network architecture shown as an example, the price preference gain learned by the first gain network can be used as input to the attention layer of the backbone network.

[0085] In this case, the attention layer can specifically learn features related to the network's learning objectives from the vector output by the attention convergence layer and the price preference gain output by the first gain network based on the attention mechanism.

[0086] See Figure 7 , Figure 7 This is a network architecture diagram of a gain network of a recommendation model shown in an embodiment of this specification.

[0087] It should be noted that both the first gain network and the second gain network can be Figure 7 The network architecture shown here differs only in the input data.

[0088] like Figure 7 As shown in FIG, the gain network may specifically include an embedding layer, an attention layer, and a feedforward network. Since the gain network only needs to take the information difference as input data, the attention convergence layer may not be included in the network structure of the gain network.

[0089] Among them, the above-mentioned Embedding layer can still be used to encode the input data into a vector, and further input the encoded vector into the attention layer.

[0090] The attention layer can still be used to learn features related to the network's learning objectives from the vector output by the Embedding layer based on the attention mechanism.

[0091] The feedforward network can be used to perform nonlinear transformation on the features learned by the attention layer to obtain the preference gain learned by the network.

[0092] For example, in one example, one can introduce A nonlinear transformation is performed on the features learned by the attention layer, thereby converting the features learned by the attention layer into a value between 0 and 1; in this case, the preference gain learned by the above-mentioned gain network can also be a value between 0 and 1 that can be used as a probability value.

[0093] The specific type of the feedforward network used in the gain network is not particularly limited in this specification.

[0094] See Figure 8 , Figure 8 This is a network architecture diagram of a gain network of another recommendation model shown in an embodiment of this specification.

[0095] like Figure 8 As shown, in practical applications, a reward network based on reinforcement learning can be used as the above-mentioned feedforward network. In this case, the preference gain learned by the above-mentioned gain network can be a reward value learned by the above-mentioned reward network using reinforcement learning.

[0096] The specific process by which the reward network learns the reward value using reinforcement learning will not be described in detail in this specification.

[0097] For example, in one example, when the above-mentioned reward value is learned using reinforcement learning, the information difference input to the reward network can be used as the state of reinforcement learning, the similar object to be recommended can be used as the action, and whether the user will operate on the similar object (such as click) can be used as the reward signal for reinforcement learning, thereby predicting the reward value of the similar object under a given information difference (i.e., the click probability).

[0098] It should be emphasized that since the above-mentioned gain network needs to further learn the preference gain corresponding to the user's preference based on multiple information differences (such as the price difference and context information difference mentioned above), in practical applications, the above-mentioned gain network can specifically adopt a neural network based on a multi-head attention mechanism.

[0099] See Figure 9 , Figure 9 This is a network architecture diagram of another recommendation model shown in an embodiment of this specification.

[0100] like Figure 9 As shown, the output layer of the recommendation model may specifically include a scoring network and a classifier.

[0101] The scoring network is used to integrate the user preferences learned by the backbone network, the price preference gain learned by the first gain network, and the context information gain learned by the second gain network to perform scoring processing and obtain a recommendation score.

[0102] For example, Figure 9 As shown, the user preference learned by the backbone network, the price preference gain learned by the first gain network, and the context information gain learned by the second gain network can be spliced ​​into a complete vector and then further input into the above-mentioned scoring network for scoring processing.

[0103] It should be noted that in the process of integrating the user preferences learned by the backbone network, the price preference gain learned by the first gain network, and the context information gain learned by the second gain network, new information can be introduced into the information to be integrated based on specific needs, which is specially limited in this specification.

[0104] For example, Figure 9 As shown, in one example, the user information of the user can also be used as part of the data that needs to be integrated. It can be spliced ​​into a complete vector with the user preference learned by the backbone network, the price preference gain learned by the first gain network, and the context information gain learned by the second gain network, and then further input into the above-mentioned scoring network for scoring processing.

[0105] The scoring network may be any type of network and is not particularly limited in this specification.

[0106] For example, Figure 9As shown in the figure, in one example, the scoring network can specifically adopt a network structure consisting of an FC layer (fully connected layer) and a Relu (rectified linear unit). In this network structure, the output of the FC layer will serve as the input of the Relu, and the output of the Relu is the recommendation score output by the scoring network.

[0107] The above classifier can be used to classify similar products included in the historical operation behavior data sample based on the recommendation score obtained by the scoring network.

[0108] The classification results of this classification process may include a first classification result indicating that the user will perform an action on the similar product, and a second classification result indicating that the user will not perform an action on the similar product. It should be noted that the classification result output by the classifier is the recommendation model's recommendation result for similar products. If the classification result output by the classifier is the first classification result, it indicates that the user is likely to perform an action on the similar product, and the similar product can be recommended; otherwise, the similar product will not be recommended.

[0109] The above-mentioned classifier may be any type of classifier and is not particularly limited in this specification.

[0110] For example, Figure 9 As shown, in one example, the above-mentioned classifier can specifically be a binary classifier. When classifying based on the recommendation score, the classifier can compare the recommendation score with a preset threshold; if the recommendation score is greater than the preset threshold, a classification label 1 can be output as the above-mentioned first classification result; otherwise, a classification label 0 can be output as the above-mentioned second classification result.

[0111] It should be noted that the aforementioned user operations on the similar products may include any form of user operations that can reflect the user's intentions. For example, the aforementioned operations may include click operations, payment operations, forwarding operations, favorite operations, etc.; these operations are not listed one by one in this specification.

[0112] In this specification, when training the recommendation model, a training sample set for training the recommendation model may be obtained. The training sample set may include several historical operation behavior data samples generated by users. The historical operation behavior data samples may include attribute information of historical objects retrieved by the user, attribute information of recommended similar objects corresponding to the historical objects, and data labels. The data labels are used to indicate whether the user has performed an operation on the similar objects included in the historical operation behavior data samples.

[0113] For example, the data label may specifically include a first label indicating that the user has performed an operation on a similar object included in the historical operation behavior data sample, and a second label indicating that the user has not performed an operation on a similar object included in the historical operation behavior data sample.

[0114] Step 204: Inputting the historical operation behavior data samples in the training sample set into the backbone network for supervised learning and training, and calculating the user preference loss of the user preference learned by the backbone network based on the output result of the backbone network; and calculating the information difference between the attribute information of the historical object and the attribute information of the similar object contained in the historical operation behavior data samples, further inputting the calculated information difference into the gain network for supervised learning and training, and calculating the preference gain loss of the preference gain learned by the gain network based on the output result of the gain network;

[0115] After the training sample set is obtained, the historical operation behavior data samples in the training sample set can be used as training samples and input into the recommendation model to train the recommendation model.

[0116] For example, in practical applications, the historical operation behavior data samples in the training sample set can be input into the recommendation model one by one to train the recommendation model, or the historical operation behavior data samples in the training sample set can be divided into several batches and then input into the recommendation model one by one to train the recommendation model.

[0117] It should be noted that, compared with the traditional recommendation model that only includes a backbone network, the above-mentioned recommendation model has at least one gain network expanded on the basis of the backbone network; therefore, the training method used in training the above-mentioned recommendation model will also be different from the traditional training method.

[0118] Specifically, when training the above recommendation model, on the one hand, the historical operation behavior data samples in the training sample set can be input into the backbone network for supervised learning training, and then the user preference loss of the user preferences learned by the backbone network can be calculated based on the output results of the backbone network;

[0119] Among them, the user preference loss of the user preference learned by the backbone network can be specifically represented by the loss of the recommendation result output by the recommendation model.

[0120] For example, see Figure 9 In one example, if the output layer of the recommendation model consists of Figure 9The scoring network and classifier shown are composed of the following: the classification result output by the classifier is the recommendation result of the recommendation model for similar products. In this case, the user preference loss of the user preference learned by the backbone network can be expressed by the cross-entropy loss (Cross-Entropy Loss) of the classification task performed by the classifier. Among them, the cross-entropy loss is usually used to measure the difference between the probability distribution of the classification result output by the classifier and the probability distribution of the classification result indicated by the true label of the training sample. The specific process of calculating the cross-entropy loss of the classification task will not be described in detail in this manual.

[0121] On the other hand, the attribute information contained in the historical operation behavior data sample can also be preprocessed to calculate the information difference between the attribute information of the historical object and the attribute information of the similar object contained in the historical operation behavior data sample, and then the calculated information difference can be further input into the gain network for supervised learning and training. Then, based on the output result of the gain network, the preference gain loss of the preference gain learned by the gain network can be calculated.

[0122] It should be noted that, in the process of training the above-mentioned recommendation model, the training of the backbone network and the training of the gain network can be performed in parallel or in a certain order, which is not particularly limited in this description.

[0123] In one embodiment shown, the object to be recommended by the recommendation model is still a commodity as an example. Figure 4 As shown, the information differences used to learn the preference gain can specifically include price differences and contextual information differences of the products, and the gain network can specifically include the first gain network and the second gain network. Accordingly, the loss generated by training the gain network can include two parts of loss: one part is the price preference gain loss of the price preference gain learned by the first gain network; the other part is the context information gain loss of the context information gain learned by the second gain network.

[0124] In this case, when training the above-mentioned gain network, on the one hand, the price difference between the price information of historical commodities and the price information of similar commodities contained in the input historical operation behavior data sample can be calculated, and the calculated price difference can be further input into the first gain network for supervised learning training. Then, based on the output result of the first gain network, the price preference gain loss of the price preference gain learned by the first gain network can be calculated;

[0125] On the other hand, the context information difference between the context information of historical products and the context information of similar products contained in the input historical operation behavior data sample can also be calculated, and the calculated context information can be further input into the second gain network for supervised learning training. Then, based on the output result of the second gain network, the context information gain loss of the context information gain learned by the second gain network can be calculated.

[0126] It should be noted that, in order to improve the learning effect of the above-mentioned gain network, in the process of training the above-mentioned first gain network and the above-mentioned second gain network, a difference data sample corresponding to the historical operation behavior data can be introduced as a reference based on the input historical operation behavior data sample.

[0127] In one embodiment shown, before calculating the price difference between the price information of historical commodities contained in the input historical operation behavior data sample and the price information of similar commodities, a difference data sample corresponding to the input historical operation behavior data sample can be searched in the training sample set.

[0128] The difference data sample may specifically be a data sample generated before the input historical operation behavior data sample;

[0129] For example, in one example, the difference data sample can specifically be the data sample generated most recently before the input historical operation behavior data sample; for example, a data sample closest to the generation time of the historical operation behavior data sample; or, the difference data sample can specifically be a data sample generated before the input historical operation behavior data sample and generated at a time similar to the generation time of the historical operation behavior data sample; for example, a data sample generated in the same period or at the same time on different dates.

[0130] In addition, the data content contained in the difference data sample may be completely identical to the input historical operation behavior data sample, but with different data labels.

[0131] Since the data label is specifically used to indicate whether the user has performed an operation on similar products contained in the data sample, the difference data sample and the input historical operation behavior data sample have different data labels, which means that the user has performed different operation behaviors on the similar products contained in the difference data sample and the similar products contained in the input historical operation behavior data sample.

[0132] In one embodiment shown, after the difference data sample is found, when the first gain network is trained, in addition to calculating the first price difference between the price information of the historical commodities contained in the input historical operation behavior data sample and the price information of similar commodities, the second price difference between the price information of the historical commodities contained in the difference data sample and the price information of similar commodities can also be calculated as a reference for the first price difference; then, the calculated first price difference and second price difference can be further input into the first gain network for supervised learning training, and based on the output result of the first gain network, the price preference gain loss of the price preference gain learned by the first gain network can be calculated.

[0133] It should be noted that if the calculated first price difference and second price difference are simultaneously input into the first gain network for supervised learning and training, the output result of the first gain network may generally include a first price preference gain learned based on the first price difference and a second price preference gain learned based on the second price difference.

[0134] The price preference gain loss of the price preference gain learned by the first gain network is usually related to the optimization target designed when the first gain network is learned and trained.

[0135] In one embodiment shown, the optimization goal of learning and training the first gain network may be to maximize the gap between the first price preference gain and the second price preference gain.

[0136] At this point, the optimization goal can be specifically expressed by the following mathematical expression:

[0137]

[0138] Among them, in the above expression, represents the first price preference gain mentioned above; represents the second price preference gain mentioned above.

[0139] It should be explained that the training process of the neural network is essentially a mathematical problem of finding the optimal minimum value; therefore, when designing the optimization target, we can introduce The function converts the optimization goal of maximizing the difference between the first price preference gain and the second price preference gain into a mathematically equivalent problem of minimizing the difference between the second price preference gain and the first price preference gain.

[0140] Based on the above optimization objectives, the loss function of the first gain network can be expressed as follows:

[0141]

[0142] In the above expression, Represents the Sigmoid function represents the price preference gain loss of the first gain network mentioned above; represents the first price preference gain mentioned above; represents the second price preference gain mentioned above.

[0143] In one embodiment shown, in order to further improve the learning effect of the first gain network, when designing the optimization target for the first gain network, in addition to considering the gap between the above-mentioned first price preference gain and the second price preference gain, the difference between the recommendation results output by the above-mentioned recommendation model for similar products contained in the input historical operation behavior data sample and the recommendation results indicated by the data labels contained in the historical operation behavior data sample can also be considered.

[0144] In this case, the optimization goal of learning and training the first gain network can be to simultaneously maximize the gap between the above-mentioned first price preference gain and the second price preference gain, as well as the gap between the classification result (i.e., recommendation result) obtained by the recommendation model by classifying similar products contained in the input historical operation behavior data sample, and the classification result indicated by the data label contained in the historical operation behavior data sample.

[0145] For example, when the data label contained in the historical operation behavior data sample indicates that the user has operated on similar products contained in the historical operation behavior data sample, the classification result indicated by the data label is the above-mentioned first classification result; when the data label contained in the historical operation behavior data sample indicates that the user has operated on similar products contained in the historical operation behavior data sample, the classification result indicated by the data label is the above-mentioned second classification result.

[0146] At this point, the optimization goal can be specifically expressed by the following mathematical expression:

[0147]

[0148] Among them, in the above expression, Still represents the first price preference gain mentioned above; Still represents the second price preference gain mentioned above; It represents the classification result obtained by the recommendation model by classifying similar products contained in the input historical operation behavior data sample; Indicates the classification result indicated by the data label contained in the historical operation behavior data sample.

[0149] Based on the above optimization objectives, the loss function of the first gain network can be expressed as follows:

[0150]

[0151] Among them, in the above expression, It still represents the price preference gain loss of the first gain network, and the meanings of other parameters remain unchanged.

[0152] In one embodiment shown, if the optimization objective of the first gain network also takes into account the difference between the recommendation results output by the above-mentioned recommendation model for similar products contained in the input historical operation behavior data sample and the recommendation results indicated by the data labels contained in the historical operation behavior data sample, an alignment function can also be introduced into the loss function of the first gain network.

[0153] At this point, the loss function of the first gain network can be expressed as follows:

[0154]

[0155] Among them, if The classification result represented is the first classification result mentioned above, then The value of is 1. If The classification result represented is the second classification result mentioned above, then The value of is -1.

[0156] For example, assuming that the first classification result is represented by the classification label 1 and the second classification result is represented by the classification label 0, then The value of can be expressed as follows:

[0157]

[0158] During the training of the first gain network, the input historical operation behavior data samples typically include positive samples (i.e., samples indicating that the classification result is the first classification result) and negative samples (i.e., samples indicating that the classification result is the second classification result). If the optimization objective of the first gain network also considers the difference between the recommendation results output by the recommendation model for similar products contained in the input historical operation behavior data samples and the recommendation results indicated by the data labels contained in the historical operation behavior data samples, then inputting negative samples into the first gain network for training typically causes the parameters of the first gain network to be adjusted in a direction opposite to the optimization objective, thereby resulting in a slower convergence rate of learning and training. Therefore, by introducing an alignment function into the loss function of the first gain network, the training process of the first gain network can be adjusted when negative samples are input into the first gain network for training, causing the network to adjust the parameters of the first gain network in the direction consistent with the optimization objective, thereby reducing the negative impact of negative samples on training and improving the convergence rate of the first gain network.

[0159] In another embodiment shown, after the difference data sample is found, when the second gain network is trained, in addition to calculating the first context information difference between the context information of the historical product and the context information of similar products contained in the input historical operation behavior data sample, the second context information difference between the context information of the historical product and the context information of similar products contained in the difference data sample can also be calculated as a reference for the first context information difference; then, the calculated first context information difference and second context information difference can be further input into the second gain network for supervised learning training, and based on the output result of the second gain network, the context information gain loss of the context information gain learned by the second gain network is calculated.

[0160] It should be noted that if the calculated first context information difference and second context information difference are simultaneously input into the second gain network for supervised learning and training, the output result of the second gain network may generally include the first context information gain learned based on the first context information difference and the second context information gain learned based on the second context information.

[0161] The contextual information gain loss of the contextual information gain learned by the second gain network is usually also related to the optimization target designed when the second gain network is learned and trained.

[0162] In one embodiment shown, the optimization goal of learning and training the second gain network may be to maximize the gap between the first context information gain and the second context information gain.

[0163] At this point, the optimization goal can also be expressed by the following mathematical expression:

[0164]

[0165] Among them, in the above expression, represents the first context information gain; represents the second context information gain.

[0166] Based on the above optimization objectives, the loss function of the second gain network can be expressed as follows:

[0167]

[0168] Among them, in the above expression; Represents the Sigmoid function represents the contextual information gain loss of the second gain network above; represents the first context information gain; represents the second context information gain; m represents the number of historical commodities queried by the user contained in the input historical operation behavior data sample.

[0169] It should be emphasized that the loss functions of the first gain network and the second gain network listed above use the BPR (Bayesian Personalized Ranking) loss function. The BPR loss function is a common loss function for learning user personalized preferences in recommendation systems. In practical applications, other types of loss functions can also be used to construct training losses for the first gain network and the second gain network, which will not be listed one by one in this specification.

[0170] In one embodiment shown, if the price preference gain learned by the first gain network is also used as an input to the backbone network, the backbone network can specifically learn the user's preference for similar objects based on the historical operation behavior data and the price preference gain learned by the first gain network.

[0171] In this case, before the backbone network is subjected to supervised learning and training, the first gain network can be trained first, and the price preference gain output by the first gain network during the training process can be obtained. Then, the historical operation behavior data samples in the training sample set and the obtained price preference gain can be input into the backbone network together for supervised learning and training.

[0172] It should be noted that, if, in the process of training the first gain network and the second gain network, based on the input historical operation behavior data samples, the difference data samples corresponding to the historical operation behavior data are introduced as references, Figure 8 The input portion of the gain network structure shown usually also requires adjustment.

[0173] See Figure 10 , Figure 10 This is a network architecture diagram of a gain network of another recommendation model shown in an embodiment of this specification.

[0174] Still taking the above-mentioned object as a product as an example, in the scenario of similar product recommendation, the attribute information of the historical product queried by the user usually includes the product query index entered by the user when querying the historical product.

[0175] For example, taking the above-mentioned product as an example of an air ticket sold to a user by a travel transaction service platform, in this case, the attribute information of the historical air tickets queried by the user usually includes the ODD information entered by the user in the query interface provided by the travel transaction service platform to the user.

[0176] In a neural network based on the attention mechanism, the input to the attention layer is usually in the form of a QKV vector, which is obtained by performing linear transformations such as embedding on the input data. The Q vector in the QKV vector is called the query vector, the K vector in the QKV vector is called the key vector, and the V vector in the QKV vector is called the value vector.

[0177] Based on this, Figure 10 As shown, when the price differences or context information differences between historical commodities and similar commodities contained in the historical operation behavior data samples and the difference data samples are input into the gain network for supervised learning training, the commodity query index contained in the historical operation behavior data samples and the difference data samples, as well as the above-mentioned price differences and context information differences can be input into the gain network as input data.

[0178] On the one hand, the product query index contained in the input historical operation behavior data sample and the difference data sample, after being processed by the embedding layer, can be further input into the attention layer as the key vector;

[0179] On the other hand, the price differences or context information differences between historical commodities and similar commodities contained in the input historical operation behavior data samples and difference data samples can be processed by the embedding layer to obtain vectors, which can be used as query vectors of the attention layer to continue to be input into the attention layer.

[0180] It should be noted that, in the scenario of similar product recommendations, the above Key vector and Query vector already contain sufficient features related to the network's learning objectives. Therefore, in practical applications, the QKV vector input to the attention layer can be simplified and the Value vector can be omitted.

[0181] Please continue to see Figure 10 After inputting the above Key vector and Query vector into the attention layer, features related to the network's learning objectives can be learned from these vectors, and then the learned features are further input into the reward network, and the above reward value is learned by reinforcement learning as the price preference gain or context information gain learned by the gain network. Then, based on the above loss function, the price preference gain loss of the price preference gain learned by the gain network and the context information gain loss of the context information gain learned by the gain network can be calculated respectively.

[0182] Step 206: Using the user preference loss and the preference gain loss as the training losses of the recommendation model, adjusting the model parameters contained in the recommendation model to complete the learning and training of the recommendation model; wherein the trained recommendation model is used to recommend similar objects corresponding to the objects queried by the user to the user based on the user preferences learned by the backbone network and the preference gains learned by the gain network.

[0183] After calculating the user preference loss of the user preferences learned by the above-mentioned backbone network and the preference gain loss of the preference gain learned by the above-mentioned gain network, the user preference loss and the preference gain loss can be used as the global training loss of the recommendation model, and the model parameters contained in the recommendation model can be adjusted to complete the learning and training of the recommendation model.

[0184] In one embodiment shown, taking the example in which the above-mentioned preference gain includes the price preference gain learned based on the above-mentioned price difference and the context information gain learned based on the above-mentioned context information difference, the user preference loss of the backbone network, the price preference gain loss of the first gain network, and the context information gain loss of the second gain network can be used as the global training loss of the recommendation model to adjust the model parameters contained in the recommendation model.

[0185] For example, in one example, a weighted calculation can be performed on the user preference loss of the backbone network, the price preference gain loss of the first gain network, and the contextual information gain loss of the second gain network to obtain the global training loss of the recommendation model. It should be noted that when performing weighted calculations on the user preference loss, price preference gain loss, and contextual information gain loss, the weight coefficients of the user preference loss, price preference gain loss, and contextual information gain loss can be flexibly set based on specific needs and are not specifically limited in this specification.

[0186] The specific implementation process of adjusting the parameters of each model included in the recommendation model based on the global training loss will not be described in detail in this specification. Those skilled in the art can refer to the records in the relevant technology.

[0187] For example, in practical applications, in a round of training for a recommendation model, after calculating the global training loss of this round of iteration, the gradients for adjusting each model parameter can be calculated based on the training loss, and then the calculated gradients can be backpropagated to adjust each model parameter based on the calculated gradients.

[0188] After the recommendation model is trained according to the training method shown above, the recommendation model can be used to recommend similar objects corresponding to the object queried by the user to the user.

[0189] For example, similar objects corresponding to the object queried by the user can be recommended to the user based on the user preferences learned by the backbone network of the recommendation model and the preference gains learned by the gain network of the recommendation model.

[0190] See Figure 11 , Figure 11 This is a flowchart of a method for recommending an object based on a recommendation model, which includes the following execution process:

[0191] Step 1102: Obtain the object queried by the user, and construct a recommendation data sample based on the object queried by the user and similar objects corresponding to the object to be recommended; wherein the recommendation data sample includes attribute information of the object queried by the user and attribute information of the similar objects to be recommended;

[0192] The training process of the above recommendation model will not be described in detail.

[0193] When using a trained recommendation model to recommend similar objects to a user's query object, the user's query object can be first obtained, and a similar object to be recommended can be obtained from a preset set of similar objects to be recommended. Then, based on the user's query object and the obtained similar object to be recommended, a recommendation data sample is constructed. The recommendation data sample can specifically include attribute information of the user's query object and attribute information of the similar object to be recommended.

[0194] For example, taking the above-mentioned objects as commodities, in this case, when using the trained recommendation model to recommend similar commodities corresponding to the commodities queried by the user to the user, after the user queries the commodity, a similar commodity to be recommended can be further obtained from the set of similar objects to be recommended, and then a recommendation data sample can be constructed based on the attribute information of the commodity queried by the user and the attribute information of the similar commodity to be recommended.

[0195] Step 1104: Input the recommended data sample into the backbone network for learning and calculation, and obtain the user preference of the user for the similar object output by the backbone network; and calculate the information difference between the attribute information of the object queried by the user contained in the recommended data sample and the attribute information of the similar object to be recommended, further input the calculated information difference into the gain network for learning and calculation, and obtain the preference gain corresponding to the user preference output by the gain network;

[0196] After constructing the recommendation data sample, on the one hand, the recommendation data sample can be further input into the backbone network of the recommendation model for learning calculation, and the user preference of the user for the similar objects to be recommended contained in the recommendation data sample output by the backbone network can be obtained.

[0197] On the other hand, the attribute information contained in the recommendation data sample can also be preprocessed to calculate the information difference between the attribute information of the object queried by the user contained in the recommendation data sample and the attribute information of the similar object to be recommended, and then the calculated information difference is further input into the gain network of the recommendation model for learning calculation, and the preference gain output by the gain network corresponding to the user preference learned by the backbone network is obtained.

[0198] In one embodiment shown, the object to be recommended by the recommendation model is still a product, and the information difference used to learn the preference gain includes the price difference and context information difference of the product. The above-mentioned gain network can specifically include the above-mentioned first gain network and the above-mentioned second gain network; wherein, the functions and network structures of the above-mentioned first gain network and the above-mentioned second gain network are not repeated.

[0199] In this scenario, on the one hand, the price difference between the price information of the product queried by the user and the price information of the similar product to be recommended contained in the recommendation data sample can be calculated, and the calculated price difference can be further input into the first gain network for learning calculation, and the price preference gain output by the first gain network can be obtained;

[0200] On the other hand, the difference between the context information of the product queried by the user and the context information of the similar product to be recommended contained in the recommendation data sample can also be calculated, and the calculated context information difference can be further input into the second gain network for learning calculation, and the context information gain output by the second gain network can be obtained.

[0201] In order to improve the learning effect of the above-mentioned gain network, in the process of training the above-mentioned first gain network and the above-mentioned second gain network, a difference data sample corresponding to the historical operation behavior data can be introduced as a reference based on the input historical operation behavior data sample. The details of introducing the difference data sample as a reference to train the above-mentioned gain network will not be repeated.

[0202] In this case, before calculating the information difference between the attribute information of the product queried by the user and the attribute information of the similar product to be recommended contained in the recommendation data sample, a difference data sample corresponding to the recommendation data sample can also be searched in the user's historical operation data sample;

[0203] The difference data sample may still be a data sample generated before the recommended data sample;

[0204] For instance, in one example, the difference data sample may specifically be the most recent data sample generated before the recommended data sample; for example, a data sample closest to the time when the recommended data sample was generated; or, the difference data sample may specifically be a data sample generated before the recommended data sample and at a time similar to the time when the recommended data sample was generated; for example, a data sample generated in the same period or at the same time on different dates.

[0205] In addition, the data content contained in the difference data sample can be exactly the same as the recommended data sample, but with different data labels. For example, the difference data sample can also specifically include the attribute information of the same historical product queried by the user and the attribute information of the same similar products that have been recommended and correspond to the historical product.

[0206] It should be noted that when using the recommendation model to recommend similar products to users, although the generated recommendation data samples do not contain data labels, the generated recommendation data samples can be used as positive samples by default, and the queried difference data samples can be used as negative samples.

[0207] In this case, the data label contained in the queried difference data sample corresponding to the recommended data sample may specifically be a data label indicating that the user has not performed any operation on a similar product contained in the difference data sample.

[0208] In one embodiment shown, after the difference data sample is found, on the one hand, the first price difference between the price information of the product queried by the user contained in the recommendation data sample and the price information of the similar product to be recommended can be further calculated, and the second price difference between the price information of the historical product contained in the difference data sample and the price information of the similar product can be calculated; then, the calculated first price difference and second price difference are further input into the first gain network for learning calculation to obtain the first price preference gain learned based on the first price difference; and the second price preference gain learned based on the second price difference.

[0209] Finally, only the first price preference gain learned based on the first price difference output by the first gain network can be obtained as the price preference gain finally learned by the first gain network. As for the second price preference gain learned based on the second price difference output by the first gain network, it is itself a price preference gain learned based on the difference data sample. When using the recommendation model to recommend similar products to users, this part of the price preference gain can be ignored;

[0210] On the other hand, a first context information difference between the context information of the product queried by the user and the context information of similar products contained in the recommendation data sample can also be calculated, and a second context information difference between the context information of the historical product and the context information of the similar product contained in the difference data sample can be calculated; then, the calculated first context information difference and second context information difference are further input into the second gain network for learning calculation, thereby obtaining a first context information gain learned based on the first context information difference; and a second price preference gain learned based on the second context information difference;

[0211] Finally, only the first context information gain learned based on the first context information difference output by the second gain network can be obtained as the context information gain learned by the second gain network.

[0212] Step 1106 : Determine whether to recommend the similar object to the user based on the user preference learned by the backbone network and the preference gain learned by the gain network.

[0213] After obtaining the user preferences learned by the backbone network and the preference gains learned by the gain network, it can be determined whether to recommend similar objects to be recommended contained in the recommendation data sample to the user based on the user preferences and the preference gains.

[0214] In one embodiment shown, taking the example where the above-mentioned preference gain includes the price preference gain learned based on the above-mentioned price difference and the context information gain learned based on the above-mentioned context information difference, it can be determined whether to recommend the similar objects to be recommended contained in the recommendation data sample to the user based on the user preferences learned by the backbone network, the price preference gain learned by the first gain network, and the context information gain learned by the second gain network.

[0215] In one embodiment shown, if the output layer of the recommendation model consists of Figure 8 The scoring network and classifier shown are composed of the following: the classification result output by the classifier is the recommendation result of the recommendation model for similar products.

[0216] The above-mentioned scoring network can be specifically used to integrate the user preferences learned by the backbone network, the price preference gain learned by the first gain network, and the context information gain learned by the second gain network to perform scoring processing to obtain a recommendation score.

[0217] The above classifier can be used to classify similar products included in the recommendation data sample based on the recommendation score obtained by the scoring network.

[0218] In this case, the user preference learned by the backbone network, the first price preference gain learned by the first gain network, and the first context information gain learned by the second gain network can be spliced ​​into a complete vector and then further input into the scoring network for scoring processing.

[0219] When the scoring network completes the scoring process and obtains the recommendation score, the recommendation score can be further input into the above-mentioned classifier, and the classifier classifies similar products included in the recommendation data sample based on the recommendation score.

[0220] The classification result of the classification process may include a first classification result in which the user will perform an operation on the similar product and a second classification result in which the user will not perform an operation on the similar product.

[0221] For example, Figure 9 As shown, in one example, the above-mentioned classifier can specifically be a binary classifier. When classifying based on the recommendation score, the classifier can compare the recommendation score with a preset threshold; if the recommendation score is greater than the preset threshold, a classification label 1 can be output as the above-mentioned first classification result; otherwise, a classification label 0 can be output as the above-mentioned second classification result.

[0222] It should be noted that the classification result output by the classifier is the recommendation result of the recommendation model for similar products.

[0223] If the classification result output by the classifier is the first classification result mentioned above, the similar product to be recommended contained in the recommendation data sample can be recommended to the user; otherwise, the similar product to be recommended will not be recommended to the user.

[0224] In the above technical solution, in the scenario where similar objects corresponding to the objects queried by the user are recommended to the user based on the recommendation model, at least one gain network is expanded on the basis of the backbone network of the recommendation model, and a preference gain that can reflect the user's preference for similar objects is learned in the gain network based on the information difference contained in the historical objects queried by the user and the recommended similar objects corresponding to the historical objects. This allows the recommendation model to recommend similar objects with a higher probability of the user performing operations to the user based on the user preferences learned by the backbone network and the preference gains learned by the gain network, thereby improving the recommendation accuracy of the recommendation model.

[0225] Corresponding to the embodiments of the aforementioned method, this specification also provides embodiments of an apparatus, an electronic device, and a storage medium.

[0226] Figure 12 This is a schematic structural diagram of an electronic device provided by an exemplary embodiment. Figure 12At the hardware level, the device includes a processor 1202, an internal bus 1204, a network interface 1206, a memory 1208, and a non-volatile memory 1210, and may also include other required hardware. One or more embodiments of this specification can be implemented based on software, such as the processor 1202 reading the corresponding computer program from the non-volatile memory 1210 into the memory 1208 and then running it. Of course, in addition to software implementation, one or more embodiments of this specification do not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0227] like Figure 13 As shown, Figure 13 This is a block diagram of an optimization device for a recommendation model according to an exemplary embodiment of the present specification. The device can be applied to Figure 12 The electronic device shown in the figure implements the technical solution of this specification. The recommendation model is used to recommend similar objects corresponding to the object queried by the user to the user; the recommendation model includes a backbone network and at least one gain network; the backbone network is used to learn the user's preference for the similar objects; the gain network is used to learn the preference gain corresponding to the user's preference based on the information difference between the historical objects queried by the user and the recommended similar objects corresponding to the historical objects; the preference gain represents the degree of preference of the user's preference; the device 130 includes:

[0228] The first acquisition module 1301 acquires a training sample set for training the recommendation model; wherein the training sample set includes a plurality of historical operation behavior data samples; the historical operation behavior data samples include attribute information of historical objects queried by the user, attribute information of recommended similar objects corresponding to the historical objects, and a label indicating whether the user has performed an operation on the similar objects;

[0229] The first calculation module 1302 inputs the historical operation behavior data samples in the training sample set into the backbone network for supervised learning and training, and calculates the user preference loss of the user preference learned by the backbone network based on the output result of the backbone network; and calculates the information difference between the attribute information of the historical object and the attribute information of the similar object contained in the historical operation behavior data sample, further inputs the calculated information difference into the gain network for supervised learning and training, and calculates the preference gain loss of the preference gain learned by the gain network based on the output result of the gain network;

[0230] The adjustment module 1303 uses the user preference loss and the preference gain loss as the training loss of the recommendation model, and adjusts the model parameters contained in the recommendation model to complete the learning and training of the recommendation model; wherein the trained recommendation model is used to recommend similar objects corresponding to the objects queried by the user to the user based on the user preferences learned by the backbone network and the preference gains learned by the gain network.

[0231] like Figure 14 As shown, Figure 14 This is a block diagram of an object recommendation device based on a recommendation model according to an exemplary embodiment of the present specification; the device can also be applied to Figure 12 The electronic device shown in the figure implements the technical solution of this specification. The recommendation model is used to recommend similar objects corresponding to the object queried by the user to the user; the recommendation model includes a backbone network and at least one gain network; the backbone network is used to learn the user's preference for the similar objects; the gain network is used to learn the preference gain corresponding to the user's preference based on the information difference between the historical objects queried by the user and the recommended similar objects corresponding to the historical objects; the preference gain represents the degree of preference of the user's preference; the device 140 includes:

[0232] The second acquisition module 1401 acquires the object queried by the user and constructs a recommendation data sample based on the object queried by the user and similar objects corresponding to the object to be recommended; wherein the recommendation data sample includes attribute information of the object queried by the user and attribute information of the similar objects to be recommended;

[0233] The second calculation module 1402 inputs the recommendation data sample into the backbone network for learning calculation, and obtains the user preference of the user for the similar object output by the backbone network; and calculates the information difference between the attribute information of the object queried by the user contained in the recommendation data sample and the attribute information of the similar object to be recommended, further inputs the calculated information difference into the gain network for learning calculation, and obtains the preference gain of the information difference output by the gain network with respect to the user preference;

[0234] The recommendation module 1403 determines whether to recommend the similar object to the user based on the user preference learned by the backbone network and the preference gain learned by the gain network.

[0235] Accordingly, this specification also provides an electronic device, which includes a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the steps in all the method flows described above.

[0236] Accordingly, this specification also provides a computer-readable storage medium on which executable computer program instructions are stored; wherein, when the instructions are executed by a processor, the steps in all the method flows described above are implemented.

[0237] Accordingly, this specification also provides a computer program product having executable computer program instructions stored thereon; wherein, when the computer program instructions are executed by a processor, the steps in all the method flows described above are implemented.

[0238] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a server system. Of course, it is not excluded that with the future development of computer technology, the computer implementing the functions of the above embodiments may be, for example, a personal computer, a laptop computer, an in-vehicle human-computer interaction device, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0239] Although one or more embodiments of this specification provide method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-inventive means. The order of steps listed in the embodiments is only one way of executing the steps among many, and does not represent the only execution order. When a device or terminal product is actually executed, the methods shown in the embodiments or figures may be executed sequentially or in parallel (for example, in a parallel processor or multi-threaded processing environment, or even a distributed data processing environment). The terms "comprise," "include," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, product, or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, product, or device. Without further limitation, it does not exclude that the process, method, product, or device that includes the elements may also have other identical or equivalent elements. For example, if words such as first and second are used to indicate names, they do not indicate any particular order.

[0240] For the convenience of description, the above devices are described in terms of functions divided into various modules. Of course, when implementing one or more of the present specifications, the functions of each module can be implemented in the same or multiple software and / or hardware, or the module that implements the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0241] The present invention is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0242] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0243] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.

[0244] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0245] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0246] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage, graphene storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0247] Those skilled in the art will appreciate that one or more embodiments of this specification may be provided as methods, systems, or computer program products. Thus, one or more embodiments of this specification may take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.

[0248] One or more embodiments of this specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In distributed computing environments, program modules may be located in local and remote computer storage media, including storage devices.

[0249] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between the various embodiments can be referenced across them. Each embodiment focuses on the differences from the other embodiments. In particular, since the system embodiments are generally similar to the method embodiments, their description is relatively simple. For relevant parts, reference can be made to the description of the method embodiments. Throughout this specification, reference to the terms "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples. Furthermore, those skilled in the art may combine and integrate the different embodiments or examples, and features of different embodiments or examples, described in this specification, without conflict.

[0250] The foregoing is merely an example of one or more embodiments of this specification and is not intended to limit the one or more embodiments of this specification. It will be apparent to those skilled in the art that various modifications and variations may be made to one or more embodiments of this specification. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this specification shall be included within the scope of the claims.

Claims

1. A method for optimizing a recommendation model, wherein the recommendation model is used to recommend similar objects corresponding to the objects queried by the user to a user; the recommendation model comprises a backbone network and at least one gain network; wherein: The backbone network is used to learn the user's preference for the similar objects; the gain network is used to learn the preference gain corresponding to the user's preference according to the difference in information contained in the historical objects queried by the user and the recommended similar objects corresponding to the historical objects; The preference gain indicates the preference degree of the user's preference; the method comprises: Acquire a training sample set for training the recommendation model; wherein the training sample set includes a number of historical operation behavior data samples; the historical operation behavior data samples include attribute information of historical objects queried by the user, attribute information of similar objects that have been recommended and corresponding to the historical objects, and a label indicating whether the user has performed an operation on the similar objects; Inputting the historical operation behavior data samples in the training sample set into the backbone network for supervised learning training, and calculating the user preference loss of the user preference learned by the backbone network based on the output result of the backbone network; and calculating the information difference between the attribute information of the historical object and the attribute information of the similar object contained in the historical operation behavior data samples, further inputting the calculated information difference into the gain network for supervised learning training, and calculating the preference gain loss of the preference gain learned by the gain network based on the output result of the gain network; The user preference loss and the preference gain loss are used as the training losses of the recommendation model, and the model parameters contained in the recommendation model are adjusted to complete the learning and training of the recommendation model; wherein the trained recommendation model is used to recommend similar objects corresponding to the objects queried by the user to the user based on the user preferences learned by the backbone network and the preference gains learned by the gain network.

2. The method according to claim 1, wherein the object comprises a commodity; the attribute information comprises price information of the commodity and context information of the commodity; The at least one gain network comprises: A first gain network and a second gain network; wherein the first gain network is used to learn a price preference gain corresponding to the user preference according to a price difference between a historical commodity queried by the user and a recommended similar commodity corresponding to the historical commodity; and the second gain network is used to learn a context information gain corresponding to the user preference according to a context information difference between a historical commodity queried by the user and a recommended similar commodity corresponding to the historical commodity; Calculating the information difference between the attribute information of the historical object and the attribute information of the similar object contained in the historical operation behavior data sample, further inputting the calculated information difference into the gain network for supervised learning training, and calculating the preference gain loss of the preference gain learned by the gain network based on the output result of the gain network, including: Calculating the price difference between the price information of the historical commodities and the price information of the similar commodities contained in the historical operation behavior data sample, further inputting the calculated price difference into the first gain network for supervised learning training, and calculating the price preference gain loss of the price preference gain learned by the first gain network based on the output result of the first gain network; Calculate the context information difference between the context information of the historical commodity and the context information of the similar commodity contained in the historical operation behavior data sample, further input the calculated context information into the second gain network for supervised learning training, and calculate the context information gain loss of the context information gain learned by the second gain network based on the output result of the second gain network.

3. The method according to claim 2, before calculating the price difference between the price information of the historical commodity and the price information of the similar commodity contained in the historical operation behavior data sample, further comprising: Searching for a difference data sample corresponding to the historical operation behavior data sample in the training sample set; wherein the difference data sample has a label different from that of the historical operation behavior data sample; and the difference data sample is a data sample generated before the historical operation behavior data sample and contains data content that is completely the same as that of the historical operation behavior data sample; Calculating the price difference between the price information of the historical commodity and the price information of the similar commodity contained in the historical operation behavior data sample, and further inputting the calculated price difference into the first gain network for supervised learning training, including: Calculating a first price difference between the price information of the historical commodity and the price information of the similar commodity contained in the historical operation behavior data sample, and calculating a second price difference between the price information of the historical commodity and the price information of the similar commodity contained in the difference data sample; The calculated first price difference and the second price difference are further input into the first gain network for supervised learning training; Calculating the context information difference between the context information of the historical commodity and the context information of the similar commodity contained in the historical operation behavior data sample, and further inputting the calculated context information into the second gain network for supervised learning training, including: Calculating a first context information difference between the context information of the historical commodity and the context information of the similar commodity contained in the historical operation behavior data sample, and calculating a second context information difference between the context information difference of the historical commodity and the context information difference of the similar commodity contained in the difference data sample; The calculated first context information difference and the second context information difference are further input into the second gain network for supervised learning and training.

4. The method according to claim 2, wherein the user preference loss and the preference gain loss are used as training losses of the recommendation model, and the model parameters included in the recommendation model are adjusted, comprising: The user preference loss, the price preference gain loss and the context information gain loss are used as training losses of the recommendation model, and the model parameters included in the recommendation model are adjusted.

5. The method according to claim 4, wherein the recommendation model further comprises a scoring network; the scoring network is used to integrate the user preference learned by the backbone network, the price preference gain learned by the first gain network, and the context information gain learned by the second gain network to perform scoring processing to obtain a recommendation score; The recommendation model is used to classify the similar commodities contained in the historical operation behavior data sample based on the recommendation score obtained by the scoring network; wherein, The classification result of the classification process includes a first classification result in which the user will perform an operation on the similar products and a second classification result in which the user will not perform an operation on the similar products.

6. The method of claim 5, wherein the optimization goal of learning and training the first gain network is to maximize the gap between the first price preference gain and the second price preference gain; The first price preference gain is a price preference gain learned by the first gain network based on the first price difference; The second price preference gain is the price preference gain learned by the first gain network based on the second price difference; The optimization objective of learning and training the first gain network is expressed by the following expression: The loss function of the first gain network is expressed as follows: Among them, in the above expression, Represents the Sigmoid function represents the price preference gain loss of the first gain network mentioned above; represents the first price preference gain mentioned above; Represents the second price preference gain mentioned above.

7. The method of claim 6, wherein the optimization goal of learning and training the first gain network is to maximize the gap between the first price preference gain and the second price preference gain; and the gap between the classification result obtained by the recommendation model by performing the classification processing on the similar commodities included in the historical operation behavior data sample and the classification result indicated by the label of the historical operation behavior data sample; The optimization objective of learning and training the first gain network is expressed by the following expression: The loss function of the first gain network is expressed as follows: in, In the above expression, represents a classification result obtained by the recommendation model performing the classification process on the similar commodities included in the historical operation behavior data sample; The classification result indicated by the label of the historical operation behavior data sample is represented.

8. The method as claimed in claim 5, wherein the optimization goal of learning and training the second gain network is to maximize the difference between the first context information gain and the second context information gain; the first context information gain is the context information gain learned by the second gain network based on the first context information; the second context information gain is the context information gain learned by the second gain network based on the difference of the second context information; The optimization objective of learning and training the second gain network is expressed by the following expression: The loss function of the second gain network is expressed as follows: in, In the above expression Represents the Sigmoid function represents the contextual information gain loss of the second gain network; represents the first context information gain; represents the second context information gain; m represents the number of historical commodities queried by the user contained in the historical operation behavior data sample.

9. An object recommendation method based on a recommendation model, wherein the recommendation model is used to recommend similar objects corresponding to the object queried by the user to the user; the recommendation model includes a backbone network and at least one gain network; wherein, The backbone network is used to learn the user's preference for the similar objects; the gain network is used to learn the preference gain corresponding to the user's preference according to the difference in information contained in the historical objects queried by the user and the recommended similar objects corresponding to the historical objects; The preference gain indicates the preference degree of the user's preference; the method comprises: Acquire the object queried by the user, and construct a recommendation data sample based on the object queried by the user and the similar objects corresponding to the object to be recommended; wherein the recommendation data sample includes the attribute information of the object queried by the user and the attribute information of the similar objects to be recommended; Inputting the recommended data sample into the backbone network for learning calculation, and obtaining the user preference of the user for the similar object output by the backbone network; and calculating the information difference between the attribute information of the object queried by the user contained in the recommended data sample and the attribute information of the similar object to be recommended, further inputting the calculated information difference into the gain network for learning calculation, and obtaining the preference gain corresponding to the user preference output by the gain network; Based on the user preference learned by the backbone network and the preference gain learned by the gain network, it is determined whether to recommend the similar object to the user.

10. A recommendation model based on a neural network, wherein: The recommendation model is used to recommend similar objects corresponding to the object queried by the user to the user, including: A backbone network, used for learning the user's preference for the similar objects based on the user's historical operation behavior data for the similar objects; at least one gain network, configured to learn a preference gain corresponding to the user preference according to a difference in information between a historical object queried by the user and a recommended similar object corresponding to the historical object; the preference gain indicates a preference degree of the user preference; The training loss of the recommendation model includes the user preference loss of the user preference learned by the backbone network; and the preference gain loss of the preference gain learned by the gain network.

11. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 are implemented.

12. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Recommendation model training method and device, recommendation method and device, equipment and medium

    CN112632403A

  • Object recommendation method and device, computer equipment and storage medium

    CN117251622A