Object recommendation method and device, equipment, storage medium and program product

By obtaining user behavior data and using feature extraction and recommendation models, the problem of over-focusing and homogeneous recommendation results in the existing recommendation system is solved, and accurate recommendations of the objects of interest to users are achieved and user experience is improved.

CN119988728APending Publication Date: 2025-05-13BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510059329.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the content business and object processing business, the recommendation results are too focused and homogeneous, so that the existing recommendation system cannot accurately and effectively recommend objects.

Method used

By obtaining user behavior data, including content browsing behavior and object operation behavior, the trained feature extraction model generates behavior feature embeddings, and the recommendation model is used to determine the object recommendation results based on behavior feature embeddings and object feature embeddings.

Benefits of technology

It realizes object recommendation based on user historical behavior, and can connect the user's content browsing behavior with object operation behavior, thereby identifying relevant objects that are of interest to the user and improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988728A_ABST
    Figure CN119988728A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an object recommendation method and device, equipment, a storage medium and a program product. The method comprises the following steps: acquiring behavior data of a user, wherein the behavior data comprises first behavior data associated with a content browsing behavior and second behavior data associated with an object operation behavior; determining a content type corresponding to the first behavior data and an object type corresponding to the second behavior data; utilizing a trained feature extraction model to generate behavior feature embedding for the user based on the first behavior data, the content type corresponding to the first behavior data, the second behavior data and the object type corresponding to the second behavior data; and using the trained recommendation model to determine an object recommendation result for the user based on behavior feature embedding and object feature embedding, the object feature embedding being associated with the object type of the at least one object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments of the present disclosure generally relate to the field of computer technology, and more particularly, to methods, devices, apparatuses, computer-readable storage media, and computer program products for object recommendation. Background Art

[0002] With the development of machine learning technology, some algorithms and deep models for learning general representations have emerged in recommendation systems. For example, they can be applied to some downstream tasks by pre-training some features to a certain extent. These downstream tasks may include cross-domain recommendations. However, at present, whether it is a recommendation system for content services (such as media content services) or a recommendation system for object processing services, the recommendation results are usually too focused and too homogeneous, and cannot accurately and effectively recommend objects. Summary of the invention

[0003] In a first aspect of the present disclosure, a method for object recommendation is provided. The method includes: obtaining user behavior data, the behavior data including first behavior data associated with content browsing behavior and second behavior data associated with object operation behavior; determining the content type corresponding to the first behavior data and the object type corresponding to the second behavior data; using a trained feature extraction model, based on the first behavior data, the content type corresponding to the first behavior data, the second behavior data and the object type corresponding to the second behavior data, generating a behavior feature embedding for the user; and using a trained recommendation model, based on the behavior feature embedding and the object feature embedding, determining an object recommendation result for the user, the object feature embedding being associated with the object type of at least one object.

[0004] In a second aspect of the present disclosure, a device for object recommendation is provided. The device includes: a behavior data acquisition module, configured to acquire the behavior data of a user, the behavior data including first behavior data associated with content browsing behavior and second behavior data associated with object operation behavior; a type determination module, configured to determine the content type corresponding to the first behavior data and the object type corresponding to the second behavior data; a feature embedding generation module, configured to generate a behavior feature embedding for the user based on the first behavior data, the content type corresponding to the first behavior data, the second behavior data and the object type corresponding to the second behavior data using a trained feature extraction model; and a recommendation result determination module, configured to determine an object recommendation result for the user based on the behavior feature embedding and the object feature embedding using a trained recommendation model, the object feature embedding being associated with the object type of at least one object.

[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory, the at least one memory is coupled to the at least one processing unit and stores instructions for execution by the at least one processing unit. When the instructions are executed by the at least one processing unit, the device executes the method of the first aspect.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the medium, and when the computer program is executed by a processor, the method of the first aspect is implemented.

[0007] In a fifth aspect of the present disclosure, a computer program product is provided, wherein the computer program product comprises a computer program, and when the computer program is executed by a processor, the method of the first aspect is implemented.

[0008] It should be understood that the content described in this section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0010] Figure 1 A schematic diagram showing an example environment in which embodiments of the present disclosure can be implemented;

[0011] Figure 2 A schematic diagram showing an example architecture for object recommendation at the inference stage according to some embodiments of the present disclosure;

[0012] Figure 3 A schematic diagram showing an example training architecture of a feature extraction model and a recommendation model in a training phase according to some example embodiments of the present disclosure;

[0013] Figure 4 A schematic diagram showing an example architecture for object recommendation according to some embodiments of the present disclosure;

[0014] Figure 5 A block diagram showing an example apparatus for object recommendation according to some embodiments of the present disclosure; and

[0015] Figure 6 A block diagram of an electronic device capable of implementing one or more embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0016] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0017] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below.

[0018] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.

[0019] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0020] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information, so that the user can independently choose whether to provide personal information to software or hardware such as electronic devices, applications, servers or storage media that perform operations of the technical solution of the present disclosure based on the prompt message.

[0021] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information is sent to the user in a manner such as a pop-up window, in which the prompt information can be presented in text form. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0022] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0023] As used herein, the term "model" can learn the association between the corresponding input and output from the training data, so that after the training is completed, the corresponding output can be generated for a given input. The generation of the model can be based on machine learning technology. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multi-layer processing units. A neural network model is an example of a model based on deep learning. In this article, "model" may also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms are used interchangeably in this article.

[0024] A "neural network" is a machine learning network based on deep learning. A neural network is capable of processing inputs and providing corresponding outputs, and typically includes an input layer and an output layer and one or more hidden layers between the input layer and the output layer. Neural networks used in deep learning applications typically include many hidden layers, thereby increasing the depth of the network. The layers of a neural network are connected in sequence so that the output of the previous layer is provided as input to the next layer, where the input layer receives the input of the neural network and the output of the output layer serves as the final output of the neural network. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each of which processes input from the previous layer.

[0025] Generally, machine learning can be roughly divided into three stages, namely the training stage, the testing stage, and the application stage (also called the inference stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter values ​​are continuously updated iteratively until the model can obtain consistent inferences that meet the expected goals from the training data. Through training, the model can be considered to be able to learn the association from input to output (also called the mapping of input to output) from the training data. The parameter values ​​of the trained model are determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, thereby determining the performance of the model. In the application stage, the model can be used to process the actual input based on the parameter values ​​obtained from the training to determine the corresponding output.

[0026] Figure 1 A schematic diagram of an example environment 100 in which an embodiment of the present disclosure can be implemented is shown. The environment 100 may include a server device 110. The server device 110 is configured to use a feature extraction model 112 and a recommendation model 114 to perform a prediction result about an object in a predetermined prediction task. The prediction task about the object can be defined in different scenarios. For example, in a commodity recommendation scenario, the object may be, for example, a commodity, and the prediction task may include a prediction of an execution operation such as a click or purchase of the commodity. It should be understood that in other scenarios, the prediction task may be any appropriate type of task that is suitable for the scenario.

[0027] The environment 100 may also include a terminal device 120, which may communicate with the server device 110. The terminal device 120 may support interaction with the user 140 and generate interaction behavior data. The terminal device 120 or the target application 125 in the terminal device 120 may receive behavior data including, for example, clicks, browses, purchases, likes, comments, and the like. The interaction behavior data of the terminal device 120 may be transmitted to the server device 110 with the authorization of the user and in compliance with relevant laws and regulations.

[0028] In some embodiments, the input of the feature extraction model 112 may include user behavior data, and the output may include feature embedding, such as behavioral feature embedding for the user. The input of the recommendation model 114 may include the above-mentioned feature embedding, and the output may include the prediction result for the object. The user behavior data may be, for example, behavior data generated based on the interaction between the user 140 and the target application 125. The corresponding prediction task can be performed based on these behavior data, and the recommendation result can be generated.

[0029] In addition, the terminal device 120 may support interaction with the user 140 through the interface 130 . Moreover, the terminal device 120 may display the recommendation result on the interface 130 .

[0030] According to application requirements, the feature extraction model 112 and the recommendation model 114 can be constructed to include one or more appropriate types of model architectures. The embodiments of the present disclosure do not limit the specific types and structures of the feature extraction model 112 and the recommendation model 114.

[0031] In environment 100, terminal device 130 can be any type of mobile terminal, fixed terminal or portable terminal, including mobile phone, desktop computer, laptop computer, notebook computer, netbook computer, tablet computer, media computer, multimedia tablet, personal communication system (PCS) device, personal navigation device, personal digital assistant (PDA), audio / video player, digital camera / camcorder, positioning device, television receiver, radio broadcast receiver, e-book device, gaming device or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, terminal device 130 can also support any type of interface for the user (such as "wearable" circuit, etc.). Server device 110, for example, can be implemented in various types of computing systems / servers that can provide computing capabilities, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, and the like.

[0032] It should be understood that the structure and function of the various elements in the environment 100 are described for exemplary purposes only and do not imply any limitation on the scope of the present disclosure.

[0033] At present, with the development of machine learning technology, a variety of classic visual models have emerged, and universal language / visual representation capabilities have been realized. The learned language representation can be applied to a variety of downstream tasks. Recently, some algorithms and deep models for learning universal user representations have also appeared in the recommendation system, that is, by pre-training user features to a certain extent and then adapting them to some downstream tasks. These downstream tasks can include cross-domain recommendations, for example. This type of technology can be called user representation recognition.

[0034] At present, in the recommendation systems of content businesses (such as media content businesses) and e-commerce businesses, when modeling user product preferences, the public desensitization behavior of users in the e-commerce domain has been captured relatively fully, resulting in recommendations that are too focused and too homogeneous. However, the real interest range and needs of users are not only reflected in past e-commerce purchase behaviors. The important feature of e-commerce products related to browsing content is to help users better discover and meet their needs through rich content. Therefore, we hope to help users discover possible e-commerce purchase interests through the rich content behaviors in public scenes such as short videos or live broadcasts, so that recommendations are richer and in line with the user's interest range.

[0035] Taking the U-BERT structure (Pre-training User Representations for Improved Recommendation) as an example, users' cross-category preferences can be modeled based on the same interests shown in their comments on different product types (hereinafter also referred to as categories). For example, the comments under toys and electric vehicle accessories all emphasize "fast logistics and good experience." In terms of the model, U-BERT introduces a comment encoder based on a multi-layer Transformer and a user encoder to model the comment text and construct a comment-enhanced user representation. The U-BERT structure proposes two pre-training tasks: predicting masked comment words and predicting review scores; in the fine-tuning stage, U-BERT uses a product encoder to represent products and a comment connection matching (co-match) layer to capture the semantic relevance between user comments and product reviews.

[0036] However, U-BERT has the following shortcomings. First, U-BERT only identifies the interests of the same user between different categories in the e-commerce domain (for example, migrating from electric vehicle accessories to the toy category), and does not solve the problem of cross-domain interest identification from the content domain to the e-commerce domain. Secondly, U-BERT uses the BERT pre-training framework, and the core pre-training task is to predict masked words (Predict masked word). The input of this model is global information, that is, the user's desensitized information regardless of time sequence and state. Therefore, this does not take into account the order of user behavior, and it is impossible to predict the next moment's behavior based on the user's previous known behavior.

[0037] Therefore, it is expected that when recommending objects (such as commodities) to users, effective and accurate object recommendations can be made for users.

[0038] In view of this, an embodiment of the present disclosure proposes an improved scheme for object recommendation. Specifically, in an embodiment of the present disclosure, the user's behavior data is obtained, and the content type corresponding to the first behavior data and the object type corresponding to the second behavior data are determined. The behavior data includes first behavior data associated with content browsing behavior and second behavior data associated with object operation behavior. Using a trained feature extraction model, a behavior feature embedding for the user is generated based on the first behavior data, the content type corresponding to the first behavior data, the second behavior data, and the object type corresponding to the second behavior data. Then, using the trained recommendation model, based on the behavior feature embedding and the object feature embedding associated with the object type of at least one object, an object recommendation result for the user is determined.

[0039] In this way, it is possible to provide users with recommendation results of related objects based on the user's historical content browsing behavior and object operation behavior. This can establish a connection between the user's rich content browsing behavior and the predicted object operation behavior, and predict the user's interested related objects from the content browsing behavior. In this way, the user's operation needs for the object of interest can be met, thereby improving the user experience.

[0040] Some example embodiments of the present disclosure will be described below with reference to the accompanying drawings.

[0041] Figure 2 FIG. 2 is a schematic diagram of an example architecture 200 for object recommendation in the inference phase according to some embodiments of the present disclosure. Figure 1 These embodiments are described in detail with reference to the environment 100 of FIG. In some examples, these embodiments can be implemented in Figure 1 In some other examples, these embodiments can also be implemented in Figure 1The server device 110 and the terminal device 120 may be used to implement the server device 110, or the server device 110 and the terminal device 120 may cooperate to implement the server device 110. The following specific embodiments are implemented at the server device 110 as an example.

[0042] like Figure 2 As shown, in the application phase (also called the inference phase), the server device 110 obtains the behavior data 210 of the user 140. As an example, the server device 110 may obtain the behavior data 210 of the user 140 from the terminal device 120. The terminal device 120 may support interaction with the user 140, thereby generating data related to the interaction behavior of the user 140. It should be understood that the user's behavior data obtained through the terminal device 120 is anonymized and authorized by the user, and complies with the requirements of relevant laws, regulations and relevant provisions.

[0043] Here, the behavior data 210 includes first behavior data 212 associated with content browsing behavior and second behavior data 214 associated with object operation behavior. As an example, the target application 125 installed in the terminal device 120 can support the user 140 to browse the content in the target application 125 and operate the objects in the target application 125.

[0044] In some embodiments, the content provided by the terminal device 120 or the target application 120 for the user to browse may include but is not limited to media content. Media content may include, for example, videos, images, animations, texts, etc., which are not limited here. Accordingly, the content browsing behavior may include, for example, the behavior of the user 140 browsing the media content such as videos, animations, or other content displayed on the interface 130 of the terminal device 120. The first behavior data 212 may include data related to such content browsing behavior.

[0045] As an example, the object may include an online commodity, or any other suitable object that can be recommended to the user through the terminal device 120. Accordingly, the object operation behavior may include, for example, the user 140's operation behavior such as clicking or purchasing the online commodity or other object displayed on the terminal device 120. In other examples, the object operation behavior may also include other suitable operation behaviors on the object. Accordingly, the second behavior data 214 may include data related to these object operation behaviors.

[0046] In some embodiments, the behavior data 210 of the user 140 may include the behavior data of the user 140 within a predetermined time period. Accordingly, the behavior data 210 may include first behavior data 212 associated with content browsing behavior within a predetermined time period, and second behavior data 214 associated with object operation behavior within the predetermined time period. The predetermined time period may be, for example, 5 minutes, 1 day, 1 week, or other time periods. As an example, in order to be able to recommend objects of interest to the user, the predetermined time period may be a shorter time period before the current time, such as 5 minutes, half a day, 1 day, or other shorter time periods before the current time. It should be understood that the predetermined time period may be set according to the specific application scenario, and no limitation is intended here.

[0047] Further, the server device 110 determines the content type corresponding to the first behavior data 212 and the object type corresponding to the second behavior data 214. Taking the content including video content as an example, the corresponding content type may include types related to the content scene, such as travel, food, beauty, etc. For the object type, for example, types related to the object attributes may be included, such as dresses, mobile phones, water cups, etc., which can all be used as object types. It should be understood that the division of content types and object types can be set according to specific application situations, and no limitation is intended here.

[0048] Continue to refer Figure 2 In some embodiments, in order to determine the content type corresponding to the first behavior data 212 and the object type corresponding to the second behavior data 214, the server device 110 may determine the content type 232 corresponding to the first behavior data 212 from the content library 222, and determine the object type 234 corresponding to the second behavior data 214 from the object library 224. Here, the content library 222 may include multiple candidate content types, and the object library 224 may include multiple candidate object types.

[0049] For example, assuming that the content library includes multiple candidate content types such as travel, food, and beauty, and assuming that the user 140 browses a makeup video through the terminal device 120, it can be determined that the content type corresponding to the browsing behavior data for the content is beauty. For another example, assuming that the object library includes multiple candidate object types such as dresses, mobile phones, and water cups, and assuming that the user 140 clicks on a water cup through the terminal device 120, it can be determined that the object type corresponding to the data of such operation behavior is a water cup.

[0050] As an example, labels may be used to represent each candidate content type and each candidate object type in the content library 222 and the object library 224. In other examples, any other suitable characters, symbols, etc. that can play an identification role may also be used to represent the candidate content type and the candidate object type.

[0051] Thus, through the content library 222 and the object library 224, the respective word libraries of the content domain and the object domain can be established to express the user's preferences in different categories. This is also conducive to expanding and classifying the categories of content or objects.

[0052] Further, the server device 110 uses the trained feature extraction model 112 to generate a behavior feature embedding 240 for the user 140 based on the first behavior data 212, the content type 232 corresponding to the first behavior data, the second behavior data 214, and the object type 234 corresponding to the second behavior data. In some embodiments, the feature extraction model 112 can be configured to extract specific features from the user's behavior data and the corresponding type, and map these features to vector representations. The feature extraction model 112 can be constructed as a deep neural network including multiple network layers. The multiple network layers may include, for example, a masked multi-head self-attention mechanism, layer normalization, forward feedback, etc. The following will be combined with Figure 3 Let's discuss the training process of the feature extraction model 112 in detail.

[0053] Furthermore, the server device 110 uses the trained recommendation model 114 to determine an object recommendation result 260 for the user 140 based on the behavior feature embedding 240 and the object feature embedding 250. The object feature embedding 250 may be associated with an object type of at least one object.

[0054] In some embodiments, in order to determine the object recommendation result for the user 140, the server device 110 can use the trained recommendation model 114 to generate a predicted execution result of performing a target operation on at least one object based on the behavior feature embedding 240 and the object feature embedding 250. The predicted execution result can indicate the possibility of the user 140 performing the target operation on each of the at least one object. Then, based on the predicted execution result, the server device 110 can determine the object recommendation result 260 for the user 140.

[0055] In some embodiments, the recommendation model 114 may be configured to generate corresponding prediction results based on specific feature embeddings. As an example, the recommendation model 114 may be constructed based on a feedforward neural network (FeedForward) architecture. In other examples, the recommendation model 114 may also be constructed based on any other appropriate machine learning model. Figure 3 The training process of the recommendation model 114 is discussed in detail.

[0056] In some embodiments, the target operation may include at least one of a click operation and a purchase operation. Accordingly, the predicted execution result may indicate the possibility of user 140 performing a click operation on each of at least one object, and may also indicate the possibility of user 140 performing a purchase operation on each object. The possibility may represent a probability. As an example, when predicting the probability of user 140 performing a click operation on an object, if user 140 performs a click operation, the probability may be 1. Conversely, if user 140 does not perform a click operation, the probability may be 0. In addition, the probability may also be other probabilities between 0 and 1. In other words, it is possible to predict how likely user 140 is to perform a click operation. Of course, when predicting the probability of user 140 performing a purchase operation on an object, it is similar to the aforementioned example and will not be repeated here.

[0057] Furthermore, based on the possibility of the user 140 performing the target operation on each of the at least one object indicated by the predicted execution result, it can be determined to recommend one or more of the at least one object to the user 140. As an example, the predicted probability of the user 140 performing the target operation on a certain object can be compared with a preset probability threshold to determine whether to recommend the object to the user 140. For example, when the predicted probability exceeds the preset probability threshold, it can be determined to recommend the object to the user 140. Conversely, when the predicted probability does not exceed the preset probability threshold, it can be determined not to recommend the object to the user 140. The preset probability threshold can be a larger probability threshold, such as 0.6, 0.7 or other larger probability thresholds. In this way, it is possible to recommend objects of interest to the user 140.

[0058] The object feature embedding 250 may be determined in a variety of ways. For example, the server device 110 may determine the object type corresponding to at least one object from the object library 224, and generate the object feature embedding 250 for the at least one object based on the object type corresponding to the at least one object. As an example, the at least one object may include a related object in the second behavior data 214 of the user 140, or may include an object selected from the most recent popular objects (e.g., popular online products or other popular objects).

[0059] In some embodiments, a feedforward neural network architecture and layer normalization may be used to generate an object feature embedding 250 based on an object type corresponding to at least one object. Specifically, the object type corresponding to at least one object may be input into the feedforward neural network architecture, and then the object feature embedding 250 may be output after layer normalization. The object feature embedding 250 may be a multidimensional vector capable of characterizing at least one object.

[0060] Therefore, the embodiments of the present disclosure can provide users with recommendation results of related objects based on the user's content browsing behavior. For example, assuming that the user has previously browsed some makeup tutorials, then cosmetics-related products can be recommended to the user, even if the user has not purchased cosmetics in the history. Therefore, the embodiments of the present disclosure can establish a connection between the user's rich content browsing behavior and the object operation behavior, and identify the user's interest in related objects from the content browsing behavior. This can meet the user's various needs for objects of interest and improve the user experience.

[0061] The above discusses the reasoning process of the feature extraction model 112 and the recommendation model 114 in the application phase. The following will discuss in detail the training process of the feature extraction model 112 and the recommendation model 114 in the training phase.

[0062] Figure 3 A schematic diagram of an example training architecture 300 of the feature extraction model 112 and the recommendation model 114 in the training phase according to some example embodiments of the present disclosure is shown. In the training architecture 300, the training phase of the feature extraction model 112 may also be referred to as a pretraining phase 310. The training phase of the recommendation model 114 may also be referred to as a fine-tuning phase 320.

[0063] In some embodiments, the feature extraction model 112 may be trained based on a first training data set. The first training data set may be constructed using unsupervised autoregressive corpus.

[0064] In some embodiments, the first training data set can be determined in the following manner: the server device 110 can obtain the first sample behavior data, and the first sample behavior data can include data associated with the content browsing behavior of the first group of users and data associated with the object operation behavior of the first group of users. The content browsing behavior and the object operation behavior have been discussed above and will not be repeated here. Then, based on the first sample behavior data and the content library 222 and the object library 224, the server device 110 can determine the first training data set.

[0065] As an example, the server device 110 can determine the content type corresponding to the data associated with the content browsing behavior of the first group of users from the content library 222, and determine the object type corresponding to the data associated with the object operation behavior of the first group of users from the object library 224. Thus, the first sequence 312 of each user in the first group of users and the second sequence 314 corresponding to the first sequence 312 can be determined, and the first training data set is determined based on the first sequence 312 and the second sequence 314 of each user. Assuming that ContDict is used to represent the content library 222, EcomDict is used to represent the object library 224, and PretrainInput1 is used to represent the first sequence 312, and PretrainInput2 is used to represent the second sequence 314, the content library 222, the object library 224, the first sequence 312, and the second sequence 314 can be represented as the following sets, respectively.

[0066] ContDict={conCate1,conCate2,...leatCate N}, N = first predetermined number

[0067] EcomDict={leafCate1,leafCate2,...leatCate M}, M = second predetermined number

[0068] PretrainInput1 = {Token1, Token2, Token3, ..., TokenK}, K = a third predetermined number

[0069] PretrainInput2={TokenDomain1, TokenDomain2, TokenDomain3,…,TokenDomainK},

[0070] K = third predetermined number

[0071] In the above four sets, conCate can represent content type, leafCate can represent object type, Token can represent feature unit, TokenDomain can represent type unit, the first predetermined number N represents the number of content types, the second predetermined number M represents the number of object types, and the third predetermined number K represents the number of feature units.

[0072] In some embodiments, the first training data set may include a first sequence 312 for each user in the first group of users and a second sequence 314 corresponding to the first sequence 312. The first sequence 312 may include a plurality of feature units, each of which includes a content type determined from the content library 222 or an object type determined from the object library 224. The second sequence 314 includes a plurality of type units corresponding to the plurality of feature units, each of which may indicate whether the feature unit corresponding to the type unit in the first sequence 312 belongs to a content type or an object type. For example, if Token1 indicates a certain object type, TokenDomain1 may indicate an object. If Token2 indicates a certain content type, TokenDomain2 may indicate content. If Token3 indicates another object type, TokenDomain3 may indicate an object, and so on.

[0073] Therefore, the feature extraction model 112 of the embodiment of the present disclosure uses the sum of feature unit embedding (token embedding) and domain embedding (domain embedding) to encode the input by inputting the first sequence and the second sequence, which is different from the traditional position embedding (position embedding) and segment embedding (segment embedding) at the input end. The feature extraction model 112 maps the Token to one of {EcomDict, ContDict} at the input end, and maps the Domain to the type of the domain (one of the content domain and the object domain), thereby focusing on representing the difference between the two domains of the content stream and the object stream, and fitting the difference between the two domains in the model.

[0074] In some embodiments, the feature extraction model 112 can be constructed using a Transformer framework. Figure 3 The Transformer framework used in the training architecture 300 may include a multi-layer network structure combination of a masked multi-head self-attention mechanism 316, layer normalization 317, feedforward 318, and layer normalization 319. It should be understood that the feature extraction model 112 may also be constructed using any other appropriate neural network architecture that can be used to extract features from input data.

[0075] In some embodiments, the feature extraction model 112 may be trained in the following manner: the server device 110 may use the feature extraction model 112 to be trained, based on the first sequence 312 and the second sequence 314, to generate a first prediction result corresponding to the first sequence 312 and a second prediction result corresponding to the second sequence 314, respectively. Here, the feature extraction model 112 may be made to adopt an autoregressive form, so as to predict the next feature unit (313) according to the input first sequence 312, and predict the next type unit (315) according to the input second sequence 314. Then, the predicted feature unit and the first sequence 312, and the predicted type unit and the second sequence 314 are used as inputs of the feature extraction model 112, and then the next feature unit and type unit are predicted, and so on.

[0076] Furthermore, the server device 110 may determine a first loss corresponding to the first prediction result and a second loss corresponding to the second prediction result. As an example, the first loss and the second loss may be constructed by maximizing the likelihood estimation function. The first loss function (which may be L t to represent) and the second loss function (which can be represented by L d The example formulas for expressing ) are as follows:

[0077] L t =∑ i logP(t i |t0,t1,...,t i-1 ) (1)

[0078] L d =∑ i logP(d i |d0,d1,...,d i-1 ) (2) Among them, P can represent probability; (t0, t1, ..., t i-1 ) can represent the 0th to i-1th feature units; t i It can represent the i-th feature unit, that is, the next feature unit after the i-1-th feature unit. (d0, d1, ..., d i-1 ) represents the 0th to i-1th type units; d i It can represent the i-th type unit, that is, the next type unit after the i-1-th type unit.

[0079] Further, the feature extraction model 112 may be trained based on the first loss and the second loss. As an example, the feature extraction model 112 may be trained based on the sum of the first loss and the second loss. That is, the feature extraction model 112 is trained by minimizing the sum of the first loss and the second loss as a training goal. In other examples, the feature extraction model 112 may also be trained based on other operations of the first loss and the second loss.

[0080] In some embodiments, the recommendation model 114 may be trained based on a second training data set. The second training data set may be constructed based on randomly sampled desensitized behavioral data related to object operations. The behavioral data related to object operations may include, for example, data related to behaviors such as clicking on objects and purchasing objects. In some embodiments, the second training data set may include multiple positive sample pairs and multiple negative sample pairs. Each positive sample pair may include a first user and a first sample object corresponding to the first user, and the first sample object may be an object on which the first user has performed a target operation. Each negative sample pair may include a second user and a second sample object corresponding to the second user, and the second sample object may be an object on which the second user has not performed a target operation.

[0081] For example, the first sample object may be a sample object on which the first user has performed a click operation, a purchase operation, or other target operations, and the second sample object may be a sample object on which the second user has not performed a click operation, a purchase operation, or other target operations.

[0082] Continue to refer Figure 3 In some embodiments, the recommendation model 114 can be trained in several ways. For example, in one implementation, the server device 110 can use the feature extraction model 112 to determine the sample behavior feature embedding 311 based on the second sample behavior data and the content library 222 and the object library 224. Here, the second sample behavior data may include data associated with the content browsing behavior of the second group of users and data associated with the object operation behavior of the second group of users. Then, based on the sample objects in the second training data set and the object library 224, the sample object feature embedding 321 corresponding to the sample object can be determined, wherein the sample object includes at least one of the first sample object and the second sample object. Then, the recommendation model 114 to be trained can be used to generate a predicted execution result of performing a target operation on the sample object based on the sample behavior feature embedding 311 and the sample object feature embedding 321. Then, the recommendation model 114 can be trained based on the difference between the predicted execution result and the actual execution result of the sample object.

[0083] During the training process of the recommendation model 114, the feature extraction model 112 may be trained. However, it should be understood that during the training process of the recommendation model 114, the feature extraction model 112 may continue to be updated. In this way, the training efficiency of the entire model system may be improved.

[0084] In some embodiments, based on whether each sample object is in a positive sample pair or a negative sample pair, it can be determined whether the actual execution result of each sample object has executed the target operation or has not executed the target operation. Taking the target operation corresponding to each sample object as a click operation or a purchase operation as an example, the corresponding predicted execution result can be, for example, a prediction 328 of a click rate, a prediction 329 of a purchase rate, or other prediction results.

[0085] In some embodiments, the training loss of the recommendation model 114 can be defined based on the difference between the predicted execution result and the actual execution result. The training loss (Loss) can be expressed as any loss function that can determine the difference between the predicted execution result and the actual execution result. As an example, the loss function can be, for example, a cross entropy classification loss function. The formula of the cross entropy classification loss function (which can be expressed as BCELoss) can be, for example:

[0086]

[0087] Where y can represent the actual execution result, It can represent the prediction execution result. N can be equal to 2, indicating 2 categories, and i can be 1 or 2. For example, it can be set that the first category corresponds to the negative sample pair and the second category corresponds to the positive sample pair.

[0088] In some embodiments, to determine the sample object feature embedding 321, the server device 110 may determine a first object type corresponding to the first sample object and / or a second object type corresponding to the second sample object from the object library 224. Then, the sample object feature embedding 321 may be generated based on at least one of the first object type and the second object type.

[0089] refer to Figure 3 In such an embodiment, the sample object A (which may include the first sample object and / or the second sample object) may be mapped to a corresponding object type in the object library 224, and then the embedding of the feature unit of the corresponding object type (such as TokenA 324) may be used to characterize the sample object A. As an example, the sample object feature embedding 321 may be generated based on the embedding of TokenA 324 through forward feedback 326 and layer normalization 327.

[0090] Thus, by training the feature extraction model 112 and the recommendation model 114, the trained feature extraction model 112 and the recommendation model 114 can predict the user's intention based on the user's content browsing behavior and object operation behavior, and recommend the object that the user is interested in. Moreover, the probability of the user operating the recommended object can be increased.

[0091] Figure 4 400 according to some embodiments of the present disclosure. The method 400 may be implemented, for example, in Figure 1 The server device 110 is at the server device 110. Figure 1 The method 400 is described with reference to the environment 100 of FIG.

[0092] In block 410 , the server device 110 obtains user behavior data, where the behavior data includes first behavior data associated with content browsing behavior and second behavior data associated with object operation behavior.

[0093] In block 420 , the server device 110 determines a content type corresponding to the first behavior data and an object type corresponding to the second behavior data.

[0094] In box 430, the server device 110 uses the trained feature extraction model to generate a behavior feature embedding for the user based on the first behavior data, the content type corresponding to the first behavior data, the second behavior data, and the object type corresponding to the second behavior data.

[0095] In block 440 , the server device 110 uses the trained recommendation model to determine an object recommendation result for the user based on the behavior feature embedding and the object feature embedding, where the object feature embedding is associated with an object type of at least one object.

[0096] In some embodiments, determining the content type corresponding to the first behavior data and the object type corresponding to the second behavior data includes: determining the content type corresponding to the first behavior data from a content library, the content library including multiple candidate content types; and determining the object type corresponding to the second behavior data from an object library, the object library including multiple candidate object types.

[0097] In some embodiments, the object feature embedding is determined by: determining an object type corresponding to at least one object from an object library, the object library including multiple candidate object types; and generating an object feature embedding for at least one object based on the object type corresponding to the at least one object.

[0098] In some embodiments, determining an object recommendation result for a user includes: utilizing a trained recommendation model to generate a predicted execution result of performing a target operation on at least one object based on behavioral feature embedding and object feature embedding, the predicted execution result indicating the likelihood of the user performing the target operation on each of the at least one object; and determining an object recommendation result for the user based on the predicted execution result.

[0099] In some embodiments, the user's behavior data includes the user's behavior data within a predetermined time period.

[0100] In some embodiments, the feature extraction model is trained based on a first training data set, and the first training data set is determined in the following manner: obtaining first sample behavior data, the first sample behavior data including data associated with content browsing behavior of a first group of users and data associated with object operation behavior of the first group of users; and determining the first training data set based on the first sample behavior data and a content library and an object library, wherein the content library includes multiple candidate content types and the object library includes multiple candidate object types, wherein the first training data set includes a first sequence for each user in the first group of users and a second sequence corresponding to the first sequence, the first sequence including multiple feature units, each feature unit including a content type determined from the content library or an object type determined from the object library, the second sequence including multiple type units corresponding to the multiple feature units, respectively, each type unit indicating whether the feature unit in the first sequence corresponding to the type unit belongs to a content type or an object type.

[0101] In some embodiments, the feature extraction model is trained in the following manner: using the feature extraction model to be trained, based on the first sequence and the second sequence, respectively generate a first prediction result corresponding to the first sequence and a second prediction result corresponding to the second sequence; determine a first loss corresponding to the first prediction result and a second loss corresponding to the second prediction result; and based on the first loss and the second loss, train the feature extraction model.

[0102] In some embodiments, the recommendation model is trained based on a second training data set, and the second training data set includes: multiple positive sample pairs, each positive sample pair includes a first user and a first sample object corresponding to the first user, the first sample object is an object on which the first user has performed a target operation; and multiple negative sample pairs, each negative sample pair includes a second user and a second sample object corresponding to the second user, the second sample object is an object on which the second user has not performed a target operation.

[0103] In some embodiments, the recommendation model is trained in the following manner: using a feature extraction model, based on second sample behavior data and a content library and an object library, determining a sample behavior feature embedding, wherein the second sample behavior data includes data associated with content browsing behavior of a second group of users and data associated with object operation behavior of a second group of users, wherein the content library includes multiple candidate content types and the object library includes multiple candidate object types; based on sample objects and the object library in a second training data set, determining a sample object feature embedding corresponding to the sample object, wherein the sample object includes at least one of a first sample object and a second sample object; using the recommendation model to be trained, based on the sample behavior feature embedding and the sample object feature embedding, generating a predicted execution result of performing a target operation on the sample object; and training the recommendation model based on the difference between the predicted execution result and the actual execution result of the sample object.

[0104] In some embodiments, determining the sample object feature embedding includes: determining a first object type corresponding to the first sample object and / or a second object type corresponding to the second sample object from an object library; and generating the sample object feature embedding based on at least one of the first object type and the second object type.

[0105] Figure 5 A schematic structural block diagram of an example apparatus 500 for object recommendation according to some embodiments of the present disclosure is shown. The apparatus 500 may be implemented as or included in the server device 110. Each module / component in the apparatus 500 may be implemented by hardware, software, firmware, or any combination thereof.

[0106] As shown in the figure, the device 500 includes a behavior data acquisition module 510, which is configured to acquire the user's behavior data, and the behavior data includes first behavior data associated with content browsing behavior and second behavior data associated with object operation behavior; a type determination module 520, which is configured to determine the content type corresponding to the first behavior data and the object type corresponding to the second behavior data; a feature embedding generation module 530, which is configured to use a trained feature extraction model to generate a behavior feature embedding for the user based on the first behavior data, the content type corresponding to the first behavior data, the second behavior data and the object type corresponding to the second behavior data; and a recommendation result determination module 540, which is configured to use a trained recommendation model to determine the object recommendation result for the user based on the behavior feature embedding and the object feature embedding, and the object feature embedding is associated with the object type of at least one object.

[0107] In some embodiments, the type determination module 520 is further configured to determine the content type corresponding to the first behavior data from a content library, the content library including multiple candidate content types; and determine the object type corresponding to the second behavior data from an object library, the object library including multiple candidate object types.

[0108] In some embodiments, the device 500 also includes an object feature embedding determination module, which is configured to determine an object type corresponding to at least one object from an object library, the object library including multiple candidate object types; and generate an object feature embedding for at least one object based on the object type corresponding to the at least one object.

[0109] In some embodiments, the recommendation result determination module 540 is further configured to utilize the trained recommendation model to generate a predicted execution result of performing a target operation on at least one object based on behavioral feature embedding and object feature embedding, wherein the predicted execution result indicates the possibility of the user performing the target operation on each of the at least one object; and determine an object recommendation result for the user based on the predicted execution result.

[0110] In some embodiments, the user's behavior data includes the user's behavior data within a predetermined time period.

[0111] In some embodiments, the feature extraction model is trained based on a first training data set, and the device 500 also includes a first training data set determination module, which is configured to obtain first sample behavior data, the first sample behavior data including data associated with the content browsing behavior of the first group of users and data associated with the object operation behavior of the first group of users; and determine the first training data set based on the first sample behavior data and the content library and the object library, wherein the content library includes multiple candidate content types and the object library includes multiple candidate object types, wherein the first training data set includes a first sequence for each user in the first group of users and a second sequence corresponding to the first sequence, the first sequence including multiple feature units, each feature unit including a content type determined from the content library or an object type determined from the object library, and the second sequence including multiple type units corresponding to the multiple feature units, respectively, each type unit indicating whether the feature unit in the first sequence corresponding to the type unit belongs to a content type or an object type.

[0112] In some embodiments, the device 500 also includes a feature extraction model training module, which is configured to use the feature extraction model to be trained to generate a first prediction result corresponding to the first sequence and a second prediction result corresponding to the second sequence based on the first sequence and the second sequence, respectively; determine a first loss corresponding to the first prediction result and a second loss corresponding to the second prediction result; and train the feature extraction model based on the first loss and the second loss.

[0113] In some embodiments, the recommendation model is trained based on a second training data set, and the second training data set includes: multiple positive sample pairs, each positive sample pair includes a first user and a first sample object corresponding to the first user, the first sample object is an object on which the first user has performed a target operation; and multiple negative sample pairs, each negative sample pair includes a second user and a second sample object corresponding to the second user, the second sample object is an object on which the second user has not performed a target operation.

[0114] In some embodiments, the device 500 also includes a recommendation model training module, which is configured to use a feature extraction model to determine a sample behavior feature embedding based on second sample behavior data and a content library and an object library, wherein the second sample behavior data includes data associated with the content browsing behavior of the second group of users and data associated with the object operation behavior of the second group of users, wherein the content library includes multiple candidate content types and the object library includes multiple candidate object types; based on the sample objects and the object library in the second training data set, determine a sample object feature embedding corresponding to the sample object, wherein the sample object includes at least one of a first sample object and a second sample object; using the recommendation model to be trained, based on the sample behavior feature embedding and the sample object feature embedding, generate a predicted execution result of performing a target operation on the sample object; and train the recommendation model based on the difference between the predicted execution result and the actual execution result of the sample object.

[0115] In some embodiments, the apparatus 500 is further configured to determine, from the object library, a first object type corresponding to the first sample object and / or a second object type corresponding to the second sample object; and generate a sample object feature embedding based on at least one of the first object type and the second object type.

[0116] Figure 6 1 shows a block diagram of an electronic device 600 in which one or more embodiments of the present disclosure may be implemented. It should be understood that Figure 6 The electronic device 600 shown is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. Figure 6 The electronic device 600 shown can be used to implement Figure 1 The server device 110 or Figure 5 Device 500.

[0117] like Figure 6As shown, the electronic device 600 is in the form of a general-purpose computing device. The components of the electronic device 600 may include, but are not limited to, one or more processors or processing units 610, a memory 620, a storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. The processing unit 610 may be an actual or virtual processor and is capable of performing various processes according to a program stored in the memory 620. In a multi-processor system, multiple processing units execute computer executable instructions in parallel to improve the parallel processing capabilities of the electronic device 600.

[0118] The electronic device 600 typically includes a plurality of computer storage media. Such media may be any available media accessible to the electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 620 may be a volatile memory (e.g., a register, a cache, a random access memory (RAM)), a non-volatile memory (e.g., a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 630 may be a removable or non-removable medium, and may include a machine-readable medium, such as a flash drive, a disk, or any other medium, which may be capable of being used to store information and / or data (e.g., training data for training) and may be accessed within the electronic device 600.

[0119] The electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Figure 6 As shown in , a disk drive for reading or writing from a removable, non-volatile disk (e.g., a "floppy disk") and an optical drive for reading or writing from a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to the bus (not shown) by one or more data media interfaces. The memory 620 may include a computer program product 625 having one or more program modules that are configured to perform various methods or actions of various embodiments of the present disclosure.

[0120] The communication unit 640 implements communication with other electronic devices through a communication medium. Additionally, the functions of the components of the electronic device 600 can be implemented with a single computing cluster or multiple computing machines that can communicate through a communication connection. Therefore, the electronic device 600 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.

[0121] The input device 650 may be one or more input devices, such as a mouse, a keyboard, a tracking ball, etc. The output device 660 may be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 600 may also communicate with one or more external devices (not shown) through the communication unit 640 as needed, such as a storage device, a display device, etc., communicate with one or more devices that allow a user to interact with the electronic device 600, or communicate with any device that allows the electronic device 600 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0122] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.

[0123] Various aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of the methods, devices, equipment, and computer program products implemented according to the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions.

[0124] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0125] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0126] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to multiple implementations of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of a module, program segment or instruction includes one or more executable instructions for realizing the logical function of the specification. In some implementations as replacements, the function marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.

[0127] The above descriptions of various implementations of the present disclosure are exemplary, non-exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The selection of terms used herein is intended to best explain the principles of the implementations, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the various implementations disclosed herein.

Claims

1. A method for object recommendation, comprising: Acquiring user behavior data, the behavior data including first behavior data associated with content browsing behavior and second behavior data associated with object operation behavior; Determine a content type corresponding to the first behavior data and an object type corresponding to the second behavior data; Generate, using the trained feature extraction model, a behavior feature embedding for the user based on the first behavior data, the content type corresponding to the first behavior data, the second behavior data, and the object type corresponding to the second behavior data; as well as An object recommendation result for the user is determined using the trained recommendation model based on the behavior feature embedding and the object feature embedding, wherein the object feature embedding is associated with an object type of at least one object.

2. The method according to claim 1, wherein determining the content type corresponding to the first behavior data and the object type corresponding to the second behavior data comprises: Determining a content type corresponding to the first behavior data from a content library, wherein the content library includes a plurality of candidate content types; as well as An object type corresponding to the second behavior data is determined from an object library, where the object library includes a plurality of candidate object types.

3. The method according to claim 1, wherein the object feature embedding is determined by: Determining an object type corresponding to the at least one object from an object library, the object library comprising a plurality of candidate object types; and Based on the object type corresponding to the at least one object, the object feature embedding for the at least one object is generated.

4. The method according to claim 1, wherein determining the object recommendation result for the user comprises: generating, using the trained recommendation model, a predicted execution result of performing a target operation on the at least one object based on the behavior feature embedding and the object feature embedding, the predicted execution result indicating a likelihood that the user will perform the target operation on each of the at least one object; as well as Based on the prediction execution result, the object recommendation result for the user is determined. The method according to claim 1 , wherein the behavior data of the user comprises behavior data of the user within a predetermined time period.

6. The method according to claim 1, wherein the feature extraction model is trained based on a first training data set, and the first training data set is determined by: Acquiring first sample behavior data, the first sample behavior data comprising data associated with content browsing behavior of a first group of users and data associated with object operation behavior of the first group of users; and determining the first training data set based on the first sample behavior data and a content library and an object library, wherein the content library includes a plurality of candidate content types and the object library includes a plurality of candidate object types, The first training data set includes a first sequence for each user in the first group of users and a second sequence corresponding to the first sequence, The first sequence includes a plurality of feature units, each feature unit includes a content type determined from the content library or an object type determined from the object library, The second sequence includes a plurality of type units corresponding to the plurality of feature units respectively, and each type unit indicates whether the feature unit corresponding to the type unit in the first sequence belongs to a content type or an object type.

7. The method according to claim 6, wherein the feature extraction model is trained by: Using the feature extraction model to be trained, based on the first sequence and the second sequence, respectively generate a first prediction result corresponding to the first sequence and a second prediction result corresponding to the second sequence; Determining a first loss corresponding to the first prediction result and a second loss corresponding to the second prediction result; as well as The feature extraction model is trained based on the first loss and the second loss.

8. The method according to claim 1, wherein the recommendation model is trained based on a second training data set, and the second training data set comprises: A plurality of positive sample pairs, each positive sample pair comprising a first user and a first sample object corresponding to the first user, the first sample object being an object on which a target operation has been performed by the first user; as well as A plurality of negative sample pairs, each negative sample pair comprising a second user and a second sample object corresponding to the second user, the second sample object being an object on which the target operation is not performed by the second user.

9. The method according to claim 8, wherein the recommendation model is trained by: Determining, using the feature extraction model, sample behavior feature embedding based on second sample behavior data and a content library and an object library, wherein the second sample behavior data includes data associated with content browsing behavior of a second group of users and data associated with object operation behavior of the second group of users, wherein the content library includes a plurality of candidate content types and the object library includes a plurality of candidate object types; Determining, based on the sample objects in the second training data set and the object library, a sample object feature embedding corresponding to the sample object, wherein the sample object includes at least one of the first sample object and the second sample object; Using the recommendation model to be trained, based on the sample behavior feature embedding and the sample object feature embedding, generating a predicted execution result of executing the target operation on the sample object; and The recommendation model is trained based on the difference between the predicted execution result and the actual execution result of the sample object.

10. The method of claim 9, wherein determining the sample object feature embedding comprises: Determine, from the object library, a first object type corresponding to the first sample object and / or a second object type corresponding to the second sample object; as well as The sample object feature embedding is generated based on at least one of the first object type and the second object type.

11. A device for object recommendation, comprising: A behavior data acquisition module, configured to acquire behavior data of a user, wherein the behavior data includes first behavior data associated with content browsing behavior and second behavior data associated with object operation behavior; a type determination module, configured to determine a content type corresponding to the first behavior data and an object type corresponding to the second behavior data; a feature embedding generation module configured to generate a behavior feature embedding for the user based on the first behavior data, the content type corresponding to the first behavior data, the second behavior data, and the object type corresponding to the second behavior data using a trained feature extraction model; as well as The recommendation result determination module is configured to use the trained recommendation model to determine the object recommendation result for the user based on the behavior feature embedding and the object feature embedding, wherein the object feature embedding is associated with the object type of at least one object.

12. An electronic device comprising: at least one processing unit; as well as At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the device to perform the method according to any one of claims 1 to 10 when executed by the at least one processing unit.

13. A computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the method according to any one of claims 1 to 10 when executed by a processor.

14. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 10.