Content recommendation method and device, equipment, medium and product
By obtaining and processing the behavior sequence encoding of the target object and the content sequence encoding of candidate recommendation content, and combining feature interaction processing to generate interest scores, it solves the problem that existing recommendation systems are difficult to achieve efficient and accurate recommendations, and achieves personalized and accurate content recommendation effects.
Patent Information
- Application Number
- CN202510123928.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-27
AI Technical Summary
It is difficult for existing recommendation systems to efficiently and accurately recommend content of interest to objects, especially when object needs are diversified and content types are explosive.
By obtaining the behavior sequence encoding of the target object and the content sequence encoding of the candidate recommendation content, linear transformation processing is performed to obtain the query vector and the key vector, and feature interaction processing is performed in combination with the behavior mask and the content mask to generate the interest score of the candidate recommendation content of the target object.
It realizes efficient and accurate content recommendations, which can effectively learn the deep correlation between behavior sequences and content sequences, ensuring the personalization and accuracy of recommendations.
Smart Images

Figure CN120045787A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of recommendation technologies, and more particularly, to a content recommendation method, a content recommendation device, an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] Recommendation systems have been widely used in industries such as e-commerce services, online games, and video recommendations. With the diversification of object needs and the explosive growth of content types, how to timely and effectively recommend interesting content to an object is an important challenge faced by recommendation systems. Summary of the Invention
[0003] Embodiments of this application provide a content recommendation method, a content recommendation device, an electronic device, a computer-readable storage medium, and a computer program product, which can achieve efficient and accurate content recommendation.
[0004] Other features and advantages of this application will become apparent from the following detailed description, or will be partially learned through the practice of this application.
[0005] According to one aspect of the embodiments of this application, a content recommendation method is provided, including: obtaining a behavior sequence encoding corresponding to a behavior sequence of a target object, and a content sequence encoding corresponding to a content sequence of a candidate recommended content; respectively performing linear transformation processing on the behavior sequence encoding and the content sequence encoding to obtain a query vector of the behavior sequence encoding, and a key vector and a value vector of the content sequence encoding; obtaining a behavior mask of the behavior sequence and a content mask of the content sequence, and performing feature interaction processing on the query vector of the behavior sequence encoding, the key vector of the content sequence encoding, and the value vector according to the behavior mask and the content mask to obtain an interest score of the target object for the candidate recommended content; and recommending a target recommended content to the target object according to the interest scores of the target object for each of the candidate recommended contents.
[0006] According to one aspect of the embodiments of the present application, a content recommendation device is provided, including: an acquisition module, configured to acquire a behavior sequence encoding corresponding to a behavior sequence of a target object, and a content sequence encoding corresponding to a content sequence of candidate recommended content; a linear transformation module, configured to perform linear transformation processing on the behavior sequence encoding and the content sequence encoding respectively to obtain a query vector of the behavior sequence encoding, and a key vector and a value vector of the content sequence encoding; a feature interaction module, configured to acquire a behavior mask of the behavior sequence and a content mask of the content sequence, and perform feature interaction processing on the query vector of the behavior sequence encoding, the key vector of the content sequence encoding, and the value vector according to the behavior mask and the content mask to obtain an interest score of the target object for the candidate recommended content; a recommendation module, configured to recommend target recommended content to the target object according to the interest scores of the target object for each of the candidate recommended content.
[0007] In an embodiment of the present application, the feature interaction module is further configured to perform union processing on the behavior mask and the content mask to obtain a target mask; perform feature interaction processing on the query vector of the behavior sequence encoding, the key vector of the content sequence encoding, and the value vector according to the target mask to obtain an interaction sequence encoding; and perform linear transformation processing on the interaction sequence encoding to obtain the interest score.
[0008] In an embodiment of the present application, the feature interaction module is further configured to calculate a similarity between each query vector and each key vector to obtain an attention score; obtain a target attention score according to the target mask and the attention score, and perform weighted summation on each value vector according to the target attention score to obtain an attention output encoding; and generate the interaction sequence encoding according to the attention output encoding and the behavior sequence encoding.
[0009] In an embodiment of the present application, the feature interaction module is further configured to perform feature fusion processing on the attention output encoding and the behavior sequence encoding to obtain a fusion encoding, and perform normalization processing on the fusion encoding to obtain a target encoding; and process the target encoding through a feed-forward neural network to generate the interaction sequence encoding.
[0010] In an embodiment of the present application, the recommendation module is further configured to obtain the comprehensive attribute information of the target object and the content information of the candidate recommended content, where the comprehensive attribute information includes basic attribute information; generate a recommendation score of the target object for the candidate recommended content according to the comprehensive attribute information and the content information; fuse the interest score and the recommendation score according to the basic attribute information of the target object to obtain a target recommendation score of the candidate recommended content; and recommend target recommended content to the target object according to the target recommendation scores of each candidate recommended content.
[0011] In an embodiment of the present application, the recommendation module is further configured to perform feature extraction on the basic attribute information of the target object to obtain object basic attribute features; map the object basic attribute features to an interest weight corresponding to the interest score and a recommendation weight corresponding to the recommendation score through a neural network layer based on gradient descent; and perform weighted summation processing on the interest score and the recommendation score according to the interest weight and the recommendation weight to obtain the target recommendation score.
[0012] In an embodiment of the present application, the recommendation module is further configured to map the object basic attribute features to the interest weight and the recommendation weight through a multi-layer perceptron formed by stacking at least two linear layers; or, obtain basic content in the content information of the candidate recommended content, and perform feature extraction processing on the basic content to obtain basic content features; and map the object basic attribute features and the basic content features to the interest weight and the recommendation weight through a factorization machine.
[0013] In an embodiment of the present application, the recommendation module is further configured to obtain a recommendation type of a recommendation task for the target object, and determine a weight bias factor according to the recommendation type, where the weight bias factor is used to indicate the influence of the recommendation type on content recommendation; adjust the interest weight and the recommendation weight according to the weight bias factor to obtain a target recommendation weight and a target recommendation weight; and perform weighted summation processing on the interest score and the recommendation score according to the target recommendation weight and the target recommendation weight to obtain the target recommendation score.
[0014] In an embodiment of the present application, the recommendation module is further configured to perform feature extraction on the comprehensive attribute information and the content information respectively to obtain object features and content features; obtain cross information between the target object and the candidate recommended content, and perform feature extraction on the cross information to obtain cross features; and generate the recommendation score according to the object features, the content features, and the cross features.
[0015] In one embodiment of the present application, the recommendation module is further configured to obtain the current environmental information, extract features from the environmental information to obtain environmental features; perform feature interaction processing on the object features and content features to obtain interaction features; perform feature fusion on the interaction features, cross features and environmental features to obtain full-scale features, and generate the recommendation score according to the full-scale features.
[0016] In one embodiment of the present application, the acquisition module is further configured to perform encoding processing on the behavior sequence through a multi-head attention mechanism to obtain the behavior sequence encoding; if the content attribute information of the candidate recommended content is a content identifier, construct the content sequence encoding according to the feature dimension and preset sequence length of the content identifier; if the content attribute information of the candidate recommended content includes a content title, generate the content sequence according to the content title, and perform processing on the content sequence through the multi-head attention mechanism to obtain the content sequence encoding.
[0017] According to one aspect of the embodiments of the present application, embodiments of the present application provide an electronic device, including one or more processors; a storage device for storing one or more computer programs, and when the one or more computer programs are executed by the one or more processors, the electronic device implements the content recommendation method as described above.
[0018] According to one aspect of the embodiments of the present application, embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor of an electronic device, the electronic device executes the content recommendation method as described above.
[0019] According to one aspect of the embodiments of the present application, embodiments of the present application provide a computer program product, including a computer program, the computer program is stored in a computer-readable storage medium, and a processor of an electronic device reads and executes the computer program from the computer-readable storage medium, so that the electronic device executes the content recommendation method as above.
[0020] In the technical solution provided by the embodiments of the present application, a behavior sequence encoding corresponding to the behavior sequence of the target object and a content sequence encoding corresponding to the content sequence of the candidate recommended content are obtained; linear transformation processing is respectively performed on the behavior sequence encoding and the content sequence encoding to obtain a query vector of the behavior sequence encoding, and a key vector and a value vector of the content sequence encoding. Through the linear transformation processing, the feature spaces of the behavior sequence and the content sequence can be aligned; a behavior mask of the behavior sequence and a content mask of the content sequence are obtained, and according to the behavior mask and the content mask, feature interaction processing is performed on the query vector of the behavior sequence encoding, the key vector and the value vector of the content sequence encoding to obtain the interest score of the target object for the candidate recommended content, that is, through the deep interaction between the query vector of the behavior sequence and the key vector and the value vector of the content sequence, the deep association between the behavior sequence and the content sequence is effectively learned, and the deep interaction process of the corresponding vectors of the sequence is controlled by the behavior mask and the content mask to ensure that only effective behaviors and content features are concerned, so that the deep interaction is more efficient and accurate, so as to fully integrate the behavior sequence and the content sequence to accurately obtain the interest score of the target object for the candidate recommended content, and then recommend the target recommended content to the target object according to the interest scores of the target object for each candidate recommended content, so as to provide a better recommendation service for the object, realizing efficient and accurate content recommendation.
[0021] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the accompanying drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings according to these drawings without creative efforts. In the drawings.
[0023] Figure 1 is a schematic diagram of an implementation environment related to the present application.
[0024] Figure 2 is a flowchart of a content recommendation method shown in an exemplary embodiment of the present application.
[0025] Figure 3 is a schematic flowchart of another content recommendation method shown in an exemplary embodiment of the present application.
[0026] Figure 4 is a flowchart of another content recommendation method shown in an exemplary embodiment of the present application.
[0027] Figure 5It is a flowchart of another content recommendation method shown in an exemplary embodiment of the present application.
[0028] Figure 6 It is a flowchart of another content recommendation method shown in an exemplary embodiment of the present application.
[0029] Figure 7 It is a schematic structural diagram of a recommendation model shown in an exemplary embodiment of the present application.
[0030] Figure 8 It is a schematic structural diagram of an object sequence encoding module shown in an exemplary embodiment of the present application.
[0031] Figure 9 It is a schematic structural diagram of a sequence cross module shown in an exemplary embodiment of the present application.
[0032] Figure 10 It is a schematic structural diagram of a personalized dynamic weight adjustment module shown in an exemplary embodiment of the present application.
[0033] Figure 11 It is a structural block diagram of a content recommendation device shown in an exemplary embodiment of the present application.
[0034] Figure 12 It shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. Detailed implementation manners
[0035] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are only examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0036] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0037] The flowcharts shown in the drawings are only exemplary descriptions and do not necessarily include all the contents and operations, nor do they necessarily execute in the described order. For example, some operations can be decomposed, and some operations can be combined or partially combined. Therefore, the actual execution order may change according to the actual situation.
[0038] It should also be noted that: "a plurality of" mentioned in this application means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0039] The technical solutions of the embodiments of the present application will be introduced in detail below.
[0040] Please refer to Figure 1 , Figure 1 which is a schematic diagram of an implementation environment involved in this application. This implementation environment includes a terminal 10 and a server 20, and a recommendation system runs in the server 20.
[0041] The terminal 10 is used to obtain the behavior sequence of the target object and send the obtained information to the server 20.
[0042] The server 20 is used to determine the candidate recommended content for the target object, obtain the behavior sequence encoding corresponding to the behavior sequence of the target object, and the content sequence encoding corresponding to the content sequence of the candidate recommended content; perform linear transformation processing on the behavior sequence encoding and the content sequence encoding respectively to obtain the query vector of the behavior sequence encoding, and the key vector and value vector of the content sequence encoding; then obtain the behavior mask of the behavior sequence and the content mask of the content sequence, and perform feature interaction processing on the query vector of the behavior sequence encoding, the key vector of the content sequence encoding, and the value vector according to the behavior mask and the content mask to obtain the interest score of the target object for the candidate recommended content, and then recommend the target recommended content to the target object according to the interest scores of the target object for each candidate recommended content.
[0043] The server can also send the target recommended content to the terminal so that the terminal can display the target recommended content to the target object.
[0044] In some embodiments, the server 20 can also obtain the behavior sequence of the target object, and then through the linear transformation processing of the sequence encoding and the feature interaction processing of the corresponding vectors, obtain the interest score of the target object for the candidate recommended content, and then perform content recommendation according to the interest score.
[0045] In some embodiments, the terminal 10 can also independently implement the process of content recommendation, that is, the terminal 10 obtains the relevant information of the target object and the relevant information of the candidate recommended content, generates the interest score of the target object for the candidate recommended content, and then performs content recommendation according to the interest score.
[0046] Among them, the aforementioned terminal 10 may be an electronic device such as a smart phone, a tablet, a notebook computer, a computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, etc. The server 20 may be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. This is not restricted here.
[0047] The terminal 10 and the server 20 are pre-connected through a network to establish a communication connection, so that the terminal 10 and the server 20 can communicate with each other through the network. The network can be a wired network or a wireless network, which is not restricted here either.
[0048] It should be noted that in the specific implementation manner of this application, at least one of the behavior sequence of the target object and the candidate recommended video is related to the object. When the embodiments of this application are applied to specific products or technologies, the permission or consent of the object needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0049] The following elaborates in detail on various implementation details of the technical solutions of the embodiments of this application.
[0050] As Figure 2 shown, Figure 2 is a flowchart of a content recommendation method shown in an embodiment of this application. This method can be applied to the Figure 1 shown implementation environment. This method can be executed by the terminal or the server, or jointly executed by the terminal and the server. In the embodiments of this application, taking the execution of this method by the server as an example, this content recommendation method may include S210 to S240, which are introduced in detail as follows.
[0051] S210. Obtain the behavior sequence encoding corresponding to the behavior sequence of the target object, and the content sequence encoding corresponding to the content sequence of the candidate recommended content.
[0052] In the embodiments of the present application, the behavior sequence of the target object refers to a sequence composed of at least two historical behaviors triggered by the target object. For example, based on the time sequence, if the target object triggers click, browse, and purchase behaviors, then according to the behavior sequence composed of click, browse, and purchase. Among them, the behavior sequence can include multiple characterization dimensions, such as the number of behaviors, the occurrence time of each behavior, the time from now, the degree of continuity, etc. The historical behavior can be the behavior of the target object within a period of time; the candidate recommended content can be a video, text, or image, which is not limited here. The content sequence of the candidate recommended content refers to a sequence composed of the content information of the candidate recommended content, such as a sequence composed of content titles.
[0053] In one example, the behavior sequence of the target object is encoded to extract the behavioral feature representation of the object, obtaining the behavior sequence encoding; the content sequence of the candidate recommended content is encoded to extract the feature representation of the recommended content, obtaining the content sequence encoding; among them, the encoding process of the behavior sequence and the encoding process of the content sequence can be the same or different.
[0054] S220. Linearly transform the behavior sequence encoding and the content sequence encoding respectively to obtain the query vector of the behavior sequence encoding, and the key vector and value vector of the content sequence encoding.
[0055] In the embodiments of the present application, the query vector Q of the behavior sequence encoding is obtained by linearly transforming the behavior sequence encoding through the learnable weight matrix W_Q, and the key vector K and value vector V of the content sequence encoding are obtained by linearly transforming the content sequence encoding through the learnable weight matrices W_K and W_V. Then, the query vector Q, the key vector K, and the value vector V are subjected to feature interaction processing. That is to say, since the query vector Q, the key vector K, and the value vector V come from different sequences respectively, important features in both the behavior sequence and the content sequence can be simultaneously concerned during feature interaction.
[0056] S230. Obtain the behavior mask of the behavior sequence and the content mask of the content sequence, and based on the behavior mask and the content mask, perform feature interaction processing on the query vector of the behavior sequence encoding, the key vector of the content sequence encoding, and the value vector to obtain the interest score of the target object for the candidate recommended content.
[0057] In the embodiments of the present application, this interest score is a score generated based on the historical behavior sequence of the target object, and by analyzing the behavior of the object in the past period of time and the content of the recommended content, the interest degree of the object in the candidate recommended content is predicted.
[0058] It should be noted that the sequence lengths of the query vectors corresponding to the behavior sequence encoding, and the key vectors and value vectors corresponding to the content sequence encoding may be of variable length. In order to learn effective weights for the interaction between the two sequences, in the embodiments of the present application, the interaction range of the variable-length sequences is controlled by the masks of the sequences to ensure that only the valid sequence parts are concerned during the feature interaction process, so as to provide accuracy and robustness.
[0059] In the embodiments of the present application, the mask is used to mask the invalid parts in the sequence, such as the filled blank positions. Among them, the mask is a matrix, and the values inside are all 0 or 1. For example, [1, 1, 1, 0, 0] represents a sequence with a length of 5, where the first three slots have content and the last two elements are empty.
[0060] For example, the behavior sequence: [Click on product A, Browse product B, Purchase product C, Fill, Fill]; the content sequence: [Fill, Product X, Fill, Product Z, Fill]. The actual lengths of the behavior sequence and the content sequence are 3 and 2 respectively, and the filled parts are for making the sequence lengths consistent; then the behavior mask of the behavior sequence is [1, 1, 1, 0, 0], where 1 represents the valid part and 0 represents the filled part, and the content mask of the content sequence is [0, 1, 0, 1, 0].
[0061] In one example, the query vector of the behavior sequence, the key vector and value vector encoded by the content sequence are interacted through the cross-attention mechanism, and according to the behavior mask and the content mask, the calculation range of the attention mechanism is restricted, which can effectively adapt to the variable-length behavior sequence and content sequence, and is more efficient and accurate when processing long sequences or sparse sequences, improving the adaptability to complex scenarios.
[0062] In another example, dot product interaction can be performed on the query vector of the behavior sequence, the key vector and value vector encoded by the content sequence for interaction to capture local cross features. Outer product can also be used to perform deeper cross of the behavior sequence encoding and the content sequence encoding to capture high-order features, and then the range of dot product interaction or outer product interaction is restricted according to the behavior mask and the content mask.
[0063] S240. Recommend the target recommended content to the target object according to the interest scores of the target object for each candidate recommended content.
[0064] It can be understood that, as Figure 3As shown, for each candidate recommended content, an interest score can be obtained through steps S210 to S230, and then content recommendation is made to the target object according to the interest scores of each candidate recommended content. For example, directly sort each candidate recommendation according to the size of the interest score, and use the top n candidate recommended contents with higher rankings as the target recommended contents, or use the candidate recommended contents with interest scores greater than the preset threshold as the target recommended contents, and recommend the target recommended contents to the target object.
[0065] In an embodiment of the present application, the behavior sequence of the target object and the content sequence of the candidate recommended content can be input into the content recommendation model, and the interest score of the candidate recommended content can be obtained through the content recommendation model.
[0066] In an embodiment of the present application, linear transformation processing is respectively performed on the behavior sequence encoding and the content sequence encoding to obtain the query vector of the behavior sequence encoding, as well as the key vector and value vector of the content sequence encoding. Through linear transformation processing, the feature spaces of the behavior sequence and the content sequence can be aligned; obtain the behavior mask of the behavior sequence and the content mask of the content sequence, and according to the behavior mask and the content mask, perform feature interaction processing on the query vector of the behavior sequence encoding, the key vector of the content sequence encoding, and the value vector to obtain the interest score of the target object for the candidate recommended content, that is, through the in-depth interaction between the query vector of the behavior sequence and the key vector and value vector of the content sequence, to effectively learn the in-depth association between the behavior sequence and the content sequence, and control the in-depth interaction process of the corresponding vectors of the sequence through the behavior mask and the content mask to ensure that only effective behaviors and content features are concerned, making the in-depth interaction more efficient and accurate, so as to fully integrate the behavior sequence and the content sequence to accurately obtain the interest score of the target object for the candidate recommended content, and then recommend the target recommended content to the target object according to the interest scores of the target object for each candidate recommended content, so as to provide better recommendation services for the object and achieve efficient, accurate and flexible content recommendation.
[0067] In an embodiment of the present application, another content recommendation method is provided, and this content recommendation method can be applied to Figure 1 the implementation environment shown. This method can be executed by the terminal or the server, or jointly executed by the terminal and the server. In an embodiment of the present application, taking the example that this method is executed by the server for illustration, as Figure 3 shown, this content recommendation method is based on S210 to S240 shown in Figure 2 and expands S230 to S410 to S430; S410 to S430 are introduced in detail as follows.
[0068] S410. Obtain the behavior mask of the behavior sequence and the content mask of the content sequence, and perform union processing on the behavior mask and the content mask to obtain the target mask.
[0069] S420. Perform feature interaction processing on the query vector encoded by the behavior sequence, the key vector encoded by the content sequence, and the value vector according to the target mask to obtain an interaction sequence encoding.
[0070] S430. Perform linear transformation processing on the interaction sequence encoding to obtain an interest score.
[0071] For example, if the behavior mask is [1, 1, 1, 0, 0] and the content mask is [0, 1, 0, 1, 0], perform a union operation. Among them, the union logic is that for each element, as long as there is a value of 1 in any one of the masks, the value of the corresponding element in the result mask is 1. For example, the obtained target mask is [1, 1, 1, 1, 0]. This target mask can effectively cover the valid elements in the behavior sequence and the content sequence. Then, perform attention interaction on the query vector encoded by the sequence, the key vector encoded by the content sequence, and the value vector according to the target mask to ensure that the attention interaction only focuses on the valid sequence part and obtain a more accurate interaction sequence encoding.
[0072] In an example, obtaining the interaction sequence encoding according to the target mask includes: for each query vector, calculate the similarity between the query vector and each key vector to obtain an attention score; obtain a target attention score according to the target mask and the attention score, and perform weighted summation on each value vector according to the target attention score to obtain an attention output encoding; generate an interaction sequence encoding according to the attention output encoding and the behavior sequence encoding.
[0073] Exemplarily,
[0074]
[0075] M is the target mask, and dk is the dimension of the key vector.
[0076] For each query vector (element in the behavior sequence), calculate its dot product with all key vectors (elements in the content sequence) to obtain an attention score matrix. Assume the attention score matrix is: [[0.8, 0.6, 0.4, 0.2, 0.1], [0.7, 0.5, 0.3, 0.1, 0.0], [0.9, 0.7, 0.5, 0.3, 0.2], [0.0, 0.0, 0.0, 0.0, 0.0], [0.0, 0.0, 0.0, 0.0, 0.0]]. Apply the target mask to the attention score matrix to mask out the attention scores of the padded parts to obtain the target attention scores. The target attention scores are [[0.8, 0.6, 0.4, 0.2, -∞], [0.7, 0.5, 0.3, 0.1, -∞], [0.9, 0.7, 0.5, 0.3, -∞], [0.0, 0.0, 0.0, 0.0, -∞], [0.0, 0.0, 0.0, 0.0, -∞]]. Normalize the target attention scores through the Softmax function to obtain a weight matrix. The weight matrix is [[0.4, 0.3, 0.2, 0.1, 0.0], [0.4, 0.3, 0.2, 0.1, 0.0], [0.4, 0.3, 0.2, 0.1, 0.0], [0.4, 0.3, 0.2, 0.1, 0.0], [0.4, 0.3, 0.2, 0.1, 0.0]]. Then, perform a weighted sum of all value vectors (elements in the content sequence) according to the weight matrix to obtain an attention output encoding, which is the encoding output by the attention layer.
[0077] In an example of the embodiment of the present application, the attention output encoding and the behavior sequence encoding can be feature concatenated to retain the information in the original behavior sequence, and then an interaction sequence encoding is generated. The interaction sequence encoding contains more information dimensions, that is, it can utilize both the feature interaction information captured by the attention mechanism and retain the original object behavior characteristics, thereby improving the accuracy of recommendation.
[0078] In another example, generating the interaction sequence encoding includes: performing a feature fusion process on the attention output encoding and the behavior sequence encoding to obtain a fusion encoding, and performing a normalization process on the fusion encoding to obtain a target encoding; processing the target encoding through a feed-forward neural network to generate the interaction sequence encoding.
[0079] Among them, the fusion encoding can be directly obtained by concatenating the attention output encoding and the behavior sequence encoding, or the fusion encoding can be obtained by performing a weighted sum of the attention output encoding and the behavior sequence encoding, where the weight corresponding to the attention output encoding is greater than the weight of the behavior sequence encoding; then, the target encoding is obtained by performing a normalization process on the fusion encoding, which can ensure that the numerical values of different feature dimensions are in the same range, thereby improving stability.
[0080] In the embodiments of the present application, the target encoding is processed by a feed-forward neural network, which can capture the complex relationships and high-order feature interactions between the target encodings to generate an interaction sequence encoding.
[0081] Optionally, after processing the target encoding by the feed-forward neural network, the target encoding and the encoding output by the feed-forward neural network can also be subjected to residual connection and normalization processing to obtain the interaction sequence encoding.
[0082] It should be noted that Figure 4 For other detailed introductions of S210, S230 to S240 shown in Figure 2 please refer to S210, S230 to S240 shown in
[0083] In the embodiments of the present application, the query vector Q, key vector K, and value vector V input to the attention mechanism respectively belong to two different sequences, enabling the attention mechanism to capture cross-sequence relationships, and being able to simultaneously focus on important features in the behavior sequence and the content sequence. A mask is used to control the attention mechanism process. The mask can effectively mask the attention scores of the padded parts, ensuring that the attention mechanism only focuses on the valid sequence parts, obtaining a more accurate interaction sequence encoding, and thus improving the accuracy and reliability of the generated first interest score.
[0084] It should be noted that in other embodiments of the present application, the behavior sequence encoding and the content sequence encoding are of variable length, and directly using a mask will introduce redundant calculations. In the embodiments of the present application, the length of the sequence can be adjusted by a dynamic truncation strategy to reduce the amount of calculation, while avoiding processing redundant or irrelevant data, further optimizing the calculation efficiency. Among them, dynamic truncation cuts the sequence according to the length of the actual valid data in the sequence, so that only the valid part is processed subsequently, saving computing resources.
[0085] In another example, after obtaining the behavior mask of the behavior sequence of the object and the content mask of the content sequence of the candidate recommended content, the behavior sequence encoding is truncated according to the behavior mask to obtain the target behavior sequence encoding, the content sequence encoding is truncated according to the content mask to obtain the target content sequence encoding, and the corresponding target behavior mask and target content mask are regenerated according to the target behavior sequence encoding and the target content sequence encoding. Furthermore, feature interaction processing is performed on the query vector of the target behavior sequence encoding, the key vector and value vector of the target content sequence encoding according to the target behavior mask and the target content mask to obtain the interaction sequence encoding.
[0086] Among them, as described above, 1 in the mask represents the valid part, and 0 represents the padding part. The effective length of the behavior sequence encoding can be calculated based on the behavior mask. According to the calculated effective length, the sequence is dynamically cropped. For example, for the behavior sequence encoding, the first effective length of valid data is retained, and the rest is truncated to obtain the target behavior sequence encoding. Similarly, the content sequence encoding is truncated to obtain the target content sequence encoding, and new target behavior masks and target content masks are generated. The feature interaction feature is executed on the target behavior sequence encoding and the target content sequence encoding, and the scope of attention calculation is controlled by the target behavior mask and the target content mask. For specific details, please refer to the foregoing embodiments and will not be elaborated herein.
[0087] An embodiment of the present application provides another content recommendation method, which can be applied to Figure 1 the implementation environment shown in the figure. This method can be executed by a terminal or a server, or jointly executed by a terminal and a server. In the embodiments of the present application, taking the execution of this method by the server as an example for illustration, as Figure 5 shown in the figure, this content recommendation method is based on Figure 2 the one shown in the figure, and expands S240 shown in Figure 2 the figure to S510 - S540. S510 - S540 are introduced in detail as follows.
[0088] S510. Obtain the comprehensive attribute information of the target object and the content information of the candidate recommended content. The comprehensive attribute information includes basic attribute information.
[0089] The comprehensive attribute information of the target object refers to the fused multi - dimensional attribute information, which includes basic attribute information and also includes extended attribute information, forming a comprehensive object data structure. Among them, the basic attribute information is such as age, gender, etc., and the extended attribute information includes object interest tags, etc.; the content information of the candidate recommended content refers to the fused multi - dimensional content information, which includes basic content and also includes extended content. The basic content includes the title; the extended content includes content tags and content context information.
[0090] S520. Generate a recommendation score of the target object for the candidate recommended content according to the comprehensive attribute information and the content information.
[0091] In the embodiments of the present application, the recommendation score is a score generated based on multi - dimensional comprehensive information, which does not depend on the historical behavior sequence of the object, but uses the comprehensive full - volume information to predict the recommendation degree of the candidate recommended content for the object.
[0092] In one example, the feature extraction process is respectively performed on the comprehensive attribute information and the content information of the target object, and the extracted features are fused, and then the recommendation score is generated according to the fused features.
[0093] In another example, generating a recommendation score of a target object for candidate recommended content includes: extracting object features and content features by performing feature extraction on comprehensive attribute information and content information respectively; obtaining cross-information between the target object and the candidate recommended content, and performing feature extraction on the cross-information to obtain cross features; generating a recommendation score according to the object features, content features, and cross features.
[0094] Exemplarily, perform embedding processing on the features of the comprehensive attribute information and content information, and convert heterogeneous features into low-dimensional dense vectors to obtain the object features corresponding to the target object and the content features of the recommended content.
[0095] In the embodiments of the present application, the cross-information is the correlation information between the object and the recommended content, which is generated through the historical interaction or context association between the target object and the candidate recommended content. The historical interaction includes whether the object has ever clicked / purchased / collected the recommended content, and the context association includes the time correlation between the object and the recommended content (such as the object prefers evening content and the release time of the recommended content is evening). The context association also includes the matching degree between the interest tags of the object and the tags of the recommended content; perform embedding processing on the features of the correlation information to obtain cross features.
[0096] In one example, the object features, content features, and cross features can be fused to obtain global features, and then a linear transformation is performed on the global features to obtain a recommendation score; wherein, fusing the object features, content features, and cross features can be splicing or weighted summation.
[0097] In another example, generating a recommendation score according to the object features, content features, and cross features includes: obtaining the current environmental information, and performing feature extraction on the environmental information to obtain environmental features; performing feature interaction processing on the object features and content features to obtain interaction features; performing feature fusion on the interaction features, cross features, and environmental features to obtain full-scale features, and generating a recommendation score according to the full-scale features.
[0098] The interest preferences of the object will change dynamically with scenarios, events, etc. Therefore, the environment is used to assist in determining the impact on the recommended content; wherein, the current environmental information includes but is not limited to weather, holidays, event promotion information, or hot events, and embedding processing is performed on the environmental information to obtain environmental features.
[0099] The direct correlation between the object and the recommended content can be obtained through the cross-information between the target object and the candidate recommended content, and the implicit correlation between the object and the recommended content can be learned by performing feature interaction processing on the object features and content features. Among them, performing feature interaction processing on the object features and content features can be simple weighted summation, or the interaction relationship between features can be captured through FM.
[0100] Feature splicing is performed on interaction features, cross features, and environmental features to obtain full features. These full features can determine the interest of the object in the recommended content from multiple dimensions, construct more comprehensive portrait features, and then perform a linear transformation on the full features to obtain a recommendation score.
[0101] In the embodiments of the present application, based on the comprehensive attribute information of the object and the content information of the recommended content, combined with the cross information and environmental information between the object and the recommended content, a more comprehensive portrait is constructed to accurately capture the needs of the object and improve the accuracy and reliability of the recommendation score.
[0102] S530. According to the basic attribute information of the target object, the interest score and the recommendation score are fused to obtain the target recommendation score of the candidate recommended content.
[0103] It should be noted that the interest score can capture the characteristics of the behavior time series, but cannot utilize the full features. The recommendation score can utilize the full features, but cannot utilize the sequence characteristics. Therefore, the two scores are personalized and fused to improve the prediction accuracy of the recommendation.
[0104] It can be understood that for different objects, the importance of the interest score obtained through sequence prediction is different. For example, for a new object without historical sequence information, the value of the interest score is very low. Therefore, in the embodiments of the present application, the basic attribute information of the target object can reflect the individual characteristics of the object, and then the interest score and the recommendation score are personalized and dynamically fused to obtain the target recommendation score of the candidate recommended content.
[0105] For example, interest weights corresponding to the interest score and recommendation weights corresponding to the recommendation score are generated according to the basic attribute information; the basic attribute information includes the object ID. By knowing the historical activity level of the target object through the target object ID, if the historical activity level is greater than the preset degree threshold, the interest weight is greater than the recommendation weight; otherwise, the interest weight is less than the recommendation weight. Another example is that the basic attribute information includes the object age. If the object age is greater than the preset age threshold, the object is an elderly person. Since the elderly have less historical sequence information due to operation reasons, the interest weight is less than the recommendation weight.
[0106] In one example, fusing the interest score and the recommendation score to obtain the target recommendation score of the candidate recommended content includes: extracting object basic attribute features from the basic attribute information of the target object; and mapping the object basic attribute features to the interest weight corresponding to the interest score and the recommendation weight corresponding to the recommendation score through a neural network layer based on gradient descent.
[0107] Among them, the basic attribute information includes different types of information, such as discrete information corresponding to the object ID, numerical information corresponding to the object age, textual information corresponding to the object gender, etc. The Embedding layer can perform feature extraction processing on different types of information in the basic attribute information respectively to obtain the information features of each object. The Embedding layer maps different types of information to the same vector space, making subsequent processing more unified. Then, the information features of each object are fused to obtain the basic attribute features of the object. Among them, the basic attribute features of the object can be obtained by concatenating the information features of each object, or by weighted summation of the information features of each object. Among them, the weights corresponding to the information features of each object can be flexibly adjusted according to the actual situation, such as the weight corresponding to the object ID is the largest, and the weight corresponding to the age is the second largest, etc.
[0108] In the embodiment of the present application, the neural network layer based on gradient descent can perform feature fusion processing on the information features of each object to obtain the basic attribute features of the object, and also map the basic attribute features of the object to the interest weight and the recommendation weight.
[0109] Among them, the neural network layer based on gradient descent includes MLP (multi-layer perceptron), Factorization Machine (FM), SeNet (Squeeze-and-Excitation Networks), LHUC (Learning Hidden Unit Contributions), etc.
[0110] In an example, different types of neural networks can be selected according to the actual situation. Specifically, in the case where there is only basic attribute information, the multi-layer perceptron formed by stacking at least two linear layers maps the basic attribute features of the object to the interest weight and the recommendation weight; among them, the multi-layer perceptron is composed of at least two linear layers and an activation function, and the output of each linear layer is used as the input of the next linear layer. The activation function of the linear layer can be arbitrarily selected, such as ReLu (RectifiedLinear Unit), SeLu (Scaled Exponential Linear Unit), etc.
[0111] In other embodiments of the present application, in addition to the basic attribute information of the object, scene features of the recommendation can also be introduced. The scene features describe the environment and context in which the object is currently located, including time and location. Through the scene features, the preferences of the object in different scenarios can be better captured, and then the scoring weights can be dynamically adjusted by combining the basic attribute information of the object and the recommendation scenario. For example, the scene features and the basic attribute features of the object are fused to obtain the target object features, and then the target object features are mapped to interest weights and recommendation weights based on the neural network layer of gradient descent.
[0112] In the embodiments of the present application, in addition to the basic attribute information on the object side, the basic content in the content information of the candidate recommended content can also be introduced, and the interaction between the basic attribute information of the object and the basic content (such as the age of the object and the category of the commodity) has a greater impact on the recommendation effect. At this time, the basic content in the content information of the candidate recommended content is obtained, and feature extraction processing is performed on the basic content to obtain basic content features. Based on the factorization machine, the basic attribute features of the object and the basic content features are mapped to interest weights and recommendation weights.
[0113] Among them, for the acquisition of the basic content features, reference can be made to the acquisition process of the basic attribute features of the object, which will not be elaborated here. The FM is used to capture the interaction relationship between features. Therefore, the basic attribute features of the object and the basic content features are input into the FM, and the FM outputs second-order interaction features, and then the second-order interaction features are mapped to interest weights and recommendation weights.
[0114] In one example, the second-order interaction features can be input into the MLP to capture high-order interactions through the MLP, and then the captured high-order interaction features are mapped to interest weights and recommendation weights based on the MLP.
[0115] In one example, the outputs of the FM and the MLP are fused, that is, the basic attribute features of the object are converted into output features based on the MLP, and the second-order interaction features output by the FM are simply weighted and summed with the output features, and then the fused features are input into the fully connected layer to calculate the final interest weights and recommendation weights.
[0116] After obtaining the corresponding weights, weighted summation processing is performed on the interest scores and recommendation scores according to the interest weights and recommendation weights to obtain the target recommendation score.
[0117] In one example of the embodiments of the present application, the interest scores and recommended interest scores can be directly weighted and summed according to the interest weights and recommendation weights to obtain the target recommendation score.
[0118] In another example, based on the interest weight and the recommendation weight, the weight can be further adjusted based on the recommendation type of the recommendation task, and then the weighted sum is performed based on the adjusted weight to obtain the target recommendation score. Specifically, it includes: obtaining the recommendation type of the recommendation task for the target object, and determining the weight bias factor according to the recommendation type; adjusting the interest weight and the recommendation weight according to the weight bias factor to obtain the target interest weight and the target recommendation weight; performing a weighted sum process on the interest score and the recommended interest score according to the target interest weight and the target recommendation weight to obtain the target recommendation score.
[0119] Among them, the recommendation types of the recommendation tasks for the target object include popularity recommendation and long-tail recommendation. Popularity recommendation refers to preferentially recommending the currently most popular content, and the influence of the object behavior sequence feature on popularity recommendation is relatively small because hot content usually attracts most objects; long-tail recommendation refers to mining and recommending long-tail content (i.e., non-popular content) that the object may be interested in, and the interaction records of long-tail products are less, and the recommendation depends more on the historical behavior sequence feature of the object; the contribution of the full-scale feature to the recommendation of long-tail products is small. Therefore, the weight bias factor can be determined when the recommendation type is known, and this weight bias factor is used to indicate the influence of the recommendation type on content recommendation; for different recommendation types, their corresponding weight bias factors are different. For example, for popularity recommendation, the weight bias factor corresponding to the recommendation score is greater than the weight bias factor corresponding to the interest score; for long-tail recommendation, the weight bias factor corresponding to the recommendation score is less than the weight bias factor corresponding to the interest score.
[0120] After determining the weight bias factor, adjust the interest weight and the recommendation weight determined by the basic attribute information to obtain the target interest weight and the target recommendation weight; for example, the interest weight is 0.4 and the recommendation weight is 0.6. Assuming the recommendation type is popularity recommendation, the weight bias factor of the interest score is -0.2, and the weight bias factor corresponding to the recommendation score is 0.2. Then the target interest weight is 0.2, and the target recommendation weight is 0.8. Then perform a weighted sum of the weight and the interest score to obtain the target recommendation score.
[0121] Among them, the recommendation type of the recommendation task for the target object can be determined by the server, or by the target object, or by analyzing the object behavior. For example, if the object behavior sequence is relatively rich, it tends to long-tail recommendation; if the object behavior sequence is relatively single, it tends to popularity recommendation; for another example, if it is determined by the object ID that the activity degree of the object is less than the preset degree threshold, it tends to long-tail recommendation; in addition, it can also be determined by the scenario. For example, during holidays, it tends to popularity recommendation.
[0122] S540. Recommend the target recommended content to the target object according to the target recommendation scores of each candidate recommended content.
[0123] For each candidate recommended content, the target recommendation score can be obtained through steps S510 to S530, and then content is recommended to the target object according to the high and low of the target recommendation scores of each candidate recommended content.
[0124] It should be noted that Figure 5 For the detailed introduction of S210 to S230 shown in Figure 2 please refer to S210 to S230 shown in
[0125] In the embodiments of the present application, the interest weight corresponding to the interest score and the recommendation weight corresponding to the recommendation score are dynamically calculated using the characteristics corresponding to the object's basic information, and the two weights are used to perform weighted fusion on the scores, making up for the deficiencies of a single information source, comprehensively depicting the object's personalized interest in the candidate recommended content to obtain a more accurate target recommendation score, and then recommending the target recommended content to the target object according to the recommendation scores of each candidate recommended content, improving the personalization and accuracy of the recommendation.
[0126] The embodiments of the present application also provide another content recommendation method, which can be applied to Figure 1 the implementation environment shown in Figure 6 This method can be executed by the terminal or the server, or jointly executed by the terminal and the server. In the embodiments of the present application, taking the example that this method is executed by the server, as Figure 2 shown, on the basis of what is shown in Figure 2 S210 shown in
[0127] S610. Encode the behavior sequence through the multi-head attention mechanism to obtain the behavior sequence encoding.
[0128] In the embodiments of the present application, the multi-head attention mechanism can capture the global dependency relationship between each behavior in the behavior sequence. Through the multi-head attention mechanism, each behavior in the behavior sequence can fully interact with other behaviors to obtain the behavior sequence encoding.
[0129] Among them, the behavior sequence input into the multi-head attention mechanism is [behavior sequence length, feature vector length]. The query vector, key vector, and value vector are obtained by linearly transforming the behavior sequence through three different weight matrices. The dot product of the query vector and the key vector is calculated, and then the Softmax function is applied to obtain the attention weights. The attention weights are multiplied by the value vector to obtain the output encoding. In the multi-head attention mechanism, the above process is performed in parallel multiple times, and each time a different linear transformation weight matrix is used to generate multiple different attention heads. The output results of different attention heads are concatenated and then linearly transformed again to generate the final behavior sequence encoding, and the behavior sequence encoding is [behavior sequence length, encoding vector length].
[0130] S620. If the content attribute information of the candidate recommended content is a content identifier, then use the content identifier as the content sequence, and construct the candidate content sequence encoding according to the feature dimension and the preset sequence length of the content sequence.
[0131] In the embodiments of the present application, in some recommendation system datasets, the candidate recommended content is only represented by a content identifier and cannot be constructed into a sequence. Then, a content sequence encoding needs to be constructed, which is [preset sequence length, ID vector length]. The preset sequence length can be 1 because the sequence length is 1 and there is no need for interaction of sequence features.
[0132] S630. If the content attribute information of the candidate recommended content includes a content title, then generate a content sequence according to the content title, and process the content sequence through the multi-head attention mechanism to obtain the content sequence encoding.
[0133] In some recommendation system datasets, there is at least one title for the candidate recommended content. At this time, a content sequence can be generated according to each character of the title. It can be understood that in addition to the title, there is also rich content information such as description, category, etc. A content sequence generated by adding a description can also be generated on the basis of the title.
[0134] The process of processing the content sequence through the multi-head attention mechanism to obtain the content sequence encoding is specifically the same as the process for the behavior sequence, which will not be elaborated here.
[0135] It should be noted that Figure 6 For other detailed introductions of S220 - S240 shown in Figure 2 please refer to S220 - S240 shown in
[0136] In the embodiments of the present application, the multi-head attention mechanism can capture the global dependencies between all positions in the behavior sequence and ensure the accuracy of extraction. For the content sequence, different encoding methods are adopted to adapt to various scenarios.
[0137] For ease of understanding, an embodiment of the present application also provides a content recommendation method, which is illustrated by taking the recommended items of a browser as an example. The content recommendation method proposed in the embodiment of the present application encodes the object behavior sequence by using the transformer global self-attention attention mechanism; uses the transformer cross attention to combine the object behavior sequence and the candidate item sequence to obtain a sequence score and an interest score; in addition, a personalized dynamic weight adjustment module is proposed to dynamically combine the interest score obtained by scoring the sequence structure and the recommendation score estimated by the main model to obtain a more accurate target recommendation score, so as to provide better recommendation services for the object according to the target recommendation score.
[0138] Among them, the content recommendation method is executed by a recommendation model, see Figure 7 As shown, the recommendation model includes an object sequence encoding module, an optional candidate item encoding module, a sequence cross module, a personalized dynamic weight adjustment module, and a recommendation main tower.
[0139] Among them, there are various choices for the object sequence encoding module. In the embodiment of the present application, a BERT (Bidirectional Encoder Representations from Transformers) encoder is selected. The core of this encoder uses Transformer global self-attention, which can fully cross the input feature sequence and can perform parallel computing with relatively high speed. The module structure is as Figure 8 shown. The object clicks on the browsing items shown in the browser to obtain a behavior sequence, and then the behavior sequence is input into the object sequence encoding module. First, it is processed by a multi-head attention layer to capture the relationships between different parts of the sequence. The output of the multi-head attention mechanism is added to the input through Add&Norm, and then normalized; the normalized output is further processed by a feed-forward neural network; the output of the feed-forward neural network is added to the input again, and then normalized; the above process will be looped Nx times to enhance the expression ability of the model, and finally the encoding result of the object behavior sequence is output.
[0140] It should be noted that other sequence layers such as a recurrent neural network (RNN) or a long short-term memory network (LSTM) can also be selected for the object sequence encoding module to achieve a similar purpose.
[0141] The candidate item encoding module and the object sequence encoding module have the same structure, but this is optional here because in many publicly available datasets of recommendation systems, there is only one ID for the candidate item and it cannot be constructed into a sequence. However, in most real application scenarios, there is at least one title on the item side, and this title can be regarded as the candidate item sequence.
[0142] The embodiments of the present application are general for the above two situations. Specifically, when there is only an ID on the item side, the shape of the input matrix on the item side will be constructed as [1, item vector length]. At this time, since the sequence length is 1, the candidate item encoding module is not needed. If the item side is a sequence, the shape of the input matrix on the item side will be constructed as [item sequence length, item vector length], and the candidate item encoding module is used to encode the candidate item.
[0143] It should be noted that in the sequence cross module, Transformer cross Attention is proposed to perform feature interaction on the behavior sequence encoding and the item content sequence encoding. A common practice in the industry is to directly connect the features on the object side and the candidate item side, and through a structure similar to the above encoding module, perform interaction on the features of both sides. The problem with this approach is that the sequence length is usually variable. When directly connecting the two sequences for model training, it is difficult for the model to distinguish the boundaries of the two sequences. Therefore, the model needs to continuously adjust the parameters to adapt to the variable-length sequences, and it is difficult to learn effective weights for the cross of the two sequences. In the sequence cross module, Transformer cross Attention is used to perform feature cross (interaction) on the two sequences, and its structure is as Figure 9 shown. This structure is a bit similar to the encoding module. The difference is that in the input of the multi-head attention layer, Q and K, V belong to two different sequences respectively.
[0144] The object behavior sequence and the item content sequence are not always of the same length. For example, one object may have 10 points of interest, and another object may have 8 points of interest. The same is true for the item sequence. However, model training requires a fixed matrix shape. If calculated directly, there will be problems. For example, if the set matrix sequence length is 10, but the object sequence length is only 8, then the remaining two element slots will both be 0 or a certain fixed value. If these fixed values participate in the calculation, the finally calculated result will not match the expectation. Therefore, two sequence masks of the two sequences need to be used to control the attention mechanism process respectively to achieve the purpose of enabling the interaction of the two variable-length sequences.
[0145] A mask is a matrix where all values are either 0 or 1. For example, [1, 1, 1, 0, 0] represents a sequence of length 5, where the first three slots have content and the last two elements are empty. By taking the union of the mask of the object behavior sequence and the item sequence and participating in the calculation process of the intersection of the two sequences, it is possible to control that only the part of the sequence with content participates in the calculation, thus ensuring that the calculation proceeds as expected.
[0146] The common practice in the industry is that only one sequence is variable-length. For example, GPT (Generative Pre-trained Transformer) actually has only one sequence, and the mask is only for this one sequence. In the sequence intersection module, the union of the two masks will be used, which can enable the two sequences to interact, so that the sequence features on the object side and the item side can have direct and sufficient interaction. For the specific processing process of the sequence intersection module, please refer to the foregoing embodiments.
[0147] Through the sequence intersection module, an effective interest score can be obtained using sequence features. In practice, there is a main recommendation tower in the recommendation model that uses all features. Although this main recommendation tower cannot utilize the sequence information in the sequence features, it can effectively use various heterogeneous information for recommendation scoring. Among them, the object-side features are the comprehensive attribute information of the object, the item-side features are the content information, the cross features are the cross information between the object and the item, and other features such as environmental information; in one example, the main recommendation tower includes an embedding layer, an interaction layer, a fusion layer, and an output layer; in another example, the main recommendation tower can be DeepFM (Deep Factorization Machine), Transformer-based Models, etc.
[0148] An important problem that arises is how to fuse the interest score obtained from sequence scoring and the recommendation score obtained from the main tower scoring. If the two interest scores are simply added together or a weight is given to the two interest scores based on some logic and then added, neither can achieve the effect that 1 + 1 is greater than 2.
[0149] Since an important feature of the recommendation model is to provide personalized experiences for different users, the importance of sequence scoring for different objects varies. For example, for a new object without historical sequence information, the value of sequence scoring is very low, and the sequence scoring weight should be very low. On the other hand, for a senior user with consistent interests, their historical sequence dominates the user information, and the corresponding sequence scoring weight should be higher. However, the richness of the object's historical sequence information can be characterized by multiple dimensions, such as the quantity of historical information, the time elapsed since the present, and the degree of continuity. In addition, other dimensions such as the user's activity level can also be used to determine the weight of the object's sequence information. Neural networks have the ability to obtain a set of suitable parameters during backpropagation learning, fuse different parameters, and map them to the target space.
[0150] Therefore, the personalized dynamic weight adjustment module embedded in the main model, as Figure 10 shown, includes three parts: module input, Embedding layer, and personalized dynamic weight mapping network. This module uses the Embedding layer to map object features to a vector space, that is, different values of each feature are represented by a vector. Then, through a personalized dynamic weight mapping network, these features are fused and finally input into the interest weight of sequence scoring and the recommendation weight of the main tower scoring.
[0151] Among them, the richness of the object's historical sequence information is a relatively useful feature, but it is not easy to characterize. In the embodiments of this application, at the minimum, only the object ID as the input of the personalized dynamic weight module can achieve good results. This is due to the memory ability of the neural network. On the basis of the object ID, further adding feature information on the object side as the module input can further improve the model's performance.
[0152] Embedding layer: The input of the model may be heterogeneous data. For example, the ID is a discrete variable, while the age is a numerical variable. The Embedding layer is an effective method to map different heterogeneous data to the same space.
[0153] Personalized Dynamic Weight Mapping Network: The purpose of this network is to fuse various features of an object and finally map them to the recommended weights for sequence scoring and the recommended weights for main tower scoring. Any neural network layer based on gradient descent can achieve this purpose. Therefore, the choice of network can be very diverse. Optionally, an MLP stacked by Linear layers can be used. The number of layers of the MLP can be arbitrarily selected, but it is found in the application embodiments that 1 to 3 layers can achieve better results. The activation function of the Linear layer can also be arbitrarily selected, such as ReLu, SeLu, etc. It is found by comparison in the application embodiments that their effects are not very different. In addition to the structure composed of Linear, other modules can also be selected for the personalized dynamic weight mapping network, such as the FM model, or more complex networks such as SeNet and LHUC network.
[0154] As shown in Table 1 below, it shows the comparison effect of a part of sequence scoring in this application.
[0155] Solution sauc_ctr gauc_ctr auc_ctr Reference Solution 0.637306366 0.615315 0.74699272 The Solution of this Application 0.639306366 0.616915 0.74829272
[0156] Table 1
[0157] Among them, sauc_ctr refers to the auc index grouped by session for click-through rate. A session is a real display of an object. Sauc can better represent the trend of real online indicators compared with auc and gauc. auc_ctr refers to the global auc index for click-through rate; gauc_ctr refers to the auc index grouped by object for click-through rate.
[0158] The reference scheme involves a method of directly connecting object behaviors and target items into a sequence and then directly putting it into the bottom layer of the main tower for training. As shown in Table 1 above, the scheme of this application has increased by 0.200%, 0.160% and 0.130% respectively in sauc_ctr, gauc_ctr and auc_ctr compared with the reference scheme.
[0159] As shown in Table 2, it shows the comparison effect of a scoring fusion in this application.
[0160]
[0161]
[0162] Table 2
[0163] Among them, solution a is a control model, which is a common solution in the industry; solution b fuses the two scores through a common fusion scoring scheme: sigmoid(logit_main + logit_seq), where logit_main is the logit value output by the main tower, and logit_seq is the logit value output by the sequence module part. The sum of the two values is mapped to the probability space through the sigmoid function to obtain a score. The sauc_ctr, gauc_ctr, and auc_ctr can only be improved by 0.036%, 0.092%, and 0.052% respectively.
[0164] Solution c is the solution proposed in this application. The input of the personalized dynamic weight adjustment module only uses the object ID. The sauc_ctr, gauc_ctr, and auc_ctr are improved by 0.444%, 0.330%, and 0.715% respectively, which is 12 times that of solution b and far exceeds the empirical significant point of 0.1%. Solution d further enriches the input of the personalized weight adjustment module, and the improvements of sauc_ctr, gauc_ctr, and auc_ctr can be further increased to 0.491%, 0.452%, and 0.687% respectively.
[0165] Here, the device embodiments of this application are introduced, which can be used to execute the content recommendation method in the above embodiments of this application. For the details not disclosed in the device embodiments of this application, please refer to the embodiments of the content recommendation method above of this application.
[0166] The embodiments of this application provide a content recommendation device, as Figure 11 shown, the device includes.
[0167] An acquisition module 1110, configured to acquire a behavior sequence encoding corresponding to a behavior sequence of a target object and a content sequence encoding corresponding to a content sequence of a candidate recommended content;
[0168] A linear transformation module 1120, configured to perform linear transformation processing on the behavior sequence encoding and the content sequence encoding respectively to obtain a query vector of the behavior sequence encoding, and a key vector and a value vector of the content sequence encoding;
[0169] A feature interaction module 1130, configured to acquire a behavior mask of the behavior sequence and a content mask of the content sequence, and perform feature interaction processing on the query vector of the behavior sequence encoding, the key vector and the value vector of the content sequence encoding according to the behavior mask and the content mask to obtain an interest score of the target object for the candidate recommended content;
[0170] A recommendation module 1140, configured to recommend target recommended content to the target object according to the interest scores of the target object for each candidate recommended content.
[0171] In one embodiment of the present application, based on the foregoing solution, the feature interaction module is further configured to perform a union process on the behavior mask and the content mask to obtain a target mask; perform feature interaction processing on the query vector encoded by the behavior sequence, the key vector and the value vector encoded by the content sequence according to the target mask to obtain an interaction sequence encoding; and perform a linear transformation process on the interaction sequence encoding to obtain the interest score.
[0172] In one embodiment of the present application, based on the foregoing solution, the feature interaction module is further configured to calculate the similarity between each query vector and each key vector to obtain an attention score; obtain a target attention score according to the target mask and the attention score, and perform a weighted sum on each value vector according to the target attention score to obtain an attention output encoding; generate the interaction sequence encoding according to the attention output encoding and the behavior sequence encoding.
[0173] In one embodiment of the present application, based on the foregoing solution, the feature interaction module is further configured to perform feature fusion processing on the attention output encoding and the behavior sequence encoding to obtain a fusion encoding, and perform normalization processing on the fusion encoding to obtain a target encoding; process the target encoding through a feed-forward neural network to generate the interaction sequence encoding.
[0174] In one embodiment of the present application, the recommendation module is further configured to obtain the comprehensive attribute information of the target object and the content information of the candidate recommended content, where the comprehensive attribute information includes basic attribute information; generate a recommendation score of the target object for the candidate recommended content according to the comprehensive attribute information and the content information; fuse the interest score and the recommendation score according to the basic attribute information of the target object to obtain a target recommendation score of the candidate recommended content; and recommend the target recommended content to the target object according to the target recommendation scores of each candidate recommended content.
[0175] In one embodiment of the present application, based on the foregoing solution, the recommendation module is further configured to perform feature extraction on the basic attribute information of the target object to obtain an object basic attribute feature; map the object basic attribute feature to an interest weight corresponding to the interest score and a recommendation weight corresponding to the recommendation score through a neural network layer based on gradient descent; and perform a weighted sum process on the interest score and the recommendation score according to the interest weight and the recommendation weight to obtain the target recommendation score.
[0176] In one embodiment of the present application, based on the foregoing solution, the recommendation module is further configured to map the object basic attribute features to the interest weight and the recommendation weight based on a multi-layer perceptron formed by stacking at least two linear layers; or, obtain the basic content in the content information of the candidate recommended content, and perform feature extraction processing on the basic content to obtain basic content features; map the object basic attribute features and the basic content features to the interest weight and the recommendation weight based on a factorization machine.
[0177] In one embodiment of the present application, based on the foregoing solution, the recommendation module is further configured to obtain the recommendation type of the recommendation task for the target object, and determine a weight bias factor according to the recommendation type, where the weight bias factor is used to indicate the influence of the recommendation type on content recommendation; adjust the interest weight and the recommendation weight according to the weight bias factor to obtain a target interest weight and a target recommendation weight; perform a weighted summation process on the interest score and the recommendation score according to the target interest weight and the target recommendation weight to obtain the target recommendation score.
[0178] In one embodiment of the present application, based on the foregoing solution, the recommendation module is further configured to perform feature extraction on the comprehensive attribute information and the content information respectively to obtain object features and content features; obtain the cross information between the target object and the candidate recommended content, and perform feature extraction on the cross information to obtain cross features; generate the recommendation score according to the object features, the content features, and the cross features.
[0179] In one embodiment of the present application, based on the foregoing solution, the recommendation module is further configured to obtain the current environmental information, and perform feature extraction on the environmental information to obtain environmental features; perform feature interaction processing on the object features and the content features to obtain interaction features; perform feature fusion on the interaction features, the cross features, and the environmental features to obtain full-scale features, and generate the recommendation score according to the full-scale features.
[0180] In one embodiment of the present application, based on the foregoing solution, the acquisition module is further configured to perform encoding processing on the behavior sequence through a multi-head attention mechanism to obtain the behavior sequence encoding; if the content attribute information of the candidate recommended content is a content identifier, construct the content sequence encoding according to the feature dimension and the preset sequence length of the content identifier; if the content attribute information of the candidate recommended content includes a content title, generate the content sequence according to the content title, and perform processing on the content sequence through the multi-head attention mechanism to obtain the content sequence encoding.
[0181] It should be noted that the device provided in the above embodiment and the method provided in the above embodiment belong to the same concept. The specific manners in which each module and unit perform operations have been described in detail in the method embodiment, and will not be elaborated here.
[0182] The device provided in the above embodiment can be provided in a terminal or in a server.
[0183] An embodiment of the present application further provides an electronic device, including one or more processors and a storage device. The storage device is used to store one or more computer programs. When the one or more computer programs are executed by the one or more processors, the electronic device implements the content recommendation method as described above.
[0184] Figure 12 The structural schematic diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown.
[0185] It should be noted that Figure 12 The computer system 1200 of the electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0186] As Figure 12 shown, the computer system 1200 includes a processor (Central Processing Unit, CPU) 1201, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1202 or the program loaded from the storage section 1208 into the random access memory (RAM) 1203, such as executing the method in the above embodiment. In the RAM 1203, various programs and data required for system operations are also stored. The CPU 1201, ROM 1202, and RAM 1203 are connected to each other through a bus 1204. The input / output (Input / Output, I / O) interface 1205 is also connected to the bus 1204.
[0187] In some embodiments, the following components are connected to the I / O interface 1205: an input section 1206 including a keyboard, a mouse, etc.; an output section 1207 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1208 including a hard disk, etc.; and a communication section 1209 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1209 performs communication processing via a network such as the Internet. The drive 1210 is also connected to the I / O interface 1205 as needed. A removable medium 1212, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1210 as needed so that a computer program read therefrom can be installed into the storage section 1208 as needed.
[0188] Specifically, according to an embodiment of the present application, the processes described above with reference to the flowcharts can be implemented as a computer program. For example, an embodiment of the present application includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1209, and / or installed from the removable medium 1212. When the computer program is executed by a processor (CPU) 1201, various functions defined in the system of the present application are executed.
[0189] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The computer program contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0190] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of apparatuses, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and a computer program.
[0191] The units or modules involved in the embodiments described in this application can be implemented in software or in hardware, and the described units or modules can also be provided in a processor. Among them, the names of these units or modules do not, in some cases, constitute a limitation on the units or modules themselves.
[0192] On the other hand, this application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the content recommendation method as described above is implemented. The computer-readable storage medium may be included in the electronic device described in the above embodiments, or may exist separately without being assembled into the electronic device.
[0193] On the other hand, this application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. The processor of the electronic device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the electronic device executes the content recommendation method as described in the above embodiments.
[0194] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of this application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0195] Those skilled in the art will readily think of other implementation schemes of this application after considering the specification and practicing the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application, which follow the general principles of this application and include the common general knowledge or conventional technical means in the technical field not disclosed in this application.
[0196] The above content is only a preferred exemplary embodiment of this application and is not used to limit the implementation scheme of this application. Those of ordinary skill in the art can easily make corresponding changes or modifications according to the main concept and spirit of this application. Therefore, the protection scope of this application should be subject to the protection scope required by the claims.
Claims
1. A content recommendation method, characterized in that: include: Obtaining a behavior sequence code corresponding to the behavior sequence of the target object and a content sequence code corresponding to the content sequence of the candidate recommended content; Performing linear transformation processing on the behavior sequence code and the content sequence code respectively to obtain a query vector of the behavior sequence code, and a key vector and a value vector of the content sequence code; Acquire a behavior mask of the behavior sequence and a content mask of the content sequence, and perform feature interaction processing on a query vector encoded by the behavior sequence and a key vector and a value vector encoded by the content sequence according to the behavior mask and the content mask to obtain an interest score of the target object for the candidate recommended content; Recommend target recommended content to the target object according to the interest score of the target object for each of the candidate recommended content.
2. The method according to claim 1, characterized in that The step of performing feature interaction processing on the query vector encoded by the behavior sequence, the key vector and the value vector encoded by the content sequence according to the behavior mask and the content mask to obtain the interest score of the target object for the candidate recommended content includes: Performing a union process on the behavior mask and the content mask to obtain a target mask; Performing feature interaction processing on the query vector of the behavior sequence encoding, the key vector and the value vector of the content sequence encoding according to the target mask to obtain an interaction sequence encoding; The interaction sequence encoding is linearly transformed to obtain the interest score.
3. The method according to claim 2, characterized in that Performing feature interaction processing on the query vector of the behavior sequence encoding, the key vector and the value vector of the content sequence encoding according to the target mask to obtain an interaction sequence encoding, including: For each of the query vectors, calculating the similarity between the query vector and each key vector to obtain an attention score; Obtaining a target attention score according to the target mask and the attention score, and performing weighted summation on each of the value vectors according to the target attention score to obtain an attention output code; The interaction sequence code is generated according to the attention output code and the behavior sequence code.
4. The method according to claim 3, characterized in that The step of generating the interaction sequence code according to the attention output code and the behavior sequence code comprises: Performing feature fusion processing on the attention output code and the behavior sequence code to obtain a fusion code, and normalizing the fusion code to obtain a target code; The target code is processed by a feedforward neural network to generate the interaction sequence code.
5. The method according to any one of claims 1 to 4, characterized in that: The step of recommending target recommended content to the target object according to the target object's interest score for each of the candidate recommended content comprises: Acquire comprehensive attribute information of the target object and content information of the candidate recommended content, wherein the comprehensive attribute information includes basic attribute information; Generating a recommendation score of the target object for the candidate recommended content according to the comprehensive attribute information and the content information; According to the basic attribute information of the target object, the interest score and the recommendation score are merged to obtain a target recommendation score for the candidate recommended content; The target recommended content is recommended to the target object according to the target recommendation score of each of the candidate recommended content.
6. The method according to claim 5, characterized in that The step of fusing the interest score and the recommendation score according to the basic attribute information of the target object to obtain a target recommendation score for the candidate recommended content includes: Extracting features from the basic attribute information of the target object to obtain basic attribute features of the object; The neural network layer based on gradient descent maps the basic attribute features of the object into an interest weight corresponding to the interest score and a recommendation weight corresponding to the recommendation score; The interest score and the recommendation score are weighted and summed according to the interest weight and the recommendation weight to obtain the target recommendation score.
7. The method according to claim 6, characterized in that The gradient descent-based neural network layer maps the basic attribute features of the object into the interest weight and the recommendation weight, including: Mapping the basic attribute features of the object into the interest weight and the recommendation weight based on a multilayer perceptron formed by stacking at least two linear layers; Or, obtaining basic content in the content information of the candidate recommended content, and performing feature extraction processing on the basic content to obtain basic content features; The object basic attribute features and the basic content features are mapped into the interest weight and the recommendation weight based on a factor decomposition machine.
8. The method according to claim 6, characterized in that The step of performing weighted sum processing on the interest score and the recommendation score according to the interest weight and the recommendation weight to obtain the target recommendation score includes: Acquire a recommendation type for the recommendation task for the target object, and determine a weight bias factor according to the recommendation type, wherein the weight bias factor is used to indicate the influence of the recommendation type on the content recommendation; Adjusting the interest weight and the recommendation weight according to the weight bias factor to obtain a target interest weight and a target recommendation weight; The interest score and the recommendation score are weighted and summed according to the target interest weight and the target recommendation weight to obtain the target recommendation score.
9. The method according to claim 5, characterized in that The step of generating a recommendation score of the target object for the candidate recommended content according to the comprehensive attribute information and the content information includes: Extracting features from the comprehensive attribute information and the content information to obtain object features and content features respectively; Acquire cross information between the target object and the candidate recommended content, and perform feature extraction on the cross information to obtain cross features; The recommendation score is generated according to the object feature, the content feature and the cross-feature.
10. The method according to claim 9, characterized in that Generating the recommendation score according to the object feature, the content feature and the cross feature includes: Obtain current environmental information, and extract features from the environmental information to obtain environmental features; Performing feature interaction processing on the object feature and the content feature to obtain an interaction feature; The interactive features, cross features and environmental features are fused to obtain full features, and the recommendation score is generated according to the full features.
11. The method according to claim 1, characterized in that: The step of obtaining the behavior sequence code corresponding to the behavior sequence of the target object and the content sequence code corresponding to the content sequence of the candidate recommended content includes: The behavior sequence is encoded by a multi-head attention mechanism to obtain the behavior sequence encoding; If the content attribute information of the candidate recommended content is a content identifier, constructing the candidate content sequence code according to the feature dimension of the content identifier and a preset sequence length; If the content attribute information of the candidate recommended content includes a content title, the content sequence is generated according to the content title, and the content sequence is processed through the multi-head attention mechanism to obtain the content sequence encoding.
12. A content recommendation device, characterized in that: include: An acquisition module, used to acquire a behavior sequence code corresponding to a behavior sequence of a target object and a content sequence code corresponding to a content sequence of a candidate recommended content; A linear transformation module, used to perform linear transformation processing on the behavior sequence code and the content sequence code respectively, to obtain a query vector of the behavior sequence code, and a key vector and a value vector of the content sequence code; a feature interaction module, configured to obtain a behavior mask of the behavior sequence and a content mask of the content sequence, and perform feature interaction processing on a query vector encoded by the behavior sequence and a key vector and a value vector encoded by the content sequence according to the behavior mask and the content mask to obtain an interest score of the target object for the candidate recommended content; The recommendation module is used to recommend target recommended content to the target object according to the interest score of the target object for each of the candidate recommended content.
13. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the electronic device to execute the method according to any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor of an electronic device, the electronic device executes the method according to any one of claims 1 to 11.
15. A computer program product, characterized in that The computer program product comprises a computer program, wherein the computer program is stored in a computer-readable storage medium, and a processor of an electronic device reads and executes the computer program from the computer-readable storage medium, so that the electronic device executes the method according to any one of claims 1 to 11.