Content recommendation methods, devices, equipment and media
By using a deep learning model for content recommenders and box triggers, user behavior and interests are analyzed in real time, solving the problem of poor recommendation performance in existing technologies and achieving efficient exposure of relevant content and improved user experience.
Patent Information
- Application Number
- CN202010719551.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-23
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2040-08-27
AI Technical Summary
In existing technologies, recommendation algorithms cannot guarantee that the second article will be exposed to the user, especially since the user may not scroll to the bottom of the screen while reading the first article, resulting in poor recommendation performance.
A content recommender and box triggers using deep learning models analyze user reading behavior and interests in real time to determine the insertion of relevant content into the information stream. The content recommender sorts the content and the box triggers determine the timing of insertion, ensuring that the recommended content is visible when the user returns to the information stream.
By inserting relevant content in real time after the user reads the first piece of content, the problem of users having difficulty seeing recommended content is solved, thereby improving the exposure of recommendations and the user experience.
Smart Images

Figure CN111831917B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and in particular to a content recommendation method, apparatus, device, and medium. Background Technology
[0002] Related recommendations are a way to expand reading. After a user has read the first article, related articles are recommended to them.
[0003] Related technologies offer a large number of deep learning-based recommendation algorithms, primarily driven by Click-Through-Rate (CTR). The first article a user has already read is considered a seed article. A candidate article set is retrieved based on the seed article. A deep learning model then sorts these candidate articles in descending order of CTR, selecting one or more of the highest-ranking candidate articles as the second article. The recommended second article is displayed at the bottom of the first article's content screen on the client-side interface.
[0004] Since users may not scroll to the bottom of the screen while reading the first article, it is difficult to guarantee that the second article will be exposed to them. Summary of the Invention
[0005] This application provides a content recommendation method, apparatus, device, and medium, which can provide a real-time related content recommendation scheme, inserting a second article into the information stream in real time after a user reads the first article. The technical solution is as follows:
[0006] According to one aspect of this application, a content recommendation method is provided, the method comprising:
[0007] Sending an information stream to the client, the information stream including first content;
[0008] In response to the client displaying the content interface of the first content, a deep learning model is invoked to determine the second content, which is recommended as related content to the first content;
[0009] The second content is sent to the client, and the second content is used to add the content to the information stream after a return operation is triggered on the content interface of the first content.
[0010] According to one aspect of this application, a content recommendation method is provided, the method comprising:
[0011] Displays the first content in the information stream;
[0012] In response to the triggering operation of the first content, the content interface of the first content is displayed;
[0013] In response to the return operation of the content interface, a second piece of content is added to the information stream, which is recommended as related content to the first content.
[0014] According to another aspect of this application, a content recommendation device is provided, the device comprising:
[0015] The sending module is used to send an information stream to the client, the information stream including first content;
[0016] The calling module is used to respond to the client displaying the content interface of the first content, and to call a deep learning model to determine the second content, which is recommended as related content to the first content;
[0017] The sending module is further configured to send the second content to the client, wherein the second content is added to the information stream after a return operation is triggered on the content interface of the first content.
[0018] According to another aspect of this application, a content recommendation device is provided, the device comprising:
[0019] The display module is used to display the first content in the information stream;
[0020] An interaction module is used to respond to the triggering operation of the first content and display the content interface of the first content;
[0021] The interaction module is also configured to respond to a return operation of the content interface by adding and displaying second content in the information stream, wherein the second content is recommended as related content to the first content.
[0022] According to another aspect of this application, a computer device is provided, the computer device including a processor and a memory, the memory storing at least one instruction, at least one program, code set or instruction set, the at least one instruction, the at least one program, the code set or instruction set being loaded and executed by the processor to implement the content recommendation method as described above.
[0023] According to another aspect of this application, a computer-readable storage medium is provided, wherein at least one instruction, at least one program, code set, or instruction set is stored therein, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the content recommendation method as described above.
[0024] According to another aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method recommended in the above aspect.
[0025] The beneficial effects of the technical solutions provided in this application include at least the following:
[0026] By inserting a second piece of content related to the first content into the information stream after the user reads the first content, the second content can be viewed when the user returns to the information stream user interface. This solves the problem in related technologies where the second content is inserted at the bottom of the first article's content interface, and the user may not necessarily scroll to the bottom of the interface, thus making it difficult to ensure that the second article is exposed to the user. Attached Figure Description
[0027] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 This is a flowchart of a content recommendation method provided in an exemplary embodiment of this application;
[0029] Figure 2 This is a block diagram of a computer system provided in an exemplary embodiment of this application;
[0030] Figure 3 This is a schematic diagram of the framework of a content recommendation method provided in another exemplary embodiment of this application;
[0031] Figure 4 This is a flowchart of a content recommendation method provided in another exemplary embodiment of this application;
[0032] Figure 5 This is a schematic diagram of a content recommender model provided in an exemplary embodiment of this application;
[0033] Figure 6 This is a flowchart of a content recommendation method provided in another exemplary embodiment of this application;
[0034] Figure 7 This is a block diagram of a content recommendation apparatus provided in an exemplary embodiment of this application;
[0035] Figure 8 This is a block diagram of a content recommendation apparatus provided in an exemplary embodiment of this application;
[0036] Figure 9 This is a block diagram of a computer device provided in an exemplary embodiment of this application. Detailed Implementation
[0037] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0038] First, a brief introduction to some of the terms used in this application is provided:
[0039] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0040] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0041] Machine Learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instruction-based learning.
[0042] Information feed: Also known as a news feed, it is a data stream where multiple pieces of content are displayed after being sorted according to a sorting method. Sorting methods include, but are not limited to, sorting by timeline, sorting by interest, etc. Users can scroll up and down in the information feed to view multiple pieces of content.
[0043] Content (Item): Refers to a single unit of information in the information stream. Content includes, but is not limited to: news, articles, pictures, videos, short videos, and articles combining text and images. In this embodiment, we will use an article as an example. Each article includes: article title, article body, article author, publication time, and one or more of the accompanying images. When displaying a piece of content in the information stream, a box is generally used to display the summary information of the content, such as the article title, article description, article author, and accompanying images. After the user clicks on the box, they are redirected to the content interface, which displays the detailed information of the content. The box can be represented as a list item, a box, a menu, etc.
[0044] This application provides a real-time relevant recommendation (R3S) scheme. This scheme is implemented using a deep learning network, which includes an item recommender (IR) and a box trigger (BT). Figure 1 As shown, an information stream 10 is displayed on the client side, including first content 12 and other content 14. When a user clicks on first content 12 and enters its content interface to read, the content recommender sorts multiple related content items (including second content 16) in the background, assuming the second content is ranked first. When the user finishes reading first content 12 and exits its content interface, the box trigger decides whether to insert second content 16 into information stream 10 in real time based on the user's preference for the first content and latency costs. When insertion is decided, the box trigger inserts second content 16 into information stream 10, displaying it between first content 12 and other content 14.
[0045] Figure 2 This illustration shows a structural block diagram of a computer system 100 provided in an exemplary embodiment of this application. The computer system 100 can be an instant messaging system, a news push system, a shopping system, an online video system, a short video system, a social client that aggregates users based on topics, channels, or circles, or other client systems with social attributes; this embodiment does not limit the scope of the application. The computer system 100 includes: a first terminal 120, a server cluster 140, and a second terminal 160.
[0046] The first terminal 120 is connected to the server cluster 120 via a wireless or wired network. The first terminal 120 can be at least one of a smartphone, game console, desktop computer, tablet computer, e-book reader, MP3 player, MP4 player, and laptop computer. The first device 120 has a client installed and running that supports information recommendation. This client can be any of an instant messaging system, news push system, shopping system, online video system, short video system, social client that aggregates users based on topics, channels, or circles, or other client systems with social attributes. The first terminal 120 is the terminal used by the first user, and the client running on the first terminal 120 has a first account logged in.
[0047] The first terminal 120 is connected to the server 140 via a wireless network or a wired network.
[0048] Server cluster 140 includes at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. Server cluster 140 provides backend services to clients supporting information recommendation. Optionally, server cluster 140 undertakes the primary computing work, while the first terminal 120 and the second terminal 160 undertake secondary computing work; or, server cluster 140 undertakes secondary computing work, while the first terminal 120 and the second terminal 160 undertake the primary computing work; or, server cluster 140, the first terminal 120, and the second terminal 160 collaborate in a distributed computing architecture.
[0049] Optionally, the server cluster 140 includes an access server 142 and an information recommendation server 144. The access server 142 provides access services and information recommendation services to the first terminal 120 and the second terminal 160, and sends recommended related information (at least one of articles, images, audio, and video) from the information recommendation server 144 to the terminals (first terminal 120 or second terminal 160). The information recommendation server 144 can be one or more. When there are multiple information recommendation servers 144, at least two information recommendation servers 144 may be used to provide different services, and / or at least two information recommendation servers 144 may be used to provide the same service, such as providing the same service in a load-balanced manner; this embodiment does not limit this. The information recommendation server 144 is equipped with a content recommender and a box trigger.
[0050] The second terminal 160 has a client installed and running that supports information recommendations. This client can be any of the following: an instant messaging system, a news push system, a shopping system, an online video system, a short video system, a social client that aggregates users based on topics, channels, or circles, or any other client system with social attributes. The second terminal 160 is the terminal used by the second user. A second account is logged into the client on the second terminal 120.
[0051] Optionally, the first account and the second account reside in a virtual social network, which includes a social relationship chain between the first account and the second account. This virtual social network can be provided by the same social platform, or it can be collaboratively provided by multiple social platforms with related relationships (such as authorized login relationships). This application embodiment does not limit the specific form of the virtual social network. Optionally, the first account and the second account can belong to the same team, the same organization, have a friend relationship, or have temporary communication permissions. Optionally, the first account and the second account can also be strangers. In summary, this virtual social network provides a one-way or two-way message transmission path between the first account and the second account.
[0052] Optionally, the clients installed on the first terminal 120 and the second terminal 160 are the same, or the clients installed on the two terminals are the same type of client on different operating system platforms, or the clients installed on the two terminals are different but support information exchange. Different operating systems include: Apple operating system, Android operating system, Linux operating system, Windows operating system, etc.
[0053] The first terminal 120 can refer to one of multiple terminals, and the second terminal 160 can also refer to one of multiple terminals. This embodiment only uses the first terminal 120 and the second terminal 160 as examples. The first terminal 120 and the second terminal 160 may be of the same or different types. These terminal types include at least one of the following: smartphones, game consoles, desktop computers, tablets, e-book readers, MP3 players, MP4 players, and laptop computers. The following embodiment uses the example of the first terminal 120 and / or the second terminal 140 being smartphones, and the existence of a friend relationship chain between the first account and the second account.
[0054] Those skilled in the art will understand that the number of terminals can be more or less. For example, there may be only one terminal, or there may be dozens or hundreds of terminals, or even more. In this case, the computer system may also include other terminals 180, where one or more of these other terminals have a second account logged in that is a friend of the first account. This application does not limit the number or type of terminals in its embodiments.
[0055] Figure 3 A flowchart illustrating a content recommendation method provided in an exemplary embodiment of this application is shown. This embodiment applies the method to... Figure 2 The following example demonstrates the execution of the method in a client. The method includes:
[0056] Step 320: Display the first content in the information stream;
[0057] The client displays an information feed. The feed includes multiple pieces of content ordered by recommendation, each occupying a separate display box. For example, in an article list using a list layout, each piece of content occupies a list item, displaying at least one of the following: article title, article summary, article image, author, and update time.
[0058] In one example, the client displays the first and third content in the information stream. The third content is the other content in the current information stream that is after the first content, that is, the third content is after the first content.
[0059] Step 340: In response to the triggering operation of the first content, display the content interface of the first content;
[0060] When a user shows interest in the first piece of content, a trigger action is used to display the content interface of that first piece of content. Trigger actions include, but are not limited to, at least one of the following: mouse operation, touchscreen operation, eye-tracking operation, and voice control operation. For example, if a user clicks on the box containing the first piece of content, it triggers a jump from the main information feed interface to the content interface of that first piece of content.
[0061] The content interface for the first item displays the details of that item. When the first item is an article, it displays the main text of the article. The content interface for the first item is the interface that appears after clicking on a list item in the first item, switching from the main information feed interface.
[0062] Step 360: In response to the return operation of the content interface, add and display second content in the information stream. The second content is recommended as related content to the first content.
[0063] The second content is recommended by the server or client. The second content and the first content are of the same type, or the second content and the first content are of different types. The second content is recommended as related content to the first content.
[0064] In one example, in response to a back action on the content interface, the user is redirected from the first content's interface to the main feed interface, inserting the second content between the first and third content in the feed. In another example, the third content in the feed is replaced with the second content.
[0065] The box occupied by the second content is located after the box occupied by the first content.
[0066] In summary, the method provided in this embodiment inserts second content related to the first content into the information stream for display after the user reads the first content in the information stream. When the user returns to the user interface of the information stream, they can view the second content. This solves the problem in related technologies where the second content is inserted at the bottom of the content interface of the first article, and the user may not necessarily scroll to the bottom of the interface, thus making it difficult to ensure that the second article is exposed to the user.
[0067] Figure 4 A flowchart of a content recommendation method provided in another exemplary embodiment of this application is shown. This embodiment applies the method to... Figure 2 The example shown is executed on the server shown. The method includes:
[0068] Step 420: Send an information stream to the client, the information stream including the first content;
[0069] The information feed includes multiple content items (content items or list items) sorted in recommendation order. Taking an article list with a list layout as an example, each content item is a list item, and the list item displays at least one of the following: article title, article summary, article image, author, and update time.
[0070] In one example, the information flow includes a first content and a third content. The third content is other content in the current information flow that follows the first content, that is, the third content is located after the first content.
[0071] Step 440: In response to the client displaying the content interface of the first content, determine the second content, which is recommended as related content to the first content;
[0072] The content interface is a separate user interface from the main information feed interface. When a user triggers the box for the first piece of content on the main interface, they are redirected to the content interface of that first piece of content; conversely, when a user triggers a back action on the content interface, they are redirected back to the main interface.
[0073] After the user clicks on the box containing the first content, and the client displays the content interface of the first content, the client sends a click event to the server. This click event carries the identifier of the first content, and the server determines that the client has displayed the content interface of the first content based on the click event.
[0074] Optionally, the server invokes a deep learning model to determine the second content. The deep learning model includes a mechanism for determining the second content and a mechanism for deciding whether to add it to the information stream for display. The second content is recommended as related content to the first content.
[0075] Optionally, the deep learning model includes a content recommender and a box trigger. In response to the client displaying the first content, the server invokes the content recommender in the deep learning model to determine the second content from multiple candidate content items corresponding to the first content; the server then invokes the box trigger in the deep learning model to determine whether to send the second content, based at least on the delay cost, which is the cost of adding the second content to the information stream.
[0076] Step 460: Send the second content to the client. The second content is used to add the content to the information stream after the return operation is triggered on the content interface of the first content.
[0077] In one example, the second content is used to be inserted between the first and third content in the information stream after a back action is triggered on the content interface of the first content. In another example, the second content is used to replace the third content in the information stream after a back action is triggered on the content interface of the first content.
[0078] In summary, the method provided in this embodiment inserts second content related to the first content into the information stream for display after the user reads the first content in the information stream. When the user returns to the user interface of the information stream, they can view the second content. This solves the problem in related technologies where the second content is inserted at the bottom of the content interface of the first article, and the user may not necessarily scroll to the bottom of the interface, thus making it difficult to ensure that the second article is exposed to the user.
[0079] The deep learning model described above is introduced below. This deep learning model can also be called a Real-time Relevant Recommendation Suggestion (R3S) model. This deep learning model includes a content recommender and a box trigger. Illustratively, the content recommender and the box trigger have exactly the same network structure, but their input features and training loss functions differ.
[0080] refer to Figure 5 The diagram illustrates the structure of a content recommender 50 in an exemplary embodiment of this application.
[0081] The input features of the content recommender 50 include: seed features, item features, user features, and context features. Seed features are the features of the primary content, i.e., the content features that the user clicks on and reads. Item features are the features of candidate content related to the primary content. User features are the features of the user account using the client, such as account name, gender, age, interests, and region. Context features are the features of the client's operating environment, such as mobile phone model, operating system type, network type, and network region.
[0082] Optionally, the candidate content of the first content is the content recalled based on a fast recall mechanism, such as using the DeepFM model for fast recall to obtain the candidate content of the first content.
[0083] The content recommender 50 consists of: n multi-expert networks 52 and m multi-commenter networks 54, where n and m are both integers greater than 1. For example, n = 3 and m = 4.
[0084] The illustrative multi-expert network 52 includes: Feature Interaction Network (FINet), Similarity Network (SimNet), and Information Gain Network (GNet).
[0085] A multi-expert network 52 performs feature interaction on the input features, outputting n feature matrices. A multi-critic network 54 then fuses these n feature matrices across multiple critics, disciplines, and experts, outputting a feature output vector for each candidate content. The candidate content is then ranked based on these feature output vectors. At least one candidate content ranked highest is designated as the second content.
[0086] The box trigger and the content recommender have the exact same network structure. The difference lies in that the box trigger has more input features, and the two have different loss functions during training.
[0087] Based on Figure 4 In an optional embodiment, since the content recommender includes n multi-expert networks and m multi-commenter networks, step 442 above includes the following sub-steps, such as... Figure 6 As shown:
[0088] Step 442a: Invoke n expert subnetworks to perform feature interaction on the first content, candidate content and additional features to obtain n feature matrices;
[0089] In one example, the n expert subnetworks include at least one of the following: a feature interaction network, a similarity network, and an information gain network.
[0090] 1. The feature interaction network is invoked to perform feature interaction on the first content, candidate content and additional features to obtain the first feature matrix used to represent attention relevance.
[0091] 2. Call the similarity network to perform feature interaction on the first content, candidate content and additional features to obtain the second feature matrix used to represent semantic relevance.
[0092] 3. The information gain network is invoked to perform feature interaction on the first content, candidate content, and additional features to obtain the third feature matrix used to represent information gain.
[0093] The additional features include at least one of user features and context features. The user features are the characteristics of the user account using the client, and the context features are the characteristics of the client's operating environment.
[0094] In this embodiment, the input features are divided into four groups: F U F S F I and F C F S F represents the first content (seed). I Represents candidate content (targetitem), F U F represents user characteristics. C Represents contextual features.
[0095] Referring to the feature domain partitioning method of the DeepFM model, the above input features are divided into features of different feature domains, resulting in multiple sets of feature vectors f. U f S f I and f C .
[0096] To F S Taking groups as an example, F S =Concat(f S 1,…,f S k Each feature group comprises k feature domains. `Concat(·)` is the join operation. `k` is the number of feature domains. Illustratively, a query function `f` is used. i =L(f i ) for each sparse feature f i Projected onto a dense feature vector of dimension d.
[0097] Feature Interaction Network: For sub-step 1 in step 442a, the server calls the feature interaction network to calculate the attention matrix by using a multi-head self-attention mechanism on the first content and candidate content; after combining the user features, context features and attention matrix, the first feature matrix is obtained.
[0098] The input feature matrix of the feature interaction network is F = {f1, ..., f2}. 2k The input feature matrix comprises a combination (or concatenation) of feature matrices of the first content and candidate content. That is, the feature combination F... S and F I A multi-head self-attention mechanism is used to model the feature combination of the first content and candidate content to generate a feature matrix.
[0099] Q j =W j Q F, K j =W j K F, V j =W j V F
[0100] Among them, W j Q W j K W j V Belongs to R d’×d Let `d` be the projection matrix corresponding to the j-th query, key, and value, respectively. `d` represents the dimension `d`, and `d' = d / h` is the distance between the feature domain space and the query. `head` is the j-th output head in the multi-head attention mechanism. j yes:
[0101] head j =Softmax(Q j ·K j V j ;
[0102] Where j ranges from 1 to 2k, and is the input feature matrix. Connecting all the output heads of the multi-head self-attention system, we get:
[0103]
[0104] In this embodiment, an additional original feature f is added to the i-th feature domain, where i can take values from 1 to k. i Short connected components are used to generate feature maps.
[0105]
[0106] in, and It belongs to the feature space The weighted matrix is d, where d is the dimension. ReLU(·) is the non-linear activation function, and Concat(·) is the concatenation operation. The value of I ranges from 1 to 2k, with 2k concatenation operations. Form an attention matrix.
[0107] Finally, this embodiment also includes the feature vector f of user features and context features. U and f C This is incorporated into the final self-network output of the feature interaction network:
[0108]
[0109] Among them, h F It is the first feature matrix output by the feature interaction network. It is a weighted matrix.
[0110] Similarity Network: For sub-step 2 in step 442a, the server calls the similarity network to calculate the first similarity between the first content and the candidate content at the element level through element-wise product; it calls the similarity network to calculate the second similarity between the first content and the candidate content at the feature domain level through inner product; after combining and calculating the user features, context features, first similarity, and second similarity, the second feature matrix is obtained.
[0111] The input features of a similarity network include: primary content and candidate content. The features of the primary content are represented as F. S and F I .
[0112] First, calculate f using element-wise product. S and f I First similarity at the element-level:
[0113]
[0114] Secondly, f is also calculated using the inner product. S and f I Second similarity at the field-level:
[0115]
[0116] Finally, this embodiment also includes the feature vector f of user features and context features. U and f C This is incorporated into the final self-network output of the similarity network:
[0117]
[0118] in, It is the second feature matrix, which is the final output of the similarity network. It is a weighted matrix, ReLU(·) is a non-linear activation function, d1 is the output dimension of ReLU(·), and Concat(·) is the concatenation operation.
[0119] Information Gain Network: For sub-step 3 in step 442a, the server calls the information gain network to calculate the information gain of the first content and candidate content under different feature domain types; after combining and calculating the user features, context features, and information gain, the third feature matrix is obtained.
[0120] The input features of an information gain network include: first content and candidate content. The features of the first content are represented as F. S and F I .
[0121] In this context, the first content and the candidate content share the same elements in the i-th feature domain. An information gain function is defined to represent the information gain from the first content to the candidate content in the i-th feature domain. Assume the feature domain types include: category F. cat and continuous F con The information gain function is as follows:
[0122]
[0123] Among them, f I i f is the sparse feature set of the candidate content in the i-th feature domain. S i Let F be the sparse feature set of the first content in the i-th feature domain, and L(F) be the lookup function from all sparse feature sets in the set F to dense features. Sum(·) is vector addition. For a category domain, first calculate F... I i and F S i The difference set is used to project each sparse feature in the difference set onto its dense features. The sum of these dense features is considered as the information gain from the first content to the candidate content.
[0124] Finally, this embodiment also includes the feature vector f of user features and context features. U and f C This is incorporated into the final self-network output of the information gain network:
[0125]
[0126] in, It is the third feature matrix of the final output of the information gain network, used to represent the additional diversified information brought by the candidate content. It is a weighted matrix, ReLU(·) is a non-linear activation function, and Concat(·) is a connection operation.
[0127] Step 442b: Invoke m critic networks to fuse n feature matrices to obtain a feature output vector from multiple experts and critics;
[0128] Multi-commentator network: The server generates a first gate vector based on the first content and user features; it calls m commentator networks to generate commentator vectors based on n feature matrices and the first gate vector; and it generates multi-expert, multi-commentator output vectors based on the first content, candidate content, user features, context features, and commentator vectors.
[0129] Since the semantic relevance and information gain of different candidate contents vary, the weights of the three expert subnetworks are not entirely consistent under different circumstances. Therefore, this application designs a multi-critic, multi-gate mixture of experts strategy. Compared to the multi-task learning framework (MMoE), this embodiment provides a novel multi-critic, multi-gate, multi-expert learning framework (M3oE). Unlike the traditional multi-task learning framework MMoE, M3oE designs a multi-headed commenting strategy, commenting on different expert subnetworks from different perspectives.
[0130] First, in this embodiment, each gate is designed to be associated only with user features and the first content. The gate vector in the j-th person's head is generated using the first content and user features:
[0131]
[0132] Among them, W x j Let x be the weight matrix of the j-th head, and Concat(·) is the concatenation operation. j This represents the importance of the j-th person. The value of j ranges from 1 to m.
[0133] This embodiment uses a single logistic regression layer (Softmax layer) to fuse the feature matrices output by multiple expert networks. A total of m gates are used to fuse the feature matrices from the n expert outputs. Let the j-th gate correspond to the j-th subnetwork among the n expert subnetworks, and let the j-th gate vector be as follows:
[0134]
[0135] Among them, W G jIt is the weight matrix of the j-th gate. It is a gate vector, in which each element corresponds to the importance of each expert.
[0136] Gate vector g j (x j By combining the feature matrices output from the three expert subnetworks (first feature matrix, second feature matrix, and third feature matrix), a commentator vector is generated.
[0137]
[0138] Among them, c j It is a commenter vector that integrates three expert subnetworks.
[0139] Combining feature f a =Concat(f U f S f I f C Generate a vector containing all experts and critics:
[0140]
[0141] in, It is a feature extracted collaboratively by multiple experts and commentators, W M It is a weight matrix. d c The number of critics, in this application, is d. c Let's take m=4 as an example.
[0142] Finally, the final output is obtained through a two-layer fully connected layer:
[0143] h f =MLP(h0);
[0144] Where MLP represents a fully connected layer, h f This is the final output of the content recommender.
[0145] Step 442c: Sort the candidate contents according to the feature output vector to determine the second content.
[0146] In one example, the content recommender described above is trained based on Time-on-item (TOI). TOI is the time a user spends on the first item. This is because traditional techniques use CTR as the output target of the ranking model, but since CTR is easily manipulated by information publishers, this embodiment uses TOI as the training target to train the content recommender.
[0147] Optionally, the loss function is defined as:
[0148]
[0149] Where y represents the discretized user features and the TOI (reading time) of the first content, N is the overall sample set, and W... T RR It is a weight vector. N a That is, all samples.
[0150] During the prediction phase, the content recommender outputs a feature vector for each candidate content. Based on these feature vectors, the candidate content is ranked to determine the second content. For example, the top three candidate content items are selected as the second content.
[0151] Step 444a: The server calls the box trigger to generate a second gate vector using the first content, the second content, the user's interaction features when reading the first content, and the third content;
[0152] The purpose of box triggers is to determine whether the system should insert relevant boxes in real-time within the information stream, considering overall performance. When the information stream uses a list representation, list item-based triggering is an implementation method in R3S. Because inserting a list item in real-time delays other list items after the first item, excessive delays should be avoided. These delayed list items may miss the opportunity to leave a lasting impression on the user. If the user does not continue reading, it will severely damage the user experience within the information stream.
[0153] In addition to the four features used by the content recommender, the box trigger introduces two additional features:
[0154] User interaction characteristics when reading the first content f US This interaction feature represents the interest characteristics between the user and the primary content, such as reading time, whether a comment was made, whether a like was given, the number of comments, and comment sentiment analysis. In this embodiment, reading time is used to represent the interaction feature.
[0155] The third content is the content that follows the first content in the information stream. It is also the content that is squeezed out of the information stream after the second content is inserted, denoted as f. D .
[0156] First, the box trigger generates a second gate vector using the first content, candidate content, user interaction features while reading the first content, and the third content:
[0157]
[0158] The second gate vector is calculated in the same way as the first gate vector, the difference being the addition of feature f. USand f D . It is a weighted matrix.
[0159] Step 444b: Call the box trigger to calculate the loss of sending the second content based on the second gate vector, the feature output vector of the second content, the interaction features, and the third content. The loss includes: delay cost and loss of user interest in the first content.
[0160] The box trigger calculates the loss for adding candidate content based on the second gate vector, feature output vector, interaction features, and third content. This loss includes: delay cost and user interest loss on the first content.
[0161] h′ f =MLP(Concat(h′0,f D f US ));
[0162] Where h′0 is the feature output vector obtained by the box trigger through multi-expert, multi-sect, and multi-commentator fusion calculation based on the first content, the second content, user features, and contextual features. The calculation method of h′0 is the same as h o Similarly, referring to the above introduction to content recommenders, it will not be repeated here. MLP represents a fully connected layer, h′ f It is the final output of the box flip-flop.
[0163] Step 444c: Determine whether to send the second content based on the loss.
[0164] If the loss of at least one second content is lower than a preset condition, the second content is sent to the client and inserted into the information stream for display.
[0165] In one example, the box trigger is trained based on click-through rate (CTR) loss and delay cost.
[0166] The box-level loss based on CTR can be expressed as:
[0167]
[0168] Among them, W T BT It is a weight vector, T is the transpose, and σ(·) is a sigmoid growth function.
[0169] Unlike the traditional CTR loss function, the box trigger loss function combines the CTR loss function with the delay cost. Taking the delay cost into account, the box trigger loss function can be expressed as:
[0170]
[0171] Where, Nd This is a new negative sample set used to measure latency costs. In N, the second content (inserted content) corresponding to a sample was not clicked, but other content following the first content of the sample was clicked—this is precisely the situation this application aims to avoid. N p It is the second positive sample of the inserted sample that was clicked, N n This refers to the negative samples whose second content was not clicked after insertion. λp, λn, and λd are the hyperparameters of the loss weights. The total loss of R3S is L. RR and L BT The sum of.
[0172] In summary, the method provided in this embodiment allows the content recommender to employ a multi-expert, multi-commentator M3oE mechanism to select superior second content. O E integrates the characteristics of three expert subnetworks, taking into account not only the feature correlation between the first and second content, but also semantic similarity and information gain considerations, thereby providing a second content that is more relevant, semantically similar, and can provide more diverse information.
[0173] Box triggers can combine box-level CTR and latency costs to decide whether to insert second content in real time. This allows for the reasonable insertion of second content without significantly impacting the overall performance of the information stream, thus preserving a better reading experience for users. Furthermore, by considering the user's interest in the first content and latency costs, second content is only inserted when it is predicted that the user is highly likely to read it, avoiding ineffective recommendations that waste network and computing resources.
[0174] Because there is no open dataset for this task, this application constructs a new dataset, WTS-RS, for relevant recommendation suggestions in real-time online scenarios. The WTS-RS dataset randomly selects 21 million users and collects 332 million actual click data from them. A total of 43 million bounding box clicks and 47 million content clicks related to the relevant recommendation suggestion scenario were extracted. For each content click, its time to read (TOI) was also recorded for training and evaluation. The dataset was split into a training set and a test set in chronological order, resulting in 232 million training instances and 100 million test instances.
[0175] Offline training on relevant recommended offline datasets was performed, with both TOI-oriented and CTR-oriented calculations conducted. The metrics used were the area under the ROC curve (AUC) and the relative improvement of the model (RelaImpr). For comparison, the following basic model was introduced:
[0176] FM: Factorization Machine (FM) models all interactions between features using the parameters of factorization. FM is proposed as a fundamental evaluation model.
[0177] Wide & Deep: Wide & Deep consists of a wide portion of the original features and a deep portion for feature interactions.
[0178] NFM: NFM introduces a bidirectional interaction layer before the DNN layer for feature interaction.
[0179] AFM: AFM has drawn attention to the interactive features of dual-interaction layers.
[0180] DeepFM combines FM and DNN in parallel to simulate raw features and higher-order interactions.
[0181] AutoInt: AutoInt introduces a self-focused neural network for original feature interactions.
[0182] Please note that this application does not use traditional query suggestion models or sentence matching models as the base model, because they are designed for different tasks, where query title or sentence similarity is the most basic objective.
[0183] All base models follow the two-step architecture of R3S, which includes a content recommender and a box trigger. The item recommender of the base models is optimized under a TOI-based discrete MSE objective. The trigger is updated in Eq using a CTR-based cross-entropy objective (without delay loss). The content recommender is used for article TOI prediction, while the box trigger is used for box-level CTR prediction. All models and R3S use the same input features and experimental settings in evaluation.
[0184] First, the WTS-RS dataset is used to evaluate the R3S of this application compared to the base model in terms of TOI prediction. The TOI prediction task aims to predict how long a user will spend on the second content in the relevant box. TOI can be seen as an enhanced CTR-related metric because it further considers user reading time rather than being limited to clicks, which reflects the user's true satisfaction. This application ranks the predicted TOIs of all results. In the content recommender, the classic AUC metric is used for evaluation. The RelaIMPR metric is also introduced to provide a relative improvement over the base model.
[0185] The specific data shown in Table 1 is as follows:
[0186] Table 1
[0187] Model AUC RelaImpr FM 0.6949 0.00% AFM 0.7002 2.72% NFM 0.7012 3.23% Wide & Deep 0.7191 12.42% DeepFM 0.7248 15.43% AutoInt 0.7128 9.18% R3S 0.7321 19.09%
[0188] According to Table 1, we can observe that:
[0189] (1) R3S achieves the best performance compared to all base models. This application also performed a significance test to verify the significance level = 0.01.
[0190] (2) This demonstrates that R3S effectively captures multiple factors related to relevant recommendations. The improvements primarily stem from two aspects: a three-expert (i.e., sub-network) and a multi-commentator, multi-expert hybrid strategy. First, the feature interaction network, similarity network, and information gain network each consider different aspects of the feature interactions between the seed and the target item. In this case, R3S can simultaneously consider user behavior, semantic similarity between the seed and the candidate, and information gain, which are essential in relevant recommendation scenarios. Second, the M3oE strategy cleverly combines three experts with multiple commentators who have different weights on these interactions, further improving the TOI prediction performance of the article.
[0191] This application also designs a new task called Box-Level CTR Prediction to evaluate the overall performance of related boxes. Box-Level CTR Prediction aims to predict whether a user will click on a related box (secondary content or category item), which is determined by the box trigger. In R3S, the goal is to cultivate users' habit of using the application's related recommendation feature. Therefore, this application encourages users to click on more related live-inserted articles, which are secondary content inserted below the articles they click. Therefore, CTR is considered the primary evaluation metric for the box trigger. Following the same metrics, AUC and RelaIMPR are also used for item TOI prediction, as shown in Table 3 below:
[0192] Table 2
[0193] Model AUC RelaImpr FM 0.7658 0.00% AFM 0.7704 1.73% NFM 0.7724 2.48% Wide & Deep 0.7866 7.83% DeepFM 0.7901 9.14% AutoInt 0.7807 5.61% R3S 0.7953 11.10%
[0194] Table 2 shows the comparison results of all models in the box-level CTR prediction dimension, from which we can see that:
[0195] (1) In box-level CTR prediction, R3S significantly outperformed all base models, with a significance level of 0.01. R3S confirms that after a user finishes reading an article, it can decide whether to insert related articles in real time. The box trigger is designed to control the frequency of inserting related articles, which is crucial for overall recommendation performance.
[0196] (2) Box-level CTR prediction and article TOI prediction are two similar tasks, but they still have some differences. Article TOI prediction primarily evaluates the content recommender's IR. Conversely, box-level CTR prediction primarily evaluates the box trigger. The box trigger is still trained using clicks because this application considers the clicks for inserting articles to be the most important reward in the box trigger. As a suggestion, the box trigger should consider the side effects and impacts of real-time box insertion on the overall system, rather than box-level performance. Therefore, the box trigger integrates user satisfaction with the seed and uses latency cost as a penalty to balance overall performance and box-level performance. The AUC of article TOI is lower than that of box CTR because article TOI prediction is more granular and challenging.
[0197] In the recommended online A / B testing, the metrics used are overall system TOI (Time-on-item inoverall system), box-level CTR (BCTR), box-level user has-click rate (BUHR), and box-level item views (BIV). Specific data is shown below:
[0198] Table 3
[0199]
[0200]
[0201] The experimental results show the percentage improvement of R3S compared to the online deep model FM. Table 3 shows that:
[0202] (1) R3S (Content Recommender + Box Trigger) achieved significant improvements across all overall and box-level metrics, with a significance level of 0.01. This validates the effectiveness of R3S in real-world scenarios. Improvements in the TOI dimension indicate that users are more satisfied and willing to spend more time on the recommendation results. Improvements in the three box-level metrics also demonstrate that R3S can subtly control the frequency of inserting relevant articles. Therefore, the quality of relevant articles is improved (see BCTR results), and more users are willing to engage in continuous extended reading through relevant articles (see BUHR results), thus enabling R3S to achieve more user interaction with relevant articles (see BIV results).
[0203] (2) In addition to BCTR, R3S (Content-only Recommender) also outperforms traditional deep learning models on three metrics. This is because changing the training objective from CTR to TOI naturally increases the time users spend reading the overall content, while inevitably harming box-level CTR. However, the improvements in BUHR and BIV indicate that more users will utilize the recommendation feature of this application, and more related articles will be clicked.
[0204] (3) The decrease in BCTR brought about by R3S (project) reveals a serious problem of overexposure of related articles. Furthermore, the overall CTR also decreased slightly. To address this issue, this application introduces a box trigger, which can better account for latency costs. After introducing the box trigger, all box-level metrics show impressive improvements compared to R3S (project). Compared to R3S (project), the overall CTR and box-level CTR even reached 0.58% and 17.52% respectively, an improvement of 52%. This further confirms the importance of box triggers.
[0205] Figure 7 A block diagram of a content recommendation apparatus according to an illustrative embodiment of this application is shown. The apparatus can be implemented as a server, or as a module within a server, and includes:
[0206] Sending module 710 is used to send an information stream to a client, the information stream including first content;
[0207] The module 720 is used to respond to the client displaying the content interface of the first content by calling a deep learning model to determine the second content, which is recommended as related content to the first content;
[0208] The sending module 710 is used to send the second content to the client, and the second content is used to be added to the information stream after a return operation is triggered on the content interface of the first content.
[0209] In one design of this application, the deep learning model includes: a content recommender and a box trigger;
[0210] The calling module 720 is configured to, in response to the client displaying the content interface of the first content, call the content recommender in the deep learning model to determine the second content from multiple candidate contents corresponding to the first content; and call the box trigger in the deep learning model to determine whether to send the second content based at least on the delay cost, wherein the delay cost is the impact cost of adding the second content to the information stream.
[0211] In one design of this application, the deep learning model includes: a content recommender and a box trigger;
[0212] The invocation module 720 is used to respond to the client displaying the content interface of the first content, invoking the content recommender to determine the second content from multiple candidate contents corresponding to the first content; and invoking the box trigger to determine whether to send the second content based at least on the delay cost, wherein the delay cost is the impact cost of adding the second content to the information stream.
[0213] In one design of this application, the content recommender includes: n expert subnetworks and m critic networks, where n and m are both integers greater than 1;
[0214] The invocation module 720 is used to invoke the n expert sub-networks to perform feature interaction on the first content, the candidate content, and additional features to obtain n feature matrices; invoke the m critic networks to fuse the n feature matrices to obtain the feature output vector of multiple experts and multiple critics; and sort the candidate content according to the feature output vector to determine the second content.
[0215] The additional features include at least one of user features and context features, wherein the user features are the features of the user account using the client, and the context features are the features of the operating environment of the client.
[0216] In one design of this application, the n expert sub-networks include at least one of: a feature interaction network, a similarity network, and an information gain network;
[0217] The invocation module 720 is used to invoke the feature interaction network to perform feature interaction on the first content, the candidate content, and the additional features to obtain a first feature matrix representing attention relevance; invoke the similarity network to perform feature interaction on the first content, the candidate content, and the additional features to obtain a second feature matrix representing semantic relevance; and invoke the information gain network to perform feature interaction on the first content, the candidate content, and the additional features to obtain a third feature matrix representing information gain.
[0218] In one design of this application, the calling module 720 is used to call the feature interaction network to calculate the attention matrix by using a multi-head self-attention mechanism on the first content and the candidate content; and to obtain the first feature matrix by combining the user features, the context features, and the attention matrix.
[0219] In one design of this application, the calling module 720 is used to call the similarity network to calculate the first similarity between the first content and the candidate content at the element layer through element-wise product; call the similarity network to calculate the second similarity between the first content and the candidate content at the feature domain layer through inner product; and combine the user features, the context features, the first similarity, and the second similarity to obtain the second feature matrix.
[0220] In one design of this application, the calling module 720 is used to call the information gain network to calculate the information gain of the first content and the candidate content under different feature domain types; and after combining and calculating the user features, the context features, and the information gain, the third feature matrix is obtained.
[0221] In one design of this application, the calling module 720 is used to generate a first gate vector through the first content and the user features; call the network of m critics to generate critic vectors based on the n feature matrices and the first gate vector; and generate the feature output vector of multiple experts and multiple critics based on the first content, the candidate content, the user features, the context features, and the critic vectors.
[0222] In one design of this application, the content recommender is trained based on the content consumption time (TOI).
[0223] In one design of this application, the information stream displays a third content while displaying the first content;
[0224] The invocation module 720 is used to invoke the box trigger to generate a second gate vector based on the first content, the second content, the user's interaction features when reading the first content, and the third content; to invoke the box trigger to calculate the loss of sending the second content based on the second gate vector, the feature output vector of the second content, the interaction features, and the third content, wherein the loss includes the delay cost and the user's loss of interest in the first content; and to determine whether to send the second content based on the loss.
[0225] In one design of this application, the box trigger is trained based on click-through rate (CTR) loss and the delay cost.
[0226] Figure 8 A block diagram of a content recommendation apparatus according to an illustrative embodiment of this application is shown. The apparatus can be implemented as a server, or as a module within a server, and includes:
[0227] Display module 820 is used to display the first content in the information stream;
[0228] The interaction module 840 is used to display the content interface of the first content in response to the triggering operation of the first content;
[0229] The interaction module 840 is used to respond to the return operation of the content interface and add the display of second content in the information stream, the second content being recommended as related content to the first content.
[0230] In one design of this application, the display module 820 is used to display the first content and the third content in the information stream, wherein the third content is located after the first content;
[0231] The interaction module 840 is configured to, in response to a return operation of the content interface, insert the second content between the first content and the third content in the information stream; or, replace the third content in the information stream with the second content.
[0232] In one design of this application, the device further includes an acquisition module 860;
[0233] The acquisition module 860 is used to obtain the second content from the server. The second content is calculated by the server using the aforementioned deep learning model.
[0234] It should be noted that the content recommendation device provided in the above embodiments is only an example of the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the content recommendation device and the content recommendation method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0235] This application also provides a computer device (terminal or server) including a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the content recommendation method provided in the above-described method embodiments. It should be noted that the computer device may be as follows: Figure 9 The computer equipment provided.
[0236] Figure 9This illustration shows a structural block diagram of a computer device 900 provided in an exemplary embodiment of this application. The computer device 900 may be a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. The computer device 900 may also be referred to as a user device, portable computer device, laptop computer device, desktop computer device, or other names.
[0237] Typically, computer device 900 includes a processor 901 and a memory 902.
[0238] Processor 901 may include one or more processing cores, such as a quad-core processor or an octa-core processor. Processor 901 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 901 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 901 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 901 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0239] Memory 902 may include one or more computer-readable storage media, which may be non-transitory. Memory 902 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in memory 902 is used to store at least one instruction, which is executed by processor 901 to implement the content recommendation method provided in the method embodiments of this application.
[0240] In some embodiments, the computer device 900 may also optionally include a peripheral device interface 903 and at least one peripheral device. The processor 901, memory 902, and peripheral device interface 903 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 903 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 904, a touch display screen 905, a camera 906, an audio circuit 907, a positioning component 908, and a power supply 909.
[0241] Peripheral device interface 903 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 901 and memory 902. In some embodiments, processor 901, memory 902 and peripheral device interface 903 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 901, memory 902 and peripheral device interface 903 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0242] The radio frequency (RF) circuit 904 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 904 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 904 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 904 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 904 can communicate with other computer devices through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 904 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.
[0243] Display screen 905 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 905 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 901 for processing. In this case, display screen 905 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 905, positioned as the front panel of the computer device 900; in other embodiments, there may be at least two display screens 905, respectively positioned on different surfaces of the computer device 900 or in a folded design; in still other embodiments, display screen 905 may be a flexible display screen, positioned on a curved or folded surface of the computer device 900. Furthermore, display screen 905 may be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. Display screen 905 may be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).
[0244] The camera assembly 906 is used to acquire images or videos. Optionally, the camera assembly 906 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the computer device, and the rear-facing camera is located on the back of the computer device. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 906 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash is a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.
[0245] The audio circuit 907 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 901 for processing, or input to the radio frequency circuit 904 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each located in a different part of the computer device 900. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 901 or the radio frequency circuit 904 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 907 may also include a headphone jack.
[0246] Positioning component 908 is used to locate the current geographical location of computer device 900 in order to enable navigation or LBS (Location Based Service). Positioning component 908 can be a positioning component based on the US GPS (Global Positioning System), China's BeiDou system, or Russia's Galileo system.
[0247] Power supply 909 is used to supply power to the various components in computer device 900. Power supply 909 can be AC power, DC power, a disposable battery, or a rechargeable battery. When power supply 909 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, while a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0248] In some embodiments, the computer device 900 further includes one or more sensors 910. The one or more sensors 910 include, but are not limited to: an accelerometer 911, a gyroscope 912, a pressure sensor 913, a fingerprint sensor 914, an optical sensor 915, and a proximity sensor 916.
[0249] Accelerometer 911 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by computer device 900. For example, accelerometer 911 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 901 can control touchscreen display 905 to display the user interface in landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 911. Accelerometer 911 can also be used for games or for acquiring user motion data.
[0250] The gyroscope sensor 912 can detect the orientation and rotation angle of the computer device 900. The gyroscope sensor 912, in conjunction with the accelerometer sensor 911, can collect 3D motion data from the user on the computer device 900. Based on the data collected by the gyroscope sensor 912, the processor 901 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.
[0251] The pressure sensor 913 can be disposed on the side bezel of the computer device 900 and / or on the lower layer of the touch display screen 905. When the pressure sensor 913 is disposed on the side bezel of the computer device 900, it can detect the user's grip signal on the computer device 900, and the processor 901 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 913. When the pressure sensor 913 is disposed on the lower layer of the touch display screen 905, the processor 901 can control the operable controls on the UI interface based on the user's pressure operation on the touch display screen 905. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0252] The fingerprint sensor 914 is used to collect a user's fingerprint. The processor 901 identifies the user based on the fingerprint collected by the fingerprint sensor 914, or vice versa. When the user's identity is verified as trusted, the processor 901 authorizes the user to perform relevant sensitive operations, including unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings. The fingerprint sensor 914 can be located on the front, back, or side of the computer device 900. When the computer device 900 has physical buttons or a manufacturer's logo, the fingerprint sensor 914 can be integrated with the physical buttons or manufacturer's logo.
[0253] An optical sensor 915 is used to collect ambient light intensity. In one embodiment, the processor 901 can control the display brightness of the touch screen 905 based on the ambient light intensity collected by the optical sensor 915. Specifically, when the ambient light intensity is high, the display brightness of the touch screen 905 is increased; when the ambient light intensity is low, the display brightness of the touch screen 905 is decreased. In another embodiment, the processor 901 can also dynamically adjust the shooting parameters of the camera assembly 906 based on the ambient light intensity collected by the optical sensor 915.
[0254] A proximity sensor 916, also known as a distance sensor, is typically located on the front panel of a computer device 900. The proximity sensor 916 is used to detect the distance between the user and the front of the computer device 900. In one embodiment, when the proximity sensor 916 detects that the distance between the user and the front of the computer device 900 is gradually decreasing, the processor 901 controls the touchscreen display 905 to switch from a screen-on state to a screen-off state; when the proximity sensor 916 detects that the distance between the user and the front of the computer device 900 is gradually increasing, the processor 901 controls the touchscreen display 905 to switch from a screen-off state to a screen-on state.
[0255] Those skilled in the art will understand that Figure 9 The structure shown does not constitute a limitation on the computer device 900, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0256] The memory also includes one or more programs stored in the memory, and the one or more programs include a content recommendation method provided in the embodiments of this application.
[0257] This application provides a computer-readable storage medium storing at least one instruction, which is loaded and executed by the processor to implement the content recommendation method provided in the above-described method embodiments.
[0258] This application also provides a computer program product that, when run on a computer, causes the computer to execute the content recommendation method provided in the above-described method embodiments.
[0259] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0260] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0261] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A content recommendation method, characterized in that, The method includes: Sending an information stream to the client, the information stream including first content; In response to the client displaying the content interface of the first content, the following steps are taken: The content recommender in the deep learning model includes n expert sub-networks that perform feature interaction on the first content, the candidate content corresponding to the first content, and additional features to obtain n feature matrices; the content recommender includes m commentator networks that fuse the n feature matrices to obtain a multi-expert, multi-commentator feature output vector; the candidate content is sorted according to the feature output vector to determine the second content, which is recommended as related content to the first content, where n and m are both integers greater than 1; the additional features include at least one of user features and context features, where the user features are the features of the user account using the client, and the context features are the features of the client's operating environment; the box trigger in the deep learning model is invoked at least based on delay costs to determine whether to send the second content, where the delay costs are the impact costs of adding the second content to the information stream. The second content is sent to the client, and the second content is used to add the content to the information stream after a return operation is triggered on the content interface of the first content.
2. The method according to claim 1, characterized in that, The n expert sub-networks include at least one of the following: feature interaction network, similarity network, and information gain network; The content recommender in the deep learning model includes n expert subnetworks that perform feature interactions on the first content, the candidate content corresponding to the first content, and additional features to obtain n feature matrices, including: The feature interaction network is invoked to perform feature interaction on the first content, the candidate content, and the additional features to obtain a first feature matrix representing attention relevance. The similarity network is invoked to perform feature interaction on the first content, the candidate content, and the additional features to obtain a second feature matrix for representing semantic relevance. The information gain network is invoked to perform feature interaction on the first content, the candidate content, and the additional features to obtain a third feature matrix representing the information gain.
3. The method according to claim 2, characterized in that, The step of invoking the feature interaction network to perform feature interaction on the first content, the candidate content, and the additional features to obtain a first feature matrix representing attention relevance includes: The feature interaction network is invoked to calculate the attention matrix by using a multi-head self-attention mechanism on the first content and the candidate content; the user features, the context features, and the attention matrix are combined and calculated to obtain the first feature matrix; The step of invoking the similarity network to perform feature interaction on the first content, the candidate content, and the additional features to obtain a second feature matrix representing semantic relevance includes: The similarity network is invoked to calculate the first similarity between the first content and the candidate content at the element level through element-wise product; the similarity network is invoked to calculate the second similarity between the first content and the candidate content at the feature domain level through inner product; the user features, the context features, the first similarity, and the second similarity are combined and calculated to obtain the second feature matrix; The step of invoking the information gain network to perform feature interaction on the first content, the candidate content, and the additional features to obtain a third feature matrix representing information gain includes: The information gain network is invoked to calculate the information gain of the first content and the candidate content under different feature domain types; the user features, the context features, and the information gain are combined and calculated to obtain the third feature matrix.
4. The method according to claim 1, characterized in that, The method of fusing the n feature matrices using the m commentator networks included in the content recommender to obtain a multi-expert, multi-commentator feature output vector includes: A first gate vector is generated based on the first content and the user features; The m critic networks are invoked to generate critic vectors based on the n feature matrices and the first gate vector; Based on the first content, the candidate content, the user features, the context features, and the commentator vector, a feature output vector for multiple experts and multiple commentators is generated.
5. The method according to any one of claims 1 to 4, characterized in that, The content recommender is trained based on the content consumption time (TOI).
6. The method according to claim 1, characterized in that, The information stream displays the first content while also displaying the third content; The invocation of the box trigger in the deep learning model to determine whether to send the second content is based at least on delay costs, including: The box trigger is invoked to generate a second gate vector using the first content, the second content, the user's interaction features when reading the first content, and the third content; The box trigger is invoked to calculate the loss of sending the second content based on the second gate vector, the feature output vector of the second content, the interaction features, and the third content. The loss includes the delay cost and the loss of user interest in the first content. Whether to send the second content is determined based on the loss.
7. The method according to claim 1 or 6, characterized in that, The box trigger is trained based on click-through rate (CTR) loss and the delay cost.
8. A content recommendation method, characterized in that, Applied to a client, the method includes: Displays the first content in the information stream; In response to the triggering operation of the first content, the content interface of the first content is displayed; In response to the return operation of the content interface, a second piece of content is added to the information stream, which is recommended as related content to the first content; Wherein, the second content is determined by the server sorting the candidate content corresponding to the first content according to the feature output vector of multiple experts and multiple commentators; the feature output vector of multiple experts and multiple commentators is obtained by the server calling the content recommender in the deep learning model to fuse n feature matrices by m commentator networks; the n feature matrices are obtained by the server responding to the content interface of the first content displayed by the client, calling the n expert sub-networks included in the content recommender to perform feature interaction on the first content, the candidate content and additional features, where n and m are both integers greater than 1, and the additional features include at least one of user features and context features, where the user features are the features of the user account using the client, and the context features are the features of the operating environment of the client; whether the second content is sent is determined by the server calling the box trigger in the deep learning model at least based on the delay cost, where the delay cost is the impact cost of adding the second content to the information stream.
9. The method according to claim 8, characterized in that, The first content in the displayed information stream includes: Display the first content and the third content in the information stream, wherein the third content is located after the first content; The response to the return operation of the content interface, adding the display of second content in the information stream, includes: In response to a return operation of the content interface, the second content is inserted between the first content and the third content in the information stream; or, the third content in the information stream is replaced and displayed with the second content.
10. A content recommendation device, characterized in that, The device includes: The sending module is used to send an information stream to the client, the information stream including first content; The calling module, in response to the client displaying the content interface of the first content, calls the n expert sub-networks included in the content recommender of the deep learning model to perform feature interaction on the first content, the candidate content corresponding to the first content, and additional features, to obtain n feature matrices; calls the m commentator networks included in the content recommender to fuse the n feature matrices to obtain a multi-expert, multi-commentator feature output vector; sorts the candidate content according to the feature output vector to determine the second content, which is recommended as related content to the first content, where n and m are both integers greater than 1, and the additional features include at least one of user features and context features, where the user features are the features of the user account using the client, and the context features are the features of the operating environment of the client; and calls the box trigger in the deep learning model, at least based on delay cost, to determine whether to send the second content, where the delay cost is the impact cost of adding the second content to the information stream. The sending module is further configured to send the second content to the client, wherein the second content is added to the information stream after a return operation is triggered on the content interface of the first content.
11. The content recommendation device according to claim 10, characterized in that, The n expert sub-networks include at least one of the following: feature interaction network, similarity network, and information gain network; The invocation module is used to invoke the feature interaction network to perform feature interaction on the first content, the candidate content, and the additional features to obtain a first feature matrix representing attention relevance; invoke the similarity network to perform feature interaction on the first content, the candidate content, and the additional features to obtain a second feature matrix representing semantic relevance; and invoke the information gain network to perform feature interaction on the first content, the candidate content, and the additional features to obtain a third feature matrix representing information gain.
12. The content recommendation device according to claim 11, characterized in that, The invocation module is used to invoke the feature interaction network to calculate the attention matrix by using a multi-head self-attention mechanism on the first content and the candidate content; and to combine the user features, the context features, and the attention matrix to obtain the first feature matrix. The similarity network is invoked to calculate the first similarity between the first content and the candidate content at the element level through element-wise product; the similarity network is invoked to calculate the second similarity between the first content and the candidate content at the feature domain level through inner product; the user features, the context features, the first similarity, and the second similarity are combined and calculated to obtain the second feature matrix; The information gain network is invoked to calculate the information gain of the first content and the candidate content under different feature domain types; the user features, the context features, and the information gain are combined and calculated to obtain the third feature matrix.
13. The content recommendation device according to claim 10, characterized in that, The calling module is used to generate a first gate vector using the first content and the user features; call the network of m critics to generate critic vectors based on the n feature matrices and the first gate vector; and generate the feature output vector of multiple experts and multiple critics based on the first content, the candidate content, the user features, the context features, and the critic vectors.
14. The content recommendation device according to claims 10 to 13, characterized in that, The content recommender is trained based on the content consumption time (TOI).
15. The content recommendation device according to claim 10, characterized in that, The information stream displays the first content while also displaying the third content; The invocation module invokes the box trigger to generate a second gate vector using the first content, the second content, the user's interaction features when reading the first content, and the third content; The box trigger is invoked to calculate the loss of sending the second content based on the second gate vector, the feature output vector of the second content, the interaction features, and the third content. The loss includes the delay cost and the loss of user interest in the first content. Whether to send the second content is determined based on the loss.
16. The content recommendation device according to claim 10 or 15, characterized in that, The box trigger is trained based on click-through rate (CTR) loss and the delay cost.
17. A content recommendation device, characterized in that, The device includes: The display module is used to display the first content in the information stream; An interaction module is used to respond to the triggering operation of the first content and display the content interface of the first content; The interaction module is also used to respond to the return operation of the content interface by adding and displaying second content in the information stream, wherein the second content is recommended as related content to the first content; Wherein, the second content is determined by the server sorting the candidate content corresponding to the first content according to the feature output vector of multiple experts and multiple commentators; the feature output vector of multiple experts and multiple commentators is obtained by the server calling the content recommender in the deep learning model to fuse n feature matrices with m commentator networks; the n feature matrices are obtained by the server responding to the content interface of the first content displayed by the client, calling the n expert sub-networks included in the content recommender to perform feature interaction on the first content, the candidate content and additional features, where n and m are both integers greater than 1, and the additional features include at least one of user features and context features, where the user features are the features of the user account using the client, and the context features are the features of the operating environment of the client; whether the second content is sent is determined by the server calling the box trigger in the deep learning model at least based on the delay cost, where the delay cost is the impact cost of adding the second content to the information stream.
18. The content recommendation device according to claim 17, characterized in that, The display module is used to display the first content and the third content in the information stream, wherein the third content is located after the first content; The interaction module is configured to, in response to a return operation of the content interface, insert the second content between the first content and the third content in the information stream; or, replace the third content in the information stream with the second content.
19. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one instruction, at least one program, a code set, or an instruction set, the at least one instruction, the at least one program, the code set, or the instruction set being loaded and executed by the processor to implement the content recommendation method as described in any one of claims 1 to 9.
20. A computer-readable storage medium, characterized in that, The readable storage medium stores at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the content recommendation method as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Data object information providing method and device and electronic device
CN110322305A