Information recommendation method and device, electronic equipment and readable storage medium
By acquiring users' historical behavior and profile information, and combining it with contextual information, an information recommendation model is used to determine the expected benefits of the content to be displayed. This solves the problem of large discrepancies between recommended content and user expectations, and improves recommendation accuracy and benefits.
Patent Information
- Application Number
- CN202411189673.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-27
- Publication Date
- 2026-03-06
AI Technical Summary
In existing technologies, recommended content often deviates significantly from user expectations, resulting in low recommendation accuracy, poor user experience, and low revenue.
By acquiring the target user's historical behavior information, profile information, and contextual information, a pre-trained information recommendation model is used to construct candidate display content. Based on historical behavior, context, and profile information, the expected benefit of each candidate display content is determined, and the content with the highest expected benefit is selected and pushed to the user.
This improved the accuracy of recommended content, making it more likely to meet user expectations and increasing content revenue.
Smart Images

Figure CN121616364A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information recommendation technology, and more specifically, to an information recommendation method, apparatus, electronic device, and readable storage medium. Background Technology
[0002] Currently, with the development of electronic information technology, it is possible to recommend content to users. However, the current methods for determining recommended content may result in recommendations that differ significantly from user expectations, leading to low accuracy, a poor user experience, and low returns for the recommended content. Summary of the Invention
[0003] This application proposes an information recommendation method, apparatus, electronic device, and readable storage medium.
[0004] In a first aspect, embodiments of this application provide an information recommendation method, which, in response to an information recommendation request initiated by a target user, acquires the target user's historical behavior information and the target user's profile information; acquires context information corresponding to the information recommendation request, as well as content to be recommended and advertisements to be recommended that match the context information; constructs different candidate display content based on at least one content information in the candidate content and at least one advertisement information in the candidate advertisements through a pre-trained information recommendation model; determines the expected benefit corresponding to each candidate display content based on the historical behavior information, context information, and profile information through the information recommendation model; and selects the candidate display content with the highest expected benefit as the target display content and pushes the target display content to the target user.
[0005] Secondly, embodiments of this application also provide an information recommendation device, including: a response unit, an acquisition unit, a display content construction unit, an expectation determination unit, and a target display content determination unit. The response unit is used to respond to an information recommendation request initiated by a target user by acquiring the target user's historical behavior information and the target user's profile information; the acquisition unit is used to acquire context information corresponding to the information recommendation request, as well as content to be recommended and advertisements to be recommended that match the context information; the display content construction unit is used to construct different candidate display contents based on at least one content information in the candidate content and at least one advertisement information in the candidate advertisements using a pre-trained information recommendation model; the expectation determination unit is used to determine the expected benefit corresponding to each candidate display content based on the historical behavior information, context information, and profile information using the information recommendation model; the target display content determination unit is used to select the candidate display content with the highest expected benefit as the target display content and push the target display content to the target user.
[0006] Thirdly, embodiments of this application also provide an electronic device, including: one or more processors; a memory; and one or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the one or more processors, and the one or more application programs are configured to perform the method described in the first aspect.
[0007] Fourthly, embodiments of this application also provide a computer-readable storage medium storing program code that can be invoked by a processor to execute the method described in the first aspect above.
[0008] The information recommendation method, apparatus, electronic device, and readable storage medium provided in this application first respond to an information recommendation request initiated by a target user by acquiring the target user's historical behavior information and the target user's profile information; then, it acquires the context information corresponding to the information recommendation request, as well as the content to be recommended and the advertisement to be recommended that match the context information; using a pre-trained information recommendation model, it constructs different candidate display contents based on at least one content information in the candidate content and at least one advertisement information in the candidate advertisement; then, using the information recommendation model, it determines the expected benefit corresponding to each candidate display content based on the historical behavior information, context information, and profile information; thereby, it selects the candidate display content with the highest expected benefit as the target display content and pushes the target display content to the target user. If the target display content to be pushed to the user is determined solely based on the content that the user is interested in, the determined target display content may differ significantly from the user's expectations, that is, the accuracy of the pushed target display content is low, which further results in low benefit corresponding to the target display content. In this application, the expected return for each candidate display content is determined by combining the user's historical behavior information with user profile information and the context information of the information recommendation request. The target display content corresponding to the candidate display content with the highest expected return is then selected and pushed to the target user. This results in higher accuracy of the target display content pushed to the target user, making it more likely to meet user expectations and improving the return on the target display content to a certain extent.
[0009] Other features and advantages of the embodiments of this application will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the embodiments of this application. The objects and other advantages of the embodiments of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 The diagram illustrates an application scenario of the information recommendation method provided in this application embodiment.
[0012] Figure 2 A flowchart of the information recommendation method provided in an embodiment of this application is shown;
[0013] Figure 3 A flowchart of an information recommendation method according to another embodiment of this application is shown;
[0014] Figure 4 This illustration shows a structural diagram of the information recommendation model provided in an embodiment of this application;
[0015] Figure 5 A schematic diagram of the feedback interaction module provided in an embodiment of this application is shown;
[0016] Figure 6 A flowchart of an information recommendation method according to another embodiment of this application is shown;
[0017] Figure 7 This illustration shows a schematic diagram of obtaining training data according to an embodiment of this application;
[0018] Figure 8 A schematic diagram illustrating the status information provided in an embodiment of this application is shown;
[0019] Figure 9 A structural block diagram of the information recommendation system provided in an embodiment of this application is shown;
[0020] Figure 10 This illustration shows a schematic diagram of the online mixed-format service provided in an embodiment of this application;
[0021] Figure 11 A structural block diagram of the information recommendation device provided in an embodiment of this application is shown;
[0022] Figure 12 A structural block diagram of the electronic device provided in an embodiment of this application is shown;
[0023] Figure 13 A structural block diagram of a computer-readable storage medium provided in an embodiment of this application is shown. Detailed Implementation
[0024] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, and not all of them. The components of the embodiments of the present application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without inventive effort are within the scope of protection of the present application.
[0025] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0026] Currently, with the development of electronic information technology, it's possible to recommend content to users. However, current methods for determining recommended content may result in recommendations that deviate significantly from user expectations, leading to low accuracy, a poor user experience, and low revenue. Therefore, how to ensure that pushed content best meets user expectations and improves both accuracy and revenue is a pressing issue that needs to be addressed.
[0027] Currently, it's possible to mine users' historical interests and then push content to them based on those interests. Additionally, it's possible to obtain user characteristics and combine these characteristics to determine the content to be pushed.
[0028] However, the inventors discovered in their research that simply mining users' historical interests to determine push content may deviate significantly from users' expectations, resulting in low accuracy and a poor user experience. Furthermore, the currently available user characteristics are limited, leading to an insufficient comprehensiveness of features considered when determining push content.
[0029] Therefore, in order to solve or partially solve the above problems, this application provides an information recommendation method, apparatus, electronic device, and readable storage medium.
[0030] Please see Figure 1 , Figure 1The illustration shows an application scenario diagram of the information recommendation method provided in the embodiments of this application, namely information recommendation scenario 100. The information recommendation scenario 100 may include a terminal device 110 and a server 120, wherein the terminal device 110 is connected to the server 120.
[0031] Terminal device 110 can connect to server 120, which is also connected to the Internet, by accessing the Internet. Terminal device 110 can access the Internet wirelessly, such as via wireless communication technologies like Wi-Fi or Bluetooth; or it can access the Internet via a wired connection, such as via an RJ45 network cable or fiber optic cable.
[0032] Users can control terminal device 110 to initiate information recommendation requests, such as sending an information recommendation request to server 120. Server 120 can then respond to the information recommendation request, determine the target information, and push it to the user. Detailed descriptions can be found in subsequent embodiments. Server 120 can be a cloud server or a local server.
[0033] Please see Figure 2 , Figure 2 A flowchart illustrating an information recommendation method provided in an embodiment of this application is shown. This information recommendation method can be applied to... Figure 1 In the illustrated information recommendation scenario, the server's processor can be used as the execution entity for the information recommendation method. This information recommendation method may include steps S110 to S150.
[0034] Step S110: In response to the information recommendation request initiated by the target user, obtain the target user's historical behavior information and the target user's profile information.
[0035] A terminal device can run applications, allowing the target user to access various functions corresponding to those applications. For example, the terminal device can run browser applications, online video applications, or online music applications. The target user can be the user initiating the information recommendation request.
[0036] When a target user runs an application on a terminal device, they can initiate an information recommendation request through that application. This information recommendation request can represent the information the target user expects to receive. In some implementations, the application's display interface on the terminal device can show text information related to the information recommendation request, as well as corresponding operation controls. For example, the text information could be "Latest Recommendations." The target user can then initiate the information recommendation request by operating these operation controls. Specifically, if the terminal device has a touchscreen, the target user can initiate the request by touching the operation controls; or, if the terminal device has a keyboard, the target user can operate the keyboard to manipulate the operation controls, thereby initiating the information recommendation request.
[0037] Information recommendation requests initiated by target users through their terminal devices can be sent to the server, for example, by transmitting the information recommendation request through a pre-established communication connection between the terminal device and the server.
[0038] After receiving the information recommendation request, the server can respond to the information recommendation request initiated by the target user and further obtain the target user's historical behavior information and the target user's profile information.
[0039] The target user's historical behavior information may include historical advertisements and the target user's feedback information on historical advertisements, wherein the feedback information may include information representing the target user's level of liking for historical advertisements.
[0040] In some implementations, the target user's historical behavior information may include positive feedback information and negative feedback information. The positive feedback information indicates that the user consumed the displayed historical advertisement, while the negative feedback information indicates that the user did not consume the displayed historical advertisement. Therefore, positive feedback information can also be used to indicate that the user liked the displayed historical advertisement; negative feedback information can also be used to indicate that the user disliked the displayed historical advertisement.
[0041] It is understandable that the historical behavior information may include multiple historical advertisements, and thus the historical behavior information may include positive or negative feedback information corresponding to each historical advertisement.
[0042] In other implementations, historical behavior information may also include historical consultations and corresponding consultation feedback; it may also include historical information stream category preference information, historical application preference tags, and novel-type application preference tags, etc. This improves the accuracy of target information when determining which information to push to target users, making it easier to meet user expectations.
[0043] Profile information can include the target user's characteristics, such as the target user's business level, static characteristics, activity level, or video usage preferences.
[0044] Step S120: Obtain the context information corresponding to the information recommendation request, as well as the content to be recommended and the advertisement to be recommended that match the context information.
[0045] It is understandable that the information recommendation request can determine the target information that needs to be pushed to the target user. Different information recommendation requests can correspond to different context information, so the context information corresponding to the information recommendation request can be obtained first.
[0046] In some implementations, the context information may include the model of the terminal device from which the target user initiated the information recommendation request, the request time, the number of times the request was made, the method of requesting recommendations, the channel for requesting recommendations, and the timestamp at which the information recommendation request was generated.
[0047] For example, after detecting that a target user has initiated the information recommendation request, the terminal device can proactively obtain information such as the model of the terminal device that initiated the information recommendation request, the request time, the number of refreshes, the method of requesting recommendations, the channel of requesting recommendations, and the timestamp of starting the information recommendation request, thereby constructing the context information corresponding to the information recommendation request. Then, the context information can be directly sent to the server, enabling the server to obtain the context information corresponding to the information recommendation request. The terminal device can obtain the context information through methods such as embedded information and listening nodes; this embodiment does not specifically limit the methods used.
[0048] Another example is that the server requests information from the terminal device, such as the model of the terminal device that initiated the information recommendation request, the request time, the number of times the request was made, the method of requesting recommendations, the channel of requesting recommendations, and the timestamp of the information recommendation request being generated. The terminal device then sends the above information to the server, and the server then uses the obtained information to construct the context information corresponding to the information recommendation request.
[0049] After obtaining the context information, further recommended content and recommended advertisements matching the context information can be obtained. The recommended content includes at least one piece of content information, which can be pushed to target users. Target users can then view this content information through their terminal device's display interface and interact with it, such as clicking on the content information to display its corresponding content. The recommended advertisements include at least one piece of advertisement information. Similar to the content information, the advertisement information can also be pushed to target users. Target users can then view this advertisement information through their terminal device's display interface and interact with it, such as clicking on the advertisement information to download content or pay for advertising.
[0050] It is understandable that, based on different contextual information, different content to be recommended and advertisements to be recommended that match the contextual information can be obtained. In some implementations, content recommendation algorithms can be used to obtain content to be recommended that matches the contextual information, and advertisement recommendation algorithms can be used to obtain advertisements to be recommended that match the contextual information.
[0051] Optionally, each piece of content information in the recommended content may also include interface hierarchy, tag information, or title information. The interface hierarchy of the content information can be used to characterize which specific level of the application's display interface the content information belongs to. For example, for an application, the interface displayed when the application starts can be considered the first interface hierarchy. It is understood that the display interface may include interactive controls that allow users to navigate to another display interface, which can then be considered the second interface hierarchy.
[0052] Optionally, the recommended advertisements may also include the expected cost per million (ECPM) for each advertisement, its interface hierarchy, tag information, title information, advertiser identity information, or advertisement type information. Similar to the interface hierarchy of content information, the interface hierarchy of advertisement information can be used to characterize which specific level of the application's display interface the advertisement information belongs to.
[0053] Step S130: Construct different candidate display contents based on at least one content information in the content to be recommended and at least one advertisement information in the advertisement to be recommended using a pre-trained information recommendation model.
[0054] As described above, while we have obtained the content and advertisements to be recommended that match the context information, randomly selecting a portion of the content and a portion of the advertisements to generate the corresponding display content and push it to the target user may result in a low expected return on that display content, making it difficult for the display content to meet the target user's expectations. The expected return on the display content can be used to characterize the target user's expectation of consuming that display content.
[0055] Therefore, in some implementations, different candidate display contents can be constructed first, and then the candidate display content with the highest expected revenue can be found from the candidate display contents. Similarly, the expected revenue of the candidate display content can be used to characterize the target user's expectation of consuming that candidate display content.
[0056] For example, a pre-trained information recommendation model can be used to construct different candidate display contents based on at least one piece of content information in the candidate content and at least one piece of advertising information in the candidate advertisements. Specifically, the candidate content and the candidate advertisements can be used as inputs to the information recommendation model to obtain the constructed different candidate display contents. The information recommendation model can also be used to determine the expected revenue corresponding to each candidate display content; please refer to subsequent steps for a detailed explanation.
[0057] Each candidate display content may include information and the order in which each piece of information is arranged. For example, the candidate display content may include content information A1, content information A2, content information A3, and advertisement information B1, with the order of each piece of information being content information A1, advertisement information B1, content information A2, and content information A3. Alternatively, for the same information in the aforementioned example, the order of each piece of information may also be content information A1, content information A2, advertisement information B1, and content information A3. It is evident that even for the same information, different orderings of each piece of information will result in different candidate display content. It is understood that the descriptions in the above examples are merely illustrative of the embodiments of this application and do not constitute a limitation.
[0058] In some implementations, the information recommendation model may pre-store a set of content ad construction schemes. This set includes schemes for constructing multiple candidate display contents based on the content to be recommended and the ads to be recommended. For example, the content ad construction scheme set may include placing the first content information in the candidate content in a first arrangement order, and placing the first ad information in the candidate ad content in a second arrangement order, to construct a candidate content. The ad construction scheme set can be represented in matrix form.
[0059] Step S140: Using the information recommendation model, based on the historical behavior information, context information, and profile information, determine the expected revenue corresponding to each candidate display content.
[0060] After obtaining multiple candidate display contents, in order to improve the accuracy of the candidate display contents pushed to users and make the pushed candidate display contents more likely to meet the user's expectations, we can first determine the expected benefits corresponding to each candidate display contents. Then, we can select the candidate display contents with the highest expected benefits based on the expected benefits of each candidate display contents for subsequent push.
[0061] In some implementations, the expected revenue for each candidate display content can be determined using the information recommendation model described in the preceding steps. For example, historical behavior information, contextual information, and profile information can be input into the information recommendation model to obtain the expected revenue for each candidate display content output by the model.
[0062] It should be noted that, in some implementations, historical behavior information, context information, profile information, content to be recommended, and advertisements to be recommended can also be input into the information recommendation model, thereby obtaining the candidate display content in step S130 and the expected revenue corresponding to each candidate display content in step S140 through the information recommendation model.
[0063] Step S150: Select the candidate display content with the highest expected benefit as the target display content, and push the target display content to the target user.
[0064] After obtaining the expected revenue for each candidate display content, the expected revenue can be filtered to find the content with the highest expected revenue, which will then be selected as the target display content. Therefore, the candidate display content with the highest expected revenue represents the most accurate display content identified for the target user, and the target user has the highest expectation of consuming this candidate display content. In other words, the target display content is more likely to meet the user's expectations.
[0065] For example, the expected revenue corresponding to each candidate display content can be sorted by size, so that the expected revenue ranked first is taken as the largest expected revenue.
[0066] Furthermore, targeted content can be pushed to target users. This essentially means pushing the targeted content to the terminal device used by the target user that initiated the information recommendation request. Upon receiving the targeted content, the terminal device can then display it on its interface.
[0067] The information recommendation method provided in this application first responds to an information recommendation request initiated by a target user by obtaining the target user's historical behavior information and profile information; then, it obtains the context information corresponding to the information recommendation request, as well as the content to be recommended and the advertisement to be recommended that match the context information; using a pre-trained information recommendation model, it constructs different candidate display contents based on at least one content information in the candidate content and at least one advertisement information in the candidate advertisement; then, using the information recommendation model, it determines the expected benefit corresponding to each candidate display content based on the historical behavior information, context information, and profile information; thereby, it selects the candidate display content with the highest expected benefit as the target display content and pushes the target display content to the target user. If the target display content to be pushed to the user is determined solely based on the content that the user is interested in, the determined target display content may differ significantly from the user's expectations, that is, the accuracy of the pushed target display content is low, which further results in low benefit corresponding to the target display content. In this application, the expected return for each candidate display content is determined by combining the user's historical behavior information with user profile information and the context information of the information recommendation request. The target display content corresponding to the candidate display content with the highest expected return is then selected and pushed to the target user. This results in higher accuracy of the target display content pushed to the target user, making it more likely to meet user expectations and improving the return on the target display content to a certain extent.
[0068] Please see Figure 3 , Figure 3 A flowchart illustrating an information recommendation method provided in an embodiment of this application is shown. This information recommendation method can be applied to... Figure 1 In the illustrated information recommendation scenario, the server's processor can be used as the execution entity for the information recommendation method. This information recommendation method may include steps S210 to S270.
[0069] Step S210: In response to the information recommendation request initiated by the target user, obtain the target user's historical behavior information and the target user's profile information.
[0070] Step S220: Obtain the context information corresponding to the information recommendation request, as well as the content to be recommended and the advertisement to be recommended that match the context information.
[0071] Step S230: Construct different candidate display contents based on at least one content information in the content to be recommended and at least one advertisement information in the advertisement to be recommended using a pre-trained information recommendation model.
[0072] Steps S210 to S230 have been described in detail in the foregoing embodiments and will not be repeated here.
[0073] Step S240: Input the state information as an input parameter into the information recommendation model to obtain the expected state benefit corresponding to the state information, wherein the state information includes the historical behavior information, context information, profile information, content to be recommended, and advertisement to be recommended.
[0074] First, the expected state revenue corresponding to the state information can be determined using an information recommendation model. This state information includes historical behavior information, contextual information, user profile information, content to be recommended, and advertisements to be recommended. It is evident that the expected state revenue is independent of the specific ordering of the content and advertisement information.
[0075] For example, please refer to Figure 8 , Figure 8 A schematic diagram of the status information provided in an embodiment of this application is shown. Figure 8 The status information 800 shown includes profile information 810, historical behavior information 820, context information 830, content to be recommended 840, and advertisements to be recommended 850.
[0076] The profile information 810 may include the target user's business classification, static characteristics, activity level, and video usage preferences. Historical behavior information 820 may include positive feedback information, negative feedback information, historical advertisements, feedback information on historical advertisements, historical inquiries, corresponding consultation feedback information, historical information stream category preference information, historical application preference tags, and novel-type application preference tags. Context information 830 may include the model of the terminal device from which the target user initiated the information recommendation request, the request time, the number of times the request was made, the method of requesting the recommendation, the channel for requesting the recommendation, and the timestamp for generating the information recommendation request. The content to be recommended 840 may include interface hierarchy, tag information, and title information. The advertisements to be recommended 850 may include the expected advertising revenue per thousand impressions for each advertisement, interface hierarchy, tag information, title information, advertiser identity information, and advertisement type information.
[0077] It should be noted that, Figure 8 The status information shown is only an example; in actual applications, the specific information included in the status information can be flexibly adjusted according to the needs.
[0078] In the embodiments provided in this application, the information recommendation model can be a model obtained by adjusting a model using a Dueling-DQN structure.
[0079] In some implementations, state information can be input as an input parameter to the information recommendation model, thereby obtaining the expected state reward corresponding to the state information through the information recommendation model. It should be noted that the expected state reward obtained through the information recommendation model may not be the final output of the model, but rather an intermediate value obtained during its computation. Therefore, optionally, the information recommendation model can be fine-tuned, for example, by adding a fully connected layer as the output layer, to output the intermediate values obtained during the model's computation.
[0080] As described above, historical behavior information can include positive feedback information and negative feedback information. The positive feedback information indicates that the user consumed the displayed historical advertisements, while the negative feedback information indicates that the user did not consume the displayed historical advertisements. The expected reward corresponding to this state information can be obtained by combining the positive and negative feedback information. Specifically, step S240 may further include steps S241 to S244.
[0081] Step S241: Based on the positive feedback information, negative feedback information, and each content information in the content to be recommended, generate a content embedding vector corresponding to each content information, and based on the positive feedback information, negative feedback information, and each advertisement information in the advertisement to be recommended, generate an advertisement embedding vector corresponding to each advertisement information.
[0082] The degree of matching between each piece of content to be recommended and the target user can be determined based on the positive and negative feedback information. The higher the degree of matching, the higher the target user's expectation for that content. Similarly, the degree of matching between each piece of advertising to be recommended and the target user can be determined based on the positive and negative feedback information. The higher the degree of matching, the higher the target user's expectation for that advertising.
[0083] In some implementations, the degree of matching between content information and the target user can be characterized by a content embedding vector, while the degree of matching between advertising information and the target user can be characterized by an advertising embedding vector. Therefore, based on the positive feedback information, negative feedback information, and each piece of content information to be recommended, a content embedding vector corresponding to each piece of content information can be generated first; and based on the positive feedback information, negative feedback information, and each piece of advertising information to be recommended, an advertising embedding vector corresponding to each piece of advertising information can be generated.
[0084] It should be noted that in the embodiments provided in this application, there are multiple negative feedback messages. Specifically, step S241 may include steps S2411 to S2416.
[0085] Step S2411: Obtain the positive feedback embedding vector of the positive feedback information, the negative feedback embedding vector of each negative feedback information, and the candidate embedding vector of the candidate information, wherein the candidate information includes content information and advertising information.
[0086] Step S2412: Based on the positive feedback embedding vector and the candidate embedding vector, construct a third embedding vector corresponding to the candidate information, and based on each of the negative feedback embedding vectors and the candidate embedding vectors, construct multiple fourth embedding vectors corresponding to the candidate information.
[0087] Step S2413: Obtain the similarity between each of the negative feedback embedding vectors and the third embedding vector.
[0088] Step S2414: Determine the target weight of the negative feedback information based on the multiple similarities.
[0089] Step S2415: Based on the target weight, perform a weighted summation on the negative feedback embedding vector to obtain the fifth embedding vector corresponding to the candidate information.
[0090] Step S2416: Concatenate the third embedding vector, the fifth embedding vector, and the fourth embedding vector in sequence to obtain the target candidate embedding vector corresponding to the candidate information. The candidate embedding vector corresponding to the target candidate information includes the content embedding vector corresponding to the content information or the advertising embedding vector corresponding to the advertising information.
[0091] First, the positive feedback embedding vector of the positive feedback information, the negative feedback embedding vector of each negative feedback information, and the candidate embedding vector of the candidate information can be obtained, wherein the candidate information includes content information and advertising information.
[0092] Please see Figure 4 , Figure 4 A schematic diagram of the information recommendation model provided in an embodiment of this application is shown. Figure 4 The information recommendation model 400 includes a feature representation module 410 (Item Representation Module, IRM). Figure 4 As can be seen, the feature representation module 410 includes an embedding layer 411, which can input positive feedback information, each negative feedback information, and alternative information as input values. The alternative information includes content information and advertising information. Figure 4 In this context, oi1……oiNoi represents each piece of content information; ad1……adNad represents each piece of advertising information.
[0093] Thus, the positive feedback information, each negative feedback information, and the alternative information are embedded and encoded through the embedding layer 411 to obtain the positive feedback embedding vector of the positive feedback information, the negative feedback embedding vector of each negative feedback information, and the alternative embedding vector of the alternative information.
[0094] It is understandable that positive feedback information can be represented by textual information. Therefore, an embedding layer can be used to convert the textual information into an embedding vector representation, thus obtaining the positive feedback embedding vector corresponding to the positive feedback information. Similarly, negative feedback information can be represented by textual information, so an embedding layer can be used to convert the negative feedback information into an embedding vector representation, thus obtaining the negative feedback embedding vector corresponding to the negative feedback information. Likewise, alternative information can be represented by textual information, so an embedding layer can be used to convert the textual information into an embedding vector representation, thus obtaining the alternative embedding vector corresponding to the alternative information.
[0095] For some implementation methods, it is possible to... To characterize the positive feedback embedding vector; through To characterize the negative feedback eigenvector through e item This is used to characterize the candidate embedding vectors. It should be noted that the positive feedback information shown in the embodiments of this application can also be multiple, thus corresponding to multiple positive feedback embedding vectors.
[0096] Therefore, based on the positive feedback embedding vector and the candidate embedding vector, a third embedding vector corresponding to the candidate information can be constructed, and based on each of the negative feedback embedding vector and the candidate embedding vector, multiple fourth embedding vectors corresponding to the candidate information can be constructed.
[0097] Please continue reading. Figure 4 The feature representation module 410 may also include a feedback interaction module (FIM) 412, so that the output of the embedding layer 411 can be used as the input of the feedback interaction module 412, and the output of the feedback interaction module 412 includes the content embedding vector corresponding to the content information or the advertising embedding vector corresponding to the advertising information.
[0098] For details, please refer to Figure 5 , Figure 5 A schematic diagram of the feedback interaction module provided in an embodiment of this application is shown. Figure 5 The diagram shows a feedback interaction module 412, whose inputs include a positive feedback embedding vector for positive feedback information, a negative feedback embedding vector for each negative feedback information, and a candidate embedding vector for candidate information.
[0099] Please continue reading. Figure 5The feedback interaction module 412 also includes an activation unit 413 and a dot product unit 414, wherein the output of the activation unit 413 serves as the input of the corresponding dot product unit 414. The positive feedback embedding vector, each negative feedback embedding vector, and the candidate embedding vector can be input to the activation unit 413 and the dot product unit 414 respectively, thereby enabling the dot product of the candidate embedding vector with the positive feedback embedding vector and each negative feedback embedding vector.
[0100] Please continue reading. Figure 5 The feedback interaction module 412 also includes a pooling summation unit 415. The dot product between the candidate embedding vector and the positive feedback embedding vector can be output to a pooling summation unit 415 for summation, thereby outputting a third embedding vector corresponding to the candidate information. The dot product between the candidate embedding vector and the negative feedback embedding vector can be output to another pooling summation unit 415 for summation, thereby outputting multiple fourth embedding vectors corresponding to the candidate information.
[0101] For some implementation methods, it is possible to... Characterizing the third embedding vector, through This represents the fourth embedding vector. Where N... p Used to characterize the amount of positive feedback information; N n Used to characterize the amount of negative feedback information.
[0102] Optional, please continue reading Figure 5 The feedback interaction module 412 may further include a difference unit 416. The positive feedback embedding vector, each negative feedback embedding vector, and candidate embedding vectors can be input into the difference unit 416 to calculate the difference between the candidate embedding vector and each of the positive and negative feedback embedding vectors. Furthermore, Figure 5 The feedback interaction module 412 shown may further include a connection layer 417, an activation layer 418, and a linearization layer 419. The output of the interpolation unit 416, along with the positive feedback embedding vector, each negative feedback embedding vector, and candidate embedding vectors, can serve as inputs to the connection layer 417. The output of the connection layer 417 serves as input to the activation layer 418, and the output of the activation layer 418 serves as input to the linearization layer 419. Finally, the activation weight is output through the linearization layer 419. For example, this activation weight can be a self-attention weight; the activation layer 418 can use Prelude or Dice activation functions, etc.
[0103] After obtaining the third and fourth embedding vectors, more valuable information can be extracted from each negative feedback information. For example, information similar to the positive feedback information can be extracted from the negative feedback information as more valuable information.
[0104] Specifically, the similarity between each negative feedback embedding vector and the third embedding vector can be obtained first. Then, the target weight of the negative feedback information is determined based on the multiple similarities. For example, the negative feedback embedding vector with a higher similarity to the third embedding vector corresponds to a larger target weight.
[0105] Therefore, the negative feedback embedding vector can be weighted and summed based on the target weights to obtain the fifth embedding vector corresponding to the candidate information. In some implementations, this can be achieved through... This is used to characterize the fifth embedding vector.
[0106] Please continue reading. Figure 5 The feedback interaction module 412 may also include a common attention mechanism module 420 (VanlillaAttention), which extracts high-value information from each negative feedback information to obtain the fifth embedding vector.
[0107] Furthermore, the third, fifth, and fourth embedding vectors can be concatenated sequentially to obtain the target candidate embedding vector corresponding to the candidate information. In some implementations, this can be achieved through e' item ←MLP(f p ||f cross ||f n ) is used to characterize the target candidate embedding vector.
[0108] The target candidate embedding vectors corresponding to the candidate information include content embedding vectors corresponding to content information and advertising embedding vectors corresponding to advertising information. For example, the target candidate embedding vectors include content embedding vectors corresponding to content information, where the number of content information items is N. oi In this case, the content embedding vector can be represented as The target candidate embedding vectors include the ad embedding vectors corresponding to the ad information, and the number of ad information is N. ad In this case, the ad embedding vector can be represented as
[0109] Please continue reading. Figure 5The feedback interaction module 412 may further include a connection layer 422 and a multilayer perceptron (MLP) layer 423. The third embedding vector, the fourth embedding vector output by the pooling summation unit 415, and the fifth embedding vector output by the ordinary attention mechanism module 420 can be input to the connection layer 422 to complete the concatenation of the third, fifth, and fourth embedding vectors, and the target candidate embedding vector is output through the multilayer perceptron layer 423.
[0110] Step S242: Obtain the context embedding vector corresponding to the context information and the image embedding vector corresponding to the image information.
[0111] It can also obtain the context embedding vector corresponding to the context information and the image embedding vector corresponding to the image information.
[0112] For some implementation methods, please refer to [link / reference]. Figure 4 Context information and image information can be input into the embedding layer 411 to obtain the context embedding vector corresponding to the context information and the image embedding vector corresponding to the image information. For example, this can be achieved through e... c Representing the context embedding vector, through e u Represents the image embedding vector.
[0113] It is understandable that contextual information can be represented by textual information, and therefore, an embedding layer can be used to convert textual information into an embedding vector representation, thereby obtaining the contextual embedding vector corresponding to the contextual information. Similarly, profile information can be represented by textual information, and therefore, an embedding layer can be used to convert profile information into an embedding vector representation, thereby obtaining the profile embedding vector corresponding to the profile information.
[0114] Step S243: Concatenate the content embedding vector corresponding to the content information with the context embedding vector and the image embedding vector to obtain the first embedding vector corresponding to the content information; and concatenate the advertisement embedding vector corresponding to the advertisement information with the context embedding vector and the image embedding vector to obtain the second embedding vector corresponding to the advertisement information.
[0115] The aforementioned steps have yielded the content embedding vector, the advertisement embedding vector, the context embedding vector, and the portrait embedding vector. The content embedding vector corresponding to the content information can then be concatenated with the context embedding vector and the portrait embedding vector to obtain the first embedding vector corresponding to the content information. Conversely, the advertisement embedding vector corresponding to the advertisement information can be concatenated with the context embedding vector and the portrait embedding vector to obtain the second embedding vector corresponding to the advertisement information.
[0116] Please continue reading. Figure 4 The feature representation module 410 may further include a multilayer perceptron layer 424, which can use the content embedding vector and advertisement embedding vector output by the feedback interaction module 412, as well as the context embedding vector and portrait embedding vector output by the embedding layer 411, as inputs to the multilayer perceptron layer 424. The multilayer perceptron layer 424 then concatenates the content embedding vector corresponding to the content information with the context embedding vector and portrait embedding vector to obtain a first embedding vector corresponding to the content information, and concatenates the advertisement embedding vector corresponding to the advertisement information with the context embedding vector and portrait embedding vector to obtain a second embedding vector corresponding to the advertisement information. Thus, it is possible to obtain the first embedding vector corresponding to each piece of content information containing context information, portrait information, positive feedback information, and negative feedback information, and the second embedding vector corresponding to each piece of advertisement information containing context information, portrait information, positive feedback information, and negative feedback information.
[0117] For some implementation methods, it is possible to... Representing the first embedding vector; through Represents the second embedding vector. That is, Figure 4 The output of the multilayer perceptron layer 424.
[0118] Step S244: Generate the expected state revenue based on the first embedding vector corresponding to each of the content information and the second embedding vector corresponding to each of the advertising information.
[0119] Therefore, the expected state return can be generated based on the first embedding vector corresponding to each of the content information and the second embedding vector corresponding to each of the advertisement information. For example, the expected state return can be obtained by performing operations such as concatenation or dimensionality reduction on the first embedding vector. Specifically, step S244 may also include steps S2441 to S2442.
[0120] Step S2441: Sequentially use the first embedding vector corresponding to each of the content information as a row vector to construct the first matrix.
[0121] Step S2442: Sequentially use the second embedding vector corresponding to each of the advertising information as a row vector to construct a second matrix.
[0122] Step S2443: Perform dimensionality reduction on the first matrix and the second matrix respectively.
[0123] Step S2444: Concatenate the first matrix after dimensionality reduction and the second matrix after dimensionality reduction to obtain the expected state return.
[0124] The first embedding vector corresponding to each piece of content information and the second embedding vector corresponding to each piece of advertising information are obtained through the aforementioned steps. Therefore, in some embodiments, the first embedding vector corresponding to each piece of content information can be used as a row vector to construct a first matrix, and the second embedding vector corresponding to each piece of advertising information can be used as a row vector to construct a second matrix.
[0125] For example, it is possible The first matrix is represented by [the matrix name]; This is used to characterize the second matrix.
[0126] Furthermore, dimensionality reduction can be performed on the first matrix and the second matrix separately. For example, the flatten function can be used to perform dimensionality reduction, which can be used to convert a multidimensional array into a one-dimensional array.
[0127] Optionally, before performing dimensionality reduction on the first and second matrices, pruning operations can be performed on them first. This can accelerate the inference speed of the information recommendation model to some extent and improve the efficiency of determining the target display content. For example, the pruning operation on the first and second matrices can be performed using the pool function.
[0128] Then, the first matrix after dimensionality reduction and the second matrix after dimensionality reduction are concatenated to obtain the expected state return. In some implementations, this can be achieved using V(s) = MLP(flatten(pool(E)). ad )||pool(E oi The expected return of a state is represented by ))).
[0129] Please continue reading. Figure 4 The information recommendation model 400 may further include a V network 430, and the output of the feature representation module 410 can be used as the input of the V network 430, thereby obtaining the expected state return through the V network 430. The output of the feature representation module 410 may include a first embedding vector corresponding to each piece of content information and a second embedding vector corresponding to each piece of advertising information.
[0130] Step S250: Obtain the expected deviation gain resulting from each of the candidate display contents, wherein the expected deviation gain is used to characterize the deviation relative to the expected state gain.
[0131] Furthermore, the expected deviation gain generated by adopting each of the candidate display contents can be obtained, wherein the expected deviation gain is used to characterize the deviation from the expected state gain. Subsequently, the expected state gain can be corrected using the expected deviation gain, thereby determining the expected gain corresponding to each of the candidate display contents. This ensures that the expected gains corresponding to the candidate display contents have high accuracy, thereby improving the accuracy of subsequently determining the target display content.
[0132] In some implementations, features between the candidate display content and the image information and context information can be extracted by interacting with the first matrix and the second matrix respectively, thereby improving the accuracy of the expected benefit of the deviation generated in the subsequent determination of the candidate display content. Specifically, step S250 may also include steps S251 to S255.
[0133] Step S251: Determine the content matrix and the advertising matrix that match each candidate display content, wherein the content matrix is used to represent the content information to be displayed, and the advertising matrix is used to represent the advertising information to be displayed.
[0134] As described above, each candidate display content can include information and the order in which the information is arranged. Therefore, the content information and advertising information to be displayed can be represented in matrix form. Specifically, a content matrix and an advertising matrix are determined for each type of candidate display content. Thus, the content matrix represents the content information to be displayed, and the advertising matrix represents the advertising information to be displayed.
[0135] For example, the content matrix may include multiple elements, each with a first or second value. Each element corresponds to a piece of content information, and the value corresponding to that element indicates whether or not the content information should be displayed. For instance, if the element is the first value, it indicates that the content information corresponding to that element should not be displayed, while if the element is the second value, it indicates that the content information corresponding to that element should be displayed. For example, the first value could be 0, and the second value could be 1.
[0136] Similarly, an ad matrix can include multiple elements, each with a first or second numerical value. Each element corresponds to an ad message, and the numerical value of that element indicates whether or not the corresponding ad message should be displayed. For example, if an element has a first numerical value, it indicates that the ad message corresponding to that element should not be displayed, while if an element has a second numerical value, it indicates that the ad message corresponding to that element should be displayed. For example, the first numerical value could be 0, and the second numerical value could be 1.
[0137] For some implementation methods, M can be used. oi The content matrix is represented by M.ad Characterize the advertising matrix.
[0138] Please continue reading. Figure 4 , Figure 4 In this context, multiple candidate display contents are represented by a1...aNa.
[0139] Step S252: The product of the content matrix corresponding to the content to be displayed and the first matrix is used as the first intermediate matrix, and the product of the advertisement matrix corresponding to the content to be displayed and the second matrix is used as the second intermediate matrix.
[0140] Step S253: Summate the first intermediate matrix and the second intermediate matrix to obtain the cross matrix matching the candidate display content.
[0141] Furthermore, the product of the content matrix corresponding to the content to be displayed and the first matrix can be used as the first intermediate matrix, which can be obtained through M. oi E oi The first intermediate matrix is represented by the product of the ad matrix corresponding to the content to be displayed and the second matrix. This product is used as the second intermediate matrix, which can be represented by M. ad E ad Characterizes the second intermediate matrix.
[0142] However, summing the first intermediate matrix and the second intermediate matrix yields the cross matrix matching the candidate display content. That is, through M... cross =M ad E ad +M oi E oi Characterizes the cross matrix.
[0143] Please continue reading. Figure 4 , Figure 4 The information recommendation model 400 may further include a State and Action Crossing Unit (SACU) 440, wherein there can be multiple SACUs, for example, the same number as the number of content items to be displayed. Thus, the information recommendation model 400 can... Figure 4 The first and second matrices, constructed from the output of the multilayer perceptron layer 424, serve as inputs to each action interaction module 440, and each candidate display content also serves as input to its corresponding action interaction module 440. Thus, the action interaction module 440 obtains a cross-matrix matching each candidate display content.
[0144] Step S254: Adjust the weights of each element in the cross matrix using a self-attention mechanism to obtain the adjusted third matrix.
[0145] After obtaining the cross matrix, a self-attention mechanism can be used to adjust the weights of each element in the cross matrix, thus obtaining the third matrix. Because each element in the third matrix is obtained by adjusting the cross matrix using the self-attention mechanism, a more accurate expected return on the bias can be obtained subsequently through the third matrix.
[0146] In some implementations, SelfAtt can be used. (i) The function implementation uses a self-attention mechanism to adjust the weights of each element in the cross matrix, thereby obtaining a third matrix.
[0147] Step S255: Perform dimensionality reduction on the third matrix to obtain the sixth embedding vector matching the candidate display content, which serves as the expected gain from the deviation generated by the candidate display content.
[0148] Furthermore, the third matrix can be dimensionality reduced to obtain the sixth embedding vector matching the candidate display content, which serves as the expected gain from the deviation generated by the candidate display content.
[0149] Similarly, the dimensionality reduction of the third matrix can be performed using the flatten function to obtain the sixth embedding vector that matches the candidate display content, which serves as the expected gain from the deviation generated by the candidate display content.
[0150] It is understandable that different target users may have different focuses on features. For example, some target users may focus on the channels they request recommendations from, while others may focus on tag information. Therefore, in some implementations, a masking operation can be performed on the cross matrix using a masking matrix to simulate situations where different target users do not pay attention to some feature information. Specifically, step S254 may also include steps S2541 to S2543, and step S255 may also include step S2544.
[0151] Step S2541: Randomly generate a specified number of masking matrices;
[0152] Step S2542: Multiply the cross matrix with each of the masking matrices to obtain a specified number of masked matrices that match the content to be displayed;
[0153] Step S2543: Adjust the weights of each element in each masked matrix that matches the selected display content using a self-attention mechanism to obtain the fourth matrix after weight adjustment for each masked matrix.
[0154] In some implementations, the randomly generated masking matrix can be a 0-1 matrix, meaning it contains only elements 0 or 1. The specified quantity can be represented as N. cThe specified quantity is a hyperparameter that can be preset.
[0155] Furthermore, the cross matrix can be multiplied with each of the masking matrices to obtain a specified number of masked matrices that match the candidate display content. Then, a self-attention mechanism is used to adjust the weights of each element in each masked matrix that matches the candidate display content to obtain a fourth matrix corresponding to each masked matrix after weight adjustment.
[0156] Among them, it can be achieved through To represent a specified number of masking matrices, thus, a specified number of masked matrices can be represented as
[0157] Please continue reading. Figure 4 , Figure 4 The information recommendation model 400 may further include a multi-channel attention unit 450, which includes a mask channel 451 and self-attention units 452. The number of mask channels 451 and self-attention units 452 may both be the same as a specified number. Each mask channel 451 corresponds to a different randomly generated mask matrix.
[0158] Therefore, the cross matrix output by the motion interaction module 440 is sent to each masking channel 451, so that the cross matrix is multiplied by the masking matrix corresponding to that masking channel 451. Furthermore, the output of a specified number of masking channels 451 is the specified number of masked matrices that match the candidate display content.
[0159] Furthermore, the output of each masking channel 451 is used as the input of the corresponding self-attention unit 452, so as to adjust the weight of each element in each masked matrix that matches the selected display content through the self-attention mechanism, and obtain the fourth matrix after weight adjustment for each masked matrix.
[0160] For some implementation methods, it is possible to... To represent the masked matrix.
[0161] Step S2544: Perform dimensionality reduction on each of the fourth matrices and concatenate the dimensionality-reduced fourth matrices to obtain the seventh embedding vector matching the candidate display content, which serves as the expected deviation benefit generated by the candidate display content.
[0162] Therefore, each of the fourth matrices is subjected to dimensionality reduction, and the dimensionality-reduced fourth matrices are concatenated to obtain the seventh embedding vector matching the candidate display content, which serves as the expected deviation benefit generated by the candidate display content.
[0163] For example, the self-attention unit 452 described above can also perform dimensionality reduction processing on each of the acquired fourth matrices to obtain the dimensionality-reduced fourth matrices.
[0164] For some implementations, the fourth matrix after dimensionality reduction can be characterized as follows:
[0165]
[0166] Please continue reading. Figure 4 The multi-attention module 450 also includes a connection layer 453. The fourth matrices corresponding to the candidate display content can be used as inputs to the connection layer 453, so that the fourth matrices are concatenated by the connection layer 453 to obtain the seventh embedding vector matching the candidate display content.
[0167] For some implementation methods, it is possible to... This is used to characterize the seventh embedding vector of the candidate display content matching. The seventh embedding vector of the candidate display content matching can be used as the expected gain of the deviation generated by the candidate display content.
[0168] It is understood that each connection layer 453 outputs the expected benefit corresponding to the candidate display content for that connection layer 453, and there are multiple candidate display contents in this embodiment. Therefore, please continue to refer to... Figure 4 The information recommendation model 400 may also include an A network 460, where the output of each connection layer 453 can be used as the input of the A network 460, thereby allowing the A network 460 to determine the expected return of each candidate display content.
[0169] For example, it can be done through The expected return for each candidate display content is represented by the deviation.
[0170] Step S260: Based on the expected revenue of the state and the expected revenue of the deviation, determine the expected revenue corresponding to each of the candidate display contents.
[0171] The expected revenue of the state and the expected revenue of the deviation of the candidate state display content have been obtained through the aforementioned steps, and the expected revenue corresponding to each candidate display content can be determined.
[0172] In some implementations, step S260 may include steps S261 to S263.
[0173] Step S261: Obtain the average value of the expected return for each of the candidate display contents as the average expected return.
[0174] Step S262: The difference between the deviation expected benefit and the average expected benefit of each candidate display content is taken as the intermediate expected benefit corresponding to that candidate display content.
[0175] Step S263: Obtain the sum of the expected revenue of the state and the intermediate expected revenue corresponding to the candidate display content, and use it as the expected revenue corresponding to the candidate display content.
[0176] First, the average of the expected returns based on the deviations of the candidate display content can be obtained as the average expected return. For example, this can be achieved through... It represents the average expected return.
[0177] Furthermore, the difference between the expected deviation return of each of the candidate display contents and the average expected return can be used as the intermediate expected return corresponding to that candidate display content. For example, this can be achieved through... This represents the intermediate expected revenue corresponding to the candidate display content.
[0178] However, the sum of the expected return of the current state and the intermediate expected return corresponding to the candidate display content can be obtained as the expected return corresponding to that candidate display content. For example, this can be achieved through... This represents the expected revenue corresponding to the selected display content.
[0179] Please continue reading. Figure 4 , Figure 4 The information recommendation model 400 also includes an expected revenue unit 470. The outputs of the V network and the A network are both input into the expected revenue unit 470, thereby the expected revenue unit 470 executes the above steps to obtain the expected revenue corresponding to each candidate display content. The expected revenue corresponding to each candidate display content can be represented as Q1…QNa.
[0180] In some implementations, the expected revenue can be determined by combining the revenue that the candidate display content can generate and the maximum revenue that the next state can generate given that the candidate display content is currently selected.
[0181] Therefore, it is possible to This is used to model and calculate the expected revenue corresponding to the candidate display content. This modeling is then used to guide, train, and fine-tune the information recommendation model. In some implementations, the determination of this expected revenue can be viewed as a unified modeling approach based on the expectations across the entire scenario. Where r(s) t ,a t This represents the revenue that the selected display content can generate. This represents the maximum potential benefit that the next state can generate given the current selection of the candidate content. γ is a long-term factor, which can be regarded as a weighted parameter of future value. When γ = 0, it means that only the potential benefit generated by the candidate content is considered, that is, the impact of the current selection of the candidate content on the subsequent events is not considered.
[0182] The r function can be used to characterize the immediate reward corresponding to the target user's feedback. In some implementations, the r function can be calculated using reward, specifically represented as reward = α·r ad +β1·r oi +β2·r ex .
[0183] Where, r ad This represents the actual revenue deducted from advertising, which is the cost that advertisers actually spend on advertising.
[0184] r oi Content exposure score is used to represent various metrics. For example, if the requested channel corresponds to a list page, the content exposure score can be the ratio of the number of clicks on the content to the number of times the content is displayed. The number of clicks refers to the number of times a target user clicks on the target content after receiving it in a push notification. Alternatively, if the requested channel corresponds to a video scenario, the average video playback time can be used as the content exposure score.
[0185] In addition, it is necessary to consider the competition for consumption resources at different interface levels. For example, if the interface level is level one, then the immediate revenue needs to be increased by the actual deduction revenue from advertising after clicking from the first interface level to the second interface level.
[0186] r ex This represents the user's experience gain. If the target user no longer browses the currently displayed interface after issuing an information recommendation request at the current moment, it means the user has left, and r... ex Record it as -1, otherwise r ex Record it as 0.
[0187] Therefore, the calculation of expected revenue in this embodiment of the application fully considers the actual advertising revenue, the exposure score of the content, the experience benefits of the target user, and the influence between the multi-level interface layers, and realizes a comprehensive consideration of expected revenue. Thus, the maximum value of expected revenue is obtained by combining various features, which has high accuracy.
[0188] Understandably, designing the form and r-function alone may not guarantee that the information recommendation model trained by reinforcement learning will meet practical needs. Therefore, when modeling and calculating expected returns, inverse reinforcement learning can be used to model a personalized reward function. By customizing the optimal reward function through inverse reinforcement learning, it is possible to find the best reward function and use it to infer expected returns, thereby improving the inference performance of the resulting information recommendation model.
[0189] Step S270: Select the candidate display content with the highest expected benefit as the target display content, and push the target display content to the target user.
[0190] Step S270 has been described in detail in the foregoing embodiments and will not be repeated here.
[0191] The information recommendation method provided in this application deeply mines the positive and negative feedback information of target users, enabling better mining of complex target user characteristics in information flow scenarios, thereby improving the accuracy of subsequent determination of target display content. Furthermore, this application embodiment constructs a relatively comprehensive information flow scenario feature framework using target user profile information, historical behavior information, contextual information, candidate recommendation advertisements, and candidate recommendation content. This framework is then used to further determine the target display content, making it more aligned with the target user's expectations. Moreover, by correcting the state expected return through deviation expected return, the expected return corresponding to each candidate display content is determined, ensuring high accuracy of the obtained expected returns for the candidate display content, thus improving the accuracy of subsequent determination of target display content. Additionally, this application embodiment can first perform pruning operations on the first and second matrices, thereby accelerating the inference speed of the information recommendation model to a certain extent and improving the efficiency of determining target display content.
[0192] Please see Figure 6 , Figure 6 A flowchart illustrating an information recommendation method provided in an embodiment of this application is shown. This information recommendation method can be applied to... Figure 1 In the illustrated information recommendation scenario, the server's processor can be used as the execution entity for the information recommendation method. This information recommendation method may include steps S310 to S370.
[0193] Step S310: In response to the information recommendation request initiated by the target user, obtain the target user's historical behavior information and the target user's profile information.
[0194] Step S310 has been described in detail in the foregoing embodiments and will not be repeated here.
[0195] It is understandable that target users generally exhibit clear interest in their search activities. Therefore, in some implementations, when a target user performs a search operation through an application on their terminal device, a recommendation request can be detected and sent to the server. The search information entered by the user can supplement the user profile, further improving the accuracy of determining the target content to display and thus increasing revenue.
[0196] Step S320: Obtain the context information corresponding to the information recommendation request.
[0197] Step S330: Input the context information into the pre-acquired content recommendation algorithm to obtain the content to be recommended.
[0198] Step S340: Input the context information into the pre-acquired advertising recommendation algorithm to obtain the advertisement to be recommended.
[0199] As described above, different recommended content and advertisements can be obtained based on different contextual information. Therefore, in some implementations, content recommendation algorithms and advertisement recommendation algorithms can be obtained in advance. The content recommendation algorithm can generate corresponding recommended content based on the input parameters; while the advertisement recommendation algorithm can generate corresponding recommended advertisements based on the input parameters.
[0200] Therefore, after obtaining the context information corresponding to the information recommendation request, the context information can be input into the content recommendation algorithm, allowing the algorithm to determine the content to be recommended based on the context information. Similarly, the context information can also be input into the advertising recommendation algorithm, allowing the algorithm to determine the advertising content to be recommended based on the context information.
[0201] Content recommendation algorithms and advertising recommendation algorithms can be obtained through the internet.
[0202] Step S350: Construct different candidate display contents based on at least one content information in the content to be recommended and at least one advertisement information in the advertisement to be recommended using a pre-trained information recommendation model.
[0203] Step S360: Using the information recommendation model, based on the historical behavior information, context information, and profile information, determine the expected revenue corresponding to each candidate display content.
[0204] Steps S350 and S360 have been described in detail in the foregoing embodiments and will not be repeated here.
[0205] In some implementations, a training dataset can be obtained first, and then the initial model can be trained using the training dataset to obtain the trained initial model as the information recommendation model.
[0206] The training dataset can include multiple training data sets. Each training data set can include training state information at a specified time, immediate reward at a specified time, and training state information for the next time step. For example, training data can be constructed as Data = {(state, action, reward, state_next)}, where state represents the training state information at a specified time; action represents the immediate reward at a specified time; and reward and state_next represent the training state information for the next time step. The training state information can include historical behavior information, contextual information, user profile information, content to be recommended, and advertisements to be recommended.
[0207] For example, the initial model can be trained offline using a temporal difference (TD) scheme, where the initial model can be based on... Figure 4 The structure of the information recommendation model shown is used to build the model. Specifically, for example, an initial model can be used to infer a loss function from the training data, and the loss function can be used to adjust parameters such as weights to train the initial model.
[0208] In some implementations, the loss function can be constructed in two parts: the first part is the loss function of a Deep Q-Network (DQN) based on the Bellman equation, and the second part is the PAE exposure constraint loss function. The final loss function is then obtained by weighted summation of the first and second parts. The first part defines the training objective of the neural network, namely, minimizing the estimation error of the Q-value.
[0209] For example, the final loss function can be represented as L(B) = L DQN (B)+αL PAE (B), where L DQN (B) The loss function used to characterize deep Q-networks based on the Bellman equation, L PAE (B) is used to characterize the PAE exposure constraint loss function.
[0210] In some implementations, the loss function of a deep Q-network based on the Bellman equation can be characterized as Here, (s,a) represents the specified time, while (s',a') represents the next time after the specified time.
[0211] Furthermore, during the initial model training process, the average percentage of ad impressions within the content to be displayed corresponding to the maximum expected revenue obtained through inference from the initial model within a training batch can be kept close to a preset value δ. Therefore, the PAE exposure constraint loss function can be characterized as follows: It is understandable that content information or advertising information that appears earlier in the order of the content to be displayed has a higher exposure probability. Therefore, each order is assigned a corresponding exposure weight, and then the PAE exposure constraint loss function is calculated by weighting.
[0212] Optionally, two identical networks, a prediction network and a target network, can be constructed. The loss function can then be calculated using the target network, thus transforming the approximate Bellman optimality operator problem into a regression problem. Subsequently, the prediction network parameters can be synchronized to the target network periodically to increase its stability. Finally, only the prediction network is retained as the information recommendation model.
[0213] In some implementation methods, advertising data, content data, and user behavior tracking data can be obtained as initial data. Then, the obtained initial data is processed to construct training data.
[0214] Please see Figure 7 , Figure 7 A schematic diagram illustrating the acquisition of training data provided in an embodiment of this application is shown. Figure 7 Initial data can be obtained through mining, for example, application server-side logs, user behavior logs, user profile information, and user behavior tracking data. Specifically, these data can be input into processing module 710 for processing. For instance, processing module 710 can be used for Spark data processing.
[0215] Furthermore, the data output by the processing module 710 can be input to the storage module 720 for storage. For example, the storage module 720 can be the Hadoop Distributed File System (HDFS).
[0216] Then, the data stored in storage module 720 can be input into offline analysis module 730, where the data can be parsed, such as for log parsing. Offline analysis module 730 can also work with database 740 for data parsing; specifically, offline analysis module 730 can query database 740. Database 740 can be a remote dictionary server (Redis).
[0217] Furthermore, after parsing the data, the offline analysis module 730 can input the parsed data into the data cleaning module 750, whereby the data cleaning module 750 cleans the parsed data. Cleaning may include filtering operations such as dirty data filtering, illegal data filtering, or valueless data filtering, and then the cleaned data is stored in the storage module 760.
[0218] For example, the storage module 760 can be a module that stores data in the TFRecord data storage format, thereby improving the efficiency of subsequent data preprocessing. After obtaining the cleaned data, the aforementioned multiple training data can be constructed based on the cleaned data to obtain the training dataset.
[0219] Step S370: Find the candidate display content that meets the pre-acquired strategy requirements and has the greatest expected benefit, select it as the target display content, and push the target display content to the target user.
[0220] In some implementation methods, after obtaining the expected revenue for each candidate display content, the candidate display content can be filtered based on the pre-obtained strategy requirements.
[0221] Strategy requirements can include multiple factors. For example, they may include filtering rules for the expected advertising revenue per thousand impressions (EPCM), pre-processing rules for advertisements, ad interval rules, ad quantity rules, rules to prevent advertisements from appearing at the end of the ad session, rules to prevent advertisements from appearing in the first N positions, special device strategies, or policy rules.
[0222] Therefore, each candidate display content can be filtered according to the strategy requirements, so that the filtered candidate display content meets every item in the strategy requirements.
[0223] Furthermore, among the expected benefits of the filtered candidate display content, the one with the highest expected benefit is selected as the target display content.
[0224] The target content can then be pushed to the target user.
[0225] The information recommendation method provided in this application does not determine the candidate display content with the highest expected benefit for each candidate display content. Instead, it first filters each candidate display content based on strategy requirements, ensuring that all filtered candidate display content meets every requirement of the strategy. Then, it searches for the candidate display content with the highest expected benefit among the filtered candidate display content, thus determining the candidate display content with the highest expected benefit as the target display content. This allows the target display content pushed to the target user to be constrained by strategy requirements.
[0226] Please see Figure 9 , Figure 9 A structural block diagram of the information recommendation system provided in an embodiment of this application is shown. The information push system 900 includes an online mixed-ranking module 910, a model training module 920, a data processing module 930, and a terminal side 940. The mixed-ranking model 910 is connected to both the model training module 920 and the terminal side 940, and the data processing module 930 is also connected to both the model training module 920 and the terminal side 940.
[0227] The data processing module 930 includes a processing component 931, a storage component 932, an offline analysis component 933, a database 934, and a data cleaning component 935. The storage component 932 is connected to both the processing component 931 and the offline analysis component 933, and the offline analysis component 933 is also connected to both the database 934 and the data cleaning component 935.
[0228] The terminal side 940 includes the target user's terminal device 941. Through the terminal device 941, server-side logs from the application, user behavior logs, user profile information, and embedded user behavior data can be collected. These data are then sent to the data processing component 931. The offline analysis component 935, in conjunction with the database 934, parses the data. The parsed data is then input into the data cleaning component 935 for data cleaning, and further training data is constructed. For a detailed description, please refer to the aforementioned... Figure 7 The relevant details will not be repeated here.
[0229] Optionally, the terminal device 941 can also be used to initiate an information recommendation request, so that the offline analysis component 933 can also be used to respond to the information recommendation request and analyze the status information based on the data provided by the terminal device 941.
[0230] The model training module 920 includes an initial model training component 921 and an offline experimental verification component 922, wherein the training component 921 and the experimental verification component 922 are connected. The initial model training component 921 can acquire training data constructed by the data processing module 930 and train the initial model using this training data. Then, the experimental verification component 922 verifies the trained initial model. In some implementations, after successful verification, the verified initial model can be used as an information recommendation model.
[0231] The online mixed-sorting module 910 includes a mixed-sorting scheduling component 911, a database 912, an operator service component 913, and a model deployment component 914. The mixed-sorting scheduling component 911 is connected to the database 912 and the operator service component 913, and the operator service component 913 is also connected to the model deployment component 914.
[0232] The information recommendation request initiated by terminal device 941 can be transmitted to the mixed ranking scheduling component 911. Mixed ranking scheduling component 911 can run the mixed ranking scheduling service mixRanker, and then communicate with database 912 in response to the information recommendation request to determine the strategy requirements matching the request. The obtained strategy requirements can then be transmitted to operator service component 913. Database 913 can store content understanding tags and advertising understanding tags. Additionally, operator service component 913 can also obtain status information matching the information recommendation request sent by offline analysis component 935. Then, operator service component 913 can request model deployment component 914 to perform inference through the information recommendation model to obtain the expected revenue corresponding to each candidate display content. Operator service component 913 then performs data post-processing on the expected revenue corresponding to each candidate display content. Specifically, operator service component 913 can filter each candidate display content according to the strategy requirements, ensuring that the filtered candidate display content meets every item in the strategy requirements. Then, among the expected benefits of the filtered candidate display content, the content with the highest expected benefit is selected as the target display content. After obtaining the target display content, it can be sent to the mixed scheduling component 911, which then pushes it to the terminal device 941. For a detailed introduction to the policy requirements, please refer to the preceding description; it will not be repeated here.
[0233] The operator service component 913 can run uniform operator services, which can be written in languages such as Python. The model deployment component 914 can use TensorFlow Serving to deploy the recommended model based on the information received from the initial model training component 921, for example, by CPU deployment.
[0234] Optionally, when responding to an information recommendation request and determining the strategy requirements matching the request, the mixed scheduling component 911 can also choose whether to use an information recommendation model or a virtual pricing scheme to determine the target content to be pushed, in order to conduct experimental comparisons and determine the effectiveness of using an information recommendation model to determine the target content. For example, when receiving an information recommendation request, the mixed scheduling component 911 can also obtain an experiment ID matching the request, and then use the experiment ID to determine whether to use an information recommendation model or a virtual pricing scheme.
[0235] For specific examples, please refer to Figure 10 , Figure 10 A schematic diagram of the online mixed-format service provided in an embodiment of this application is shown. Figure 10 The mixed-display scheduling component 911 shown can determine the target display content based on the experiment ID, either through a virtual pricing scheme or an information recommendation model. Specifically, determining the target display content through an information recommendation model can also be referred to as a model scheme. Furthermore, if it is determined that the target display content will be determined through an information recommendation model, the strategy requirement can be further transmitted to the operator service component 913. The operator service component 913 then requests the model deployment component 914 to perform inference through the information recommendation model to obtain the expected revenue corresponding to each candidate display content, and further determine the target display content based on the strategy requirement. For details, please refer to the description of the aforementioned implementation.
[0236] Based on the foregoing embodiments and experimental verification, the following analyses were conducted: adding only a feedback interaction module to the information push model; adding a feedback interaction module and combining it with a unified modeling approach based on full-scenario expectations; and adding a feedback interaction module and combining it with a unified modeling approach based on full-scenario expectations, further analyzing the target user's historical behavior information. The experimental data shown in Table 1 are as follows.
[0237] Table 1 Experimental Data
[0238]
[0239]
[0240] In Table 1, model-1 corresponds to adding only the feedback interaction module to the information push model; model-2 corresponds to adding the feedback interaction module and combining it with a unified modeling approach based on full-scenario expectations; and model-3 corresponds to adding the historical behavior information of the target user in addition to adding the feedback interaction module and combining it with a unified modeling approach based on full-scenario expectations.
[0241] As shown in Table 1, model-3 has the best average revenue per user (ARPU), average content clicks per user, and average information feed duration.
[0242] Please see Figure 11 , Figure 11 The diagram shows a structural block diagram of an information recommendation device provided in an embodiment of this application, which is applied to an electronic device. The information recommendation device 1100 includes: a response unit 1110, an acquisition unit 1120, a display content construction unit 1130, an expectation determination unit 1140, and a target display content determination unit 1150.
[0243] The response unit 1110 is used to respond to an information recommendation request initiated by a target user and obtain the target user's historical behavior information and the target user's profile information.
[0244] The acquisition unit 1120 is used to acquire the context information corresponding to the information recommendation request, as well as the content to be recommended and the advertisement to be recommended that match the context information.
[0245] Optionally, the acquisition unit 1120 can also be used to acquire context information corresponding to the information recommendation request; input the context information into a pre-acquired content recommendation algorithm to obtain the content to be recommended; and input the context information into a pre-acquired advertising recommendation algorithm to obtain the advertisement to be recommended.
[0246] The content construction unit 1130 is used to construct different candidate display contents based on at least one content information in the content to be recommended and at least one advertisement information in the advertisement to be recommended, using a pre-trained information recommendation model.
[0247] The expectation determination unit 1140 is used to determine the expected revenue corresponding to each candidate display content based on the historical behavior information, context information and profile information through the information recommendation model.
[0248] Optionally, the expectation determination unit 1140 can also be used to input state information as an input parameter into the information recommendation model to obtain the expected state benefit corresponding to the state information, wherein the state information includes the historical behavior information, context information, profile information, content to be recommended, and advertisement to be recommended; obtain the deviation expected benefit generated by each of the candidate display contents, wherein the deviation expected benefit is used to characterize the deviation from the state expected benefit; and determine the expected benefit corresponding to each of the candidate display contents based on the state expected benefit and the deviation expected benefit.
[0249] Optionally, the expectation determination unit 1140 can also be used to generate a content embedding vector corresponding to each content information based on the positive feedback information, negative feedback information, and each content information in the content to be recommended, and to generate an ad embedding vector corresponding to each ad information based on the positive feedback information, negative feedback information, and each ad information in the ad to be recommended; obtain the context embedding vector corresponding to the context information and the portrait embedding vector corresponding to the portrait information; concatenate the content embedding vector corresponding to the content information with the context embedding vector and the portrait embedding vector to obtain a first embedding vector corresponding to the content information, and concatenate the ad embedding vector corresponding to the ad information with the context embedding vector and the portrait embedding vector to obtain a second embedding vector corresponding to the ad information; and generate the expected state benefit based on the first embedding vector corresponding to each content information and the second embedding vector corresponding to each ad information.
[0250] Optionally, the expectation determination unit 1140 can also be used to construct a first matrix by sequentially using the first embedding vector corresponding to each of the content information as a row vector; construct a second matrix by sequentially using the second embedding vector corresponding to each of the advertising information as a row vector; perform dimensionality reduction processing on the first matrix and the second matrix respectively; and concatenate the first matrix after dimensionality reduction processing and the second matrix after dimensionality reduction processing as the expected state benefit.
[0251] Optionally, the expectation determination unit 1140 can also be used to obtain the positive feedback embedding vector of the positive feedback information, the negative feedback embedding vector of each negative feedback information, and the candidate embedding vector of the candidate information, wherein the candidate information includes content information and advertising information; construct a third embedding vector corresponding to the candidate information based on the positive feedback embedding vector and the candidate embedding vector, and construct multiple fourth embedding vectors corresponding to the candidate information based on each negative feedback embedding vector and the candidate embedding vector; obtain the similarity between each negative feedback embedding vector and the third embedding vector; determine the target weight of the negative feedback information based on the multiple similarities; perform a weighted summation of the negative feedback embedding vectors based on the target weight to obtain a fifth embedding vector corresponding to the candidate information; and concatenate the third embedding vector, the fifth embedding vector, and the fourth embedding vector in sequence to obtain the candidate embedding vector corresponding to the candidate information, wherein the candidate embedding vector corresponding to the candidate information includes the content embedding vector corresponding to the content information and the advertising embedding vector corresponding to the advertising information.
[0252] Optionally, the expectation determination unit 1140 can also be used to determine a content matrix and an advertising matrix that match each candidate display content, wherein the content matrix is used to represent the content information to be displayed, and the advertising matrix is used to represent the advertising information to be displayed; the product of the content matrix corresponding to the candidate display content and the first matrix is used as the first intermediate matrix, and the product of the advertising matrix corresponding to the candidate display content and the second matrix is used as the second intermediate matrix; the first intermediate matrix and the second intermediate matrix are summed to obtain the cross matrix matching the candidate display content; the weights of each element in the cross matrix are adjusted through a self-attention mechanism to obtain the adjusted third matrix; the third matrix is dimensionality reduced to obtain the sixth embedding vector matching the candidate display content, which is used as the expected gain of the deviation generated by the candidate display content.
[0253] Optionally, the expectation determination unit 1140 can also be used to randomly generate a specified number of masking matrices; multiply the cross matrix with each of the masking matrices respectively to obtain a specified number of masked matrices that match the candidate display content; adjust the weight of each element in each masked matrix that matches the candidate display content through a self-attention mechanism to obtain a fourth matrix after weight adjustment for each masked matrix; perform dimensionality reduction on each of the fourth matrices and concatenate the dimensionality-reduced fourth matrices to obtain the seventh embedding vector that matches the candidate display content, which is used as the expected gain of the deviation generated by the candidate display content.
[0254] Optionally, the expectation determination unit 1140 can also be used to obtain the average value of the deviation expected benefit of each of the candidate display contents as the average expected benefit; to take the difference between the deviation expected benefit of each of the candidate display contents and the average expected benefit as the intermediate expected benefit corresponding to the candidate display content; and to obtain the sum of the state expected benefit and the intermediate expected benefit corresponding to the candidate display content as the expected benefit corresponding to the candidate display content.
[0255] The target display content determination unit 1150 is used to select the candidate display content with the highest expected benefit as the target display content and push the target display content to the target user.
[0256] Optionally, the target display content determination unit 1150 can also be used to find candidate display content that meets the pre-acquired strategy requirements and has the greatest expected benefit, as the target display content, and push the target display content to the target user.
[0257] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the above-described apparatus and unit can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0258] In the several embodiments provided in this application, the coupling between the units can be electrical, mechanical, or other forms of coupling. Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0259] Please see Figure 12 , Figure 12 This diagram illustrates a structural block diagram of an electronic device according to an embodiment of this application. The electronic device 1200 may be a desktop computer, an in-vehicle computer, a cloud server, or a tablet computer, etc. The electronic device 1200 in this application may include one or more of the following components: a processor 1201, a memory 1202, and one or more application programs, wherein the processor 1201 is electrically connected to the memory 1202, and the one or more programs are configured to execute the methods described in the foregoing embodiments.
[0260] Processor 1201 may include one or more processing cores. Processor 1201 connects to various parts within the electronic device 1200 using various interfaces and lines, and performs various functions and processes data of the electronic device 1200 by running or executing instructions, programs, code sets, or instruction sets stored in memory 1202, and by calling data stored in memory 1202. Optionally, processor 1201 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). Processor 1201 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and computer programs; the GPU is responsible for rendering and drawing the displayed content; and the modem is used for wireless communication. It is understood that the modem may also not be integrated into processor 1201 and may be implemented separately through a communication chip. Specifically, the methods described in the foregoing embodiments can be executed by one or more processors 1201.
[0261] In some implementations, memory 1202 may include random access memory (RAM) or read-only memory (ROM). Memory 1202 can be used to store instructions, programs, code, code sets, or instruction sets. Memory 1202 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function, instructions for implementing the various method embodiments described below, etc. The data storage area may also store data created by the electronic device 1200 during use.
[0262] Please see Figure 13 This diagram illustrates a structural block diagram of a computer-readable storage medium provided in an embodiment of this application. The computer-readable medium 1300 stores program code that can be called by a processor to execute the methods described in the above method embodiments.
[0263] The computer-readable storage medium 1300 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, the computer-readable storage medium 1300 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 1300 has storage space for program code 1310 that performs any of the method steps described above. This program code can be read from or written to one or more computer program products. The program code 1310 may be compressed, for example, in a suitable form.
[0264] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. An information recommendation method characterized by comprising: The method comprises: obtaining historical behavior information of a target user and portrait information of the target user in response to an information recommendation request initiated by the target user; obtaining context information corresponding to the information recommendation request, and obtaining to-be-recommended content and to-be-recommended advertisements matched with the context information; constructing different to-be-displayed contents based on at least one content information in the to-be-recommended content and at least one advertisement information in the to-be-recommended advertisements by using a pre-trained information recommendation model; determining an expected income corresponding to each to-be-displayed content based on the historical behavior information, the context information and the portrait information by using the information recommendation model; pushing a target display content, which has the maximum expected income, to the target user.
2. The method of claim 1, wherein, The method comprises: inputting state information as an input parameter into the information recommendation model to obtain a state expected income corresponding to the state information, wherein the state information comprises the historical behavior information, the context information, the portrait information, the to-be-recommended content and the to-be-recommended advertisements; obtaining a deviation expected income generated by adopting each to-be-displayed content, wherein the deviation expected income is used to represent a deviation compared with the state expected income; determining the expected income corresponding to each to-be-displayed content based on the state expected income and the deviation expected income.
3. The method of claim 2, wherein, The historical behavior information comprises positive feedback information and negative feedback information, wherein the positive feedback information is used to represent that a user consumes a displayed historical advertisement, and the negative feedback information is used to represent that the user does not consume the displayed historical advertisement. The method comprises: generating a content embedding vector corresponding to each content information based on the positive feedback information, the negative feedback information and each content information in the to-be-recommended content, and generating an advertisement embedding vector corresponding to each advertisement information based on the positive feedback information, the negative feedback information and each advertisement information in the to-be-recommended advertisements; obtaining a context embedding vector corresponding to the context information and obtaining a portrait embedding vector corresponding to the portrait information; splicing the content embedding vector corresponding to the content information with the context embedding vector and the portrait embedding vector to obtain a first embedding vector corresponding to the content information, and splicing the advertisement embedding vector corresponding to the advertisement information with the context embedding vector and the portrait embedding vector to obtain a second embedding vector corresponding to the advertisement information; 4. The method of claim 3, wherein, generating the state expected income based on the first embedding vector corresponding to each content information and the second embedding vector corresponding to each advertisement information. The method comprises: The first embedding vector corresponding to each of the content information is used as a row vector to construct a first matrix; The second embedding vector corresponding to each of the aforementioned advertising information is used as a row vector to construct a second matrix; Dimensionality reduction is performed on the first matrix and the second matrix respectively; The first matrix after dimensionality reduction and the second matrix after dimensionality reduction are concatenated to obtain the expected return of the state.
5. The method of claim 3, wherein, The negative feedback information is multiple. The step of generating a content embedding vector corresponding to each content information based on the positive feedback information, negative feedback information, and each content information in the content to be recommended, and generating an ad embedding vector corresponding to each ad information based on the positive feedback information, negative feedback information, and each ad information in the ad to be recommended, includes: Obtain the positive feedback embedding vector of the positive feedback information, the negative feedback embedding vector of each negative feedback information, and the candidate embedding vector of the candidate information, wherein the candidate information includes content information and advertising information; Based on the positive feedback embedding vector and the candidate embedding vector, a third embedding vector corresponding to the candidate information is constructed, and based on each of the negative feedback embedding vectors and the candidate embedding vectors, multiple fourth embedding vectors corresponding to the candidate information are constructed. Obtain the similarity between each negative feedback embedding vector and the third embedding vector; The target weight of the negative feedback information is determined based on multiple similarities. The negative feedback embedding vector is weighted and summed based on the target weight to obtain the fifth embedding vector corresponding to the candidate information. The third, fifth, and fourth embedding vectors are concatenated sequentially to obtain the target candidate embedding vector corresponding to the candidate information. The target candidate embedding vector corresponding to the candidate information includes the content embedding vector corresponding to the content information and the advertising embedding vector corresponding to the advertising information.
6. The method of claim 4, wherein, The step of obtaining the expected gain from the deviation generated by each of the candidate display contents includes: A content matrix and an advertising matrix are determined to match each type of candidate display content, wherein the content matrix is used to represent the content information to be displayed, and the advertising matrix is used to represent the advertising information to be displayed; The product of the content matrix corresponding to the candidate display content and the first matrix is used as the first intermediate matrix, and the product of the advertisement matrix corresponding to the candidate display content and the second matrix is used as the second intermediate matrix. Summing the first intermediate matrix and the second intermediate matrix yields the cross matrix matching the candidate display content; The weights of each element in the cross matrix are adjusted using a self-attention mechanism to obtain the adjusted third matrix. The third matrix is dimensionality reduced to obtain the sixth embedding vector that matches the candidate display content, which is used as the expected gain of the deviation generated by the candidate display content.
7. The method of claim 6, wherein, The step of adjusting the weights of each element in the cross matrix using a self-attention mechanism to obtain the adjusted third matrix includes: Randomly generate a specified number of masking matrices; The cross matrix is multiplied by each of the masking matrices to obtain a specified number of masked matrices that match the content to be displayed. The self-attention mechanism is used to adjust the weight of each element in each masked matrix matched with the candidate display content, to obtain a fourth matrix corresponding to each masked matrix after weight adjustment; The third matrix is dimensionally reduced to obtain a sixth embedding vector of the candidate display content matching, as a deviation expected return generated by the candidate display content, including: Each fourth matrix is dimensionally reduced, and the fourth matrix after dimensional reduction is spliced to obtain a seventh embedding vector of the candidate display content matching, as a deviation expected return generated by the candidate display content.
8. The method of claim 2, wherein Based on the state expected return and the deviation expected return, the expected return corresponding to each candidate display content is determined, including: An average value of the deviation expected return of each candidate display content is obtained as an average expected return; The difference between the deviation expected return of each candidate display content and the average expected return is taken as the intermediate expected return corresponding to the candidate display content; The sum of the state expected return and the intermediate expected return corresponding to the candidate display content is taken as the expected return corresponding to the candidate display content.
9. The method of claim 1, wherein, The context information corresponding to the information recommendation request, and the recommended content and the recommended advertisement matched with the context information are obtained, including: The context information corresponding to the information recommendation request is obtained; The context information is input into a pre-obtained content recommendation algorithm to obtain the recommended content; The context information is input into a pre-obtained advertisement recommendation algorithm to obtain the recommended advertisement.
10. The method of claim 1, wherein, The candidate display content with the maximum expected return is taken as the target display content, including: The candidate display content satisfying the pre-obtained strategy requirement and having the maximum expected return is found as the target display content, and the target display content is pushed to the target user.
11. An information recommendation device characterized by comprising: Including: A response unit is configured to obtain historical behavior information of a target user and portrait information of the target user in response to an information recommendation request initiated by the target user; An obtaining unit is configured to obtain context information corresponding to the information recommendation request, and recommended content and recommended advertisement matched with the context information; A display content construction unit is configured to construct different candidate display contents based on at least one content information in the recommended content and at least one advertisement information in the recommended advertisement through a pre-trained information recommendation model; An expected determination unit is configured to determine an expected return corresponding to each candidate display content based on the historical behavior information, context information, and portrait information through the information recommendation model; A target display content determination unit is configured to take the candidate display content with the maximum expected return as the target display content, and push the target display content to the target user.
12. An electronic device, comprising: Including: One or more processors; Memory; One or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the method of any one of claims 1-10.
13. A computer-readable storage medium, characterized in that, The computer readable storage medium stores program codes, which can be invoked by the processor to execute the method of any one of claims 1-10.