Content recommendation method, apparatus, device, and storage medium
By acquiring the channel and content preference features of the target object and using a recommendation model trained with reinforcement learning algorithms, the recommendation is made by comprehensively considering channel and content information, which solves the problem of poor recommendation effect in the existing technology and improves click-through rate and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-29
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies only consider content-level information in content recommendation, resulting in poor recommendation performance, low click-through rates, and a poor user interaction experience.
By acquiring the channel preference features and content preference features of the target object, a recommendation model trained based on reinforcement learning algorithm is used to obtain the target channel sequence and target content sequence, and recommendations are made by comprehensively considering channel and content information.
It improved the effectiveness of content recommendations, increased click-through rates, and enhanced the user's interactive experience.
Smart Images

Figure CN111552888B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a content recommendation method, apparatus, device, and storage medium. Background Technology
[0002] With the rapid development of artificial intelligence (AI) technology, an increasing number of applications are leveraging AI to recommend personalized content (videos, news, articles, etc.) to users, thereby enhancing their interactive experience. In many applications, the content to be recommended can come from multiple channels, such as video channels, news channels, and article channels.
[0003] In the process of recommending content to users, related technologies first predict the click-through rate (CTR) of each candidate item, then rank the candidates based on the predicted CTR, and recommend the top-ranked items to the user. This method of recommending content directly ranks each candidate item, considering only content-level information. This limited information leads to poor content recommendation performance, lower CTRs, and a less-than-ideal user experience. Summary of the Invention
[0004] This application provides a content recommendation method, apparatus, device, and storage medium, which can be used to improve the effectiveness of content recommendation. The technical solution is as follows:
[0005] On the one hand, embodiments of this application provide a content recommendation method, the method comprising:
[0006] Obtain the channel preference features, content preference features, and candidate content set corresponding to the target object. The candidate content set includes at least one candidate content, and each candidate content corresponds to a candidate channel.
[0007] Based on the channel preference features and the candidate channel set corresponding to the candidate content set, a target channel sequence is obtained, wherein the candidate channel set includes the candidate channel corresponding to each candidate content in the candidate content set.
[0008] Based on the content preference features and the candidate content set, obtain the target content sequence corresponding to the target channel sequence;
[0009] The target content sequence is recommended to the target object.
[0010] In one possible implementation, the step of invoking the second target recommendation model to perform recommendation processing on the content preference features to obtain a content recommendation result includes:
[0011] The content preference features are input into the second target recommendation model to obtain the first content recommendation sub-result output by the second target recommendation model;
[0012] obtaining an updated content preference feature based on the first content recommendation sub-result, inputting the updated content preference feature into the second target recommendation model, and obtaining a second content recommendation sub-result output by the second target recommendation model;
[0013] continuing to obtain content recommendation sub-results based on the second content recommendation sub-result;
[0014] in response to the number of content recommendation sub-results being equal to the number threshold, arranging each content recommendation sub-result in the order of obtaining as the content recommendation result.
[0015] A content recommendation method is also provided, and the method comprises:
[0016] obtaining a target hierarchical recommendation model, the target hierarchical recommendation model comprising a first target recommendation model and a second target recommendation model;
[0017] calling the first target recommendation model to perform recommendation processing on a channel preference feature corresponding to a target object, and obtaining a channel recommendation result; and based on the channel recommendation result and a candidate channel set corresponding to the target object, obtaining a target channel sequence;
[0018] calling the second target recommendation model to perform recommendation processing on a content preference feature corresponding to the target object, and obtaining a content recommendation result; and based on the content recommendation result and a candidate content set, obtaining a target content sequence corresponding to the target channel sequence;
[0019] recommending the target content sequence to the target object.
[0020] In another aspect, a content recommendation device is provided, and the device comprises:
[0021] a first obtaining unit, configured to obtain a channel preference feature, a content preference feature and a candidate content set corresponding to a target object, the candidate content set comprising at least one candidate content, and any candidate content corresponding to a candidate channel;
[0022] a second obtaining unit, configured to obtain a target channel sequence based on the channel preference feature and a candidate channel set corresponding to the candidate content set, the candidate channel set comprising a candidate channel corresponding to each candidate content in the candidate content set;
[0023] a third obtaining unit, configured to obtain a target content sequence corresponding to the target channel sequence based on the content preference feature and the candidate content set;
[0024] a recommendation unit, configured to recommend the target content sequence to the target object.
[0025] In a possible implementation, the second obtaining unit is configured to invoke a first target recommendation model to perform recommendation processing on the channel preference feature, to obtain a channel recommendation result composed of at least one channel recommendation sub-result; for any channel recommendation sub-result, a target channel matching the channel recommendation sub-result is obtained from the candidate channel set; and a target channel sequence is obtained based on the target channels respectively matching the channel recommendation sub-results.
[0026] In a possible implementation, the third obtaining unit is configured to invoke a second target recommendation model to perform recommendation processing on the content preference feature, to obtain a content recommendation result composed of at least one content recommendation sub-result; for any content recommendation sub-result, a target content matching the content recommendation sub-result is obtained from a target candidate content set composed of candidate contents in the candidate content set that meet a condition, the candidate content meeting the condition including a candidate content whose corresponding candidate channel is a specified channel, the specified channel including a target channel whose arrangement position in the target channel sequence and the arrangement position of the content recommendation sub-result in the content recommendation result are consistent; and a target content sequence corresponding to the target channel sequence is obtained based on the target contents respectively matching the content recommendation sub-results.
[0027] In a possible implementation, the second obtaining unit is further configured to input the channel preference feature into the first target recommendation model, to obtain a first channel recommendation sub-result output by the first target recommendation model; obtain an updated channel preference feature based on the first channel recommendation sub-result, input the updated channel preference feature into the first target recommendation model, to obtain a second channel recommendation sub-result output by the first target recommendation model; continue to obtain channel recommendation sub-results based on the second channel recommendation sub-result; and in response to the number of obtained channel recommendation sub-results being equal to a quantity threshold, arrange each channel recommendation sub-result in a sequence according to the order of obtaining, as the channel recommendation result.
[0028] In a possible implementation, the third obtaining unit is further configured to input the content preference feature into the second target recommendation model, to obtain a first content recommendation sub-result output by the second target recommendation model; obtain an updated content preference feature based on the first content recommendation sub-result, input the updated content preference feature into the second target recommendation model, to obtain a second content recommendation sub-result output by the second target recommendation model; continue to obtain content recommendation sub-results based on the second content recommendation sub-result; and in response to the number of content recommendation sub-results being equal to a quantity threshold, arrange each content recommendation sub-result in a sequence according to the order of obtaining, as the content recommendation result.
[0029] In a possible implementation, the first obtaining unit is configured to obtain a historical recommendation content sequence corresponding to the target object; based on the historical recommendation content sequence, obtain a to-be-processed channel feature sequence and a to-be-processed content feature sequence; call the first processing model to process the to-be-processed channel feature sequence, to obtain a channel preference feature corresponding to the target object; and call the second processing model to process the to-be-processed content feature sequence, to obtain a content preference feature corresponding to the target object.
[0030] In a possible implementation, the historical recommendation content sequence is composed of at least one historical recommendation content arranged in sequence; the first obtaining unit is further configured to, for any historical recommendation content, obtain corresponding basic information, to-be-processed channel information, and to-be-processed content information of the any historical recommendation content, wherein the basic information includes at least one of object attribute information and environment information, the to-be-processed channel information includes at least one of channel information and accumulated channel information, and the to-be-processed content information includes at least one of content information and accumulated content information; perform fusion processing on the basic information and the to-be-processed channel information corresponding to the any historical recommendation content, to obtain to-be-processed channel features corresponding to the any historical recommendation content; perform fusion processing on the basic information and the to-be-processed content information corresponding to the any historical recommendation content, to obtain to-be-processed content features corresponding to the any historical recommendation content; arrange the to-be-processed channel features corresponding to each historical recommendation content in the arrangement order of each historical recommendation content in the historical recommendation content sequence, to obtain the to-be-processed channel feature sequence; and arrange the to-be-processed content features corresponding to each historical recommendation content in the arrangement order of each historical recommendation content in the historical recommendation content sequence, to obtain the to-be-processed content feature sequence.
[0031] Also provided is a content recommendation apparatus, which comprises:
[0032] a first obtaining unit configured to obtain a target hierarchical recommendation model, wherein the target hierarchical recommendation model comprises a first target recommendation model and a second target recommendation model;
[0033] a second obtaining unit configured to call the first target recommendation model to perform recommendation processing on a channel preference feature corresponding to a target object, to obtain a channel recommendation result; and based on the channel recommendation result and a candidate channel set corresponding to the target object, obtain a target channel sequence;
[0034] a third obtaining unit configured to call the second target recommendation model to perform recommendation processing on a content preference feature corresponding to the target object, to obtain a content recommendation result; and based on the content recommendation result and a candidate content set corresponding to the target object, obtain a target content sequence corresponding to the target channel sequence;
[0035] The recommendation unit is configured to recommend the target content sequence to the target object.
[0036] In a possible implementation, the apparatus further includes:
[0037] The fourth obtaining unit is configured to obtain a training sample set, the training sample set including at least one training sample, any training sample including sample channel features, sample content features, and feedback information corresponding to a sample recommended content sequence;
[0038] The training unit is configured to train a first initial recommendation model and a second initial recommendation model in an initial hierarchical recommendation model based on the sample channel features, the sample content features, and the feedback information corresponding to the sample recommended content sequence in the training sample, to obtain a target hierarchical recommendation model.
[0039] In a possible implementation, the first initial recommendation model includes a first initial recommendation submodel and a first initial evaluation submodel, and the second initial recommendation model includes a second initial recommendation submodel and a second initial evaluation submodel; the training unit is configured to obtain a first set of enhancement values and a second set of enhancement values based on the feedback information corresponding to the sample recommended content sequence in the training sample; input the sample channel features in the training sample into the first initial recommendation model, to obtain an initial channel recommendation result output by the first initial recommendation submodel and a first set of evaluation values for the initial channel recommendation result output by the first initial evaluation submodel; input the sample content features in the training sample into the second initial recommendation model, to obtain an initial content recommendation result output by the second initial recommendation submodel and a second set of evaluation values for the initial content recommendation result output by the second initial evaluation submodel; update parameters of the first initial recommendation submodel based on the first set of evaluation values; update parameters of the second initial recommendation submodel based on the second set of evaluation values; obtain a channel loss function based on the first set of enhancement values and the first set of evaluation values; obtain a content loss function based on the second set of enhancement values and the second set of evaluation values; calculate a target loss function based on the channel loss function and the content loss function; and update parameters of the first initial evaluation submodel and the second initial evaluation submodel based on the target loss function.
[0040] In a possible implementation, the training sample further includes a sample recommended content sequence; and the training unit is further configured to obtain at least one of a click rate loss function and a similarity loss function based on the initial content recommendation result and the sample recommended content sequence in the training sample; and calculate a target loss function based on at least one of the click rate loss function and the similarity loss function, and the channel loss function and the content loss function.
[0041] In a possible implementation, the training unit is further configured to: acquire at least one of reading duration information, diversity information, and novelty information of any sample recommended content in the sample recommended content sequence and click information of the any sample recommended content based on the feedback information in the training sample; acquire a first enhancement value corresponding to the any sample recommended content based on the click information of the any sample recommended content; acquire a second enhancement value corresponding to the any sample recommended content based on the at least one of the reading duration information, the diversity information, and the novelty information of the any sample recommended content and the click information of the any sample recommended content; and set a set of the first enhancement values corresponding to each sample recommended content as a first enhancement value set and set a set of the second enhancement values corresponding to each sample recommended content as a second enhancement value set.
[0042] In another aspect, a computer device is provided, which includes a processor and a memory, and the memory stores at least one program code, which is loaded and executed by the processor to implement the content recommendation method described above.
[0043] In another aspect, a computer readable storage medium is also provided, which stores at least one program code, which is loaded and executed by a processor to implement the content recommendation method described above.
[0044] The technical solutions provided by the embodiments of the present application at least bring the following beneficial effects:
[0045] The target channel sequence is first acquired based on the channel preference feature, and then the target content sequence corresponding to the target channel sequence is acquired based on the content preference feature, and then the target content sequence is recommended to the target object. In the content recommendation process, the channel preference feature reflects the information of the channel aspect, the content preference feature reflects the information of the content aspect, the content recommendation process considers both the information of the channel aspect and the information of the content aspect, the considered information is more comprehensive, which is conducive to improving the effect of content recommendation, the click rate of the recommended content is higher, and the user's interactive experience is better. BRIEF DESCRIPTION OF DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0047] Figure 1 is a process schematic diagram of reinforcement learning provided by the embodiments of the present application;
[0048] Figure 2 is a schematic diagram of an implementation environment of a content recommendation method provided by an embodiment of the present application;
[0049] Figure 3 is a flowchart of a content recommendation method provided by an embodiment of the present application;
[0050] Figure 4 is a schematic diagram of a process of displaying a recommendation page on a terminal screen provided by an embodiment of the present application;
[0051] Figure 5 is a schematic diagram of a process of obtaining a target content sequence provided by an embodiment of the present application;
[0052] Figure 6 is a flowchart of a content recommendation method provided by an embodiment of the present application;
[0053] Figure 7 is a flowchart of a method of training a first initial recommendation model and a second initial recommendation model in an initial hierarchical recommendation model provided by an embodiment of the present application;
[0054] Figure 8 is a schematic diagram of a content recommendation device provided by an embodiment of the present application;
[0055] Figure 9 is a schematic diagram of a content recommendation device provided by an embodiment of the present application;
[0056] Figure 10 is a schematic diagram of a content recommendation device provided by an embodiment of the present application;
[0057] Figure 11 is a schematic diagram of a structure of a server provided by an embodiment of the present application;
[0058] Figure 12 is a schematic diagram of a structure of a terminal provided by an embodiment of the present application. DETAILED DESCRIPTION
[0059] In order to make the purpose, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the drawings.
[0060] Artificial Intelligence (AI) is the theory, method, technology and application system of using digital computers or digital computer controlled machines to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.
[0061] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technologies generally include technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, machine learning and natural language processing technology, etc.
[0062] Among them, machine learning (Machine Learning, ML) is a multi-disciplinary subject, involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It is a subject that studies how computers simulate or implement human learning behavior to acquire new knowledge or skills, and reorganize existing knowledge structure to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental approach to making computers intelligent, and its applications are widespread in various fields of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and teaching learning technologies. Among them, reinforcement learning is a field of machine learning that emphasizes how to act based on the environment to achieve maximum expected benefits. Deep reinforcement learning is a combination of deep learning and reinforcement learning, which uses deep learning techniques to solve reinforcement learning problems.
[0063] Reinforcement learning is to learn an optimal strategy that can allow the agent to make actions based on the current state in a specific environment, so as to obtain the maximum reward.
[0064] Reinforcement learning can be modeled simply by the four-tuple <A, S, R, P>. A represents Action, which is the action issued by the Agent; State is the state of the world that the Agent can perceive; Reward is a real number value representing reward or punishment; P is the environment that the Agent interacts with.
[0065] The influence relationship between the four-tuple <A, S, R, P> is as follows:
[0066] Action space: A, i.e. all actions A constitute the action space.
[0067] State space: S, i.e. all states S constitute the state space.
[0068] Reward: S*A*S’->R, i.e. after performing action A in the current state S, the current state becomes S’, and the reward R corresponding to action A is obtained.
[0069] Transition: S*A->S’, i.e. after performing action A in the current state S, the current state becomes S’.
[0070] In fact, the process of reinforcement learning is a process of continuous iteration, as shown in Figure 1 , in the process of continuous iteration, for the agent, after obtaining the state o t and the reward r t of the environment feedback, action a t is performed; for the environment, after accepting the action a t performed by the agent, the state o t+1 and the reward r t+1 of the environment feedback are output. The recommendation model used in the embodiments of the present application is trained based on a reinforcement learning algorithm.
[0071] With the rapid development of artificial intelligence technology, more and more application scenarios use artificial intelligence technology to recommend personalized content (video, news, article, etc.) for users to improve the user’s interactive experience. In many application scenarios, the content to be recommended can come from multiple channels, for example, a video channel, a news channel, an article channel, etc.
[0072] To this end, the embodiments of the present application provide a content recommendation method, please refer to Figure 2 , which shows a schematic diagram of an implementation environment of the content recommendation method provided by the embodiments of the present application. The implementation environment can include a terminal 21 and a server 22.
[0073] The terminal 21 is installed with an application program or webpage capable of recommending content to the target object, and the application program or webpage is capable of recommending content to the target object based on the method provided in the embodiments of the present application. In the embodiments of the present application, the content available for recommendation includes but is not limited to long videos, short videos, news, articles, etc., and the number of content recommended to the target object can be one or more. In the process of recommending content to the target object, the terminal 21 can obtain the channel preference feature, the content preference feature and the candidate content set corresponding to the target object, and then obtain the target content sequence and recommend it to the target object; of course, the server 22 can also obtain the channel preference feature, the content preference feature and the candidate content set corresponding to the target object, and then obtain the target content sequence, and after obtaining the target content sequence, the server 22 can send the target content sequence to the terminal 21, and the terminal 21 recommends the target content sequence to the target object.
[0074] In a possible implementation manner, the terminal 21 can be any electronic product capable of human-computer interaction with a user through one or more of a keyboard, a touchpad, a touch screen, a remote controller, voice interaction or a handwriting device, for example, a PC (Personal Computer), a mobile phone, a smart phone, a PDA (Personal Digital Assistant), a wearable device, a PPC (Pocket PC), a tablet computer, a smart car machine, a smart television, a smart speaker, etc. The server 22 can be a server or a server cluster composed of multiple servers, or a cloud computing service center. The terminal 21 and the server 22 establish a communication connection through a wired or wireless network.
[0075] Those skilled in the art should understand that the terminal 21 and the server 22 described above are only examples, and other existing or future terminal or server, such as those applicable to the present application, should also be included in the protection scope of the present application and are hereby included by reference.
[0076] Based on the above Figure 2 The embodiments of the present application provide a content recommendation method. As shown in Figure 3 The content recommendation method provided by the embodiments of the present application can include the following steps:
[0077] In step 301, the channel preference feature, the content preference feature and the candidate content set corresponding to the target object are obtained.
[0078] The candidate content set includes at least one candidate content, and any candidate content corresponds to a candidate channel.
[0079] The target object refers to any interactive object that needs the terminal to recommend content. It should be noted that in the embodiments of the present application, the content available for recommendation includes but is not limited to short videos, long videos, news, articles and other structured content. Different structures of content come from different channels. For example, short video structured content comes from a short video channel, and news type content comes from a news channel. Content from the same channel is homogeneous content, and content from different channels is heterogeneous content. In the embodiments of the present application, the content available for recommendation can include both homogeneous content and heterogeneous content. When there is heterogeneous content available for recommendation, different structures of content can be recommended to the target object, improving the diversity of recommended content and the interactive experience of the target object. When there is heterogeneous content available for recommendation, the recommendation process is referred to as a comprehensive recommendation process, and comprehensive recommendation refers to recommending heterogeneous content from different channels to the target object.
[0080] The terminal is installed with an application program or a webpage capable of recommending content, and when the target object opens the application program or the webpage, a recommendation content acquisition request can be sent in the application program or the webpage to acquire the recommended content and browse the recommended content. It should be noted that the process of acquiring recommended content in the embodiments of the present application can be executed based on a recommendation content acquisition request, or can be executed based on a pre-set trigger condition, which is not limited in the embodiments of the present application. The pre-set trigger condition can refer to a pre-set trigger time interval, etc.
[0081] In the process of recommending content, the terminal first acquires the channel preference feature, the content preference feature and the candidate content set corresponding to the target object. It should be noted that the channel preference feature, the content preference feature and the candidate content set acquired here are all acquired for the target object. That is, the process of recommending content is a personalized recommendation process for the interactive object.
[0082] The channel preference feature is used to represent the preference of the target object in the channel, and the content preference feature is used to represent the preference of the target object in the content. In one possible implementation, the process of acquiring the channel preference feature and the content preference feature corresponding to the target object includes the following steps 1 to 3:
[0083] Step 1: Acquire the historical recommendation content sequence corresponding to the target object.
[0084] The historical recommendation content sequence is composed of at least one historical recommendation content arranged in sequence. The historical recommendation content refers to content that has been recommended to the target object. The historical recommendation content can be obtained from a historical behavior log of the target object. It should be noted that the number of historical recommendation contents used to constitute the historical recommendation content sequence, the conditions that the historical recommendation contents need to meet, and the arrangement order of each historical recommendation content can be set according to experience, or can be flexibly adjusted according to application scenarios, and the embodiments of the present application do not limit them.
[0085] Exemplarily, the number of historical recommendation contents can be set to 50, and the condition that the historical recommendation contents need to meet can be set to a time interval between a recommendation time stamp and a current time stamp not exceeding a time interval threshold. In the case that the number of historical recommendation contents is 50, by adjusting the time interval threshold, the historical recommendation contents can be limited to the 50 historical recommendation contents that have been recently recommended to the target object.
[0086] The arrangement order of each historical recommendation content can refer to the chronological order of the recommendation time stamps of the historical recommendation contents. It should be noted that for the case that multiple historical recommendation contents are recommended at the same time, the front-back order of the positions of the multiple historical recommendation contents on the terminal screen can be used as the arrangement order of the multiple historical recommendation contents with the same recommendation time stamp.
[0087] Step 2: Based on the historical recommendation content sequence, obtain a to-be-processed channel feature sequence and a to-be-processed content feature sequence.
[0088] The to-be-processed channel feature sequence is composed of at least one to-be-processed channel feature arranged in sequence, and the to-be-processed content feature sequence is composed of at least one to-be-processed content feature arranged in sequence. It should be noted that the number of to-be-processed channel features and the number of to-be-processed content features are the same as the number of historical recommendation contents. That is, based on each historical recommendation content, a to-be-processed channel feature and a to-be-processed content feature are obtained. Exemplarily, the to-be-processed channel feature sequence can be represented as wherein, represents the to-be-processed channel feature sequence, m (m is an integer not less than 1) represents the number of historical recommendation contents, represents the to-be-processed channel feature located at the mth arrangement position in the to-be-processed channel feature sequence. Exemplarily, the to-be-processed content feature sequence can be represented as wherein, represents the to-be-processed content feature sequence, m represents the number of historical recommendation contents, represents the to-be-processed content feature located at the mth arrangement position in the to-be-processed content feature sequence.
[0089] In a possible implementation manner, based on the historical recommended content sequence, the process of obtaining the to-be-processed channel feature sequence and the to-be-processed content feature sequence includes the following steps a to d:
[0090] Step a: For any historical recommended content, obtain the basic information, the to-be-processed channel information, and the to-be-processed content information corresponding to the historical recommended content, the basic information includes at least one of object attribute information and environment information, the to-be-processed channel information includes at least one of channel information and accumulated channel information, and the to-be-processed content information includes at least one of content information and accumulated content information.
[0091] The historical recommended content sequence is composed of at least one historical recommended content arranged in sequence, that is, each historical recommended content in the historical recommended content sequence has a sequence position. For any historical recommended content in each historical recommended content constituting the historical recommended content sequence, the related information corresponding to the historical recommended content is obtained, so as to further obtain the to-be-processed channel feature and the to-be-processed content feature corresponding to the historical recommended content by using the related information of the historical recommended content.
[0092] In the embodiment of the application, the related information corresponding to any historical recommended content includes but is not limited to the basic information, the to-be-processed channel information, and the to-be-processed content information. The basic information includes at least one of object attribute information and environment information, the to-be-processed channel information includes at least one of channel information and accumulated channel information, and the to-be-processed content information includes at least one of content information and accumulated content information. Next, the object attribute information, the environment information, the channel information, the accumulated channel information, the content information, and the accumulated content information corresponding to any historical recommended content are introduced respectively:
[0093] The object attribute information is obtained based on the object attribute of the target object. The object attribute information can include basic attribute information (for example, age, gender, home address, position, social relationship, etc.) of the target object, interest preference information (for example, favorite theme, tag, category), and cross information, etc. In the process of continuously interacting with the terminal, the terminal will build and continuously update the object attribute of the target object. The object attribute information corresponding to any historical recommended content can be extracted from the object attribute of the target object that has been built by the terminal when the any historical recommended content is recommended.
[0094] The environment information refers to information of a recommendation environment when the any historical recommendation content is recommended. The environment information includes, but is not limited to, a terminal device type (for example, an IOS mobile phone, an Android mobile phone, a computer, and the like), a network type (for example, a 4G network, a WiFi network, and the like), a time factor (for example, a recommendation time stamp, and the like), and a current location of the terminal, and the like. The environment information corresponding to any historical recommendation content can be acquired and stored when the any historical recommendation content is recommended. At this time, the environment information corresponding to the any historical recommendation content can be directly extracted from the storage.
[0095] The channel information refers to information of the any historical recommendation content at a channel level, and is used to indicate a channel from which the any historical recommendation content is derived. For example, when the any historical recommendation content is a short video introducing a certain product, the channel information corresponding to the any historical recommendation content is used to indicate a short video channel. The channel information can refer to a feature of the channel from which the any historical recommendation content is derived. The embodiments of the present application do not limit this.
[0096] The content information refers to information of the any historical recommendation content at a content level. In a possible implementation manner, the content information corresponding to the any historical recommendation content includes, but is not limited to, classification information (for example, a tag, a category, a theme, a channel, and the like) of the any historical recommendation content, popularity information, time sensitivity, a content providing program, cross information, and the like. The content information corresponding to the any historical recommendation content can be stored together with the any historical recommendation content. When the any historical recommendation content is acquired, the content information corresponding to the any historical recommendation content can be acquired.
[0097] The cumulative channel information is used to represent, to a certain extent, a preference of the target object in the channel. In a possible implementation manner, the cumulative channel information corresponding to the any historical recommendation content is acquired in the following manner: a historical recommendation content arranged before the any historical recommendation content in the sequence of historical recommendation contents is regarded as a preceding historical recommendation content corresponding to the any historical recommendation content; and the cumulative channel information corresponding to the any historical recommendation content is acquired based on a triggering condition of a channel corresponding to the preceding historical recommendation content. The triggering condition of the channel is used to indicate whether the target object triggers the channel. Based on the triggering condition of the channel corresponding to the preceding historical recommendation content, the process of acquiring the cumulative channel information corresponding to the any historical recommendation content can be: based on the triggering condition of the channel corresponding to the preceding historical recommendation content, at least one of a number of times each channel is triggered and a proportion of times each channel is triggered is counted, and the counted information is regarded as the cumulative channel information corresponding to the any historical recommendation content.
[0098] The accumulated content information is used to represent the preference of the target object in the content to some extent. In a possible implementation, the accumulated content information corresponding to any historical recommended content is obtained in the following manner: historical recommended contents arranged before the any historical recommended content in the sequence of historical recommended contents are regarded as the preceding historical recommended contents corresponding to the any historical recommended content; and the accumulated content information corresponding to the any historical recommended content is obtained based on the triggering condition of the preceding historical recommended contents. The triggering condition of the content is used to indicate whether the target object triggers the content. Based on the triggering condition of the preceding historical recommended contents, the process of obtaining the accumulated content information corresponding to the any historical recommended content can be: based on the triggering condition of the preceding historical recommended contents, at least one of the number of times and the proportion of each content label being triggered is counted, and the counted information is taken as the accumulated content information corresponding to the any historical recommended content. The content label is used to represent the category, theme and other related information of the recommended content, and one historical recommended content can correspond to one or more content labels.
[0099] Exemplarily, when the any historical recommended content is arranged at the i th (i is an integer not less than 1 and not greater than m) position in the sequence of historical recommended contents (i.e., at the i th position), the object attribute information, the environment information, the channel information, the content information, the accumulated channel information and the accumulated content information corresponding to the any historical recommended content can be respectively represented as: and The basis information corresponding to the any historical recommended content includes at least one of and The to-be-processed channel information corresponding to the any historical recommended content includes at least one of and The to-be-processed content information corresponding to the any historical recommended content includes at least one of and The basis information and the to-be-processed channel information corresponding to the any historical recommended content are used to obtain the to-be-processed channel feature corresponding to the any historical recommended content, and the obtaining process is described in step b; and the basis information and the to-be-processed content information corresponding to the any historical recommended content are used to obtain the to-be-processed content feature corresponding to the any historical recommended content, and the obtaining process is described in step c.
[0100] Step b: performing fusion processing on the basis information and the to-be-processed channel information corresponding to the any historical recommended content to obtain the to-be-processed channel feature corresponding to the any historical recommended content.
[0101] The basic information corresponding to any historical recommended content and the to-be-processed channel information are original channel information, and by fusing the basic information corresponding to any historical recommended content and the to-be-processed channel information, the original channel information can be fully utilized. The feature obtained after the fusion processing is taken as the to-be-processed channel feature corresponding to the any historical recommended content.
[0102] The fusion processing process is not limited in the embodiments of the present application, as long as the fusion feature can be obtained by comprehensively considering various information. Exemplarily, the manner of fusing the basic information corresponding to any historical recommended content and the to-be-processed channel information to obtain the to-be-processed channel feature corresponding to the any historical recommended content can be: constructing a first feature matrix based on the basic information corresponding to any historical recommended content and the to-be-processed channel information; extracting a first parameter, a second parameter and a third parameter based on the first feature matrix; calculating a first header information by using the first parameter, the second parameter and the third reference; and calculating the to-be-processed channel feature corresponding to the any historical recommended content based on the first header information.
[0103] It should be noted that the number of the first parameter, the second parameter, the third parameter and the first header information is the same, and each of them can be one or more, and the dimensions of the first parameter, the second parameter and the third parameter are the same. Assuming that the first feature matrix constructed based on the basic information corresponding to any historical recommended content and the to-be-processed channel information is The process of extracting the first parameter, the second parameter and the third parameter based on the first feature matrix can be implemented based on formula 1, the process of calculating the first header information by using the first parameter, the second parameter and the third parameter can be implemented based on formula 2, and the process of calculating the to-be-processed channel feature corresponding to the any historical recommended content based on the first header information can be implemented based on formula 3:
[0104]
[0105]
[0106] wherein Q j represents the j(th) (an integer not less than 1) first parameter; K j represents the j(th) second parameter; V j represents the j(th) third parameter; head j represents the j(th) first header information; and represents the projection matrix of the j(th) first header information; d h represents the dimension of the first parameter; softmax represents a function; f i lrepresents the to-be-processed channel feature corresponding to the historical recommended content located at the i-th position in the historical recommended content sequence; MultiHead represents a multi-head self-attention feature interaction operation; concat represents a merging operation; w represents a weight vector O represents a weight vector (w O belongs to a d h dimensional Euclidean space, i.e., ).
[0107] Step c: fusing the basis information corresponding to any historical recommended content and the to-be-processed content information to obtain the to-be-processed content feature corresponding to any historical recommended content.
[0108] The basis information corresponding to any historical recommended content and the to-be-processed content information are original information in terms of content. By fusing the basis information corresponding to any historical recommended content and the to-be-processed content information, the original information in terms of content can be fully utilized. The feature obtained after the fusion is taken as the to-be-processed content feature corresponding to the historical recommended content.
[0109] Exemplarily, the manner of fusing the basis information corresponding to any historical recommended content and the to-be-processed content information to obtain the to-be-processed content feature corresponding to any historical recommended content can be as follows: constructing a second feature matrix based on the basis information corresponding to any historical recommended content and the to-be-processed content information; extracting a fourth parameter, a fifth parameter and a sixth parameter based on the second feature matrix; calculating second header information by using the fourth parameter, the fifth parameter and the sixth parameter; and calculating the to-be-processed content feature corresponding to any historical recommended content based on the second header information. The implementation process can be referred to step b, which will not be described herein again. After this process, the to-be-processed content feature corresponding to any historical recommended content can be represented as wherein, f i h represents the to-be-processed content feature corresponding to the historical recommended content located at the i-th position in the historical recommended content sequence; represents the second feature matrix; MultiHead represents a multi-head self-attention feature interaction operation.
[0110] Step d: arranging the to-be-processed channel features corresponding to respective historical recommended contents according to the arrangement order of the historical recommended contents in the historical recommended content sequence to obtain a to-be-processed channel feature sequence; and arranging the to-be-processed content features corresponding to respective historical recommended contents according to the arrangement order of the historical recommended contents in the historical recommended content sequence to obtain a to-be-processed content feature sequence.
[0111] According to the above step a and the above step b, the to-be-processed channel features corresponding to the respective historical recommendation contents can be obtained, and then the to-be-processed channel feature sequence is obtained based on the to-be-processed channel features corresponding to the respective historical recommendation contents. The process of obtaining the to-be-processed channel feature sequence based on the to-be-processed channel features corresponding to the respective historical recommendation contents is: arranging the to-be-processed channel features corresponding to the respective historical recommendation contents according to the arrangement order of the respective historical recommendation contents in the historical recommendation content sequence to obtain the to-be-processed channel feature sequence. That is, the to-be-processed channel feature at a certain arrangement position in the to-be-processed channel feature sequence corresponds to the historical recommendation content at the same arrangement position in the historical recommendation content sequence.
[0112] According to the above step a and the above step c, the to-be-processed content features corresponding to the respective historical recommendation contents can be obtained, and then the to-be-processed content feature sequence is obtained based on the to-be-processed content features corresponding to the respective historical recommendation contents. The process of obtaining the to-be-processed content feature sequence based on the to-be-processed content features corresponding to the respective historical recommendation contents is: arranging the to-be-processed content features corresponding to the respective historical recommendation contents according to the arrangement order of the respective historical recommendation contents in the historical recommendation content sequence to obtain the to-be-processed content feature sequence. That is, the to-be-processed content feature at a certain arrangement position in the to-be-processed content feature sequence corresponds to the historical recommendation content at the same arrangement position in the historical recommendation content sequence.
[0113] According to the above steps a to d, the to-be-processed channel feature sequence and the to-be-processed content feature sequence can be obtained, and step 3 is executed.
[0114] Step 3: calling the first processing model to process the to-be-processed channel feature sequence to obtain the channel preference feature corresponding to the target object; and calling the second processing model to process the to-be-processed content feature sequence to obtain the content preference feature corresponding to the target object.
[0115] The to-be-processed channel feature sequence is used to obtain the channel preference feature corresponding to the target object, and the first processing model is used to process the to-be-processed channel feature sequence. It should be noted that, since the to-be-processed channel feature sequence is composed of at least one to-be-processed channel feature arranged in sequence, in the process of processing the to-be-processed channel feature sequence, not only the respective to-be-processed channel features are considered, but also the association relationship between the respective to-be-processed channel features. The structure of the first processing model is not limited in the embodiments of the present application. For example, the first processing model can be a GRU (Gated Recurrent Unit) model. The process of calling the first processing model to process the to-be-processed channel feature sequence to obtain the channel preference feature corresponding to the target object can be implemented based on formula 4:
[0116]
[0117] wherein, denotes the channel preference feature corresponding to the target object; GRU l denotes the first processing model; denotes the sequence of content features to be processed.
[0118] The sequence of content features to be processed is used to obtain the content preference feature corresponding to the target object, and the second processing model is used to process the sequence of content features to be processed. It should be noted that, since the sequence of content features to be processed is composed of at least one content feature to be processed arranged in sequence, in the process of processing the sequence of content features to be processed, not only the content features to be processed are considered, but also the association relationship between the content features to be processed is considered. The structure of the second processing model is not limited by the embodiments of the present application. Illustratively, the second processing model can also be a GRU (Gated Recurrent Unit, Gated Recurrent Unit) model. The process of calling the second processing model to process the sequence of content features to be processed to obtain the content preference feature corresponding to the target object can be implemented based on formula 5:
[0119]
[0120] wherein, denotes the content preference feature corresponding to the target object; GRU h denotes the second processing model; denotes the sequence of content features to be processed.
[0121] It should be noted that, when the structures of the first processing model and the second processing model are the same, the parameters of the first processing model and the second processing model can be the same or different, and the embodiments of the present application do not limit this.
[0122] The above steps 1 to 3 introduce the process of obtaining the channel preference feature and the content preference feature corresponding to the target object, and the process of obtaining the candidate content set corresponding to the target object is introduced as follows:
[0123] The candidate content set includes at least one candidate content, and the process of obtaining the candidate content set is the process of obtaining each candidate content. In a possible implementation, the process of obtaining the candidate content set corresponding to the target object can include the following steps. Based on the historical behavior information of the target object, all contents in the content library are preliminarily screened, the contents obtained through the preliminary screening are grouped according to the source channels, and the content groups corresponding to each channel are obtained. In the content group corresponding to each channel, each content is sorted according to the matching degree with the target object. The reference number of contents with high ranking in each content group are taken as candidate contents. The set of candidate contents is taken as the candidate content set. It should be noted that the reference number can be set differently or uniformly for different content groups. For example, the reference number is uniformly set to 200 for different content groups, and the 200 contents with high ranking in each content group are taken as candidate contents. The embodiment of the present application does not limit the preliminary screening rule and the way of calculating the matching degree with the target object, and can be flexibly set according to the application scenario.
[0124] In step 302, the target channel sequence is obtained based on the channel preference feature and the candidate channel set corresponding to the candidate content set.
[0125] The candidate channel set includes the candidate channel corresponding to each candidate content in the candidate content set.
[0126] The channel preference feature is used to obtain the target channel sequence, and the target channel sequence is composed of at least one target channel arranged in sequence. The target channel in the target channel sequence is used to constrain the source of each content that needs to be recommended to the target object. The process of obtaining the target channel sequence can be regarded as a coarse-grained recommendation process, and this recommendation process only recommends channels. It should be noted that the channel recommended by this coarse-grained recommendation process is only used to constrain the next content recommendation process, and is not directly recommended to the target object. Through this process, the task of recommending content to the target object can be divided into two sub-tasks. The first sub-task is to recommend channels, and the second sub-task is to recommend contents under the constraint of the recommended channels. This way not only considers the content preference of the target object, but also considers the channel preference of the target object, which is beneficial to improve the content recommendation effect and improve the long-term interaction experience of the target object.
[0127] The candidate channel set is a set of candidate channels corresponding to each candidate content in the candidate content set. It should be noted that different candidate contents can correspond to the same candidate channel, and the candidate channels included in the candidate channel set are different candidate channels. Based on the channel preference feature and the candidate channel set corresponding to the candidate content set, the process of obtaining the target channel sequence can be regarded as a process of selecting candidate channels from the candidate channel set to form the target channel sequence based on the channel preference feature. In one possible implementation, based on the channel preference feature and the candidate channel set corresponding to the candidate content set, the process of obtaining the target channel sequence includes the following steps 3021 to step 3023:
[0128] Step 3021: calling the first target recommendation model to perform recommendation processing on the channel preference feature to obtain a channel recommendation result, the channel recommendation result being composed of at least one channel recommendation sub-result.
[0129] The first target recommendation model is a model pre-trained to output a channel recommendation result based on a channel preference feature. The channel recommendation result is composed of at least one channel recommendation sub-result, and any channel recommendation sub-result is used to indicate a virtual channel. The present embodiment does not limit the form of any channel recommendation sub-result, which can be represented by a feature vector, for example, based on which a virtual channel is indicated. It should be noted that the virtual channel here is relative to the real candidate channel in the candidate channel set, and the virtual channel can be consistent with any candidate channel or inconsistent with each candidate channel. In one possible implementation, the channel recommendation result is composed of at least one channel recommendation sub-result arranged in sequence.
[0130] Step 3022: for any channel recommendation sub-result, obtaining a target channel matching the any channel recommendation sub-result in the candidate channel set.
[0131] The channel recommendation sub-result is used to indicate a virtual channel, and the real channel should be actually recommended for the target object, so after obtaining the channel recommendation sub-result, a target channel matching the channel recommendation sub-result needs to be obtained in the candidate channel set.
[0132] In a possible implementation, for any channel recommendation sub-result, the process of obtaining a target channel matching the any channel recommendation sub-result from the candidate channel set includes: converting each candidate channel in the candidate channel set into the same form as the any channel recommendation sub-result; calculating the similarity between each candidate channel and the any channel recommendation sub-result based on the converted form; and taking the candidate channel with the highest similarity in the candidate channel set as the target channel matching the any channel recommendation sub-result. It should be noted that, since the form of the candidate channel in the candidate channel set can be different from the form of the channel recommendation sub-result, the form needs to be converted first to facilitate the calculation of the similarity. For example, when the form of the channel recommendation sub-result is a feature vector, each candidate channel needs to be converted into the form of the feature vector. The similarity between two vectors can be calculated in various ways, for example, the cosine similarity between two vectors is taken as the similarity between the two vectors.
[0133] For each channel recommendation sub-result in the channel recommendation result, the target channel matching the channel recommendation sub-result can be obtained through step 3022, and then step 3023 is performed.
[0134] In a possible implementation, the process of obtaining the channel recommendation result can be a loop process, each loop obtains one channel recommendation sub-result output by the first target recommendation model, and the channel recommendation sub-result obtained in each loop process is associated with the channel recommendation sub-result obtained in the previous loop process, and the channel recommendation result obtained in this way has better effect. In this case, step 3021 can be performed crosswise with step 3022, that is, each time a channel recommendation sub-result is obtained, the target channel matching the channel recommendation sub-result is obtained. In a possible implementation, the process of obtaining the channel recommendation result by calling the first target recommendation model to perform recommendation processing on the channel preference feature includes the following four steps.
[0135] Step 1: inputting the channel preference feature into the first target recommendation model to obtain a first channel recommendation sub-result output by the first target recommendation model.
[0136] The first target recommendation model outputs only the first channel recommendation sub-result based on the channel preference feature. In a possible implementation, the first target recommendation model includes a first target recommendation sub-model, and the first target recommendation model outputs the first channel recommendation sub-result by using the first target recommendation sub-model. The structure of the first target recommendation sub-model is not limited in the embodiments of the present application, for example, the first target recommendation sub-model can be a fully connected layer. The process of outputting the first channel recommendation sub-result by using the first target recommendation sub-model can be implemented based on formula 6.
[0137]
[0138] wherein, represents the first channel recommendation sub-result, may be a vector; tanh represents an activation function; represents the weight of the first target recommendation sub-model; represents the bias of the first target recommendation sub-model; represents the channel preference feature.
[0139] Step 2, based on the first channel recommendation sub-result, an updated channel preference feature is obtained, and the updated channel preference feature is input into the first target recommendation model to obtain a second channel recommendation sub-result output by the first target recommendation model.
[0140] In a possible implementation manner, the process of obtaining the updated channel preference feature based on the first channel recommendation sub-result is as follows: a target channel matching the first channel recommendation sub-result is obtained from the candidate channel set; a to-be-processed channel feature corresponding to the target channel is obtained, the to-be-processed channel feature is added after the last to-be-processed channel feature in the existing to-be-processed channel feature sequence to obtain an updated to-be-processed channel feature sequence; the first processing model is called to process the updated to-be-processed channel feature sequence to obtain the updated channel preference feature. The process of obtaining the to-be-processed channel feature corresponding to the target channel can be referred to step 301, which will not be described herein.
[0141] After obtaining the updated channel preference feature, the updated channel preference feature is input into the first target recommendation model, and the channel recommendation sub-result output by the first target recommendation model is taken as the second channel recommendation sub-result.
[0142] Step 3, based on the second channel recommendation sub-result, a channel recommendation sub-result is continuously obtained.
[0143] The process of continuously obtaining a channel recommendation sub-result based on the second channel recommendation sub-result is a loop process, and each loop process obtains a channel recommendation sub-result according to the manner of step 2. After obtaining a channel recommendation sub-result, it is determined whether the number of channel recommendation sub-results is less than the number threshold. When the number of channel recommendation sub-results is less than the number threshold, the next channel recommendation sub-result is continuously obtained, until the number of channel recommendation sub-results is equal to the number threshold, and step 4 is executed.
[0144] The quantity threshold is used to limit the maximum number of channel recommendation sub-results output by the first target recommendation model. The quantity threshold can be set according to experience or adjusted flexibly according to application scenarios, and the embodiments of the present application do not limit this. For example, the quantity threshold can be set to 10. It should be noted that since the target channels in the target channel sequence are matched one by one with the channel recommendation sub-results, the number of target channels in the target channel sequence is the same as the number of channel recommendation sub-results. Therefore, the quantity threshold is also used to limit the number of target channels in the target channel sequence.
[0145] It should be noted that as the number of obtained channel recommendation sub-results increases, the number of to-be-processed channel features in the to-be-processed channel feature sequence used to obtain the updated channel features also increases. For example, for the process of obtaining the tthchannel recommendation sub-result, the to-be-processed channel feature sequence can be represented as wherein, represents the to-be-processed channel feature sequence required to obtain the tth(t is an integer not less than 1) channel recommendation sub-result; m (m is an integer not less than 1) represents the number of historical recommended contents; (t-1) represents the number of channel recommendation sub-results that have been obtained; represents the to-be-processed channel feature obtained based on the (t-1)thchannel recommendation sub-result, which is located at the (m+t-1)tharrangement position in the to-be-processed channel feature sequence.
[0146] Step 4, in response to the number of channel recommendation sub-results being equal to the quantity threshold, arranging each channel recommendation sub-result arranged in sequence according to the acquisition order as the channel recommendation result.
[0147] When the number of channel recommendation sub-results is equal to the quantity threshold, it means that the required number of channel recommendation sub-results has been obtained. At this time, each channel recommendation sub-result arranged in sequence according to the acquisition order is arranged as the channel recommendation result. In this way, the channel recommendation result is obtained.
[0148] In the process of obtaining the channel recommendation result based on the above steps 1 to 4, each time a channel recommendation sub-result is obtained, the target channel matched with the channel recommendation sub-result is obtained. After all channel recommendation sub-results are obtained, the target channels matched with each channel recommendation sub-result are obtained, and then step 3023 is executed. It should be noted that the process of obtaining the channel recommendation result based on the above steps 1 to 4 is only an exemplary description of the case where the quantity threshold is greater than 2. In the case where the quantity threshold is 1, the channel recommendation result is obtained based on one channel recommendation sub-result obtained based on step 1. In the case where the quantity threshold is 2, the channel recommendation result is obtained based on two channel recommendation sub-results obtained based on steps 1 and 2.
[0149] Step 3023: Obtain a target channel sequence based on the target channels respectively matched with the channel recommendation sub-results.
[0150] After obtaining the target channels respectively matched with the channel prediction sub-results, a target channel sequence is obtained based on the target channels respectively matched with the channel recommendation sub-results. The target channel sequence is a sequence of channels corresponding to the content that needs to be recommended to the target object.
[0151] In a possible implementation, the manner of obtaining the target channel sequence based on the target channels respectively matched with the channel recommendation sub-results is as follows: the target channels respectively matched with the channel recommendation sub-results are arranged according to the arrangement order of the channel recommendation sub-results in the channel recommendation result, to obtain the target channel sequence. After obtaining the target channel sequence based on this manner, the target channel located at a certain arrangement position in the target channel sequence is matched with the channel recommendation sub-result located at the same arrangement position in the channel recommendation result.
[0152] In step 303, a target content sequence corresponding to the target channel sequence is obtained based on the content preference feature and the candidate content set.
[0153] The content preference feature is used to represent the preference of the target object in terms of content, the target channels in the target channel sequence are used to constrain the sources of the content that needs to be recommended to the target object, and the candidate content set includes candidate content available for recommendation. Based on the content preference feature and the candidate content set, a target content sequence corresponding to the target channel sequence is obtained, and the target content in the target content sequence is the content that needs to be recommended to the target object.
[0154] In a possible implementation, the process of obtaining the target content sequence corresponding to the target channel sequence based on the content preference feature and the candidate content set includes the following steps 3031 to 3033.
[0155] Step 3031: A second target recommendation model is invoked to perform recommendation processing on the content preference feature, to obtain a content recommendation result composed of at least one content recommendation sub-result.
[0156] The second target recommendation model is a pre-trained model for outputting a content recommendation result based on a content preference feature. The content recommendation result is composed of at least one content recommendation sub-result, and any content recommendation sub-result is used to indicate a virtual content. Embodiments of the present application do not limit the form of any content recommendation sub-result, and any content recommendation sub-result can be represented by a feature vector, for example, which indicates a virtual content based on the feature vector. It should be noted that the virtual content herein is relative to the real candidate content in the candidate content set, and the virtual content can be consistent with any candidate content, or can be inconsistent with each candidate content. In one possible implementation, the content recommendation result is composed of at least one channel recommendation sub-result arranged in sequence.
[0157] Step 3032: For any content recommendation sub-result, a target content matching the any content recommendation sub-result is obtained from the target candidate content set.
[0158] The content recommendation sub-result is used to indicate a virtual content, and the real content should be actually recommended for the target object, so after obtaining the content recommendation sub-result, a target content matching the content recommendation sub-result needs to be obtained from the candidate content set.
[0159] For any content recommendation sub-result, a target content matching the any content recommendation sub-result is obtained from the target candidate content set. The target candidate content set is composed of candidate contents in the candidate content set that meet a condition, and the candidate content that meets the condition includes a candidate content whose corresponding candidate channel is a specified channel. The specified channel includes a target channel in the target channel sequence and a target channel in the content recommendation result. That is, according to the constraint of the target channel in the target channel sequence, the target candidate content set is determined, and then the target content matching the any content recommendation sub-result is obtained from the target candidate content set.
[0160] In a possible implementation, for any content recommendation sub-result, the process of obtaining a target content matching the any content recommendation sub-result in the target candidate content set is: converting each candidate content in the target candidate content set into the same form as the any content recommendation sub-result; calculating the similarity between each candidate content and the any content recommendation sub-result based on the converted form; and taking the candidate content with the highest similarity in the target candidate content set as the target content matching the any content recommendation sub-result. It should be noted that, since the form of the candidate content in the target candidate content set can be different from the form of the content recommendation sub-result, the form needs to be converted first to facilitate the calculation of the similarity. For example, when the form of the content recommendation sub-result is a feature vector, each candidate content needs to be converted into the form of the feature vector. The similarity between two vectors can be calculated in various ways, for example, the cosine similarity between two vectors is taken as the similarity between the two vectors.
[0161] For each channel recommendation sub-result in the content recommendation result, the matched target content can be obtained through step 3032, and then step 3033 is performed.
[0162] In a possible implementation, the process of obtaining the content recommendation result can be a loop process, each loop obtains a content recommendation sub-result output by the second target recommendation model, and the content recommendation sub-result obtained in each loop process is associated with the content recommendation sub-result obtained in the previous loop process, and the content recommendation result obtained in this way has better effect. In this case, step 3031 can be performed crosswise with step 3032, that is, after obtaining a content recommendation sub-result, the target content matching the content recommendation sub-result is obtained. In a possible implementation, the process of obtaining the content recommendation result by calling the second target recommendation model to perform recommendation processing on the content preference feature includes the following four steps.
[0163] Step 1: input the content preference feature into the second target recommendation model to obtain a first content recommendation sub-result output by the second target recommendation model.
[0164] The second target recommendation model outputs only the first content recommendation sub-result based on the content preference feature. In a possible implementation, the second target recommendation model includes a second target recommendation sub-model, and the second target recommendation model outputs the first content recommendation sub-result by using the second target recommendation sub-model. The second target recommendation sub-model is not limited in the embodiments of the present application. For example, the second target recommendation sub-model can be a fully connected layer. It should be noted that when the structures of the first target recommendation sub-model and the second target recommendation sub-model are the same, the parameters of the first target recommendation sub-model and the second target recommendation sub-model are different because the first target recommendation sub-model and the second target recommendation sub-model are used to recommend different aspects of results. In a possible implementation, the process of outputting the first content recommendation sub-result by using the second target recommendation sub-model can be implemented based on formula 7:
[0165]
[0166] wherein, represents the first content recommendation sub-result, may be a vector; and tanh represents an activation function; represents a weight of the second target recommendation sub-model; represents a bias of the second target recommendation sub-model; represents the content preference feature.
[0167] Step 2: Based on the first content recommendation sub-result, an updated content preference feature is obtained, and the updated content preference feature is input into the second target recommendation model to obtain a second content recommendation sub-result output by the second target recommendation model.
[0168] In a possible implementation, the process of obtaining the updated content preference feature based on the first content recommendation sub-result includes: obtaining a target content matching the first content recommendation sub-result in the candidate content set; obtaining a to-be-processed content feature corresponding to the target content, adding the to-be-processed content feature to the last to-be-processed content feature in the existing to-be-processed content feature sequence to obtain an updated to-be-processed content feature sequence; and calling the second processing model to process the updated to-be-processed content feature sequence to obtain the updated content preference feature. The process of obtaining the to-be-processed content feature corresponding to the target content can be referred to step 301, which is not described herein again.
[0169] After obtaining the updated content preference feature, the updated content preference feature is input into the second target recommendation model, and a content recommendation sub-result output by the second target recommendation model is taken as the second content recommendation sub-result.
[0170] Step 3: Based on the second content recommendation sub-result, a content recommendation sub-result is continuously obtained.
[0171] The process of obtaining content recommendation sub-results based on the second content recommendation sub-result is a loop process, and each loop process obtains a content recommendation sub-result according to the manner of step 2. After obtaining each content recommendation sub-result, it is determined whether the number of content recommendation sub-results is less than the number threshold. When the number of content recommendation sub-results is less than the number threshold, the next content recommendation sub-result is obtained until the number of content recommendation sub-results is equal to the number threshold, and step 4 is executed.
[0172] The number threshold is used to limit the maximum number of content recommendation sub-results output by the second target recommendation model, and the number threshold is the same as the number threshold used to limit the maximum number of channel recommendation sub-results output by the first target recommendation model. It should be noted that since the target contents in the target content sequence are matched with the content recommendation sub-results one by one, the number of target contents in the target content sequence is the same as the number of content recommendation sub-results, and the number threshold is also used to limit the number of target contents in the target content sequence.
[0173] It should be noted that as the number of obtained content recommendation sub-results increases, the number of to-be-processed content features in the to-be-processed content feature sequence used to obtain the updated content features also increases. Exemplarily, for the process of obtaining the tthcontent recommendation sub-result, the to-be-processed content feature sequence can be represented as wherein, represents the to-be-processed content feature sequence required for obtaining the tthcontent recommendation sub-result (t is an integer not less than 1); m (m is an integer not less than 1) represents the number of historical recommended contents; (t-1) represents the number of obtained content recommendation sub-results; represents the to-be-processed content feature based on the (t-1)thcontent recommendation sub-result, which is located at the (m+t-1)tharrangement position in the to-be-processed content feature sequence.
[0174] Step 4: In response to the number of content recommendation sub-results being equal to the number threshold, each content recommendation sub-result arranged in sequence according to the obtaining order is taken as the content recommendation result.
[0175] When the number of content recommendation sub-results is equal to the number threshold, it means that the required number of content recommendation sub-results has been obtained, and at this time, each content recommendation sub-result arranged in sequence according to the obtaining order is taken as the content recommendation result. In this way, the content recommendation result is obtained.
[0176] In the process of obtaining the content recommendation result based on the steps 1 to 4, each time a content recommendation sub-result is obtained, the target content matched by the content recommendation sub-result is obtained, and after all the content recommendation sub-results are obtained, the target content matched by each content recommendation sub-result is obtained, and then the step 3033 is performed. It should be noted that the process of obtaining the content recommendation result based on the steps 1 to 4 is only an exemplary description of the case where the number threshold is greater than 2. In the case where the number threshold is 1, the content recommendation result is obtained based on one content recommendation sub-result obtained in the step 1. In the case where the number threshold is 2, the content recommendation result is obtained based on two content recommendation sub-results obtained in the steps 1 and 2.
[0177] The step 3033: obtaining the target content sequence corresponding to the target channel sequence based on the target content matched by each content recommendation sub-result.
[0178] After the target content matched by each content prediction sub-result is obtained, the target content sequence corresponding to the target channel sequence is obtained based on the target content matched by each content recommendation sub-result. The target content sequence is a sequence of the content constituting the final content to be recommended to the target object.
[0179] In a possible implementation manner, the manner of obtaining the target content sequence corresponding to the target channel sequence based on the target content matched by each content recommendation sub-result is that the target content matched by each content recommendation sub-result is arranged according to the arrangement order of each content recommendation sub-result in the content recommendation result to obtain the target content sequence. After the target content sequence is obtained based on this manner, the target content located at a certain arrangement position in the target content sequence is matched with the content recommendation sub-result located at the same arrangement position in the content recommendation result.
[0180] It should be noted that before the content recommendation task is implemented by using the first target recommendation model and the second target recommendation model, the target hierarchical recommendation model including the first target recommendation model and the second target recommendation model needs to be obtained by training. The process of obtaining the target hierarchical recommendation model by training is described in detail in the embodiment shown in Figure 6 and will not be described here in detail.
[0181] In the step 304, the target content sequence is recommended to the target object.
[0182] After the target content sequence is obtained, the target content sequence is recommended to the target object for browsing and viewing by the target object. It should be noted that since the target content sequence is constituted by at least one target content arranged in sequence, recommending the target content sequence to the target object means recommending each target content constituting the target content sequence to the target object.
[0183] In a possible implementation, the manner in which the target content sequence is recommended to the target object is that the target content sequence is recommended to the target object based on a recommendation content acquisition request of the target object. The embodiments of the present application do not limit the manner in which the recommendation content acquisition request is acquired. For example, the manner in which the recommendation content acquisition request is acquired can be based on a downward sliding gesture of the target object, or the recommendation content acquisition request is acquired automatically based on a successful login instruction of the target object.
[0184] In a possible implementation, the process in which the target content sequence is recommended to the target object is that the target contents are page laid according to the arrangement order in the target content sequence, a recommendation page is obtained, and the recommendation page is displayed on the terminal screen. It should be noted that the embodiments of the present application do not limit the page layout rule, as long as the target contents located at the front position in the target content sequence are still located at the front position in the laid page. In addition, the size of the recommendation page can be greater than the screen visible area. At this time, the process in which the recommendation page is displayed on the terminal screen is that the target area of the recommendation page is displayed in the screen visible area, and other areas of the recommendation page are displayed according to the sliding instruction of the target object. The target area of the recommendation page can be the upper area of the recommendation page, or the upper left corner area of the recommendation page, and the embodiments of the present application do not limit this.
[0185] For example, the process in which the recommendation page is displayed on the terminal screen can be as shown in Figure 4 For example, the process in which the recommendation page is displayed on the terminal screen can be as shown in
[0186] After the target content sequence is recommended to the target object, the feedback of the target object can be collected, for example, the clicking condition of the target object on the target contents in the target content sequence, the reading time length, and the like, so as to further adjust the recommendation model according to the feedback of the target object in the subsequent process, to further improve the recommendation effect of the model.
[0187] For example, the process in which the target content sequence is acquired can be as shown in Figure 5 For example, the process in which the target content sequence is acquired can be as shown in Figure 5In the process of acquiring the target content d located at the t-th position in the target content sequence, the channel preference feature t and the content preference feature t used for acquiring the target content d t are first acquired. The channel preference feature is input into the first target recommendation model to obtain a channel recommendation sub-result located at the t-th position output by the first target recommendation model Further, a target channel matching the channel recommendation sub-result is acquired from the candidate channel set The content preference feature is input into the second target recommendation model to obtain a content recommendation sub-result located at the t-th position output by the second target recommendation model Further, based on the constraint of the target channel , a target content d t matching the content recommendation sub-result is acquired from the candidate content set After the target content sequence is recommended to the target object, the recommendation system (environment) can collect feedback of the target object on each target content, and generate feedback information corresponding to each target channel and each target content according to the feedback of the target object on each target content, which is used for subsequent adjustment of the first target recommendation model and the second target recommendation model. For example, according to the feedback of the target object on the target content located at the t-th position, feedback information corresponding to the target channel located at the t-th position and feedback information corresponding to the target content located at the t-th position are generated. The feedback information is fed back to the first target recommendation model and the second target recommendation model.
[0188] It should be noted that the application examples of the present application do not limit the application scenarios of content recommendation. Exemplarily, the application scenario can be a feed stream recommendation scenario. The feed stream is an information stream that is continuously updated and presented to the interactive object. Feed stream recommendation is a kind of content recommendation of aggregated information, and through the feed stream, dynamic and real-time dissemination can be given to the subscriber, which is an effective way for the interactive object to obtain the information stream. Of course, the application examples of the present application can not only be applied to the comprehensive recommendation of the feed stream, but also can be applied to other recommendation scenarios containing heterogeneous content. At the same time, the main idea here is to use the hierarchical recommendation method to split the comprehensive recommendation problem containing heterogeneous content into two parts, and recommend the channel by using the first target recommendation model and recommend the content by using the second target recommendation model.
[0189] The comprehensive recommendation faces the following challenges: 1. Heterogeneous content from different channels usually has different features and ranking strategies, which makes the ranking scores of different content incomparable. 2. The interactive object not only has personalized preferences for different content, but also has personalized preferences for different channels. This coarse-grained preference for channels is also meaningful to the recommendation system. 3. The online comprehensive recommendation in the industry pays great attention to the robustness and stability of the system. A small fluctuation in one channel can have a great impact on the performance of the entire recommendation system.
[0190] At present, most comprehensive recommendations use CTR (Click-Through-Rate, click-through rate) to guide the common ranking of heterogeneous content, or make recommendations based on rules. However, CTR guidance will homogenize channels and content, affecting the long-term experience of the interactive object, and using experience to set rules will inevitably reduce the personalization of recommendations. In the embodiments of the present application, the comprehensive recommendation is divided into two sub-tasks to recommend channels and content respectively: the first target recommendation model serves as a channel selector to generate a personalized channel sequence; and the first target recommendation model serves as a content recommender to recommend corresponding content in a specific channel to generate a final target content sequence. By efficiently and flexibly capturing the personalized preferences of the interactive object for channels and content, the above problems are solved, and the overall effect of the comprehensive recommendation is optimized.
[0191] In the embodiments of the present application, the target channel sequence is first obtained based on channel preference features, and then the target content sequence corresponding to the target channel sequence is obtained based on content preference features, and the target content sequence is recommended to the target object. In this content recommendation process, the channel preference features reflect the information of the channel, the content preference features reflect the information of the content, the content recommendation process considers both the information of the channel and the information of the content, and the information considered is more comprehensive, which is conducive to improving the effect of content recommendation, the click rate of the recommended content is higher, and the interactive experience of the user is better.
[0192] Based on the implementation environment shown in Figure 2 The embodiments of the present application provide a content recommendation method. The method is applied to the terminal 21 as an example. As shown in Figure 6 The content recommendation method provided by the embodiments of the present application can include the following steps:
[0193] In step 601, a target hierarchical recommendation model is obtained, and the target hierarchical recommendation model includes a first target recommendation model and a second target recommendation model.
[0194] The target hierarchical recommendation model refers to a trained model used to implement content recommendation. The target hierarchical recommendation model can be trained by a terminal or a server, and embodiments of the present application do not limit this. For the case where the target hierarchical recommendation model is trained by the terminal, the terminal can directly obtain the target hierarchical recommendation model; for the case where the target hierarchical recommendation model is trained by the server, the terminal obtains the target hierarchical recommendation model from the server. Embodiments of the present application take the case where the target hierarchical recommendation model is trained by the terminal as an example for description.
[0195] Before obtaining the target hierarchical recommendation model, the target hierarchical recommendation model needs to be trained. In a possible implementation manner, the process of training the target hierarchical recommendation model includes the following steps 6011 and 6012:
[0196] Step 6011: Obtain a training sample set, the training sample set including at least one training sample, any training sample including sample channel features, sample content features, and feedback information corresponding to a sample recommendation content sequence.
[0197] The training sample is obtained based on historical recommendation content of multiple interaction objects. For any interaction object, in a content recommendation scenario, the interaction object sends one or more content recommendation requests in an application program or a webpage capable of content recommendation, and the recommendation system recommends a content sequence for each content recommendation request, and each content recommendation sequence includes one or more historical recommendation contents. All content sequences recommended for one or more content recommendation requests of an interaction object constitute a session. Embodiments of the present application do not limit the manner in which the interaction object sends the content recommendation request. For example, the interaction object sends the content recommendation request through a downward swipe gesture on the screen. Embodiments of the present application obtain the training sample based on the historical actually recommended content sequence. Embodiments of the present application do not limit the number of interaction objects, the number of sessions of the interaction objects, the number of recommendation instances extracted in the session, and the number of clicks involved in the recommendation instance when obtaining the training sample. For example, the number of interaction objects is 22.5 million, the number of sessions of the interaction objects is 141 million, the number of recommendation instances extracted in the session is 3.8 billion, and 355 million clicks are involved in the 3.8 billion recommendation instances.
[0198] Each training sample includes sample channel features, sample content features, and feedback information corresponding to a sample recommended content sequence. The acquisition process of the sample channel features and the sample content features in each training sample can refer to the process of acquiring the channel preference features and the content preference features of the target object in step 301, which will not be repeated here. The sample recommended content sequence refers to the actual recommended content sequence based on the sample channel features and the sample content features. The feedback information corresponding to the sample recommended content sequence includes, but is not limited to, actual operation information of each sample recommended content in the sample recommended content sequence by an interactive object after the sample recommended content sequence is recommended to the interactive object, and recommendation characteristic information of the sample recommended content itself, etc. The operation of the interactive object on each sample recommended content includes, but is not limited to, a click operation, a reading operation, etc.
[0199] Step 6012: Based on the sample channel features, the sample content features, and the feedback information corresponding to the sample recommended content sequence in the training sample, the first initial recommendation model and the second initial recommendation model in the initial hierarchical recommendation model are trained to obtain a target hierarchical recommendation model.
[0200] After obtaining the training sample set, the training samples in the training sample set are used to train the initial hierarchical recommendation model to obtain a target hierarchical recommendation model. In a possible implementation manner, in the process of training the target hierarchical recommendation model, the logic of a reinforcement learning algorithm is used to update the model parameters. The application does not limit which logic of the reinforcement learning algorithm is used. Exemplarily, the logic of a DDPG (Deep Deterministic Policy Gradient, deep deterministic policy gradient) algorithm, the logic of a DQN (Deep Q-Learning Network, deep Q-learning network) algorithm, and the logic of an A3C (Asynchronous Advantage Actor-Critic, asynchronous advantage actor-critic) algorithm, etc. can be used.
[0201] In a possible implementation manner, the first initial recommendation model includes a first initial recommendation submodel and a first initial evaluation submodel, and the second initial recommendation model includes a second initial recommendation submodel and a second initial evaluation submodel. Referring to Figure 7 Based on the sample channel features, the sample content features, and the feedback information corresponding to the sample recommended content sequence in the training sample, the method for training the first initial recommendation model and the second initial recommendation model in the initial hierarchical recommendation model includes the following steps 60121 to 60125:
[0202] Step 60121: Based on the feedback information corresponding to the sample recommended content sequence in the training sample, a first set of enhanced values and a second set of enhanced values are obtained.
[0203] The training sample herein refers to a training sample required for training the initial hierarchical recommendation model once, and the number of training samples can be one or more, which is not limited by the embodiments of the present application. For the case of multiple training samples, the related data obtained in steps 60121 to 60123 are obtained for each training sample respectively. The embodiments of the present application introduce the process of obtaining related data in steps 60121 to 60123, taking one training sample required for training the initial hierarchical recommendation model once as an example.
[0204] The first enhancement value set refers to a set of first enhancement values, the first enhancement value refers to a channel aspect enhancement value, and the first enhancement value set is used to guide the update of the first initial recommendation model; the second enhancement value set refers to a set of second enhancement values, the second enhancement value refers to a content aspect enhancement value, and the second enhancement value set is used to guide the update of the second initial recommendation model.
[0205] In a possible implementation manner, based on the feedback information corresponding to the sample recommendation content sequence in the training sample, the process of obtaining the first enhancement value set and the second enhancement value set includes the following steps A to D:
[0206] Step A: Based on the feedback information in the training sample, obtain at least one of the reading time information, the diversity information and the novelty information of each sample recommendation content in the sample recommendation content sequence, and the click information of each sample recommendation content.
[0207] The feedback information includes the click information of each sample recommendation content after the sample recommendation content sequence is recommended to a certain interaction object. According to the click information, the click information of each sample recommendation content in the sample recommendation content sequence can be obtained. The click information is used to indicate whether the sample recommendation content is clicked.
[0208] In addition to the click information of each sample recommendation content of the interaction object, the feedback information can also include the reading information of each sample recommendation content of the interaction object. According to the reading information, the reading time information of each sample recommendation content in the sample recommendation content sequence can be obtained. The reading time information is used to indicate the reading time of the sample recommendation content. It should be noted that the reading in the embodiments of the present application can refer to reading news, articles and the like, or can refer to watching video content.
[0209] The diversity information is used to evaluate the diversity of the sample recommended content, and the novelty information is used to evaluate the novelty of the sample recommended content. In a possible implementation manner, in addition to the click information of the interaction object on each sample recommended content, the feedback information can further include information of content labels corresponding to each sample recommended content in the sample recommended content sequence, and the information of the content labels is used to indicate which content labels correspond to the sample recommended content. For any sample recommended content, a manner of obtaining the diversity information of the sample recommended content can be as follows: content labels corresponding to each sample recommended content arranged before the sample recommended content in the sample content recommended sequence are counted, the content labels corresponding to the sample recommended content and the previous content labels are compared, an increment of repeated content labels in the content labels corresponding to the sample recommended content is calculated, and the increment of the repeated content labels is taken as the diversity information of the sample recommended content.
[0210] In a possible implementation manner, in addition to the click information of the interaction object on each sample recommended content, the feedback information can further include a user interest label. For any sample recommended content, a manner of obtaining the novelty information of the sample recommended content can be as follows: the content labels corresponding to the sample recommended content are compared with the user interest label, an increment of new content labels in the content labels corresponding to the sample recommended content is calculated, and the increment of the new content labels is taken as the novelty information of the sample recommended content.
[0211] After obtaining at least one of the reading duration information, the diversity information and the novelty information of any sample recommended content and the click information of the sample recommended content, steps B and C are performed.
[0212] Step B: based on the click information of any sample recommended content, a first enhancement value corresponding to the sample recommended content is obtained.
[0213] The first enhancement value refers to a channel aspect enhancement value, and the click information of any sample recommended content can be regarded as click information of a sample recommended channel corresponding to the sample recommended content. Then, the first enhancement value of the channel aspect corresponding to the sample recommended content can be obtained according to the click information of the sample recommended content.
[0214] In a possible implementation, the process of obtaining the first enhancement value corresponding to the any-sample recommended content based on the click information of the any-sample recommended content is: searching for a score corresponding to the click information of the any-sample recommended content, and taking the score corresponding to the any-sample recommended content as the first enhancement value corresponding to the any-sample recommended content. The correspondence between the click information and the score can be pre-set and stored, and the score corresponding to the click information of the any-sample recommended content is searched based on the correspondence between the click information and the score, and then the first enhancement value corresponding to the any-sample recommended content is obtained.
[0215] Step C: obtaining the second enhancement value corresponding to the any-sample recommended content based on at least one of the reading duration information, the diversity information and the novelty information of the any-sample recommended content and the click information of the any recommended content.
[0216] The second enhancement value corresponding to the any-sample recommended content is obtained based on all the information of the any recommended content obtained in step A. The second enhancement value refers to an enhancement value in terms of content.
[0217] In a possible implementation, for the case that all the information of the any recommended content obtained in step A includes the click information, the reading duration information, the diversity information and the novelty information of the any recommended content, the process of obtaining the second enhancement value corresponding to the any-sample recommended content is: obtaining the second enhancement value corresponding to the any-sample recommended content based on the click information, the reading duration information, the diversity information and the novelty information of the any-sample recommended content.
[0218] In a possible implementation, the process of obtaining the second enhancement value corresponding to the any-sample recommended content based on the click information, the reading duration information, the diversity information and the novelty information of the any-sample recommended content is: converting the click information of the any-sample recommended content into a click enhancement value corresponding to the any-sample recommended content; converting the reading duration information of the any-sample recommended content into a reading enhancement value corresponding to the any-sample recommended content; converting the diversity information of the any-sample recommended content into a diversity enhancement value corresponding to the any-sample recommended content; converting the novelty information of the any-sample recommended content into a novelty enhancement value corresponding to the any-sample recommended content; and determining the second enhancement value corresponding to the any-sample recommended content based on the click enhancement value, the reading enhancement value, the diversity enhancement value and the novelty enhancement value.
[0219] The click enhancement value is used to optimize the click rate of the content recommended by the model, the reading enhancement value is used to learn the real reading preference of the interactive object, the diversity enhancement value is used to measure diversity, and the novelty enhancement value is used to measure novelty. The diversity enhancement value and the novelty enhancement value are beneficial to improving the long-term experience of the interactive object.
[0220] The process of converting information into an enhancement value can refer to a process of looking up a corresponding score as an enhancement value. In one possible implementation, the process of determining the second enhancement value corresponding to the any sample recommended content based on the click enhancement value, the reading enhancement value, the diversity enhancement value, and the novelty enhancement value can be completed based on formula 8:
[0221]
[0222] wherein, represents the second enhancement value corresponding to the sample recommended content at the t-th position in the sequence of sample recommended contents, represents the i-th (i is an integer not less than 1 and not greater than 4) of the click enhancement value, the reading enhancement value, the diversity enhancement value, and the novelty enhancement value, represents the deviation of the i-th (i is an integer not less than 1 and not greater than 4) enhancement value, represents the weight of the i-th (i is an integer not less than 1 and not greater than 4) enhancement value.c t represents the sample recommended channel corresponding to the sample recommended content at the t-th position. That is, the weight of the i-th (i is an integer not less than 1 and not greater than 4) enhancement value is set based on the sample recommended channel corresponding to the sample recommended content. The set of wherein, represents the click enhancement value; represents the reading enhancement value; represents the diversity enhancement value; represents the novelty enhancement value.
[0223] According to the above steps A to C, the first enhancement value corresponding to each sample recommended content in the sequence of sample recommended contents, and the second enhancement value corresponding to each sample recommended content can be obtained, and then step D is executed.
[0224] Step D: taking the set of the first enhancement value corresponding to each sample recommended content as the first enhancement value set; and taking the set of the second enhancement value corresponding to each sample recommended content as the second enhancement value set.
[0225] After obtaining the first enhancement value corresponding to each sample recommended content, the set of the first enhancement value corresponding to each sample recommended content is taken as the first enhancement value set. Thus, the first enhancement value set is obtained. After obtaining the second enhancement value corresponding to each sample recommended content, the set of the second enhancement value corresponding to each sample recommended content is taken as the second enhancement value set. Thus, the second enhancement value set is obtained.
[0226] Step 60122: input the sample channel feature in the training sample into the first initial recommendation model to obtain an initial channel recommendation result output by a first initial recommendation submodel and a first evaluation value set for the initial channel recommendation result output by a first initial evaluation submodel.
[0227] The first initial recommendation model includes the first initial recommendation submodel and the first initial evaluation submodel. The first initial recommendation submodel is configured to output the initial channel recommendation result based on the sample channel feature, and the first initial evaluation submodel is configured to evaluate the initial channel recommendation result output by the first initial recommendation submodel and output the first evaluation value set for the initial channel recommendation result. The first evaluation value set is used to guide parameter updating of the first initial recommendation model.
[0228] It should be noted that the initial channel recommendation result is composed of at least one initial channel recommendation subresult arranged in sequence, and each initial channel recommendation subresult corresponds to a first evaluation value. Therefore, the first evaluation value set is a set of first evaluation values corresponding to each initial channel recommendation subresult.
[0229] In a possible implementation, a model structure of the first initial recommendation model is an Actor-Critic structure. Based on this, the first initial recommendation submodel in the first initial recommendation model is an Actor model, and the first initial evaluation submodel is a Critic model. The first initial evaluation submodel can be a fully connected layer.
[0230] A calculation formula of the first theoretical evaluation value used to evaluate the initial channel recommendation subresult is shown in formula 9. In an actual Critic model, the first theoretical evaluation value is predicted by using formula 10. The first evaluation value involved in the embodiments of the present application refers to the first evaluation value predicted by the first initial evaluation submodel.
[0231]
[0232] wherein, represents the first theoretical evaluation value used to evaluate the t th< initial channel recommendation subresult; represents the first enhancement value corresponding to the t th< sample recommendation content in the sample recommendation content sequence; γ represents a discount factor; represents the first theoretical evaluation value used to evaluate the (t+1) th< initial channel recommendation subresult; represents the channel feature corresponding to the t th< initial channel recommendation subresult; represents the t th< initial channel recommendation subresult. represents the first evaluation value corresponding to the t th< initial channel recommendation subresult output by the first initial evaluation submodel; ReLU represents a Rectified Linear Unit. and This represents the weight of the first initial evaluation sub-model. This indicates the bias of the first initial evaluation sub-model. and By inputting the first initial evaluation sub-model, you can obtain the first evaluation value corresponding to the t-th initial channel recommendation sub-result output by the first initial evaluation sub-model.
[0233] After obtaining the first evaluation value corresponding to each initial channel recommendation sub-result, the set of the first evaluation values corresponding to each initial channel recommendation sub-result is taken as the first evaluation value set.
[0234] Step 60123: Input the sample content features in the training samples into the second initial recommendation model to obtain the initial content recommendation result output by the second initial recommendation sub-model and the second evaluation value set output by the second initial evaluation sub-model for the initial content recommendation result.
[0235] The second initial recommendation model comprises a second initial recommendation sub-model and a second initial evaluation sub-model. The second initial recommendation sub-model outputs initial content recommendation results based on sample content features. The second initial evaluation sub-model evaluates the initial content recommendation results output by the second initial recommendation sub-model, outputting a second set of evaluation values for the initial content recommendation results. This second set of evaluation values guides the parameter updates of the second initial recommendation model.
[0236] It should be noted that the initial content recommendation result is composed of at least one initial content recommendation sub-result arranged in sequence. Each initial content recommendation sub-result corresponds to a second evaluation value. Therefore, the set of second evaluation values is the set of second evaluation values corresponding to each initial content recommendation sub-result.
[0237] In one possible implementation, the second initial recommendation model also uses an Actor-Critic structure. Based on this, the second initial recommendation sub-model is an Actor model, and the second initial evaluation sub-model is a Critic model. The second initial evaluation sub-model can be a fully connected layer.
[0238] The formula for calculating the second theoretical evaluation value used to assess the initial content recommendation sub-result is shown in Formula 11. In the actual Critic model, Formula 12 is used to predict the second theoretical evaluation value. In the embodiments of this application, the second evaluation value refers to the second evaluation value predicted by the second initial evaluation sub-model.
[0239]
[0240] in, denotes a second theoretical evaluation value for evaluating the tth initial content recommendation sub-result; denotes a second enhancement value corresponding to the tth sample recommended content in the sample recommended content sequence; γ denotes a discount factor; denotes a second theoretical evaluation value for evaluating the (t+1)th initial content recommendation sub-result; denotes a content feature corresponding to the tth initial content recommendation sub-result; denotes the tth initial content recommendation sub-result. denotes a second evaluation value of the tth initial content recommendation sub-result output by the second initial evaluation sub-model; ReLU denotes a Rectified Linear Unit (ReLU); and denotes a weight of the second initial evaluation sub-model, denotes a bias of the second initial evaluation sub-model. The and The second initial evaluation sub-model is input, and a second evaluation value corresponding to the tth initial content recommendation sub-result output by the second initial evaluation sub-model is obtained.
[0241] After obtaining the second evaluation value corresponding to each initial content recommendation sub-result respectively, a set of the second evaluation value corresponding to each initial content recommendation sub-result respectively is taken as a second evaluation value set.
[0242] Step 60124: updating the parameters of the first initial recommendation sub-model based on the first evaluation value set; and updating the parameters of the second initial recommendation sub-model based on the second evaluation value set.
[0243] It should be noted that each training sample corresponds to a first evaluation value set, that is, the number of first evaluation value sets is the same as the number of training samples used for one training of the initial hierarchical recommendation model. For the case where the number of training samples used for one training of the initial hierarchical recommendation model is multiple, the parameters of the first initial recommendation sub-model are updated based on the multiple first evaluation value sets corresponding to each training sample, and the parameters of the second initial recommendation sub-model are updated based on the multiple second evaluation value sets corresponding to each training sample. The embodiments of the present application take the case where the number of training samples used for one training of the initial hierarchical recommendation model is one as an example for description.
[0244] In a possible implementation, the process of updating the parameters of the first initial recommendation sub-model based on the first set of evaluation values is: calculating a first update gradient based on each first evaluation value in the first set of evaluation values; and updating the parameters of the first initial recommendation sub-model in a direction of maximizing the first update gradient. In a possible implementation, the process of calculating the first update gradient based on each first evaluation value in the first set of evaluation values is: calculating a first target evaluation value based on each first evaluation value in the first set of evaluation values; and calculating the first update gradient based on the first target evaluation value.
[0245] In a possible implementation, the first target evaluation value can be calculated based on each first evaluation value in the first set of evaluation values in the following manner: setting a weight for each first evaluation value respectively, and taking a weighted average of the first evaluation values as the first target evaluation value. The first update gradient can be calculated based on the first target evaluation value according to Formula 13.
[0246]
[0247] wherein, the first update gradient is denoted by the random strategy adopted by the first initial recommendation sub-model when outputting a channel recommendation result is denoted by l the parameters of the first initial recommendation sub-model are denoted by l the set of channel features involved in the process of outputting the initial channel recommendation result is denoted by l the initial channel recommendation result is denoted by the first target evaluation value is denoted by
[0248] After obtaining the first update gradient, the parameters of the first initial recommendation sub-model are updated in a direction of maximizing the first update gradient, because the optimization direction is to make the evaluation value as large as possible.
[0249] For a case where the number of training samples used for training the initial hierarchical recommendation model once is multiple, the first update gradient refers to an average value of multiple first update gradients calculated based on multiple first sets of evaluation values corresponding to the multiple training samples.
[0250] In a possible implementation, the process of updating the parameters of the second initial recommendation sub-model based on the second set of evaluation values is: calculating a second update gradient based on each second evaluation value in the second set of evaluation values; and updating the parameters of the second initial recommendation sub-model in a direction of maximizing the second update gradient. In a possible implementation, the process of calculating the second update gradient based on each second evaluation value in the second set of evaluation values is: calculating a second target evaluation value based on each second evaluation value in the second set of evaluation values; and calculating the second update gradient based on the second target evaluation value.
[0251] In a possible implementation, the manner of calculating the second target evaluation value based on each second evaluation value in the second evaluation value set can be that: a weight is set for each second evaluation value respectively, and the weighted evaluation value of each second evaluation value is taken as the second target evaluation value. The process of calculating the second update gradient based on the second target evaluation value can be performed according to formula 14:
[0252]
[0253] wherein, denotes the second update gradient; denotes a random strategy adopted when the second initial recommendation submodel outputs the content recommendation result; φ h denotes a parameter of the second initial recommendation submodel; s h denotes a set of sample content features involved in the process of outputting the initial content recommendation result; a h denotes the initial content recommendation result; denotes the second target evaluation value.
[0254] After obtaining the second update gradient, since the optimization direction is that the evaluation value is the larger the better, the parameter of the second initial recommendation submodel is updated in the direction of maximizing the second update gradient.
[0255] For the case where the number of training samples used for training the initial hierarchical recommendation model once is multiple, the second update gradient refers to an average value of multiple second update gradients calculated based on multiple second evaluation value sets corresponding to each training sample.
[0256] Step 60125: obtaining a channel loss function based on the first enhancement value set and the first evaluation value set, obtaining a content loss function based on the second enhancement value set and the second evaluation value set, calculating a target loss function based on the channel loss function and the content loss function, and updating parameters of the first initial evaluation submodel and the second initial evaluation submodel based on the target loss function.
[0257] The first enhancement value set includes a first enhancement value corresponding to each sample recommendation content. Since the sample recommendation sequence in the sample recommendation content sequence corresponds to the initial channel recommendation subresult in the initial channel recommendation result, the first enhancement value corresponding to each sample recommendation content can be considered as the first enhancement value corresponding to each initial channel recommendation subresult. The first evaluation value set includes a first evaluation value corresponding to each initial channel recommendation subresult.
[0258] In a possible implementation, based on the first set of enhancement values and the first set of evaluation values, the process of obtaining the channel loss function is: for any initial channel recommendation sub-result, obtaining the first enhancement value corresponding to the any initial channel recommendation sub-result in the first set of enhancement values, and obtaining the first evaluation value corresponding to the any initial channel recommendation sub-result in the first set of evaluation values; based on the first enhancement value and the first evaluation value corresponding to the any initial channel recommendation sub-result, obtaining the channel sub-loss function corresponding to the any initial channel recommendation sub-result; and based on the channel sub-loss functions respectively corresponding to the initial channel recommendation sub-results, obtaining the channel loss function.
[0259] For the initial channel recommendation sub-result located at the tth position in the initial channel recommendation result, the process of obtaining the channel sub-loss function corresponding to the initial recommendation sub-result based on the first enhancement value and the first evaluation value corresponding to the initial channel recommendation sub-result can be performed according to formula 15 and formula 16:
[0260]
[0261] wherein, L t (θ l ) represents the channel sub-loss function corresponding to the initial channel recommendation sub-result located at the tth position in the initial channel recommendation result; θ l and θ l' represent the parameters of the first initial evaluation sub-model, wherein θ l is constantly updated in the training process, and θ l' is fixed in each optimization process, and after a certain number of training processes are completed, the parameter copying is performed on θ l . represents the channel feature corresponding to the initial channel recommendation sub-result located at the tth position; represents the initial channel recommendation sub-result located at the tth position; represents the first evaluation value corresponding to the initial channel recommendation sub-result located at the tth position; represents the first reference evaluation value; represents the first enhancement value corresponding to the initial channel recommendation sub-result located at the tth position; γ represents a discount factor; represents the first evaluation value corresponding to the initial channel recommendation sub-result located at the (t+1)th position under the θ l' parameters; represents the channel feature corresponding to the initial channel recommendation sub-result located at the (t+1)th position; represents the initial channel recommendation sub-result located at the (t+1)th position output by the first initial recommendation sub-model.
[0262] After determining the channel sub-loss functions respectively corresponding to the initial channel recommendation sub-results, a channel loss function is obtained based on the channel sub-loss functions respectively corresponding to the initial channel recommendation sub-results. In a possible implementation, the manner of obtaining the channel loss function based on the channel sub-loss functions respectively corresponding to the initial channel recommendation sub-results is: setting a weight for each channel sub-loss function, and taking the weighted average result of each channel sub-loss function as the channel loss function.
[0263] The second enhancement value set includes second enhancement values respectively corresponding to the sample recommendation contents. Since the sample recommendation sequences in the sample recommendation content sequence correspond to the initial content recommendation sub-results in the initial content recommendation result, the second enhancement values respectively corresponding to the sample recommendation contents can be considered as the second enhancement values respectively corresponding to the initial content recommendation sub-results. The second evaluation value set includes second evaluation values respectively corresponding to the initial content recommendation sub-results.
[0264] In a possible implementation, the process of obtaining the content loss function based on the second enhancement value set and the second evaluation value set is: for any initial content recommendation sub-result, obtaining the second enhancement value corresponding to the initial content recommendation sub-result in the second enhancement value set, and obtaining the second evaluation value corresponding to the initial content recommendation sub-result in the second evaluation value set; obtaining the content sub-loss function corresponding to the initial content recommendation sub-result based on the second enhancement value and the second evaluation value corresponding to the initial content recommendation sub-result; and obtaining the content loss function based on the content sub-loss functions respectively corresponding to the initial content recommendation sub-results.
[0265] For the initial content recommendation sub-result located at the t th position in the initial content recommendation result, the process of obtaining the content sub-loss function corresponding to the initial recommendation sub-result based on the second enhancement value and the second evaluation value corresponding to the initial content recommendation sub-result can be performed according to formula 17 and formula 18:
[0266]
[0267] wherein, L t (θ h ) represents the content sub-loss function corresponding to the initial content recommendation sub-result located at the t th position in the initial content recommendation result; θ h and θ h' represent the parameters of the second initial evaluation sub-model, wherein θ h is constantly updated in the training process, and θ h' is fixed in each optimization process. After a certain number of training processes are completed, the parameters of θ h are copied; represents the content feature corresponding to the initial content recommendation sub-result located at the t th position. represents the initial content recommendation sub-result located at the tth position; represents the second evaluation value corresponding to the initial content recommendation sub-result located at the tth position; represents the second reference evaluation value; represents the second enhancement value corresponding to the initial content recommendation sub-result located at the tth position; γ represents a discount factor; represents the initial content recommendation sub-result located at the (t+1)th position under the parameter θ h' and the second evaluation value corresponding to the initial content recommendation sub-result located at the (t+1)th position; represents the content feature corresponding to the initial content recommendation sub-result located at the (t+1)th position; represents the initial content recommendation sub-result located at the (t+1)th position output by the second initial recommendation sub-model.
[0268] After determining the content sub-loss functions respectively corresponding to the initial content recommendation sub-results, the content loss function is obtained based on the content sub-loss functions respectively corresponding to the initial content recommendation sub-results. In a possible implementation, the content loss function is obtained based on the content sub-loss functions respectively corresponding to the initial content recommendation sub-results in the following manner: weights are respectively set for the content sub-loss functions, and a weighted average result of the content sub-loss functions is taken as the content loss function.
[0269] After determining the channel loss function and the content loss function, the target loss function is calculated based on the channel loss function and the content loss function. In a possible implementation, the target loss function is calculated based on the channel loss function and the content loss function, which can be implemented based on formula 19.
[0270] L = λ t L(θ l ) + λ h L(θ h ) (formula 19)
[0271] wherein L represents the target loss function; L(θ l ) represents the channel loss function; L(θ h ) represents the content loss function; λ t represents the weight of the channel loss function; and λ h represents the weight of the content loss function.
[0272] After obtaining the target loss function, the parameters of the first initial evaluation sub-model and the second initial evaluation sub-model are updated based on the target loss function.
[0273] For a case where the number of training samples used for training the initial hierarchical recommendation model once is multiple, the target loss function refers to an average result of multiple target loss functions calculated based on the training samples.
[0274] It should be noted that each time steps 60121 to 60125 are executed, the training process of the initial hierarchical recommendation model is completed. The training process of the hierarchical recommendation model is an iterative process. Each time the training process is completed, it is determined whether the training termination condition is met. When the training termination condition is not met, the hierarchical recommendation model is continuously trained according to steps 60121 to 60125; until the training termination condition is met, the hierarchical recommendation model obtained when the training termination condition is met is taken as the target hierarchical recommendation model. In a possible implementation manner, the training termination condition is met includes, but is not limited to, the following three cases:
[0275] Case 1: The number of iterations reaches a number threshold.
[0276] The number threshold can be set according to experience, or can be flexibly adjusted according to application scenarios, and the embodiments of the present application do not limit this.
[0277] Case 2: The target loss function is less than a loss threshold.
[0278] Case 3: The target loss function converges.
[0279] The convergence of the target loss function means that, with the increase of the number of iterations, the target loss function fluctuates within a reference range in the training results of a reference number. For example, it is assumed that the reference range is -10 -3 ~ 10 -3 , and the reference number is 10. If the target loss function fluctuates within -10 -3 ~ 10 -3 in the iteration training results of 10 times, it is considered that the target loss function converges.
[0280] When any of the above cases is met, it is considered that the training process of the hierarchical recommendation model meets the training termination condition, and the hierarchical recommendation model obtained at this time is taken as the target hierarchical recommendation model.
[0281] In a possible implementation manner, in the process of obtaining the target loss function for updating the parameters of the first initial evaluation sub-model and the second initial evaluation sub-model, in addition to obtaining the channel loss function and the content loss function, other loss functions can also be obtained to further increase the updating effect of the model parameters. In a possible implementation manner, the training sample further includes a sample recommendation content sequence; after obtaining the channel loss function and the content loss function, the method further includes: obtaining at least one of a click rate loss function and a similarity loss function based on the initial content recommendation result and the sample recommendation content sequence in the training sample. The click rate loss function is used to make the content recommended by the model have a better click rate, and the similarity loss function is used to make the content recommended by the model closer to the sample recommendation content.
[0282] The initial content recommendation result includes at least one initial content recommendation sub-result arranged in sequence, the sample recommendation content sequence includes at least one sample recommendation content arranged in sequence, and the initial content recommendation sub-result and the sample recommendation content in the same arrangement position correspond to each other. In a possible implementation, based on the initial content recommendation result and the sample recommendation content sequence in the training sample, the process of obtaining at least one of the click rate loss function and the similarity loss function is: based on each initial recommendation sub-result arranged in sequence in the initial recommendation result and each sample recommendation content arranged in sequence in the sample recommendation content sequence, at least one of the click rate loss function and the similarity loss function is obtained.
[0283] In a possible implementation, the process of obtaining the click rate loss function can be implemented based on formula 20:
[0284]
[0285] L click rate (a, d) = -log(p click rate (a, d)) (formula 20) c denotes the click rate loss function; denotes the sample recommendation content clicked by the interactive object; denotes the sample recommendation content not clicked by the interactive object; denotes the calculation formula of the click rate of the initial recommendation sub-result a and the sample recommendation content corresponding to the initial recommendation sub-result a; The calculation formula is shown in formula 21:
[0286] f(a, d) = σ(w f ·concat(a, d) + b f )(formula 21)
[0287] wherein w f denotes a weight vector, b f denotes a bias; σ denotes a sigmoid function; d denotes the sample recommendation content corresponding to the initial recommendation sub-result a has the same form of expression, for example, when the form of expression of the initial recommendation sub-result a is a feature vector, d denotes the sample recommendation content corresponding to the initial recommendation sub-result a has the same form of expression; concat denotes a merging operation.
[0288] In a possible implementation, the process of obtaining the similarity loss function can be implemented based on formula 22:
[0289]
[0290] L similarity (a, d) = 1 - sim(a, d) (formula 22) s denotes the similarity loss function, represents the initial recommendation sub-result a and the sample recommended content corresponding to the initial recommendation sub-result a cosine_sim(a,d) represents the initial recommendation sub-result a and the sample recommended content corresponding to the initial recommendation sub-result a has the same form.
[0291] For the case of obtaining the channel loss function and the content loss function, and further obtaining at least one of the click rate loss function and the similarity loss function, the target loss function is calculated based on the channel loss function and the content loss function, including: based on at least one of the click rate loss function and the similarity loss function, and the channel loss function and the content loss function, calculating the target loss function.
[0292] In a possible implementation, for the case of obtaining the channel loss function and the content loss function, and further obtaining the click rate loss function and the similarity loss function, the target loss function is calculated based on at least one of the click rate loss function and the similarity loss function, and the channel loss function and the content loss function, which means: based on the click rate loss function, the similarity loss function, the channel loss function and the content loss function, calculating the target loss function. The process of calculating the target loss function can be implemented based on formula 23:
[0293] L=λ t L(θ l )+λ h L(θ h )+λ c L c +λ s L s (Formula 23)
[0294] Wherein, L represents the target loss function; L(θ l ) represents the channel loss function; L(θ h ) represents the content loss function; L c represents the click rate loss function; L s represents the similarity loss function; λ t represents the weight of the channel loss function; λ h represents the weight of the content loss function; λ c represents the weight of the click rate loss function; λ s represents the weight of the similarity loss function.
[0295] It should be noted that the above is only an exemplary description of the initial hierarchical recommendation model training process. In a possible implementation manner, in the process of training the initial hierarchical recommendation model by using the training sample, the experience array can be obtained based on the training sample first, the experience array is placed in the experience pool, and then a reference number of experience arrays are randomly selected from the experience pool for model updating. It should be noted that the experience array includes data necessary for parameter updating, including but not limited to the initial channel recommendation result, the initial content recommendation result, the first enhancement value set, the second enhancement value set, the first evaluation value set, the second evaluation value set, and the like obtained based on the training sample. The obtaining process of the experience array can refer to the related processes of steps 60121 to 60123 described above, and will not be described here. This kind of way can reduce the adverse effects of the correlation between data sets, and improve the model training effect.
[0296] After obtaining the target hierarchical recommendation model, the target hierarchical recommendation model and the recommendation model in the related art are respectively tested offline and online to verify the effectiveness of the target hierarchical recommendation model compared with the recommendation model in the related art.
[0297] In the offline test, the indicators for measuring the performance of the recommendation model are AUC (Area Under Curve) and RelaImpr (the improvement rate relative to the basic recommendation model (the LR model in the related art)), and the test results are shown in Table 1:
[0298] Table 1
[0299] Model AUC RelaImpr LR 0.7311 0.00% FM 0.7585 11.86% NFM 0.7620 13.37% AFM 0.7686 16.23% Wide & Deep 0.7801 21.20% DeepFM 0.7819 21.98% AutoInt 0.7837 22.76% Target hierarchical recommendation model 0.8097 34.01%
[0300] In Table 1, LR, FM, NFM, AFM, Wide&Deep, DeepFM and AutoInt are all recommendation models in the related art. According to Table 1, the target hierarchical recommendation model is significantly better than all the recommendation models in the related art in terms of AUC, and achieves a relative improvement rate of 34.01% compared with the basic recommendation model (the LR model in the related art). The improvement of the target hierarchical recommendation model mainly comes from two aspects: (1) The structure of hierarchical recommendation separates the channel recommendation and content recommendation tasks, making the comprehensive recommendation more accurate and flexible. The trial-and-error method based on reinforcement learning also helps the target hierarchical recommendation model to effectively learn the optimal selection. (2) The second enhancement value at the content level contains four different aspects of enhancement values to reflect the accuracy, diversity and novelty of the recommendation, and to improve the short-term and long-term experience of the interactive object from different aspects.
[0301] In online testing, the indicators for measuring the performance of the recommendation model are CTR (Click-Through-Rate) and ACN (Average Click Number Per Capita). The improvement rates of CTR and ACN relative to the basic recommendation model (the LR model in the related art) are taken as the test results, and the test results are shown in Table 2:
[0302] Table 2
[0303]
[0304]
[0305] In Table 2, DQN (LR), DQN (GRU), Double-Dueling-DQN, DDPG and hierarchical DDPG are all recommendation models based on reinforcement learning in the related art. According to Table 2, the target hierarchical recommendation model is significantly better than the recommendation models based on reinforcement learning in the related art in terms of CTR and ACN. CTR measures the accuracy of recommendation, while ACN reflects the overall satisfaction of users with the recommended content. ACN is usually more concerned because a higher ACN usually means that the interactive object is more willing to browse the recommended content, that is, the target hierarchical recommendation model can recommend better content to attract the clicks of the interactive object.
[0306] After obtaining the target hierarchical recommendation model, the target hierarchical recommendation model can also be updated based on the collected feedback of the interactive object. In real industrial-level recommendation systems, the stability of the model is one of the important factors affecting the user experience. The interactive object will passively learn how to interact effectively with the recommendation system to obtain the content of interest. This learning often lasts for a period of time, forming a stable use habit that is difficult to change once determined. However, in integrated recommendation, in order to meet the diverse needs of the interactive object, heterogeneous content of multiple channels is combined together, which also brings instability. Any change in the multiple channels and the model can cause interference in the recommendation results, thereby confusing the interactive object and damaging the experience of the interactive object. In order to evaluate the stability of the model, the change in the proportion of each channel after the model is updated is studied.
[0307] The stability of the target hierarchical recommendation model in the embodiments of the present application and the DQN model in related technologies is tested. In order to reduce the bias caused by different times and dates, the proportion of video channels recommended by the two models from 00:00 on Saturday to 23:00 on Sunday of two adjacent weeks is counted. The maximum and average relative changes of the video channel proportion of DQN can reach 18.0% and 11.7%. In contrast, the maximum and average relative changes of the video channel proportion of the target hierarchical recommendation model are only 4.5% and 1.4%, and the target hierarchical recommendation model is more stable. This is because the target hierarchical recommendation model implements the channel recommendation task and the content recommendation task by using two recommendation models with different parameters and enhancement values. The target hierarchical recommendation model can successfully learn the preference of the interactive object for the channel, so as to smooth the trend jitter caused by model updating. With the help of the hierarchical reinforcement learning architecture, the target hierarchical recommendation model will remain stable during model updating, will not confuse the cognitive and usage habits of the interactive object, will increase the stickiness of the interactive object, and will enhance the long-term experience of the interactive object.
[0308] In the embodiments of the present application, the channel recommendation task and the content recommendation task are implemented by using two recommendation models with different parameters and enhancement values, the accuracy, diversity and novelty of the recommendation results are improved by designing multiple loss functions and enhancement values, and the target hierarchical recommendation model obtained based on this training manner performs content recommendation for the interactive object, which can improve the effect of content recommendation, the click rate of the recommended content is higher, and the interactive object has a better long-term and short-term experience.
[0309] In step 602, the first target recommendation model is called to perform recommendation processing on the channel preference features corresponding to the target object to obtain a channel recommendation result; and based on the channel recommendation result and a candidate channel set corresponding to the target object, a target channel sequence is obtained.
[0310] After obtaining the target hierarchical recommendation model based on step 601, the target hierarchical recommendation model includes a first target recommendation model and a second target recommendation model. The first target recommendation model is used to obtain a channel recommendation result, and the second target recommendation model is used to obtain a content recommendation result.
[0311] The implementation process of calling the first target recommendation model in the target hierarchical recommendation model to perform recommendation processing on the channel preference features corresponding to the target object to obtain a channel recommendation result, and then obtaining a target channel sequence based on the channel recommendation result and a candidate channel set corresponding to the target object can be referred to step 302 in the embodiment shown in Figure 3 Here, no longer be described.
[0312] It should be noted that before step 602 is performed, the channel preference feature corresponding to the target object and the candidate channel set corresponding to the target object need to be obtained. The process of obtaining the channel preference feature corresponding to the target object and the candidate channel set corresponding to the target object can be referred to Figure 3 Step 301 in the embodiment shown in FIG. 3, which will not be repeated here.
[0313] In step 603, a second target recommendation model is called to perform recommendation processing on the content preference feature corresponding to the target object to obtain a content recommendation result; and based on the content recommendation result and the candidate content set corresponding to the target object, a target content sequence corresponding to the target channel sequence is obtained.
[0314] The second target recommendation model is used to obtain the content recommendation result. The second target recommendation model is called to perform recommendation processing on the content preference feature corresponding to the target object to obtain the content recommendation result, and then based on the content recommendation result and the candidate content set corresponding to the target object, the target content sequence corresponding to the target channel sequence is obtained. The implementation process can be referred to Figure 3 Step 303 in the embodiment shown in FIG. 3, which will not be repeated here.
[0315] It should be noted that before step 603 is performed, the content preference feature corresponding to the target object and the candidate content set corresponding to the target object need to be obtained. The process of obtaining the content preference feature corresponding to the target object and the candidate content set corresponding to the target object can be referred to Figure 3 Step 301 in the embodiment shown in FIG. 3, which will not be repeated here.
[0316] In step 604, the target content sequence is recommended to the target object.
[0317] The implementation process of step 604 can be referred to Figure 3 Step 304 in the embodiment shown in FIG. 3, which will not be repeated here.
[0318] In the embodiment of the present application, the first target recommendation model in the target hierarchical recommendation model is used to implement the task of obtaining the target channel sequence, and then the second target recommendation model in the target hierarchical recommendation model is used to implement the task of obtaining the target content sequence corresponding to the target channel sequence, and then the target content sequence is recommended to the target object. In this content recommendation process, the first target recommendation model considers the information of the channel, the second target recommendation model considers the information of the content, the content recommendation process considers both the information of the channel and the information of the content, and the information considered is more comprehensive, which is conducive to improving the effect of content recommendation, the click rate of the recommended content is higher, and the user's interactive experience is better.
[0319] Referring to Figure 8 The embodiment of the present application provides a content recommendation device, which comprises:
[0320] The first obtaining unit 801 is configured to obtain a channel preference feature, a content preference feature and a candidate content set corresponding to a target object, the candidate content set including at least one candidate content, and any candidate content corresponding to a candidate channel;
[0321] The second obtaining unit 802 is configured to obtain a target channel sequence based on the channel preference feature and a candidate channel set corresponding to the candidate content set, the candidate channel set including candidate channels corresponding to each candidate content in the candidate content set;
[0322] The third obtaining unit 803 is configured to obtain a target content sequence corresponding to the target channel sequence based on the content preference feature and the candidate content set.
[0323] The recommendation unit 804 is configured to recommend the target content sequence to the target object.
[0324] In a possible implementation, the second obtaining unit 802 is configured to call a first target recommendation model to perform recommendation processing on the channel preference feature, to obtain a channel recommendation result, and the channel recommendation result is composed of at least one channel recommendation sub-result; for any channel recommendation sub-result, a target channel matching the channel recommendation sub-result is obtained from the candidate channel set; and a target channel sequence is obtained based on the target channels respectively matching each channel recommendation sub-result.
[0325] In a possible implementation, the third obtaining unit 803 is configured to call a second target recommendation model to perform recommendation processing on the content preference feature, to obtain a content recommendation result, and the content recommendation result is composed of at least one content recommendation sub-result; for any content recommendation sub-result, a target content matching the content recommendation sub-result is obtained from a target candidate content set, the target candidate content set is composed of candidate contents in the candidate content set that meet a condition, the candidate contents that meet the condition include candidate contents whose corresponding candidate channels are specified channels, and the specified channels include target channels whose arrangement positions in the target channel sequence and arrangement positions of the content recommendation sub-result in the content recommendation result are consistent; and a target content sequence corresponding to the target channel sequence is obtained based on the target contents respectively matching each content recommendation sub-result.
[0326] In a possible implementation, the second obtaining unit 802 is further configured to input the channel preference feature into a first target recommendation model to obtain a first channel recommendation sub-result output by the first target recommendation model; obtain an updated channel preference feature based on the first channel recommendation sub-result, input the updated channel preference feature into the first target recommendation model, and obtain a second channel recommendation sub-result output by the first target recommendation model; continue to obtain channel recommendation sub-results based on the second channel recommendation sub-result; and in response to a quantity of obtained channel recommendation sub-results being equal to a quantity threshold, arrange each channel recommendation sub-result arranged in a sequence according to an obtaining sequence as the channel recommendation result.
[0327] In a possible implementation, the third obtaining unit 803 is further configured to input the content preference feature into a second target recommendation model to obtain a first content recommendation sub-result output by the second target recommendation model; obtain an updated content preference feature based on the first content recommendation sub-result, input the updated content preference feature into the second target recommendation model, and obtain a second content recommendation sub-result output by the second target recommendation model; continue to obtain content recommendation sub-results based on the second content recommendation sub-result; and in response to a quantity of content recommendation sub-results being equal to a quantity threshold, arrange each content recommendation sub-result arranged in a sequence according to an obtaining sequence as the content recommendation result.
[0328] In a possible implementation, the first obtaining unit 801 is configured to obtain a historical recommendation content sequence corresponding to a target object; obtain a to-be-processed channel feature sequence and a to-be-processed content feature sequence based on the historical recommendation content sequence; call a first processing model to process the to-be-processed channel feature sequence to obtain a channel preference feature corresponding to the target object; and call a second processing model to process the to-be-processed content feature sequence to obtain a content preference feature corresponding to the target object.
[0329] In a possible implementation manner, the sequence of historical recommended contents is composed of at least one historical recommended content arranged in sequence; the first obtaining unit 801 is further configured to, for any historical recommended content, obtain corresponding basic information, to-be-processed channel information and to-be-processed content information of the any historical recommended content, wherein the basic information includes at least one of object attribute information and environment information, the to-be-processed channel information includes at least one of channel information and accumulated channel information, and the to-be-processed content information includes at least one of content information and accumulated content information; perform fusion processing on the basic information and the to-be-processed channel information corresponding to the any historical recommended content, to obtain to-be-processed channel features corresponding to the any historical recommended content; perform fusion processing on the basic information and the to-be-processed content information corresponding to the any historical recommended content, to obtain to-be-processed content features corresponding to the any historical recommended content; arrange the to-be-processed channel features corresponding to each historical recommended content in the sequence of the historical recommended contents, to obtain a sequence of to-be-processed channel features; and arrange the to-be-processed content features corresponding to each historical recommended content in the sequence of the historical recommended contents, to obtain a sequence of to-be-processed content features.
[0330] In the embodiments of the present application, the target channel sequence is obtained based on the channel preference features, and then the target content sequence corresponding to the target channel sequence is obtained based on the content preference features, and the target content sequence is recommended to the target object. In the process of content recommendation, the channel preference features reflect information about channels, the content preference features reflect information about contents, the process of content recommendation considers both information about channels and information about contents, the considered information is more comprehensive, which is conducive to improving the effect of content recommendation, the click rate of the recommended content is higher, and the interactive experience of the user is better.
[0331] Referring to Figure 9 The embodiments of the present application provide a content recommendation device, which comprises:
[0332] The first obtaining unit 901 is configured to obtain a target layered recommendation model, wherein the target layered recommendation model comprises a first target recommendation model and a second target recommendation model.
[0333] The second obtaining unit 902 is configured to call the first target recommendation model to perform recommendation processing on channel preference features corresponding to a target object, to obtain a channel recommendation result; and obtain a target channel sequence based on the channel recommendation result and a candidate channel set corresponding to the target object.
[0334] The third obtaining unit 903 is configured to call the second target recommendation model to perform recommendation processing on the content preference features corresponding to the target object, to obtain a content recommendation result; and obtain a target content sequence corresponding to the target channel sequence based on the content recommendation result and a candidate content set corresponding to the target object.
[0335] The recommendation unit 904 is configured to recommend the target content sequence to the target object.
[0336] In a possible implementation manner, referring to Figure 10 The apparatus further includes:
[0337] The fourth obtaining unit 905 is configured to obtain a training sample set, the training sample set including at least one training sample, and any training sample including sample channel features, sample content features, and feedback information corresponding to a sample recommendation content sequence.
[0338] The training unit 906 is configured to train the first initial recommendation model and the second initial recommendation model in the initial hierarchical recommendation model based on the sample channel features, the sample content features, and the feedback information corresponding to the sample recommendation content sequence in the training sample, to obtain the target hierarchical recommendation model.
[0339] In a possible implementation manner, the first initial recommendation model includes a first initial recommendation submodel and a first initial evaluation submodel, and the second initial recommendation model includes a second initial recommendation submodel and a second initial evaluation submodel; the training unit 906 is configured to obtain a first enhancement value set and a second enhancement value set based on the feedback information corresponding to the sample recommendation content sequence in the training sample; input the sample channel features in the training sample into the first initial recommendation model, to obtain an initial channel recommendation result output by the first initial recommendation submodel and a first evaluation value set for the initial channel recommendation result output by the first initial evaluation submodel; input the sample content features in the training sample into the second initial recommendation model, to obtain an initial content recommendation result output by the second initial recommendation submodel and a second evaluation value set for the initial content recommendation result output by the second initial evaluation submodel; update parameters of the first initial recommendation submodel based on the first evaluation value set; update parameters of the second initial recommendation submodel based on the second evaluation value set; obtain a channel loss function based on the first enhancement value set and the first evaluation value set; obtain a content loss function based on the second enhancement value set and the second evaluation value set; calculate a target loss function based on the channel loss function and the content loss function; and update parameters of the first initial evaluation submodel and the second initial evaluation submodel based on the target loss function.
[0340] In a possible implementation, the training sample further includes a sample recommended content sequence; the training unit 906 is further configured to obtain at least one of the click rate loss function and the similarity loss function based on the initial content recommendation result and the sample recommended content sequence in the training sample; and calculate the target loss function based on at least one of the click rate loss function and the similarity loss function, and the channel loss function and the content loss function.
[0341] In a possible implementation, the training unit 906 is further configured to obtain at least one of reading duration information, diversity information and novelty information of any sample recommended content in the sample recommended content sequence and click information of the any sample recommended content based on the feedback information in the training sample; obtain a first enhancement value corresponding to the any sample recommended content based on the click information of the any sample recommended content; obtain a second enhancement value corresponding to the any sample recommended content based on at least one of the reading duration information, the diversity information and the novelty information of the any sample recommended content and the click information of the any recommended content; and set a set of the first enhancement values corresponding to each sample recommended content as a first enhancement value set, and set a set of the second enhancement values corresponding to each sample recommended content as a second enhancement value set.
[0342] In the embodiments of the present application, the first target recommendation model in the target hierarchical recommendation model is used to implement the task of obtaining the target channel sequence, and then the second target recommendation model in the target hierarchical recommendation model is used to implement the task of obtaining the target content sequence corresponding to the target channel sequence, and then the target content sequence is recommended to the target object. In the process of content recommendation, the first target recommendation model considers the information of the channel, the second target recommendation model considers the information of the content, the process of content recommendation considers both the information of the channel and the information of the content, the considered information is more comprehensive, which is conducive to improving the effect of content recommendation, the click rate of the recommended content is higher, and the interactive experience of the user is better.
[0343] It should be noted that the apparatus provided in the above embodiments is only exemplified by the division of the above functional modules in realizing its functions, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the above described functions. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be repeated here.
[0344] Figure 11Fig. 1 is a structural schematic diagram of a server provided by an embodiment of the present application. The server can have great differences due to different configurations or performances, and can include one or more processors (Central Processing Units, CPUs) 1101 and one or more memories 1102. The one or more memories 1102 store at least one piece of program code, which is loaded and executed by the one or more processors 1101 to implement the content recommendation method provided by each of the above-mentioned method embodiments. Of course, the server can also have a wired or wireless network interface, a keyboard, an input and output interface, and other components for realizing the functions of the device, and will not be described here in detail.
[0345] Figure 12 Fig. 2 is a structural schematic diagram of a terminal provided by an embodiment of the present application. The terminal can be a smart phone, a tablet computer, a notebook computer or a desktop computer. The terminal can also be referred to as a user equipment, a portable terminal, a laptop terminal, a desktop terminal, and other names.
[0346] Generally, the terminal includes a processor 1201 and a memory 1202.
[0347] The processor 1201 can include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1201 can be implemented in at least one of a hardware form of a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), and a PLA (Programmable Logic Array). The processor 1201 can also include a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also known as a CPU (Central Processing Unit). The coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor 1201 can be integrated with a GPU (Graphics Processing Unit) that is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1201 can also include an AI (Artificial Intelligence) processor for processing machine learning-related computing operations.
[0348] The memory 1202 can include one or more computer-readable storage media. The computer-readable storage media can be non-transitory. The memory 1202 can also include high-speed random access memory and can include nonvolatile memory, such as one or more magnetic disk storage devices, optical storage devices, flash memory devices, or other nonvolatile solid-state storage devices. In some embodiments, the non-transitory computer-readable storage medium of the memory 1202 is used to store at least one instruction for execution by the processor 1201 to implement the content recommendation method provided by the method embodiments of the present application.
[0349] In some embodiments, the terminal can further optionally include a peripheral device interface 1203 and at least one peripheral device. The processor 1201, the memory 1202, and the peripheral device interface 1203 can be connected through a bus or a signal line. Each peripheral device can be connected to the peripheral device interface 1203 through a bus, a signal line, or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 1204, a touch display screen 1205, a camera assembly 1206, an audio circuit 1207, a positioning assembly 1208, and a power supply 1209.
[0350] The peripheral device interface 1203 can be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 1201 and the memory 1202. In some embodiments, the processor 1201, the memory 1202, and the peripheral device interface 1203 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1201, the memory 1202, and the peripheral device interface 1203 can be implemented on a separate chip or circuit board, and the present embodiment is not limited in this regard.
[0351] The radio frequency circuit 1204 is configured to receive and send RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1204 communicates with communication networks and other communication devices through electromagnetic signals. The radio frequency circuit 1204 converts electrical signals to electromagnetic signals for transmission, or vice versa. Optionally, the radio frequency circuit 1204 includes an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and the like. The radio frequency circuit 1204 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes, but is not limited to, a metropolitan area network, various generations of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1204 can also include NFC (Near Field Communication)-related circuitry, which is not limited in the present application.
[0352] The display screen 1205 is configured to display a UI (User Interface). The UI can include graphics, text, icons, video, and any combination thereof. When the display screen 1205 is a touch display screen, the display screen 1205 also has the ability to collect touch signals on or above the surface of the display screen 1205. The touch signals can be input as control signals to the processor 1201 for processing. At this time, the display screen 1205 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, the display screen 1205 can be one, arranged on the front panel of the terminal; in other embodiments, the display screen 1205 can be at least two, arranged on different surfaces of the terminal or in a folding design; in still other embodiments, the display screen 1205 can be a flexible display screen, arranged on a curved surface or a folding surface of the terminal. Even, the display screen 1205 can also be arranged in an irregular shape, i.e., a special-shaped screen. The display screen 1205 can be made of materials such as LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode), and the like.
[0353] The camera component 1206 is configured to capture images or videos. Optionally, the camera component 1206 includes a front camera and a rear camera. Generally, the front camera is disposed on the front panel of the terminal, and the rear camera is disposed on the back of the terminal. In some embodiments, the rear camera is at least two, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a long-focus camera, to realize the background blur function of the main camera and the depth-of-field camera, the panorama shooting and VR (Virtual Reality) shooting function of the main camera and the wide-angle camera, or other fusion shooting functions. In some embodiments, the camera component 1206 can also include a flash. The flash can be a single-color-temperature flash or a dual-color-temperature flash. The dual-color-temperature flash refers to the combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.
[0354] The audio circuit 1207 can include a microphone and a speaker. The microphone is configured to capture sound waves of a user and an environment, and convert the sound waves into an electrical signal input to the processor 1201 for processing or to the radio frequency circuit 1204 to realize voice communication. For the purpose of stereo sound collection or noise reduction, the microphone can be multiple, which are respectively disposed at different parts of the terminal. The microphone can also be an array microphone or an omnidirectional collection microphone. The speaker is configured to convert an electrical signal from the processor 1201 or the radio frequency circuit 1204 into sound waves. The speaker can be a traditional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert an electrical signal into a sound wave audible to humans, but also convert an electrical signal into an inaudible sound wave to humans for ranging purposes. In some embodiments, the audio circuit 1207 can also include a headphone jack.
[0355] The positioning component 1208 is configured to position the current geographical position of the terminal to realize navigation or LBS (Location Based Service). The positioning component 1208 can be a positioning component based on the GPS (Global Positioning System) of the United States, the Beidou system of China, the Glonass system, or the Galileo system of the European Union.
[0356] The power supply 1209 is configured to supply power to each component in the terminal. The power supply 1209 can be alternating current, direct current, a disposable battery, or a rechargeable battery. When the power supply 1209 includes a rechargeable battery, the rechargeable battery can support wired charging or wireless charging. The rechargeable battery can also be used to support fast charging technology.
[0357] In some embodiments, the terminal further includes one or more sensors 1210. The one or more sensors 1210 include, but are not limited to, an acceleration sensor 1211, a gyroscope sensor 1212, a pressure sensor 1213, a fingerprint sensor 1214, an optical sensor 1215, and a proximity sensor 1216.
[0358] The acceleration sensor 1211 can detect the acceleration magnitude in three coordinate axes of a coordinate system established by the terminal. For example, the acceleration sensor 1211 can be used to detect the components of the gravitational acceleration in three coordinate axes. The processor 1201 can control the touch display screen 1205 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signals collected by the acceleration sensor 1211. The acceleration sensor 1211 can also be used for game or user motion data collection.
[0359] The gyroscope sensor 1212 can detect the body orientation and rotation angle of the terminal. The gyroscope sensor 1212 can work with the acceleration sensor 1211 to collect the 3D motion of the user to the terminal. The processor 1201 can implement the following functions according to the data collected by the gyroscope sensor 1212: motion sensing (e.g., changing the UI according to the user's tilt operation), image stabilization when shooting, game control, and inertial navigation.
[0360] The pressure sensor 1213 can be disposed on the side frame of the terminal and / or the lower layer of the touch display screen 1205. When the pressure sensor 1213 is disposed on the side frame of the terminal, the user's holding signal to the terminal can be detected, and the left-hand or right-hand recognition or shortcut operation can be performed by the processor 1201 according to the holding signal collected by the pressure sensor 1213. When the pressure sensor 1213 is disposed on the lower layer of the touch display screen 1205, the operable control on the UI interface can be controlled by the processor 1201 according to the user's pressure operation to the touch display screen 1205. The operable control includes at least one of a button control, a scroll bar control, an icon control, and a menu control.
[0361] The fingerprint sensor 1214 is used to collect the user's fingerprint. The identity of the user can be recognized by the processor 1201 according to the fingerprint collected by the fingerprint sensor 1214, or by the fingerprint sensor 1214 according to the collected fingerprint. When the identity of the user is recognized as a trusted identity, the processor 1201 authorizes the user to perform related sensitive operations, including unlocking the screen, viewing encrypted information, downloading software, payment, and changing settings, etc. The fingerprint sensor 1214 can be disposed on the front, back, or side of the terminal. When the terminal is provided with a physical button or a manufacturer's logo, the fingerprint sensor 1214 can be integrated with the physical button or the manufacturer's logo.
[0362] The optical sensor 1215 is configured to collect ambient light intensity. In one embodiment, the processor 1201 can control the display brightness of the touch display screen 1205 according to the ambient light intensity collected by the optical sensor 1215. Specifically, when the ambient light intensity is high, the display brightness of the touch display screen 1205 is increased; when the ambient light intensity is low, the display brightness of the touch display screen 1205 is decreased. In another embodiment, the processor 1201 can also dynamically adjust the shooting parameters of the camera assembly 1206 according to the ambient light intensity collected by the optical sensor 1215.
[0363] The proximity sensor 1216, also referred to as a distance sensor, is usually arranged on the front panel of the terminal. The proximity sensor 1216 is configured to collect the distance between the user and the front of the terminal. In one embodiment, when the proximity sensor 1216 detects that the distance between the user and the front of the terminal gradually decreases, the processor 1201 controls the touch display screen 1205 to switch from the bright screen state to the screen-off state; when the proximity sensor 1216 detects that the distance between the user and the front of the terminal gradually increases, the processor 1201 controls the touch display screen 1205 to switch from the screen-off state to the bright screen state.
[0364] Those skilled in the art can understand that the structure shown in the above embodiments does not constitute a limitation on the terminal, and the terminal can include more or fewer components than those shown in the figure, or combine certain components, or adopt a different component arrangement. Figure 12
[0365] In an example embodiment, a computer device is also provided, which includes a processor and a memory having at least one program code stored therein. The at least one program code is loaded and executed by one or more processors to implement any of the above content recommendation methods.
[0366] In an example embodiment, a computer readable storage medium is also provided, which has at least one program code stored therein. The at least one program code is loaded and executed by the processor of the computer device to implement any of the above content recommendation methods.
[0367] Optionally, the above computer readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0368] It should be understood that the "multiple" mentioned herein refers to two or more than two. The "and / or" describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the three cases of A existing alone, A and B existing together, and B existing alone. The character " / " generally represents that the associated objects before and after it are in an "or" relationship.
[0369] It should be noted that the terms "first", "second" and the like (if any) in the specification and claims of the present application are used to distinguish similar objects, and do not necessarily have to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. Rather, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0370] The above description is only an exemplary embodiment of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A content recommendation method, characterized in that, The method includes: Obtain the channel preference features, content preference features, and candidate content set corresponding to the target object. The candidate content set includes at least one candidate content, and each candidate content corresponds to a candidate channel. Based on the channel preference features and the candidate channel set corresponding to the candidate content set, a target channel sequence is obtained. The candidate channel set includes candidate channels corresponding to each candidate content in the candidate content set. The target channel sequence includes target channels corresponding to at least one channel recommendation sub-result obtained cyclically. In the at least one channel recommendation sub-result, the first channel recommendation sub-result is obtained by the first target recommendation model performing recommendation processing on the channel preference features. The second channel recommendation sub-result is obtained by the first target recommendation model performing recommendation processing on the updated channel preference features obtained based on the first channel recommendation sub-result. Other channel recommendation sub-results are obtained cyclically based on the second channel recommendation sub-result. The channel recommendation sub-results obtained in each cycle are correlated with the channel recommendation sub-results obtained in the previous cycle. Based on the content preference features and the candidate content set, a target content sequence corresponding to the target channel sequence is obtained. The target content sequence is recommended to the target object.
2. The method according to claim 1, characterized in that, The step of obtaining the target channel sequence based on the channel preference features and the candidate channel set corresponding to the candidate content set includes: The first target recommendation model is invoked to perform recommendation processing on the channel preference features to obtain channel recommendation results, which are composed of at least one channel recommendation sub-result. For any channel recommendation sub-result, obtain the target channel that matches the channel recommendation sub-result from the candidate channel set; Based on the target channels that are matched with the recommendation sub-results of each channel, the target channel sequence is obtained.
3. The method according to claim 2, characterized in that, The step of obtaining the target content sequence corresponding to the target channel sequence based on the content preference features and the candidate content set includes: The second target recommendation model is invoked to perform recommendation processing on the content preference features to obtain a content recommendation result, which consists of at least one content recommendation sub-result. For any content recommendation sub-result, target content matching the content recommendation sub-result is obtained from the target candidate content set. The target candidate content set consists of candidate content that meets the conditions in the candidate content set. The candidate content that meets the conditions includes candidate content whose corresponding candidate channel is a specified channel. The specified channel includes target channels whose arrangement position in the target channel sequence is the same as the arrangement position of the content recommendation sub-result in the content recommendation result. Based on the target content that matches each content recommendation sub-result, a target content sequence corresponding to the target channel sequence is obtained.
4. The method according to claim 2, characterized in that, The step of calling the first target recommendation model to perform recommendation processing on the channel preference features to obtain channel recommendation results includes: The channel preference features are input into the first target recommendation model to obtain the first channel recommendation sub-result output by the first target recommendation model. Based on the first channel recommendation sub-result, the updated channel preference features are obtained, and the updated channel preference features are input into the first target recommendation model to obtain the second channel recommendation sub-result output by the first target recommendation model. Based on the second channel recommendation sub-result, continue to obtain channel recommendation sub-results; In response to the number of channel recommendation sub-results being equal to the number threshold, each channel recommendation sub-result arranged in the order of acquisition is taken as the channel recommendation result.
5. The method according to any one of claims 1-4, characterized in that, The acquisition of the channel preference features and content preference features corresponding to the target object includes: Obtain the historical recommended content sequence corresponding to the target object; Based on the historical recommended content sequence, obtain the channel feature sequence to be processed and the content feature sequence to be processed; The first processing model is invoked to process the channel feature sequence to be processed, thereby obtaining the channel preference features corresponding to the target object; the second processing model is invoked to process the content feature sequence to be processed, thereby obtaining the content preference features corresponding to the target object.
6. The method according to claim 5, characterized in that, The historical recommendation content sequence is composed of at least one historical recommendation content arranged sequentially; the step of obtaining the channel feature sequence and the content feature sequence to be processed based on the historical recommendation content sequence includes: For any historical recommended content, obtain the basic information, channel information to be processed, and content information to be processed corresponding to the historical recommended content. The basic information includes at least one of object attribute information and environment information. The channel information to be processed includes at least one of channel information and cumulative channel information. The content information to be processed includes at least one of content information and cumulative content information. The basic information and channel information to be processed corresponding to any historical recommended content are fused to obtain the channel features to be processed corresponding to any historical recommended content. The basic information and the content information to be processed corresponding to any historical recommended content are fused to obtain the content features to be processed corresponding to any historical recommended content. The channel features corresponding to each historical recommended content are arranged according to the order of the historical recommended content in the historical recommended content sequence to obtain the channel feature sequence to be processed; the content features corresponding to each historical recommended content are arranged according to the order of the historical recommended content in the historical recommended content sequence to obtain the content feature sequence to be processed.
7. A content recommendation method, characterized in that, The method includes: Obtain a target hierarchical recommendation model, wherein the target hierarchical recommendation model includes a first target recommendation model and a second target recommendation model; The first target recommendation model is invoked to perform recommendation processing on the channel preference features corresponding to the target object, resulting in a channel recommendation result including at least one channel recommendation sub-result. Based on the channel recommendation result and the candidate channel set corresponding to the target object, a target channel sequence is obtained. The target channel sequence includes the target channels corresponding to each of the at least one channel recommendation sub-result obtained cyclically. In the at least one channel recommendation sub-result, the first channel recommendation sub-result is obtained by the first target recommendation model performing recommendation processing on the channel preference features, the second channel recommendation sub-result is obtained by the first target recommendation model performing recommendation processing on the updated channel preference features obtained based on the first channel recommendation sub-result, and other channel recommendation sub-results are obtained cyclically based on the second channel recommendation sub-result. The channel recommendation sub-results obtained in each cycle are correlated with the channel recommendation sub-results obtained in the previous cycle. The second target recommendation model is invoked to perform recommendation processing on the content preference features corresponding to the target object, and a content recommendation result is obtained; based on the content recommendation result and the candidate content set corresponding to the target object, a target content sequence corresponding to the target channel sequence is obtained; The target content sequence is recommended to the target object.
8. The method according to claim 7, characterized in that, Before obtaining the target hierarchical recommendation model, the method further includes: Obtain a training sample set, which includes at least one training sample. Each training sample includes sample channel features, sample content features, and feedback information corresponding to the sample recommendation content sequence. Based on the sample channel features, sample content features, and feedback information corresponding to the sample recommendation content sequence in the training samples, the first initial recommendation model and the second initial recommendation model in the initial hierarchical recommendation model are trained to obtain the target hierarchical recommendation model.
9. The method according to claim 8, characterized in that, The first initial recommendation model includes a first initial recommendation sub-model and a first initial evaluation sub-model, and the second initial recommendation model includes a second initial recommendation sub-model and a second initial evaluation sub-model; The training of the first and second initial recommendation models in the initial hierarchical recommendation model based on the feedback information corresponding to the sample channel features, sample content features, and sample recommendation content sequences in the training samples includes: Based on the feedback information corresponding to the sample recommendation content sequence in the training samples, obtain the first enhancement value set and the second enhancement value set; Input the sample channel features from the training samples into the first initial recommendation model to obtain the initial channel recommendation result output by the first initial recommendation sub-model and the first evaluation value set output by the first initial evaluation sub-model for the initial channel recommendation result; Input the sample content features from the training samples into the second initial recommendation model to obtain the initial content recommendation result output by the second initial recommendation sub-model and the second evaluation value set output by the second initial evaluation sub-model for the initial content recommendation result; Update the parameters of the first initial recommendation sub-model based on the first evaluation value set; update the parameters of the second initial recommendation sub-model based on the second evaluation value set; Based on the first augmentation set and the first evaluation set, obtain the channel loss function; based on the second augmentation set and the second evaluation set, obtain the content loss function; based on the channel loss function and the content loss function, calculate the target loss function; update the parameters of the first initial evaluation sub-model and the second initial evaluation sub-model based on the target loss function.
10. The method according to claim 9, characterized in that, The training samples also include a sequence of recommended content; the calculation of the target loss function based on the channel loss function and the content loss function includes: Based on the initial content recommendation results and the sample recommendation content sequence in the training samples, at least one of the click-through rate loss function and the similarity loss function is obtained; The target loss function is calculated based on at least one of the click-through rate loss function and the similarity loss function, as well as the channel loss function and the content loss function.
11. The method according to claim 9 or 10, characterized in that, The step of obtaining a first enhancement value set and a second enhancement value set based on feedback information corresponding to the sample recommendation content sequence in the training samples includes: Based on the feedback information in the training samples, obtain at least one of the following: reading time information, diversity information, and novelty information of any sample recommended content in the sample recommended content sequence, as well as the click information of any sample recommended content; Based on the click information of the recommended content of any sample, obtain the first enhancement value corresponding to the recommended content of any sample; Based on at least one of the reading time information, diversity information, and novelty information of any sample recommended content, and the click information of any sample recommended content, obtain the second enhancement value corresponding to any sample recommended content; The set of first enhancement values corresponding to the recommended content of each sample is taken as the first enhancement value set; the set of second enhancement values corresponding to the recommended content of each sample is taken as the second enhancement value set.
12. A content recommendation device, characterized in that, The device includes: The first acquisition unit is used to acquire the channel preference features, content preference features and candidate content set corresponding to the target object. The candidate content set includes at least one candidate content, and each candidate content corresponds to a candidate channel. The second acquisition unit is used to acquire a target channel sequence based on the channel preference features and the candidate channel set corresponding to the candidate content set. The candidate channel set includes candidate channels corresponding to each candidate content in the candidate content set. The target channel sequence includes target channels corresponding to at least one channel recommendation sub-result obtained cyclically. Among the at least one channel recommendation sub-result, the first channel recommendation sub-result is obtained by the first target recommendation model performing recommendation processing on the channel preference features. The second channel recommendation sub-result is obtained by the first target recommendation model performing recommendation processing on the updated channel preference features obtained based on the first channel recommendation sub-result. Other channel recommendation sub-results are obtained cyclically based on the second channel recommendation sub-result. The channel recommendation sub-results obtained in each cycle are correlated with the channel recommendation sub-results obtained in the previous cycle. The third acquisition unit is used to acquire a target content sequence corresponding to the target channel sequence based on the content preference features and the candidate content set. The recommendation unit is used to recommend the target content sequence to the target object.
13. The apparatus according to claim 12, characterized in that, The second acquisition unit is used for: The first target recommendation model is invoked to perform recommendation processing on the channel preference features to obtain channel recommendation results, which are composed of at least one channel recommendation sub-result. For any channel recommendation sub-result, obtain the target channel that matches the channel recommendation sub-result from the candidate channel set; Based on the target channels that are matched with the recommendation sub-results of each channel, the target channel sequence is obtained.
14. The apparatus according to claim 13, characterized in that, The third acquisition unit is used for: The second target recommendation model is invoked to perform recommendation processing on the content preference features to obtain a content recommendation result, which consists of at least one content recommendation sub-result. For any content recommendation sub-result, target content matching the content recommendation sub-result is obtained from the target candidate content set. The target candidate content set consists of candidate content that meets the conditions in the candidate content set. The candidate content that meets the conditions includes candidate content whose corresponding candidate channel is a specified channel. The specified channel includes target channels whose arrangement position in the target channel sequence is the same as the arrangement position of the content recommendation sub-result in the content recommendation result. Based on the target content that matches each content recommendation sub-result, a target content sequence corresponding to the target channel sequence is obtained.
15. The apparatus according to claim 13, characterized in that, The second acquisition unit is further configured to: The channel preference features are input into the first target recommendation model to obtain the first channel recommendation sub-result output by the first target recommendation model. Based on the first channel recommendation sub-result, the updated channel preference features are obtained, and the updated channel preference features are input into the first target recommendation model to obtain the second channel recommendation sub-result output by the first target recommendation model. Based on the second channel recommendation sub-result, continue to obtain channel recommendation sub-results; In response to the number of channel recommendation sub-results being equal to the number threshold, each channel recommendation sub-result arranged in the order of acquisition is taken as the channel recommendation result.
16. The apparatus according to any one of claims 12-14, characterized in that, The first acquisition unit is used for: Obtain the historical recommended content sequence corresponding to the target object; Based on the historical recommended content sequence, obtain the channel feature sequence to be processed and the content feature sequence to be processed; The first processing model is invoked to process the channel feature sequence to be processed, thereby obtaining the channel preference features corresponding to the target object; The second processing model is invoked to process the feature sequence of the content to be processed, thereby obtaining the content preference features corresponding to the target object.
17. The apparatus according to claim 16, characterized in that, The sequence of historical recommended content consists of at least one historical recommended content arranged sequentially; the first acquisition unit is further configured to: For any historical recommended content, obtain the basic information, channel information to be processed, and content information to be processed corresponding to the historical recommended content. The basic information includes at least one of object attribute information and environment information. The channel information to be processed includes at least one of channel information and cumulative channel information. The content information to be processed includes at least one of content information and cumulative content information. The basic information and channel information to be processed corresponding to any historical recommended content are fused to obtain the channel features to be processed corresponding to any historical recommended content. The basic information and the content information to be processed corresponding to any historical recommended content are fused to obtain the content features to be processed corresponding to any historical recommended content. The channel features corresponding to each historical recommended content are arranged according to the order of the historical recommended content in the historical recommended content sequence to obtain the channel feature sequence to be processed; the content features corresponding to each historical recommended content are arranged according to the order of the historical recommended content in the historical recommended content sequence to obtain the content feature sequence to be processed.
18. A content recommendation device, characterized in that, The device includes: The first acquisition unit is used to acquire a target hierarchical recommendation model, wherein the target hierarchical recommendation model includes a first target recommendation model and a second target recommendation model; The second acquisition unit is used to call the first target recommendation model to perform recommendation processing on the channel preference features corresponding to the target object, and obtain a channel recommendation result including at least one channel recommendation sub-result; based on the channel recommendation result and the candidate channel set corresponding to the target object, acquire a target channel sequence; the target channel sequence includes the target channels corresponding to each of the at least one channel recommendation sub-result obtained in a loop, wherein the first channel recommendation sub-result is obtained by the first target recommendation model to perform recommendation processing on the channel preference features, the second channel recommendation sub-result is obtained by the first target recommendation model to perform recommendation processing on the updated channel preference features obtained based on the first channel recommendation sub-result, and other channel recommendation sub-results are obtained in a loop based on the second channel recommendation sub-result, and the channel recommendation sub-results obtained in each loop process are related to the channel recommendation sub-results obtained in the previous loop process; The third acquisition unit is used to call the second target recommendation model to perform recommendation processing on the content preference features corresponding to the target object to obtain the content recommendation result; based on the content recommendation result and the candidate content set corresponding to the target object, the target content sequence corresponding to the target channel sequence is obtained. The recommendation unit is used to recommend the target content sequence to the target object.
19. The apparatus according to claim 18, characterized in that, The device further includes: The fourth acquisition unit is used to acquire a training sample set before acquiring the target hierarchical recommendation model. The training sample set includes at least one training sample, and each training sample includes sample channel features, sample content features, and feedback information corresponding to the sample recommendation content sequence. The training unit is used to train the first and second initial recommendation models in the initial hierarchical recommendation model based on the sample channel features, sample content features and feedback information corresponding to the sample recommendation content sequence in the training samples, so as to obtain the target hierarchical recommendation model.
20. The apparatus according to claim 19, characterized in that, The first initial recommendation model includes a first initial recommendation sub-model and a first initial evaluation sub-model, and the second initial recommendation model includes a second initial recommendation sub-model and a second initial evaluation sub-model; The training unit is used for: Based on the feedback information corresponding to the sample recommendation content sequence in the training samples, obtain the first enhancement value set and the second enhancement value set; Input the sample channel features from the training samples into the first initial recommendation model to obtain the initial channel recommendation result output by the first initial recommendation sub-model and the first evaluation value set output by the first initial evaluation sub-model for the initial channel recommendation result; Input the sample content features from the training samples into the second initial recommendation model to obtain the initial content recommendation result output by the second initial recommendation sub-model and the second evaluation value set output by the second initial evaluation sub-model for the initial content recommendation result; Update the parameters of the first initial recommendation sub-model based on the first evaluation value set; Update the parameters of the second initial recommendation sub-model based on the second evaluation value set; Based on the first enhanced value set and the first evaluated value set, obtain the channel loss function; Based on the second augmentation value set and the second evaluation value set, obtain the content loss function; Calculate the target loss function based on the channel loss function and the content loss function; update the parameters of the first initial evaluation sub-model and the second initial evaluation sub-model based on the target loss function.
21. The apparatus according to claim 20, characterized in that, The training samples also include sequences of recommended content; the training unit is further used for: Based on the initial content recommendation results and the sample recommendation content sequence in the training samples, at least one of the click-through rate loss function and the similarity loss function is obtained; The target loss function is calculated based on at least one of the click-through rate loss function and the similarity loss function, as well as the channel loss function and the content loss function.
22. The apparatus according to claim 20 or 21, characterized in that, The training unit is also used for: Based on the feedback information in the training samples, obtain at least one of the following: reading time information, diversity information, and novelty information of any sample recommended content in the sample recommended content sequence, as well as the click information of any sample recommended content; Based on the click information of the recommended content of any sample, obtain the first enhancement value corresponding to the recommended content of any sample; Based on at least one of the reading time information, diversity information, and novelty information of any sample recommended content, and the click information of any sample recommended content, obtain the second enhancement value corresponding to any sample recommended content; The set of first enhancement values corresponding to the recommended content of each sample is taken as the first enhancement value set; The set of second enhancement values corresponding to the recommended content of each sample is taken as the second enhancement value set.
23. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one piece of program code, the at least one piece of program code being loaded and executed by the processor to implement the content recommendation method as described in any one of claims 1 to 6, or to implement the content recommendation method as described in any one of claims 7 to 11.
24. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one piece of program code, which is loaded and executed by a processor to implement the content recommendation method as described in any one of claims 1 to 6, or to implement the content recommendation method as described in any one of claims 7 to 11.
Citation Information
Patent Citations
Insurance product recommendation method and apparatus, computer device and storage medium
CN109165983A
Information flow recommendation method and device, computer device and storage medium
CN110825956A