Project Recommendation Method and Apparatus, Project Recommendation Model, Medium and Electronic Device

By introducing a filtering mechanism based on reinforcement learning in the project recommendation system, the problem of interfering with information in the user's browsing data affecting the accuracy of recommendation is solved, and a more efficient project recommendation effect is achieved.

CN114741586BActive Publication Date: 2025-06-20TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110025336.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-08
Publication Date
2025-06-20
Estimated Expiration
2041-01-08

AI Technical Summary

Technical Problem

The accuracy of project recommendations in the prior art is low, mainly due to the presence of interfering information in the user's browsing data, which affects the prediction effect of the model.

Method used

The project recommendation method based on reinforcement learning is adopted, and the first reinforcement learning model is used to determine whether the user browsing data needs to be filtered. If necessary, the second reinforcement learning model is used to filter out the interference information, and then the project recommendation is based on the filtered optimization project list.

Benefits of technology

It improves the accuracy of project recommendations, effectively filters interfering information, enhances the prediction ability of the recommendation model and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114741586B_ABST
    Figure CN114741586B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of artificial intelligence technology, and provides a project recommendation method and apparatus, a model, a medium, and a device based on reinforcement learning. The method includes: predicting a recommendation value for a target project according to a vector representation of the target project and historical browsing projects in a current project list, where the current project list consists of projects including user browsing behaviors; obtaining relative features between the target project and the historical browsing projects; based on a reinforcement learning model, determining whether filtering processing needs to be performed on the current project list according to the relative features and the recommendation value; in response to the need to perform filtering processing on the current project list, determining projects to be filtered in the current project list based on the relative features according to the reinforcement learning model, so as to perform project recommendation based on an optimized project list after filtering. This technical solution can quickly and effectively identify whether the current project list needs to be filtered and determine the projects to be filtered, which is beneficial to improving the accuracy of project recommendation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology. Specifically, it relates to a project recommendation method and device based on reinforcement learning, a project recommendation model, a computer-readable storage medium for implementing the above-mentioned project recommendation method based on reinforcement learning, and an electronic device. Background Art

[0002] With the development of artificial intelligence technology, a recommendation system predicts the content that a user may like through a prediction model, so as to recommend content that meets the user's personalized needs for the user. For example, recommend songs that the user may like to music listeners, and recommend content that they may like to video / short video viewers.

[0003] In related technologies, generally, historical browsing data of multiple users is first obtained, and then a prediction model is trained based on the historical browsing data. Further, the trained prediction model is used to recommend items that the target user may like to the target user.

[0004] However, when project recommendations are made through the solutions provided by related technologies, the recommendation accuracy is relatively low.

[0005] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure. Summary of the Invention

[0006] An object of the present disclosure is to provide a project recommendation method and device based on reinforcement learning, a project recommendation model, a computer-readable storage medium for implementing the above method, and an electronic device, which filter out interference information in user browsing data and perform project recommendations based on the filtered user browsing data, so as to improve the recommendation accuracy to at least a certain extent.

[0007] Other features and advantages of the present disclosure will become apparent through the following detailed description, or will be partially learned through the practice of the present disclosure.

[0008] According to one aspect of the present disclosure, there is provided a project recommendation method based on reinforcement learning, including: predicting a recommendation value for a target project according to a vector representation of the target project and historical browsing projects in a current project list, where the current project list consists of projects including user browsing behaviors; obtaining relative features between the target project and the historical browsing projects; determining, based on a first reinforcement learning model, whether it is necessary to perform filtering processing on the current project list according to the relative features and the recommendation value; and, in response to the need to perform filtering processing on the current project list, determining, based on a second reinforcement learning model, items to be filtered in the current project list according to the relative features, so as to perform project recommendations based on the optimized project list after filtering.

[0009] According to one aspect of the present disclosure, there is provided a project recommendation model, including: a recommendation value prediction model, a first reinforcement learning model, and a second reinforcement learning model.

[0010] Wherein, the above-mentioned recommendation value prediction model is configured to: predict the recommendation value for the target project according to the vector representation of the target project and the vector representation of the user, wherein the vector representation of the user is determined according to the current project list, and the current project list is composed of projects including user browsing behaviors; the above-mentioned first reinforcement learning model is configured to: determine whether to filter the current project list according to the relative feature and the recommendation value, wherein the relative feature is the relative feature between the target project and the historical browsing projects; and, the above-mentioned second reinforcement learning model is configured to: in response to the need to filter the current project list, determine the items to be filtered in the current project list according to the relative feature, so that the recommendation value prediction model performs project recommendation based on the filtered optimized project list.

[0011] According to one aspect of the present disclosure, there is provided a project recommendation device based on reinforcement learning, including: a prediction module, an acquisition module, a judgment module, and a filtering processing module.

[0012] Wherein, the above-mentioned prediction module is configured to: predict the recommendation value for the target project according to the vector representation of the target project and the historical browsing projects in the current project list, wherein the current project list is composed of projects including user browsing behaviors; the above-mentioned acquisition module is configured to: acquire the relative feature between the target project and the historical browsing projects; the above-mentioned judgment module is configured to: based on the first reinforcement learning model, determine whether to filter the current project list according to the relative feature and the recommendation value; and, the above-mentioned filtering processing module is configured to: in response to the need to filter the current project list, based on the second reinforcement learning model, determine the items to be filtered in the current project list according to the relative feature, so as to perform project recommendation based on the filtered optimized project list.

[0013] In some embodiments of the present disclosure, based on the foregoing solution, the acquisition module is specifically configured to: respectively acquire the individual relative features between the target project and the multiple historical browsing projects; and, determine the overall relative feature between the target project and the current project list according to the individual relative features.

[0014] In some embodiments of the present disclosure, based on the foregoing solution, the filtering processing module is specifically configured to: input the overall relative feature and the recommended value into the first reinforcement learning model to determine the overall relevance between the target item and the current item list based on the first reinforcement learning model; and determine whether to perform filtering processing on the current item list based on the overall relevance.

[0015] In some embodiments of the present disclosure, based on the foregoing solution, the reinforcement learning-based item recommendation device further includes: a first training module.

[0016] Wherein, the first training module is configured to: update the model parameters of the first reinforcement learning model according to the first reward of the j-th time, determine whether to perform the j-th filtering processing according to the first reinforcement learning model after updating the model parameters, where j is an integer greater than 1; in response to the need to perform the j-th filtering processing, determine the j-th recommended value for the target item according to the j-th optimized item list obtained after the j-th filtering processing; determine the (j + 1)-th first reward for the first reinforcement learning model according to the j-th recommended value and the (j - 1)-th recommended value, so as to update the model parameters of the first reinforcement learning model according to the (j + 1)-th first reward.

[0017] In some embodiments of the present disclosure, based on the foregoing solution, the filtering processing module is further specifically configured to: obtain the j-th overall relative feature between the target item and the historical browsing items in the j-th optimized item list; determine the j-th state vector according to the j-th recommended value and the j-th overall relative feature; perform a rectified linear unit processing on the j-th state vector based on the first reinforcement learning model after updating the model parameters with the first reward of the j-th time to obtain the j-th rectified vector; and process the j-th rectified vector based on the action parameters of the first reinforcement learning model to obtain the overall relevance between the target item and the current item list.

[0018] In some embodiments of the present disclosure, based on the foregoing solution, the reinforcement learning-based item recommendation device further includes: an item recommendation module.

[0019] Wherein, the item recommendation module is configured to: in response to the need not to perform the j-th filtering processing, perform item recommendation based on the (j - 1)-th optimized item list filtered by the (j - 1)-th filtering processing.

[0020] In some embodiments of the present disclosure, based on the foregoing solution, the obtaining module is specifically configured to: respectively obtain the individual relative features between the target item and the multiple historical browsing items.

[0021] In some embodiments of the present disclosure, based on the foregoing solution, the filtering processing module is specifically configured to: input the t-th individual relative feature corresponding to the t-th historical browsing item into the second reinforcement learning model to determine the individual association degree between the target item and the t-th historical browsing item based on the model parameters of the second reinforcement learning model, where t is an integer not greater than the number of historical browsing items in the t-th optimized item list; and determine whether the t-th historical browsing item is an item to be filtered based on the individual association degree.

[0022] In some embodiments of the present disclosure, based on the foregoing solution, the reinforcement learning-based item recommendation device further includes: a second training module.

[0023] Wherein, the second training module is configured to: update the model parameters of the second reinforcement learning model according to the k-th second reward, and determine the item to be filtered according to the second reinforcement learning model after updating the model parameters, where k is an integer greater than 1; determine the k-th recommendation value for the target item according to the k-th optimized item list obtained after the k-th filtering process; determine the first part of the reward according to the k-th recommendation value and the (k - 1)-th recommendation value, where the (k - 1)-th recommendation value is determined according to the (k - 1)-th optimized item list obtained after the (k - 1)-th filtering process; calculate the similarity between each historical browsing item in the k-th optimized item list and the target item respectively, and determine the second part of the reward according to the multiple similarities; and determine the (k + 1)-th second reward for the second reinforcement learning model according to the first part of the reward and the second part of the reward, so as to update the model parameters of the second reinforcement learning model according to the (k + 1)-th second reward.

[0024] In some embodiments of the present disclosure, based on the foregoing solution, the filtering processing module is further specifically configured to: obtain the t-th individual relative feature between the target item and the t-th historical browsing item; determine the t-th state vector according to the t-th individual relative feature; perform a rectified linear unit processing on the t-th state vector based on the second reinforcement learning model after updating the model parameters according to the k-th second reward to obtain the t-th rectified vector; and process the t-th rectified vector based on the action parameters of the second reinforcement learning model to obtain the individual association degree between the target item and the t-th historical browsing item.

[0025] In some embodiments of the present disclosure, based on the foregoing solution, the prediction module includes a user vector representation unit and a recommendation value determination unit.

[0026] Among them, the above user vector representation unit is configured to: determine the vector representation of the user according to the historical browsing items in the above current project list; the above recommended value determination unit is configured to: predict the recommended value for the above target project based on the vector representation of the above user and the vector representation of the above target project.

[0027] In some embodiments of the present disclosure, based on the foregoing solution, the above user vector representation unit is specifically configured to: average the vector representations of the historical browsing items in the j-th optimized project list obtained after the j-th filtering process to obtain the j-th vector representation for the user, so as to be used to predict the j-th recommended value for the above target project.

[0028] In some embodiments of the present disclosure, based on the foregoing solution, the above acquisition module is configured to: acquire one or more of the following information between the above target project and the above historical browsing items: cosine distance, dot product of vectors, and similarity between feature tags of items, to obtain the above relative features.

[0029] According to one aspect of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method for project recommendation based on reinforcement learning described in the above first aspect is implemented.

[0030] According to one aspect of the present disclosure, there is provided an electronic device, including: one or more processors; and a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method for project recommendation based on reinforcement learning described in the above first aspect.

[0031] According to one aspect of the present disclosure, there is provided a computer program product or a computer program, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method for project recommendation based on reinforcement learning provided in the above various embodiments.

[0032] It can be seen from the above technical solutions that the method for project recommendation based on reinforcement learning, the device for project recommendation based on reinforcement learning, the computer-readable storage medium, and the electronic device in the exemplary embodiments of the present disclosure at least have the following advantages and positive effects:

[0033] In the technical solutions provided by some embodiments of the present disclosure, the recommended value for the target item is first predicted through a prediction model. Then, based on the first reinforcement learning model, according to the relative features between the target item and the historical browsing items and the above-mentioned recommended value, the relevance between the target video and the current item list at the overall level is judged. If the overall relevance is small, it indicates that there is more interference information in the current item list, and then the item list needs to be filtered. Further, based on the second reinforcement learning model, according to the relative features between the target item and the historical browsing items, the relevance between the target video and each historical browsing item in the current item list at the individual level is judged. If the relevance with a certain historical browsing item is small, it indicates that the historical browsing item belongs to interference information, and the interference information is filtered out. It can be seen that the present technical solution can quickly and effectively distinguish whether the current item list needs to be filtered and determine the items to be filtered based on judging the relevance between the target video and the historical browsing items in the item list at two levels. And predicting the recommended value for the target item based on the item list after filtering out the interference information helps improve the accuracy of model prediction.

[0034] At the same time, the present technical solution performs item recommendation based on the optimized item list after filtering, that is, the present technical solution enhances the prediction model for determining the recommended value based on the above-mentioned first reinforcement learning model and the second reinforcement learning model, which is beneficial to further improving the accuracy of item recommendation.

[0035] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts. In the drawings:

[0037] Figure 1 Shows a schematic structural diagram of an item recommendation model in an exemplary embodiment of the present disclosure.

[0038] Figure 2 Shows a schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of the present disclosure can be applied.

[0039] Figure 3 Shows a schematic diagram of an item list in an exemplary embodiment of the present disclosure.

[0040] Figure 4Schematic flowchart of a project recommendation method based on reinforcement learning in an exemplary embodiment of the present disclosure.

[0041] Figure 5 Schematic flowchart of a method for predicting project recommendation values in an exemplary embodiment of the present disclosure.

[0042] Figure 6 Schematic flowchart of a project recommendation method based on reinforcement learning in another exemplary embodiment of the present disclosure.

[0043] Figure 7 Schematic flowchart of a method for training a first reinforcement learning model in an exemplary embodiment of the present disclosure.

[0044] Figure 8 Schematic flowchart of a method for training a second reinforcement learning model in an exemplary embodiment of the present disclosure.

[0045] Figure 9 Schematic structural diagram of a project recommendation device based on reinforcement learning in an exemplary embodiment of the present disclosure.

[0046] Figure 10 Schematic structural diagram of an electronic device in an exemplary embodiment of the present disclosure. Detailed implementation manners

[0047] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.

[0048] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or can be implemented using other methods, components, devices, steps, etc. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present disclosure.

[0049] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0050] The flowcharts shown in the accompanying drawings are merely illustrative and not necessarily inclusive of all content and operations / steps, nor are they necessarily to be executed in the order described. For example, some operations / steps may be decomposed, while some operations / steps may be combined or partially combined, so the actual execution order may change according to the actual situation.

[0051] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.

[0052] Artificial intelligence technology is an interdisciplinary subject that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0053] Reinforcement Learning (RL), also known as re-inforcement learning, evaluation learning, or enhancement learning, is one of the paradigms and methodologies of machine learning, used to describe and solve the problem of an agent achieving maximum reward or a specific goal through learning strategies during the interaction with the environment. It obtains learning information by receiving rewards (feedback) from the environment for actions and updates the model parameters.

[0054] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, transfer learning, inductive learning, and rote learning.

[0055] The solution provided by the embodiments of the present disclosure relates to technologies such as reinforcement learning and machine learning in artificial intelligence, and is specifically described through the following embodiments:

[0056] In the related art, generally, historical browsing data of multiple users is first obtained, and then a prediction model is trained based on the historical browsing data, that is, it depends on analyzing the historical browsing data of users to determine user interest points. However, when the user browsing volume is large or there is interference information in the user historical browsing data, it will lead to interference information in the data used to train the prediction model, resulting in unnecessary computational complexity and being not conducive to the prediction accuracy of the prediction model.

[0057] In view of the technical problems existing in the related art, the present technical solution provides a project recommendation method, device, recommendation model, medium and device based on reinforcement learning. First, Figure 1 Schematically shows a schematic diagram of a project recommendation model provided according to an exemplary embodiment of the present disclosure. Refer to Figure 1 , the project recommendation model includes: a recommendation value prediction model 11, a first reinforcement learning model 12, and a second reinforcement learning model 13.

[0058] Specifically, the above-mentioned recommendation value prediction model 11 predicts the recommendation value A for the target project according to the vector representation of the target project and the vector representation of the user. Among them, the project list consists of projects containing user browsing behaviors, and the above-mentioned vector representation of the user can be determined according to the project list.

[0059] Furthermore, the relative feature B between the target project and the historical browsing projects is obtained. In the above-mentioned first reinforcement learning model, it is determined whether to perform filtering processing on the project list according to the relative feature B and the recommendation value A for the target project. In the case of determining to perform filtering processing C on the current project list, based on the second reinforcement learning model 13, the projects to be filtered are determined from the current project list according to the relative feature B, so that the recommendation value prediction model 11 performs project recommendation based on the optimized project list D after filtering.

[0060] Figure 2 Shows a schematic diagram of an exemplary system architecture to which the technical solution of the embodiments of the present disclosure can be applied.

[0061] As Figure 2 shown, the system architecture 100 may include a terminal 110, a network 120, and a server side 130. Among them, the terminal 110 and the server side 130 are connected through the network 120.

[0062] The terminal 110 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The network 120 can be a communication medium of various connection types capable of providing a communication link between the terminal 110 and the server side 130. For example, it can be a wired communication link, a wireless communication link, or an optical fiber cable, etc. This application does not make any restrictions here. The server 130 can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0063] Specifically, the server side 130 can provide the training of the project recommendation model. For example, the training of the recommendation value prediction model 11, the training of the first reinforcement learning model 12, the training of the second reinforcement learning model 13, etc. It can also store the trained project recommendation model, and store the recommendation value prediction model 11, the first reinforcement learning model 12, and the second reinforcement learning model 13 to the server side 130.

[0064] In addition, the user can generate browsing behaviors on the terminal 120 for projects (such as videos, songs, goods, news, etc.), thereby generating a project list (such as Figure 3 ). Referring to Figure 3 , taking videos as an example of the project, referring to Figure 3 the "video browsing record" (project list) shown, which can be various types of videos that the user clicks, browses, collects, forwards, or comments on. Exemplarily, the server side 130 obtains the project containing the user's browsing behavior to obtain a project list. Furthermore, the following steps are performed on the server side 130: predicting the recommendation value for the target project based on the vector representation of the target project and the historical browsing projects in the current project list, and obtaining the relative features between the target project and the historical browsing projects; based on the first reinforcement learning model, determining whether to filter the current project list according to the relative features and the recommendation value. If it is necessary to filter the current project list, then based on the second reinforcement learning model, determining the projects to be filtered in the current project list according to the relative features, so as to perform project recommendation based on the optimized project list after filtering.

[0065] The project recommendation method based on reinforcement learning in the embodiments of the present disclosure can also be applied to the terminal. The present disclosure does not make any special limitations on this. The embodiments of the present disclosure mainly take the application of the project recommendation method based on reinforcement learning to the server side 130 as an example for illustration.

[0066] Next, the project recommendation method based on reinforcement learning provided by this technical solution is introduced. Among them, Figure 4A flowchart of a project recommendation method based on reinforcement learning in an exemplary embodiment of the present disclosure is shown. Figure 4 , the project recommendation method based on reinforcement learning provided by this embodiment includes:

[0067] Step S410, predicting a recommendation value for the target item based on a vector representation of the target item and historically browsed items in a current item list, wherein the current item list is composed of items containing user browsing behaviors;

[0068] Step S420, obtaining relative features between the target item and the historical browsing items;

[0069] Step S430, based on the first reinforcement learning model, determining whether filtering is required for the current item list according to the relative features and the recommendation value; and,

[0070] Step S440, in response to the need to filter the current project list, based on the second reinforcement learning model, determine the projects to be filtered in the current project list according to the relative features, so as to make project recommendations based on the filtered optimized project list.

[0071] Figure 4 The reinforcement learning-based item recommendation scheme provided in the illustrated embodiment is applicable to the prediction and recommendation of short videos (generally short-duration video content, such as micro-recording vlogs, short videos, etc.), videos (such as film and television works), music, commodities, and news. The "target item" in each of the following embodiments can be any item. For example, in the short video recommendation scheme, the "target item" can be any short video. This scheme first determines the correlation between the target video and the current project list at the overall level. If the overall correlation is small, it means that the current project list contains more interference information, and the project list needs to be filtered. Then, the correlation between the target video and each historical browsing item in the current project list at the individual level is determined. If the correlation with a certain historical browsing item is small, it means that the historical browsing item belongs to interference information, and the interference information is filtered out. It can be seen that the technical scheme is based on judging the correlation between the target video and the historical browsing items in the project list at two levels, and can quickly and effectively distinguish whether the current project list needs to be filtered and determine the items to be filtered. Predicting the recommended value of the target item based on the list of items that have filtered out interference information helps improve the accuracy of the model's prediction.

[0072] Meanwhile, this technical solution recommends projects based on the filtered optimized project list. That is to say, this technical solution enhances the prediction model for determining the recommendation value based on the above-mentioned first reinforcement learning model and the second reinforcement learning model, which is conducive to further improving the accuracy of project recommendation.

[0073] In the following embodiments, "short video" is taken as an example to Figure 4 elaborate in detail the specific implementation manners of each step of the illustrated embodiments:

[0074] In an exemplary embodiment, a project list is first obtained, where the project list consists of projects including user browsing behaviors. For example, the short videos that the user has performed the following operations on in the past week: clicking, browsing, favoriting, forwarding, or commenting, etc., can be used as the above-mentioned project list.

[0075] It should be noted that in the training processes of the first reinforcement learning model 12 and the second reinforcement learning model 13, the recommendation value prediction model 11 is continuously enhanced. That is to say, the "project list" used by the recommendation value prediction model 11 for prediction is continuously updated. Therefore, the "current project list" is used to refer to the original project list that has not been filtered, and can also refer to the optimized project list after the j-th filtering process (which can be denoted as the "j-th optimized project list").

[0076] In step S410, the recommendation value for the target project is predicted through the recommendation value prediction model 11. As a specific implementation manner, Figure 5 An exemplary method flow diagram for determining the recommendation value in an embodiment of the present disclosure is shown. Refer to Figure 5 , this method includes step S510 and step S520.

[0077] In step S510, the vector representation of the user is determined according to the historical browsing projects in the current project list.

[0078] In an exemplary embodiment, the vector representation of the user is constructed by learning the interaction information between the user and the historical browsing projects. For example, the vector representation of the user is determined according to the following formula:

[0079]

[0080] where q u is the vector representation of the user, t u is the number of projects in the current project list, and p t u is the vector representation of the t-th project in the current project list.

[0081] It should be noted that the specific implementation method for determining the vector representation of the user in the above step S510 is not limited to the above one, and can also be a matrix factorization method, that is, the vector representation of the user is obtained by decomposing the user video interaction matrix; it can also be graph network embedding, that is, by randomly walking on the user video interaction graph to construct a training corpus, and then learning the vector representation of the user through word2vec, etc. This embodiment does not limit this.

[0082] In step S520, based on the vector representation of the user and the vector representation of the target item, the recommended value for the target item is predicted.

[0083] In an exemplary embodiment, the recommended value for the target item c i is predicted according to the following formula:

[0084]

[0085] where ε u represents the current item list, including multiple historical browsing items, q u is the vector representation of the user, c i represents the target item, p i is the vector representation of the target item, and P(y = 1|ε u , c i ) represents the recommended value for the target item c i .

[0086] It should be noted that the specific implementation method for determining the recommended value for the target item in the above step S410 is not limited to the above one prediction model, and can also be other prediction models, such as a neural network model. This embodiment does not limit this.

[0087] In step S420, the relative features between the above target item and each historical browsing item in the current item list are obtained. Exemplarily, in this technical solution, this "relative feature" is used to measure the correlation (which can also be used as a difference degree) between the target item and the current item list / a certain historical browsing item in the list. Referring to Figure 3 , the "video browsing record" (item list) includes: video A, video B, and video F in the food category, short video C in the sports category, video E in the beauty category, and video D in the game category. Suppose the target item is a food video. Then, the correlation between this target item and video A is relatively strong, while the correlation between it and video C, video D, or video E is relatively weak. Among them, the items with a relatively strong correlation with the target item are beneficial to determining the recommended value for this target item. On the contrary, the items with a relatively weak correlation with the target item are not beneficial to determining the recommended value for this target item, that is, they can be filtered as "interference information" to improve the prediction accuracy of the recommended value for the target item.

[0088] To avoid the problem of large computational complexity caused by directly comparing the relevance between the target project and each historical browsing project, in this technical solution, the relevance between the target project and the entire project record is determined first. Specifically, the "overall relative feature" between the target project and the current project list is obtained. If this overall relative feature shows that the relevance between the target project and the current project list is weak / greatly different, it indicates that there is a lot of interfering information in the current project list, and the project list needs to be filtered.

[0089] Furthermore, the "individual relative feature" between the target project and a certain historical browsing project in the current project list is obtained. If this individual relative feature shows that the relevance between the target project and the historical browsing project is weak / greatly different, it indicates that the historical browsing project is a project to be filtered. On the contrary, if this individual relative feature shows that the relevance between the target project and the historical browsing project is strong / less different, it indicates that the historical browsing project can be used to predict the recommended value of the target project, and then it is retained.

[0090] It can be seen that this technical solution can quickly and effectively distinguish whether the current project list needs to be filtered and determine the projects to be filtered based on judging the relevance between the target video and the historical browsing projects in the project list at two levels.

[0091] In an exemplary embodiment, Figure 6 shows a schematic flowchart of a method for determining that the current project list needs to be filtered in an exemplary embodiment of the present disclosure. Refer to Figure 6 , this method includes the following steps:

[0092] In step S610, the individual relative features between the target project and the multiple historical browsing projects are obtained respectively.

[0093] In an exemplary embodiment, in order to obtain the relevance / difference degree between the target project c i (vector representation is p i ) and the historical browsing projects (vector representation is p t u , where t ranges from positive integers not greater than the number of projects in the current project list) in the current project list, one or several of the following information between the target project and the historical browsing projects can be calculated: cosine distance The dot product of vectors (p i ·p t u ) and the similarity between the feature tags of the projects, so as to obtain the individual relative feature between the target project and a certain historical browsing project.

[0094] In an exemplary embodiment, the Euclidean distance can be used to determine the similarity between feature tags. Specifically, the feature tags of the target item and any historical browsing item can be obtained separately first. Among them, the feature tags are used to reflect the characteristics of the item itself. For example, if the target item is a sports news, the keywords of the news can be used as its feature tags. For example, its feature tags can include: sports, basketball, Kobe, champion. Similarly, the keywords of the historical browsing news, that is, its feature tags are: sports, table tennis, men's singles, Ma Long. Further, calculate the Euclidean distance between the feature tags of the target item and the feature tags of the historical browsing item. Exemplarily, the smaller the value of the obtained Euclidean distance, the greater the similarity between the feature tags of the target item and the feature tags of the historical browsing item. Conversely, the smaller the similarity between the two.

[0095] In step S620, the overall relative feature between the target item and the current item list is determined according to the individual relative feature.

[0096] In an exemplary embodiment, the average value of the individual relative features corresponding to each historical browsing item in the current item list is calculated as the overall relative feature between the target item and the current item list. By taking the average, the correlation between the current item list and the target item is reflected as a whole.

[0097] In an exemplary embodiment, before performing the prediction process based on the first reinforcement learning model, first pass through Figure 7 Introduce the training process of the first reinforcement learning model. Refer to Figure 7 :

[0098] In step S710, the model parameters of the first reinforcement learning model are updated according to the first reward of the j-th time, where j is an integer greater than 1.

[0099] Among them, the above "first reward" is the reward value (reward) of the first reinforcement learning model in this iteration process. The determination of the first reward will be introduced in the corresponding embodiment part of step S740.

[0100] In step S720, it is determined whether to perform the j-th filtering process according to the first reinforcement learning model after updating the model parameters.

[0101] Exemplarily, the specific processing method for the first reinforcement learning model to determine whether to perform the j-th filtering process will be introduced in the corresponding embodiment part of step S640.

[0102] In response to the need not to perform the j-th filtering process, it indicates that the current project list has a strong correlation with the target project and there is no need to filter the current project list any further. That is, using the optimized project list after the previous filtering process for project recommendation can meet the requirements of recommendation accuracy. Then, step S730' is executed: Based on the (previous) j-1 optimized project list filtered by the (previous) j-1-th filtering process, project recommendation is performed.

[0103] In response to the need to perform the j-th filtering process, it indicates that the current project list has a weak correlation with the target project and the current project list needs to be filtered again. Then, step S730 and step S740 are executed. In step S730, according to the j-th optimized project list (denoted as ) obtained after the j-th filtering process, the j-th recommendation value (denoted as ) for the target project is determined.

[0104] Exemplarily, in the case of responding to the need to perform the 2nd (i.e., j = 2) filtering process, for example, the project list (denoted as the "j-th optimized project list") after the 2nd filtering process can be expressed as [project a, project b, project x, project y]. Then, the vector representation of [project a, project b, project x, project y] is input into the recommendation value prediction model 11, and the j-th recommendation value for the target project is determined based on formula (2).

[0105] In step S740, the first reward for the (j + 1)-th time of the first reinforcement learning model is determined according to the j-th recommendation value and the (j - 1)-th recommendation value.

[0106] Among them, the (j - 1)-th recommendation value (denoted as "P(y = 1|ε u ,c i )") is determined based on the manner shown in step S730 according to the (j - 1)-th optimized project list obtained after the (j - 1)-th filtering process.

[0107] Exemplarily, first, the sub-reward for the first reinforcement learning model is determined according to the following formula (3).

[0108]

[0109] Among them, R(s h t ,a h t ) represents the above-mentioned first reward, s h t represents the state vector of the first reinforcement learning model in the j-th iteration process, a h t represents the action parameter of the first reinforcement learning model in the j-th iteration process, represents based on the j-th optimized vector list The determined recommended value for the target item c i , P(y = 1|ε u , c i ) represents the recommended value for the target item c determined based on the (j - 1)-th optimization vector list ε u The determined recommended value for the target item c i , and t u represents the number of items in the j-th optimization vector list.

[0110] Then, according to the following formula (4), combined with the above sub-reward R(s h t , a h t ), the first reward for the (j + 1)-th time of the first reinforcement learning model is determined.

[0111]

[0112] where Θ h are the parameters of the first reinforcement learning model.

[0113] Reference Figure 7 , assign j + 1 to j, and assign j to j - 1, and continue to execute step S710 to update the model parameters of the first reinforcement learning model according to the first reward for the (j + 1)-th time.

[0114] Furthermore, based on the first reinforcement learning model after iterative processing, determine the overall relevance between the target item and the current item list: continue to refer to Figure 6 , in step S630, input the overall relative feature and the recommended value into the first reinforcement learning model to determine the overall relevance between the target item and the current item list based on the first reinforcement learning model; and, in step S640, determine whether filtering processing needs to be performed on the current item list based on the overall relevance.

[0115] In an exemplary embodiment, the specific implementation manners of step S630 and step S640 include:

[0116] S1. Obtain the j-th overall relative feature between the target item and the historical browsing items in the j-th optimized item list. The specific implementation manner of step S1 is the same as that of step S620 and will not be elaborated here.

[0117] S2. Determine the j-th state vector s based on the j-th recommended value h t and the j-th overall relative feature.

[0118] S3. After updating the model parameters based on the first reward at the j-th time, the first reinforcement learning model uses formula (5) to perform a rectified linear unit (ReLU) operation on the j-th state vector to obtain the j-th rectified vector.

[0119] H h t = ReLU(W h 1s h t + b h 1) (5)

[0120] where H h t represents the j-th rectified vector in the first reinforcement learning model, ReLU() is the rectified linear unit function, s h t represents the j-th state vector of the first reinforcement learning model, W h 1 and b h 1 are the rectification parameters updated based on the first reward at the j-th time.

[0121] S4. Using formula (6), based on the action parameter a h t of the first reinforcement learning model, process the j-th rectified vector to obtain the overall correlation degree between the target item and the current item list.

[0122] π(s h t , a h t ) = P(a h t | s h t , Θ h ) = a h t σ(W h 2H h t ) + (1 - a h t )(1 - σ(W h 2H h t )) (6)

[0123] where Θ h is the model parameter of the first reinforcement learning model, W h 2 is the activation parameter updated based on the first reward at the j-th time, σ(·) is the sigmoid function, and the action parameter a h t takes values of 0 or 1. When the action parameter a h t takes the value of 0, π(sh t , a h t ) Determine the probability that the current item list does not need to be filtered. When the action parameter a h t takes the value of 1, π(s h t , a h t ) Determine the probability that the current item list needs to be filtered.

[0124] Thus, when the action parameter a h t takes the value of 1, it can be determined whether the current item list is filtered according to the output probability value of formula (6).

[0125] When it is determined that the current item list needs to be filtered, the items to be filtered can be determined from the current item list through the above second reinforcement learning model. First, the training process of the second reinforcement learning model will be introduced below, and then the specific implementation of determining the items to be filtered from the current item list based on the second reinforcement learning model will be introduced.

[0126] Figure 8 shows a schematic flow chart of the training method of the second reinforcement learning model in an exemplary embodiment of the present disclosure. Refer to Figure 8 , the method includes steps S810 - step S850.

[0127] In step S810, update the model parameters of the second reinforcement learning model according to the k-th second reward, where k is an integer greater than 1.

[0128] Among them, the above "second reward" is the reward value (reward) for the second reinforcement learning model in this iteration process. The determination of the second reward will be introduced in steps S840 and S850.

[0129] In step S820, according to the k-th optimized item list (denoted as ) obtained after the k-th filtering process, determine the k-th recommendation value for the target item

[0130] Exemplarily, in response to the need for the 2nd (i.e., k = 3) filtering process, for example, the item list after the 3rd filtering process (denoted as "the k-th optimized item list") can be expressed as [item o, item p, item q]. Then, the vector representation of [item o, item p, item q] is input into the recommendation value prediction model 11, and the k-th recommendation value for the target item is determined based on formula (2).

[0131] In step S830, a first part of the reward is determined according to the k-th recommended value and the (k-1)-th recommended value, where the (k-1)-th recommended value (denoted as "P(y = 1|ε u ,c i )") is determined according to the (k-1)-th optimized item list obtained after the (k-1)-th filtering process.

[0132] Exemplarily, the first part of the reward is determined according to the following formula (7).

[0133]

[0134] where R(s l t ,a l t ) represents the first part of the first reward, s l t represents the state vector in the k-th iteration process, a l t represents the action parameter in the k-th iteration process, represents the recommended value of the target item c determined based on the k-th optimized vector list i of P(y = 1|ε u ,c i ) represents the recommended value of the target item c u determined based on the (k-1)-th optimized vector list ε i , t u represents the number of items in the k-th optimized vector list.

[0135] It can be seen that the first part of the reward determined by formula (7) can only be determined after the end of the k-th iteration process. The reward generated during this iteration process, that is, the second part of the reward, will be introduced below.

[0136] In step S840, the similarity between each historical browsing item in the k-th optimized item list and the target item is calculated respectively, and the second part of the reward G(s l t ,a l t ) is determined according to the multiple similarities. Suppose the k-th optimized item list is expressed as [item o, item p, item q], the cosine distances between item o, item p, and item q and the target item are calculated respectively, and the second part of the reward G(s l t ,a l t ) is determined according to the cosine distances.

[0137] The above-mentioned second part of the reward G(s lt , a l t ) It does not need to be determined after the end of this iteration, so it is beneficial to accelerate the training of the second reinforcement learning model.

[0138] In step S850, according to the following formula (8), combining the first part of the reward R(s l t , a l t ) and the second part of the reward G(s l t , a l t ) to determine the first reward for the (k + 1)-th time of the second reinforcement learning model, so as to update the model parameters of the second reinforcement learning model according to the first reward for the (k + 1)-th time.

[0139]

[0140] Among them, Θ l is the parameter of the second reinforcement learning model.

[0141] Refer to Figure 8 , assign k + 1 to k, and continue to execute step S810 to update the model parameters of the second reinforcement learning model according to the second reward for the (k + 1)-th time.

[0142] Refer to Figure 6 , including step S610, step S620', and step S630'. Among them, the specific implementation manner of step S610 will not be elaborated.

[0143] Furthermore, based on the second reinforcement learning model after iterative processing, determine the individual association degree between the target item and the historical browsing items in the current item list: continue to refer to Figure 6 , including step S610, step S620', and step S630'. Among them, the specific implementation manner of step S610 will not be elaborated.

[0144] In step S620', input the t-th individual relative feature corresponding to the t-th historical browsing item into the second reinforcement learning model to determine the individual association degree between the target item and the t-th historical browsing item based on the model parameters of the second reinforcement learning model; and, in step S630', determine whether the t-th historical browsing item is an item to be filtered based on the individual association degree.

[0145] In an exemplary embodiment, the specific implementation manners of steps S620' and S630' include:

[0146] S1. Obtain the t-th individual relative feature between the target project and the t-th historical browsing project. The specific implementation of step S1 is the same as that of step S610 and will not be elaborated here.

[0147] S2. Determine the t-th state vector s according to the t-th individual relative feature l t .

[0148] S3. Based on the second reinforcement learning model after updating the model parameters with the k-th second reward, linearly rectify the t-th state vector using formula (9) to obtain the t-th rectified vector.

[0149] H l t = ReLU(W l 1s l t + b l 1) (9)

[0150] where H l t represents the t-th rectified vector in the second reinforcement learning model, ReLU() is the linear rectification function, s l t represents the t-th state vector of the second reinforcement learning model, W l 1 and b l 1 are the rectification parameters updated based on the k-th reward.

[0151] S4. Using formula (10), based on the action parameter a of the second reinforcement learning model l t , process the t-th rectified vector to obtain the individual association degree between the target project and the t-th historical browsing project.

[0152] π(s l t , a l t ) = P(a l t | s l t , Θ l ) = a l t σ(W l 2H l t ) + (1 - a l t )(1 - σ(W l 2H l t )) (10)

[0153] where Θl Based on the model parameters of the second reinforcement learning model, W l 2 is the activation parameter updated based on the second reward at the k-th time, σ(·) is the sigmode function, and the action parameter a l t takes a value of 0 or 1. When the action parameter a l t takes a value of 0, π(s l t , a l t ) determines the probability that the t-th historical browsing item is retained. When the action parameter a l t takes a value of 1, π(s l t , a l t ) determines the probability that the t-th historical browsing item is filtered.

[0154] Thus, when the action parameter a l t takes a value of 1, it is possible to determine whether the t-th historical browsing item in the current item list is filtered according to the output probability value of formula (10). Further, the item list obtained by filtering the items to be filtered is used as the updated "current item list", and it is judged whether the "current item list" needs to be filtered in the manner shown in steps S610 - S640 as follows Figure 6 .

[0155] When the "current item list" does not contain the items to be filtered, the "current item list" can be used for item recommendation. Then, taking videos as an example, the specific implementation methods for video recommendation based on the filtered optimized item list can include the following several schemes:

[0156] Scheme A: Recommendation based on similarity

[0157] Recommendation techniques based on similarity are mainly divided into two categories: one is the recommendation method based on user similarity. For a given user, the recommendation result is given by recommending the videos browsed by users similar to this user. The other is the recommendation method based on videos. For a given user, the recommendation result is given by recommending the videos most similar to the videos browsed by this user.

[0158] Scheme B: Recommendation based on matching degree

[0159] Recommendation techniques based on matching degree mainly calculate the matching degree between a given user and the candidate set of videos, then sort the candidate set of videos according to the matching degree, and then recommend the top several videos with the highest matching degree.

[0160] Scheme C: Recommendation based on node embedding

[0161] Recommendation techniques based on node embedding obtain user and video features through methods such as collaborative filtering and graph embedding, and then use the dot product of feature vectors to predict the relationship between users and videos. And recommend several videos with the largest dot product vectors.

[0162] It should be noted that the solutions for project recommendation are not limited to the above several, and other project recommendation methods are also possible. This application does not make any limitations in this regard.

[0163] This technical solution is based on judging the relevance between the target video and the historical browsing items in the project list at two levels, and can quickly and effectively distinguish whether the current project list needs to be filtered and determine the items to be filtered. At the same time, this technical solution performs project recommendation based on the optimized project list after filtering, that is to say, this technical solution enhances the prediction model for determining the recommendation value based on the above first reinforcement learning model and the second reinforcement learning model, which is beneficial to improving the accuracy of project recommendation.

[0164] Those skilled in the art can understand that all or part of the steps for implementing the above embodiments are implemented as a computer program executed by a processor (including GPU / CPU). When this computer program is executed by the GPU / CPU, it executes the above functions defined by the above method provided by the present disclosure. The said program can be stored in a computer-readable storage medium, and this storage medium can be a read-only memory, a magnetic disk or an optical disc, etc.

[0165] In addition, it should be noted that the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, rather than for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the time sequence of these processes. Additionally, it is also easy to understand that these processes can be executed, for example, synchronously or asynchronously in multiple modules.

[0166] The following Figure 9 introduces an embodiment of the project recommendation device based on reinforcement learning of the present disclosure, which can be used to execute the above-mentioned project recommendation method based on reinforcement learning of the present disclosure.

[0167] Figure 9 shows a schematic structural diagram of the project recommendation device based on reinforcement learning in an exemplary embodiment of the present disclosure. As Figure 9 shown, the above-mentioned project recommendation device 900 based on reinforcement learning includes: a prediction module 901, an acquisition module 902, a judgment module 903, and a filtering processing module 904.

[0168] Among them, the above prediction module 901 is configured to: predict the recommendation value for the above target project according to the vector representation of the target project and the historical browsing projects in the current project list, where the above current project list consists of projects including user browsing behaviors; the above acquisition module 902 is configured to: acquire the relative features between the above target project and the above historical browsing projects; the above judgment module 903 is configured to: based on the first reinforcement learning model, determine whether to perform filtering processing on the above current project list according to the above relative features and the above recommendation value; and, the above filtering processing module 904 is configured to: in response to the need to perform filtering processing on the above current project list, based on the second reinforcement learning model, determine the projects to be filtered in the above current project list, so as to perform project recommendation based on the optimized project list after filtering.

[0169] In some embodiments of the present disclosure, based on the foregoing solution, the above acquisition module 902 is specifically configured to: respectively acquire the individual relative features between the above target project and the above multiple historical browsing projects; and, determine the overall relative features between the above target project and the above current project list according to the above individual relative features.

[0170] In some embodiments of the present disclosure, based on the foregoing solution, the above filtering processing module 904 is specifically configured to: input the above overall relative features and the above recommendation value into the above first reinforcement learning model to determine the overall relevance between the above target project and the above current project list based on the above first reinforcement learning model; and, determine whether to perform filtering processing on the above current project list based on the above overall relevance.

[0171] In some embodiments of the present disclosure, based on the foregoing solution, the above project recommendation device 900 based on reinforcement learning further includes: a first training module.

[0172] Among them, the above first training module is configured to: update the model parameters of the above first reinforcement learning model according to the first reward of the jth time, determine whether to perform the jth filtering processing according to the first reinforcement learning model after updating the model parameters, where j is an integer greater than 1; in response to the need to perform the jth filtering processing, determine the jth recommendation value for the above target project according to the jth optimized project list obtained after the jth filtering processing; determine the (j + 1)th first reward for the above first reinforcement learning model according to the above jth recommendation value and the (j - 1)th recommendation value, so as to update the model parameters of the above first reinforcement learning model according to the (j + 1)th first reward.

[0173] In some embodiments of the present disclosure, based on the foregoing solution, the filtering processing module 904 is further specifically configured to: obtain the j-th overall relative feature between the target item and the historical browsing items in the j-th optimized item list; determine the j-th state vector according to the j-th recommendation value and the j-th overall relative feature; perform a rectified linear unit (ReLU) processing on the j-th state vector based on the first reinforcement learning model after updating the model parameters with the first reward for the j-th time, to obtain the j-th rectified vector; and, based on the action parameters of the first reinforcement learning model, process the j-th rectified vector to obtain the overall correlation degree between the target item and the current item list.

[0174] In some embodiments of the present disclosure, based on the foregoing solution, the reinforcement learning-based item recommendation device 900 further includes: an item recommendation module.

[0175] Wherein, the item recommendation module is configured to: in response to not needing to perform the j-th filtering process, perform item recommendation based on the (j - 1)-th optimized item list filtered by the (j - 1)-th filtering process.

[0176] In some embodiments of the present disclosure, based on the foregoing solution, the obtaining module 902 is specifically configured to: respectively obtain the individual relative features between the target item and the multiple historical browsing items.

[0177] In some embodiments of the present disclosure, based on the foregoing solution, the filtering processing module 904 is specifically configured to: input the t-th individual relative feature corresponding to the t-th historical browsing item into the second reinforcement learning model, to determine the individual correlation degree between the target item and the t-th historical browsing item based on the model parameters of the second reinforcement learning model, where t is an integer not greater than the number of historical browsing items in the t-th optimized item list; and, determine whether the t-th historical browsing item is an item to be filtered based on the individual correlation degree.

[0178] In some embodiments of the present disclosure, based on the foregoing solution, the reinforcement learning-based item recommendation device 900 further includes: a second training module.

[0179] Among them, the second training module is configured to: update the model parameters of the second reinforcement learning model according to the second reward at the k-th time, and determine the items to be filtered according to the second reinforcement learning model after updating the model parameters, where k is an integer greater than 1; determine the k-th recommended value for the target item according to the k-th optimized item list obtained after the k-th filtering process; determine the first part of the reward according to the k-th recommended value and the (k - 1)-th recommended value, where the (k - 1)-th recommended value is determined according to the (k - 1)-th optimized item list obtained after the (k - 1)-th filtering process; calculate the similarity between each historical browsing item in the k-th optimized item list and the target item respectively, determine the second part of the reward according to the multiple similarities; and determine the second reward for the (k + 1)-th time for the second reinforcement learning model according to the first part of the reward and the second part of the reward, so as to update the model parameters of the second reinforcement learning model according to the second reward for the (k + 1)-th time.

[0180] In some embodiments of the present disclosure, based on the foregoing solution, the filtering processing module 904 is further specifically configured to: obtain the t-th individual relative feature between the target item and the t-th historical browsing item; determine the t-th state vector according to the t-th individual relative feature; perform a rectified linear unit processing on the t-th state vector based on the second reinforcement learning model after updating the model parameters according to the second reward at the k-th time, to obtain the t-th rectified vector; and process the t-th rectified vector based on the action parameters of the second reinforcement learning model, to obtain the individual correlation degree between the target item and the t-th historical browsing item.

[0181] In some embodiments of the present disclosure, based on the foregoing solution, the prediction module 901 includes a user vector representation unit and a recommended value determination unit.

[0182] Among them, the user vector representation unit is configured to: determine the vector representation of the user according to the historical browsing items in the current item list; the recommended value determination unit is configured to: predict the recommended value for the target item based on the vector representation of the user and the vector representation of the target item.

[0183] In some embodiments of the present disclosure, based on the foregoing solution, the user vector representation unit is specifically configured to: average the vector representations of the historical browsing items in the j-th optimized item list obtained after the j-th filtering process, to obtain the j-th vector representation for the user, so as to predict the j-th recommended value for the target item.

[0184] In some embodiments of the present disclosure, based on the foregoing solution, the obtaining module 902 is configured to: obtain one or several of the following information between the target item and the historical browsing item: cosine distance, dot product of vectors, and similarity between feature labels of items, to obtain the relative feature.

[0185] The specific details of each unit in the above-described project recommendation device based on reinforcement learning have been described in detail in the project recommendation method based on reinforcement learning, and thus will not be elaborated herein.

[0186] Figure 10 FIG. shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present disclosure.

[0187] It should be noted that Figure 10 the shown computer system 1000 of the electronic device is only an example, and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0188] As Figure 10 shown, the computer system 1000 includes a processor 1001, where the processor 1001 may specifically include: a graphics processing unit (GPU) and a central processing unit (CPU), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage section 1008 into a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for system operations are also stored. The processor 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0189] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, etc.; an output section 1007 including such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and speakers; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a local area network (LAN) card, a modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as required. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1010 as required, so that a computer program read from it can be installed into the storage section 1008 as required.

[0190] In particular, according to an embodiment of the present disclosure, the processes described below with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 1009, and / or installed from the removable medium 1011. When the computer program is executed by the processor 1001, various functions defined in the system of the present application are performed.

[0191] It should be noted that the computer-readable medium shown in the embodiments of the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0192] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0193] The units described in the embodiments of the present disclosure can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not constitute a limitation on the units themselves in some cases.

[0194] As another aspect, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the one or more programs are executed by an electronic device, the electronic device is caused to implement the methods described in the above embodiments.

[0195] For example, the electronic device may implement as Figure 4 shown in: Step S410, predicting a recommendation value for the target item based on the vector representation of the target item and the historical browsing items in the current item list, where the current item list consists of items including user browsing behaviors; Step S420, obtaining the relative features between the target item and the historical browsing items; Step S430, based on a first reinforcement learning model, determining whether to perform a filtering process on the current item list according to the relative features and the recommendation value; and Step S440, in response to the need to perform a filtering process on the current item list, based on a second reinforcement learning model, determining the items to be filtered in the current item list according to the relative features, so as to perform item recommendation based on the optimized item list after filtering.

[0196] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0197] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by software in combination with necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0198] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed herein. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0199] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope.

Claims

1. A project recommendation method based on reinforcement learning, characterized in that, The method includes: Predicting a recommendation value for the target item based on the vector representation of the target item and historical browsing items in the current item list, where the current item list consists of items containing user browsing behaviors; Obtaining the relative features between the target item and the historical browsing items; Based on a first reinforcement learning model, determining whether to perform filtering processing on the current item list according to the relative features and the recommendation value; In response to the need to perform filtering processing on the current item list, based on a second reinforcement learning model, determining the items to be filtered in the current item list according to the relative features, so as to perform item recommendation based on the optimized item list after filtering; Wherein, obtaining the relative features between the target item and the historical browsing items includes: respectively obtaining the individual relative features between the target item and multiple historical browsing items; Determining the overall relative features between the target item and the current item list according to the individual relative features; Based on the first reinforcement learning model, determining whether to perform filtering processing on the current item list according to the relative features and the recommendation value includes: inputting the overall relative features and the recommendation value into the first reinforcement learning model to determine the overall relevance between the target item and the current item list based on the first reinforcement learning model; determining whether to perform filtering processing on the current item list based on the overall relevance.

2. The method according to claim 1, characterized in that, The method further includes: Updating the model parameters of the first reinforcement learning model according to the first reward of the j-th time, and determining whether to perform the j-th filtering processing according to the first reinforcement learning model after updating the model parameters, where j is an integer greater than 1; In response to the need to perform the j-th filtering processing, determining the j-th recommendation value for the target item according to the j-th optimized item list obtained after the j-th filtering processing; Determining the (j + 1)-th first reward for the first reinforcement learning model according to the j-th recommendation value and the (j - 1)-th recommendation value, so as to update the model parameters of the first reinforcement learning model according to the (j + 1)-th first reward, and the (j - 1)-th recommendation value is determined according to the (j - 1)-th optimized item list obtained after the (j - 1)-th filtering processing.

3. The method according to claim 2, characterized in that, Inputting the overall relative features and the recommendation value into the first reinforcement learning model to determine the overall relevance between the target item and the current item list based on the first reinforcement learning model includes: Obtaining the j-th overall relative features between the target item and the historical browsing items in the j-th optimized item list; Determining the j-th state vector according to the j-th recommendation value and the j-th overall relative features; Performing a rectified linear unit processing on the j-th state vector based on the first reinforcement learning model after updating the model parameters according to the first reward of the j-th time, to obtain the j-th rectified vector; Processing the j-th rectified vector based on the action parameters of the first reinforcement learning model to obtain the overall relevance between the target item and the current item list.

4. The method according to claim 2, characterized in that, The method further includes: In response to the need not to perform the j-th filtering process, project recommendation is performed based on the (j - 1)-th optimized project list filtered by the (j - 1)-th filtering process.

5. The method according to claim 1, characterized in that, Based on the second reinforcement learning model, determining the item to be filtered in the current project list according to the relative feature includes: Inputting the t-th individual relative feature corresponding to the t-th historical browsing item into the second reinforcement learning model to determine the individual association degree between the target item and the t-th historical browsing item based on the model parameters of the second reinforcement learning model, where t is an integer not greater than the number of historical browsing items in the t-th optimized project list; Determining whether the t-th historical browsing item is an item to be filtered based on the individual association degree.

6. The method according to claim 5, characterized in that, The method further includes: Updating the model parameters of the second reinforcement learning model according to the second reward of the k-th time, and determining the item to be filtered according to the second reinforcement learning model with the updated model parameters, where k is an integer greater than 1; Determining the k-th recommendation value for the target item according to the k-th optimized project list obtained after the k-th filtering process; Determining the first part of the reward according to the k-th recommendation value and the (k - 1)-th recommendation value, where the (k - 1)-th recommendation value is determined according to the (k - 1)-th optimized project list obtained after the (k - 1)-th filtering process; Calculating the similarity between each historical browsing item in the k-th optimized project list and the target item respectively, and determining the second part of the reward according to the multiple similarities; Determining the second reward of the (k + 1)-th time for the second reinforcement learning model according to the first part of the reward and the second part of the reward, so as to update the model parameters of the second reinforcement learning model according to the second reward of the (k + 1)-th time.

7. The method according to claim 6, characterized in that, The step of inputting the t-th individual relative feature corresponding to the t-th historical browsing item into the second reinforcement learning model to determine the individual association degree between the target item and the t-th historical browsing item based on the model parameters of the second reinforcement learning model includes: Obtaining the t-th individual relative feature between the target item and the t-th historical browsing item; Determining the t-th state vector according to the t-th individual relative feature; Performing a rectified linear unit process on the t-th state vector based on the second reinforcement learning model with the model parameters updated according to the second reward of the k-th time to obtain the t-th rectified vector; Processing the t-th rectified vector based on the action parameters of the second reinforcement learning model to obtain the individual association degree between the target item and the t-th historical browsing item.

8. The method according to any one of claims 2 to 7, characterized in that The step of predicting the recommendation value for the target item according to the vector representation of the target item and the historical browsing items in the current project list includes: Determining the vector representation of the user according to the historical browsing items in the current project list; Predicting the recommendation value for the target item based on the vector representation of the user and the vector representation of the target item.

9. The method according to claim 8, characterized in that The step of determining the vector representation of the user according to the historical browsing items in the current project list includes: Taking the average of the vector representations of the historical browsing items in the j-th optimized project list obtained after the j-th filtering process to obtain the j-th vector representation for the user, so as to predict the j-th recommendation value for the target item.

10. The method according to any one of claims 1 to 7, characterized in that Obtaining the relative features between the target item and the historical browsing items includes: Obtaining one or several of the following information between the target item and the historical browsing items: cosine distance, dot product of vectors, and similarity between feature labels of items, to obtain the relative features.

11. A project recommendation system, characterized in that The system includes: A recommended value prediction model configured to: predict the recommended value for the target item according to the vector representation of the target item and the vector representation of the user, where the vector representation of the user is determined according to the current item list, and the current item list consists of items including user browsing behaviors; A first reinforcement learning model configured to: respectively obtain the individual relative features between the target item and multiple historical browsing items, and determine the overall relative features between the target item and the current item list according to the individual relative features; based on the individual relative features and the overall relative features, determine the overall relevance between the target item and the current item list, and determine whether to perform filtering processing on the current item list based on the overall relevance; A second reinforcement learning model configured to: in response to the need to perform filtering processing on the current item list, determine the items to be filtered in the current item list according to the relative features, so that the recommended value prediction model performs item recommendation based on the filtered optimized item list.

12. A project recommendation device based on reinforcement learning, characterized in that The device includes: A prediction module configured to: predict the recommended value for the target item according to the vector representation of the target item and the historical browsing items in the current item list, where the current item list consists of items including user browsing behaviors; An acquisition module configured to: respectively obtain the individual relative features between the target item and multiple historical browsing items, and determine the overall relative features between the target item and the current item list according to the individual relative features; A judgment module configured to: input the overall relative features and the recommended value into the first reinforcement learning model to determine the overall relevance between the target item and the current item list based on the first reinforcement learning model, and determine whether to perform filtering processing on the current item list based on the overall relevance; A filtering processing module configured to: in response to the need to perform filtering processing on the current item list, based on the second reinforcement learning model, determine the items to be filtered in the current item list according to the relative features, to perform item recommendation based on the filtered optimized item list.

13. A computer-readable storage medium, characterized in that A computer program is stored thereon; When the computer program is executed by a processor, it implements the reinforcement learning-based item recommendation method according to any one of claims 1 to 10.

14. An electronic device, characterized in that The electronic device includes: One or more processors; A storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the reinforcement learning-based item recommendation method according to any one of claims 1 to 10.

15. A computer program product, characterized in that The computer program product includes computer instructions that are stored in a computer-readable storage medium; A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the reinforcement learning-based project recommendation method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Personalized recommender with limited data availability

    CN112182360A