Resource service access recommendation method and device
By building a multi-scenario model, using the resource recommendation model and the policy network obtained by reinforcement learning, resource recommendation decisions are made based on the user's access path parameters, and the problem of low user conversion rate in online resource services is solved, and higher user conversion rate and purchase probability are achieved.
Patent Information
- Application Number
- CN202510300072.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-24
AI Technical Summary
In online resource services, how to improve user conversion rates and cope with fierce competition among institutions.
By building a multi-scenario model, the policy network obtained by using the resource recommendation model and the agent to perform reinforcement learning based on the state space and action space, the decision module is determined based on the user's access path parameters, and resource recommendation decisions are made.
It improves the conversion rate of users from access to resource purchase in resource services, and improves the user experience and purchase probability through resource recommendations for different scenarios.
Smart Images

Figure CN120196812A_ABST
Abstract
Description
Technical Field
[0001] This document relates to the technical field of data processing, and in particular, to a method and device for accessing and recommending resource services. Background Art
[0002] With the continuous development and popularization of the Internet, the scope of application of various online services provided based on the Internet is also becoming wider and wider. In this case, a method of purchasing resources online has emerged and has gradually been accepted by users. However, as the number of institutions providing online resource services increases, the competition among institutions is also intensifying. In this case, how to improve the user conversion of online resource services has become the current focus of attention. Summary of the Invention
[0003] One or more embodiments of this specification provide a method for accessing and recommending resource services, including: determining a decision-making module for resource recommendation in a multi-scenario model according to access path parameters of a resource page accessed by a user in a resource service. The decision-making module includes a resource recommendation model and a policy network obtained by an intelligent agent through reinforcement learning based on a state space and an action space. If the decision-making module is the policy network, the intelligent agent obtains the access state data of the user, and inputs the access state data into the policy network to make a resource recommendation decision to obtain a recommended action value. An action is selected in the action space according to the recommended action value, and the selected resource recommendation action is executed.
[0004] One or more embodiments of this specification provide a model training method, including: constructing an initialization network for determining a mapping relationship between a state space and an action space. The intelligent agent selects an action in the action space and executes the selected resource recommendation action. A reward is determined according to the action execution feedback of the selected resource recommendation action, and the parameters of a preset policy network are adjusted based on the reward to obtain a multi-scenario model including a policy network and a resource recommendation model after the reinforcement learning is completed. The resource recommendation model is obtained by training a to-be-trained model constructed based on the state space and the action space.
[0005] One or more embodiments of this specification provide an access recommendation device for resource services, including: A decision-making module determination module configured to determine a decision-making module for resource recommendation in a multi-scenario model according to access path parameters of a resource page accessed by a user in a resource service. The decision-making module includes a resource recommendation model and a policy network obtained by an agent through reinforcement learning based on a state space and an action space. A recommendation decision-making module configured to, if the decision-making module is the policy network, obtain the access state data of the user through the agent, and input the access state data into the policy network to make a resource recommendation decision to obtain a recommended action value. An action selection execution module configured to perform action selection in the action space according to the recommended action value, and execute the selected resource recommendation action.
[0006] One or more embodiments of this specification provide a model training device, including: A network construction module configured to construct an initialization network for determining a mapping relationship between a state space and an action space. An action selection execution module configured to perform action selection in the action space through an agent, and execute the selected resource recommendation action. A parameter adjustment module configured to determine a reward according to an action execution feedback of the selected resource recommendation action, and perform parameter adjustment on a preset policy network based on the reward, so as to obtain a multi-scenario model including a policy network and a resource recommendation model after the reinforcement learning is completed. Wherein, the resource recommendation model is obtained by training a to-be-trained model constructed based on the state space and the action space.
[0007] One or more embodiments of this specification provide an access recommendation device for resource services, including: A processor; and a memory configured to store computer-executable instructions, the computer-executable instructions, when executed, cause the processor to: Determine a decision-making module for resource recommendation in a multi-scenario model according to access path parameters of a resource page accessed by a user in a resource service. The decision-making module includes a resource recommendation model and a policy network obtained by an agent through reinforcement learning based on a state space and an action space. If the decision-making module is the policy network, obtain the access state data of the user through the agent, and input the access state data into the policy network to make a resource recommendation decision to obtain a recommended action value. Perform action selection in the action space according to the recommended action value, and execute the selected resource recommendation action.
[0008] One or more embodiments of the present specification provide a model training device, including: a processor; and a memory configured to store computer-executable instructions, which, when executed, cause the processor to: construct an initialization network for determining a mapping relationship between a state space and an action space. Select an action in the action space through an agent, and execute the selected resource recommendation action. Determine a reward based on the action execution feedback of the selected resource recommendation action, and adjust the parameters of a preset policy network based on the reward to obtain a multi-scenario model including a policy network and a resource recommendation model after the reinforcement learning is completed. Wherein, the resource recommendation model is obtained by training a to-be-trained model constructed based on the state space and the action space.
[0009] One or more embodiments of the present specification provide a computer-readable storage medium for storing computer-executable instructions, which, when executed, implement the following process: Determine a decision-making module for resource recommendation in a multi-scenario model according to access path parameters of a resource page accessed by a user. The decision-making module includes a resource recommendation model and a policy network obtained by an agent through reinforcement learning based on a state space and an action space. If the decision-making module is the policy network, obtain the access state data of the user through the agent, and input the access state data into the policy network to make a resource recommendation decision to obtain a recommended action value. Select an action in the action space according to the recommended action value, and execute the selected resource recommendation action.
[0010] One or more embodiments of the present specification provide another computer-readable storage medium for storing computer-executable instructions, which, when executed, implement the following process: construct an initialization network for determining a mapping relationship between a state space and an action space. Select an action in the action space through an agent, and execute the selected resource recommendation action. Determine a reward based on the action execution feedback of the selected resource recommendation action, and adjust the parameters of a preset policy network based on the reward to obtain a multi-scenario model including a policy network and a resource recommendation model after the reinforcement learning is completed. Wherein, the resource recommendation model is obtained by training a to-be-trained model constructed based on the state space and the action space. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in one or more embodiments of the present specification or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. Figure 1 Schematic diagram of an implementation environment of an access recommendation method for a resource service provided for one or more embodiments of this specification; Figure 2 Processing flowchart of an access recommendation method for a resource service provided for one or more embodiments of this specification; Figure 3 Processing flowchart of an access recommendation method for a resource service applied to a resource service scenario provided for one or more embodiments of this specification; Figure 4 Schematic diagram of an implementation environment of a model training method provided for one or more embodiments of this specification; Figure 5 Schematic diagram of an embodiment of an access recommendation device for a resource service provided for one or more embodiments of this specification; Figure 6 Schematic diagram of an embodiment of a model training device provided for one or more embodiments of this specification; Figure 7 Schematic diagram of the structure of an access recommendation device for a resource service provided for one or more embodiments of this specification; Figure 8 Schematic diagram of the structure of a model training device provided for one or more embodiments of this specification. Detailed implementation manners
[0012] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the following will clearly and completely describe the technical solutions in one or more embodiments of this specification with reference to the accompanying drawings in one or more embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this document.
[0013] The access recommendation method for a resource service provided for one or more embodiments of this specification is applicable to the implementation environment of a resource service background. Referring to Figure 1 , this implementation environment at least includes: The multi-scenario model 101 refers to a model that can adapt to and process different scenarios, conditions, or tasks during the process of a user accessing a resource service, or a model composed of multiple sub-models and / or networks that can work together; among them, the multi-scenario model 101 includes a resource recommendation model 101-1 and a policy network 101-2. Correspondingly, this implementation environment further includes an agent 102 for obtaining the policy network 101-2 through a reinforcement learning method; Specifically, the resource recommendation model 101-1 is used for the resource recommendation action in a scenario of resource service; the policy network 101-2 is used to determine the resource recommendation action for resource recommendation in another scenario of resource service.
[0014] In this implementation environment, during the process of resource recommendation to a user when the user accesses the resource service through the multi-scenario model, on the premise that the multi-scenario model includes two decision-making modules, namely the resource recommendation model and the policy network obtained by the agent through reinforcement learning based on the state space and the action space, according to the access path parameters of the resource page accessed by the user in the resource service, determine the decision-making module for resource recommendation to the user. If the determined decision-making module is the policy network, the agent is used to obtain the access state data of the user, and the access state data is input into the policy network for resource recommendation decision-making to obtain the recommended action value. An action is selected from the action space according to the recommended action value, and the selected resource recommendation action is executed, so as to perform targeted resource recommendation to the user during the process of the user accessing the resource service with the help of the multi-scenario model.
[0015] It should be noted that considering that relevant data such as the access state data, page access data, and access behavior data of the users involved in this specification may be the privacy of the users to a certain extent. Therefore, if you want to collect relevant data such as the access state data, page access data, and access behavior data of the users, you can obtain the authorization of the users before collecting the data, so that the operation of collecting data complies with relevant data management regulations. For example, data authorization can be carried out during the current access process of the user to the resource service, or data authorization can also be carried out during the first access process of the user to the resource service; the specific method of data authorization can be to send a user data authorization reminder to the user, and the user can obtain data collection authorization after confirming the reminder through an instruction. Or, the method of data authorization can also be to obtain data collection authorization by signing a data authorization agreement.
[0016] One or more embodiments of an access recommendation method for a resource service provided in this specification are as follows: Refer to Figure 2 , the access recommendation method for the resource service provided in this embodiment, the method specifically includes steps S202 to S206.
[0017] Step S202, determine the decision-making module for resource recommendation in the multi-scenario model according to the access path parameters of the resource page accessed by the user in the resource service.
[0018] The resource service described in this embodiment refers to the service related to resources provided within an application, such as a resource service for performing related processing such as recommendation and transaction of resource objects. The resource object can be an entitlement resource, such as tradable asset varieties (funds, stocks, bonds), etc., or can also be a capital resource or a virtual resource. The virtual resource can be a virtual resource such as carbon reduction amount, points, etc.
[0019] The resource page refers to a page for displaying detailed information, transaction, or other data information of a resource object, or a page for performing corresponding processing on a resource object; optionally, the resource page includes a resource list page, an object detail page, a resource purchase page, a buy resource page, and / or a buy resource detail page; Among them, the resource list page is used to display a list of resource objects composed of multiple resource objects. Users can independently select a resource object on the resource list page for viewing and access. Specifically, they can view and access by clicking on the area where a certain resource object is located on the resource list page, or by clicking on the resource recommendation data of a certain resource object on the resource list page. For example, when a user clicks on the first object data or the first area of a certain resource object on the resource list page, they can jump to the object detail page, or when they click on the second object data or the second area of a certain resource object on the resource list page, they can jump to the buy resource detail page; The object detail page is an object detail page for displaying the detailed information of a resource object. Users can purchase the resource object through the object detail page. For example, users can click on the purchase control configured on the object detail page to jump to the resource purchase page of the resource object to purchase the resource object; The resource purchase page is used for purchasing a resource object. Users can enter the amount of the resource object they want to purchase on the resource purchase page, and click the confirmation button to make a purchase payment for the resource object after entering the resource amount; The buy resource page is used to display all the resource objects purchased by the user. Users can select a certain purchased resource object on the buy resource page for viewing. Specifically, they can jump to the buy resource detail page to view the detailed resource data of the purchased resource object; The buy resource detail page is used to display the details of a certain resource object purchased by the user, such as a buy resource detail page for displaying the purchase time, resource amount, resource income, etc. of a certain resource object purchased by the user. Users can jump to the resource purchase page of the purchased resource object through the buy resource detail page to continue purchasing the resource object.
[0020] In practical applications, during the process of a user accessing a resource service, from the start of accessing a resource object in the resource service to the completion of purchasing the resource object, this conversion process can be achieved by accessing different resource pages. In this embodiment, the path composed of multiple resource pages accessed by the user during the process from accessing a resource object to completing the purchase of the resource object is called an access path. For example, when the resource pages include five resource pages: resource list page A, buy resource page B, object details page C, buy resource details page D, and resource purchase page E, the access paths for the user from accessing a resource object to completing the purchase of the resource object include the following five: Access path a: Resource list page A -> Object details page C -> Resource purchase page E; Access path b: Resource list page A -> Buy resource details page D -> Resource purchase page E; Access path c: Buy resource page B -> Buy resource details page D -> Resource purchase page E; Access path d: Object details page C -> Resource purchase page E; Access path e: Buy resource details page D -> Resource purchase page E.
[0021] In this embodiment, a multi-scenario model refers to a model that can adapt to and handle different scenarios, conditions, or tasks during the process of a user accessing a resource service, or a model composed of multiple sub-models and / or networks that can work together. Specifically, by setting multiple decision modules in the multi-scenario model to perform resource recommendations for different scenarios of the user accessing the resource service. Among them, different scenarios of the user accessing the resource service can be different resource pages of the user accessing the resource service, that is: corresponding resource recommendations are made for different resource pages of the user accessing the resource service through multiple decision modules in the multi-scenario model, so as to increase the probability of the user's conversion from access to resource purchase in the resource service.
[0022] Optionally, the decision module includes a resource recommendation model and a policy network obtained by an agent through reinforcement learning based on a state space and an action space. During the specific execution process, when the decision module of the multi-scenario model includes a resource recommendation model and a policy network, the resource recommendation model and the policy network for performing resource recommendations for different scenarios can be obtained through multi-task learning. The following respectively specifically describes the generation processes of the resource recommendation model and the policy network.
[0023] (1) Policy network The role of the policy network is to determine the actions that the agent should take in a given state, and to help the agent learn policies through interactions with the environment. Among them, the state can be provided by the state space, the actions can be provided by the action space, and policy learning can be carried out in the way of reinforcement learning. Optionally, the state space includes the access state data of the user; the action space includes resource recommendation actions for recommending data for at least one resource in each resource page of the preset access path.
[0024] Optionally, the access state data includes: the page access data of the previous resource page of the current resource page in the preset access path, and / or, the access behavior data.
[0025] Among them, the page access data refers to the data related to the user's behavior of accessing the resource page generated during the user's access to the resource page of the resource service. For example, the page access data includes at least one of the following: the object data of the resource object in the resource page accessed by the user, the resource pages that the user has accessed, the operation action data of the user in the accessed resource page, and the stay duration of the user in the accessed resource page.
[0026] The access behavior data can be the historical access records of the user's previous access to the resource service, or the program access records of the user's access to the host program of the resource service, or the access records composed of both the historical access records and the program access records. In addition, the access behavior data can also be composed of at least one of the historical access records and the program access records and the user data. The user data can be identity data, risk data, credit data, and / or salary data.
[0027] The resource recommendation action refers to the recommendation action of recommending resources to the user by displaying recommended data on the resource page during the user's access to the resource page of the resource service. Specifically, the resource recommendation action can be to display instantiated controls, copywriting, or other forms of recommended materials on the resource page. For example, display management suggestions for resource management on the object details page, or display purchase suggestions for the current resource object; or, display historical income reminders for the current resource object and the purchase trigger control for the resource object on the buy resource details page; The resource recommendation action can also be to display corresponding instantiated controls, copywriting, or other forms of recommended materials for one or more resource objects displayed on the resource page. For example, display value change reminders for the current resource object in the area below a certain resource object on the resource list page; or, display historical income reminders for the current resource object and the purchase trigger control for the resource object in the area below a certain resource object on the buy resource page; During the specific execution process, based on the state space and the action space, a policy network is generated by means of reinforcement learning. In an optional implementation provided in this embodiment, the policy network is obtained in the following manner: Construct an initialization network for determining the mapping relationship between the state space and the action space; The agent selects an action in the action space and executes the selected resource recommendation action; Determine the reward according to the action execution feedback of the selected resource recommendation action, and adjust the parameters of the initialization network based on the reward to obtain the policy network after the reinforcement learning is completed.
[0028] For example, in the process of training and generating a policy network using deep reinforcement learning, the goal is to maximize the probability that the user completes the conversion according to access path a, access path b, access path c, access path d, or access path e. First, determine the state space S = (A_state, B_state, C_state, D_state, E_state) for deep reinforcement learning, where A_state, B_state, C_state, D_state, and E_state respectively represent the access state data of the five resource pages: resource list page A, buy resource page B, object details page C, buy resource details page D, and resource purchase page E. The access state data A_state includes: the object data of the resource object in the resource page accessed by the user, the resource pages that the user has accessed, the operation action data of the user on the accessed resource page, and / or the residence duration of the user on the accessed resource page; the access state data B_state, access state data C_state, access state data D_state, and access state data E_state are similar to the access state data A_state, and also include the object data of the resource object in the resource page accessed by the user, the resource pages that the user has accessed, the operation action data of the user on the accessed resource page, and / or the residence duration of the user on the accessed resource page; And determine the action space A = (A_action, B_action, C_action, D_action, E_action), where A_action, B_action, C_action, D_action, and E_action respectively represent the resource recommendation actions executed on the resource list page A, buy resource page B, object details page C, buy resource details page D, and resource purchase page E. The resource recommendation action A_action ∈ {resource recommendation data 1, resource recommendation data 2,...}; On this basis, an initialization network is constructed according to the mapping relationship between the state space and the action space, and the initial parameters of the initialization network are set. During the deep reinforcement learning process, a resource recommendation action is selected in the action space according to the current state, and the reward is determined according to the action execution feedback of the selected resource recommendation action. Then, the parameters of the initialization network are adjusted based on the reward. The above deep reinforcement learning process is repeated until the policy network is obtained after meeting the preset convergence conditions. The convergence conditions can be reaching the preset number of iterations, the change in the cumulative reward being less than the preset threshold, etc.
[0029] In addition to the above-mentioned method of generating a policy network through deep reinforcement learning, the Q-Learning method can also be used for reinforcement learning to obtain a policy network. Specifically, based on the state space and the action space, a Q-table is constructed and initialized according to the mapping relationship between the state space and the action space. A resource recommendation action is selected in the action space according to the current state, and the reward is determined according to the action execution feedback of the selected resource recommendation action. Then, the Q-table is adjusted based on the reward. The above reinforcement learning process is repeated until the policy network is obtained after meeting the preset convergence conditions. The convergence conditions can be reaching the preset number of iterations, the change in the cumulative reward being less than the preset threshold, etc.
[0030] (2)Resource recommendation model In an optional implementation provided in this embodiment, the resource recommendation model is obtained by training in the following manner: Input the data sample into the model to be trained for resource recommendation decision-making to obtain resource recommendation data; Calculate the training loss based on the resource recommendation data and the data label, and adjust the parameters of the model to be trained, so as to obtain the resource recommendation model after the training is completed.
[0031] Among them, the model to be trained can be constructed based on the state space and the action space, and specifically can be obtained by modeling according to the mapping relationship between the state space and the action space. In addition to the above-mentioned implementation method of training the model to be trained in a supervised manner to obtain the resource recommendation model, the model to be trained can also be trained in an unsupervised manner to obtain the resource recommendation model.
[0032] Specifically, when the multi-scenario model includes two decision-making modules, namely the resource recommendation model and the policy network, during the process of resource recommendation when the user accesses the resource service, the decision-making module for resource recommendation to the user is determined according to the resource page accessed by the user in the resource service. In this way, different decision-making modules in the multi-scenario model can perform targeted resource recommendation for different resource pages accessed by the user in the resource service, which helps to increase the probability of the user's conversion from resource service access to resource purchase.
[0033] Specifically, in the process of determining the decision-making module for resource recommendation to the user based on the resource page accessed by the user in the resource service, the decision-making module for resource recommendation to the user can be determined according to the access path parameters of the resource page accessed by the user in the resource service. In an optional implementation manner provided in this embodiment, according to the access path parameters of the resource page accessed by the user in the resource service, the decision-making module for resource recommendation in the multi-scenario model is determined, including: If the access path parameter is the head node type, the resource recommendation model is determined as the decision-making module; If the access path parameter is the intermediate node type, the policy network is determined as the decision-making module.
[0034] Among them, the access path parameter refers to the node position or node order of the resource page currently accessed by the user in the access path. The head node type and the intermediate node type can be determined according to the node where the resource page is located in the access path; specifically, the head node type means that the resource page currently accessed by the user is the first resource page (the first resource page) in the preset access path, that is, the head path node (the first path node); the intermediate node type means that the resource page currently accessed by the user is the resource page after the first resource page in the preset access path, that is: the intermediate path node.
[0035] In practical applications, in order to improve the response speed of resource recommendation to the user during the user's access to the resource service, a processing module for determining the decision-making module for resource recommendation in the multi-scenario model can also be configured in the multi-scenario model. After the processing module determines the decision-making module, the corresponding decision-making module can also be called by the processing module to perform the corresponding decision-making process for resource recommendation. Optionally, the multi-scenario model further includes a processing module for obtaining the access path parameter and determining the decision-making module according to the access path parameter; the processing module can also be used to initiate a call to the resource recommendation model, or can also be used to initiate a call for resource recommendation decision-making through the agent.
[0036] Step S204, if the decision-making module is the policy network, obtain the access status data of the user through the agent, and input the access status data into the policy network to obtain the recommended action value for resource recommendation decision-making.
[0037] As described above, on the premise that the multi-scenario model includes two decision-making modules, namely the resource recommendation model and the policy network, in the process of resource recommendation during the user's access to the resource service, after determining the decision-making module for resource recommendation to the user according to the resource page accessed by the user in the resource service, the determined decision-making module may be the resource recommendation model or may also be the policy network. Here, if the determined decision-making module is the policy network, the access status data of the user is obtained through the agent, and the access status data is input into the policy network to obtain the recommended action value for resource recommendation decision-making.
[0038] Step S206: Select an action in the action space according to the recommended action value, and execute the selected resource recommendation action.
[0039] In specific implementation, after inputting the access status data into the policy network to obtain the recommended action value for resource recommendation decision-making, an action is selected in the action space according to the recommended action value, and the selected resource recommendation action is executed. Among them, in the process of selecting an action in the action space according to the recommended action value, in order to maximize the probability of the user's conversion from resource service access along the access path to resource purchase, in an optional implementation manner provided in this embodiment, selecting an action in the action space according to the recommended action value includes: calculating the path value of each resource page corresponding to each resource recommendation action for each preset access path according to the recommended action value; selecting the target resource recommendation action corresponding to the resource page of the target access path whose path value meets the condition in the action space. Among them, the path value meeting the condition may be sorting the path values of each preset access path in descending order and selecting the preset access path ranked first.
[0040] As described above, on the premise that the multi-scenario model includes two decision-making modules, namely the resource recommendation model and the policy network, in the process of resource recommendation during the user's access to the resource service, after determining the decision-making module for resource recommendation according to the resource page accessed by the user during the resource service access, the determined decision-making module may be the resource recommendation model or the policy network. Here, if the determined decision-making module is the resource recommendation model, resource recommendation processing can be performed through the resource recommendation model. Specifically, in an optional implementation manner provided in this embodiment, if the decision-making module is the resource recommendation model, input the user's access behavior data into the resource recommendation model to obtain resource recommendation data for resource recommendation decision-making, and generate a resource page containing the resource recommendation data; or input the user's access behavior data into the resource recommendation model to obtain a resource recommendation action, and execute the obtained resource recommendation action.
[0041] In summary, in the resource service access recommendation method provided in this embodiment, during the process of resource recommendation when a user accesses a resource service through a multi-scenario model, on the premise that the multi-scenario model includes two decision-making modules, namely a resource recommendation model and a policy network obtained by an agent through reinforcement learning based on a state space and an action space, according to the access path parameters of the resource page accessed by the user in the resource service, the decision-making module for resource recommendation to the user is determined. On the one hand, if the determined decision-making module is the policy network, the agent obtains the user's access status data and inputs the access status data into the policy network for resource recommendation decision-making to obtain a recommended action value. An action is selected in the action space according to the recommended action value, and the selected resource recommendation action is executed. In this way, targeted resource recommendation is made to the user in a specific scenario of the user accessing the resource service by means of the policy network obtained by reinforcement learning, which helps to increase the possibility of the user converting from accessing the resource service along the access path to purchasing the resource. On the other hand, if the determined decision-making module is the resource recommendation model, the user's access behavior data is input into the resource recommendation model for resource recommendation decision-making to obtain resource recommendation data, and a resource page containing the resource recommendation data is generated; or, the user's access behavior data is input into the resource recommendation model for resource recommendation decision-making to obtain a resource recommendation action, and the obtained resource recommendation action is executed. In this way, through the cooperation of the resource recommendation model and the policy network, corresponding resource recommendations are made in different scenarios, so as to further increase the possibility of the user converting from accessing the resource service along the access path to purchasing the resource.
[0042] The following takes the application of the resource service access recommendation method provided in this embodiment in the resource service scenario as an example, and in combination with Figure 3 , the resource service access recommendation method provided in this embodiment is further described. See Figure 3 , the resource service access recommendation method applied to the resource service scenario specifically includes the following steps.
[0043] Step S302, read the access path parameters of the resource page accessed by the user in the resource service.
[0044] Step S304, determine the decision-making module for resource recommendation in the multi-scenario model according to the access path parameters; Optionally, the decision-making module includes a resource recommendation model and a policy network obtained by an agent through reinforcement learning based on a state space and an action space; If the decision-making module is the policy network, execute steps S306 to S310; If the decision-making module is the policy network, execute steps S312 to S314.
[0045] Step S306: The agent obtains the user's access status data and inputs the access status data into the policy network for resource recommendation decision-making to obtain the recommended action value.
[0046] Step S308: Calculate the path value of each resource page corresponding to each resource recommendation action for each preset access path according to the recommended action value.
[0047] Step S310: Select the target resource recommendation action corresponding to the resource page of the target access path whose path value meets the condition in the action space and execute the action.
[0048] Step S312: Input the user's access behavior data into the resource recommendation model for resource recommendation decision-making to obtain the resource recommendation action.
[0049] Step S314: Execute the obtained resource recommendation action.
[0050] It should be noted that any one step or any combination of steps from Step S302 to Step S314 can be combined with any one step or any combination of the above steps S202 to S206 according to the needs of implementation and deployment to form a new implementation method; in addition, according to the actual deployment needs, any one or any combination of technical features in Step S302 to Step S314 can be combined with any one or more technical features provided by the above steps S202 to S206 to form a new implementation method; or, any one or any combination of technical features in Step S302 to Step S314 can also be replaced by any one or more technical features provided by the above steps S202 to S206 according to the actual deployment needs to form a new implementation method, which will not be elaborated here one by one.
[0051] One or more embodiments of a model training method provided in this specification are as follows: Refer to Figure 4 In this embodiment, the model training method specifically includes steps S402 to S406.
[0052] Step S402: Construct an initialization network for determining the mapping relationship between the state space and the action space; Step S404: The agent selects an action in the action space and executes the selected resource recommendation action; Step S406: Determine the reward according to the action execution feedback of the selected resource recommendation action, and adjust the parameters of the preset policy network based on the reward to obtain a multi-scenario model including the policy network and the resource recommendation model after the reinforcement learning is completed.
[0053] Optionally, the resource recommendation model is obtained by training a pre-constructed model to be trained, and the model to be trained can be constructed based on a state space and an action space.
[0054] In this embodiment, the multi-scenario model refers to a model that can adapt to and process different scenarios, conditions, or tasks during the user's access to the resource service, or a model composed of multiple sub-models and / or networks that can work together. Specifically, multiple decision modules are set in the multi-scenario model to perform resource recommendation for different scenarios of the user's access to the resource service. Among them, different scenarios of the user's access to the resource service can be different resource pages of the user's access to the resource service, that is: corresponding resource recommendations are made for different resource pages of the user's access to the resource service through multiple decision modules in the multi-scenario model, so as to increase the probability of the user's conversion from access to resource purchase in the resource service.
[0055] Optionally, the decision module includes a resource recommendation model and a policy network obtained by the intelligent agent through reinforcement learning based on the state space and the action space. In the specific execution process, when the decision module of the multi-scenario model includes a resource recommendation model and a policy network, the resource recommendation model and the policy network for resource recommendation for different scenarios can be obtained through multi-task learning. The generation processes of the resource recommendation model and the policy network are specifically described below.
[0056] (1) Policy network The role of the policy network is to determine the actions that the intelligent agent should take in a given state, and help the intelligent agent perform policy learning through interaction with the environment. Among them, the state can be provided by the state space, the action can be provided by the action space, and the policy learning can be performed by means of reinforcement learning. Optionally, the state space includes the user's access state data; the action space includes a resource recommendation action for recommending data for at least one resource in each resource page of the preset access path.
[0057] Optionally, the access state data includes: page access data of the previous resource page of the current resource page in the preset access path, and / or access behavior data.
[0058] Among them, the page access data refers to data related to the user's behavior of accessing the resource page generated during the user's access to the resource page of the resource service. For example, the page access data includes at least one of the following: object data of the resource object in the resource page accessed by the user, the resource pages already accessed by the user, operation action data of the user on the accessed resource page, and the residence time of the user on the accessed resource page.
[0059] The access behavior data can be the historical access records of the user's previous access to resource services, or the program access records of the host program for the user to access resource services, or the access records composed of both the historical access records and the program access records. In addition, the access behavior data can also be composed of at least one of the historical access records and the program access records and the user data. The user data can be identity data, risk data, credit data, and / or salary data.
[0060] The resource recommendation action refers to the recommendation action of recommending resources to the user by displaying recommended data on the resource page during the process of the user accessing the resource page of the resource service. Specifically, the resource recommendation action can be to display instantiated controls, copywriting, or other forms of recommended materials on the resource page. For example, display management suggestions for resource management on the object details page, or display purchase suggestions for the current resource object; for another example, display historical revenue reminders for the current resource object and purchase trigger controls for the resource object on the buy resource details page. The resource recommendation action can also be to display corresponding instantiated controls, copywriting, or other forms of recommended materials for one or more resource objects displayed on the resource page. For example, display value change reminders for the current resource object in the area below a certain resource object on the resource list page; for another example, display historical revenue reminders for the current resource object and purchase trigger controls for the resource object in the area below a certain resource object on the buy resource page. In the specific execution process, based on the state space and the action space, a policy network is generated by means of reinforcement learning. In an optional implementation manner provided in this embodiment, the policy network is obtained in the following manner: Construct an initialization network for determining the mapping relationship between the state space and the action space; The agent selects an action in the action space and executes the selected resource recommendation action; Determine the reward according to the action execution feedback of the selected resource recommendation action, and adjust the parameters of the initialization network based on the reward to obtain the policy network after the reinforcement learning is completed.
[0061] For example, in the process of training a generation policy network using deep reinforcement learning, the goal is to maximize the probability that a user completes conversion according to access path a, access path b, access path c, access path d, or access path e. First, determine the state space S = (A_state, B_state, C_state, D_state, E_state) for deep reinforcement learning. Among them, A_state, B_state, C_state, D_state, and E_state respectively represent the access status data of the five resource pages: resource list page A, purchased resource page B, object details page C, purchased resource details page D, and resource purchase page E. The access status data A_state includes: the object data of the resource object in the resource page accessed by the user, the resource pages already accessed by the user, the operation action data of the user on the accessed resource page, and / or the residence duration of the user on the accessed resource page; the access status data B_state, access status data C_state, access status data D_state, and access status data E_state are similar to the access status data A_state, and also include the object data of the resource object in the resource page accessed by the user, the resource pages already accessed by the user, the operation action data of the user on the accessed resource page, and / or the residence duration of the user on the accessed resource page. And, determine the action space A = (A_action, B_action, C_action, D_action, E_action). Among them, A_action, B_action, C_action, D_action, and E_action respectively represent the resource recommendation actions executed on resource list page A, purchased resource page B, object details page C, purchased resource details page D, and resource purchase page E. The resource recommendation action A_action ∈ {resource recommendation data 1, resource recommendation data 2,...}. On this basis, construct an initial network according to the mapping relationship between the state space and the action space, and set the initial parameters of the initial network. In the process of deep reinforcement learning, select a resource recommendation action in the action space according to the current state, determine the reward according to the action execution feedback of the selected resource recommendation action, and adjust the parameters of the initial network based on the reward. Repeat the above deep reinforcement learning process until a policy network is obtained after meeting the preset convergence conditions. The convergence conditions can be reaching a preset number of iterations, the change in cumulative reward being less than a preset threshold, etc.
[0062] In addition to generating the policy network by using the deep reinforcement learning method provided above, the Q-Learning method can also be used for reinforcement learning to obtain the policy network. Specifically, based on the state space and the action space, a Q-table is constructed and initialized based on the mapping relationship between the state space and the action space. A resource recommendation action is selected from the action space according to the current state, and the reward is determined according to the action execution feedback of the selected resource recommendation action. The Q-table is adjusted based on the reward, and the above reinforcement learning process is repeated until the policy network is obtained after meeting the preset convergence conditions. The convergence conditions can be reaching the preset number of iterations, the change in the cumulative reward being less than the preset threshold, etc.
[0063] (2) Resource recommendation model In an alternative implementation provided in this embodiment, the resource recommendation model is obtained by training in the following manner: Input the data sample into the model to be trained for resource recommendation decision-making to obtain resource recommendation data; Calculate the training loss based on the resource recommendation data and the data label and adjust the parameters of the model to be trained, so as to obtain the resource recommendation model after the training is completed.
[0064] Among them, the model to be trained can be constructed based on the state space and the action space, and can be specifically obtained by modeling based on the mapping relationship between the state space and the action space. In addition to the above-mentioned implementation manner of training the model to be trained in a supervised manner to obtain the resource recommendation model, the model to be trained can also be trained in an unsupervised manner to obtain the resource recommendation model.
[0065] It should be noted that the descriptions of steps S402 to S406 above are only illustrative. For the specific descriptions of the corresponding contents involved in steps S402 to S406, reference can also be made to the corresponding contents provided in steps S202 to S206 of the above method embodiment, which will not be elaborated here in this embodiment.
[0066] An embodiment of an access recommendation device for a resource service provided in this specification is as follows: In the above embodiment, an access recommendation method for a resource service is provided. Correspondingly, an access recommendation device for a resource service is also provided, which will be described below with reference to the drawings.
[0067] Refer to Figure 5 , which shows a schematic diagram of an embodiment of an access recommendation device for a resource service provided in this embodiment.
[0068] Since the device embodiment corresponds to the method embodiment, the description is relatively simple. For the relevant parts, please refer to the corresponding description of the method embodiment provided above. The device embodiments described below are only illustrative.
[0069] This embodiment provides an access recommendation device for resource services. The device includes: A decision module determination module 502, configured to determine a decision module for resource recommendation in a multi-scenario model according to access path parameters of a resource page accessed by a user; the decision module includes a resource recommendation model and a policy network obtained by an agent through reinforcement learning based on a state space and an action space; A recommendation decision module 504, configured to, if the decision module is the policy network, obtain access state data of the user through the agent, and input the access state data into the policy network to make a resource recommendation decision to obtain a recommended action value; An action selection execution module 506, configured to select an action in the action space according to the recommended action value, and execute the selected resource recommendation action.
[0070] An embodiment of a model training device provided in this specification is as follows: In the above embodiment, a model training method is provided. Correspondingly, a model training device is also provided. The following is described with reference to the accompanying drawings.
[0071] Refer to Figure 6 , which shows a schematic diagram of an embodiment of a model training device provided in this embodiment.
[0072] Since the device embodiment corresponds to the method embodiment, the description is relatively simple. For the relevant parts, please refer to the corresponding description of the method embodiment provided above. The device embodiments described below are merely illustrative.
[0073] This embodiment provides a model training device. The device includes: A network construction module 602, configured to construct an initialization network for determining a mapping relationship between a state space and an action space; An action selection execution module 604, configured to select an action in the action space through an agent, and execute the selected resource recommendation action; A parameter adjustment module 606, configured to determine a reward according to an action execution feedback of the selected resource recommendation action, and adjust parameters of a preset policy network based on the reward to obtain a multi-scenario model including a policy network and a resource recommendation model after reinforcement learning is completed; Wherein, the resource recommendation model is obtained by training a to-be-trained model constructed based on the state space and the action space.
[0074] An embodiment of a model training device provided in this specification is as follows: Corresponding to the model training method described above, based on the same technical concept, one or more embodiments of this specification also provide a model training device, which is used to execute the model training method provided above. Figure 7 It is a schematic structural diagram of a model training device provided by one or more embodiments of this specification.
[0075] An access recommendation device for resource services provided in this embodiment includes: As Figure 7 shown, the access recommendation device for resource services may vary greatly due to configuration or performance, and may include one or more processors 701 and a memory 702. One or more application programs or data may be stored in the memory 702. Among them, the memory 702 may be transient storage or persistent storage. The application programs stored in the memory 702 may include one or more modules (not shown in the figure), and each module may include a series of computer-executable instructions in the access recommendation device for resource services. Further, the processor 701 may be configured to communicate with the memory 702 and execute a series of computer-executable instructions in the memory 702 on the access recommendation device for resource services. The access recommendation device for resource services may also include one or more power supplies 703, one or more wired or wireless network interfaces 704, one or more input / output interfaces 705, one or more keyboards 706, etc.
[0076] In a specific embodiment, the access recommendation device for resource services includes a memory and one or more programs, where one or more programs are stored in the memory, and one or more programs may include one or more modules, and each module may include a series of computer-executable instructions in the access recommendation device for resource services, and is configured to be executed by one or more processors. The one or more programs include the following computer-executable instructions: Determine a decision-making module for resource recommendation in the multi-scenario model according to the access path parameters of the resource page accessed by the user; the decision-making module includes a resource recommendation model and a policy network obtained by the agent through reinforcement learning based on the state space and the action space; If the decision-making module is the policy network, obtain the access state data of the user through the agent, and input the access state data into the policy network to make a resource recommendation decision to obtain a recommended action value; Select an action in the action space according to the recommended action value, and execute the selected resource recommendation action.
[0077] An embodiment of the model training device provided in this specification is as follows: Corresponding to the model training method described above, based on the same technical concept, one or more embodiments of this specification also provide a model training device, which is used to execute the model training method provided above. Figure 8 It is a schematic structural diagram of a model training device provided by one or more embodiments of this specification.
[0078] A model training device provided in this embodiment includes: As Figure 8 shown, the model training device may vary greatly due to configuration or performance, and may include one or more processors 801 and a memory 802. One or more applications or data may be stored in the memory 802. Among them, the memory 802 may be short-term storage or persistent storage. The application programs stored in the memory 802 may include one or more modules (not shown in the figure), and each module may include a series of computer-executable instructions in the model training device. Further, the processor 801 may be set to communicate with the memory 802 and execute a series of computer-executable instructions in the memory 802 on the model training device. The model training device may also include one or more power supplies 803, one or more wired or wireless network interfaces 804, one or more input / output interfaces 805, one or more keyboards 806, etc.
[0079] In a specific embodiment, the model training device includes a memory and one or more programs, where one or more programs are stored in the memory, and one or more programs may include one or more modules, and each module may include a series of computer-executable instructions in the model training device, and is configured to be executed by one or more processors. The one or more programs include the following computer-executable instructions: Construct an initialization network for determining the mapping relationship between the state space and the action space; Select an action in the action space through an agent and execute the selected resource recommendation action; Determine a reward based on the action execution feedback of the selected resource recommendation action, and adjust the parameters of the preset policy network based on the reward to obtain a multi-scenario model including a policy network and a resource recommendation model after reinforcement learning is completed; Among them, the resource recommendation model is obtained by training a model to be trained constructed based on the state space and the action space.
[0080] An embodiment of a computer-readable storage medium provided in this specification is as follows: Corresponding to the access recommendation method for a resource service described above, based on the same technical concept, one or more embodiments of this specification also provide a computer-readable storage medium.
[0081] The computer-readable storage medium provided in this embodiment is used to store computer-executable instructions, and when the computer-executable instructions are executed, the following process is implemented: According to the access path parameters of the resource page accessed by the user in the resource service, determine the decision-making module for resource recommendation in the multi-scenario model; the decision-making module includes a resource recommendation model and a policy network obtained by an agent through reinforcement learning based on the state space and the action space; If the decision-making module is the policy network, obtain the access status data of the user through the agent, and input the access status data into the policy network to make a resource recommendation decision to obtain a recommended action value; Perform action selection in the action space according to the recommended action value, and execute the selected resource recommendation action.
[0082] It should be noted that the embodiment of the computer-readable storage medium in this specification and the embodiment of the access recommendation method for a resource service in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the corresponding method described above, and the repeated parts will not be elaborated.
[0083] Another embodiment of the computer-readable storage medium provided in this specification is as follows: Corresponding to another model training method described above, based on the same technical concept, one or more embodiments of this specification also provide another computer-readable storage medium.
[0084] The computer-readable storage medium provided in this embodiment is used to store computer-executable instructions, and when the computer-executable instructions are executed, the following process is implemented: Construct an initialization network for determining the mapping relationship between the state space and the action space; Perform action selection in the action space through the agent, and execute the selected resource recommendation action; Determine the reward according to the action execution feedback of the selected resource recommendation action, and adjust the parameters of the preset policy network based on the reward to obtain a multi-scenario model including a policy network and a resource recommendation model after the reinforcement learning is completed; Among them, the resource recommendation model is obtained by training a to-be-trained model constructed based on the state space and the action space.
[0085] It should be noted that the embodiments of another computer-readable storage medium in this specification and the embodiments of a model training method in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the corresponding method described above, and the repeated parts will not be elaborated.
[0086] An embodiment of a computer program product provided in this specification is as follows: Corresponding to the above-described method for accessing and recommending resources of a resource service, based on the same technical concept, one or more embodiments of this specification also provide a computer program product.
[0087] A computer program product includes a computer program / instructions, and when the computer program / instructions are executed by a processor, the following steps are implemented: According to the access path parameters of the resource page accessed by the user in the resource service, determine the decision-making module for resource recommendation in the multi-scenario model; the decision-making module includes a resource recommendation model and a policy network obtained by an agent through reinforcement learning based on a state space and an action space; If the decision-making module is the policy network, obtain the access state data of the user through the agent, and input the access state data into the policy network to make a resource recommendation decision to obtain a recommended action value; Perform action selection in the action space according to the recommended action value, and execute the selected resource recommendation action.
[0088] It should be noted that the embodiments of a computer program product in this specification and the embodiments of a method for accessing and recommending resources of a resource service in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the corresponding method described above, and the repeated parts will not be elaborated.
[0089] Another embodiment of a computer program product provided in this specification is as follows: Corresponding to the above-described model training method, based on the same technical concept, one or more embodiments of this specification also provide a computer program product.
[0090] A computer program product includes a computer program / instructions, and when the computer program / instructions are executed by a processor, the following steps are implemented: Construct an initialization network for determining the mapping relationship between the state space and the action space; Perform action selection in the action space through the agent, and execute the selected resource recommendation action; Determine a reward according to the action execution feedback of the selected resource recommendation action, and adjust the parameters of a preset policy network based on the reward to obtain a multi-scenario model including a policy network and a resource recommendation model after the reinforcement learning is completed; Among them, the resource recommendation model is obtained by training a to-be-trained model constructed based on the state space and the action space.
[0091] It should be noted that the embodiments of a computer program product in this specification and the embodiments of a model training method in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the corresponding method described above, and the repeated parts will not be elaborated.
[0092] The embodiments in this specification are all described in a progressive manner. For the same or similar parts between the embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. For example, the device embodiments, equipment embodiments, computer-readable storage medium embodiments, and computer program product embodiments are all similar to the method embodiments, so the descriptions are relatively simple. Please refer to the relevant parts of the method embodiments for reading the relevant content in the device embodiments, equipment embodiments, computer-readable storage medium embodiments, and computer program product embodiments.
[0093] The specific embodiments of this specification are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order from that in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0094] In the 1930s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compilers used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL). There is not just one type of HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing some logical programming on the method flow using the above-mentioned several hardware description languages and programming it into the integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.
[0095] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that, in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same functions. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.
[0096] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0097] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing the embodiments of this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0098] Those skilled in the art should understand that one or more embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0099] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable resource service access recommended devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable resource service access recommended devices generate means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0100] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable resource service access recommended device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0101] These computer program instructions can also be loaded onto a computer or other programmable resource service access recommended device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0102] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0103] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0104] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer-readable storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0105] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of further restrictions, the elements defined by the sentence "includes at least one ..." do not exclude the presence of other identical elements in the process, method, commodity or device including the elements.
[0106] One or more embodiments of the present specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the present specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0107] The above description is only an embodiment of this document and is not intended to limit this document. For those skilled in the art, this document may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this document should be included in the scope of the claims of this document.
Claims
1. A method for recommending access to a resource service, comprising: Determine a decision module for resource recommendation in a multi-scenario model based on access path parameters of a resource page accessed by a user in a resource service; The decision-making module includes a resource recommendation model and a strategy network obtained by the agent through reinforcement learning based on the state space and action space; If the decision module is the policy network, the access status data of the user is obtained through the agent, and the access status data is input into the policy network to make a resource recommendation decision to obtain a recommended action value; An action is selected in the action space according to the recommended action value, and the selected resource recommended action is executed.
2. According to the access recommendation method for resource services described in claim 1, the state space includes the user's access state data; the action space includes a resource recommendation action for recommending data for at least one resource object in each resource page of a preset access path.
3. The resource service access recommendation method according to claim 2, wherein the access status data comprises: The page access data and access behavior data of the resource page before the current resource page in the preset access path.
4. According to the resource service access recommendation method of claim 1, the policy network is obtained in the following manner: Constructing an initialization network for determining a mapping relationship between the state space and the action space; Selecting an action in the action space by the agent and executing the selected resource recommendation action; A reward is determined according to the action execution feedback of the selected resource recommendation action, and parameters of the initialized network are adjusted based on the reward to obtain the policy network after reinforcement learning is completed.
5. The resource service access recommendation method according to claim 1, wherein the step of selecting an action in the action space according to the recommended action value comprises: Calculate the path value of each resource recommended action corresponding to each resource page of each preset access path according to the recommended action value; A target resource recommended action corresponding to a resource page of a target access path whose path value satisfies a condition is selected in the action space.
6. The access recommendation method for resource services according to claim 1, after the step of determining the decision module for resource recommendation in the multi-scenario model according to the access path parameters of the resource page accessed by the user in the resource service is executed, it also includes: If the decision module is the resource recommendation model, the user's access behavior data is input into the resource recommendation model to make a resource recommendation decision to obtain resource recommendation data; A resource page including the resource recommendation data is generated.
7. According to the resource service access recommendation method of claim 4, the resource recommendation model is trained and obtained in the following manner: Input the data sample into the model to be trained to make resource recommendation decisions and obtain resource recommendation data; The model to be trained is constructed based on the state space and the action space; The training loss is calculated based on the resource recommendation data and the data labels, and the parameters of the model to be trained are adjusted to obtain the resource recommendation model after the training is completed.
8. The access recommendation method for resource services according to claim 1, wherein the decision module for making resource recommendations in the multi-scenario model is determined according to the access path parameters of the resource page accessed by the user in the resource service, comprising: If the access path parameter is a first node type, determining the resource recommendation model as the decision module; If the access path parameter is an intermediate node type, determining the policy network as the decision module; The first node type and the intermediate node type are determined according to the node where the resource page is located in the access path.
9. The access recommendation method for resource services according to claim 1, wherein the multi-scenario model further comprises a processing module for acquiring the access path parameters and determining the decision module according to the access path parameters; The processing module is further used to initiate a call to the resource recommendation model, or to initiate a call to make a resource recommendation decision through the agent.
10. A model training method, comprising: Construct an initialization network for determining the mapping relationship between the state space and the action space; Selecting an action in the action space through an intelligent agent and executing the selected resource recommendation action; Determine a reward according to the action execution feedback of the selected resource recommendation action, and adjust parameters of a preset policy network based on the reward to obtain a multi-scenario model including a policy network and a resource recommendation model after reinforcement learning is completed; The resource recommendation model is obtained by performing model training on a to-be-trained model constructed based on the state space and the action space.
11. A model training device, comprising: A decision module determination module is configured to determine a decision module for resource recommendation in a multi-scenario model according to access path parameters of a resource page accessed by a user in a resource service; the decision module includes a resource recommendation model and a strategy network obtained by an intelligent agent through reinforcement learning based on a state space and an action space; A recommendation decision module is configured to obtain the user's access status data through the agent if the decision module is the policy network, and input the access status data into the policy network to make a resource recommendation decision to obtain a recommended action value; The action selection and execution module is configured to select an action in the action space according to the recommended action value and execute the selected resource recommendation action.
12. A resource service access recommendation device, comprising: A network construction module is configured to construct an initialization network for determining a mapping relationship between a state space and an action space; an action selection and execution module, configured to select an action in the action space through an agent and execute the selected resource recommendation action; a parameter adjustment module configured to determine a reward according to the action execution feedback of the selected resource recommendation action, and to adjust parameters of a preset policy network based on the reward, so as to obtain a multi-scenario model including a policy network and a resource recommendation model after reinforcement learning is completed; The resource recommendation model is obtained by performing model training on a to-be-trained model constructed based on the state space and the action space.
13. A resource service access recommendation device, comprising: processor; and a memory configured to store computer executable instructions that, when executed, cause the processor to: Determine a decision module for resource recommendation in a multi-scenario model according to access path parameters of a resource page accessed by a user in a resource service; the decision module includes a resource recommendation model and a strategy network obtained by an intelligent agent through reinforcement learning based on a state space and an action space; If the decision module is the policy network, the access status data of the user is obtained through the agent, and the access status data is input into the policy network to make a resource recommendation decision to obtain a recommended action value; An action is selected in the action space according to the recommended action value, and the selected resource recommended action is executed.
14. A model training device, comprising: processor; and a memory configured to store computer executable instructions that, when executed, cause the processor to: Construct an initialization network for determining the mapping relationship between the state space and the action space; Selecting an action in the action space through an intelligent agent and executing the selected resource recommendation action; Determine a reward according to the action execution feedback of the selected resource recommendation action, and adjust parameters of a preset policy network based on the reward to obtain a multi-scenario model including a policy network and a resource recommendation model after reinforcement learning is completed; The resource recommendation model is obtained by performing model training on a to-be-trained model constructed based on the state space and the action space.
15. A computer-readable storage medium for storing computer-executable instructions, wherein the computer-executable instructions implement the steps of the method of claim 1 or 10 when executed.