Model training method, information release method and device and electronic equipment
通过双层强化学习模型选择和生成投放节点及策略,解决了社交平台信息投放系统中精准度和传播效果不佳的问题,实现了更高效的信息传播。
Patent Information
- Application Number
- CN202510306386.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-11
AI Technical Summary
The existing information delivery system lacks a deep understanding of the complex network structure and user behavior of social platforms, resulting in poor accuracy of information delivery and poor communication effect, making it difficult to achieve accurate target user coverage and efficient dissemination.
The two-layer reinforcement learning model architecture is adopted, and the delivery node is selected through the first reinforcement learning model, and the second reinforcement learning model generates delivery strategies, simulates the information dissemination process and optimizes the model parameters to adapt to user behavior characteristics and network structure characteristics.
It improves the accuracy and dissemination effect of information delivery, ensures that information can reach the target users to the greatest extent, realizes the synergy between delivery channels and strategies, and reduces training complexity and cost.
Smart Images

Figure CN120297436A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of artificial intelligence technology, and particularly to a method for training a model, a method for information delivery, an apparatus, and an electronic device. Background Art
[0002] With the rapid development of social media, the ways and means of information dissemination have undergone great changes, and enterprises and advertisers have turned to social platforms for information delivery and marketing.
[0003] Currently, most information delivery systems adopt a static analysis method based on the basic characteristics of users (such as gender, age, region, etc.) and historical behaviors, relying on simple rules or data mining algorithms to determine target users and dissemination strategies. However, this method does not fully consider the dynamic changes and complex user relationships in the information dissemination on social platforms, lacks a deep understanding of the complex network structure and user behaviors on social platforms, resulting in poor accuracy of information delivery, ineffective information dissemination, and difficulty in achieving precise target user coverage and efficient dissemination. Summary of the Invention
[0004] To overcome the problems existing in the related art, this specification provides a method for training a model, a method for information delivery, an apparatus, and an electronic device.
[0005] According to a first aspect of an embodiment of this specification, a method for training a model is provided, and the method includes:
[0006] Selecting at least one simulated delivery node from candidate delivery nodes in an information dissemination network through a first reinforcement learning model; wherein, the information dissemination network is constructed according to a target social platform, the nodes of the information dissemination network include candidate delivery nodes and forwarding nodes, and the edges of the information dissemination network are used to represent the association relationships between nodes;
[0007] Generating a simulated delivery strategy for each simulated delivery node through a second reinforcement learning model;
[0008] Generating simulated release information respectively according to the simulated delivery strategies of each simulated delivery node, and simulating the forwarding process of the forwarding nodes based on the simulated release information delivered by each simulated delivery node to generate a simulated dissemination result of the information dissemination network;
[0009] Updating the parameters of the first reinforcement learning model and the second reinforcement learning model according to the number of forwarding times of the forwarding nodes in the simulated dissemination result as a reward.
[0010] According to a second aspect of an embodiment of this specification, a method for information delivery is provided, and the method includes:
[0011] Obtain a first reinforcement learning model and a second reinforcement learning model trained for a target social platform, where the first reinforcement learning model and the second reinforcement learning model are obtained by the model training method described in the first aspect;
[0012] Select at least one actual placement node from the candidate placement nodes in the information dissemination network through the first reinforcement learning model;
[0013] Generate an actual placement strategy for each actual placement node through the second reinforcement learning model, where the actual placement strategy is used to indicate generating the actual release information content to be placed for the actual placement node.
[0014] According to the third aspect of the embodiments of this specification, there is provided a model training device, where the device includes:
[0015] A simulated node selection module, configured to select at least one simulated placement node from the candidate placement nodes in the information dissemination network through the first reinforcement learning model; where the information dissemination network is constructed according to the target social platform, the nodes of the information dissemination network include candidate placement nodes and forwarding nodes, and the edges of the information dissemination network are used to represent the association relationship between the nodes;
[0016] A simulated strategy generation module, configured to generate a simulated placement strategy for each simulated placement node through the second reinforcement learning model;
[0017] A simulated dissemination module, configured to respectively generate simulated release information according to the simulated placement strategies of the respective simulated placement nodes, and simulate the forwarding process of the forwarding nodes based on the simulated release information placed by the respective simulated placement nodes to generate the simulated dissemination result of the information dissemination network;
[0018] A parameter update module, configured to update the parameters of the first reinforcement learning model and the second reinforcement learning model according to the number of forwarding times of the forwarding nodes in the simulated dissemination result as a reward.
[0019] According to the fourth aspect of the embodiments of this specification, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, it implements the steps of the method described in the first aspect or the second aspect.
[0020] According to the fifth aspect of the embodiments of this specification, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of the method described in the first aspect or the second aspect.
[0021] The technical solutions provided by the embodiments of this specification may include the following beneficial effects:
[0022] In the embodiments of this specification, this solution uses a first reinforcement learning model to select simulated placement nodes from candidate placement nodes in the information dissemination network, and uses a second reinforcement learning model to generate simulated placement strategies for each simulated placement node. Then, after each simulated placement node generates simulated release information according to the simulated placement strategy, the information dissemination process between nodes in the information dissemination network is simulated. Finally, according to the simulated dissemination results, the parameters of the first reinforcement learning model and the second reinforcement learning model are optimized to adapt to the user behavior characteristics and network structure characteristics of the target social platform, so as to ultimately achieve a better dissemination effect when performing real placement on the target social platform using the decision-making results of the first reinforcement learning model and the second reinforcement learning model.
[0023] It can be seen that, first of all, the first reinforcement learning model and the second reinforcement learning model trained by this solution can be used to assist manual selection of actual placement nodes and generation of actual placement strategies. Since the first reinforcement learning model and the second reinforcement learning model are adjusted through feedback according to the network dissemination effect multiple times during the simulation phase, their decision-making results take into account user behavior characteristics and network structure characteristics. Therefore, the placement tasks generated by using them can ensure that information can reach the target users to the greatest extent, thereby improving the accuracy of information placement.
[0024] Secondly, the first reinforcement learning model designed in this solution is used to select placement nodes, and the second reinforcement learning model is used to generate placement strategies for the placement nodes selected by the first reinforcement learning model. Due to this two-layer learning architecture, the collaboration between the two can optimize both the placement channels and the placement strategies simultaneously, and can better exert the synergy of different placement channels, which is conducive to improving the information dissemination effect.
[0025] Finally, the second reinforcement learning model in this solution does not directly generate release information for the placement nodes, but generates placement strategies, and then generates simulated release information according to the placement strategies. On the one hand, this reduces the complexity and training cost of training the second reinforcement learning model; on the other hand, it can also achieve the purpose of collaborative dissemination of placement strategies between different placement nodes, and each placement node can also better exert its subjective initiative according to the placement strategies to generate personalized release content, thereby achieving a better dissemination effect.
[0026] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with this specification, and are used together with the specification to explain the principles of this specification.
[0028] Figure 1It is a flowchart of a method for training a model shown in this specification according to an exemplary embodiment.
[0029] Figure 2 It is a flowchart of a preferred method for training a model shown in this specification according to an exemplary embodiment.
[0030] Figure 3 It is a flowchart of a method for information delivery shown in this specification according to an exemplary embodiment.
[0031] Figure 4 It is a schematic structural diagram of an electronic device shown in this specification according to an exemplary embodiment.
[0032] Figure 5 It is a block diagram of a device for training a model shown in this specification according to an exemplary embodiment.
[0033] Figure 6 It is a block diagram of a device for information delivery shown in this specification according to an exemplary embodiment. Detailed implementation manners
[0034] With the rapid development of social media, the ways and means of information dissemination have undergone great changes. Enterprises and advertisers have turned to social platforms for information delivery and marketing. Among them, social platforms can refer to platforms that connect users through the network and provide functions such as information, interaction, and sharing. These platforms enable users to interact and communicate through content publishing, commenting, liking, forwarding, etc. Enterprises and advertisers can select appropriate target social platforms according to their brand characteristics, utilize the interactivity of the target social platforms, and expand brand influence by posting interactive content on the target social platforms, launching activities and attracting user participation.
[0035] Currently, most information delivery systems adopt static analysis methods based on users' basic characteristics (such as gender, age, region, etc.) and historical behaviors, relying on simple rules or data mining algorithms to determine target users and information delivery strategies. For example, according to the established user portraits, through in-depth analysis of users, personalized advertising content is designed to ensure that the delivered advertisements can effectively reach target users.
[0036] However, this method does not fully consider the dynamic changes in information dissemination and complex user relationships in social platforms, lacks in-depth understanding of the complex network structure and user behaviors in social platforms, resulting in poor accuracy of information delivery, ineffective information dissemination, and difficulty in achieving precise target user coverage and efficient dissemination.
[0037] To address the above technical problems, this specification proposes a method for training a model and a method for information delivery. By deeply analyzing the complex network structure of the social platform and user behavior, it aims to achieve precise target user coverage and efficient dissemination.
[0038] Next, the embodiments of this specification will be described in detail.
[0039] Figure 1 is a flowchart of a method for training a model shown according to an exemplary embodiment of this specification. As Figure 1 shown, it includes steps 101 - 104:
[0040] Step 101: Select at least one simulated delivery node from the candidate delivery nodes in the information dissemination network through a first reinforcement learning model; wherein, the information dissemination network is constructed based on the target social platform, the nodes of the information dissemination network include candidate delivery nodes and forwarding nodes, and the edges of the information dissemination network are used to represent the association relationship between nodes.
[0041] Step 102: Generate a simulated delivery strategy for each simulated delivery node through a second reinforcement learning model.
[0042] Step 103: Generate simulated release information respectively according to the simulated delivery strategies of each simulated delivery node, and simulate the forwarding process of the forwarding nodes based on the simulated release information delivered by each simulated delivery node to generate the simulated dissemination result of the information dissemination network.
[0043] Step 104: Update the parameters of the first reinforcement learning model and the second reinforcement learning model according to the number of forwarding times of the forwarding nodes in the simulated dissemination result as a reward.
[0044] The target social platform can be one selected from numerous social platforms. For example, it can be selected according to the matching degree between social platform users and the brand, the delivery budget, etc. This specification does not impose any restrictions on the type of the target social platform.
[0045] Obtain the relevant information of the candidate cooperation accounts from the target social platform. The account selection can be made according to the brand party's preference. If there is no preference, the candidate cooperation accounts can be sorted in descending order according to the number of fans of the cooperation accounts in the brand-related field, and the top n are selected. After selecting the candidate cooperation accounts, the relevant information of the ordinary user accounts associated with the candidate cooperation accounts can be obtained. This association relationship enables the ordinary user accounts to obtain the information released by the candidate cooperation accounts and forward it. For example, this association relationship can be that the ordinary user account follows the candidate cooperation account, or of course, the ordinary user account and the candidate cooperation account are friends.
[0046] The information dissemination network can be constructed based on the relevant information of candidate cooperation accounts and the relevant information of ordinary user accounts obtained. The nodes of the information dissemination network include candidate placement nodes and forwarding nodes, and the edges of the information dissemination network are used to represent the association relationships between the nodes. The association relationships can be following relationships, friend relationships, etc., and this specification does not impose any restrictions on the association relationships. The candidate placement nodes can represent the relevant information of candidate cooperation accounts, and the forwarding nodes can represent the relevant information of ordinary user accounts. The edges between different nodes can be represented by association relationships. For example, if there is an association relationship between node A and node B, there can be an edge between node A and node B; if there is no association relationship between node A and node B, there may not be an edge between node A and node B. The directionality of the edge can be set according to the type of association relationship. For example, if node A follows node B, the direction of the edge between node A and node B can be from B to A, indicating that the information published by node B can flow to node A. If node A and node B are friends with each other, the edge may have no directionality, indicating that the information published by node A and node B can flow to each other.
[0047] When representing the information dissemination network, a matrix can be used to record the association relationships between candidate placement nodes and forwarding nodes, as well as the relevant information of each node. This specification does not limit the network representation method. Exemplarily, taking the association relationship as a following relationship as an example, the relevant information of candidate placement nodes can be saved in a two-dimensional list, and the relevant information includes anonymized id, field, number of fans, time of publishing topics in the recent k days, content, number of forwards, and the simulated placement strategy adopted for publishing simulated release information in this round of simulation. Similarly, the relevant information of forwarding nodes can be saved in a two-dimensional list, and the relevant information includes anonymized id, published and forwarded topic content in the recent k days, and the information forwarding decision in this round of simulation. The relevant information of the nodes can be queried through the anonymized id. In addition, a separate memory data table can be configured for each anonymized id to store the simulation process in the form of a list.
[0048] After the information dissemination network is constructed, for the training process of any period:
[0049] At least one simulated placement node can be selected from the candidate placement nodes in the information dissemination network through the first reinforcement learning model. This specification does not limit the model architecture of the first reinforcement learning model and the reinforcement learning algorithm adopted by the first reinforcement learning model. The simulated placement node can be a node used for placing and publishing information in the information dissemination of this period. The reward target of the first reinforcement learning model can be that the selected simulated placement node can achieve the optimal information dissemination effect in the simulated information dissemination of this period, such as the largest amount of information forwarding in the network.
[0050] In one embodiment, the state space of the first reinforcement learning model may include the state representation of the information dissemination network, and the action space of the first reinforcement learning model may be to select at least one simulated placement node from the candidate placement nodes. The state representation of the information dissemination network may be input into the first reinforcement learning model, and at least one simulated placement node output by the first reinforcement learning model may be obtained. Exemplarily, the state representation of the information dissemination network may be a comprehensive representation of the network structure features and each node feature of the information dissemination network. This specification does not impose any restrictions on the state representation of the information dissemination network. Preferably, it may specifically be obtained by concatenating the network structure features of the information dissemination network, the candidate placement node vectors, the candidate set feature vectors, and the total limit scalar of the simulated placement nodes. Among them, if a candidate placement node is not selected as a simulated placement node, the vector of this candidate placement node may be 0; if a candidate placement node is selected as a simulated placement node, the vector of this candidate placement node may be the vector representation of the simulated placement strategy. The total limit scalar of the simulated placement nodes may be the maximum threshold of the total number of simulated placement nodes selected by the first reinforcement learning model. When the total number of simulated placement nodes selected by the first reinforcement learning model exceeds the total limit scalar of the simulated placement nodes, the selection may stop. The setting of the total limit scalar of the simulated placement nodes is to optimize the dissemination effect under the condition of limited budget.
[0051] In one embodiment, the cumulative number of selected nodes in the action space may not be greater than the total limit s of the simulated placement nodes. Each action may be represented by a group of vectors, and each element in the vector corresponds to a candidate placement node, where 1 indicates selection and 0 indicates non-selection. The optional action space is various combinations where the number of elements with a value of 1 in the vector is less than s. For example, if there are 3 candidate placement nodes and s = 1, the optional action space is Option 1 (1, 0, 0), Option 2 (0, 1, 0), Option 3 (0, 0, 1), and this action space may be represented by a three-dimensional vector a0 t where the value of each element represents the probability of selecting the corresponding option. For example, (0.5, 0.2, 0.4) represents the probability of selecting Option 1 as 0.5, Option 2 as 0.2, and Option 3 as 0.4.
[0052] After at least one simulated placement node is selected from the candidate placement nodes in the information dissemination network through the first reinforcement learning model, a simulated placement strategy may be generated for each simulated placement node through the second reinforcement learning model.
[0053] Since traditional reinforcement learning methods are usually optimized only for a single task and lack overall optimization of the multi-level decision-making process, traditional reinforcement learning methods cannot consider multiple aspects such as the selection of placement nodes and the generation of placement strategies simultaneously. The first reinforcement learning model designed in this solution is used to select placement nodes, and the second reinforcement learning model is used to generate a placement strategy for the placement nodes selected by the first reinforcement learning model. By adopting this two-layer learning architecture, the two can cooperate to optimize the placement channels and placement strategies simultaneously, better exert the synergy of different placement channels, and contribute to improving the information dissemination effect. Secondly, the second reinforcement learning model of this solution does not directly generate release information for the placement nodes, but generates a placement strategy, and then generates simulated release information according to the placement strategy. On the one hand, it reduces the complexity and training cost of training the second reinforcement learning model; on the other hand, it can also achieve the purpose of using the placement strategy collaboration between different placement nodes for dissemination, and each placement node can also better exert its subjective initiative according to the placement strategy to generate personalized release content, so as to achieve a better dissemination effect.
[0054] In one embodiment, the second reinforcement learning model can directly adopt the model architecture of the existing prompt learning model, which can accurately complete the task according to the given prompt words. Specifically, for each simulated placement node, the first prompt word corresponding to the simulated placement node is input into the second reinforcement learning model, and the simulated placement strategy selected by the second reinforcement learning model according to the first prompt word is obtained. Among them, the first prompt word is used to instruct the second reinforcement learning model to select a simulated placement strategy from the preset placement strategy set. Compared with the traditional technical means that rely on a large number of experiments to optimize the placement strategy, this embodiment applies the prompt learning model to the information placement field and combines the reinforcement learning technology to generate a placement strategy for the placement nodes, which can avoid excessive exploration and resource waste.
[0055] The state information S1 of the second reinforcement learning model t can be obtained by embedding the feature information of the simulated placement node and the information to be promoted into the first prompt word template. Among them, the feature information of the simulated placement node can be the fields actually followed by the cooperative account of the simulated placement node, the posting time, posting content, and effects (such as the number of likes and reposts) of the most recent k posts, and the number of fans owned by the cooperative account.
[0056] Among them, an example of the first prompt word template is as follows:
[0057] Your identity: You are a blogger on the target social platform in the X field. The content of the post you recently published is: 1. You published X and received A likes and B reposts......
[0058] Your task: You have now accepted the promotion task of a product of brand X, which is mainly used for..., and currently the brand promotion has (just started / already has 1% of users forwarding / 10% of users forwarding / ). You need to send a relevant post using your target social media platform account to get the most forwards.
[0059] Please select a placement strategy:
[0060] 1. Post a trial experience
[0061] 2. Use popular topic tags
[0062] 3. Post a poll on relevant questions ......
[0064] Output requirements: Do not output other information, only reply with the option number to indicate your choice.
[0065] As described above, 1. Post a trial experience; 2. Use popular topic tags; 3. Post a poll on relevant questions, etc. are the preset placement strategy sets. Of course, those skilled in the art can set other preset placement strategy sets according to actual placement needs, and this specification does not impose any restrictions on the content of this preset placement strategy set. In addition, those skilled in the art can also customize the content of the first prompt word so that the second reinforcement learning model can select a simulated placement strategy according to the first prompt word, and this specification does not impose any restrictions on the content of the first prompt word.
[0066] The information output by the second reinforcement learning model according to the first prompt word can be the probability a1 of selecting different placement strategies from the preset placement strategy set t 。
[0067] In one embodiment, after generating a simulated placement strategy for each simulated placement node through the second reinforcement learning model, simulated release information can be generated respectively according to the simulated placement strategies of each simulated placement node. Exemplarily, for each simulated placement node, the second prompt word corresponding to the simulated placement node can be input into the network simulation model to obtain the simulated placement information to be placed by the simulated placement node generated by the network simulation model according to the second prompt word; wherein, the second prompt word is used to instruct the network simulation model to generate simulated placement information according to the simulated placement strategy of the simulated placement node.
[0068] Of course, multiple candidate simulated release information sets can also be generated respectively for each simulated placement strategy, and any target simulated release information can be randomly selected from the candidate simulated release information set corresponding to the selected simulated placement strategy as the simulated release information generated by the simulated placement node for the selected simulated placement strategy.
[0069] Among them, the network simulation model can be a prompting learning model, which can embed the feature information of the simulated delivery node and the selected simulated delivery strategy into the second prompt word template to obtain the second prompt word. By inputting the second prompt word into the network simulation model, the generated simulated delivery information can be output. The simulated delivery information can be stored in the corresponding memory data table for subsequent simulation processes.
[0070] An example of the second prompt word template is as follows:
[0071] Your identity: You are a blogger in the X field, and the content of your most recent post is: 1. You posted X before X and received A likes and B reposts.
[0072] Your task: Now you have accepted a promotion task for the X brand. The product is mainly used for... and hopes to get the most reposts. Please form a text of no more than 100 characters around your feelings about the trial use of this product. ......
[0074] Requirements for the output content: Do not output other information, just send the text content.
[0075] It should be noted that this specification does not limit the specific content of the second prompt word, and those skilled in the art can design the second prompt word according to actual needs and the characteristics of the prompting learning model.
[0076] In one embodiment, after the simulated release information is generated respectively according to the simulated delivery strategies of each simulated delivery node, the forwarding process of the forwarding node can be simulated based on the simulated release information delivered by each simulated delivery node to generate the simulated propagation result of the information dissemination network. Exemplarily, this simulation process can adopt the traditional simulation methods involved in network analysis and graph theory to simulate the simulation process of the information dissemination network. Exemplarily, for each forwarding node, the third prompt word corresponding to the forwarding node is input into the network simulation model, and the forwarding decision result output by the network simulation model according to the third prompt word is obtained; wherein, the third prompt word is used to instruct the network simulation model to simulate the forwarding decision process of the forwarding node to generate the forwarding decision result; the forwarding decision results of all forwarding nodes constitute the simulated propagation result, and the forwarding decision result of each forwarding node indicates whether the forwarding node forwards the simulated release information.
[0077] Specifically, traverse the forwarding nodes. If a forwarding node has already forwarded the simulated delivery information published by the simulated delivery node, skip it. If not, the forwarding node can traverse and record the simulated publication information published by the simulated delivery nodes associated with it and the simulated publication information forwarded by other forwarding nodes, embed the information forwarded or published by these related nodes into the third prompt template to generate a third prompt, and input the third prompt into the network simulation model. Simulate the forwarding decision-making process of the forwarding node through this network simulation model to generate a forwarding decision result, which indicates whether the forwarding node forwards the simulated publication information.
[0078] Among them, an example of the third prompt template is as follows:
[0079] You are X, and the content of the post you recently published is: 1. At what time did you publish X. 2......
[0080] The content of the post you recently forwarded is: 1. At what time did you forward the content of X's post
[0081] Your last decision was: A, B published X information, E forwarded Y information...... You chose not to forward the information.
[0082] What you observed this time is: A, B, D published X information, E, F forwarded Y information......
[0083] Your task: Whether to forward a certain piece of information. If you forward, which one to choose to forward?
[0084] Output requirements: If you choose not to forward, output 0; if you choose to forward, give the content to be forwarded. Do not output other information, just send the text content.
[0085] Among them, the content of the third prompt can include the content forwarding and posting situation of the forwarding node in the target social platform in the recent k days, the forwarding decision result of the previous time step, and the information of adjacent nodes observed in the current time step. The propagation situation of information can be counted at each time step.
[0086] In this embodiment, compared with the traditional simulation method involving network analysis and graph theory for simulating the information propagation network. By inputting the behavior information of the forwarding node and the behavior information of its adjacent nodes into the prompt learning model and letting the prompt learning model simulate the decision-making process of users, it is possible to capture the complex influencing factors of the network dependence relationship suffered by each forwarding node in a specific environment, so that the simulated decision result can better reflect the dynamically changing environment in the information propagation process.
[0087] At the end of the current cycle simulation, the parameters of the first reinforcement learning model and the second reinforcement learning model can be updated as rewards based on the number of forwarding times of the forwarding nodes in the simulation propagation results. The reward mechanism can be determined according to the total forwarding volume, and the total forwarding volume is positively correlated with the reward; the reward mechanism can also be determined based on the propagation range. If the simulated release information released by the simulated placement node is forwarded by more different forwarding nodes, a higher reward will be given. Those skilled in the art can flexibly set the reward mechanisms of the first reinforcement learning model and the second reinforcement learning model according to the propagation situation of the forwarding nodes in the simulation propagation results, and this specification does not impose any restrictions on this.
[0088] In one embodiment, for the simulation propagation results at each time step, the parameters of the first reinforcement learning model can be updated as rewards based on the increment of the number of forwarding times at the current time step relative to the previous time step. Among them, the size of the reward is positively correlated with the increment of the number of forwarding times. Of course, a penalty term can also be added to the reward mechanism. If the simulated placement nodes selected in this round do not include the simulated placement nodes selected in the previous round, a penalty positively correlated with the non-included quantity will be given.
[0089] In one embodiment, for each simulated placement node in the simulation propagation results at each time step, the number of forwarding times with the simulated placement node as the source point can be set as a reward to update the parameters of the second reinforcement learning model. For example, the logarithm of the weighted sum of the number of forwarding times with the simulated placement node as the source and the total information forwarding volume can be used as the reward.
[0090] Steps 101-104 show the training process of any cycle. If the information propagation situation in the current cycle changes compared with the information propagation situation in the previous cycle, steps 101-104 are re-executed; if there is no change and both the first reinforcement learning model and the second reinforcement learning model converge, the first reinforcement model and the second reinforcement learning model are trained.
[0091] Next, a training method for a preferred model will be described:
[0092] Figure 2 is a flowchart of a training method for a preferred model shown in this specification according to an exemplary embodiment. As Figure 2 shown, it specifically includes steps 201-209:
[0093] Step 201: Initialize the simulation environment. This includes initializing the first reinforcement learning model, the second reinforcement learning model, and the network simulation model, as well as initializing the information propagation network. Among them, the information propagation network is constructed according to the target social platform. The nodes of the information propagation network include candidate placement nodes and forwarding nodes, and the edges of the information propagation network are used to represent the association relationships between the nodes.
[0094] Step 202: Output a simulated placement node through the first reinforcement learning model.
[0095] Step 203: Generate a simulated placement strategy for the simulated placement node through the second reinforcement learning model.
[0096] Step 204: Generate a simulated release message and a simulated propagation process for the simulated placement node through the network propagation model. Specifically, it includes generating a simulated release message for each simulated placement node according to the simulated placement strategy of each simulated placement node through the network propagation model, and simulating the forwarding process of the forwarding nodes based on the simulated release messages placed by each simulated placement node to generate the simulated propagation result of the information dissemination network.
[0097] Step 205: Determine whether the simulation stop condition for this cycle is satisfied. If not, go to Step 202; if satisfied, go to Step 206. Among them, the simulation stop condition for this cycle can be that the current time step is greater than or equal to T or all adjacent nodes of the forwarding nodes have completed the forwarding decision. Among them, T is the time step threshold.
[0098] Step 206: Statistically analyze the simulated propagation result, and reverse infer the action rewards of each model at each time step to form training samples, and store the training samples in the replay buffer.
[0099] Step 207: Obtain training samples from the replay buffer, and train the first reinforcement learning model and the second reinforcement learning model according to the training samples.
[0100] Step 208: Determine whether the model training completion condition is satisfied. If satisfied, go to Step 209; if not satisfied, go to Step 202. Among them, the condition for satisfying the model training completion can be that the information dissemination situation in the current cycle has no change compared with the information dissemination situation in the previous cycle and both the first reinforcement learning model and the second reinforcement learning model converge. The condition for not satisfying the model training completion can be that the information dissemination situation in the current cycle has a change compared with the information dissemination situation in the previous cycle or the first reinforcement learning model and the second reinforcement learning model do not converge.
[0101] Step 209: Obtain the trained first reinforcement learning model and the second reinforcement learning model.
[0102] Next, this specification introduces the exemplary algorithm implementation processes of the first reinforcement learning model and the second reinforcement learning model:
[0103] S1. Initialize the first reinforcement learning model. Among them, the reinforcement learning algorithm adopted is exemplified by the DQN algorithm, and the first reinforcement learning model is exemplified by the Q-network M0:
[0104] Initialize the Q-network M0(S0 t ,a0 t; θ), M0 receives the state vector S0 t as input and outputs the action vector a0 t (Each element represents the Q-value of choosing an action plan). Set up an experience replay pool to store the tuples of state, action, reward, and next state Replay0(S0 t , a0 t , R0 t , S0 t+1 ).
[0105] S2. Initialize the second reinforcement learning model M1. Among them, the reinforcement learning algorithm adopted is illustrated by the PPO algorithm:
[0106] M1(S1 t , a1 t ; θ). Among them, M1 receives the state S1 t as input and outputs the action vector a1 t (Each element represents the probability of choosing a placement strategy). Set up an experience replay pool to store the tuples of state, action, and reward Replay1(S1 t , a1 t , R1 t ).
[0107] S3. Obtain the state S0 t =(P t , E t ), obtain the Q-value of each action output by M0, and use the ε-greedy policy to select an action: Select a random action (exploration) with probability ε. Select the action with the largest Q(S0 t , a0 t ; θ) value with probability 1-ε (exploitation). Output an action vector a0, and convert the action into the selected node E t+1 . Among them, Diff = E t+1 -E t , and the nodes corresponding to the elements greater than 0 are the newly added simulated placement nodes, forming a set of newly added simulated placement nodes Ncand. At the same time, calculate the number of elements less than 0, denoted as WrC.
[0108] S4. Traverse Ncand, and form the input (state) of the first reinforcement learning model according to the first prompt word. The output is the probability vector P i t j (For example, the probability of the option number can be taken as the probability of each action). Select a random action (exploration) with probability ε. Select the action with the largest F(S1 t , a1 t ; θ) value with probability 1-ε (exploitation). Form (S0 t , a0t ) The simulated delivery strategy under t , the newly added simulated delivery nodes, and the selected simulated delivery strategy.
[0109] S5.1. Traverse the simulated delivery nodes obtained this time, and generate the simulated release content of the simulated delivery nodes based on the selected simulated delivery strategy.
[0110] S5.2. Traverse the forwarding nodes. If a forwarding node has already forwarded relevant information, skip it. If not, the forwarding node will traverse and record the release information of its associated nodes (including the forwarding information of other forwarding nodes) to form a third prompt word, and decide whether to forward the content according to the third prompt word.
[0111] S5.3. Repeat step 5.2 for j times. After repetition, calculate the number of forwards of the simulated release information in this cycle * coefficient to obtain R0 t , and the number of forwards of the release information of each newly added simulated delivery node * coefficient to obtain R1 t . And obtain the result status S1 t+1 . If Wrc > 0, R0 t = R0 t - Wrc * coefficient. Comprehensively obtain (S0 t , a0 t , R0 t , S0 t+1 ) and the corresponding (S1 tj , a1 tj , Pi tj , R1 tj ) for each newly added cooperation node, and store them in Replay0 and Replay1.
[0112] S6. Model training. When the number of elements in Replay0 > Nr, perform model training; take out a group of (S0 t , a0 t , R0 t , S0 t+1 ) and (S1 t , a1 t , R1 t ) from Replay0 and Replay1; calculate the loss function of M0, L0(θ) = (y - Q(S0 t , a0 t ; θ))2, where y = r + γmaxQ(S0 t+1 , a0 t+1 ; θ) and update the parameters of model M0 through the backpropagation method. Calculate the loss function of M1, L1(θ) = R1 t * log(Pi tj ) and update the parameters of model M1 through the backpropagation method.
[0113] Figure 3 This is a flowchart of a method for information delivery shown in accordance with an exemplary embodiment of this specification. As Figure 3 shown, it includes steps 301 - 303:
[0114] Step 301: Obtain a first reinforcement learning model and a second reinforcement learning model trained for a target social platform.
[0115] Step 302: Select at least one actual delivery node from candidate delivery nodes in the information dissemination network through the first reinforcement learning model.
[0116] Step 303: Generate an actual delivery strategy for each actual delivery node through the second reinforcement learning model, where the actual delivery strategy is used to indicate generating actual release information content to be delivered for the actual delivery node.
[0117] Among them, Figure 3 the first reinforcement learning model and the second reinforcement learning model in Figure 1 can be obtained by the training method of
[0118] In an exemplary application scenario, assume that a brand party hopes to deliver advertisements to a target social platform, but is troubled by not knowing which cooperative accounts on the target social platform to cooperate with, and after selecting cooperative accounts, which delivery strategy to adopt for each cooperative account can achieve better information dissemination effects and more accurately reach the target users. The first reinforcement learning model and the second reinforcement learning model trained for the target social platform can be used to meet the above application requirements. By using the first reinforcement learning model to select appropriate target cooperative accounts from candidate cooperative accounts, and using the second reinforcement learning model to generate appropriate target information delivery strategies for each target cooperative account, the brand party can generate actual release information according to the target information delivery strategy. It can be seen that the trained first reinforcement learning model and second reinforcement learning model can be used to assist manual selection of actual delivery nodes and generation of actual delivery strategies. Since the first reinforcement learning model and the second reinforcement learning model are adjusted through feedback according to network propagation effects multiple times during the simulation stage, their decision results consider user behavior characteristics and network structure characteristics. Therefore, the delivery tasks generated by using them can ensure that information can reach the target users to the greatest extent, thereby improving the accuracy of information delivery.
[0119] Corresponding to the foregoing method embodiments, this specification also provides embodiments of a device and a terminal to which it is applied.
[0120] Figure 4 This is a schematic structural diagram of an electronic device shown in accordance with an exemplary embodiment of this specification. As Figure 4As shown in the figure, at the hardware level, the electronic device 400 includes a processor 402, an internal bus 404, a network interface 406, a memory 408, and a non-volatile memory 410. Of course, it may also include other hardware required for other services. One or more embodiments of this specification can be implemented in software. For example, the processor 402 reads the corresponding computer program from the non-volatile memory 410 into the memory 408 and then runs it. Of course, in addition to the software implementation method, one or more embodiments of this specification do not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic module, and can also be hardware or logic devices.
[0121] Figure 5 is a block diagram of a training device for a model shown according to an exemplary embodiment of this specification. As Figure 5 shown, this device can be applied to the electronic device 400 as Figure 4 shown to implement the technical solution of this specification. The device includes:
[0122] A simulation node selection module 502, configured to select at least one simulated placement node from candidate placement nodes in the information propagation network through a first reinforcement learning model; wherein, the information propagation network is constructed according to a target social platform, the nodes of the information propagation network include candidate placement nodes and forwarding nodes, and the edges of the information propagation network are used to represent the association relationship between nodes.
[0123] A simulation strategy generation module 504, configured to generate a simulation placement strategy for each simulated placement node through a second reinforcement learning model.
[0124] A simulation propagation module 506, configured to respectively generate simulation release information according to the simulation placement strategies of each simulated placement node, and simulate the forwarding process of the forwarding nodes based on the simulation release information placed by each simulated placement node to generate a simulation propagation result of the information propagation network;
[0125] A parameter update module 508, configured to update the parameters of the first reinforcement learning model and the second reinforcement learning model according to the number of forwarding times of the forwarding nodes in the simulation propagation result as a reward.
[0126] Optionally, the state space of the first reinforcement learning model includes a state representation of the information propagation network, and the action space of the first reinforcement learning model is to select at least one simulated placement node from the candidate placement nodes. The simulation node selection module 502 is specifically configured to input the state representation of the information propagation network into the first reinforcement learning model and obtain at least one simulated placement node output by the first reinforcement learning model.
[0127] Optionally, the simulation strategy generation module 504 is specifically configured to, for each simulated placement node, input the first prompt word corresponding to the simulated placement node into the second reinforcement learning model, and obtain the simulated placement strategy selected by the second reinforcement learning model according to the first prompt word; wherein, the first prompt word is used to instruct the second reinforcement learning model to select a simulated placement strategy from a preset set of placement strategies.
[0128] Optionally, the simulation propagation module 506 is specifically configured to, for each simulated placement node, input the second prompt word corresponding to the simulated placement node into the network simulation model, and obtain the simulated placement information to be placed at the simulated placement node generated by the network simulation model according to the second prompt word; wherein, the second prompt word is used to instruct the network simulation model to generate simulated placement information according to the simulated placement strategy of the simulated placement node.
[0129] Optionally, the simulation propagation module 506 is specifically configured to, for each forwarding node, input the third prompt word corresponding to the forwarding node into the network simulation model, and obtain the forwarding decision result output by the network simulation model according to the third prompt word; wherein, the third prompt word is used to instruct the network simulation model to simulate the forwarding decision-making process of the forwarding node to generate a forwarding decision result; the forwarding decision results of all forwarding nodes constitute the simulation propagation result, and the forwarding decision result of each forwarding node indicates whether the forwarding node forwards the simulated release information.
[0130] Optionally, the parameter update module 508 is specifically configured to, for the simulation propagation result of each time step, update the parameters of the first reinforcement learning model based on the increment of the total number of forwards in the current time step relative to the total number of forwards in the previous time step as a reward; and, for each simulated placement node, update the parameters of the second reinforcement learning model with the number of forwards with the simulated placement node as the source point as a reward.
[0131] Figure 6 is a block diagram of an information placement device shown in accordance with an exemplary embodiment of the present specification. As Figure 6 shown, the device can be applied to an electronic device 400 as shown in Figure 4 to implement the technical solution of the present specification. The device includes:
[0132] A model acquisition module 602, configured to acquire a first reinforcement learning model and a second reinforcement learning model trained for a target social platform, where the first reinforcement learning model and the second reinforcement learning model are obtained by Figure 1 the training method described above.
[0133] The actual node selection module 604 is configured to select at least one actual placement node from candidate placement nodes in the information dissemination network through the first reinforcement learning model.
[0134] The actual policy generation module 606 is configured to generate an actual placement policy for each actual placement node through the second reinforcement learning model, where the actual placement policy is used to indicate the actual published information content to be placed for the actual placement node.
[0135] For the implementation processes of the functions and roles of each module in the above device, refer to the implementation processes of the corresponding steps in the above method for details, which will not be elaborated here.
[0136] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described above is only illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution in this specification. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0137] This specification also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of any one of the foregoing model training methods or information placement methods provided in this application are implemented.
[0138] Specifically, computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices (such as EPROM, EEPROM, and flash memory devices), magnetic disks (such as internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks.
Claims
1. A training method for a model, characterized in that, The method includes: Selecting at least one simulated placement node from candidate placement nodes in the information dissemination network through a first reinforcement learning model; wherein, the information dissemination network is constructed based on a target social platform, nodes of the information dissemination network include candidate placement nodes and forwarding nodes, and edges of the information dissemination network are used to represent the association relationship between nodes; Generating a simulated placement strategy for each simulated placement node through a second reinforcement learning model; Generating simulated release information respectively according to the simulated placement strategies of each simulated placement node, and simulating the forwarding process of the forwarding nodes based on the simulated release information placed by each simulated placement node to generate a simulated dissemination result of the information dissemination network; Updating parameters of the first reinforcement learning model and the second reinforcement learning model according to the number of forwarding times of the forwarding nodes in the simulated dissemination result as a reward.
2. The method according to claim 1, wherein The state space of the first reinforcement learning model includes a state representation of the information dissemination network, and the action space of the first reinforcement learning model is to select at least one simulated placement node from the candidate placement nodes; The step of selecting at least one simulated placement node from candidate placement nodes in the information dissemination network through the first reinforcement learning model includes: Inputting the state representation of the information dissemination network into the first reinforcement learning model, and obtaining at least one simulated placement node output by the first reinforcement learning model.
3. The method according to claim 1, wherein The step of generating a simulated placement strategy for each simulated placement node through the second reinforcement learning model includes: For each simulated placement node, inputting a first prompt word corresponding to the simulated placement node into the second reinforcement learning model, and obtaining a simulated placement strategy selected by the second reinforcement learning model according to the first prompt word; wherein, the first prompt word is used to instruct the second reinforcement learning model to select a simulated placement strategy from a preset set of placement strategies.
4. The method according to claim 1, wherein The step of generating simulated release information respectively according to the simulated placement strategies of each simulated placement node includes: For each simulated placement node, inputting a second prompt word corresponding to the simulated placement node into the network simulation model, and obtaining simulated placement information to be placed by the simulated placement node generated by the network simulation model according to the second prompt word; wherein, the second prompt word is used to instruct the network simulation model to generate simulated placement information according to the simulated placement strategy of the simulated placement node.
5. The method according to claim 1, wherein The step of simulating the forwarding process of the forwarding nodes based on the simulated placement information placed by each simulated placement node to generate a simulated dissemination result of the information dissemination network includes: For each forwarding node, inputting a third prompt word corresponding to the forwarding node into the network simulation model, and obtaining a forwarding decision result output by the network simulation model according to the third prompt word; wherein, the third prompt word is used to instruct the network simulation model to simulate the forwarding decision process of the forwarding node to generate a forwarding decision result; the forwarding decision results of all forwarding nodes constitute the simulated dissemination result, and the forwarding decision result of each forwarding node indicates whether the forwarding node forwards the simulated release information.
6. The method according to claim 1, characterized in that, Updating the parameters of the first reinforcement learning model and the second reinforcement learning model with the number of forwarding times of the forwarding nodes in the simulated propagation result as a reward includes: For the simulated propagation result at each time step, updating the parameters of the first reinforcement learning model with the increment of the total number of forwarding times at the current time step relative to the total number of forwarding times at the previous time step as a reward; and for each simulated placement node, updating the parameters of the second reinforcement learning model with the number of forwarding times with the simulated placement node as the source as a reward.
7. A method for information delivery, characterized in that, The method includes: Obtaining a first reinforcement learning model and a second reinforcement learning model trained for a target social platform, where the first reinforcement learning model and the second reinforcement learning model are obtained by the training method described in claim 1; Selecting at least one actual placement node from the candidate placement nodes in the information propagation network through the first reinforcement learning model; Generating an actual placement strategy for each actual placement node through the second reinforcement learning model, where the actual placement strategy is used to indicate the actual release information content to be released for the actual placement node.
8. A training device for a model, characterized in that, The device includes: A simulated node selection module, configured to select at least one simulated placement node from the candidate placement nodes in the information propagation network through the first reinforcement learning model; where the information propagation network is constructed according to the target social platform, the nodes of the information propagation network include candidate placement nodes and forwarding nodes, and the edges of the information propagation network are used to represent the association relationship between the nodes; A simulated strategy generation module, configured to generate a simulated placement strategy for each simulated placement node through the second reinforcement learning model; A simulated propagation module, configured to respectively generate simulated release information according to the simulated placement strategies of the respective simulated placement nodes, and simulate the forwarding process of the forwarding nodes based on the simulated release information placed by the respective simulated placement nodes to generate the simulated propagation result of the information propagation network; A parameter update module, configured to update the parameters of the first reinforcement learning model and the second reinforcement learning model with the number of forwarding times of the forwarding nodes in the simulated propagation result as a reward.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the steps of the method described in any one of claims 1-7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the steps of the method described in any one of claims 1-7 are implemented.
Citation Information
Cited By
Content delivery method and device of social platform, storage medium and equipment
CN121412460A