Information guidance strategy generation method, device and equipment in multi-body interaction scene
By constructing a social network model and a hierarchical cognitive mechanism, and combining it with a policy pool response evangelism framework, a multi-agent policy iteration framework is generated. This solves the problem of poor information guidance in multi-agent interaction scenarios, achieves accurate matching of information guidance strategies and resource optimization, and adapts to the needs of multiple scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-17
- Publication Date
- 2026-04-14
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies lack a strategy pool construction and iteration mechanism for multi-subject collaborative adaptation in multi-subject interaction scenarios. Node selection is an exclusive design that does not consider node reuse and cost trade-offs between subjects, making it difficult to adapt to scenarios without historical data support. Furthermore, the information guidance process fails to achieve accurate adaptation and dynamic adjustment, resulting in poor guidance effects.
We construct a basic model of social networks, set up a viewpoint update mechanism based on hierarchical cognition, build a multi-agent policy iteration framework through policy pool response evangelism, generate a policy interaction matrix, solve for equilibrium policies and iteratively optimize them, allow agents to flexibly select nodes, optimize resource allocation through a cost multiplier mechanism, and achieve precise matching and dynamic adjustment of information guidance policies.
It enhances the robustness and relevance of information guidance strategies, enabling flexible responses to complex scenarios involving dynamic adjustments by multiple stakeholders, optimizing resource utilization, adapting to information guidance needs across multiple scenarios, breaking free from reliance on historical data, and accommodating interactions under multi-stakeholder collaboration and budget constraints.
Smart Images

Figure CN121860164A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of reinforcement learning technology, and in particular to a method, apparatus, and device for generating information guidance strategies in multi-agent interaction scenarios. Background Technology
[0002] With the rapid development of basic social network models, information dissemination in networks is characterized by high speed, wide range, and multi-party participation. How to generate efficient information guidance strategies in multi-party interaction scenarios has become an important research direction in the field of basic social network model analysis.
[0003] Currently, related technologies mostly focus on optimizing information dissemination for a single objective or node selection in specific scenarios, and have not yet formed a complete technical solution adapted to the needs of multi-agent collaboration and dynamic interaction. Existing decision-making methods for basic social network models based on reinforcement learning lack the construction and iteration mechanism for policy pools adapted to multi-agent collaboration, and node selection is mostly an exclusive design, failing to consider the actual needs of node reuse and cost trade-offs among agents, while ignoring the impact of dynamic updates of viewpoints on the guidance effect. Methods relying on user historical behavior sequences are limited by data availability and are difficult to adapt to scenarios without historical data support, and the iteration process does not incorporate policy interaction logic, making it unable to flexibly respond to policy adjustments among multiple agents. While technologies based on greedy algorithms or Monte Carlo simulations can optimize node selection through community structure or marginal gains, they lack the ability of agents to learn autonomously and optimize policies, making it difficult to adapt to complex scenarios of multi-agent dynamic interaction.
[0004] Furthermore, existing technologies often focus on node state transitions or the propagation effect of a single objective, failing to construct a mechanism for the evolution of viewpoints in multi-party participation, and thus unable to achieve precise adaptation and dynamic adjustment of user viewpoints during information guidance. These issues result in poor information guidance effectiveness of existing methods in multi-party interaction scenarios, making it difficult to balance guidance efficiency, resource constraints, and the needs of multiple parties. Summary of the Invention
[0005] Therefore, it is necessary to provide a method, apparatus, and device for generating information guidance strategies in multi-agent interaction scenarios that can integrate multi-agent collaborative strategies, dynamic viewpoint updates, and resource optimization allocation, in order to address the aforementioned technical problems.
[0006] A method for generating information guidance strategies in a multi-subject interaction scenario, the method comprising:
[0007] Step 1: Construct a basic model of a social network, identify the multiple entities involved in guiding information, and set up a viewpoint update mechanism based on hierarchical cognition; Step 2: Build a multi-agent policy iteration framework based on policy pool response arbitrage, and then initialize the policy pool of each executing agent; Step 3: Based on the social network basic model and the policies in the policy pool, generate a policy interaction matrix through multiple rounds of simulation; solve the equilibrium policy of each executing entity under the policy interaction matrix, train the optimal response policy for the equilibrium policy, and add the optimal response policy to the policy pool. Step 4: Iterate through Step 3 until the preset number of iterations is reached to obtain a set of optimized strategies that adapt to multi-subject interactions; Step 5: Each executing entity selects a seed node from the social network basic model according to the set of optimization strategies to deploy the information propagation starting point; wherein, the subsequent executing entity obtains the seed node already selected by the previous executing entity according to the requirements and by a set cost multiple; Step 6: After all the execution entities have exhausted their budgets, complete the evolution of the viewpoints of all nodes in the basic social network model according to the viewpoint update mechanism, and output the information guidance strategy that allows the first execution entity to obtain the maximum support.
[0008] On the other hand, an information guidance strategy generation device for multi-subject interaction scenarios is also provided, including: The social network modeling and opinion mechanism configuration module is used to build a basic model of a social network, identify multiple executive entities involved in information guidance, and set up an opinion update mechanism based on hierarchical cognition. The multi-agent policy framework building module is used to build a multi-agent policy iteration framework based on policy pool response arbitrage, and then initialize the policy pool of each executing agent. The optimal response strategy training module is used to generate a strategy interaction matrix through multiple rounds of simulation based on the social network basic model and the strategies in the strategy pool; solve the equilibrium strategy of each executing entity under the strategy interaction matrix; train the optimal response strategy for the equilibrium strategy; and supplement the optimal response strategy to the strategy pool. The strategy iteration and optimization module is used to iteratively execute the process in the best response strategy training module until the preset number of iterations is reached, and obtain an optimized strategy set that adapts to multi-subject interaction. The seed node selection module is used by each execution entity to select a seed node from the social network basic model according to the set of optimization strategies to deploy the information propagation starting point; wherein, the subsequent execution entity obtains the seed node already selected by the previous execution entity according to the requirements and by a set cost multiple; The optimal strategy output module is used to complete the evolution of the viewpoints of all nodes in the basic model of the social network according to the viewpoint update mechanism after all the execution entities have exhausted their budgets, and output the information guidance strategy that the first execution entity obtains the maximum support.
[0009] On another front, a computer device is also provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the above-mentioned information guidance strategy generation method in a multi-subject interaction scenario.
[0010] Compared with existing technologies, the information guidance strategy generation method, apparatus, and device for multi-subject interaction scenarios provided by this invention have the following beneficial effects: 1. This method builds a multi-agent policy iteration framework based on Policy Pool Response Oracle (PSRO). Through a closed-loop mechanism of initializing multi-agent policy pools, generating policy interaction matrices through multiple rounds of simulation, solving for equilibrium policies, and iteratively supplementing the best response policies, it can accurately capture the policy interaction relationships between multiple agents, making the generated guidance policies more robust and flexibly able to cope with complex scenarios of dynamic adjustment of multiple agents.
[0011] 2. A viewpoint update mechanism based on hierarchical cognition was constructed, which can capture the dynamic changes of node viewpoints in real time, making the information guidance process accurately match the evolution of user viewpoints and improving the pertinence and effectiveness of guidance strategies.
[0012] 3. This method designs a mechanism in which "the subsequent executing entity can obtain the seed node selected by the preceding executing entity by setting a cost multiplier". This allows the entity to make flexible decisions between independently selecting new nodes and reusing existing nodes according to actual needs, thereby optimizing the allocation of node resources. The cost multiplier constraint balances the reuse requirements and cost trade-offs, improving the utilization rate of limited resources. It is especially suitable for multi-entity interaction or collaboration scenarios under budget constraints.
[0013] 4. This method generates effective guidance strategies without relying on historical user behavior data by constructing a basic social network model, initializing and iteratively optimizing the strategy pool. This eliminates the dependence on historical data, broadens the applicable scenarios of the technology, and adapts to information guidance needs in multiple scenarios. Attached Figure Description
[0014] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention, and those skilled in the art can obtain other related drawings based on these drawings without creative effort.
[0015] Figure 1 This is a flowchart illustrating the information guidance strategy generation method in a multi-subject interaction scenario in Example 1. Figure 2 This is a flowchart illustrating the policy pool response argument algorithm in Example 1; Figure 3 This is a structural block diagram of the information guidance strategy generation device in the multi-subject interaction scenario of Example 2; Figure 4 This is a diagram of the internal structure of the computer device in Example 3.
[0016] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0018] It should be noted that in this invention, the use of terms such as "first," "second," etc., is for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0019] It is understood that the technical solutions of the various embodiments of the present invention can be combined with each other, but only if they are based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.
[0020] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0021] Example 1 like Figure 1 As shown, this embodiment provides a method for generating information guidance strategies in a multi-subject interaction scenario, including the following steps: Step 1: Construct a basic model of a social network, identify the multiple entities involved in guiding information, and set up a viewpoint update mechanism based on hierarchical cognition.
[0022] Step 2: Build a multi-agent policy iteration framework based on policy pool response ideologies, and then initialize the policy pools of each executing agent.
[0023] Step 3: Based on the social network basic model and the policies in the policy pool, generate the policy interaction matrix through multiple rounds of simulation; solve the equilibrium policy of each executing entity under the policy interaction matrix, train the best response policy for the equilibrium policy, and supplement the best response policy to the policy pool.
[0024] Step 4: Iterate through Step 3 until the preset number of iterations is reached to obtain a set of optimized strategies that adapt to multi-subject interactions.
[0025] Step 5: Each executing entity selects a seed node from the basic social network model based on the set of optimization strategies to deploy the information propagation starting point; among them, the subsequent executing entity obtains the seed node already selected by the preceding executing entity according to the set cost multiple based on its needs.
[0026] Step 6: After all the execution entities have exhausted their budgets, complete the evolution of the viewpoints of all nodes in the basic social network model according to the viewpoint update mechanism, and output the information guidance strategy that allows the first execution entity to obtain the maximum support.
[0027] The information guidance strategy generation method for multi-subject interaction scenarios provided by this invention effectively solves the problems of poor guidance effect, low resource utilization and poor adaptability of existing technologies in multi-subject interaction scenarios by integrating multi-subject collaborative strategies, dynamic viewpoint updates and resource optimization allocation.
[0028] In the specific implementation of step 1, the basic model of social networks is based on social graphs. It means that among them Represents a set of nodes. Denotes the set of edges. This represents the influence weight between nodes.
[0029] The entities involved in information guidance include the entities that first implement the guidance. and subsequent execution entity Among them, the first subject to execution The core entity that initiates information guidance, and the subsequent implementing entity. Other entities participating in the interaction; each entity is subject to a preset budget. Constraints. First, execute the main body. and subsequent execution entity Requires a pre-set budget From the node set Select a seed node set At the same time, for each node Configuration cost function ,Right now Indicates the selected node The required budget dictates that, initially, each implementing entity must meet budget constraints. .
[0030] Information dissemination can be divided into three stages: the initial guidance period, the diffusion period, and the plateau period. This method focuses on the initial guidance period, that is, the generation of guidance strategies in the early stage of information evolution. At this time, the other party's information dissemination has not yet reached a certain scale, and the guidance cost is lower and the effect is better.
[0031] The viewpoint update mechanism based on hierarchical cognition is specifically designed as follows: Set the range of opinion values as follows , where the interval Indicates support for the implementing entity The smaller the value, the stronger the support for the view; range Indicates support for the entity that will execute first. The higher the value of the viewpoint, the stronger the support. A viewpoint value of 0 indicates that the node is in a neutral state, meaning it does not support any viewpoint. Since this state may be due to disagreement with any party, not being influenced by any party, or the node being in an information-closed state during information evolution, this invention does not include a viewpoint value of 0 for any party.
[0032] The nodes in the basic model of social networks are divided into hierarchical levels. Seed nodes selected by the executing entity are determined to be high-level nodes, while non-seed nodes are basic-level nodes.
[0033] For high-level nodes, the average viewpoint of social neighbors that are one hop below their own thinking level can be obtained and integrated into their own viewpoint for updating. That is, high-level nodes first collect the viewpoint values of all social neighbors within their two-hop range and calculate the average viewpoint value; then, they perform weighted fusion processing on their current viewpoint value and the average viewpoint value to adjust their own viewpoint value and enhance their guiding role.
[0034] For basic-level nodes, the viewpoint is updated based on their current viewpoint value and the viewpoint influence of adjacent nodes, in accordance with the base-level user update rules of the CHOD model.
[0035] This step, through a hierarchical viewpoint update logic, not only strengthens the guiding role of seed nodes but also aligns with the viewpoint evolution patterns of ordinary users, making information guidance more precise.
[0036] In the specific implementation of step 2, the aim is to construct a multi-agent policy iteration framework adapted to dynamic multi-agent interactions, thereby achieving optimized generation of information-guided strategies. During the first round of decision-making, the first agent cannot know the strategy of the subsequent agent; this is a multi-agent interaction environment with incomplete information. Therefore, it is necessary to first predict the possible strategies of the subsequent agent for derivation. Specifically, in a multi-agent interaction scenario, the information-guided strategy generation process is as follows: assuming the first agent… To take the initiative, after resetting the initial environment, execute the main action first. The observations form the basis of the entire social network model. Its action is to select a node from the network. (Can be marked as a friendly node); then transition to the main execution body. The rounds, the subsequent implementing entities Observable first execution subject The selection result is based on the current state, deciding whether to proceed. Double the cost to acquire the first execution subject Selected Nodes Alternatively, select a node from the remaining nodes at the original price; the final execution entity... The selected nodes can be marked accordingly. It's worth noting that if a node is initially... Select, currently selected Acquiring, or conversely, can be distinguished by different markers.
[0037] In this embodiment, a multi-agent policy iteration framework adapted to dynamic multi-agent interactions is constructed based on Policy-space Response Oracles (PSRO). Its difference from Independent Reinforcement Learning (InRL) lies in the fact that in InRL, each agent considers other agents as part of the environment, and the actions of other agents cause irregular changes in the environment's state transition matrix, leading to non-stationarity. PSRO, on the other hand, is a unified framework for solving policy exploration and equilibrium problems in multi-agent policy interaction scenarios (especially policy interactions under imperfect information). By combining Empirical Game-Theoretic Analysis (EGTA) with deep reinforcement learning, an iterative policy pool expansion and equilibrium solution mechanism is constructed. Figure 2 As shown, the policy pool response idempotency algorithm mainly comprises a policy pool maintenance module, a meta-policy interaction solution module, and an optimal response calculation module. These are the three core modules. Specifically, the policy pool maintenance module manages the primary policy pool that executes first. With the post-execution subject strategy pool After selecting a strategy from two strategy pools, the simulation begins; the meta-policy interactive solver module evaluates the benefits based on the simulation results. The policy distribution is obtained through the meta-solver M; the optimal response calculation module trains the oracle O based on the payoff information, generates the optimal response policy and updates it to the policy pool, realizing the iterative expansion of the policy.
[0038] In this step, the policy pool is first constructed and initialized. The policy pool starts with different types of basic policies, providing a diverse foundation for subsequent iterative optimization. This invention can use algorithms including, but not limited to, randomized policies, greedy policies as benchmarks, and heuristics as policies.
[0039] Because maximizing the effectiveness of information guidance is a strategic interaction process with a specific order between the two parties, the current network state and selection history need to be considered when formulating the strategy pool. Based on this, the main body should be executed first. Initialize the policy pool Different strategies can be considered, such as a random strategy or a greedy strategy that prioritizes high-value, low-cost nodes. The policy pool of the execution entity is executed after initialization. Similarly, one could consider always seizing or never seizing, that is... In the formula, This indicates that the first node selected by the main body will be executed first. One strategy, ; This indicates the strategy of the subsequent executing entity preempting the seed node of the preceding executing entity; This indicates that the executing entity will only select seed nodes from the remaining nodes.
[0040] This step addresses the problem of policy exploration and equilibrium solution in multi-agent interactions using the PSRO framework. The initial policy pool covers multiple types of algorithm policies, providing a rich and diverse starting point for subsequent iterative optimization and ensuring the adaptability of policies to multi-agent dynamic interaction scenarios.
[0041] In the specific implementation of step 3, based on the social network basic model and the policies in the policy pool, through... In a round of simulation, the payoff values corresponding to any two policy combinations are calculated, and a policy interaction matrix is constructed, represented as follows: ; In the formula, Represents the payoff matrix; This represents the activation function, used to calculate the range of influence that the seed node can generate.
[0042] In maximizing the effectiveness of information guidance, a conservative optimization strategy is often adopted, that is, finding a strategy that guarantees maximum benefit even in the worst-case scenario. Furthermore, because the policy pools of both parties are relatively small in the early stages, directly solving for the Nash equilibrium may lead to overfitting. Therefore, based on the policy interaction matrix, a maximum-minimum benefit approach is used to solve for the equilibrium strategy of each executor, expressed as: ; In the formula, Indicates the first Always prioritize executing the balancing strategy from the pool of available strategies available to the main body. Indicates the first After a certain time, the balancing strategy from the optional strategy pool of the main body is executed; For the first Round strategy pool The corresponding strategy distribution set; The strategy sampling is performed first by the subject. For policy sampling of subsequent execution entities, This is the expected computation function. This solution method avoids the overfitting problem caused by directly solving for the equilibrium when the policy pool is small in the early stages, and the results are more stable.
[0043] In maximizing the information guidance effect, training the optimal response (i.e., Oracles) means finding an approximately optimal strategy for the current balancing strategy. Therefore, training the optimal response strategy for the balancing strategy includes: training the optimal response strategy for the first executing subject, expressed as: ; The expression for training the optimal response strategy for the subsequent execution subject is: ; In the formula, Indicates the first The best response strategy for the first executing entity; Indicates the first Equilibrium strategies for post-round execution entities; This indicates that the subject will be executed first from the policy pool. Candidate strategies selected from the pool; This indicates the policy pool that executes first. Indicates the activation function; Indicates the first The optimal response strategy for the post-execution entity; Represents a node in the basic model of a social network; Indicates an indicator function; Indicates the cost multiple; Indicates the selected node The budget required; The expected computation function; Indicates the first The strategy distribution corresponding to the equilibrium strategy of the first-to-first-execute entity; Indicates in On this basis, no preemptive action will be taken.
[0044] Furthermore, during the training process of the primary optimal response policy, the gradient expression for updating the policy network parameters is: ; In the formula, Describe the policy objective function Policy network parameters The gradient; Indicates the strategy Selected node Take the expected value; Representation Strategy Select node log probability of parameters The gradient; Indicates the subject that executes first. Value function.
[0045] In this step, through multiple rounds of simulation and policy interaction, the policy can be accurately adapted to the interaction relationship between multiple subjects, and the supplementation of the best response policy continuously optimizes the diversity and effectiveness of the policy pool.
[0046] In the specific implementation of step 4, each round generates a new policy interaction matrix based on the policy pool of the previous round, solves the equilibrium policy, trains the best response policy, and adds the new best response policy to the original policy pool to realize the dynamic expansion of the policy pool.
[0047] The iteration termination condition is reaching the preset number of iterations. At this point, the strategies in the strategy pool have undergone multiple rounds of strategy interaction optimization, forming an optimized strategy set that adapts to dynamic interactions among multiple subjects.
[0048] This step uses an iterative mechanism to continuously adapt the strategy to the dynamic changes in multi-agent interactions, avoiding the limitations of a single strategy and improving the robustness of the strategy.
[0049] In the specific implementation of step 5, the goal of information guidance is to achieve the first execution subject Maximize profit, that is ,in, This indicates the set of seed nodes selected by the execution entity based on the budget. This represents the benefit of the strategy of the first executing entity; then, through different information propagation models, the evolution occurs, and the benefit is the amount that supports the first executing entity in the steady state. Number of users expressing an opinion.
[0050] In the initial stage, the node view value held by the executing entity is first executed. The node view value held by the subsequent executing entity is Both parties will deploy their own viewpoints based on the optimized strategy set.
[0051] Seed node selection must follow these rules: Rule 1: Each node in a social network corresponds to a cost; the entity that executes the task first has the budget it can spend as needed. Acquire a node as your own seed node and promote the spread of your own viewpoint to that node.
[0052] Rule 2: The implementing entity can spend the budget as needed. Seed nodes can be selected from the remaining nodes after the initial selection of the execution subject, or by paying a cost multiplier. Obtain the selected seed node of the first execution subject to achieve flexible reuse of node resources.
[0053] Rule 3: When a seed node initiates view evolution, it only holds the view of its own executing entity. That is, once a node is acquired by an executing entity, it will directly hold the view of that entity. Rule 4: The first executing entity and the second executing entity take turns selecting the seed node until both entities run out of budget. If one entity runs out of budget first, but the other entity still has budget, then the entity enters the independent selection phase until its budget is completely used up.
[0054] Among them, the first-running entity is in the first round of sequential selection and can adopt three core strategies: prioritize high-value, low-cost nodes to reduce the possibility of them being acquired by subsequent entities; predict the nodes that subsequent entities may choose and occupy them in advance to increase their costs; and adopt a "pre-selection" strategy to select nodes that subsequent entities need but whose value and cost are higher than average.
[0055] The subsequent executor, being in a later round of sequential selection, can fully observe and consider the selection information of the earlier executors, and weigh whether it is necessary to pay a cost multiplier. Retrieve the selected seed node from the main body of execution, or from the remaining node set. Select a set of high-value nodes.
[0056] Therefore, in the first round, the main body is executed first. The policy pool is The cost function for each node is: , Then the main body will be executed first. The first round of budget constraints is: ; At this time, the executing entity From the remaining node set Choose the original price, where \ represents finding the difference set. You can also use The cost of persuading individuals in the first-mover strategy set is such that when a node needs to be selected from the set of seed nodes already selected by the first-mover, the corresponding constraint formula is: ; In the formula, This indicates the pre-set budget of each implementing entity; Represents the set of remaining nodes; Indicates the cost multiple.
[0057] In the second round, the main body will be executed first. Being able to see the opponent's information, therefore, the main body should be executed first. The returns are: ; In the formula, Indicates the executing entity The set of seed nodes; Indicates the subject to be executed first. The set of seed nodes; This indicates that the entity holding the first execution right after co-evolution... Users with opinions.
[0058] Post-execution entity The returns are: ; In the formula, Indicates the entity that holds and executes the shares after co-evolution. Users with opinions; Indicates the executing entity The first implementing entity obtained after the decision The set of seed nodes.
[0059] Therefore, for the subject that executes first Its optimization objective is: ; In the formula, This indicates that the first-round revenue of the executing entity will be used. This indicates that the entity holding the first execution right after co-evolution... Users with opinions; This indicates that the first round of strategy pool for the main body will be executed first; The first round of strategy pool will be executed after the main body is indicated; This indicates the pre-set budget of each implementing entity.
[0060] Post-execution entity The optimization objective is: ; When it is necessary to select a node from the set of seed nodes already selected by the executing entity, the cost paid is... The corresponding constraint formula is: ; In the formula, This indicates that the second round of revenue will be distributed to the entity that will receive the benefits first. Indicates the entity that holds and executes the shares after co-evolution. Users with opinions; This represents the set of seed nodes selected by the main body in the second round; The second round of strategy pool will be executed after the main body is notified; Represents the set of remaining nodes in the second round; This indicates that the second round of strategy pool will be executed first. This indicates the pre-set budget of each implementing entity; Represents the set of remaining nodes; Indicates the cost multiple.
[0061] This step breaks away from the traditional exclusive design of node selection. It achieves node reuse and cost trade-off through a cost multiplier mechanism, thereby improving the utilization rate of limited resources. At the same time, the first-mover strategy of the first-mover subject and the flexible decision-making of the subsequent-mover subject are both in line with the needs of actual interaction scenarios.
[0062] In the specific implementation of step 6, after all implementing entities deploy seed nodes, the nodes in the social network dynamically evolve according to the hierarchical cognitive viewpoint update mechanism set in step 1: higher-level nodes strengthen their guiding role by merging the average viewpoints of their two-hop neighbors, while basic-level nodes adjust their own viewpoints under the influence of neighboring nodes. After the evolution terminates, the number of nodes supporting each entity is counted, and an information guidance strategy that maximizes the support of the first implementing entity is output. The information guidance strategy includes the seed node selection scheme, the strategy execution order, etc.
[0063] In this step, the hierarchical guidance of seed nodes can effectively improve the efficiency of opinion dissemination. The seed nodes selected by the first executing entity through the optimization strategy set are more likely to gain an information guidance advantage in the interaction with the subsequent executing entity.
[0064] It should be understood that, although this embodiment Figure 1 The steps are shown sequentially as indicated by the arrows, but they are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order in which these steps are performed; they can be executed in other orders. Figure 1 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.
[0065] Example 2 Based on the information guidance strategy generation method in a multi-subject interaction scenario in Embodiment 1, this embodiment discloses an information guidance strategy generation device in a multi-subject interaction scenario, such as... Figure 3 As shown, the information guidance strategy generation device in a multi-agent interaction scenario includes: a social network modeling and opinion mechanism configuration module 401, a multi-agent strategy framework construction module 402, an optimal response strategy training module 403, a strategy iteration optimization module 404, a seed node selection module 405, and an optimal strategy output module 406, wherein: The social network modeling and opinion mechanism configuration module 401 is used to build a basic model of a social network, identify multiple executive entities involved in information guidance, and set an opinion update mechanism based on hierarchical cognition. The multi-agent policy framework building module 402 is used to build a multi-agent policy iteration framework based on policy pool response arbitrage, and then initialize the policy pool of each executing agent; The optimal response strategy training module 403 is used to generate a strategy interaction matrix through multiple rounds of simulation based on the social network basic model and the strategies in the strategy pool; solve the equilibrium strategy of each executing entity under the strategy interaction matrix; train the optimal response strategy for the equilibrium strategy; and supplement the optimal response strategy to the strategy pool. The strategy iteration optimization module 404 is used to iteratively execute the process in the best response strategy training module 403 until the preset number of iterations is reached, so as to obtain an optimized strategy set that adapts to multi-subject interaction. The seed node selection module 405 is used by each execution entity to select a seed node from the social network basic model according to the set of optimization strategies to deploy the information propagation starting point; wherein, the subsequent execution entity obtains the seed node already selected by the previous execution entity according to the requirements and by a set cost multiple; The optimal strategy output module 406 is used to complete the view evolution of all nodes in the basic model of the social network according to the view update mechanism after all the execution subjects have exhausted their budgets, and output the information guidance strategy that the first execution subject obtains the maximum support.
[0066] In this embodiment, the specific working process and working principle of the social network modeling and opinion mechanism configuration module 401, the multi-agent policy framework construction module 402, the optimal response policy training module 403, the policy iteration optimization module 404, the seed node selection module 405, and the optimal policy output module 406 are the same as those in Embodiment 1, and therefore will not be described again in this embodiment. Each unit module can be implemented entirely or partially through software, hardware, or a combination thereof. Each unit module can be embedded in or independent of the processor in a computer device in hardware form, or it can be stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each of the above unit modules.
[0067] Example 3 like Figure 4 The diagram illustrates a computer device disclosed in this embodiment, including a transmitter, a receiver, a memory, and a processor. The transmitter is used to send instructions and data, the receiver is used to receive instructions and data, the memory is used to store computer execution instructions, and the processor is used to execute the computer execution instructions stored in the memory to implement the method in Embodiment 1 above.
[0068] It is important to note that the aforementioned memory can be either standalone or integrated with the processor. When the memory is set up independently, the terminal device also includes a bus for connecting the memory and the processor.
[0069] Example 4 This embodiment discloses a computer-readable storage medium storing computer-executable instructions. When a processor executes the computer-executable instructions, it implements the method in Embodiment 1 above.
[0070] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0071] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0072] The embodiments described above are merely examples of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention.
Claims
1. A method for generating information guidance strategies in a multi-subject interaction scenario, characterized in that, The method includes: Step 1: Construct a basic model of a social network, identify the multiple entities involved in guiding information, and set up a viewpoint update mechanism based on hierarchical cognition; Step 2: Build a multi-agent policy iteration framework based on policy pool response arbitrage, and then initialize the policy pool of each executing agent; Step 3: Based on the social network basic model and the policies in the policy pool, generate a policy interaction matrix through multiple rounds of simulation; solve the equilibrium policy of each executing entity under the policy interaction matrix, train the optimal response policy for the equilibrium policy, and add the optimal response policy to the policy pool. Step 4: Iterate through Step 3 until the preset number of iterations is reached to obtain a set of optimized strategies that adapt to multi-subject interactions. Step 5: Each executing entity selects a seed node from the social network basic model according to the set of optimization strategies to deploy the information propagation starting point; wherein, the subsequent executing entity obtains the seed node already selected by the previous executing entity according to the requirements and by a set cost multiple; Step 6: After all the execution entities have exhausted their budgets, complete the evolution of the viewpoints of all nodes in the basic social network model according to the viewpoint update mechanism, and output the information guidance strategy that allows the first execution entity to obtain the maximum support.
2. The method for generating information guidance strategies in a multi-subject interaction scenario according to claim 1, characterized in that, In step 1, a viewpoint update mechanism based on hierarchical cognition is established, including: Set the range of opinion values as follows , where the interval This indicates support for the viewpoint of the implementing entity, within a certain range. This indicates a viewpoint that supports the subject that executes first; a viewpoint value of 0 indicates that the node is in a neutral state. The nodes in the basic model of social networks are hierarchically divided, and the seed nodes selected by the executing entity are determined to be high-level nodes, while non-seed nodes are basic-level nodes. For high-level nodes, first collect the opinion values of all social neighbors within their two-hop range and calculate the average opinion value; then, perform a weighted fusion process between their current opinion value and the average opinion value to adjust their own opinion value. For basic-level nodes, the viewpoint is updated based on their current viewpoint value and the viewpoint influence of adjacent nodes, in accordance with the base-level user update rules of the CHOD model.
3. The method for generating information guidance strategies in a multi-subject interaction scenario according to claim 1, characterized in that, In step 3, based on the social network basic model and the policies in the policy pool, a policy interaction matrix is generated through multiple rounds of simulation, including: Based on the aforementioned social network infrastructure model and the policies in the policy pool, through In a round of simulation, the payoff values corresponding to any two policy combinations are calculated, and a policy interaction matrix is constructed, represented as follows: ; In the formula, Represents the payoff matrix; Indicates the activation function; This indicates that the first node selected by the main body will be executed first. One strategy, ; This indicates the strategy of the subsequent executing entity preempting the seed node of the preceding executing entity; This indicates that the executing entity will only select seed nodes from the remaining nodes.
4. The method for generating information guidance strategies in multi-subject interaction scenarios according to any one of claims 1 to 3, characterized in that, In step 3, the equilibrium strategy of each executing entity under the strategy interaction matrix is solved, and the expression is: ; In the formula, Indicates the first Always prioritize executing the balancing strategy from the pool of available strategies available to the main body. Indicates the first After a certain time, the balancing strategy from the optional strategy pool of the main body is executed; For the first Round strategy pool The corresponding strategy distribution set; The strategy sampling is performed first by the subject. For policy sampling of subsequent execution entities, This is the function to be computed.
5. The method for generating information guidance strategies in multi-subject interaction scenarios according to any one of claims 1 to 3, characterized in that, Step 3, training the optimal response strategy for the equilibrium strategy, includes: training the optimal response strategy for the first executing subject, expressed as: ; The expression for the optimal response strategy trained by the subsequent execution entity is: ; In the formula, Indicates the first The best response strategy for the first executing entity; Indicates the first Equilibrium strategies for post-round execution entities; This indicates that the subject will be executed first from the policy pool. Candidate strategies selected from the pool; This indicates the policy pool that executes first. Indicates the activation function; Indicates the first The optimal response strategy for the post-execution entity; Represents a node in the basic model of a social network; Indicates an indicator function; Indicates the cost multiple; Indicates the selected node The budget required; The expected computation function; Indicates the first The strategy distribution corresponding to the equilibrium strategy of the first-to-first-execute entity; Indicates in On this basis, no preemptive action will be taken.
6. The method for generating information guidance strategies in a multi-subject interaction scenario according to claim 5, characterized in that, In step 3, during the initial training of the optimal response policy, the gradient expression for updating the policy network parameters is as follows: ; In the formula, Describe the policy objective function Policy network parameters The gradient; Indicates the strategy Selected node Take the expected value; Representation Strategy Select node log probability of parameters The gradient; Indicates the subject that executes first. Value function.
7. The method for generating information guidance strategies in a multi-subject interaction scenario according to claim 5, characterized in that, In step 5, when each executing entity selects a seed node from the social network basic model to deploy the information propagation starting point according to the set of optimization strategies, the first executing entity is in the first round of sequential selection, and the optimization objective is: ; In the formula, This indicates that the first-round revenue of the executing entity will be used. This indicates that the entity holding the first execution right after co-evolution... Users with opinions; This indicates that the first round of strategy pool for the main body will be executed first; The first round of strategy pool will be executed after the main body is indicated; This indicates the pre-set budget of each implementing entity.
8. The method for generating information guidance strategies in a multi-subject interaction scenario according to claim 7, characterized in that, In step 5, when each executing entity selects a seed node from the social network basic model based on the set of optimization strategies to deploy the information propagation starting point, the subsequent executing entity is in a later round of sequential selection, and the optimization objective is: ; In the formula, This indicates that the second round of revenue will be distributed to the entity that will receive the benefits first. Indicates the entity that holds and executes the shares after co-evolution. Users with opinions; This represents the set of seed nodes selected by the main body in the second round; The second round of strategy pool will be executed after the main body is notified; Represents the set of remaining nodes in the second round; This indicates that the second round of strategy pool will be executed first. This indicates the pre-set budget of each implementing entity; Represents the set of remaining nodes; Indicates the cost multiple.
9. An information guidance strategy generation device for multi-subject interaction scenarios, characterized in that, The device includes: The social network modeling and opinion mechanism configuration module is used to build a basic model of a social network, identify multiple executive entities involved in information guidance, and set up an opinion update mechanism based on hierarchical cognition. The multi-agent policy framework building module is used to build a multi-agent policy iteration framework based on policy pool response arbitrage, and then initialize the policy pool of each executing agent. The optimal response strategy training module is used to generate a strategy interaction matrix through multiple rounds of simulation based on the social network basic model and the strategies in the strategy pool; solve the equilibrium strategy of each executing entity under the strategy interaction matrix; train the optimal response strategy for the equilibrium strategy; and supplement the optimal response strategy to the strategy pool. The strategy iteration and optimization module is used to iteratively execute the process in the best response strategy training module until the preset number of iterations is reached, and obtain an optimized strategy set that adapts to multi-subject interaction. The seed node selection module is used by each execution entity to select a seed node from the social network basic model according to the set of optimization strategies to deploy the information propagation starting point; wherein, the subsequent execution entity obtains the seed node already selected by the previous execution entity according to the requirements and by a set cost multiple; The optimal strategy output module is used to complete the evolution of the viewpoints of all nodes in the basic model of the social network according to the viewpoint update mechanism after all the execution entities have exhausted their budgets, and output the information guidance strategy that the first execution entity obtains the maximum support.
10. A computer device, comprising a memory and a processor, characterized in that, The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the information guidance strategy generation method in any one of claims 1 to 8.