A social robot-detecter adversarial simulation method and apparatus

By modeling the adversarial process between social robots and detectors as a Markov decision process, and optimizing the behavior of social robots using multi-type agents and user caution prediction models, the problems of realism and collaborative manipulation in social robot detection are solved, enabling more complex social behavior simulation and cluster control.

CN121173690BActive Publication Date: 2026-06-02UNIV OF SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
UNIV OF SCI & TECH OF CHINA
Filing Date
2025-11-18
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies for social robot detection suffer from problems such as lack of realism, inability to perform large-scale collaborative control, rudimentary modeling of influence propagation, limited actions that social robots can perform, difficulty in selecting target users, and inability to simulate social clusters.

Method used

The adversarial process between social robots and detectors is modeled as a Markov decision process. By utilizing multi-type intelligent agent cooperative control and combining a user caution prediction model, a profile editor, and a social cluster reward function, the behavioral strategy of social robots is optimized.

Benefits of technology

It improves the realism of the simulated environment, supports more complex social behavior patterns, reduces training difficulty, and enables effective control of social robot swarms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121173690B_ABST
    Figure CN121173690B_ABST
Patent Text Reader

Abstract

The application provides a social robot-detecter confrontation simulation method and device, which can be applied to the technical field of network space security. The social robot-detecter confrontation simulation method comprises the following steps: modeling the confrontation process between a social robot and a social robot detector in a virtual social environment as a Markov decision process, and utilizing multiple types of intelligent agents to cooperatively control the interaction between the social robot and the virtual social environment in the Markov decision process; utilizing a trained user caution degree prediction model to predict the reply probability of the social robot in multiple types of interactive scenes; based on the editable attribute features of the social robot, utilizing a trained personal profile editor and a target preselection mechanism to optimize and improve the social robot in the Markov decision process; utilizing a social cluster reward function to overall optimize the interactive state of a social robot cluster composed of multiple social robots in the Markov decision process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of cyberspace security technology, and specifically to a social robot-detector adversarial simulation method and apparatus. Background Technology

[0002] Social bots are accounts on online social networking platforms controlled by computer programs rather than real people. The operators behind these bots control a large number of them. These bots attempt to impersonate normal users to gain trust and control the narrative. The key to the social bot offensive and defensive game lies in the analysis and countermeasures of account data. Defenders (e.g., social platforms and researchers) use multimodal data mining and machine learning techniques to build detection models based on account characteristics, attempting to identify social bots from a massive user base.

[0003] However, there are currently few technical solutions researched from an offensive and defensive perspective, and existing simulated social environment technologies suffer from a lack of realism and an inability to perform large-scale collaborative control of social robots. Furthermore, offensive research faces challenges such as rudimentary impact propagation modeling, limited executable actions for social robots, difficulty in selecting target users, and the inability to simulate social clusters. Summary of the Invention

[0004] In view of the above problems, the present invention provides a social robot-detector adversarial simulation method and apparatus.

[0005] According to a first aspect of the present invention, a social robot-detector adversarial simulation method is provided, comprising: modeling the adversarial process between a social robot and a social robot detector in a virtual social environment as a Markov decision process; utilizing multi-type intelligent agents to collaboratively control the interaction between the social robot and the virtual social environment during the Markov decision process; using a trained user caution prediction model to predict the return-to-follow probability of the social robot in multi-type interaction scenarios; optimizing and improving the social robot during the Markov decision process based on the editable attribute features of the social robot using a trained profile editor and target pre-selection mechanism; and using a social cluster reward function to comprehensively optimize the interaction state of a social robot cluster composed of multiple social robots during the Markov decision process.

[0006] According to an embodiment of the present invention, the above-described modeling of the adversarial process between the social robot and the social robot detector in a virtual social environment as a Markov decision process includes: constructing a virtual social environment based on a social directed graph and a multi-dimensional user feature matrix to simulate the topology of a real social network; setting a reward function and action spaces and observation spaces for multiple types of intelligent agents in the virtual social environment, and initializing the social robot and the social robot detector; training the social robot under the cooperative control of multiple types of intelligent agents using a proximal policy optimization algorithm based on the action spaces and observation spaces of multiple types of intelligent agents and the reward function; performing adversarial identification on the social robot using the social robot detector during the training process; and modeling the adversarial identification process between the social robot and the social robot detector, and the interaction process between the social robot and the virtual social environment as a Markov decision process.

[0007] According to an embodiment of the present invention, the aforementioned social directed graph includes a follow network, a forwarding network, and a mention network; wherein, the action space of the multi-type intelligent agents is a set of actions used to simulate user action selection and target selection, and the observation space of the multi-type intelligent agents is constructed based on a multi-dimensional user feature matrix; wherein, the interaction between the social robot and the virtual social environment is controlled collaboratively by the multi-type intelligent agents includes: during the training process of the social robot, using action selection intelligent agents and target selection intelligent agents to collaboratively control the action selection and target selection of the social robot.

[0008] According to an embodiment of the present invention, the above-mentioned prediction of the follow-back probability of a social robot in multiple types of interaction scenarios using a trained user caution prediction model includes: training the user caution prediction model using user account data samples in a single interaction scenario to obtain a trained user caution prediction model; and in a virtual social environment, processing the embedded account data of the social robot using the trained user caution prediction model to obtain the follow-back probability of the target social robot in multiple types of interaction scenarios.

[0009] According to embodiments of the present invention, the aforementioned single interaction scenario includes a single follow-back scenario, a single mention-back scenario, or a single reply-back scenario; wherein, the user caution prediction model includes a multi-layer follow-back prediction feedforward neural network, a multi-layer mention-back scenario prediction feedforward neural network, and a multi-layer reply-back scenario prediction feedforward neural network; wherein, the social robot's account data embedding includes the social robot's profile data, the social robot's generated content, and the social robot's social relationship network data; wherein, the multi-type interaction scenario includes a follow-back scenario, a mention-back scenario, and / or a reply-back scenario.

[0010] According to an embodiment of the present invention, the above-mentioned optimization and improvement of a social robot in a Markov decision-making process based on the editable attribute features of a social robot, using a trained profile editor and a target pre-selection mechanism, includes: training the profile editor based on generative-adversarial methods using user editable attribute feature samples and state information samples to obtain a trained profile editor; and optimizing and improving a controlled social robot in a Markov decision-making process based on the editable attribute features of a controlled social robot, using the trained profile editor to obtain an optimized and improved controlled social robot.

[0011] According to embodiments of the present invention, the above-mentioned optimization and improvement of the social robot in the Markov decision-making process based on the editable attribute features of the social robot, using a trained profile editor and a target pre-selection mechanism, further includes: performing interest mining on the current editable attribute features of other users in the virtual social environment to obtain the interest features of other users; analyzing the current editable attribute features of other users using a principal component analysis algorithm to obtain analysis results; clustering the analysis results and interest features using a K-means algorithm, and mapping the clustering results of other users to independent actions to obtain an action space for the controlled social robot to perform target selection; and based on the action space, using a social robot detector to detect all users in the virtual social environment, and filtering all users based on the detection results to achieve the target pre-selection mechanism.

[0012] According to an embodiment of the present invention, the above-described generative-adversarial training of a profile editor using user editable attribute feature samples and state information samples to obtain a trained profile editor includes: processing the user's state information samples using the profile editor's generator to obtain a first robot rating prediction result for the user; processing the user's editable attribute feature samples and the first robot rating prediction result using the profile editor's substitution detector to obtain a second robot rating prediction result for the user, wherein the substitution detector is constructed based on a social robot detector; processing the first robot rating prediction result and the second robot rating prediction result using a preset loss function to obtain a loss value, and updating the generator parameters using the loss value; iteratively executing the generator training operation, the substitution detector data processing operation, the loss value calculation operation, and the generator parameter update operation until the preset training conditions are met to obtain a trained profile editor.

[0013] According to an embodiment of the present invention, the above-mentioned optimization of the interaction state of a social robot cluster composed of multiple social robots in a Markov decision-making process using a social cluster reward function includes: constructing a social cluster reward function based on a policy similarity penalty term based on KL divergence and the reward function of each social robot; using multiple action selection agents and multiple target selection agents to collaboratively control multiple social robots to perform action selection and target selection in the social robot cluster; introducing multiple types of interactive actions between multiple social robots into the discrete action space of the action selection agent, and expanding the target selection action space of the target selection agent based on a target pre-selection mechanism; and using the social cluster reward function to optimize the interaction state of the social robot cluster composed of multiple social robots in a Markov decision-making process.

[0014] A second aspect of the present invention provides a social robot-detector adversarial simulation device, comprising: an interaction modeling module for modeling the adversarial interaction between a social robot and a social robot detector in a virtual social environment as a Markov decision process, and utilizing multi-type intelligent agents to collaboratively control the interaction between the social robot and the virtual social environment during the Markov decision process; a follow-up probability prediction module for predicting the follow-up probability of the social robot in multi-type interaction scenarios using a trained user caution prediction model; an optimization and improvement module for optimizing and improving the social robot during the Markov decision process based on the editable attribute features of the social robot, using a trained profile editor and target pre-selection mechanism; and a social cluster optimization module for optimizing the overall interaction state of a social robot cluster composed of multiple social robots during the Markov decision process using a social cluster reward function.

[0015] A third aspect of the present invention provides an electronic device comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.

[0016] A fourth aspect of the present invention also provides a computer-readable storage medium having a computer program or instructions stored thereon, wherein the computer program or instructions, when executed by a processor, implement the steps of the above-described method.

[0017] The social robot-detector adversarial simulation method provided by this invention models the adversarial process between social robots and social robot detectors as a Markov decision process, which can more realistically simulate the decision-making and counter-decision-making processes in dynamic environments. By utilizing the collaborative control of multiple types of intelligent agents to control the interaction between social robots and the virtual environment, more complex and diverse social behavior patterns can be simulated. Through the user caution prediction model, the ease with which different users are influenced by social robots can be better reflected, improving the realism of the simulation environment and providing an important basis for target user selection. The profile editor can dynamically adjust the profile information of social robots. The target pre-selection mechanism realizes the dimensionality reduction of the large discrete action space of target user selection, reducing the training difficulty. At the same time, the research on social robot action strategies is extended from focusing on a single social robot to considering multiple social robots simultaneously, providing support for the swarm action control of social robots. Attached Figure Description

[0018] The above-described features, other objects, and advantages of the present invention will become clearer from the following description of embodiments of the invention with reference to the accompanying drawings, in which:

[0019] Figure 1 This is an application scenario diagram of the social robot-detector adversarial simulation method according to an embodiment of the present invention.

[0020] Figure 2 This is a flowchart of a social robot-detector adversarial simulation method according to an embodiment of the present invention.

[0021] Figure 3 This is a diagram illustrating the interaction process between a single social robot and a virtual social environment according to an embodiment of the present invention.

[0022] Figure 4 This is a simulation framework diagram of a social robot cluster according to an embodiment of the present invention.

[0023] Figure 5 This is a structural block diagram of a social robot-detector adversarial simulation device according to an embodiment of the present invention.

[0024] Figure 6 This is a block diagram of an electronic device suitable for implementing a social robot-detector adversarial simulation method according to an embodiment of the present invention. Detailed Implementation

[0025] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the invention. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the invention for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concept of the invention.

[0026] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0027] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0028] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0029] For social bots, account data is primarily influenced by action strategies. Reasonable actions, such as periodically changing usernames and posting comments at opportune moments, can confuse social bot detection models. Training agents for social bot action decisions can improve their stealth and influence, breaking through the defenses of opposing detection models and increasing our voice in relevant games. Against this backdrop, researching intelligent decision-making models and training methods for social bots is of great significance. However, current research from the attacker's perspective is limited, and the application of social bots lacks theoretical guidance. Research on real social platforms is time-consuming and laborious, and can easily lead to adverse effects. Furthermore, current simulated experimental environments lack realism, especially in complex information warfare scenarios, where existing methods struggle to achieve large-scale collaborative manipulation. The reasons for this, therefore, lie in the following key challenges faced by attacker research:

[0030] (1) The modeling of influence propagation is rudimentary: Current research uses a simple heuristic based on the number of followers of the user and the number of times the social robot interacts with the user when modeling the probability that the social robot will influence the propagation to a specific user, but this cannot reflect the real situation.

[0031] (2) Limited Actions for Social Robots: Current models typically simplify the action space of social robots to four basic operations: "follow / post / forward / mention," while ignoring other possible actions. Editing account attributes (such as modifying usernames or profile pictures) has been proven to be an effective means of enhancing the detection avoidance capabilities of social robots. However, existing reinforcement learning-based methods have failed to achieve this function due to the difficulty in mapping such discrete actions to the action space.

[0032] (3) Difficulty in selecting target users: Existing studies usually use dedicated agents with discrete action spaces for user selection. The dimension N of the action space is usually set to around 1000, with each action corresponding to a specific user. However, as N increases, this method leads to the problem of high-dimensional discrete action spaces, which means that the agent needs to explore the optimal strategy from a larger set of actions, thus significantly increasing the training cost.

[0033] (4) Inability to simulate social robot swarms: Current research is mostly limited to the behavior optimization of individual social robots, neglecting scenarios where multiple social robots collaborate. In real social networks, social robots often exist in groups, and their collaboration may bring new challenges and opportunities. If the social relationships between multiple social robots are too close (such as mechanically following each other), the community they form may form an abnormal subgraph structure, which can be identified by the current mainstream graph neural network-based detection methods. Reasonable collaboration strategies may actually help social robots evade such detection methods.

[0034] To address these issues, there is an urgent need to construct a realistic social platform simulation environment and a social robot action strategy model. Such an environment should be able to accurately simulate the influence of social robots and their interactions with users. The social robot action strategy model should be able to support more complex behavioral strategies and employ a more reasonable action space representation method to enhance its scalability.

[0035] It should be noted that in the embodiments of the present invention, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of the present invention. However, they do not mean that the applicant has used or necessarily used the solution.

[0036] It should be noted that, in the technical solution of this invention, the user information (including but not limited to user personal information, user image information, user device information, such as location information) and data (including but not limited to data used for analysis, stored data, and displayed data) involved are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of related data all comply with relevant laws, regulations, and standards, necessary confidentiality measures have been taken, they do not violate public order and good morals, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0037] Meanwhile, in scenarios involving automated decision-making using personal information, the methods, devices, and systems provided in this invention all offer users corresponding entry points for choosing to agree to or reject the automated decision-making results; if the user chooses to reject, the process proceeds to the expert decision-making stage. Here, "automated decision-making" refers to the activity of automatically analyzing and evaluating an individual's behavioral habits, interests, or economic, health, and credit status through computer programs, and then making a decision. Here, "expert decision-making" refers to the activity of making decisions by personnel who specialize in a particular field, possess specialized experience, knowledge, and skills, and have reached a certain level of professional expertise.

[0038] Figure 1 This is an application scenario diagram of the social robot-detector adversarial simulation method according to an embodiment of the present invention.

[0039] like Figure 1 As shown, the application scenario 100 according to this embodiment may include scenarios such as cyberspace security. Network 104 serves as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables, etc.

[0040] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).

[0041] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0042] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0043] It should be noted that the social robot-detector adversarial simulation method provided in this embodiment of the invention can generally be executed by server 105. Correspondingly, the social robot-detector adversarial simulation device provided in this embodiment of the invention can generally be located in server 105. The social robot-detector adversarial simulation method provided in this embodiment of the invention can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the social robot-detector adversarial simulation device provided in this embodiment of the invention can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.

[0044] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0045] The following will be based on Figure 1 The described scene, through Figures 2-4 The social robot-detector adversarial simulation method of the disclosed embodiments is described in detail.

[0046] Figure 2 This is a flowchart of a social robot-detector adversarial simulation method according to an embodiment of the present invention.

[0047] like Figure 2 As shown, the social robot-detector adversarial simulation in this embodiment includes operations S210 to S240.

[0048] In operation S210, the adversarial process between the social robot and the social robot detector in the virtual social environment is modeled as a Markov decision process, and in the Markov decision process, the interaction between the social robot and the virtual social environment is controlled by multiple types of intelligent agents in a collaborative manner.

[0049] Modeling the adversarial process between social robots and social robot detectors as a Markov decision process allows for a more realistic simulation of decision-making and counter-decision-making processes in dynamic environments. This modeling approach considers state transitions and reward mechanisms over time, enabling the social robot's behavioral strategies to be dynamically adjusted based on feedback from the social robot detector, thus more effectively evading detection.

[0050] The aforementioned types of intelligent agents include action-selection agents and goal-selection agents. These two different types of agents work collaboratively to control the social robot's strategy selection (e.g., action selection and goal selection). The action-selection agent controls the social robot's behavior, such as making statements, while the goal-selection agent controls the social robot's choice of which social accounts to follow, such as choosing to follow, not follow, or unfollow a particular social media account. Different types of agents can focus on different interaction strategies. This division of labor and cooperation can generate a more realistic swarm effect, increasing the detection difficulty for social robot detectors.

[0051] When operating S220, the trained user caution prediction model is used to predict the follow-back probability of the social robot in various interaction scenarios.

[0052] By using a trained user caution prediction model, social bots can adjust their interaction strategies to make their behavior more consistent with that of real users, thereby reducing the risk of being identified by social bot detectors.

[0053] In operation S230, based on the editable attribute features of the social robot, the social robot in the Markov decision-making process is optimized and improved by using the trained profile editor and target pre-selection mechanism.

[0054] The target pre-selection mechanism can filter out target users from the virtual social environment, and by clustering the target users, the results of the clustering process can be used to expand the action space of multiple types of intelligent agents.

[0055] Using a profile editor enables social bots to quickly adapt to the environments and user characteristics of different social platforms, thereby maintaining the effectiveness of their human-like features.

[0056] In operation S240, the social cluster reward function is used to optimize the overall interaction state of a social robot cluster consisting of multiple social robots in the Markov decision process.

[0057] By leveraging the collaborative behavior of multiple agents within a social bot swarm, more complex and realistic social interaction patterns can be simulated, making them harder for social bot detectors to identify as social bots. Optimizing the social bot swarm as a whole, rather than optimizing each individual bot, helps the swarm maintain stable adversarial capabilities in dynamically changing virtual social environments.

[0058] Operations S210 to S240 above together realize the adversarial simulation between the social robot and the social robot detector. Specifically, operation S210 realizes the function of modeling the adversarial simulation process as a Markov decision process; operation S220 realizes the simulation of the social robot's follow-up function; operation S230 realizes the simulation of the social robot's self-data (or profile) editing function; and operation S240 realizes the simulation and optimization of the social cluster.

[0059] The social robot-detector adversarial simulation method provided by this invention models the adversarial process between social robots and social robot detectors as a Markov decision process, which can more realistically simulate the decision-making and counter-decision-making processes in dynamic environments. By utilizing the collaborative control of multiple types of intelligent agents to control the interaction between social robots and the virtual environment, more complex and diverse social behavior patterns can be simulated. Through the user caution prediction model, the ease with which different users are influenced by social robots can be better reflected, improving the realism of the simulation environment and providing an important basis for target user selection. The profile editor can dynamically adjust the profile information of social robots. The target pre-selection mechanism realizes the dimensionality reduction of the large discrete action space of target user selection, reducing the training difficulty. At the same time, the research on social robot action strategies is extended from focusing on a single social robot to considering multiple social robots simultaneously, providing support for the swarm action control of social robots.

[0060] The following detailed description, in conjunction with the accompanying drawings, further illustrates each operation of the social robot-detector adversarial simulation method provided in this embodiment.

[0061] According to an embodiment of the present invention, the above-described modeling of the adversarial process between the social robot and the social robot detector in a virtual social environment as a Markov decision process includes: constructing a virtual social environment based on a social directed graph and a multi-dimensional user feature matrix to simulate the topology of a real social network; setting a reward function and action spaces and observation spaces for multiple types of intelligent agents in the virtual social environment, and initializing the social robot and the social robot detector; training the social robot under the cooperative control of multiple types of intelligent agents using a proximal policy optimization algorithm based on the action spaces and observation spaces of multiple types of intelligent agents and the reward function; performing adversarial identification on the social robot using the social robot detector during the training process; and modeling the adversarial identification process between the social robot and the social robot detector, and the interaction process between the social robot and the virtual social environment as a Markov decision process.

[0062] According to an embodiment of the present invention, the aforementioned social directed graph includes a follow network, a forwarding network, and a mention network; wherein, the action space of the multi-type intelligent agents is a set of actions used to simulate user action selection and target selection, and the observation space of the multi-type intelligent agents is constructed based on a multi-dimensional user feature matrix; wherein, the interaction between the social robot and the virtual social environment is controlled collaboratively by the multi-type intelligent agents includes: during the training process of the social robot, using action selection intelligent agents and target selection intelligent agents to collaboratively control the action selection and target selection of the social robot.

[0063] The following describes specific implementation methods in conjunction with appendices. Figure 3 The operation S210 involved in the above embodiments of the present invention will be described in further detail.

[0064] Figure 3 This is a diagram illustrating the interaction process between a single social robot and a virtual social environment according to an embodiment of the present invention.

[0065] The operation S210 involved in the above embodiment mainly models the interaction process between the social robot and the virtual social environment as a Markov decision process. The interaction between the social robot and the virtual social environment includes the adversarial recognition process between the social robot and the social robot detector, the interaction process between the social robot and other social robots, and the interaction process between the social robot and non-social robots (such as real users).

[0066] The virtual social environment described in this invention is as follows: Figure 3As shown. To simulate the development process of a social robot within a reasonable timeframe and avoid interfering with normal users, this invention constructs a virtual social environment based on real data. For example, user class 1 and user class 2 are constructed in the target selection agent to simulate different users in a real social platform. User class 1 and user class 2 are used as selectable target sequences for the target selection agent. Those skilled in the art can select other types or numbers of users according to actual needs. An action sequence is constructed for the action selection agent, including publishing an article, modifying a profile, replying, following, and waiting for one day. When modifying a profile, the action selection agent can call the profile survivor to modify the social robot's profile. Interaction with the social robot is achieved through the above two types of agents, and the data account embedding of the social robot is updated. This virtual social environment is implemented through a social directed graph. (in This indicates an interest in the internet. Indicates forwarding network, (Indicates mention of network) and user feature matrix (i.e., a multi-dimensional user feature matrix) is used to simulate the core interactive functions of a real social platform. The node set of the social directed graph... (in, Focus on the set of nodes in the network. This represents the set of nodes in the forwarding network. The set of nodes (representing the network mentioned) represents users in the network that are followed, the network that is forwarded, and the network that is mentioned, respectively. The edge set... (in, Focus on the edge sets of the network. Represents the edge set of the forwarding network. The edge set of the network mentioned corresponds to the specific social behavior connection. 3D user feature matrix Each user's attributes are fully described using feature vectors, where This indicates the total number of users in a virtual social environment (or virtual social platform). The number of features is denoted by . This virtual social environment recreates the network topology from real social datasets and retains only active accounts that directly interact with the controlled social bot to improve computational efficiency. When the social bot performs actions such as following / forwarding / mentioning, or receives interactions from other accounts (other uncontrolled social bots or real users), the topological relationships of the entire graph are updated in real time. This design provided by the present invention ensures both the realism of the simulation method and the efficiency of the computation process.

[0067] This invention uses multiple types of intelligent agents to make decisions for the actions of social robots, such as Figure 3The diagram shows the action selection agent and the target selection agent, respectively. The social bot's account data embedding is input into multiple agent types as the current state, yielding the next action: posting an article, interacting with other users, modifying profile data, or waiting for a period of time. If interacting with other users is chosen, the target selection agent selects a target user class and then chooses a target user based on the user's level of caution. The simulation environment then randomly generates the interaction result based on the user's level of caution, i.e., whether the user will follow the social bot. If modifying profile data is chosen, a profile data generator generates a suitable user profile based on the current account data embedding. After completing an action, the account data embedding is updated based on the action result and input into the social bot detector. The detection result determines whether to continue the simulation. In the simulation environment, a new account is created, and reinforcement learning is used to allow the agent to control the actions of the new account, continuously interacting with the environment and updating the account data until it is detected by the detection model. This simulates the development process of a social bot account.

[0068] This invention employs the Proximal Policy Optimization (PPO) algorithm based on a multilayer perceptron strategy to train a social robot policy model. The behavior of the social robot is collaboratively controlled by two agents, one responsible for action selection and the other for target selection.

[0069] Observation Space: Due to the experimental scale limitations provided by this invention, a complete virtual social environment state cannot be input. Therefore, the input to the social robot policy model is limited to the current state features (i.e., the features used to train the random forest social robot detection model). The two agents share a dimension of... The same observation space.

[0070] Action Space: To enable social robots to interact in a virtual environment, this invention defines a set of actions that simulate the behavior of users on real social platforms, such as... Figure 3 As shown, it includes:

[0071] (1) Post a tweet.

[0072] (2) Update personal profile.

[0073] (3) Forward / mention / follow other accounts.

[0074] (4) Create or delete lists.

[0075] (5) Wait one day.

[0076] Therefore, the action space of the action selection agent is a discrete space containing 8 actions.

[0077] To reflect the two objectives of anonymity and influence of social bots, this invention sets the reward function for reinforcement learning as follows: the gains, or the impact, of a social bot can be measured by its number of followers. Assume a social bot, at time... The state at that time is The number of fans is Combining the two objectives of anonymity and influence of social robots, a reward function can be proposed. As shown in formula (1):

[0078] (1).

[0079] in, For social robot detectors, This indicates that the social bot detector inputs account data. The probability that the account output is a social bot. This is an adjustable parameter. The larger the size, the more the social robot's training will tend to expand its influence; conversely, the smaller the size, the more it will tend to evade detection.

[0080] According to an embodiment of the present invention, the above-mentioned prediction of the follow-back probability of a social robot in multiple types of interaction scenarios using a trained user caution prediction model includes: training the user caution prediction model using user account data samples in a single interaction scenario to obtain a trained user caution prediction model; and in a virtual social environment, processing the embedded account data of the social robot using the trained user caution prediction model to obtain the follow-back probability of the target social robot in multiple types of interaction scenarios.

[0081] According to embodiments of the present invention, the aforementioned single interaction scenario includes a single follow-back scenario, a single mention-back scenario, or a single reply-back scenario; wherein, the user caution prediction model includes a multi-layer follow-back prediction feedforward neural network, a multi-layer mention-back scenario prediction feedforward neural network, and a multi-layer reply-back scenario prediction feedforward neural network; wherein, the social robot's account data embedding includes the social robot's profile data, the social robot's generated content, and the social robot's social relationship network data; wherein, the multi-type interaction scenario includes a follow-back scenario, a mention-back scenario, and / or a reply-back scenario.

[0082] The following detailed description of the operation S220 involved in the above embodiments of the present invention will be provided through specific implementation methods.

[0083] It is generally believed that the spread of influence by social bots occurs when normal users follow the bot. Current research uses a simple heuristic based on the number of followers and interactions between the user and the bot when modeling the probability of a social bot's influence reaching a specific user, but this fails to reflect reality. In fact, different users vary greatly in their susceptibility to social bot influence: some users are more cautious when facing strangers, while others are less discerning. Such differences cannot be simply summarized by the number of followers. Therefore, it is necessary to develop a more comprehensive model based on user account data to predict the probability of a user being influenced by a social bot. Specifically, this invention will utilize mature account data embedding technology in social bot detection, based on user profile data, user-generated content, and social relationship network data, to predict the probability that a user will follow the bot when followed / mentioned / replied to by it. Modeling this probability can effectively improve the realism of the simulated environment.

[0084] First, this invention embeds the following account data based on user data:

[0085] (1) Account profile data: Using traditional feature engineering methods, the main features extracted are account age, number of followers, number of followings, number of posts, and whether the default profile background image is used.

[0086] (2) User-generated content embedding: The embedding of each of its posts is obtained using an improved pre-trained language model (e.g., the RoBERTa model) and the average embedding vector of all its posts is calculated.

[0087] (3) Social Relationship Network Embedding: The above two account data embeddings are concatenated to obtain the embedding vector of each user. Each user is regarded as a node. Based on its social relationship graph, the embedding vector is passed through a two-layer graph convolutional neural network to obtain the account data embedding containing its social relationship graph information.

[0088] Next, based on the account data embeddings obtained for each user, this invention uses three three-layer feedforward neural networks to predict the probability that a user will follow the social bot when followed / mentioned / replied by it. To avoid these three types of interactions from interfering with each other, only users exposed to a single type of interaction with the social bot will be used during training. This embedding will also be used in step one and subsequent steps to represent an account at time [time]. state of time .

[0089] According to an embodiment of the present invention, the above-mentioned optimization and improvement of a social robot in a Markov decision-making process based on the editable attribute features of a social robot, using a trained profile editor and a target pre-selection mechanism, includes: training the profile editor based on generative-adversarial methods using user editable attribute feature samples and state information samples to obtain a trained profile editor; and optimizing and improving a controlled social robot in a Markov decision-making process based on the editable attribute features of a controlled social robot, using the trained profile editor to obtain an optimized and improved controlled social robot.

[0090] According to embodiments of the present invention, the above-mentioned optimization and improvement of the social robot in the Markov decision-making process based on the editable attribute features of the social robot, using a trained profile editor and a target pre-selection mechanism, further includes: performing interest mining on the current editable attribute features of other users in the virtual social environment to obtain the interest features of other users; analyzing the current editable attribute features of other users using a principal component analysis algorithm to obtain analysis results; clustering the analysis results and interest features using a K-means algorithm, and mapping the clustering results of other users to independent actions to obtain an action space for the controlled social robot to perform target selection; and based on the action space, using a social robot detector to detect all users in the virtual social environment, and filtering all users based on the detection results to achieve the target pre-selection mechanism.

[0091] According to an embodiment of the present invention, the above-described generative-adversarial training of a profile editor using user editable attribute feature samples and state information samples to obtain a trained profile editor includes: processing the user's state information samples using the profile editor's generator to obtain a first robot rating prediction result for the user; processing the user's editable attribute feature samples and the first robot rating prediction result using the profile editor's substitution detector to obtain a second robot rating prediction result for the user, wherein the substitution detector is constructed based on a social robot detector; processing the first robot rating prediction result and the second robot rating prediction result using a preset loss function to obtain a loss value, and updating the generator parameters using the loss value; iteratively executing the generator training operation, the substitution detector data processing operation, the loss value calculation operation, and the generator parameter update operation until the preset training conditions are met to obtain a trained profile editor.

[0092] The operation S230 involved in the above embodiments of the present invention will be further described in detail below through specific implementation methods.

[0093] Editing account attributes (such as modifying usernames or profile pictures) has proven to be an effective means of enhancing the detection avoidance capabilities of social bots. However, existing reinforcement learning-based methods have failed to achieve this functionality due to the difficulty in mapping such discrete actions to an action space. To overcome this technical limitation, this invention proposes implementing a dedicated account attribute editor module, thereby eliminating reliance on agent-based methods.

[0094] This account attribute editor takes the current state of the controlled social bot as input, and its core objective is to adjust account-related characteristics to minimize suspiciousness. It's important to note that while editing certain characteristics (such as account verification status or privacy settings) might significantly reduce anomaly scores, it could lead to increased costs, difficulties in spreading influence, and other challenges; therefore, these have been excluded from the editable scope. After systematic review, the editable characteristics include:

[0095] (1) Username Length. Considering the naming rules of a certain real social platform (usernames must be longer than 4 characters and no more than 15 characters, only letters, numbers, and underscores are allowed, and spaces are prohibited) and the flexibility requirements for the number of digits in the username, this invention assumes that usernames only allow letters and underscores, thus each character has 27 possible schemes. Given that usernames are unique and the number of real users exceeds 300 million, the username length range is set as follows: .

[0096] (2) The number of digits in the username.

[0097] (3) Display name length: The display name can be up to 50 characters long and can be repeated with other users, so its value range is: .

[0098] (4) Whether to use the default profile background image.

[0099] (5) Length of personal profile. Its value range is... .

[0100] (6) Does the personal profile include a URL?

[0101] The profile editor's structure resembles a generative adversarial network, consisting of two three-layer multilayer perceptrons and a sigmoid layer. It's primarily used for editing the social bot's own profile (or archive), thus simulating the profile editing functionality of a real social media account. This is because social bots cannot directly access the social bot detector. Therefore, an alternative detector is used. Fit its output to achieve adversarial training. Another acts as a generator. Multilayer perceptron with As input, a 6-dimensional vector is generated, which can be mapped to the value range of the six features mentioned above. During the training phase, a social bot detector is first used to predict social bot ratings for all users in the dataset; then, an alternative model is trained to predict user-specific ratings. Predict social bot ratings; finally optimize the generator using the following loss function. To minimize the user's social bot rating, as shown in formula (2):

[0102] (2).

[0103] To reduce the dimensionality of the action space during target user selection, this invention proposes user clustering before constructing the action space. Specifically, the user selection process is heuristically decomposed into two sequential steps: initial selection based on interests and final selection based on reputation. This invention first uses tweet embedding vectors to capture user interest features. Given the high dimensionality of the original embeddings, this invention employs Principal Component Analysis (PCA) to reduce data sparsity and computational overhead, setting the number of principal components to 15 (cumulative variance explained slightly above 0.95). Subsequently, the K-Means algorithm is applied to cluster users with similar interests, mapping each cluster to an independent action. Then, a social bot detection model is used to predict the suspiciousness of all users, selecting the user with the lowest suspiciousness (i.e., the highest reputation) as the target user for that cluster.

[0104] According to an embodiment of the present invention, the above-mentioned optimization of the interaction state of a social robot cluster composed of multiple social robots in a Markov decision-making process using a social cluster reward function includes: constructing a social cluster reward function based on a policy similarity penalty term based on KL divergence and the reward function of each social robot; using multiple action selection agents and multiple target selection agents to collaboratively control multiple social robots to perform action selection and target selection in the social robot cluster; introducing multiple types of interactive actions between multiple social robots into the discrete action space of the action selection agent, and expanding the target selection action space of the target selection agent based on a target pre-selection mechanism; and using the social cluster reward function to optimize the interaction state of the social robot cluster composed of multiple social robots in a Markov decision-making process.

[0105] The following describes specific implementation methods in conjunction with appendices. Figure 4 The overall optimization process of the social robot cluster involved in operating S240 is explained in further detail.

[0106] Figure 4 This is a simulation framework diagram of a social robot cluster according to an embodiment of the present invention.

[0107] In practical applications, attackers often control a large number of social bots simultaneously to maximize their influence. However, existing research has only discussed the case of controlling a single social bot, neglecting the issue of how to ensure cooperation and mutual benefit among social bots when the number of bots expands. Therefore, this invention proposes a social bot swarm simulation framework and, based on the interaction process between a single social bot and its environment described in the previous steps, modifies the reward function and the target selection action space as follows.

[0108] The simulation framework for social robot swarms proposed in this invention is as follows: Figure 4 As shown. During the training process, as... Figure 4 As shown, accounts for multiple social bots (i.e., social bot 1, social bot 2, social bot 3, and social bot N shown in the figure) will be deployed in a simulated environment to simulate the swarm behavior of social bots. This invention maintains a pool of available social bots and trains multiple agents (i.e.,...) Figure 4 Agent 1, Agent 2, and Agent n (where N and n are positive integers) control a subset of social robots in the pool. Training multiple agents allows the social robots to offer different action strategy choices (i.e., Figure 4 The modified profile, reply, follow, and wait one day shown make the social robot cluster more differentiated. The account data embedding of the social robot cluster is updated through the above action strategy selection results, and in this process, the social robot detector is used to remove the social robots that have been identified in the available social robot pool. To this end, the present invention further introduces a policy similarity penalty term based on KL divergence into the reward function, as shown in formula (3):

[0109] (3).

[0110] in, Indicates the first Action strategy of a controlled social robot Indicates the first Action strategy for a controlled social robot. This invention proposes to introduce the interactive actions between social robots, such as following / replying to other social robots, into the action space of the agent, expanding the optimization object of the agent from the state of a single social robot to the state of a social robot cluster, thereby achieving overall optimization of the interactive actions between social robots. Therefore, the final target selection agent action space is a discrete action space that maps the target selected in operation S230 to the actions of the friendly social robots. The reward function is the sum of the actions of all friendly social robots. The sum of the policy similarity penalty term.

[0111] The social robot-detector adversarial simulation method provided by this invention, by introducing normal user return probability modeling, better reflects the degree of difficulty of different users being influenced by social robots, improves the realism of the simulation environment, and provides an important basis for target user selection; it improves the optional actions of social robots and innovatively introduces a profile editor generated by adversarial training; it proposes to use user clustering to reduce the dimensionality of the large discrete action space of target user selection, thereby reducing the training difficulty; at the same time, it expands the research on social robot action strategies from focusing on a single social robot to considering multiple social robots simultaneously, providing support for the control of social robot cluster actions.

[0112] The following detailed experiment further illustrates the social robot-detector adversarial simulation method provided by this invention.

[0113] Given a dataset containing real user data from a social platform, this invention first extracts a directed social graph from it according to operation S210. and multi-dimensional user feature matrix A simulation environment is constructed using a reinforcement learning library, and an opponent detection model to be simulated is trained. Then, a user caution prediction model is trained according to operation S220. Next, following operations S230 and S240, the agents, observation space, and action space of the action policy model are determined, and a social robot swarm action policy model based on multi-agent reinforcement learning is established. Finally, the simulation environment can be used to... Figure 3 The interaction process in the process involves training a social robot action strategy model or simulating a pre-trained social robot swarm (such as...). Figure 4 (As shown) its performance on social media platforms.

[0114] Based on the aforementioned social robot-detector adversarial simulation method, this invention also provides a social robot-detector adversarial simulation device. The following will combine... Figure 5 The device is described in detail.

[0115] Figure 5 This is a structural block diagram of a social robot-detector adversarial simulation device according to an embodiment of the present invention.

[0116] like Figure 5 As shown, the social robot-detector adversarial simulation device 500 of this embodiment includes an interaction modeling module 510, a return probability prediction module 520, an optimization and improvement module 530, and a social cluster optimization module 540.

[0117] The interaction modeling module 510 is used to model the adversarial relationship between the social robot and the social robot detector in the virtual social environment as a Markov decision process, and to use multiple types of intelligent agents to collaboratively control the interaction between the social robot and the virtual social environment during the Markov decision process; in one embodiment, the interaction modeling module 510 can be used to perform the operation S210 described above, which will not be repeated here.

[0118] The follow-up probability prediction module 520 is used to predict the follow-up probability of the social robot in various interaction scenarios using a trained user caution prediction model. In one embodiment, the follow-up probability prediction module 520 can be used to perform the operation S220 described above, which will not be repeated here.

[0119] The optimization and improvement module 530 is used to optimize and improve the social robot in the Markov decision-making process based on the editable attribute features of the social robot, using a trained profile editor and a target pre-selection mechanism. In one embodiment, the optimization and improvement module 530 can be used to perform the operation S230 described above, which will not be repeated here.

[0120] The social cluster optimization module 540 is used to optimize the overall interaction state of a social robot cluster consisting of multiple social robots in a Markov decision process using a social cluster reward function. In one embodiment, the social cluster optimization module 540 can be used to execute the operation S240 described above, which will not be repeated here.

[0121] According to embodiments of the present invention, any plurality of modules among the interaction modeling module 510, the return probability prediction module 520, the optimization and improvement module 530, and the social cluster optimization module 540 can be merged into one module, or any one of these modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in one module. According to embodiments of the present invention, at least one of the interaction modeling module 510, the return probability prediction module 520, the optimization and improvement module 530, and the social cluster optimization module 540 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging circuitry, or implemented in hardware or firmware, or in any one of software, hardware, and firmware implementations, or in a suitable combination of any of these. Alternatively, at least one of the interaction modeling module 510, the return probability prediction module 520, the optimization and improvement module 530, and the social cluster optimization module 540 can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.

[0122] Figure 6 This is a block diagram of an electronic device suitable for implementing a social robot-detector adversarial simulation method according to an embodiment of the present invention.

[0123] like Figure 6 As shown, an electronic device 600 according to an embodiment of the present invention includes a processor 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage portion 608 into a random access memory (RAM) 603. The processor 601 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 601 may also include onboard memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0124] RAM 603 stores various programs and data required for the operation of electronic device 600. Processor 601, ROM 602, and RAM 603 are interconnected via bus 604. Processor 601 executes various operations of the method flow according to embodiments of the present invention by executing programs in ROM 602 and / or RAM 603. It should be noted that the programs may also be stored in one or more memories other than ROM 602 and RAM 603. Processor 601 may also execute various operations of the method flow according to embodiments of the present invention by executing programs stored in said one or more memories.

[0125] According to an embodiment of the present invention, the electronic device 600 may further include an input / output (I / O) interface 605, which is also connected to a bus 604. The electronic device 600 may also include one or more of the following components connected to the input / output (I / O) interface 605: an input section 606 including a keyboard, mouse, etc.; an output section 607 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the input / output (I / O) interface 605 as needed. A removable medium 611, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 610 as needed so that computer programs read from it can be installed into the storage section 608 as needed.

[0126] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of the present invention.

[0127] According to embodiments of the present invention, a computer-readable storage medium may be a non-volatile computer-readable storage medium, such as including, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of the present invention, a computer-readable storage medium may include ROM 602 and / or RAM 603 and / or one or more memories other than ROM 602 and RAM 603 described above.

[0128] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0129] Those skilled in the art will understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, the features described in the various embodiments of the present invention can be combined and / or combined in various ways without departing from the spirit and teachings of the present invention. All such combinations and / or combinations fall within the scope of the present invention.

[0130] The embodiments of the present invention have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of the invention. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of the invention, and all such substitutions and modifications should fall within the scope of the invention.

Claims

1. A social robot-detector adversarial simulation method, characterized in that, The method includes: The adversarial process between the social robot and the social robot detector in the virtual social environment is modeled as a Markov decision process, and in the Markov decision process, the interaction between the social robot and the virtual social environment is controlled by multiple types of intelligent agents in a collaborative manner. The method involves using a trained user caution prediction model to predict the follow-back probability of the social robot in various interaction scenarios. This includes: training the user caution prediction model using user account data samples from a single interaction scenario to obtain the trained model; and processing the social robot's account data embedding using the trained user caution prediction model in the virtual social environment to obtain the follow-back probability of the target social robot in various interaction scenarios. The user caution prediction model includes a multi-layer follow-back prediction feedforward neural network, a multi-layer mention-follow-back prediction feedforward neural network, and a multi-layer reply-follow-back prediction feedforward neural network. The social robot's account data embedding includes the social robot's profile data, the social robot's generated content, and the social robot's social relationship network data. The generated content is obtained using a pre-trained language model, and the social relationship network data is obtained using a two-layer graph convolutional neural network. Based on the editable attribute features of social robots, the social robots in the Markov decision-making process are optimized and improved by using a trained profile editor and a target pre-selection mechanism. The social cluster reward function is used to optimize the overall interaction state of the social robot cluster consisting of multiple social robots in the Markov decision process. In the social robot cluster, the multi-type agent can control multiple social robots.

2. The method according to claim 1, characterized in that, Modeling the adversarial process between social robots and social robot detectors in virtual social environments as a Markov decision process includes: A virtual social environment is constructed based on social directed graphs and multi-dimensional user feature matrices to simulate the topology of real social networks; In the virtual social environment, a reward function and the action space and observation space of the multi-type intelligent agents are set, and the social robot and the social robot detector are initialized; Based on the action space and observation space of the multi-type intelligent agents and the reward function, the training of the social robot under the cooperative control of the multi-type intelligent agents is realized by using the proximal policy optimization algorithm. During the training process of the social robot, the social robot detector is used to perform adversarial identification on the social robot; The adversarial identification process between the social robot and the social robot detector, and the interaction process between the social robot and the virtual social environment are modeled as the Markov decision process.

3. The method according to claim 2, characterized in that, The social directed graph includes a follow network, a forwarding network, and a mention network; The action space of the multi-type intelligent agent is a set of actions used to simulate user action selection and target selection, and the observation space of the multi-type intelligent agent is constructed based on the multi-dimensional user feature matrix. The method of using multiple types of intelligent agents to collaboratively control the interaction between the social robot and the virtual social environment includes: during the training process of the social robot, using an action selection agent and a target selection agent to collaboratively control the action selection and target selection of the social robot.

4. The method according to claim 1, characterized in that, The single interaction scenario includes a single follow-follow scenario, a single mention-follow-follow scenario, or a single reply-follow-follow scenario. The various interactive scenarios include follow-follow-back scenarios, mention-follow-back scenarios, and / or reply-follow-back scenarios.

5. The method according to claim 1, characterized in that, Based on the editable attribute features of social robots, the optimization and improvement of social robots in the Markov decision-making process are carried out using a trained profile editor and a target pre-selection mechanism, including: The profile editor is trained using generative-adversarial training based on user editable attribute feature samples and state information samples to obtain the trained profile editor. Based on the editable attribute features of the controlled social robot, the trained profile editor is used to optimize and improve the controlled social robot in the Markov decision process, resulting in an optimized and improved controlled social robot.

6. The method according to claim 5, characterized in that, Also includes: Interest mining is performed on the current editable attribute features of other users in the virtual social environment to obtain the interest features of the other users; The principal component analysis algorithm is used to analyze the current editable attribute features of the other users to obtain the analysis results; The K-means algorithm is used to cluster the analysis results and the interest features, and the clustering results of other users are mapped to independent actions to obtain the action space for the controlled social robot to select targets. Based on the action space, the social robot detector is used to detect all users in the virtual social environment, and the detection results are used to filter all users to achieve the target pre-selection mechanism.

7. The method according to claim 5, characterized in that, The profile editor is trained using user-editable attribute feature samples and state information samples, resulting in a trained profile editor comprising: The user's status information sample is processed using the generator of the profile editor to obtain the user's first robot rating prediction result; The user's editable attribute feature samples and the first robot rating prediction result are processed using the alternative detector of the profile editor to obtain the user's second robot rating prediction result, wherein the alternative detector is constructed based on the social robot detector; The first robot rating prediction result and the second robot rating prediction result are processed using a preset loss function to obtain a loss value, and the generator parameters are updated using the loss value. The generator training operation, the alternative detector data processing operation, the loss value calculation operation, and the generator parameter update operation are executed iteratively until the preset training conditions are met, and the trained profile editor is obtained.

8. The method according to claim 1, characterized in that, The overall optimization of the interaction state of a social robot cluster consisting of multiple social robots in the Markov decision process using a social cluster reward function includes: A social cluster reward function is constructed based on a policy similarity penalty term using KL divergence and a reward function for each social robot. Multiple action selection agents and multiple target selection agents are used to collaboratively control multiple social robots to perform action and target selection in the social robot cluster; Multiple types of interactive actions between the social robots are introduced into the discrete action space of the action selection agent, and the target selection action space of the target selection agent is expanded based on the target pre-selection mechanism. The social cluster reward function is used to optimize the overall interaction state of the social robot cluster, which consists of multiple social robots, in the Markov decision process.

9. A social robot-detector adversarial simulation device, characterized in that, The device includes: The interaction modeling module is used to model the adversarial relationship between the social robot and the social robot detector in the virtual social environment as a Markov decision process, and to use multiple types of intelligent agents to collaboratively control the interaction between the social robot and the virtual social environment in the Markov decision process. The follow-back probability prediction module is used to predict the follow-back probability of the social robot in multiple interaction scenarios using a trained user caution prediction model. This includes: training the user caution prediction model using user account data samples from a single interaction scenario to obtain the trained user caution prediction model; and processing the social robot's account data embedding using the trained user caution prediction model in the virtual social environment to obtain the follow-back probability of the target social robot in multiple interaction scenarios. The user caution prediction model includes a multi-layer follow-back prediction feedforward neural network, a multi-layer mention-follow-back prediction feedforward neural network, and a multi-layer reply-follow-back prediction feedforward neural network. The social robot's account data embedding includes the social robot's profile data, the social robot's generated content, and the social robot's social relationship network data. The generated content is obtained using a pre-trained language model, and the social relationship network data is obtained using a multi-layer graph convolutional neural network. An optimization and improvement module is used to optimize and improve the social robot in the Markov decision-making process based on the editable attribute features of the social robot, using a trained profile editor and a target pre-selection mechanism. The social cluster optimization module is used to optimize the overall interaction state of the social robot cluster consisting of multiple social robots in the Markov decision process using the social cluster reward function, wherein the multi-type agent can control multiple social robots in the social robot cluster.