Lightweight large language model social simulation method and device
Patent Information
- Application Number
- CN202611356800.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-09-03
- Publication Date
- 2026-09-29
AI Technical Summary
[0004]本申请提供一种轻量级大语言模型社会模拟方法及装置,能够解决轻量级模型在社会模拟中行为漂移及误差不可控累积的问题
本申请实施例提供的轻量级大语言模型社会模拟方法,通过获取目标事件及与目标事件相关联的真实行为序列,按时间顺序将智能体分批进行模拟,智能体用于针对目标事件生成模拟行为文本;对于当前批次中的每一个智能体,获取智能体的决策相关信息;将决策相关信息输入预评估模型,由预评估模型对当前智能体的行为进行预测,输出一个结构化行为状态向量;结构化行为状态向量包含多个行为维度的类别信息;将结构化行为状态向量转化为文本形式的生成约束;将生成约束与决策相关信息的至少一部分进行组合后,输入大语言模型,由大语言模型在生成约束的约束下生成当前智能体的模拟行为文本;根据当前批次生成的全部模拟行为文本更新群体状态和各智能体的历史发言记录,将更新后的群体状态用于后续批次的模拟。
Smart Images

Figure CN122840106A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of multi-agent social simulation technology, and in particular to a lightweight method and apparatus for social simulation of large language models. Background Technology
[0002] Social simulation studies individual behavior and group interactions through computational modeling, and has significant value in public opinion analysis and public policy deduction. In recent years, large language model-driven agents have provided a new technical path for social simulation: compared with traditional rule-driven agent modeling (ABM), language model agents can understand the event context and user profiles described in natural language, and generate text that closely resembles the expressions of real users.
[0003] Existing social simulation methods (such as OASIS and Generative Agents) allow agents to receive natural language context such as event text, user profiles, historical posts, and popular comments, with the language model directly generating comments. Mean-Field LLM further introduces mean-field theory, using group-level natural language summaries to update group dynamics, forming a closed loop of "individual decision-making—group dynamics update—re-decision-making" to reduce the computational complexity of large-scale interactions. However, existing technologies still have the following shortcomings: when the simulation scale increases and the model needs to be called frequently under limited computing power, the inference cost of large-parameter models is high, while lightweight models for multi-agent systems have not been specifically designed. Specifically, existing methods concatenate various social signals such as event topics, group mean fields, and popular comments into unstructured natural language as prompt input. When lightweight models face fuzzy summaries, they need to infer stance and emotion on their own, which can easily lead to unstable behavioral directions. The one-step "context-comment" generation coupled behavioral decisions and language expression in the same model call makes lightweight models prone to behavioral deviations, and there is no intermediate state to check before generation. Summary of the Invention
[0004] This application provides a lightweight large language model social simulation method and apparatus, which can solve the problems of behavioral drift and uncontrollable error accumulation in lightweight models in social simulation.
[0005] To achieve the above objectives, this application adopts the following technical solution: In a first aspect, this application provides a lightweight large language model social simulation method, characterized in that the method includes: Acquire a target event and the real behavior sequence associated with the target event, and simulate the intelligent agent in batches according to time sequence. The intelligent agent is used to generate simulated behavior text for the target event. For each agent in the current batch, obtain the decision-related information of that agent; The decision-related information is input into the pre-evaluation model, which predicts the behavior of the current agent and outputs a structured behavior state vector; the structured behavior state vector contains category information of multiple behavior dimensions; The structured behavior state vector is transformed into a textual form of generation constraints; The generation constraints are combined with at least a portion of the decision-related information and then input into a large language model, which generates simulated behavioral text of the current agent under the constraints of the generation constraints. The group state and the historical speech records of each agent are updated based on all simulated behavioral texts generated in the current batch, and the updated group state is used for simulation in subsequent batches.
[0006] In one embodiment, the real behavior sequence is a chain of comments on a social media platform, and the simulated behavior text is a simulated comment; In the batch simulation of the intelligent agents, the number of intelligent agents in each batch is N, where N is a preset positive integer; the first W time steps of the simulation are the warm-up stage, and the formal simulation stage begins from the (W+1)th time step. The decision-related information includes at least one of the following: the original post text of the event, a summary of the group situation at the current time step, popular comments before the current time step, the user profile of the agent, and the historical speaking records of the agent; The behavioral dimensions in the structured behavioral state vector include: behavioral type, stance, degree of belief, emotional state, emotional tendency, intention category, subjectivity / objectivity, expression style, expression intensity, and information source type; each dimension is selected from its respective predefined candidate category set. The group state is a group situation summary in natural language form, which includes at least a description of the group's dominant position, sentiment distribution, and discussion focus.
[0007] In one embodiment, converting the structured behavior state vector into a text-based generation constraint includes: According to the preset hierarchy, each behavior dimension in the structured behavior state vector is divided into at least two priority levels, and the behavior dimensions of each level are mapped to the corresponding generation constraints in order of priority from high to low. When there is a conflict between generation constraints at different levels, the generation constraints corresponding to the level with higher priority shall prevail.
[0008] In one embodiment, the preset hierarchy includes: The first level consists of four dimensions: behavior type, stance, degree of belief, and emotional tendency, which are mapped to the core requirement class to generate constraints. The second level consists of two dimensions: emotional state and intention category, which are mapped to supplementary requirement classes to generate constraints. The third level consists of four dimensions: subjectivity and objectivity, expression style, expression intensity, and information source type, which are mapped to style class generation constraints. The first level has a higher priority than the second level, and the second level has a higher priority than the third level.
[0009] In one embodiment, the pre-evaluation model is a multi-task classification model consisting of a Transformer-based encoder and multiple parallel classification heads; The encoder encodes the input decision-related information to obtain a shared semantic representation vector; the multiple parallel classification heads share the shared semantic representation vector, each classification head corresponds to a behavior dimension, and outputs the category prediction result of the corresponding dimension; the outputs of each classification head together constitute the structured behavior state vector. The training of the pre-evaluation model includes: during the training phase, randomly masking at least one of the group situation summary, local popular comments, and historical speaking records in the decision-related information, and inputting the masked decision-related information as training samples into the pre-evaluation model; The random masking includes at least one of the following: randomly deleting some short sentences in the group situation summary, randomly deleting some comment lines in popular comments, and randomly deleting some comments in the historical comment records.
[0010] In one embodiment, the training loss function of the pre-evaluation model is:
[0011] Where N is the batch size, K is the total number of behavioral dimensions, and CE is the cross-entropy loss function. For the i-th sample k VI's true label For the i-th sample k Dimension's predicted label, For the first k The weighting coefficients of the dimension; The weighting coefficient The calculation method is as follows:
[0012] in, For the first k Preset importance weights for dimensions For the first k The frequency inverse weight of category c in dimension; The frequency inverse weight The calculation method is as follows:
[0013] in, For the first k The total number of training samples in dimension, For the first k The total number of categories of dimensions For the first k The number of training samples for class c in dimension. This is the smoothing coefficient.
[0014] In one embodiment, before combining the generation constraints with at least a portion of the decision-related information, the method further includes: Determine whether the structured behavior state vector satisfies the preset short comment triggering condition; If satisfied, the corresponding preset comment text is directly retrieved from the preset short comment library as the simulated behavior text of the current agent, without performing the step of combining the generation constraints with at least a portion of the decision-related information; If not satisfied, then the step of combining the generation constraints with at least a portion of the decision-related information and inputting it into the large language model is performed.
[0015] In one embodiment, updating the group state based on all simulated behavioral texts generated in the current batch includes: The group state before the update and all simulated behavior texts generated in this batch are input into the large language model. The large language model integrates the group state before the update and the newly added simulated behavior texts in this batch to output the updated group state.
[0016] In one embodiment, the method further includes: The simulated behavior text is input into the semantic evaluation model, which predicts the behavior dimension of the simulated behavior text and outputs a behavior state annotation vector; the behavior state annotation vector and the structured behavior state vector are in the same behavior label space; Calculate a pre-evaluation accuracy value, which is used to measure the consistency between the structured behavior state vector and the actual behavior state label; Calculate the prediction-generation consistency value, which is used to measure the consistency between the structured behavior state vector and the behavior state index vector; When the pre-evaluation accuracy value is lower than the first preset threshold, it is determined that the simulation error originates from the pre-evaluation model; When the prediction-generation consistency value is lower than the second preset threshold, it is determined that the simulation error originates from the failure of the generation of the large language model.
[0017] A second aspect of this application provides a lightweight large language model social simulation device, the device comprising: The first acquisition module is used to acquire a target event and a sequence of real behaviors associated with the target event, and to simulate the intelligent agent in batches according to time order. The intelligent agent is used to generate simulated behavior text for the target event. The second acquisition module is used to acquire decision-related information of each agent in the current batch; The prediction module is used to input the decision-related information into the pre-evaluation model, which then predicts the behavior of the current agent and outputs a structured behavior state vector. The structured behavior state vector contains category information for multiple behavior dimensions. A conversion module is used to convert the structured behavior state vector into a text-based generation constraint; The simulation module is used to combine the generation constraints with at least a portion of the decision-related information and input it into a large language model, so that the large language model can generate simulated behavior text of the current agent under the constraints of the generation constraints. The update module is used to update the group state and the historical speech records of each agent based on all simulated behavioral texts generated in the current batch, and to use the updated group state for simulation in subsequent batches.
[0018] A third aspect of this application provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, it implements the lightweight large language model social simulation method described in the first aspect of this application.
[0019] A fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the lightweight large language model social simulation method described in the first aspect of this application.
[0020] The beneficial effects of the technical solutions provided in this application include at least the following: The lightweight large language model social simulation method provided in this application obtains a target event and the real behavior sequence associated with the target event, and simulates agents in batches according to time sequence. The agents generate simulated behavior text in response to the target event. For each agent in the current batch, decision-related information of the agent is obtained. The decision-related information is input into a pre-evaluation model, which predicts the behavior of the current agent and outputs a structured behavior state vector. The structured behavior state vector contains category information of multiple behavior dimensions. The structured behavior state vector is converted into generation constraints in text form. The generation constraints are combined with at least a part of the decision-related information and input into a large language model. The large language model generates the simulated behavior text of the current agent under the constraints of the generation constraints. The group state and the historical speech records of each agent are updated according to all the simulated behavior text generated in the current batch, and the updated group state is used for simulation in subsequent batches.
[0021] This application explicitly predicts behavior label vectors (containing multiple dimensions such as behavior type, stance, and emotional tendency) in the pre-evaluation model and transforms these label vectors into explicit generation constraints, thereby locking in the agent's behavioral direction before text generation. The lightweight large language model no longer needs to infer "which stance and emotion to use" during the execution phase; it only needs to complete the language expression within the given constraints. Experimental results show that compared to existing direct generation methods, the JS divergence of this application's method is reduced from 0.227 to 0.133 (a reduction of 41.5%), and the Trajectory Score is increased from 0.750 to 0.842 (an improvement of 12.3%), demonstrating that the introduction of the structured behavioral state intermediate layer in this application effectively reduces behavioral drift.
[0022] Furthermore, this application generates structured behavior state vectors, making the behavior pre-evaluation results an observable, recordable, and verifiable intermediate decision-making interface. The real behavior labels, the pre-evaluation model's predicted labels, and the generated comment label are all located in the same behavior label space. This allows for the quantification of pre-evaluation accuracy (consistency between real and predicted labels) and prediction-generation consistency (consistency between predicted and comment labels), thus decomposing the simulation error into two independent components: "behavior pre-evaluation error" and "generation compliance error." If the former is low, the problem lies in the pre-evaluation model; if the latter is low, the problem lies in the generation and execution stage of the lightweight large language model, eliminating the need for manual review of each comment text to infer the cause of the failure. Attached Figure Description
[0023] Figure 1 A flowchart illustrating a lightweight large language model social simulation method provided in this application embodiment; Figure 2A structural diagram of a lightweight large language model social simulation device provided in this application embodiment; Figure 3 This is a schematic diagram of the internal structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0024] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0025] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of this disclosure, unless otherwise stated, "a plurality of" means two or more.
[0026] In addition, the use of “based on” or “according to” implies openness and inclusivity, because processes, steps, calculations or other actions “based on” or “according to” one or more conditions or values can in practice be based on additional conditions or values beyond those conditions.
[0027] This application provides a lightweight social simulation method using a large language model, illustrating the simulation scenario with a public event comment chain on a social media platform. However, those skilled in the art will understand that this method is also applicable to social simulation scenarios with temporal behavioral sequences, such as forum discussion areas and news comment sections. Figure 1 As shown, the method includes the following steps: Step 101: Obtain the target event and the real behavior sequence associated with the target event, and simulate the intelligent agent in batches according to the time sequence. The intelligent agent is used to generate simulated behavior text for the target event. Before starting the simulation, we first obtain the target event and the sequence of real-world behaviors associated with it. Taking a social media platform as an example, the target event T is an original post containing an event title and a body description, such as "A large-scale delay occurred on Metro Line 2 in a certain city during the morning rush hour, with many passengers stranded on the platform; no official statement has been released yet." The sequence of real-world behaviors is a chain of user comments related to this event, arranged chronologically by posting time, represented as... M represents the total number of comments, arranged in chronological order. This application uses the target event "A large-scale delay occurred on Metro Line 2 in a certain city during the morning rush hour, with a large number of passengers stranded on the platform, and no official statement has yet been released" as an example for illustration.
[0028] An agent is a virtual entity that simulates the behavior of users on a social network. Each agent is used to generate simulated comments in response to a target event. Specifically, agents are simulated in batches according to time sequence, and the number of agents in each batch is denoted as N (N=16 in this embodiment).
[0029] Step 102: For each agent in the current batch, obtain the decision-related information of the agent.
[0030] For each agent i in the current batch, the system constructs a decision context for it. Decision context It includes the following five types of information: the original post text T, a summary of the group's situation at the current time step, and so on. Popular comments before the current time step User profile of the intelligent agent and the agent's historical speech records In the embodiments of this application, the methods for obtaining each component are as follows.
[0031] The original event post text T is a fixed input, containing the event title and body description, and is shared by all agents.
[0032] Group Situation Summary This is a natural language description of the public opinion atmosphere of the group up to the current time step t, generated by the mean field update module after the previous batch of simulations is completed. In the first time step, the group situation summary can be set to the initial default description, such as "No comments yet, group attitude unknown". After the warm-up phase, the group situation summary contains sufficiently rich information, such as "Early comments are mainly questions and information seeking, some users express dissatisfaction, but there are no clear voices questioning the official stance, and the main emotion is anxiety".
[0033] Popular comments in certain sections The top K comments in terms of interaction volume (including at least one of likes, replies, and shares; in this embodiment, the weighted sum of the three) among those generated before the current time step are represented by K, where K is a preset positive integer (K=3 in this embodiment). These popular comments are used to simulate the impact of high-exposure content on subsequent user behavior in real-world scenarios.
[0034] Specifically, the calculation method for the popularity score of the local popular comments is as follows:
[0035] Reposts, Comments, and Attitudes represent the number of reposts, comments, and likes for each comment, respectively.
[0036] User Profile : This is the set of static attributes for agent i, including but not limited to the number of followers, total number of posts, and registration years, such as "ordinary user, 120 followers, registered in 2020, 600 posts". The user profile is pre-configured for each agent before the simulation begins. It can be randomly generated to simulate different user types, or it can be sampled based on the distribution of real user data.
[0037] Historical speech record : This represents the L most recent simulated comments generated by agent i before the current time step. In this embodiment, L=3. For example, if the user posted a comment "Hope traffic can be restored as soon as possible" in the first round, this information will be retained as their historical comment record. The method for updating the historical comment record will be detailed in subsequent steps.
[0038] The above five categories of decision-related information are serialized into plain text format according to a preset fixed field order. That is, to obtain decision-related information for each agent. During serialization, it is necessary to ensure that the serialized data does not contain real comments that have not yet been published in this batch or simulated comments that have just been generated in this batch and have not yet entered the next time step, so as to ensure the temporal causal consistency during the simulation process.
[0039] Step 103: Input the decision-related information into the pre-evaluation model, and the pre-evaluation model predicts the behavior of the current agent and outputs a structured behavior state vector; the structured behavior state vector contains category information of multiple behavior dimensions.
[0040] Serialized decision-related information Input pre-evaluation model The pre-evaluation model predicts the agent's behavior at the current time step and outputs a structured behavior state vector. This structured behavior state vector contains category information across multiple behavior dimensions. In this embodiment, it specifically includes multiple behavior dimensions and is generated, observable, and recorded before the simulated behavior text is generated. This design makes the behavior pre-evaluation result an intermediate decision-making interface that can be checked, recorded, and intervened upon before generation, rather than hiding the behavior judgment within the internal reasoning process of the language model as in existing technologies.
[0041] The pre-evaluation model is structured as follows: one BERT encoder and ten parallel softmax classification heads. The BERT encoder processes the input serialized text. Encode the hidden vector at position [CLS]. As a shared semantic representation, the ten classifiers share this latent vector. Each classifier corresponds to a behavior dimension and outputs the category probability distribution for that dimension. The category with the highest probability is taken as the predicted value for that dimension. The outputs of all classifiers are concatenated to form a ten-dimensional behavior label prediction vector. .
[0042] In this embodiment, taking an agent at the 101st time step of a subway delay event as an example, the group situation summary in its decision context is "Most comments question the delayed official response, doubt the authenticity of the information, and are predominantly negative." Popular comments include "Delayed again? It was the same rhetoric last time [angry]" and "Wait for the official announcement, don't spread rumors." The user profile is "Ordinary user, 320 followers, registered in 2019," and the historical comment record is "Hope the reason can be announced as soon as possible." The ten-dimensional behavioral labels output by the pre-evaluation model for this agent are: {Behavior type = comment, stance = opposition, degree of belief = doubt, emotional state = anger, emotional tendency = negative, intention category = expressing opinion, subjectivity = subjective, expression style = rhetorical question, expression intensity = high, information source type = none}.
[0043] This structured behavior state vector serves as an intermediate layer, rewriting the traditional one-step generation process of "context → comment" into a two-stage process of "context → structured behavior label → controlled comment generation". This allows the behavior judgment results to be recorded and verified separately before generation, for use in subsequent error diagnosis.
[0044] Step 104: Convert the structured behavior state vector into a text-based generation constraint.
[0045] Taking the aforementioned ten-dimensional tags as an example, the three-layer generation constraints after conversion are as follows: the core requirement is "questioning the authenticity of the news, expressing an opposing stance, and using a negative tone"; the supplementary requirement is "expressing opinions with anger"; and the style suggestion is "using rhetorical questions, using a strong tone, and expressing subjective opinions." Among the generation constraints at each layer, the behavioral direction in the core requirement (such as "expressing an opposing stance" and "questioning the authenticity of the news") must be strictly followed in the subsequent generation stage, while the style suggestion (such as "using rhetorical questions") is an optimization of the expression style without changing the core behavioral direction.
[0046] Step 105: Combine the generation constraints with at least a portion of the decision-related information and input them into a large language model, which then generates the simulated behavior text of the current agent under the constraints of the generation constraints.
[0047] The transformed text form is used to generate constraints and at least a portion of decision-related information to form a controlled generation prompt. In this embodiment, the controlled generated prompt includes the original event post text T and some popular comments. User profile Historical speech records And three layers of generation constraints (core requirements, supplementary requirements, and style suggestions). Group Situation Summary These are deliberately excluded from the controlled generation of prompts and used only in the behavior pre-evaluation stage to ensure that the simulated comments generated in this batch are not used for the pre-evaluation of the current time step, thus guaranteeing temporal causality.
[0048] For example, the specific content of the controlled generated prompt is as follows: "Target event: A large-scale delay occurred on Metro Line 2 in a certain city during the morning rush hour, with a large number of passengers stranded on the platform. No official statement has been released yet." Popular comments: "What's wrong with Line 2? I've been waiting for 20 minutes and there's still no movement," "Delayed again? It was the same last week. Can't you do something useful?" User profile: Regular user, 320 followers, registered in 2019; Historical statement: "We hope the reasons can be announced as soon as possible." Key requirements: Question the veracity of the information, express opposition, and use a negative tone; Additional requirement: Expressing opinions with anger; Style suggestion: Use rhetorical questions with a strong tone; Please generate a comment. Input the above controlled generated prompts into the large language model (In this embodiment, a lightweight large language model with 2B parameters is used.) The large language model generates the simulated behavior text of the current agent under the constraints of the generation constraints. .
[0049] Under the above input in this embodiment, the simulated comment output by the lightweight large language model is: "Another delay? They always say they're verifying it, can we still trust them?" This comment meets the core requirements of "expressing opposition" and "questioning the authenticity of the information," meets the supplementary requirements of "anger," and meets the style suggestions of "rhetorical question" and "strong tone." This shows that the controlled generation mechanism can effectively constrain the lightweight model to generate text within a predetermined behavioral direction, avoiding the stance drift and emotional distortion caused by the lightweight model's own inference of behavioral direction in traditional one-step methods.
[0050] Step 106: Update the group state and the historical speech records of each agent based on all the simulated behavior texts generated in the current batch, and use the updated group state for simulation in subsequent batches.
[0051] After all N agents in the current batch have generated simulated comments, perform a state update operation.
[0052] First, update the historical speech records of each agent. For each agent i, update its new comments generated at the current time step. Added to his / her history of statements In the middle, it is truncated into the nearest L (L=3 in this embodiment), that is .
[0053] Taking the aforementioned agent as an example, its updated historical speech record reads "Hope the reason can be announced as soon as possible" + "Another delay? They always say they're verifying it, can we still trust them?", preserving the emotional progression from expectation to disappointment. The role of individual memory updating is to ensure that the behavior of the same agent at different time steps has emotional coherence and logical consistency, avoiding behavioral personality splits caused by the agent starting from zero in each round.
[0054] Secondly, update the group state. The group state is a group situation summary in natural language form. It should at least include a description of the group's dominant stance, emotional distribution, and the focus of discussion. (The previous group situation summary should be updated.) With the complete set of simulated behavioral texts generated in this batch Input large language model (Mean field update function in this embodiment) (Implemented by a large language model), which integrates the group situation before the update with the newly added simulated comments in this batch, and outputs a summary of the updated group situation. In this embodiment, the group situation summary before the update was "Most comments questioned the delayed official response, doubted the authenticity of the information, and the sentiment was negative," while after the update it was "Questioning voices continue to increase, opposition and skepticism dominate, anger is still prominent, and the focus of discussion remains on the timeliness of the official response."
[0055] Updated group status The simulation is used for subsequent batches (time step t+1), serving as a component of the "group situation summary" in the decision context of each agent in the next round of simulation. The pre-evaluation at the current time step only uses... Do not use the simulated comments generated in this batch. This ensures temporal causal consistency during the simulation process, meaning that the pre-evaluation of an agent's behavior at each step is based solely on historical information available up to the current time step, without peeking at comments not yet published in the current batch. Through this update mechanism, the system forms a causal closed loop of "individual behavior → group situation update → influencing subsequent individual behavior," enabling the simulation to evolve continuously over hundreds of time steps.
[0056] The lightweight large language model social simulation method provided in this application determines the core behavioral attributes of the agent, such as stance, belief, emotion and intention, before text generation by using the structured behavioral state vector output by the behavioral pre-evaluation model. It also explicitly transforms these attributes into observable and recordable intermediate decision states, thus avoiding behavioral drift caused by the lightweight language model inferring behavioral direction from fuzzy natural language summaries.
[0057] Comparative experiments conducted on a test set containing 59 events and 11,170 comments demonstrate that the method in this application reduces the JS divergence from 0.227 to 0.133 (a reduction of 41.5%) and increases the Trajectory Score from 0.750 to 0.842 (an improvement of 12.3%) compared to the direct generation method in the prior art. This proves that the introduction of the structured behavioral state intermediate layer in this application effectively reduces behavioral drift.
[0058] This application uses a structured behavior state vector as the sole exit point for behavior judgment, placing the behavior pre-evaluation result, the actual behavior label, and the generated return label in the same behavior label space. This provides a quantifiable and comparable coordinate reference system for error tracing. When simulation deviations occur, the source of error can be located in the pre-evaluation stage or the generation execution stage by separately calculating the pre-evaluation accuracy and prediction-generation consistency, thus solving the problem of indistinguishable error sources in the one-step generation paradigm of existing technologies.
[0059] This application, through a dual update mechanism of group status and individual history, enables the simulation system to not only track the evolution of group public opinion at the macro level, but also ensure the continuity of individual behavior at the micro level, thus realizing a sustainable multi-round social simulation closed loop and providing a feasible computational path for long-term public opinion evolution analysis.
[0060] Optionally, the target event T includes metadata such as topic tags, content, publication time, and publisher information; the actual behavior sequence. It includes all user comments sorted by posting time. Each comment contains the comment text, posting time, poster information, and interaction volume (likes, replies, shares). Simulated comments are generated according to the chronological order of real comments. For example, the real comment chain spans 120 minutes from 08:00 to 10:00. The system divides the time into batches: 08:00-08:20 is batch 1 (16 simulated comments), 08:20-08:40 is batch 2 (16 simulated comments), and so on. The collective evolution trajectory of simulated comments in terms of behavioral dimensions such as stance, emotion, and intention is used as the evaluation object and compared with the corresponding trajectory of the real comment chain to verify the simulation effect.
[0061] In the agent batch simulation, the number of agents in each batch is N, where N is a preset positive integer. In this embodiment, N=16. The value of the batch size N affects the computational efficiency and statistical stability of group behavior: when N is too small (e.g., N=1), the computational efficiency is low and it is difficult to reflect the group interaction effect; when N is too large (e.g., N=100), the time granularity may be lost, and it is impossible to accurately reproduce the group attitude change that occurs in the comment chain within a short time window. Experiments have verified that N=16 can achieve a good balance between computational efficiency and statistical stability of group behavior.
[0062] It should be noted that the first W time steps of the simulation are the warm-up phase, and the formal simulation phase begins from the (W+1)th time step. In this embodiment, W=100, that is, the first 100 batches (a total of 1600 simulated comments) are the warm-up phase, and the formal simulation phase begins from the 101st batch.
[0063] The core purpose of the warm-up phase is to address the cold start problem in social simulations. Specifically, at the very beginning of the simulation, there are no comments or group situation summaries in the system. Empty or contains only the initial default value, local popular comments An empty set containing the historical speech records of each agent. It is also empty. Under these conditions, the first batch of agents cannot obtain any input information about the group environment. Their behavior pre-evaluation mainly relies on the original post of the event and user profiles, which is inconsistent with the situation in real scenarios where users can see existing comments. Through 100 time steps of warm-up, the system gradually accumulates a sufficient amount of initial comments. The group situation summary, the popular comment pool, and the historical speaking records of each agent all reach a state of relatively saturated information. When entering the formal simulation stage, the agents can make decisions based on relatively complete information about the public opinion environment. The comments generated in the warm-up stage do not participate in the final behavior trajectory evaluation, but are only used to construct the initial simulation environment.
[0064] The structured behavioral state vector includes ten behavioral dimensions: behavior_type, stance, belief_degree, sentiment_state, sentiment_tendency, intent_classification, subjectivity, expression_style, expression_intensity_level, and information_mode. Each dimension is selected from its predefined set of candidate categories. The definitions of the ten dimensions and their candidate category sets are shown in Table 1.
[0065] Table 1. Behavioral Dimensions in the Structured Behavioral State Vector
[0066] The aforementioned ten-dimensional behavioral labeling system comprehensively depicts the behavioral characteristics of a social media comment in seven aspects: "what to do" (behavior type), "what stance to take" (stance), "belief or disbelief" (degree of belief), "emotional state" (emotional state + emotional tendency), "why to participate" (intention), "how to express" (subjectivity and objectivity + expression style + intensity), and "based on" (information source). This provides a structured foundation for subsequent generation constraint transformation and error diagnosis.
[0067] Group status is a summary of group situation in natural language. The group situation summary should include at least a description of the group's dominant stance, sentiment distribution, and focus of discussion. The advantage of using natural language in the group situation summary is its good readability and comprehensibility, facilitating manual review of the simulated state and enabling large language models to directly understand and update it.
[0068] For example, a typical example of a group situation summary is as follows: Early stage of the pre-heating phase: "Currently, there are few comments, and the group's attitude has not yet formed a clear trend. Some users are asking about the details of the incident." Mid-stage of the pre-heating phase: "Approximately 60% of the comments express doubt or dissatisfaction, 20% of the comments call for rational waiting for official announcements, and 20% of the comments are neutral inquiries. The focus of discussion is on the cause of the incident and the timeliness of the official response." Formal simulation phase: "Voices of doubt continue to increase, opposition and skepticism dominate, anger remains prominent, and the focus of discussion remains on the timeliness of the official response." The updating of the group situation summary is performed by a large language model.
[0069] Optionally, the structured behavior state vector is converted into a text-based generation constraint, including: According to the preset hierarchy, each behavior dimension in the structured behavior state vector is divided into at least two priority levels, and the behavior dimensions of each level are mapped to the corresponding generation constraints in order of priority from high to low.
[0070] For example, there can be 3 preset levels, dividing the ten behavioral dimensions into three priority levels, as follows: The first level (core behavior layer) consists of four dimensions: behavior type, stance, degree of belief, and emotional tendency. These four dimensions collectively determine the basic direction of the agent's behavior—behavior type defines the nature of the behavior (posting a comment or forwarding), stance defines the agent's attitude towards the event (support, neutrality, or opposition), degree of belief defines the agent's judgment of the information's authenticity (belief, uncertainty, or doubt), and emotional tendency defines the overall emotional polarity of the behavior (positive, neutral, or negative). These four dimensions together constitute the "skeleton" of the agent's behavior and must be strictly followed.
[0071] The second level (participation style layer) consists of two dimensions: emotional state and intention category. Emotional state is the concretization of emotional tendencies (e.g., anger is a specific negative emotion), while intention category defines the agent's main purpose for participating in the discussion (e.g., asking questions, expressing opinions, calling for action, etc.). These two dimensions add specific behavioral postures to the "skeleton" defined in the first level.
[0072] The third level (expression style layer) consists of four dimensions: subjectivity / objectivity, expression style, expression intensity, and source type. These four dimensions define how the agent organizes language—expression style determines sentence structure (declarative, rhetorical, exclamatory, slogan, or ironic), expression intensity determines the intensity of the language, subjectivity / objectivity determines whether the language is based on facts or feelings, and the source type determines the type of basis for the comment. These four dimensions only modify the way language is expressed and must not overturn the behavioral direction already determined by the first and second levels.
[0073] In this embodiment, the specific mapping method from each level to the natural language generation constraints adopts a combination of predefined table lookup rules and template filling, specifically as follows: The first-level mapping concatenates the four dimensions of behavior type, stance, degree of belief, and emotional tendency as an index, and matches the corresponding natural language expression in a predefined rule base. For example, "behavior type = comment, stance = opposition, degree of belief = doubt, emotional tendency = negative" matches the rule "questioning the authenticity of the message, expressing an opposing stance, and having a negative tone", and the output is a constraint generated as a core requirement (Core) class.
[0074] The second-level mapping uses the concatenation of the emotion state and intent category dimensions as an index to match the corresponding natural expression in a predefined rule base. For example, "emotion state = anger, intent category = expressing opinion" matches the rule "expressing opinion with anger," and the output is a supplementary requirement (Aux) class constraint.
[0075] The third-level mapping uses a concatenated index of four dimensions—expression style, subjectivity / objectivity, expression intensity, and information source type—to match the corresponding natural expression in a predefined rule base. For example, "expression style = rhetorical question, subjectivity / objectivity = subjective, expression intensity = high, source type = none" matches the rule "rhetorical question format is available, strong tone, based on subjective feelings," and the output is a style-class generation constraint.
[0076] The mapping sequence described above is executed sequentially according to the priority order of Level 1 → Level 2 → Level 3. The mapping process at each level follows the principle of "higher-level anchoring, lower-level modification": once the Level 1 mapping is completed, the core behavioral direction is locked in; the mapping result of Level 2 must not conflict with Level 1 (e.g., if Level 1 sets "emotional tendency = negative," then Level 2's "emotional state" must be selected from negative emotion candidates such as anger, sadness, and fear, and "happy" cannot be selected); Level 3 only provides suggestions on expression methods and does not change the behavioral direction.
[0077] It should be noted that when there are conflicts between generation constraints at different levels, the generation constraints corresponding to the higher-priority level shall prevail. The specific methods for determining and resolving conflicts are as follows.
[0078] The types of conflict mainly include the following two: The first is directional conflict: the behavioral direction required by lower-level constraints contradicts the behavioral direction locked by higher-level constraints. For example, the first level (core requirement) sets "express an opposing stance with a negative tone," while the third level (style suggestion) sets "express in a lighthearted, joking, and positive manner"—the two directly conflict in terms of emotional inclination. In this case, the first level takes precedence, abandoning the "positive" emotional inclination and maintaining the "negative tone." In this embodiment, in any situation where there is a conflict with the first level in the two core dimensions of stance and emotional inclination, the first level shall prevail.
[0079] The second type is execution feasibility conflict: the specific behaviors required by lower-level constraints cannot be executed within the framework defined by higher-level constraints. For example, the second level sets "intention = questioning", but the first level sets "position = opposition and degree of belief = doubt"—the intention to ask a question can be neutral verification-oriented or questioning-oriented, and the two are not necessarily in conflict; however, if the specific execution methods required by the lower level are incompatible with the behavioral direction of the higher level, the conflicting part of the lower level should be adjusted or discarded.
[0080] In this embodiment, the priority rule for conflict resolution is: Level 1 (core requirements) > Level 2 (supplementary requirements) > Level 3 (style suggestions). That is, when any lower-level constraint conflicts with a Level 1 constraint in terms of stance, belief, or emotional inclination, the Level 1 constraint has absolute priority; when the Level 3 constraint conflicts with the Level 2 constraint in a scope that does not involve the core dimension of the Level 1 constraint, the Level 2 constraint takes precedence over the Level 3 constraint.
[0081] Through the aforementioned hierarchical mapping and conflict resolution mechanism, the core design goal of "behavioral direction locking" can be guaranteed during the transformation of generation constraints. That is, the core stance, belief and emotional tendency are "anchored" before entering the large language model, and the model cannot overturn or weaken the high-level established behavioral direction through low-level means such as style modification when performing generation.
[0082] Optionally, the pre-evaluation model is a multi-task classification model consisting of a Transformer-based encoder and multiple parallel classification heads. In this embodiment, the encoder adopts a BERT-based architecture, containing 12 Transformer encoding layers, 768-dimensional latent vectors, and 12 attention heads (i.e., the standard configuration of BERT-based-uncased), with a parameter size of 110M. The encoder provides serialization decision-related information about the input. Encode the hidden vector at position [CLS]. As a shared semantic representation vector The dimension is 768.
[0083] Multiple parallel classification heads share this shared semantic representation vector. In this embodiment, ten parallel classification heads are used, each corresponding to one of the ten behavioral dimensions. Each classification head is a linear layer (weight matrix). ,in It consists of the total number of categories in this dimension plus the softmax activation function, i.e. Each classifier head outputs the probability distribution of the corresponding dimension, and the class with the highest probability is taken as the predicted value for that dimension; the outputs of all classifier heads together constitute a structured behavior state vector. .
[0084] The training process for the BERT pre-evaluation model is as follows: Training Data Construction: Training samples are extracted from real comment propagation chains. The input to each training sample contains only contextual information prior to the appearance of the target comment (including the original post, a summary of the group's behavior before the target comment, popular comments before the target comment, user profiles, and the target user's historical posts), excluding the target comment itself and later comments from the same batch, to ensure temporal causal consistency in the training data. The training dataset contains approximately 900,000 samples, divided into training and test sets based on events. Comments from the same event do not appear in either the training or test sets (i.e., event-level division) to evaluate the model's generalization ability.
[0085] Training process: Serializing the training samples into text. The BERT encoder is input, and the [CLS] latent vector is then fed into ten classifier heads. Each classifier head outputs the class prediction for the corresponding dimension. The prediction results for each dimension are compared with the ground truth labels, either manually labeled or labeled with the assistance of a large-parameter language model. The multi-task cross-entropy loss is calculated, and the parameters of the BERT encoder and each classifier head are updated through backpropagation. The design of the specific training loss function is detailed in subsequent embodiments.
[0086] Optionally, the training loss function for the pre-evaluation model is a weighted sum of the multi-task cross-entropy losses:
[0087] Where N is the batch size, K is the total number of behavioral dimensions (K=10 in this example), and CE is the cross-entropy loss function. For the i-th sample k The true labels of the dimension (i.e., the true categories of manually labeled or large-parameter LLM-assisted labels). For the i-th sample k The predicted label of dimension (i.e., the category corresponding to the highest probability in the probability distribution output by the model). For the first k The weighting coefficients of the dimension.
[0088] Cross-entropy loss function The definition of is: for the th k The samples in dimension whose true class is c, ,in Predict the probability that the sample belongs to class c for the model.
[0089] The multi-task learning framework allows ten classification tasks to share the semantic representation of the same BERT encoder, thereby fully utilizing the correlations between different behavioral dimensions. For example, there is a natural correlation between emotional state (anger) and emotional tendency (negative), and there is also a correlation between emotional state (anger) and expression intensity (high)—sharing representations allows these related dimensions to provide learning signals to each other, improving the prediction accuracy of each dimension when the training samples are limited.
[0090] Weighting coefficient The calculation is performed by combining the dimensional importance coefficient with the inverse weight of category frequency.
[0091] in, For the first k Preset importance weights for dimensions For the first k The frequency inverse weight of class c in dimension (i.e., for a sample belonging to class c, its i-th frequency inverse weight is given by the frequency inverse weight of the sample belonging to class c). k The loss weight for a dimension is the inverse frequency weight of that category.
[0092] Preset importance weights The weighting principle is as follows: core behavioral dimensions (behavior type, stance, degree of belief, emotional tendency) have higher weights, participation method dimensions (emotional state, intention category) have medium weights, and expression style dimensions (subjectivity / objectivity, expression style, expression intensity, information source type) have lower weights. In this embodiment, The value can be: first level dimension Second level dimension The third level dimension .
[0093] Frequency Inverse Weight The calculation method is as follows:
[0094] in, For the first k The total number of training samples in dimension, For the first k The total number of categories of dimensions For the first k The number of training samples for class c in dimension. Smoothing coefficient (in this embodiment) , used to avoid (Time denominator is zero and to prevent excessive weighting of rare categories).
[0095] The core function of frequency inverse weighting is to address class imbalance in training data. In real-world commentary data, certain categories naturally occur more frequently than others: for example, "stance = neutral" may occur more frequently than "stance = support" or "stance = oppose," and "expression style = statement" may occur far more frequently than "expression style = irony" or "expression style = slogan." Without class weighting, the model will tend to prioritize learning high-frequency categories, significantly reducing the prediction accuracy for low-frequency categories. Frequency inverse weighting, by assigning higher loss weights to less frequent categories, forces the model to give more attention to low-frequency categories during training, thereby improving the model's prediction performance on a minority of categories.
[0096] The frequency inverse weights are calculated before each training round based on the class distribution of the current training set, and the weight values remain fixed during training. For each class in each dimension, its weight is inversely proportional to its frequency of occurrence in the training set: classes with higher frequencies have lower weights, and classes with lower frequencies have higher weights, with all weight values greater than 0. In the training loss, The value is dynamically selected based on the category to which the current sample belongs during each forward propagation, and then multiplied by the preset importance weight of that dimension. Thus, the actual loss weight of the k-th dimension of the sample is obtained.
[0097] Optionally, during the training phase of the pre-evaluation model, the input decision-related information is randomly masked: at least one of the following is randomly deleted with a preset probability: some short sentences in the group situation summary, some comment lines in the popular comments, and some comments in the historical comment records. The masked data is then used as training samples to input into the pre-evaluation model. Random masking enables the model to learn to stably predict behavioral states from incomplete or noisy contexts during training, improving its robustness in real-world simulation scenarios such as incomplete group situation summary information, large fluctuations in popular comments, or short user history.
[0098] Optionally, before combining the generation constraints with at least a portion of the decision-related information, the method further includes: Determine whether the structured behavior state vector satisfies the preset short comment triggering condition; if it does, directly call the corresponding preset comment text from the preset short comment library as the simulated behavior text of the current agent, skipping the step of calling the large language model; if it does not satisfy the condition, perform the step of combining the generation constraint with at least a part of the decision-related information and inputting it into the large language model.
[0099] In actual implementation, a pre-defined short comment rule matching step is included before combining at least a portion of the generated constraints and decision-related information. The core purpose of this step is as follows: Social media comments contain a large number of highly repetitive short comments (such as "forwarding Weibo", "[angry]", "support", "haha", etc.). For these comments, calling a large language model to generate them one by one is not only costly, but the model output is often standardized pre-defined text. Therefore, by using pre-defined short comment rule matching, the pre-defined comment text can be directly returned when the triggering condition is met, which can significantly reduce the number of calls to the large language model, save computing resources, and improve simulation efficiency.
[0100] Determine whether the structured behavior state vector meets the preset short comment trigger conditions. The trigger conditions are determined by the combination of the values of the behavior type dimension and the values of other dimensions, specifically including the following two types: Category 1: Behavior type is "Retweet". Regardless of the values of other dimensions, when the behavior type is "Retweet", the preset text "Retweet Weibo" or "Retweet" is returned directly. The system's preset short comment library stores multiple synonym variations for the "Retweet" entry (such as "Retweet Weibo", "Retweet Content", "Retweet", etc.). This embodiment uses "Retweet Weibo" by default.
[0101] The second category: multi-dimensional combination conditions. For example, if the behavior type is comment, the stance is supportive, the emotional tendency is positive, the intention category is expressing emotion, the expression intensity is low, and the expression style is exclamation, and the combined weight score is lower than a preset threshold, the preset text "Support" or "Support" will be returned directly. The trigger threshold for multi-dimensional combination conditions is calculated by the combined weight score of each dimension: when the score is lower than the preset threshold, a short comment is returned; when it is higher than or equal to the threshold, the normal controlled generation process begins.
[0102] If any of the above triggering conditions are met, the corresponding preset comment text is directly retrieved from the preset short comment library as the simulated behavior text for the current agent, skipping the step of calling the large language model. The short comment texts in the preset short comment library are indexed according to the triggering conditions, supporting fast matching. When the triggering conditions are met, step 104 is not executed, and the preset text is directly output. In this embodiment, the preset contents of the preset short comment library are shown in Table 2.
[0103] Table 2. Preset Contents of the Preset Short Review Library
[0104] If the preset short comment triggering conditions are not met, then the step of combining at least a portion of the generated constraint and decision-related information and inputting it into the large language model is executed, i.e., following the process of steps 104 and 105.
[0105] The default execution position for short comment rule matching is after the generation of the structured behavior state vector and before the invocation of the large language model. Since the structured behavior state vector has already been generated and can be observed and recorded, even if the invocation of the large language model is skipped by rule matching, this intermediate state is still fully recorded and can be used by the subsequent error diagnosis module.
[0106] Optionally, updating the group state based on all simulated behavior texts generated in the current batch specifically involves: inputting the group state before the update and all simulated behavior texts generated in this batch into a large language model, and having the large language model integrate the group state before the update and the newly added simulated behavior texts in this batch to output the updated group state.
[0107] Group state update function In this embodiment, it is implemented using a large language model. Specifically, the group situation summary before the update is... With the N simulated comments generated in this batch Assemble update prompts according to a preset template, input them into a large language model, and have it output the updated population situation summary. .
[0108] In this embodiment, the template for the update prompt is: "[Current Group Situation Summary] {MF_t}"
New Comments in This Batch
[0109] The updated summary should include: the dominant group position, the main sentiment distribution, and the focus of the discussion.
[0110] The abstract should be 50-80 words long and written in natural language. For example, taking a subway delay event as an example, the input for a certain update operation is: "Summary of the group situation before the update". "Most comments questioned the delayed official response, doubted the authenticity of the information, and were generally negative." This batch of 16 new comments included phrases such as "Another delay?", "Let's wait for the official announcement," "They just restored it saying it was a signal failure," and "Heh, signal failure is a universal excuse." The large language model outputs an updated group situation summary. "Questioning voices continue to increase, with opposition and skepticism dominating, and anger remaining prominent. The focus of discussion remains on the timeliness of official responses, and a few comments have begun to mention that the cause of the malfunction has been confirmed."
[0111] During the update process, the input to the large language model only contains a summary of the group situation before the update. And the newly added simulated comments in this batch It does not contain any information that has not yet been generated for the next time step, ensuring the consistency of information timing during the update process. Furthermore, the pre-evaluation for the current time step only uses... The simulated comments generated in this batch will not be used; that is, the comments generated in this batch will be used to generate [the simulated comments]. However, the pre-evaluation of each agent at the current time step has already been completed. Based on this foundation, the comments in this batch will not provide feedback on behaviors that affect the current time step. This design ensures the causal constraint of "not peeking into future information".
[0112] The group situation summary, as a natural language description of the distribution of group behavior, evolves continuously over time. During the warm-up phase, the group situation summary gradually transitions from "few comments, unclear situation" to a statistically significant distribution description. In the formal simulation phase, the group situation summary is continuously updated with each new batch of comments, recording the evolution of group opinion. (Updated group situation summary) It is written into the system's global state for use in constructing the decision context of each agent at the next time step (step t+1).
[0113] Optionally, since the structured behavior state vector can be observed and recorded before generation, it makes it possible to diagnose simulation errors. This embodiment provides an error diagnosis mechanism that splits the error sources into two independent components, "pre-evaluation error" and "generation compliance error," by sharing a behavior label space.
[0114] The method further includes: inputting the simulated behavioral text into a semantic evaluation model, whereby the semantic evaluation model predicts the behavioral dimension of the simulated behavioral text and outputs a behavioral state annotation vector; the behavioral state annotation vector and the structured behavioral state vector are in the same behavioral label space; calculating a pre-evaluation accuracy value, which measures the consistency between the structured behavioral state vector and the real behavioral state labels; calculating a prediction-generation consistency value, which measures the consistency between the structured behavioral state vector and the behavioral state annotation vector; when the pre-evaluation accuracy value is lower than a first preset threshold, determining that the simulation error originates from the pre-evaluation model; when the prediction-generation consistency value is lower than a second preset threshold, determining that the simulation error originates from the failure of the large language model's generation adherence.
[0115] In actual implementation, simulated behavioral text will be used. The input is a semantic evaluation model, which predicts the behavior dimension of the simulated behavior text and outputs a behavior state back label vector. The behavior state index vector and the structured behavior state vector They are in the same behavioral label space, that is, they also contain multiple behavioral dimensions (behavior type, stance, degree of belief, emotional state, emotional tendency, intention category, subjectivity and objectivity, expression style, expression intensity, information source type), and the candidate category set of each dimension is exactly the same.
[0116] The semantic evaluation model can adopt the same model structure as the BERT pre-evaluation model (BERT encoder + multiple parallel classification heads), or it can use a large-parameter language model (such as Qwen3.6-35B) to annotate the simulated behavioral text with behavioral dimensions. Regardless of the implementation method, the output format of the semantic evaluation model must be consistent with the output format of the BERT pre-evaluation model to ensure... and They are in the same behavioral label space and are comparable.
[0117] The core diagnostic principle of this embodiment is as follows: Real Behavior Status Label (i.e., the labels corresponding to real comments), and the structured behavioral state vectors predicted by the BERT pre-evaluation model. With behavioral state backscalar vector All three exist within the same behavioral label space, sharing a ten-dimensional behavioral label coordinate system. Based on this shared space, two key diagnostic indicators can be quantified separately.
[0118] Pre-evaluation accuracy: used to measure the structured behavior state vector With real behavior status tags The difference lies in whether the BERT pre-evaluation model correctly interprets social signals and makes accurate behavioral judgments. The specific calculation method is: dimensional comparison. and To determine consistency across behavioral dimensions, the proportion of consistent dimensions to the total number of dimensions can be used; alternatively, a weighted calculation can be performed based on the pre-defined importance weights of each dimension, giving core dimensions (such as stance and emotional inclination) a higher weight in the accuracy calculation. Pre-assessment accuracy reflects whether the error originates from the pre-assessment stage—low pre-assessment accuracy indicates that the BERT pre-assessment model failed to correctly infer the behavioral direction from decision-related information.
[0119] Prediction-Generation Consistency: A measure of structured behavioral state vectors With behavioral state backscalar vector The difference lies in whether the large language model adheres to the generation constraints set by the BERT pre-evaluation model during the generation phase. The specific calculation method is the same as above: dimensional comparison. and To determine consistency across behavioral dimensions, the proportion of consistent dimensions to the total number of dimensions can be used; alternatively, a weighted calculation can be performed based on the pre-defined importance weights of each dimension. Prediction-generation consistency reflects whether the error originates from the generation stage—if prediction-generation consistency is low, it indicates that although the BERT pre-evaluation model correctly predicted the behavioral direction, the large language model failed to correctly implement that behavioral direction into the text when generating the comment.
[0120] Based on the two diagnostic indicators mentioned above, the source of simulation error can be independently determined: When the pre-evaluation accuracy is lower than the first preset threshold, it is determined that the simulation error mainly originates from the pre-evaluation model. That is, the BERT pre-evaluation model fails to correctly infer the agent's behavioral direction from the decision-related information. Possible reasons include: insufficient decision-related information itself, insufficient training of the BERT pre-evaluation model, or inaccurate category prediction for a certain behavioral dimension.
[0121] When the prediction-generation consistency value is lower than the second preset threshold, it is determined that the simulation error mainly stems from the failure of the large language model to follow the generation rules. That is, although the BERT pre-evaluation model correctly predicted the behavior direction, the large language model failed to correctly implement that behavior direction into the text when generating comments. Possible reasons include: the expression of the generation constraints is not clear enough for the large language model, the instruction-following ability of this lightweight large language model is insufficient, or there is ambiguity in the expression of a certain priority level in the generation constraints.
[0122] The specific values of the first and second preset thresholds can be configured according to the actual application scenario and accuracy requirements. In this embodiment, both are set to 0.85.
[0123] Through the above error diagnosis mechanism, the true label BERT pre-evaluation model predicts labels With generating return index All three components share the same behavior label space, allowing for the quantification of pre-evaluation accuracy and prediction-generation consistency within a unified coordinate reference system. Pre-evaluation accuracy reveals biases at the "behavior understanding" level, while prediction-generation consistency reveals biases at the "behavior execution" level. High prediction-generation consistency coupled with low pre-evaluation accuracy indicates a bottleneck in the behavior pre-evaluation module, requiring optimization of the BERT pre-evaluation model's training strategy or input features. Conversely, high pre-evaluation accuracy coupled with low prediction-generation consistency suggests a bottleneck in the generator's failure to adhere to constraints, necessitating adjustments to the expression template for generation constraints or fine-tuning the instruction strategy of the lightweight large language model. Researchers can quickly pinpoint the performance bottleneck of the simulation system without relying on manual review of each comment text, enabling targeted optimization.
[0124] The lightweight large language model social simulation method provided in this application adopts a behavioral state intermediate layer. The core behavioral attributes of the agent, such as stance, beliefs, emotions, and intentions, are determined before the comment text is generated, rather than relying on a lightweight, large language model to extract fuzzy group situation summaries. Implicit inference of behavioral direction. Unlike Mean-Field LLM, which directly inputs the mean-field summary into the policy model and allows the policy model to complete behavioral judgment and text generation in one go, this application splits social simulation into two sequential stages: "behavior pre-evaluation" and "controlled generation." This allows the behavioral judgment results to be recorded and tested separately as independent and observable intermediate states. This design allows lightweight large language models to avoid the burden of complex social signal understanding and inference tasks, and only need to complete text expression under given behavioral constraints. This effectively solves the behavioral drift problem caused by insufficient semantic understanding capabilities of lightweight models in existing methods, while providing an interpretable intermediate decision interface for the system.
[0125] In this application, the real behavior label Behavioral pre-assessment model predicts labels With generating comment tag All three share the same ten-dimensional behavioral label space, providing a quantifiable and comparable coordinate reference system for error tracing. Based on this shared space, two independent diagnostic indicators can be calculated: pre-assessment accuracy (measuring the accuracy of the true label). With predictive labels (Differences between them) and prediction-generation consistency (measures of prediction labels) With the return label (Differences between them). If the prediction-generation consistency is high but the pre-evaluation accuracy is low, it indicates that the simulation error mainly comes from the behavior pre-evaluation stage, and the BERT pre-evaluation model failed to correctly infer the behavior direction from social signals; if the pre-evaluation accuracy is high but the prediction-generation consistency is low, it indicates that the simulation error mainly comes from the generation stage, and the lightweight large language model failed to effectively follow the behavior constraints set by the pre-evaluation model when generating comments.
[0126] This application further employs Oracle experiments (using real-behavior labels) Replace BERT pre-evaluation model for predicting labels The effectiveness of the above diagnostic logic was verified by the generation constraints: the Trajectory Score in the Oracle experiment reached 0.950, significantly higher than the 0.842 achieved in the actual operation of this invention. This confirms that the upper bound of the framework in this application is much higher than the actual level of the current BERT pre-evaluation model. The main room for improvement lies in the prediction accuracy of the pre-evaluation module, rather than the instruction compliance capability of the generator. This error diagnosis mechanism allows developers to quickly locate the performance bottleneck of the simulation system without relying on manual review of each comment text, providing a clear technical direction for subsequent targeted optimization.
[0127] This application provides a lightweight large language model social simulation device, such as Figure 2 As shown, the device includes: The first acquisition module 11 is used to acquire a target event and a sequence of real behaviors associated with the target event, and to simulate the intelligent agent in batches according to time order. The intelligent agent is used to generate simulated behavior text for the target event. The second acquisition module 12 is used to acquire decision-related information of each agent in the current batch; The prediction module 13 is used to input the decision-related information into the pre-evaluation model, which then predicts the behavior of the current agent and outputs a structured behavior state vector. The structured behavior state vector contains category information for multiple behavior dimensions. The conversion module 14 is used to convert the structured behavior state vector into a text-based generation constraint; Simulation module 15 is used to combine the generation constraints with at least a portion of the decision-related information and input it into a large language model, so that the large language model can generate simulated behavior text of the current agent under the constraints of the generation constraints. The update module 16 is used to update the group state and the historical speech records of each agent based on all the simulated behavior texts generated in the current batch, and to use the updated group state for simulation in subsequent batches.
[0128] In one embodiment, the real behavior sequence is a chain of comments on a social media platform, and the simulated behavior text is a simulated comment; The decision-related information includes at least one of the following: the original post text of the event, a summary of the group situation at the current time step, popular comments before the current time step, the user profile of the agent, and the historical speaking records of the agent; The behavioral dimensions in the structured behavioral state vector include: behavioral type, stance, degree of belief, emotional state, emotional tendency, intention category, subjectivity / objectivity, expression style, expression intensity, and information source type; each dimension is selected from its respective predefined candidate category set. The group state is a group situation summary in natural language form, which includes at least a description of the group's dominant position, sentiment distribution, and discussion focus.
[0129] In one embodiment, the conversion module 14 is specifically used for: According to the preset hierarchy, each behavior dimension in the structured behavior state vector is divided into at least two priority levels, and the behavior dimensions of each level are mapped to the corresponding generation constraints in order of priority from high to low. When there is a conflict between generation constraints at different levels, the generation constraints corresponding to the level with higher priority shall prevail.
[0130] In one embodiment, the preset hierarchy includes: The first level consists of four dimensions: behavior type, stance, degree of belief, and emotional tendency, which are mapped to the core requirement class to generate constraints. The second level consists of two dimensions: emotional state and intention category, which are mapped to supplementary requirement classes to generate constraints. The third level consists of four dimensions: subjectivity and objectivity, expression style, expression intensity, and information source type, which are mapped to style class generation constraints. The first level has a higher priority than the second level, and the second level has a higher priority than the third level.
[0131] In one embodiment, the pre-evaluation model is a multi-task classification model consisting of a Transformer-based encoder and multiple parallel classification heads; The encoder encodes the input decision-related information to obtain a shared semantic representation vector; The multiple parallel classification heads share the shared semantic representation vector, each classification head corresponds to a behavior dimension, and outputs the category prediction result of the corresponding dimension; The outputs of each classification head together constitute the structured behavior state vector; The training of the pre-evaluation model includes: during the training phase, randomly masking at least one of the group situation summary, local popular comments, and historical speaking records in the decision-related information, and inputting the masked decision-related information as training samples into the pre-evaluation model; The random masking includes at least one of the following: randomly deleting some short sentences in the group situation summary, randomly deleting some comment lines in popular comments, and randomly deleting some comments in the historical comment records.
[0132] In one embodiment, the training loss function of the pre-evaluation model is:
[0133] Where N is the batch size, K is the total number of behavioral dimensions, and CE is the cross-entropy loss function. For the i-th sample k VI's true label For the i-th sample k Dimension's predicted label, For the first k The weighting coefficients of the dimension; The weighting coefficient The calculation method is as follows:
[0134] in, For the first k Preset importance weights for dimensions For the first k The frequency inverse weight of category c in dimension; The frequency inverse weight The calculation method is as follows:
[0135] in, For the first k The total number of training samples in dimension, For the first k The total number of categories of dimensions For the first k The number of training samples for class c in dimension. This is the smoothing coefficient.
[0136] In one embodiment, the conversion module 14 is further configured to: Determine whether the structured behavior state vector satisfies the preset short comment triggering condition; If the conditions are met, the corresponding preset comment text is directly retrieved from the preset short comment library as the simulated behavior text of the current agent, skipping the step of calling the large language model; If not satisfied, then the step of combining the generation constraints with at least a portion of the decision-related information and inputting it into the large language model is performed.
[0137] In one embodiment, the update module 16 is further configured to: The group state before the update and all simulated behavior texts generated in this batch are input into the large language model. The large language model integrates the group state before the update and the newly added simulated behavior texts in this batch to output the updated group state.
[0138] In one embodiment, the prediction module 13 is further configured to: The simulated behavior text is input into the semantic evaluation model, which predicts the behavior dimension of the simulated behavior text and outputs a behavior state annotation vector; the behavior state annotation vector and the structured behavior state vector are in the same behavior label space; Calculate a pre-evaluation accuracy value, which is used to measure the consistency between the structured behavior state vector and the actual behavior state label; Calculate the prediction-generation consistency value, which is used to measure the consistency between the structured behavior state vector and the behavior state index vector; When the pre-evaluation accuracy value is lower than the first preset threshold, it is determined that the simulation error originates from the pre-evaluation model; When the prediction-generation consistency value is lower than the second preset threshold, it is determined that the simulation error originates from the failure of the generation of the large language model.
[0139] The lightweight large language model social simulation device provided in this application embodiment can execute the above-described lightweight large language model social simulation method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0140] Specific limitations regarding the lightweight large language model social simulation device can be found in the limitations of the lightweight large language model social simulation method described above, and will not be repeated here. Each module in the aforementioned lightweight large language model social simulation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor of the electronic device in hardware form, or stored in the memory of the electronic device in software form, so that the processor can call and execute the corresponding operations of each module.
[0141] The lightweight large language model social simulation method provided in this application embodiment can be executed by an electronic device, which can be a processor, processing chip, computer equipment, terminal equipment, server or server cluster. This application embodiment does not specifically limit this.
[0142] Figure 3 This is a schematic diagram of the internal structure of an electronic device provided in an embodiment of this application. Figure 3 As shown, the electronic device includes a processor and a memory connected via a system bus. The processor provides computational and control capabilities. The memory may include a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. These computer programs can be executed by the processor to implement the steps of the lightweight large language model social simulation method provided in the various embodiments above. The internal memory provides a cached runtime environment for the operating system and computer programs in the non-volatile storage medium.
[0143] Those skilled in the art will understand that Figure 3 The diagram shown is an internal structure diagram of an electronic device, but it is only a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. Specific electronic devices may include those that are more advanced than those described above. Figure 3 The diagram shows more or fewer components, or combinations of certain components, or different component arrangements.
[0144] In another embodiment of this application, a computer-readable storage medium is also provided, on which a computer program is stored, wherein when the computer program is executed by a processor, the steps of the lightweight large language model social simulation method as described in the embodiments of this application are implemented.
[0145] In another embodiment of this application, a computer program product is also provided, which includes computer instructions that, when executed on a lightweight large language model social simulation device, cause the lightweight large language model social simulation device to perform each step of the lightweight large language model social simulation method in the method flow shown in the above method embodiment.
[0146] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software programs, implementation can be, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When these computer instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device containing one or more servers, data centers, etc., that can be integrated with the medium. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks (SSDs)).
[0147] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0148] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A lightweight social simulation method for large language models, characterized in that, The method includes: Acquire a target event and the real behavior sequence associated with the target event, and simulate the intelligent agent in batches according to time sequence. The intelligent agent is used to generate simulated behavior text for the target event. For each agent in the current batch, obtain the decision-related information of that agent; The decision-related information is input into the pre-evaluation model, which predicts the behavior of the current agent and outputs a structured behavior state vector; the structured behavior state vector contains category information of multiple behavior dimensions; The structured behavior state vector is transformed into a textual form of generation constraints; The generation constraints are combined with at least a portion of the decision-related information and then input into a large language model, which generates simulated behavioral text of the current agent under the constraints of the generation constraints. The group state and the historical speech records of each agent are updated based on all simulated behavioral texts generated in the current batch, and the updated group state is used for simulation in subsequent batches.
2. The method according to claim 1, characterized in that, The real behavior sequence is a chain of comments on a social media platform, and the simulated behavior text is a simulated comment; The decision-related information includes at least one of the following: the original post text of the event, a summary of the group situation at the current time step, popular comments before the current time step, the user profile of the agent, and the historical speaking records of the agent; The behavioral dimensions in the structured behavioral state vector include: behavioral type, stance, degree of belief, emotional state, emotional tendency, intention category, subjectivity and objectivity, expression style, expression intensity, and information source type. Each dimension is selected from its respective predefined set of candidate categories; The group state is a group situation summary in natural language form, which includes at least a description of the group's dominant position, sentiment distribution, and discussion focus.
3. The method according to claim 1, characterized in that, The process of converting the structured behavior state vector into textual generation constraints includes: According to the preset hierarchy, each behavior dimension in the structured behavior state vector is divided into at least two priority levels, and the behavior dimensions of each level are mapped to the corresponding generation constraints in order of priority from high to low. When there is a conflict between generation constraints at different levels, the generation constraints corresponding to the level with higher priority shall prevail.
4. The method according to claim 3, characterized in that, The preset hierarchy includes: The first level consists of four dimensions: behavior type, stance, degree of belief, and emotional tendency, which are mapped to the core requirement class to generate constraints. The second level consists of two dimensions: emotional state and intention category, which are mapped to supplementary requirement classes to generate constraints. The third level consists of four dimensions: subjectivity and objectivity, expression style, expression intensity, and information source type, which are mapped to style class generation constraints. The first level has a higher priority than the second level, and the second level has a higher priority than the third level.
5. The method according to claim 1, characterized in that, The pre-evaluation model is a multi-task classification model consisting of a Transformer-based encoder and multiple parallel classification heads; The encoder encodes the input decision-related information to obtain a shared semantic representation vector; the multiple parallel classification heads share the shared semantic representation vector, each classification head corresponds to a behavior dimension, and outputs the category prediction result of the corresponding dimension; the outputs of each classification head together constitute the structured behavior state vector. The training of the pre-evaluation model includes: during the training phase, randomly masking at least one of the group situation summary, local popular comments, and historical speaking records in the decision-related information, and inputting the masked decision-related information as training samples into the pre-evaluation model; The random masking includes at least one of the following: randomly deleting some short sentences in the group situation summary, randomly deleting some comment lines in popular comments, and randomly deleting some comments in the historical comment records.
6. The method according to claim 5, characterized in that, The training loss function of the pre-evaluation model is: Where N is the batch size, K is the total number of behavioral dimensions, and CE is the cross-entropy loss function. For the i-th sample k VI's true label For the i-th sample k Dimension's predicted label, For the first k The weighting coefficients of the dimension; The weighting coefficient The calculation method is as follows: in, For the first k Preset importance weights for dimensions For the first k The frequency inverse weight of category c in dimension; The frequency inverse weight The calculation method is as follows: in, For the first k The total number of training samples in dimension, For the first k The total number of categories of dimensions For the first k The number of training samples for class c in dimension. This is the smoothing coefficient.
7. The method according to claim 1, characterized in that, Before combining the generation constraints with at least a portion of the decision-related information, the method further includes: Determine whether the structured behavior state vector satisfies the preset short comment triggering condition; If satisfied, the corresponding preset comment text is directly retrieved from the preset short comment library as the simulated behavior text of the current agent, without performing the step of combining the generation constraints with at least a portion of the decision-related information; If not satisfied, then the step of combining the generation constraints with at least a portion of the decision-related information and inputting it into the large language model is performed.
8. The method according to claim 1, characterized in that, The step of updating the group state based on all simulated behavioral texts generated in the current batch includes: The group state before the update and all simulated behavior texts generated in this batch are input into the large language model. The large language model integrates the group state before the update and the newly added simulated behavior texts in this batch to output the updated group state.
9. The method according to claim 1, characterized in that, The method further includes: The simulated behavior text is input into the semantic evaluation model, which predicts the behavior dimension of the simulated behavior text and outputs a behavior state annotation vector; the behavior state annotation vector and the structured behavior state vector are in the same behavior label space; Calculate a pre-evaluation accuracy value, which is used to measure the consistency between the structured behavior state vector and the actual behavior state label; Calculate the prediction-generation consistency value, which is used to measure the consistency between the structured behavior state vector and the behavior state index vector; When the pre-evaluation accuracy value is lower than the first preset threshold, it is determined that the simulation error originates from the pre-evaluation model; When the prediction-generation consistency value is lower than the second preset threshold, it is determined that the simulation error originates from the failure of the generation of the large language model.
10. A lightweight large-scale language model social simulation device, characterized in that, The device includes: The first acquisition module is used to acquire a target event and a sequence of real behaviors associated with the target event, and to simulate the intelligent agent in batches according to time order. The intelligent agent is used to generate simulated behavior text for the target event. The second acquisition module is used to acquire decision-related information of each agent in the current batch; The prediction module is used to input the decision-related information into the pre-evaluation model, which then predicts the behavior of the current agent and outputs a structured behavior state vector. The structured behavior state vector contains category information for multiple behavior dimensions. A conversion module is used to convert the structured behavior state vector into a text-based generation constraint; The simulation module is used to combine the generation constraints with at least a portion of the decision-related information and input it into a large language model, so that the large language model can generate simulated behavior text of the current agent under the constraints of the generation constraints. The update module is used to update the group state and the historical speech records of each agent based on all simulated behavioral texts generated in the current batch, and to use the updated group state for simulation in subsequent batches.