Non-independent identically distributed data simulation method based on large model agent
Through the creator simulator based on big model agents, simulating the creator's behavior and information cognition process, the problem that existing simulators are difficult to capture the long-term dynamic changes of the recommendation system is solved, and the accuracy and effectiveness of long-term evaluation of the recommendation system is achieved.
Patent Information
- Application Number
- CN202510095786.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-22
AI Technical Summary
Existing recommendation system simulators are difficult to effectively capture the interactive behavior patterns between creators and recommendation systems, and cannot fully simulate the long-term dynamic changes of the recommendation platform, resulting in the inability to accurately evaluate the long-term impact of the recommendation algorithm.
The creator simulator based on large-model agents is adopted to initialize the creator's inherent characteristics and behavioral patterns through the personal data module, memory module, belief module and creation module, simulate the creator's feedback memory and creative memory, perceive the creator's information cognitive process under limited information, and generate content consistent with the real creator through the creative module combining fast and slow thinking.
It realizes effective simulation of the long-term dynamic changes of the recommendation system, enhances the long-term evaluation ability of the recommendation system, can more accurately capture the interactive behavior patterns between the creator and the recommendation system, and improves the accuracy of the long-term impact evaluation of the recommendation algorithm.
Smart Images

Figure CN120011636A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computers, and more specifically, to a non-independent and identically distributed data simulation enhancement method based on a large model intelligent agent. Background Art
[0002] Nowadays, recommendation systems have become an important part of online content platforms. As platforms pay more and more attention to long-term sustainable development, how to effectively evaluate the long-term impact and performance of recommendation algorithms has become an important issue. When we consider the long-term impact of recommendation algorithms on platforms, content creators are a role that cannot be ignored. Because they constantly reshape the platform's product pool by uploading new products, and their behaviors (such as leaving the platform, uploading products) are also indirectly affected by recommendation algorithms, thus affecting the long-term development of the platform. In addition, although online testing can be evaluated in an environment with creators, due to the high overhead of long-term online testing, simulators are considered to be a more efficient solution. Therefore, in order to more efficiently evaluate recommendation algorithms in the long term, it is crucial to build a simulator that simulates the interactive behavior between creators and platforms.
[0003] The key to evaluating the long-term impact of recommendation systems is to consider the interactive behavior patterns between creators and recommendation systems, which cannot be captured well by existing modeling methods. Currently, most simulators based on reinforcement learning do not model the behavior of creators and focus mainly on evaluating reinforcement learning algorithms. The few simulators that consider creator behavior cannot align with real-world creator behavior due to their rule-based approach. For the first time, the ability of large models to simulate human behavior is used to introduce memory, behavior and other modules to simulate user behaviors such as viewing and liking.
[0004] Although existing simulators have achieved excellent performance in user simulation, most of them ignore the modeling of creator behavior, making it impossible to fully capture the long-term impact of real-world recommendation systems on the platform. The few simulators that consider creator behavior have a certain gap with the strategic creation behavior of real-world creators due to their rule-based approach. Summary of the invention
[0005] The purpose of the embodiments of the present disclosure is to provide a non-independent and identically distributed data simulation enhancement method based on a large model agent. A creator simulator based on a large model agent is proposed to address the long-term evaluation problem in the recommendation system, so as to better help the creator agent understand the limited user feedback information and enhance its analysis and creation capabilities.
[0006] In general, a non-independent and identically distributed data simulation method based on a large model intelligent agent is provided. The intelligent agent consists of four modules: a personal profile module, a memory module, a belief module, and a creation module. The intelligent agent is deployed in a system environment through training, and the intelligent agent is accessed by a user terminal to obtain an output result. The operation of the user terminal constitutes a feedback memory of the feedback information input into the memory module. The personal profile module initializes the creator's social identity, creative intrinsic motivation, and creative activity through a real data set, controls the creative frequency of the creator simulator by the activity, and inputs the social identity and intrinsic motivation as text prompts into the creation module each time the creation is performed to assist the large model creation.
[0007] During the offline long-term evaluation of the recommendation system, it is very important to simulate the dynamic changes of the content (product) pool of the recommendation system. Such dynamic changes need to be achieved by continuously injecting simulated content data into the product pool. At the same time, such simulated content data needs to be independent and identically distributed with real-world data (text content, content categories), and at the same time, it needs to be consistent with the content created by real individual creators. Traditional recommendation simulators do not implement the simulation of such a dynamic product pool, resulting in the inability to simulate the long-term dynamic changes of the recommendation platform, which creates difficulties for the long-term evaluation of the recommendation system. Therefore, we proposed a large-model intelligent agent creator simulator to achieve such independent and identically distributed data simulation. The construction process of the intelligent agent is to first initialize the inherent characteristics of the creator through the personal data module; the memory module is used to reflect the feedback memory and creative memory of the real creator, wherein the creative memory retrieves the fast thinking module in the creative module according to the relevance and timeliness, and inputs the most recent creative instance in the feedback memory into the slow thinking module in the creative module; the belief module perceives the creator's information cognition update under limited user feedback through the feedback memory and the information update of the creative memory, and then serves as the input of the slow thinking module of the creative module; the creative module uses a combination of fast and slow thinking to restore the creative process of the real world creator, including a fast thinking module and a slow thinking module, and the processing result of the slow thinking module is applied to the fast thinking module, and finally the processing result is obtained and fed back to the user. Through such simulated creation, the constructed creator intelligent agent can effectively generate text content that is independent and identically distributed with the real platform data, and at the same time, it can well maintain the content and category consistency with the content created by the real creator. By injecting it into the recommendation simulator, it can effectively simulate the long-term dynamics of the recommendation system and enhance the long-term simulation and evaluation of the recommendation system.
[0008] The implementation of the profile module is as follows: a profile is initialized with a pre-collected real-world data set, and the profile includes three dimensions: social identity, intrinsic motivation, and creative activity;
[0009] Among them, social identity is obtained by analyzing the creator's creative history and basic information, intrinsic motivation is summarized using a large language model, and creative activity represents the average number of items created by each creator per day.
[0010] The memory module constructs two kinds of memory: feedback memory and creative memory.
[0011] The feedback memory is constructed by: expressing the feedback memory of creator c as Due to its position in the information asymmetry state, at the end of each time step n, the feedback memory It will be updated based on some of the user feedback information provided by the platform regarding historical creations.
[0012]
[0013] The creation memory Stores the historical creation product information of the creator agent.
[0014] The implementation method of the belief module is as follows: constructing skill beliefs and audience beliefs, where the skill belief represents the creator's confidence in their ability to create each type of product, which is defined as the proportion of each type of content they create. At the beginning of each time step n, the skill confidence of creator c for type g will be updated according to the creation memory;
[0015]
[0016] Audience beliefs represent the creator’s internal understanding and expectations of the preferences of each type of users. At the beginning of each time step n, the audience belief of creator c in type g will be calculated based on the information stored in the feedback memory. Updated based on user feedback;
[0017]
[0018] The implementation method of the creation module is as follows: the thinking process is divided into two stages: slow thinking is used for strategic planning and analysis, and fast thinking is based on experience and instinct to quickly generate content;
[0019] The slow thinking stage reflects that in each creation process, the user’s feedback on the most recently created product will directly affect the creator’s judgment on whether to continue or change his current creation strategy. At the beginning of each time step n, the creator receives three key factors as input: (1) the utility of the most recently created product, i.e., z i (n), (2) Skill beliefs and audience beliefs (3) Social identity and intrinsic motivation Subsequently, three key factors are input into the large language model for slow thinking through the designed prompt P1:
[0020]
[0021] The fast thinking stage is that after generating the analysis results, the creator agent discovers the generated content based on the analysis results. The generated content is mainly divided into four parts: product title, product type, product label and product description. Before creating content, the creator agent generates content based on the action From creating memories Search for creative experience To assist fast thinkers in their creative process;
[0022]
[0023] The training method is:
[0024] The Proximal Policy Optimization (PPO) algorithm is used to fine-tune the Creator Agent, and the reward formula is: After obtaining the final user-item prediction scores, pairwise ranking loss is used to optimize all trainable parameters θ.
[0025] The technical effects to be achieved by the embodiments of the present invention are:
[0026] A novel non-independent and identically distributed data simulator based on a large model agent is proposed to model the behavior under information asymmetry between the platform and the creator in the recommendation platform; the personal profile module is initialized using real platform data to restore the creative preferences of the creator agent, and feedback and creative memory are combined to assist the creative process of the creator agent; at the same time, inspired by relevant research in game theory and cognitive science, modules such as belief and creation are used to simulate the cognitive process of creators under limited information.
[0027] The present invention can be integrated into various offline recommendation simulation environments to evaluate various existing recommendation algorithms and has a wide range of application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The above and other objects and features of the present disclosure will become more apparent from the following description in conjunction with the accompanying drawings.
[0029] Figure 1 It is a schematic diagram showing the architecture of a non-independent and identically distributed data simulation enhancement method based on a large model intelligent agent according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0030] The following specific embodiments are provided to help the reader obtain a comprehensive understanding of the methods, devices and / or systems described herein. However, after understanding the disclosure of the present application, various changes, modifications and equivalents of the methods, devices and / or systems described herein will be clear. For example, the order of operations described herein is only an example and is not limited to those orders set forth herein, but can be changed as will be clear after understanding the disclosure of the present application, except for operations that must occur in a specific order. In addition, for greater clarity and simplicity, the description of features known in the art may be omitted.
[0031] The features described herein can be implemented in different forms and should not be construed as being limited to the examples described herein. Rather, the examples described herein have been provided to illustrate only some of the many possible ways to implement the methods, devices, and / or systems described herein, which will be clear after understanding the disclosure of the present application.
[0032] As used herein, the term "and / or" includes any one of the associated listed items and any combination of any two or more.
[0033] Although terms such as "first", "second", and "third" may be used herein to describe various members, components, regions, layers, or portions, these members, components, regions, layers, or portions should not be limited by these terms. Instead, these terms are only used to distinguish one member, component, region, layer, or portion from another member, component, region, layer, or portion. Therefore, without departing from the teachings of the examples described herein, the first member, first component, first region, first layer, or first portion referred to in the examples may also be referred to as the second member, second component, second region, second layer, or second portion.
[0034] In the specification, when an element (such as a layer, a region or a substrate) is described as being “on”, “connected to” or “coupled to” another element, the element may be directly “on”, “connected to” or “coupled to” another element, or one or more other elements may be present therebetween. Conversely, when an element is described as being “directly on”, “directly connected to” or “directly coupled to” another element, there may be no other elements present therebetween.
[0035] The terms used herein are only used to describe various examples and are not intended to limit the disclosure. Unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. The terms "comprise", "include" and "have" indicate the presence of the described features, quantities, operations, components, elements and / or combinations thereof, but do not exclude the presence or addition of one or more other features, quantities, operations, components, elements and / or combinations thereof.
[0036] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by a person of ordinary skill in the art to which the present disclosure belongs after understanding the present disclosure. Unless explicitly defined as such herein, terms (such as those defined in a general dictionary) should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and should not be interpreted in an idealized or overly formal manner.
[0037] Furthermore, in the description of examples, when it is considered that a detailed description of a well-known related structure or function would cause vague interpretation of the present disclosure, such a detailed description will be omitted.
[0038] Figure 1 It is a schematic diagram showing the architecture of a non-independent and identically distributed data simulation enhancement method based on a large model intelligent agent according to an embodiment of the present disclosure.
[0039] The creator simulator based on the large model agent consists of four modules: personal profile module, memory module, belief module, and creation module.
[0040] The personal data module initializes the inherent characteristics of the creator, the memory module reflects the feedback memory and creative memory of the real creator, the belief module perceives the creator's information cognition under limited user feedback, and the creation module uses a combination of fast and slow thinking to restore the creative process of real-world creators.
[0041] Profile Module
[0042] In this module, we use a pre-collected real-world dataset to initialize a profile, which is crucial to align the agent’s behavior with real human behavior. The profile consists of three elements: social identity, intrinsic motivation, and creative activity.
[0043] Social identity: As a social behavior, content creation is not only driven by economic interests, but also influenced by social identity, a factor that is often overlooked in behavior modeling. To address this problem, a large language model is used to identify the social identity of the creator by analyzing the creator's creation history and basic information.
[0044] Intrinsic motivation: Intrinsic motivation is a key factor that influences creators’ creative behavior. I use the Big Language Model to help summarize each creator’s intrinsic motivation, which is extracted from their creative history and frequency as part of their profile.
[0045] Creative activity: In addition to the two economic characteristics mentioned above, creative activity (i.e., the frequency of creation of creators) is also an important characteristic of creators. The collected dataset is used to initialize the intrinsic activity level of each creator. It represents the average number of items created by each creator per day.
[0046] Memory Module
[0047] Considering the two key pieces of information that creators prioritize and store in reality, we construct two types of memories for the creator agent: feedback memory and creation memory.
[0048] Feedback memory: Different from the memory modeling of user agents, the user feedback received by each creator is usually recorded and maintained on the platform for a long time. The feedback memory of creator c is represented as Due to its position in the information asymmetry state, at the end of each time step n, the feedback memory It will be updated based on some of the user feedback information provided by the platform regarding historical creations.
[0049]
[0050] Creative Memory: Creative Memory It stores the historical product information of the creator agent. In order to better reflect the human time decay memory mechanism, we introduced the power function forgetting rate.
[0051]
[0052] Before each creation, the creator will retrieve the most important and recent creative experience from the creative memory. i and r i is a normalized timeliness and importance score in the interval (0.0, 1.0), where a larger value indicates that the memory is more recent and important. We want to make it possible to recall important memories that were made long ago.
[0053] Belief Module
[0054] Creators’ beliefs are usually formed based on the initial information provided by the platform and are updated over time and with the acquisition of new information. Specifically, creators’ strategic creative behavior is mainly driven by two beliefs: skill beliefs and audience beliefs.
[0055] Skill belief: With limited information, creators have incomplete knowledge of item types. They will gradually acquire more information about types during the creation process and improve their proficiency in these types. Skill confidence represents the creator's confidence in their ability to create each type of item, which is defined as the proportion of each type of content they have created. At the beginning of each time step n, creator c's skill confidence for type g will be updated based on the creation memory.
[0056]
[0057] Audience Belief: Audience belief represents the creator’s internal understanding and expectation of the preferences of users in each category. At the beginning of each time step n, the audience belief of creator c in category g will be stored in the feedback memory according to Updated based on user feedback from .
[0058]
[0059] Creation Module
[0060] In this module, the fast and slow thinking mechanism is applied to give the creative self-agent the analytical and creative capabilities similar to humans. The thinking process consists of two stages: slow thinking for strategic planning and analysis, and fast thinking for rapid content generation based on experience and instinct.
[0061] Slow thinking: During each creation process, users’ feedback on the most recently created product will directly affect the creator’s judgment on whether to continue or change his current creation strategy. At the beginning of each time step n, the creator receives three key factors as input, which affect his creation strategy: (1) the utility of the most recently created product, i.e., z i (n), (2) Skill beliefs and audience beliefs and (3) social identity and intrinsic motivation These inputs are then fed into the large language model for slow thinking via the designed prompt P1:
[0062]
[0063] Quick thinking: After generating the analysis results, the creator agent will generate content based on these findings. The content is mainly divided into four parts: product title, product type, product tags and product description. Before creating content, the creator agent will generate content based on the action From creating memories Search for creative experience To assist fast thinkers in their creative process.
[0064]
[0065] Model training:
[0066] The Proximal Policy Optimization (PPO) algorithm is used to fine-tune the Creator Agent. By simulating the real-world cycle of creating, receiving rewards, analyzing, and creating again, our goals are to: (1) improve the Creator Agent's understanding of the Creator's limited information state and user feedback; and (2) enhance the Creator Agent's analysis and creation capabilities.
[0067] To avoid the large language model strategy π θ (parameter is θ) and the initial reference strategy If the deviation is too far, we follow and introduce the KL divergence penalty into the current reward function. Therefore, the final reward formula is:
[0068] After obtaining the final user-item prediction scores, we use pairwise ranking loss to optimize all trainable parameters θ:
[0069]
[0070] Although some embodiments of the present disclosure have been shown and described, it will be appreciated by those skilled in the art that modifications may be made to the embodiments without departing from the principles and spirit of the present disclosure, the scope of which is defined by the claims and their equivalents.
Claims
1. A non-independent and identically distributed data simulation method based on a large model agent, characterized in that: The intelligent agent consists of four modules: personal profile module, memory module, belief module, and creation module. The intelligent agent is deployed in the system environment through training. The intelligent agent is accessed by the user end and the output result is obtained. The operation of the user end will constitute the feedback memory of the feedback information input memory module; the personal profile module initializes the creator's social identity, creative intrinsic motivation and creative activity through the real data set, controls the creation frequency of the creator simulator with the activity, and inputs the social identity and intrinsic motivation as text prompts into the creation module to assist the creation of the large model each time; Input the text content, content category data of the real world, and the individual creative content of real creators; the construction process of the intelligent body first initializes the inherent characteristics of the creator through the personal data module; the memory module is used to reflect the feedback memory and creative memory of the real creator, wherein the creative memory retrieves the fast thinking module in the creative module according to the relevance and timeliness, and inputs the most recent creative instance in the feedback memory into the slow thinking module in the creative module; the belief module perceives the creator's information cognition update under limited user feedback through the feedback memory and the information update of the creative memory, and then serves as the input of the slow thinking module of the creative module; the creative module uses a combination of fast and slow thinking to restore the creative process of the real-world creator, including a fast thinking module and a slow thinking module, and the processing result of the slow thinking module is applied to the fast thinking module, and finally the processing result is obtained and fed back to the user, and the feedback content is the consistency between the content generated by the constructed creator intelligent body and the text content of the independent and identically distributed real platform data.
2. A non-independent and identically distributed data simulation method based on a large model agent as claimed in claim 1, characterized in that: The implementation of the profile module is as follows: a profile is initialized with a pre-collected real-world data set, and the profile includes three dimensions: social identity, intrinsic motivation, and creative activity; Among them, social identity is obtained by analyzing the creator's creative history and basic information, intrinsic motivation is summarized using a large language model, and creative activity represents the average number of items created by each creator per day.
3. A non-independent and identically distributed data simulation method based on a large model agent as claimed in claim 2, characterized in that: The memory module constructs two kinds of memory: feedback memory and creative memory.
4. A non-independent and identically distributed data simulation method based on a large model agent as claimed in claim 3, characterized in that: The feedback memory is constructed by: expressing the feedback memory of creator c as Due to its position in the information asymmetry state, at the end of each time step n, the feedback memory It will be updated based on some of the user feedback information provided by the platform regarding historical creations. The creation memory Stores the historical creation product information of the creator agent.
5. A non-independent and identically distributed data simulation method based on a large model agent as claimed in claim 4, characterized in that: The implementation method of the belief module is as follows: constructing skill beliefs and audience beliefs, where the skill belief represents the creator's confidence in their ability to create each type of product, which is defined as the proportion of each type of content they create. At the beginning of each time step n, the skill confidence of creator c for type g will be updated according to the creation memory; Audience beliefs represent the creator’s internal understanding and expectations of the preferences of users in each category. At the beginning of each time step n, the audience belief of creator c in category g will be calculated based on the information stored in the feedback memory. Updated based on user feedback; 6. A non-independent and identically distributed data simulation method based on a large model agent as claimed in claim 5, characterized in that: The implementation method of the creation module is as follows: the thinking process is divided into two stages: slow thinking is used for strategic planning and analysis, and fast thinking is based on experience and instinct to quickly generate content; The slow thinking stage reflects that in each creation process, the user’s feedback on the most recently created product will directly affect the creator’s judgment on whether to continue or change his current creation strategy. At the beginning of each time step n, the creator receives three key factors as input: (1) the utility of the most recently created product, i.e., z i (n), (2) Skill beliefs and audience beliefs (3) Social identity and intrinsic motivation Subsequently, three key factors are input into the large language model for slow thinking through the designed prompt P1: The fast thinking stage is that after generating the analysis results, the creator agent discovers the generated content based on the analysis results. The generated content is mainly divided into four parts: product title, product type, product label and product description. Before creating content, the creator agent generates content based on the action From creating memories Search for creative experience To assist fast thinkers in their creative process; 7. A non-independent and identically distributed data simulation method based on a large model agent as claimed in claim 6, characterized in that: The training method is: The Proximal Policy Optimization (PPO) algorithm is used to fine-tune the Creator Agent, and the reward formula is: After obtaining the final user-item prediction scores, pairwise ranking loss is used to optimize all trainable parameters θ.
Citation Information
Patent Citations
Writing conception auxiliary system and network system based on enhanced intelligence
CN113220901A
Rational thinking method
CN118821941A
Memory enhanced artist agent autonomous learning and creation system and method
CN119294470A
Method of and system for controlling the qualities of musical energy embodied in and expressed by digital music to be automatically composed and generated by an automated music composition and generation engine
US20190237051A1
Cited By
Novel large model autonomous reasoning system based on fast and slow thinking
CN120654821A