Model training method and device based on transaction decision and electronic equipment
Through the multi-model training method, global and internal state parameters are used to generate decision results, solving the problem of inaccurate decision-making in traditional marketing strategies, and achieving more efficient marketing strategy optimization and user attraction.
Patent Information
- Application Number
- CN202510865356.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-08-08
AI Technical Summary
Traditional marketing strategies rely on static data analysis and rule engines, which are difficult to adapt to dynamic changes in market and user behavior, resulting in inefficient marketing budgets, inaccurate decision-making, and ineffective users.
By obtaining model sets and sample user characteristics, using global state parameters, internal model state parameters and behavior rules to generate sample transaction decision results, determine the true selection results, and update the model set based on decision reward and punishment values, implementing multi-model training to improve decision adaptability and accuracy.
It improves the decision accuracy and adaptability of the target transaction model in a dynamic environment, optimizes the effectiveness of marketing strategies, and enhances the ability to respond to market and user behavior.
Smart Images

Figure CN120450086A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and in particular to a transaction decision-based model training method, device, and electronic device. Background Art
[0002] With the development of technology, user needs are becoming increasingly personalized, information dissemination is rapid, and market competition is fierce. Marketing strategies can help businesses stay competitive in this complex and changing market environment and attract and retain customers.
[0003] Traditional marketing strategies often rely on static data analysis and rule engines, making them difficult to adapt to dynamic changes in the market and user behavior. This leads to inefficient use of marketing budgets, inaccurate decision-making, and difficulty in attracting users. Summary of the Invention
[0004] The main purpose of this specification is to provide a model training method, device, and electronic device based on transaction decision-making, aiming to improve the adaptability and accuracy of the model-generated transaction decision results. The technical solution is as follows:
[0005] In a first aspect, the embodiments of this specification provide a model training method based on transaction decision-making, including:
[0006] Acquire a model set and sample user features of sample users; the model set includes a target transaction party model and at least one associated transaction party model;
[0007] Inputting the sample user characteristics into each model in the model set respectively, and generating a sample transaction decision result corresponding to each model based on the global state parameter, the sample user characteristics, the internal state parameter and behavior rules of each model;
[0008] Determining a real selection result of the sample user for the sample transaction decision result of the target transaction party model;
[0009] Determining a decision reward or penalty value for each of the models based on the actual selection result;
[0010] The model set is updated according to the objective function of each model, the decision reward and punishment value, and the sample transaction decision result, and the step of obtaining the model set and the sample user characteristics of the sample user is executed until the simulation loop end condition is reached, thereby obtaining a trained model set. The target transaction party model in the trained model set is used to generate a transaction decision result based on the target user characteristics of the target user.
[0011] In a second aspect, an embodiment of this specification provides a model training device based on transaction decision-making, including:
[0012] An acquisition unit, configured to acquire a model set and sample user features of sample users; the model set includes a target transaction party model and at least one associated transaction party model;
[0013] a decision unit, configured to input the sample user characteristics into each model in the model set, and generate a sample transaction decision result corresponding to each model based on a global state parameter, the sample user characteristics, an internal state parameter of each model, and a behavior rule;
[0014] A result determination unit, configured to determine a real selection result of the sample user for the sample transaction decision result of the target transaction party model;
[0015] A reward and punishment unit, configured to determine a decision reward and punishment value for each of the models based on the actual selection result;
[0016] The training unit is used to update the model set according to the objective function of each model, the decision reward and punishment value and the sample transaction decision result, and proceed to execute the step of obtaining the model set and the sample user characteristics of the sample user until the simulation loop end condition is reached, thereby obtaining a trained model set. The target transaction party model in the trained model set is used to generate a transaction decision result based on the target user characteristics of the target user.
[0017] In a third aspect, an embodiment of this specification provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the steps of the above method when executed by the processor.
[0018] In a fourth aspect, an embodiment of this specification provides a storage medium having a computer program stored thereon, and the computer program implements the steps of the above method when executed by a processor.
[0019] In a fifth aspect, an embodiment of this specification provides a computer program product, comprising: a computer program, which, when executed by a processor of an electronic device, enables the processor to at least implement the method described in the first aspect.
[0020] In an embodiment of the present specification, by obtaining a model set and sample user features of sample users, each model in the model set represents a subject participating in a transaction, the model set includes a target transaction party model and at least one associated transaction party model, the target transaction party model is a model corresponding to the target transaction party, and is used to help the target transaction party generate a transaction decision result, the sample user features are respectively input into each model in the model set, and the sample transaction decision result corresponding to each model is generated based on the global state parameters, sample user features, internal state parameters and behavioral rules of each model, the actual selection result of the sample user for the sample transaction decision result of the target transaction party model is determined, the decision reward and punishment value for each model is determined based on the actual selection result, the model set is updated according to the objective function, decision reward and punishment value and sample transaction decision result of each model, and the step of obtaining the model set and sample user features of the sample user is executed until the simulation loop end condition is reached, and a trained model set is obtained, and the target transaction party model in the trained model set is used to generate a transaction decision result based on the target user features of the target user. By training multiple models together, the target transaction party model can interact with the environment and the associated transaction party models, so that the target transaction party model can learn how to make better decisions in different environments and in the presence of other transaction parties, thereby improving the adaptability and effectiveness of the model's decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0022] Figure 1 This is an example diagram of a transaction decision-making-based model training method provided in an embodiment of this specification;
[0023] Figure 2 This is an example diagram of a transaction decision-making-based model training method provided in an embodiment of this specification;
[0024] Figure 3 This is a flowchart of a transaction decision-making-based model training method provided in an embodiment of this specification;
[0025] Figure 4 This is a flowchart of a transaction decision-making-based model training method provided in an embodiment of this specification;
[0026] Figure 5 This is a flowchart of a transaction decision-making-based model training method provided in an embodiment of this specification;
[0027] Figure 6 This is a schematic diagram of the architecture of a transaction decision-making-based model training method provided in an embodiment of this specification;
[0028] Figure 7 This is a schematic diagram of the steps of a transaction decision-making-based model training method provided in an embodiment of this specification;
[0029] Figure 8 This is a structural diagram of a transaction decision-based model training device provided in an embodiment of this specification;
[0030] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this specification. DETAILED DESCRIPTION
[0031] The following will be combined with the drawings in the embodiments of this specification to clearly and completely describe the technical solutions in the embodiments of this specification. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this specification.
[0032] In related technologies, taking the marketing of internet financial products as an example, the use of various online benefits such as coupons, points, and discounts is one of the most common and important means for internet consumer finance institutions to attract users. However, when making decisions about the issuance of marketing benefits, there are the following problems: the budget for issuing online benefits is limited, and if the market and users are not clearly understood, it is easy to waste the budget; traditional marketing strategies often rely on static data analysis and rule engines, which are difficult to cope with the dynamic changes in the market and user behavior; there are many different types of benefits, and the decision-making method is single, only considering the product itself, ignoring the existence of multiple competitors in the market and the phenomenon that competitors' strategies will affect user choices. See Figure 1 , Figure 1 This document provides an example diagram of a model training method based on transaction decision-making. In related technologies, when making marketing decisions, the target transaction party typically relies on static historical data analysis, which is unable to capture real-time changes in the market and user behavior. This phenomenon can lead to low utilization of current marketing budgets, inaccurate marketing decisions, and a failure to attract users.
[0033] Based on the above problems, an embodiment of the present specification provides a model training method based on transaction decision-making, which obtains sample user features of a model set and sample users, each model in the model set represents a transaction party participating in the transaction activity, and the model set includes a target transaction party model and at least one associated transaction party model. The sample user features are respectively input into each model in the model set, and the sample transaction decision results corresponding to each model are generated based on the global state parameters, sample user features, internal state parameters and behavioral rules of each model. The actual selection result of the sample user for the sample transaction decision result of the target transaction party model is determined, and the decision reward and punishment value for each model is determined based on the actual selection result. The model set is updated according to the objective function, decision reward and punishment value and sample transaction decision result of each model, and the step of obtaining the sample user features of the model set and the sample user is executed until the simulation loop end condition is reached, thereby obtaining a trained model set. The target transaction party model in the trained model set is used to generate a transaction decision result based on the target user features of the target user. By interacting with the associated transaction party model and continuously optimizing the associated transaction party model and the target transaction party model, the model effect of the associated transaction party model can be improved simultaneously, so that the target transaction party model can adapt to how to make better decisions under different global state parameters (environments) and when associated transaction parties exist, so as to obtain more effective decision results for the target transaction party, thereby improving the accuracy of the target transaction party model's judgment on the transaction decision results.
[0034] Please also see Figure 2 , an example schematic diagram of a model training method based on transaction decision-making is provided for the embodiment of this specification. When applied to a marketing scenario, the target transaction party model is the model corresponding to the target transaction party, which is used to help the target transaction party generate transaction decision results. A corresponding target transaction party model is constructed for the target transaction party, and the associated transaction party model is a model corresponding to the associated transaction party. For example, it can be a competitor or partner of the target transaction party. The target transaction party simulates the real market environment and the decision-making process of different transaction parties at multiple time points (in different states) based on the interaction with the environment and associated transaction parties. Through multiple rounds of joint decision-making, market reactions and user behaviors can be predicted more accurately, making the decision results more accurate and improving the marketing effect of the target transaction party.
[0035] It's understandable that this target transaction model can be applied to a variety of scenarios involving competition with other transaction parties. For example, in finance, it can be used to develop marketing plans and interest rates for different users; in retail, it can be used to develop strategies for supermarket promotions; in advertising, it can be used to bid for ad space; and it can also be used in recommendation scenarios, such as video recommendations.
[0036] It is understandable that the model training device provided in the embodiments of this specification can be a terminal device such as a mobile phone, computer, tablet computer, smart watch or vehicle-mounted device, or it can be a module in the terminal device for implementing the model training method.
[0037] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in the embodiments of this specification are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the sample user characteristics, user supplementary characteristics, and target user characteristics involved in this specification are all obtained with full authorization.
[0038] The transaction decision-making model training method provided in this specification is described in detail below with reference to specific embodiments.
[0039] See Figure 3 , provides a flow chart of a model training method based on transaction decision making for the embodiment of this specification. Figure 3 As shown, the method of the embodiment of this specification may include the following steps S102-S106.
[0040] S102, obtaining a model set and sample user features of a sample user;
[0041] In one embodiment of the present specification, the generation of transaction strategies is achieved by training the model through reinforcement learning. Reinforcement learning (RL) is an important paradigm in machine learning. It focuses on how the agent learns strategies in the environment by interacting with the environment to maximize the accumulated rewards. In this process, the agent affects the state of the environment by taking a series of actions and adjusts its behavior strategy according to the rewards fed back by the environment. Similarly, reinforcement learning can also be applied to the training of the model. Corresponding to the idea of reinforcement learning, the sample transaction decision results generated by the model correspond to action, the global state parameters correspond to state, and the decision reward and punishment values correspond to reward.
[0042] Specifically, a model set and sample user characteristics of the corresponding experimental users (sample users) are obtained. The model set includes a target transaction party model and at least one associated transaction party model. Taking the financial scenario as an example, each model represents a role in the market and has its own objective function and decision logic (behavioral rules). The model set simulates the interaction and competition between multiple entities participating in marketing activities in the market. Each model in the model set continuously optimizes its behavioral strategy through interaction with the environment and other models. The target transaction party model is the model corresponding to the target transaction party, which is the transaction party for which transaction decision results need to be generated; the associated transaction party model is the model corresponding to the associated transaction party. Associated transaction parties are transaction parties associated with the target transaction party. Typically, associated transaction parties are transaction parties in the same market or industry that provide similar products or services as the target transaction party. In the financial scenario, for example, the target transaction party may be internet bank A, and the associated transaction party may be a financial institution, internet bank B, etc. In the retail scenario, for example, the target transaction party may be shopping mall A, and the associated transaction party may be shopping mall B, etc. In the recommendation scenario, for example, the target transaction party may be video platform A, and the associated transaction party may be video platform B, etc.
[0043] Sample users are a subset of real users selected for the experiment, such as consumers of the target party's products. For example, stratified sampling can be used to stratify all users of the target party based on certain characteristics (such as age, gender, and region), and then samples are drawn from each stratum to ensure representativeness.
[0044] Sample user characteristics refer to basic information that can describe a user's personal circumstances, such as name, age, sex, education level, salary level, location, and marital status. They may also include sample user preferences and interaction data regarding transactions provided by the target transaction party. For example, this may include a user's historical video click data (for video recommendation transactions) or historical product purchases (for product recommendation transactions). Typically, sample user data is anonymized and desensitized to ensure compliance with data protection regulations.
[0045] S104, inputting the sample user characteristics into each model in the model set, and generating a sample transaction decision result corresponding to each model based on the global state parameter, the sample user characteristics, the internal state parameters and behavior rules of each model;
[0046] In one embodiment of the present specification, the acquired sample user features are given to each model in the model set. Each model can make decisions based on its own internal state parameters and behavioral rules, combined with the input sample user features and global state parameters, to obtain sample transaction decision results. The sample transaction decision results are the decision results made for the sample users and the required transactions, and the transaction decision results corresponding to different transactions are different. Taking the financial scenario as an example, the transaction can be marketing, and the sample transaction decision result is a marketing plan. Taking the advertising bidding scenario as an example, the transaction can be a bid, and the sample transaction decision result is the bid price. Taking the video recommendation scenario as an example, the transaction can be a video recommendation, and the sample transaction decision result is a recommended video.
[0047] Global state parameters describe the overall state of the environment for all models in the system at a given moment. For example, global state parameters include resource distribution and the state of each model. In the financial scenario, for example, global state parameters might include total resource availability, the current economic cycle, interest rate fluctuations, and policies.
[0048] A model's internal state parameters refer to all information the model holds internally that aids in decision-making, in addition to the external environment and global state. For example, internal state parameters can include historical information and preferences. Historical information refers to the experience accumulated by the model from past actions, including its behavior and environmental responses. Historical information is typically a time series, reflecting the model's decisions, states, and rewards at specific moments in the past. For example, in a financial scenario, internal state parameters may include its defined initial state parameters, such as initial resources. Model A has a budget of 100,000 yuan and issues 1,000 coupons; Model B has a budget of 500,000 yuan and issues 500 coupons; and so on. After each decision, the initial state parameters are modified. For example, if XX yuan was spent on a previous marketing campaign, the resource parameters in the initial state parameters are updated based on the specific marketing expenditure.
[0049] Behavioral rules are policies or guidelines that define how a model responds to external signals (environmental states or the behavior of other models) based on these inputs. These rules can be based on a policy function that determines its actions. A policy can be a mapping based on the state of the environment (e.g., which action should the model choose given the current state). In reinforcement learning, policies can be trained and optimized to determine the preferred behavior by maximizing a certain reward.
[0050] In one feasible implementation, the model can be generated based on a pre-trained large language model. By inputting relevant information about the target transaction party and related transaction parties, the large language model can gain a certain understanding of each transaction party, learn the characteristics of different transaction parties, and generate a model corresponding to each transaction party. Furthermore, historical user behavior data can be input into the large language model to learn the user's likely behavior in this transaction party, and then fine-tune the model to generate a model corresponding to each transaction party. Each model generates transaction decision results based on its own understanding of the environment and users.
[0051] It should be noted that the models can generate sample transaction decision results sequentially according to a preset order. For example, after generating a sample transaction decision result, the first model updates the global state parameters based on the sample transaction decision result, and the next model can make a decision based on the updated global state. Optionally, the preset order can be that the associated transaction party model makes the decision first, and after the associated transaction party model completes the decision, the target transaction party model makes the selection. In this way, the target transaction party model can fully reference the decisions made by the associated transaction party model and its impact on the environment to generate its own decisions, thereby improving the target transaction party model's understanding of various factors in the decision-making process.
[0052] S106, determining the actual selection result of the sample user for the sample transaction decision result of the target transaction party model;
[0053] In one embodiment of the present specification, after all model decisions are completed, whether the sample user selects the marketing plan provided by the sample transaction decision result of the target transaction party model is obtained to obtain the real selection result. It can be understood that the related transaction party model is used to simulate the decision-making actions of related transaction parties such as competing transaction parties or cooperative transaction parties, but does not actually provide the sample user with the sample transaction decision result (decision plan) of the related transaction party model. Therefore, only information on whether the sample user selects the sample transaction decision result provided by the target transaction party model can be obtained. Specifically, the real selection result may include whether the user selects the sample transaction decision result provided by the target transaction party model and the selected decision content. For example, if the decision plan provides the user with multiple rights and interests, the real selection result may include the rights and interests that the user ultimately receives.
[0054] S108, determining a decision reward or penalty value for each of the models based on the actual selection result;
[0055] In one embodiment of the present specification, the decision reward and penalty values for the decision actions made by each model can be determined based on whether the sample user selects the strategy provided by the target transaction party model. It is understandable that if the actual selection result is that the sample user selects the strategy of the target transaction party model, it indicates that the decision action made by the target transaction party model this time is effective, and it can be rewarded to encourage it to choose similar behaviors in the future; correspondingly, it indicates that the decision actions made by other models (related transaction party models) are not effective enough, and corresponding penalties need to be given to encourage them to improve their decision-making methods.
[0056] Specifically, a reward function can be used to evaluate the effectiveness of model behavior. Based on the reward function and the actual choice results, the rewards and penalties each model receives after performing a certain behavior are determined, thereby guiding the model's decision-making behavior.
[0057] S110, update the model set according to the objective function of each model, the decision reward and punishment value and the sample transaction decision result, and proceed to execute the step of obtaining the model set and the sample user characteristics of the sample user until the simulation loop end condition is reached to obtain the trained model set.
[0058] In one embodiment of this specification, each model can be assigned an objective function, such as maximizing revenue or minimizing cost. Each model then updates itself based on its objective function, decision-making reward and penalty values, and sample transaction decision results, optimizing its decision-making approach. This allows the model to determine optimal promotional timing, the most effective reward form, and other sample transaction decision-making outcomes. This optimization approach can employ, for example, a policy gradient method, which directly updates the policy by optimizing its gradient, enabling it to generate higher expected returns in future decisions. Suppose a multi-model system is used for advertising. The objective function is: each model's goal is to maximize click-through rate (CTR) while minimizing ad budget waste. Decision-making reward and penalty values: Each time a model makes an ad placement decision, it receives a reward based on the number of clicks on the ad. For example, the more clicks on an ad, the higher the reward; ads that waste budget or have low CTRs receive a penalty. At each decision step, the model evaluates the quality of its current behavior (e.g., click count and cost) using immediate rewards. Simultaneously, the model optimizes its future decisions using long-term objective functions. For example, it might allocate some of its advertising budget to long-tail ads, rather than focusing solely on high-traffic ads, to achieve greater click returns in the future.
[0059] Optionally, you can also retrieve sample user characteristics and all simulation data to update the model set. Simulation data refers to the sample transaction decision results and state change data (including global and internal states) made by all models in a decision simulation round. For example, you might provide the model with a prompt text: "The user has these characteristics [sample user characteristics], but ultimately did not choose me. Please analyze and update the strategy based on all [sample transaction decision results] and [state change data] made by the model set in this round."
[0060] It's important to note that in practical applications, multiple rounds of recommendations are typically used, with feedback from each round used to refine the next round until the user's preferred video is found. However, in marketing scenarios, the accuracy of decision-making is even higher. Once a marketing strategy is generated (e.g., offering a user an interest rate plan), it won't be frequently changed in the short term. If a user doesn't adopt the strategy, this could lead to user churn, for example, choosing an interest rate plan offered by another financial institution.
[0061] This optimization and update process can be performed for a preset number of rounds, such as performing multiple rounds of simulation for each sample user until a simulation cycle termination condition is reached. This method is then repeated for a large number of sample users to complete the training of the model set. The simulation cycle termination condition can be, for example, reaching a set number of cycles.
[0062] The target transaction party model in the trained model set is used to generate transaction decision results based on the target user characteristics of the target user. It should be noted that the target user can be the same as the sample user or a new user, and there is no specific limitation. In the multi-model decision game, by simultaneously optimizing the target transaction party model and its associated transaction parties (such as competitors), the adaptability and robustness of the model can be effectively improved. The strategy of optimizing competitors enables the target transaction party model to not only cope with the challenges of the current opponent, but also to still have an advantage in the face of possible strategy evolution and improvement of opponent capabilities. In this way, the performance of the target transaction party model is more robust and can cope with more complex and dynamic game environments.
[0063] Optionally, the embodiment of this specification provides a transaction decision method, which may include the following steps S1002 to S1004.
[0064] S1002, obtaining target user characteristics of the target user for the target transaction;
[0065] S1004, input the target user characteristics into the target transaction party model, and generate the transaction decision result of the target user for the target transaction based on the global state parameters, the sample user characteristics, the internal state parameters and behavioral rules of the target transaction party model; wherein, the target transaction party model is a model corresponding to the target transaction party in the model set, and the model set also includes at least one associated transaction party model; the model set is updated according to the objective function, decision reward and punishment value and sample transaction decision result of each model; the reward and punishment decision value of each model is determined based on the actual selection result of the sample user for the sample transaction decision result of the target transaction party model; the sample transaction decision result of each model is generated by inputting the sample user characteristics of the sample user into each model in the model set respectively, and each model is generated based on the global state parameters, the sample user characteristics, the internal state parameters and behavioral rules of each model.
[0066] The target transaction party model is a trained target transaction party model that already includes the target transaction party model's internal state parameters and behavioral rules. Furthermore, the target transaction party model may also include the current global state parameters. That is, the behavioral rules, internal state parameters, and global state parameters are pre-stored in the target transaction party model as prior knowledge. The target user features consist of the target user's own characteristics and the target user's characteristics related to the target transaction. The target user is the object for which decision-making is required, and different target users may have different personal characteristics. Therefore, it is necessary to obtain their target user features as model input. It is understandable that the target user may be an existing user or a new user who has not participated in the target transaction. In this case, the target user features may also only include the target user's own characteristics. The target transaction is the service provided by the target transaction party, such as marketing, advertising bidding, video recommendations, etc.
[0067] In one embodiment, when the target transaction party model outputs the transaction decision result, it can generate the target user's transaction decision result for the target transaction based on the target user's characteristics, current global state parameters (for example, including current environmental information, transaction background information, characteristic information of other transaction parties, etc.), internal state parameters, and behavioral rules. Specifically, when updating the model set based on the objective function of each model, the decision reward and penalty value, and the sample transaction decision result, a loop training method is adopted. If it is determined that the simulation loop end condition has not been met, the step of obtaining the model set and the sample user's sample user characteristics is executed, the sample user characteristics are input into each model in the model set, and the sample transaction decision result corresponding to each model is generated based on the global state parameters, the sample user characteristics, the internal state parameters of each model, and the behavioral rules. The actual selection result of the sample user for the sample transaction decision result of the target transaction party model is determined, and the decision reward and penalty value for each model is determined based on the actual selection result. The model set is updated based on the objective function of each model, the decision reward and penalty value, and the sample transaction decision result until the simulation loop end condition is met.
[0068] By adopting this method of joint training with the associated transaction party model, the trained target transaction party model can improve the decision-making accuracy of the target transaction party on the target transaction while engaging in bargaining or cooperation with the associated transaction party. In addition, the target transaction party model learns the constantly changing environmental state, so that when the environmental state changes, the target transaction party model can make transaction decisions that are more in line with the current environmental state, improving decision-making adaptability.
[0069] In an embodiment of the present specification, by obtaining the characteristic data of a model set and sample users, the sample user characteristics are input into each model in the model set, and the marketing decision of each model is generated based on the global state parameters, sample user characteristics, internal state parameters and behavioral rules of the model. Subsequently, by determining the actual selection results of the sample users on the marketing decision of the target transaction party model, the decision reward and punishment values of each model are calculated and determined. According to these decision reward and punishment values, the objective function of the model and the sample transaction decision results, the model set is updated, and the step of obtaining the sample user characteristics is returned until the conditions for the end of the simulation cycle are met, and finally a trained model set is obtained. The trained target transaction party model will be used to generate transaction decision results based on the target user characteristics. By training multiple models together, the target transaction party model can interact with the environment and the associated transaction party model, thereby optimizing decisions in the context of a constantly changing environment and diverse associated transaction parties, and improving the decision adaptability and effectiveness of the model.
[0070] See Figure 4 , provides a flow chart of a model training method based on transaction decision making for the embodiment of this specification. Figure 4 As shown, the method of the embodiment of this specification may include the following steps S202-S220.
[0071] S202, fine-tuning the pre-trained large language model based on the target transaction party's dataset to obtain a target large language model;
[0072] In one embodiment of the present specification, the target transaction party model can be constructed based on a fine-tuned large language model. Specifically, a suitable pre-trained large language model is selected, the target transaction party's data set is obtained, and the target transaction party's data set is input into the large language model. After the large language model learns, the target large language model is obtained. Taking the target transaction party as a bank as an example, the data set may include historical promotional activities, advertising placement, user responsiveness, service features, product pricing, user complaints, profits, salaries, balance sheets, number of customers, etc. In addition, based on the field of the marketing affairs involved, field-related knowledge can be obtained and input into the large language model for learning. For example, relevant data of related transaction parties can be provided so that the target transaction party can understand the relevant characteristics of the related transaction parties to make better decisions.
[0073] S204, fine-tuning the pre-trained large language model based on the data set of the related transaction party to obtain a competing large language model;
[0074] In one embodiment of the present specification, the associated transaction party model can be constructed based on the fine-tuned large language model, obtain the data set of the competing transaction party, and fine-tune the pre-trained large language model based on the data set to obtain the competing large language model.
[0075] Taking the retail scenario as an example, the target transaction party may be shopping mall A, and the data set may include data such as shopping mall A’s popular products, brands, unique brands, shopping mall positioning, shopping environment, shopping mall traffic, and geographical location; similarly, if the competing transaction party is shopping mall B, the data set may include data such as shopping mall B’s popular products, brands, unique brands, shopping mall positioning, shopping environment, shopping mall traffic, and geographical location.
[0076] S206: Generate a target transaction party model corresponding to the target transaction party based on the target large language model, generate an associated transaction party model corresponding to the competitor transaction party based on the competitor large language model, and obtain a model set based on the target transaction party model and the associated transaction party model;
[0077] In one embodiment of the present specification, the target transaction party model is obtained by setting its objective function, internal state parameters, and behavioral rules based on the determined target large language model. The target transaction party model is then obtained by setting its objective function, internal state parameters, and behavioral rules for the competing large language model. The associated transaction party model and the target transaction party model are combined to form a model set.
[0078] Optionally, when setting internal state parameters, you can set a more favorable initial state for the associated transaction model. This increases the difficulty of competition and improves the decision-making ability of the target transaction model. For example, suppose the target transaction model is given a budget of 5 million yuan and the associated transaction model is given a budget of 10 million yuan. The goal is for the target transaction model to learn how to persuade users to choose the target transaction when the target's budget is low (i.e., at a disadvantage).
[0079] S208, inputting the sample user features into each model in the model set respectively;
[0080] S210, generating a first sample transaction decision result of the first model based on the global state parameter, the sample user characteristics, the internal state parameter of the first model and the behavior rule;
[0081] In one embodiment of the present specification, taking how a model in a model set generates a sample transaction decision result as an example, the first model is any model in the model set. After the sample user features are input into each model in the model set respectively, the first model makes a decision based on the global state parameters, sample user features, the internal state parameters and behavioral rules of the first model to obtain a first sample transaction decision result.
[0082] S212: Based on the first sample transaction decision result, the global state parameter, and the state transition function, determine the global state parameter at the next moment and update the number of decision rounds of the first model. The global state parameter at the next moment is determined as the global state parameter, and the process proceeds to executing the step of generating the first sample transaction decision result of the first model based on the global state parameter, the sample user characteristics, the internal state parameter of the first model, and the behavior rule, until all models in the model set complete their decisions and the number of decision rounds reaches a preset number, thereby obtaining the sample transaction decision results corresponding to each model.
[0083] In one embodiment of the present specification, after generating the first sample transaction decision result, the first model determines the global state parameters at the next moment based on the first sample transaction decision result, the current global state parameters and the state transition function. That is, after each model makes a decision, the global state will change and needs to be updated immediately so that the next model to make a decision can make a decision based on the most recent environmental state, thereby improving the accuracy of the decision. In addition, the number of decision rounds of the first model can also be updated. Usually, the number of decision rounds for each model is consistent. When all models in the model set complete the decision and the number of decision rounds reaches the preset number, all sample transaction decision results made by the model in this simulation are obtained. For example, the preset number is 1, and after all models have made a decision once, the final user's selection result for the transaction decision result given by the target transaction party model is obtained.
[0084] Optionally, in one embodiment of the present specification, determining the actual selection result of the sample user's sample transaction decision result for the target transaction party model includes:
[0085] S2122, determining the sample transaction decision result of the target transaction party model in the last round of decision;
[0086] In one embodiment of the present specification, the target transaction party model may perform multiple rounds of decision-making, obtain the sample transaction decision result generated in the last round of decision-making, and provide the last sample transaction decision result to the sample user.
[0087] S2124, obtaining the actual selection results of the sample users for the sample transaction decision results in the last round of decision;
[0088] In one embodiment of the present specification, the actual selection result of the sample user for the sample transaction decision result in the last round of decision-making is determined. It is understood that the actual selection result of the user, such as whether to cancel the issued benefits, can be determined by collecting user behavior data on the target transaction party.
[0089] S214, determining a simulated selection result of each of the associated transaction party models based on the actual selection result;
[0090] In one embodiment of the present specification, the simulated selection results of the associated transaction party model are inferred based on the user's actual selection results for the target transaction party model. It is understood that because the associated transaction party model is fictitious and its transaction decision results are not provided to the user, it is impossible to determine whether the user will select the strategy provided by the associated transaction party model. The only way to infer this is based on the user's actual selection results for the target transaction party model. If the user selects the strategy of the target transaction party model, then the user has not selected strategies from other models. Similarly, if the user does not select the strategy of the target transaction party model, then the user may have selected strategies from other models.
[0091] S216, determining a decision reward or penalty value for the associated transaction party model based on the simulation selection result;
[0092] In one embodiment of this specification, the decision reward and penalty values for each associated transaction party model's decision action (sample transaction decision result) are determined based on the simulated selection results of each associated transaction party model. Specifically, a reward function can be used to define how to quantify the quality of each action. The decision reward and penalty values for each associated transaction party model are then determined based on the reward function and the simulated selection results. For example, if the equity is claimed, the reward is +10, and if the equity is not claimed, the reward is -5.
[0093] S218, determining a decision reward or penalty value for the target transaction party model based on the actual selection result;
[0094] In one embodiment of the present specification, the decision reward and penalty value of the target transaction party model is determined by the actual selection result. Specifically, the decision reward and penalty value of the target transaction party model can also be determined according to the corresponding reward function and the actual selection result.
[0095] S220, update the model set according to the objective function of each model, the decision reward and punishment value and the sample transaction decision result, and proceed to the step of obtaining the model set and the sample user characteristics of the sample user until the simulation loop end condition is reached to obtain the trained model set.
[0096] In one embodiment of this specification, the model set is updated based on the decision reward and penalty values, the model's objective function, and the sample transaction decision results. Sample user features are then retrieved and the multi-model game steps are repeated until the simulation loop ends, ultimately resulting in a trained model set. The trained target transaction party model is used to generate transaction decision results based on the target user features. It should be noted that for any portions not fully elaborated, please refer to step S110 in the above embodiment of the specification.
[0097] In the embodiments of the present specification, a target large language model is obtained by fine-tuning a pre-trained large language model based on a data set of a target transaction party, a competing large language model is obtained by fine-tuning a pre-trained large language model based on a data set of a competing transaction party, a target transaction party model corresponding to the target transaction party is generated based on the target large language model, and an associated transaction party model corresponding to the competing transaction party is generated based on the competing large language model. By establishing a model based on a large language model, the amount of model training can be reduced, so that the model not only contains rich world knowledge, but also can obtain unique knowledge of the corresponding transaction party. A model set is obtained based on the target transaction party model and the associated transaction party model, and the sample user characteristics are respectively input into each model in the model set. The first sample transaction decision result of the first model is generated based on the global state parameters, sample user characteristics, internal state parameters of the first model and behavioral rules. The global state parameters of the next moment are determined based on the first sample transaction decision result, the global state parameters and the state transfer function, and the number of decision rounds of the first model is updated. The global state parameters of the next moment are determined as the global state parameters, so that each model can update the global state parameters in time after generating the sample transaction decision result, thereby ensuring that the next model can generate a strategy according to the real-time state change when generating a strategy, and then enter the execution based on the global state parameters, sample user characteristics, internal state parameters of the first model and behavioral rules. The step of generating a first sample transaction decision result of the first model, until all models in the model set complete decision-making and the number of decision rounds reaches a preset number, obtaining the sample transaction decision result corresponding to each model, determining the simulated selection result of each related transaction party model based on the actual selection result, determining the decision reward and punishment value for the related transaction party model based on the simulated selection result, determining the decision reward and punishment value for the target transaction party model based on the actual selection result, updating the model set according to the objective function, decision reward and punishment value and sample transaction decision result of each model, and executing the step of obtaining sample user characteristics of the model set and sample users until the simulation loop end condition is reached, obtaining the trained model set, and automatically deciding the transaction decision result according to the target transaction party model obtained by simulation, reducing manual intervention, and improving transaction completion efficiency and transaction decision effect.
[0098] See Figure 5 , provides a flow chart of a model training method based on transaction decision making for the embodiment of this specification. Figure 5 As shown, the method of the embodiment of this specification may include the following steps S302-S306.
[0099] S302, obtaining user supplementary features collected by the target transaction party corresponding to the target transaction party model for the sample user;
[0100] In one embodiment of this specification, since the sample user is a user already in the target transaction party model, more feature data about the sample user can be obtained. This data will be input into the target transaction party model as supplementary user features along with the sample user features. The supplementary user features can be based on additional personal information provided by the sample user to the target transaction party, or information about interactions between the sample user and the target transaction party, such as transaction records. This information is unique and available only to the target transaction party.
[0101] S304, inputting the user supplementary features and the sample user features into the target transaction party model;
[0102] S306 : Generate a sample transaction decision result corresponding to the target transaction party model based on the global state parameters, the sample user characteristics, the user supplementary characteristics, the internal state parameters and behavior rules of the target transaction party model.
[0103] In one embodiment of the present specification, the target transaction party model generates a sample transaction decision result corresponding to the target transaction party model based on the acquired global state parameters, sample user characteristics, user supplementary characteristics, internal state parameters and behavioral rules of the target transaction party model. By obtaining the user supplementary characteristics, the target transaction party model can better understand the sample user and thus improve the decision-making effect of the target transaction party model.
[0104] Optionally, before generating the sample transaction decision results, environmental change information may be obtained, and global state parameters and internal state parameters and behavioral rules of each model may be updated based on the environmental change information. Sample transaction decision results for each model are generated based on the updated global state parameters, internal state parameters, and behavioral rules of each model.
[0105] It is understandable that in order to improve the model's adaptability to dynamic environmental changes, environmental changes can be simulated during training. By obtaining current environmental change information, such as interest rate fluctuations, policy adjustments, economic cycles, and other market change information for financial scenarios, global state parameters are adjusted based on this environmental change information. The environmental change information and the adjusted global state parameters are then input into the model, allowing the model to self-adjust based on the environmental change information and the adjusted global state parameters, thereby obtaining adjusted internal state parameters and behavioral rules. For example, assuming a total of 10 decision rounds, environmental change information can be set in the 5th round and input into the multi-model system.
[0106] See Figure 6 , Figure 6This is a schematic diagram of the architecture of a model training method based on transaction decision-making provided in an embodiment of this specification. This embodiment of this specification provides a model training system. Taking the financial scenario as an example, the main components include: a large language model (LLM) module, a multi-model module (multi-agent model module), a simulation engine module, a data processing module, and an analysis and recommendation system module. The large language module includes a number of fine-tuned LLMs, each LLM corresponds to a financial institution; each model in the multi-model module represents a participant (financial institution) in the market. The simulation engine module is responsible for running the above-mentioned multi-model model and generating a large number of virtual transaction records, including sample transaction decision results generated by each model and the actual selection results of sample users. Through multiple iterations, the dynamic changes of the market and the strategy generation method of continuously updating the model are simulated. The data processing module collects various types of information generated during the simulation process, and cleans and organizes it for subsequent analysis. The analysis and recommendation system module uses statistical and machine learning techniques to extract patterns and trends from historical data, such as user behavior patterns, market fluctuation patterns, etc., predict future trends, and provide marketing plan suggestions based on this.
[0107] See Figure 7 , Figure 7 This is a schematic diagram of the steps of a transaction decision-based model training method provided by an embodiment of this specification, which provides an overall overview of the model training method. Figure 7 The serial numbers marked in the figure are an indication of the execution order. First, by fine-tuning the large language model, we obtain the Self agent, that is, the target transaction party model, which represents its own organization, and the agents A~N (related transaction party models A~N). Each related transaction party model represents a financial institution, that is, corresponding to financial institutions 1~N. Then, user features (sample user features) are generated based on the user's profile and input into each model. Among them, the target transaction party model can also obtain a unique user profile, and then extract the user's supplementary features for input. Model A generates a decision action (sample transaction decision result) based on the user features, updates the global state based on the decision action made by model A, and feeds the updated global state back to model B. Model B can then adjust the state or behavior rules based on the feedback and then make a decision. Similarly, the global state is updated based on the decision, and so on, until model N completes the decision and feeds the state back to the global. At this time, the global state includes the decision results of all related transaction party models, that is, the competitive decision results in the figure. The target transaction party model updates the state and adjusts the behavior rules based on the competitive decision results, and obtains the decision action. At this time, it can be obtained whether the sample user has selected the target transaction party model's user's final choice z s(Real user selection), output the result to complete a simulation. Simultaneously, the global state needs to be updated based on the real user selection, and simulated user selections z1-z3 need to be determined. Complete data for this simulation is obtained, and decision rewards and penalties are calculated based on this data, as well as the model set is updated.
[0108] In an embodiment of the present specification, the target transaction party model is configured to obtain user supplementary features collected by the target transaction party for sample users, input the user supplementary features and sample user features into the target transaction party model, and generate a sample transaction decision result corresponding to the target transaction party model based on the global state parameters, sample user features, user supplementary features, internal state parameters of the target transaction party model, and behavioral rules. By incorporating these unique user features, the target transaction party model can better understand the unique needs and behavior patterns of each user, thereby providing more personalized recommendations or predictions, helping the target transaction party model better capture subtle differences between users and reduce the risk of overfitting. In particular, when user data is diversified, it can improve the generalization ability of the target transaction party model.
[0109] The following will be combined with the Figure 8 , the model training device based on transaction decision-making provided by the embodiment of this specification is introduced in detail. Figure 8 The model training device in this manual is used to execute Figure 2-Figure 7 For the convenience of explanation, only the part related to the embodiment of this specification is shown. For the specific technical details not disclosed, please refer to this specification. Figure 2-Figure 7 The embodiment shown.
[0110] See Figure 8 , which shows a schematic diagram of the structure of a transaction decision-based model training device provided in an exemplary embodiment of this specification. The transaction decision-based model training device can be implemented as all or part of a device through software, hardware, or a combination of both. The device 1 includes an acquisition unit 11, a decision unit 12, an outcome determination unit 13, a reward and punishment unit 14, and a training unit 15.
[0111] An acquisition unit 11 is configured to acquire a model set and sample user features of a sample user; the model set includes a target transaction party model and at least one associated transaction party model;
[0112] A decision unit 12 is configured to input the sample user characteristics into each model in the model set, and generate a sample transaction decision result corresponding to each model based on a global state parameter, the sample user characteristics, the internal state parameters of each model, and the behavior rules;
[0113] A result determination unit 13 is used to determine the actual selection result of the sample user for the sample transaction decision result of the target transaction party model;
[0114] A reward and punishment unit 14 is used to determine a decision reward and punishment value for each of the models based on the actual selection result;
[0115] The training unit 15 is used to update the model set according to the objective function of each model, the decision reward and punishment value and the sample transaction decision result, and proceed to execute the step of obtaining the model set and the sample user characteristics of the sample user until the simulation loop end condition is reached, thereby obtaining a trained model set. The target transaction party model in the trained model set is used to generate a transaction decision result based on the target user characteristics of the target user.
[0116] Optionally, the acquisition unit 11 is further configured to fine-tune the pre-trained large language model based on the target transaction party's data set to obtain a target large language model;
[0117] Fine-tune the pre-trained large language model based on the competitor's dataset to obtain a competitive large language model;
[0118] A target transaction party model corresponding to the target transaction party is generated based on the target large language model, an associated transaction party model corresponding to the competitor transaction party is generated based on the competitor large language model, and a model set is determined based on the target transaction party model and the associated transaction party model.
[0119] Optionally, the decision unit 12 is specifically configured to input the sample user features into each model in the model set respectively;
[0120] generating a first sample transaction decision result of the first model based on the global state parameter, the sample user characteristics, the internal state parameter and behavior rule of the first model; the first model is any model in the model set;
[0121] Based on the first sample transaction decision result, the global state parameter and the state transfer function, the global state parameter at the next moment is determined and the number of decision rounds of the first model is updated. The global state parameter at the next moment is determined as the global state parameter, and the step of generating the first sample transaction decision result of the first model based on the global state parameter, the sample user characteristics, the internal state parameter and behavior rules of the first model is executed until all models in the model set complete the decision and the number of decision rounds reaches a preset number, thereby obtaining the sample transaction decision results corresponding to each model.
[0122] Optionally, the result determination unit 13 is specifically configured to determine the sample transaction decision result of the target transaction party model in the last round of decision;
[0123] Obtain the actual selection results of the sample users for the sample transaction decision results in the last round of decision.
[0124] Optionally, the reward and punishment unit 14 is specifically configured to determine a simulated selection result of each of the associated transaction party models based on the actual selection result;
[0125] Determining a decision reward or penalty value for the associated transaction party model based on the simulation selection result;
[0126] A decision reward or penalty value for the target transaction party model is determined based on the actual selection result.
[0127] Optionally, the decision unit 12 is specifically configured to obtain user supplementary features collected by the target transaction party corresponding to the target transaction party model for the sample user;
[0128] inputting the user supplementary features and the sample user features into the target transaction party model;
[0129] A sample transaction decision result corresponding to the target transaction party model is generated based on the global state parameters, the sample user characteristics, the user supplementary characteristics, the internal state parameters and behavior rules of the target transaction party model.
[0130] Optionally, the decision unit 12 is further configured to obtain environmental change information, and update global state parameters and internal state parameters and behavior rules of each model based on the environmental change information.
[0131] It should be noted that the model training device provided in the above embodiment only uses the division of the above functional modules as an example when executing the model training method. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the model training device provided in the above embodiment and the model training method embodiment belong to the same concept. The implementation process is detailed in the method embodiment and will not be repeated here.
[0132] The serial numbers of the embodiments in this specification are for descriptive purposes only and do not represent the merits of the embodiments. In some cases, the actions or steps recited in the claims may be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0133] The embodiment of this specification also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the above Figure 2-Figure 7 The model training method based on transaction decision-making in the embodiment shown in the figure can be found in the specific execution process. Figure 2-Figure 7 The detailed description of the illustrated embodiment will not be repeated here.
[0134] Please refer to Figure 9 , which shows a schematic diagram of the structure of an electronic device provided by an exemplary embodiment of this specification. The electronic device described in this specification may include one or more of the following components: a processor 110, a memory 120, an input device 130, an output device 140, and a bus 150. The processor 110, the memory 120, the input device 130, and the output device 140 may be connected via the bus 150.
[0135] The processor 110 may include one or more processing cores. The processor 110 utilizes various interfaces and circuits to connect various components within the electronic device. It executes instructions, programs, code sets, or instruction sets stored in the memory 120, as well as accesses data stored in the memory 120, to perform various functions of the electronic device and process data. Optionally, the processor 110 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 110 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interfaces, and applications; the GPU is responsible for rendering and drawing display content; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 110 and may be implemented separately via a communications chip.
[0136] The memory 120 may include a random access memory (RAM) or a read-only memory (ROM). Optionally, the memory 120 includes a non-transitory computer-readable storage medium (Non-Transitory Computer-Readable Storage Medium). The memory 120 may be used to store instructions, programs, codes, code sets or instruction sets. The memory 120 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc. The operating system may be an Android system, including a system deeply developed based on the Android system, an iOS system developed by Apple, including a system deeply developed based on the iOS system or other systems.
[0137] The memory 120 can be divided into an operating system space and a user space. The operating system runs in the operating system space, and native and third-party applications run in the user space. In order to ensure that different third-party applications can achieve better operating results, the operating system allocates corresponding system resources to different third-party applications. However, the requirements for system resources in different application scenarios in the same third-party application are also different. For example, in the local resource loading scenario, the third-party application has higher requirements for disk reading speed; in the animation rendering scenario, the third-party application has higher requirements for GPU performance. The operating system and the third-party application are independent of each other, and the operating system often cannot perceive the current application scenario of the third-party application in a timely manner, resulting in the operating system being unable to perform targeted system resource adaptation according to the specific application scenario of the third-party application.
[0138] In order for the operating system to distinguish the specific application scenarios of third-party applications, it is necessary to open up data communication between third-party applications and the operating system so that the operating system can obtain the current scenario information of third-party applications at any time, and then perform targeted system resource adaptation based on the current scenario.
[0139] The input device 130 is used to receive input commands or data and includes, but is not limited to, a keyboard, a mouse, a camera, a microphone, or a touch-sensitive device. The output device 140 is used to output commands or data and includes, but is not limited to, a display device and a speaker. In one example, the input device 130 and the output device 140 may be combined, and the input device 130 and the output device 140 may be a touch-sensitive display.
[0140] The touch display screen can be designed as a full screen, a curved screen or a special-shaped screen. The touch display screen can also be designed as a combination of a full screen and a curved screen, or a combination of a special-shaped screen and a curved screen, which is not limited in this embodiment of the present invention.
[0141] In addition, those skilled in the art will understand that the structures of the electronic devices shown in the above figures do not limit the electronic devices. The electronic devices may include more or fewer components than shown, or may combine certain components or arrange the components differently. For example, the electronic devices may also include radio frequency circuits, input units, sensors, audio circuits, WiFi modules, power supplies, Bluetooth modules, and other components, which will not be described in detail here.
[0142] exist Figure 9 In the electronic device shown, the processor 110 may be configured to call a computer application stored in the memory 120 and specifically perform the following operations:
[0143] Acquire a model set and sample user features of sample users; the model set includes a target transaction party model and at least one associated transaction party model;
[0144] Inputting the sample user characteristics into each model in the model set respectively, and generating a sample transaction decision result corresponding to each model based on the global state parameter, the sample user characteristics, the internal state parameter and behavior rules of each model;
[0145] Determining a real selection result of the sample user for the sample transaction decision result of the target transaction party model;
[0146] Determining a decision reward or penalty value for each of the models based on the actual selection result;
[0147] The model set is updated according to the objective function of each model, the decision reward and punishment value, and the sample transaction decision result, and the step of obtaining the model set and the sample user characteristics of the sample user is executed until the simulation loop end condition is reached, thereby obtaining a trained model set. The target transaction party model in the trained model set is used to generate a transaction decision result based on the target user characteristics of the target user.
[0148] In one embodiment, before executing the step of acquiring the model set and the sample user features of the sample users, the processor 110 further performs the following operations:
[0149] Fine-tune the pre-trained large language model based on the target transaction party's dataset to obtain the target large language model;
[0150] Fine-tune the pre-trained large language model based on the competitor's dataset to obtain a competitive large language model;
[0151] A target transaction party model corresponding to the target transaction party is generated based on the target large language model, an associated transaction party model corresponding to the competitor transaction party is generated based on the competitor large language model, and a model set is determined based on the target transaction party model and the associated transaction party model.
[0152] In one embodiment, when the processor 110 inputs the sample user features into each model in the model set and generates a sample transaction decision result corresponding to each model based on the global state parameter, the sample user features, the internal state parameters of each model, and the behavior rules, the processor 110 specifically performs the following operations:
[0153] Inputting the sample user features into each model in the model set respectively;
[0154] generating a first sample transaction decision result of the first model based on the global state parameter, the sample user characteristics, the internal state parameter and behavior rule of the first model; the first model is any model in the model set;
[0155] Based on the first sample transaction decision result, the global state parameter and the state transfer function, the global state parameter at the next moment is determined and the number of decision rounds of the first model is updated. The global state parameter at the next moment is determined as the global state parameter, and the step of generating the first sample transaction decision result of the first model based on the global state parameter, the sample user characteristics, the internal state parameter and behavior rules of the first model is executed until all models in the model set complete the decision and the number of decision rounds reaches a preset number, thereby obtaining the sample transaction decision results corresponding to each model.
[0156] In one embodiment, when determining the actual selection result of the sample user's sample transaction decision result for the target transaction party model, the processor 110 specifically performs the following operations:
[0157] Determining a sample transaction decision result of the target transaction party model in the last round of decision;
[0158] Obtain the actual selection results of the sample users for the sample transaction decision results in the last round of decision.
[0159] In one embodiment, when determining the decision reward or penalty value for each model based on the actual selection result, the processor 110 specifically performs the following operations:
[0160] Determining a simulated selection result of each of the associated transaction party models based on the actual selection result;
[0161] Determining a decision reward or penalty value for the associated transaction party model based on the simulation selection result;
[0162] A decision reward or penalty value for the target transaction party model is determined based on the actual selection result.
[0163] In one embodiment, when the processor 110 inputs the sample user features into each model in the model set and generates a sample transaction decision result corresponding to each model based on the global state parameter, the sample user features, the internal state parameters of each model, and the behavior rules, the processor 110 specifically performs the following operations:
[0164] Obtaining user supplementary features collected by the target transaction party corresponding to the target transaction party model for the sample user;
[0165] inputting the user supplementary features and the sample user features into the target transaction party model;
[0166] A sample transaction decision result corresponding to the target transaction party model is generated based on the global state parameters, the sample user characteristics, the user supplementary characteristics, the internal state parameters and behavior rules of the target transaction party model.
[0167] In one embodiment, the processor 110 is further configured to perform the following operations:
[0168] Acquire environmental change information, and update global state parameters and internal state parameters and behavior rules of each model based on the environmental change information.
[0169] In an embodiment of the present specification, by obtaining the characteristic data of a model set and sample users, the sample user characteristics are input into each model in the model set, and the marketing decision of each model is generated based on the global state parameters, sample user characteristics, internal state parameters of the model and behavioral rules. Subsequently, by determining the actual selection results of the sample users on the marketing decision of the target transaction party model, the decision reward and punishment values of each model are calculated and determined. According to these decision reward and punishment values, the objective function of the model and the sample transaction decision results, the model set is updated, and the step of obtaining the sample user characteristics is returned until the conditions for the end of the simulation cycle are met, and finally a trained model set is obtained. The trained target transaction party model will be used to generate transaction decision results based on the target user characteristics. By training multiple models together, the target transaction party model can interact with the environment and the associated transaction party model, thereby optimizing decisions in a constantly changing environment and diversified competitive scenarios, and improving the decision adaptability and effectiveness of the model.
[0170] Furthermore, by fine-tuning the pre-trained big language model based on the data set of the target transaction party to obtain the target big language model, fine-tuning the pre-trained big language model based on the data set of the competing transaction party to obtain the competing big language model, generating the target transaction party model corresponding to the target transaction party based on the target big language model, and generating the associated transaction party model corresponding to the competing transaction party based on the competing big language model, by establishing a model (intelligent agent) based on the big language model, the amount of model training can be reduced, so that it not only contains rich world knowledge, but also can obtain the unique knowledge of the corresponding transaction party. A model set is obtained based on the target transaction party model and the associated transaction party model, and the sample user characteristics are respectively input into each model in the model set. The first sample transaction decision result of the first model is generated based on the global state parameters, sample user characteristics, internal state parameters of the first model and behavioral rules. The global state parameters at the next moment are determined based on the first sample transaction decision result, the global state parameters and the state transfer function, and the number of decision rounds of the first model is updated. The global state parameters at the next moment are determined as the global state parameters, so that each model can update the global state parameters in time after generating the sample transaction decision result, thereby ensuring that the next model can generate a strategy according to the real-time state change when generating a strategy, and then enter the execution based on the global state parameters, sample user characteristics, internal state parameters of the first model and behavior rules. The rule generates a step of generating a first sample transaction decision result of the first model, until all models in the model set complete the decision and the number of decision rounds reaches a preset number, and obtains the sample transaction decision result corresponding to each model, determines the simulated selection result of each related transaction party model based on the actual selection result, determines the decision reward and punishment value for the related transaction party model based on the simulated selection result, determines the decision reward and punishment value for the target transaction party model based on the actual selection result, updates the model set according to the objective function, decision reward and punishment value and sample transaction decision result of each model, and proceeds to execute the step of obtaining sample user characteristics of the model set and sample users until the simulation loop end condition is reached, obtains the trained model set, and automatically decides the transaction decision result according to the target transaction party model obtained by simulation, reduces manual intervention, and improves marketing efficiency and marketing effect.
[0171] Furthermore, by obtaining supplementary user features collected by the target transaction party for sample users corresponding to the target transaction party model, the user supplementary features and sample user features are input into the target transaction party model, and the sample transaction decision results corresponding to the target transaction party model are generated based on the global state parameters, sample user features, user supplementary features, the internal state parameters and behavioral rules of the target transaction party model. By incorporating these unique user features, the target transaction party model can better understand the unique needs and behavior patterns of each user, thereby providing more personalized recommendations or predictions. This can help the target transaction party model better capture subtle differences between users, reduce the risk of overfitting, and improve the generalization ability of the target transaction party model, especially when user data is diverse.
[0172] In addition, an embodiment of this specification provides a computer program product, which includes a computer program. When the computer program is executed by a processor of an electronic device, the processor can at least implement the above-mentioned Figures 2 to 7 The methods provided in the illustrated embodiments.
[0173] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0174] The above disclosure is only a preferred embodiment of this specification, and certainly cannot be used to limit the scope of rights of this specification. Therefore, equivalent changes made according to the claims of this specification are still within the scope covered by this specification.
Claims
1. A model training method based on transaction decision-making, the method comprising: Acquire a model set and sample user features of sample users; the model set includes a target transaction party model and at least one associated transaction party model; Inputting the sample user characteristics into each model in the model set respectively, and generating a sample transaction decision result corresponding to each model based on the global state parameter, the sample user characteristics, the internal state parameter and behavior rules of each model; Determining a real selection result of the sample user for the sample transaction decision result of the target transaction party model; Determining a decision reward or penalty value for each of the models based on the actual selection result; The model set is updated according to the objective function of each model, the decision reward and punishment value, and the sample transaction decision result, and the step of obtaining the model set and the sample user characteristics of the sample user is executed until the simulation loop end condition is reached, thereby obtaining a trained model set. The target transaction party model in the trained model set is used to generate a transaction decision result based on the target user characteristics of the target user.
2. The method according to claim 1, before obtaining the model set and the sample user features of the sample users, further comprising: Fine-tune the pre-trained large language model based on the target transaction party's dataset to obtain the target large language model; Fine-tune the pre-trained large language model based on the dataset of the related transaction party to obtain a competitive large language model; A target transaction party model corresponding to the target transaction party is generated based on the target large language model, an associated transaction party model corresponding to the associated transaction party is generated based on the competitor large language model, and a model set is determined based on the target transaction party model and the associated transaction party model.
3. The method of claim 1, wherein inputting the sample user characteristics into each model in the model set, and generating a sample transaction decision result corresponding to each model based on a global state parameter, the sample user characteristics, the internal state parameters of each model, and the behavior rules, comprises: Inputting the sample user features into each model in the model set respectively; generating a first sample transaction decision result of the first model based on the global state parameter, the sample user characteristics, the internal state parameter and behavior rule of the first model; the first model is any model in the model set; Based on the first sample transaction decision result, the global state parameter and the state transfer function, the global state parameter at the next moment is determined and the number of decision rounds of the first model is updated. The global state parameter at the next moment is determined as the global state parameter, and the step of generating the first sample transaction decision result of the first model based on the global state parameter, the sample user characteristics, the internal state parameter and behavior rules of the first model is executed until all models in the model set complete the decision and the number of decision rounds reaches a preset number, thereby obtaining the sample transaction decision results corresponding to each model.
4. The method according to claim 3, wherein determining the actual selection result of the sample user for the sample transaction decision result of the target transaction party model comprises: Determining a sample transaction decision result of the target transaction party model in the last round of decision; Obtain the actual selection results of the sample users for the sample transaction decision results in the last round of decision.
5. The method according to claim 1, wherein determining the decision reward and penalty value for each model based on the actual selection result comprises: Determining a simulated selection result of each of the associated transaction party models based on the actual selection result; Determining a decision reward or penalty value for the associated transaction party model based on the simulation selection result; A decision reward or penalty value for the target transaction party model is determined based on the actual selection result.
6. The method of claim 1, wherein inputting the sample user characteristics into each model in the model set, and generating a sample transaction decision result corresponding to each model based on a global state parameter, the sample user characteristics, the internal state parameters of each model, and the behavior rules, comprises: Obtaining user supplementary features collected by the target transaction party corresponding to the target transaction party model for the sample user; Inputting the user supplementary features and the sample user features into the target transaction party model; A sample transaction decision result corresponding to the target transaction party model is generated based on the global state parameters, the sample user characteristics, the user supplementary characteristics, the internal state parameters and behavior rules of the target transaction party model.
7. The method of claim 1, wherein inputting the sample user characteristics into each model in the model set, and generating a sample transaction decision result corresponding to each model based on a global state parameter, the sample user characteristics, the internal state parameters of each model, and the behavior rules, comprises: Acquiring environmental change information, and updating global state parameters and internal state parameters and behavior rules of each model based on the environmental change information; The sample user features are input into each model in the model set respectively, and the sample transaction decision results corresponding to each model are generated based on the sample user features, the updated global state parameters, the updated internal state parameters and behavior rules of each model.
8. A model training device based on transaction decision-making, comprising: An acquisition unit, configured to acquire a model set and sample user features of sample users; the model set includes a target transaction party model and at least one associated transaction party model; a decision unit, configured to input the sample user characteristics into each model in the model set, and generate a sample transaction decision result corresponding to each model based on a global state parameter, the sample user characteristics, an internal state parameter of each model, and a behavior rule; A result determination unit, configured to determine a real selection result of the sample user for the sample transaction decision result of the target transaction party model; A reward and punishment unit, configured to determine a decision reward and punishment value for each of the models based on the actual selection result; The training unit is used to update the model set according to the objective function of each model, the decision reward and punishment value and the sample transaction decision result, and proceed to execute the step of obtaining the model set and the sample user characteristics of the sample user until the simulation loop end condition is reached, thereby obtaining a trained model set. The target transaction party model in the trained model set is used to generate a transaction decision result based on the target user characteristics of the target user.
9. An electronic device comprising: processor and memory; The memory stores a computer program, which is suitable for being loaded by the processor and executing the steps of the method according to any one of claims 1 to 7.
10. A storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7.
11. A computer program product comprising: A computer program, when executed by a processor of an electronic device, causes the processor to perform the steps of the method according to any one of claims 1 to 7.