Intelligent agent training method and device, data configuration method and device, medium and equipment
By building an interactive environment and using reinforcement learning to train the agent, the ratio of resource provision is solved, and the problem of difficult to reasonably determine the ratio of resource provision for resource providers in financial activities is achieved, and the optimization of resource allocation and management is achieved.
Patent Information
- Application Number
- CN202510088916.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-13
AI Technical Summary
In financial activities, when multiple resource providers jointly provide resources, it is difficult to reasonably determine the resource provision ratio of each resource provider, which affects the efficiency of resource allocation and management.
By obtaining switch effect information samples and resource asset information samples, building an interactive environment, using reinforcement learning to train the agent, and optimizing the resource provision ratio.
The optimization of resource allocation and management of each resource provider has been achieved, and the efficiency and rationality of resource allocation have been improved.
Smart Images

Figure CN119990370A_ABST
Abstract
Description
Technical Field
[0001] The present specification relates to the field of computer technology, and more specifically, to an intelligent agent training method, data configuration method, apparatus, medium and device in the field of computer technology. Background Art
[0002] In financial activities, some transactions require multiple resource providers to jointly provide resources to achieve their goals. The resource provision ratios of different resource providers may vary. In order to better implement resource allocation and management of each resource provider, it is necessary to reasonably determine the resource provision ratio of each resource provider. Summary of the invention
[0003] This specification provides an intelligent agent training method, data configuration method, device, medium and equipment, which helps to optimize the resource configuration and management of each resource provider using a target intelligent agent.
[0004] In a first aspect, a method for training an intelligent agent is provided, the method comprising:
[0005] Obtaining a switch effect information sample, the switch effect information sample including a plurality of switch parameter combination samples and effect data of each of the plurality of switch parameter combination samples on a target transaction within a historical time range, the switch parameter combination sample indicating a resource provision ratio of each of the plurality of resource providers to convert resources into assets based on the target transaction;
[0006] Obtain resource asset information samples of each resource provider within a historical time range;
[0007] Building a first interactive environment based on the switch effect information sample and the resource asset information sample of each resource provider;
[0008] Based on the first interactive environment, reinforcement learning training is performed on the initial intelligent agent to obtain a target intelligent agent.
[0009] In combination with the first aspect, in some possible implementations, reinforcement learning training is performed on an initial intelligent agent based on a first interactive environment to obtain a target intelligent agent, including: determining a first solution space for the initial intelligent agent; determining a first action configuration parameter of the initial intelligent agent, the first action configuration parameter indicating an action of the initial intelligent agent, the action being selecting a switch parameter combination sample from multiple switch parameter combination samples as the current switch parameter combination sample; determining a first reward mechanism parameter of the initial intelligent agent, the first reward mechanism parameter indicating that after the initial intelligent agent performs an action, first feedback data that changes with time within a historical time range and is returned by the first interactive environment is obtained, and reward data of the initial intelligent agent under the current switch parameter combination sample is determined based on the first feedback data; based on the first solution space, the first action configuration parameter and the first reward mechanism parameter, the initial intelligent agent is iteratively trained in the first interactive environment; if a preset convergence condition is met, the initial intelligent agent is determined as the target intelligent agent.
[0010] In combination with the first aspect and the above-mentioned implementation methods, in some possible implementation methods, the first solution space of the initial intelligent agent is analyzed and obtained, including: obtaining the first transaction constraint data of each resource provider in the historical time range; determining the upper and lower asset limits of each resource provider in the historical time range based on the switch effect information sample and the resource asset information sample and the first transaction constraint data of each resource provider; and obtaining the first solution space of the initial intelligent agent based on the upper and lower asset limits of each resource provider in the historical time range.
[0011] In combination with the first aspect and the above-mentioned implementation methods, in some possible implementation methods, obtaining a switch effect information sample includes: determining multiple first switch parameter combination samples and multiple second switch parameter combination samples; obtaining the impact effect data of each first switch parameter combination sample in the multiple first switch parameter combination samples on the target transaction within a historical time range; calling a preset linear regression model based on each first switch parameter combination sample, the impact effect data of each first switch parameter combination sample, and multiple second switch parameter combination samples to obtain the impact effect data of each second switch parameter combination sample in the multiple second switch parameter combination samples output by the linear regression model on the target transaction within the historical time range; constructing a switch effect information sample based on each first switch parameter combination sample, the impact effect data of each first switch parameter combination sample, each second switch parameter combination sample, and the impact effect data of each second switch parameter combination sample.
[0012] In combination with the first aspect and the above-mentioned implementation methods, in some possible implementation methods, resource asset information samples of each resource provider in a historical time range are obtained, including: obtaining resource asset status prediction data and resource asset status observation data of each resource provider in the historical time range; and determining the resource asset information samples of each resource provider in the historical time range based on the resource asset status prediction data and resource asset status observation data of each resource provider in the historical time range.
[0013] In a second aspect, a data configuration method is provided, the method comprising:
[0014] Obtaining switch effect information, the switch effect information including multiple switch parameter combinations and effect data of each of the multiple switch parameter combinations on the target transaction within a target time range, the switch parameter combination indicating a resource provision ratio of each of the multiple resource providers to convert resources into assets based on the target transaction;
[0015] Obtain resource asset information of each resource provider within the target time range;
[0016] Building a second interactive environment based on the switch effect information and the resource asset information of each resource provider;
[0017] Invoking a target agent based on the second interactive environment so that the target agent determines a target switch parameter combination from a plurality of switch parameter combinations, wherein the target agent is trained based on the first aspect or a possible implementation of the first aspect;
[0018] The target resource provision ratio of each resource provider is configured based on the target switch parameter combination.
[0019] In combination with the second aspect, in some possible implementations, calling a target intelligent agent based on a second interactive environment so that the target intelligent agent determines a target switch parameter combination from multiple switch parameter combinations includes: determining a second solution space for the target intelligent agent; determining a second action configuration parameter for the target intelligent agent, the second action configuration parameter indicating an action of the target intelligent agent, the action being to select a switch parameter combination from multiple switch parameter combinations as the current switch parameter combination; determining a second reward mechanism parameter for the target intelligent agent, the second reward mechanism parameter indicating that after the target intelligent agent performs the action, second feedback data that changes with time within a target time range and is returned by the second interactive environment is obtained, and reward data for the target intelligent agent under the current switch parameter combination is determined based on the second feedback data; based on the second solution space, the second action configuration parameter and the second reward mechanism parameter, calling the target intelligent agent in the second interactive environment so that the target intelligent agent determines the target switch parameter combination from multiple switch parameter combinations.
[0020] In combination with the second aspect and the above-mentioned implementation methods, in some possible implementation methods, determining the second solution space of the target intelligent entity includes: obtaining the second transaction constraint data of each resource provider in the target time range; determining the upper and lower asset limits of each resource provider in the target time range based on the switch effect information and the resource asset information and the second transaction constraint data of each resource provider; and obtaining the second solution space of the target intelligent entity based on the analysis of the upper and lower asset limits of each resource provider in the target time range.
[0021] In combination with the second aspect and the above-mentioned implementation methods, in some possible implementation methods, based on the second solution space, the second action configuration parameters and the second reward mechanism parameters, the target intelligent agent is called in the second interactive environment so that the target intelligent agent determines the target switch parameter combination from multiple switch parameter combinations, including: based on the switch effect information and the resource asset information of each resource provider, and the second transaction constraint data, determining the function curve of the asset upper limit and asset lower limit of the target resource provider changing with time within the target time range; based on the function curve, determining at least one target time point of the target resource provider within the target time range; dividing the target time range based on at least one target time point to obtain multiple time periods; based on multiple time periods, the second solution space, the second action configuration parameters and the second reward mechanism parameters, calling the target intelligent agent in the second interactive environment so that the target intelligent agent determines the target switch parameter combination corresponding to each of the multiple time periods from the multiple switch parameter combinations.
[0022] In a third aspect, an intelligent agent training device is provided, the device comprising:
[0023] A first acquisition unit is used to acquire a switch effect information sample, the switch effect information sample including a plurality of switch parameter combination samples and effect data of each of the plurality of switch parameter combination samples on a target transaction within a historical time range, the switch parameter combination sample indicating a resource provision ratio of each of the plurality of resource providers to convert resources into assets based on the target transaction;
[0024] The second acquisition unit is used to acquire resource asset information samples of each resource provider within a historical time range;
[0025] A construction unit, configured to construct a first interactive environment based on the switch effect information sample and the resource asset information samples of each resource provider;
[0026] The training unit is used to perform reinforcement learning training on the initial intelligent agent based on the first interactive environment to obtain a target intelligent agent.
[0027] In a fourth aspect, a data configuration device is provided, the device comprising:
[0028] A first acquisition unit is used to acquire switch effect information, the switch effect information including a plurality of switch parameter combinations and effect data of each of the plurality of switch parameter combinations on a target transaction within a target time range, the switch parameter combination indicating a resource provision ratio of each of the plurality of resource providers to convert resources into assets based on the target transaction;
[0029] The second acquisition unit is used to acquire resource asset information of each resource provider within a target time range;
[0030] A construction unit, used to construct a second interactive environment based on the switch effect information and the resource asset information of each resource provider;
[0031] A calling unit, configured to call a target agent based on a second interactive environment, so that the target agent determines a target switch parameter combination from a plurality of switch parameter combinations, wherein the target agent is trained based on the first aspect or a possible implementation of the first aspect;
[0032] The configuration unit is used to configure the target resource provision ratio of each resource provider based on the target switch parameter combination.
[0033] In a fifth aspect, a computer-readable storage medium is provided, which stores a computer program code. When the computer program code is executed, the above method is implemented.
[0034] In a sixth aspect, an electronic device is provided, comprising: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the steps of the above method.
[0035] In a seventh aspect, a computer program product is provided, which stores at least one instruction, and when the at least one instruction is executed by a processor, the steps of the above method are implemented.
[0036] The solution provided in this specification mainly includes: first, obtaining a switch effect information sample, the switch effect information sample includes multiple switch parameter combination samples and the effect data of each switch parameter combination sample on the target transaction within a historical time range. Secondly, obtaining a resource asset information sample of each resource provider within a historical time range. A first interactive environment is constructed based on the switch effect information sample and the resource asset information sample of each resource provider, so that the first interactive environment can reflect the impact of the switch parameter combination sample on the target transaction, as well as the resource asset situation of each resource provider. Based on the first interactive environment, reinforcement learning training is performed on the initial intelligent agent to obtain a target intelligent agent, which can select a switch parameter combination according to the impact of the switch parameter combination on the target transaction and the resource asset situation of each resource provider to reasonably determine the resource provision ratio of each resource provider, which helps to optimize the resource allocation and management of each resource provider. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 This is a schematic diagram of a scenario in which a switch parameter combination provided in an embodiment of this specification affects a resource provision ratio of a resource provider;
[0038] Figure 2 It is a flowchart of an intelligent agent training method provided in an embodiment of this specification;
[0039] Figure 3 It is a flowchart of a training process for obtaining a target intelligent agent provided by an embodiment of this specification;
[0040] Figure 4 It is a schematic diagram of a flow chart of obtaining a switching effect information sample provided by an embodiment of this specification;
[0041] Figure 5 This is a flow chart of obtaining a resource asset information sample provided by an embodiment of this specification;
[0042] Figure 6 It is a flowchart of a data configuration method provided in an embodiment of this specification;
[0043] Figure 7 is a schematic diagram of a process for determining target switch parameters provided by an embodiment of this specification;
[0044] Figure 8 is a schematic diagram of a flow chart for determining a second solution space provided in an embodiment of this specification;
[0045] Fig. 9 is an example schematic diagram of a function curve provided in an embodiment of this specification;
[0046] Fig.10is an example schematic diagram of an intelligent agent interacting with an interactive environment provided in an embodiment of this specification;
[0047] Fig.11 It is a structural schematic diagram of an intelligent agent training device provided in an embodiment of this specification;
[0048] Fig.12 It is a structural schematic diagram of a data configuration device provided in an embodiment of this specification;
[0049] Fig.13 It is a structural schematic diagram of an electronic device provided in an embodiment of this specification. DETAILED DESCRIPTION
[0050] The technical solutions in this specification will be described clearly and in detail below in conjunction with the accompanying drawings. In the description of the embodiments of this specification, unless otherwise specified, " / " means or, for example, A / B can mean A or B: "and / or" in the text is only a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of this specification, "multiple" means two or more than two.
[0051] In the following, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as suggesting or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features.
[0052] In financial activities, some transactions require multiple resource providers to jointly provide resources to achieve, and the resource provision ratios of different resource providers may be different. In order to better implement the resource allocation and management of each resource provider, it is necessary to reasonably determine the resource provision ratio of each resource provider. To this end, a switch can be introduced to achieve the control of the resource provision ratio of the resource provider. Specifically, the switch mentioned in this specification is a generalized control element, for example, it can be a single switch (with two state parameters such as 0 and 1), an enumerated switch (with multiple discrete state parameters such as 0, 1, 2, etc.), or a composite switch (whose state parameters can be selected in a continuous range, such as any value from 0 to 300, and the control method is similar to a sliding rheostat). The switch parameter combination refers to the combination of the state parameters of each switch in a plurality of switches. In order to realize the control of the resource provision ratio of different resource providers to convert resources into assets based on transactions, the mapping relationship between the switch parameter combination and the resource provision ratio of each resource provider can be predetermined, that is, the resource provision ratio of each resource provider can be changed by changing the switch parameter combination. Based on this, it can be considered that the switch parameter combination indicates the resource provision ratio of each resource provider to convert resources into assets based on transactions.
[0053] See also Figure 1 , is a schematic diagram of a scenario in which a switch parameter combination provided in an embodiment of this specification affects the resource provision ratio of a resource provider. Among them, the switch set includes switch 1 to switch n, a total of n switches, and based on the different state parameters of each switch in switch 1 to switch n, a total of m switch parameter combinations, namely switch parameter combination 1 to switch parameter combination m, can be formed. Assuming that the resource providers are resource provider 1, resource provider 2, and resource provider 3, respectively, different switch parameter combinations indicate different resource provision ratios of resource provider 1, resource provider 2, and resource provider 3. For example, switch parameter combination 1 indicates that the resource provision ratio of resource provider 1 is 10%, the resource provision ratio of resource provider 2 is 20%, and the resource provision ratio of resource provider 3 is 70%; for another example, switch parameter combination 2 indicates that the resource provision ratio of resource provider 1 is 20%, the resource provision ratio of resource provider 2 is 30%, and the resource provision ratio of resource provider 3 is 50%; for another example, switch parameter combination m indicates that the resource provision ratio of resource provider 1 is 40%, the resource provision ratio of resource provider 2 is 50%, and the resource provision ratio of resource provider 3 is 10%.
[0054] The resource asset status of different resource providers will change dynamically, and the resource asset status here can be the current resource asset status or the future resource asset status. For example, some resource providers currently have fewer resources, or some resource providers expect that their resources will increase significantly in the future.
[0055] In addition, the resource asset planning of different resource providers will also change dynamically. The resource asset planning here refers to the planning proposed by the resource provider based on its own management requirements for resource assets. For example, some resource providers plan to significantly increase assets within a specific time frame, or some resource providers plan to appropriately increase the conversion rate of resources and assets within a specific time frame.
[0056] In some possible cases, the above-mentioned transactions in financial activities may be credit transactions. In credit transactions, the relevant resources are usually funds, and the assets are the claims formed through credit transactions, representing the future income rights of the resource provider.
[0057] In order to adapt to the changes in the resource asset status of different resource providers and to meet the resource asset planning of different resource providers as much as possible, it is necessary to reasonably determine the resource provision ratio of each resource provider. The above switch provides the basis for controlling the resource provision ratio of the resource provider. Therefore, by determining the appropriate switch parameter combination, the resource provision ratio of each resource provider can be reasonably determined.
[0058] In response to the above requirements, the main solutions provided by the embodiments of this specification include: first, obtaining a switch effect information sample, which includes multiple switch parameter combination samples and the effect data of each switch parameter combination sample on the target transaction within a historical time range. Secondly, obtaining a resource asset information sample of each resource provider within a historical time range. A first interactive environment is constructed based on the switch effect information sample and the resource asset information sample of each resource provider, so that the first interactive environment can reflect the impact of the switch parameter combination sample on the target transaction, as well as the resource asset situation of each resource provider. Based on the first interactive environment, reinforcement learning training is performed on the initial agent to obtain a target agent, which can select a switch parameter combination according to the impact of the switch parameter combination on the target transaction and the resource asset situation of each resource provider to reasonably determine the resource provision ratio of each resource provider, which helps to optimize the resource allocation and management of each resource provider.
[0059] based on Figure 1 The scene shown is shown below. Figure 2 - Fig.10 , the intelligent agent training method and data configuration method provided in the embodiments of this specification are introduced in detail.
[0060] See also Figure 2 , is a flow chart of an agent training method provided in the embodiment of this specification. Figure 2 As shown, the method of the embodiment of this specification may include the following steps S102 to S106.
[0061] S102, obtaining a switch effect information sample, the switch effect information sample including a plurality of switch parameter combination samples, and effect data of each of the plurality of switch parameter combination samples on a target transaction within a historical time range, the switch parameter combination sample indicating a resource provision ratio of each of the plurality of resource providers for converting resources into assets based on the target transaction.
[0062] Specifically, the switch effect information sample involved in this embodiment refers to a type of switch effect information used as training data in the reinforcement learning training process. On the one hand, the switch effect information sample includes a plurality of switch parameter combination samples, and the switch parameter combination sample refers to a group of data composed of the state parameters of each switch in a plurality of switches. On the one hand, the switch effect information sample includes the effect data of each switch parameter combination sample in a plurality of switch parameter combination samples on the target transaction within a historical time range, wherein the target transaction refers to one of the various transactions involved in financial activities, for example, the target transaction may be a credit transaction; the historical time range refers to a certain time interval in the past, such as the past day, month or several months; the effect data of each switch parameter combination sample on the target transaction within the historical time range refers to the quantitative data of the execution effect of each switch parameter combination sample on the target transaction within the historical time range.
[0063] In some possible cases, the types of impact effect data may include impact effect data of resource providers and impact effect data of resource acquirers. Resource acquirers refer to the party that acquires resources through target transactions and converts them into assets in financial activities, such as borrowers in credit transactions. For example, the impact effect data of resource providers may include quantitative data on the income, risk, evaluation, resource liquidity, etc. of resource providers under the influence of corresponding switch parameter combinations; similarly, the impact effect data of resource acquirers may include quantitative data on the income, risk, evaluation, etc. of resource providers under the influence of corresponding switch parameter combinations.
[0064] A resource provider refers to a party that provides resources for a target transaction in a financial activity. The switch parameter combination sample indicates the resource provision ratio of each resource provider among multiple resource providers to convert resources into assets based on the target transaction, which means that there is a mapping relationship between the switch parameter combination sample and the resource provision ratio of each resource provider. A specific switch parameter combination sample indicates a specific resource provision ratio of each resource provider, and different switch parameter combination samples indicate different specific resource provision ratios of each resource provider.
[0065] Regarding the process of obtaining switch effect information samples, on the one hand, it is necessary to obtain multiple switch parameter combination samples. Among them, multiple switch parameter combination samples can be obtained by analyzing all possible state parameters of the corresponding multiple switches. For example, assuming that there are 3 enumerated switches, and each enumerated switch has 3 state parameters, then by combining all possible state parameters of these 3 enumerated switches, 3 to the power of 3, that is, 27 different switch parameter combination samples can be obtained. Each switch parameter combination sample represents a specific switch state configuration, and the switch state configuration determines the resource provision ratio of each resource provider in the target transaction.
[0066] On the other hand, it is necessary to obtain the effect data of each switch parameter combination sample on the target transaction within the historical time range. The effect data of each switch parameter combination sample on the target transaction within the historical time range can be recorded within the historical time range, and the pre-recorded effect data of each switch parameter combination sample on the target transaction within the historical time range can be directly read. Alternatively, multiple switch parameter combination samples are divided into two parts, and the effect data of a part of the switch parameter combination samples on the target transaction within the historical time range is pre-recorded within the historical time range. Based on the pre-recorded effect data of a part of the switch parameter combination samples on the target transaction within the historical time range, the effect data of another part of the switch parameter combination samples on the target transaction within the historical time range can be predicted.
[0067] Based on multiple switch parameter combination samples and the effect data of each switch parameter combination sample in the multiple switch parameter combination samples on the target transaction within a historical time range, the switch effect information sample can be determined.
[0068] S104, obtaining resource asset information samples of each resource provider within a historical time range.
[0069] Specifically, the resource asset information sample involved in this embodiment refers to a type of resource asset information used as training data in the reinforcement learning training process.
[0070] The resource asset information samples of each resource provider within the historical time range refer to the quantitative data on the resource status, asset status, and conversion between resources and assets of each resource provider within the historical time range, which are used to reflect the resource asset status and operation status of the resource provider within the historical time range.
[0071] It should be noted that the data involved in the resource asset information samples of each resource provider in the historical time range may include at least one of the associated prediction data and observation data. Among them, prediction data refers to the data predicted by artificial, machine learning models or related mathematical models, reflecting the possible resource asset status of the resource provider; observation data refers to the data obtained through actual observation and recording, reflecting the actual resource asset status of the resource provider.
[0072] Regarding the process of obtaining resource asset information samples of each resource provider in the historical time range, in a possible implementation, the relevant documents of each resource provider can be obtained, and then the resource asset information samples of each resource provider in the historical time range can be extracted from the relevant documents of each resource provider. In a possible implementation, the resource asset information samples of each resource provider in the historical time range can be obtained from the data platform of each resource provider using a data interface or data crawling technology. In a possible implementation, the resource asset information samples of each resource provider in the historical time range are predetermined and stored locally, and the resource asset information samples of each resource provider in the historical time range can be directly read locally.
[0073] S106: Building a first interactive environment based on the switch effect information samples and the resource asset information samples of each resource provider.
[0074] Specifically, the interactive environment is one of the basic elements of reinforcement learning, which refers to the virtual or simulated environment in which the agent learns and makes decisions. In the interactive environment, the agent can observe the state changes of the interactive environment by performing actions and adjust its strategy according to the feedback of the environment. In this embodiment, the first interactive environment is constructed based on the switch effect information samples and the resource asset information samples of each resource provider. It simulates the impact of the switch parameter combination on the target transaction in the financial activity, as well as the resource asset status of each resource provider.
[0075] Regarding the process of constructing the first interactive environment based on the switch effect information samples and the resource asset information samples of each resource provider, in some possible implementations, the switch effect information samples and the resource asset information samples of each resource provider can be mapped to the pre-created initial interactive environment to obtain the first interactive environment. Specifically, on the one hand, the switch effect information samples are mapped to the action space of the initial interactive environment, and on the other hand, the resource asset information samples of each resource provider are mapped to the state space of the initial interactive environment, and then the initial interactive environment is determined as the first interactive environment, providing a basis for subsequent reinforcement learning training.
[0076] S108, performing reinforcement learning training on the initial intelligent agent based on the first interactive environment to obtain a target intelligent agent.
[0077] Specifically, the initial agent involved in this embodiment is a type of agent, which refers to a system or entity that can perceive the state of the environment, perform actions, and learn and make decisions based on environmental feedback in the reinforcement learning framework. The initial agent is an untrained or only basically trained agent, whose strategies and behaviors may not be optimized enough and need to be improved through reinforcement learning training.
[0078] The process of training the initial agent through reinforcement learning based on the first interactive environment to obtain the target agent is specifically as follows: first, the initial agent is placed in the first interactive environment, and it is allowed to perceive the state of the first interactive environment, perform actions, and receive the first feedback data of the first interactive environment through interaction with the first interactive environment. In this process, the initial agent will continuously try different switch parameter combination samples to observe its impact on the state of the first interactive environment and the rewards obtained. Then, based on these experiences and feedback, the initial agent will gradually adjust its strategy to optimize its future behavioral decisions. After a certain degree of training and iteration, the strategy of the initial agent will tend to be stable and show good performance, at which point it can be determined as the target agent.
[0079] In this embodiment, first, a switch effect information sample is obtained, and the switch effect information sample includes multiple switch parameter combination samples and the effect data of each switch parameter combination sample on the target transaction within a historical time range. Secondly, a resource asset information sample of each resource provider within a historical time range is obtained. A first interactive environment is constructed based on the switch effect information sample and the resource asset information sample of each resource provider, so that the first interactive environment can reflect the impact of the switch parameter combination sample on the target transaction, as well as the resource asset situation of each resource provider. Based on the first interactive environment, reinforcement learning training is performed on the initial intelligent agent to obtain a target intelligent agent, which can select a switch parameter combination according to the impact of the switch parameter combination on the target transaction and the resource asset situation of each resource provider to reasonably determine the resource provision ratio of each resource provider, which helps to optimize the resource allocation and management of each resource provider.
[0080] See also Figure 3 , which is a schematic diagram of a process for training a target intelligent agent according to an embodiment of this specification. Figure 3 As shown, the method of the embodiment of this specification may include the following steps S202 to S210, and steps S202 to S210 may be used as Figure 2 The detailed steps of step S108 of the illustrated embodiment.
[0081] S202, determining a first solution space of the initial intelligent agent;
[0082] S204, determining a first action configuration parameter of the initial agent, where the first action configuration parameter indicates an action of the initial agent, where the action is selecting a switch parameter combination sample from a plurality of switch parameter combination samples as a current switch parameter combination sample;
[0083] S206, determining a first reward mechanism parameter of the initial agent, the first reward mechanism parameter indicating that after the initial agent performs an action, first feedback data that changes over time within a historical time range returned by the first interactive environment is obtained, and reward data of the initial agent under the current switch parameter combination sample is determined according to the first feedback data;
[0084] S208, iteratively training the initial agent in the first interactive environment based on the first solution space, the first action configuration parameter, and the first reward mechanism parameter;
[0085] S210: If the preset convergence condition is met, the initial intelligent agent is determined as the target intelligent agent.
[0086] Specifically, resource providers may propose corresponding plans based on their own management requirements for resource assets. For example, some resource providers require that their resources cannot be lower than a specific threshold within a historical time range, or some resource providers plan to achieve a specific target ratio of resource conversion into assets at a certain point in the future or within a time period. These plans reflect the personalized needs and goals of resource providers for resource management, and are factors that intelligent agents need to consider when making decisions.
[0087] To this end, this embodiment proposes to determine the first solution space of the initial agent, and use the first solution space as the range limit for the initial agent to explore the switch parameter combination samples. In the first solution space, the initial agent can freely try different switch parameter combination samples, so as to ensure that the exploration process of the initial agent is carried out within a reasonable range, avoiding blind search and waste of resources.
[0088] Regarding the process of determining the first solution space of the initial agent, in some possible implementations, the first solution space of the initial agent can be obtained based on the analysis of the first transaction constraint data of each resource provider in the historical time range. Among them, the first transaction constraint data reflects the constraints and restrictions of each resource provider on the resource provision ratio, resource conversion conditions, risk management requirements, etc. of the target transaction within the historical time range. In the first solution space, the initial agent can freely try different switch parameter combination samples, observe their impact on the target transaction, and adjust its strategy and behavior according to the first feedback data of the first interactive environment. At the same time, since the first solution space is determined based on the first transaction constraint data of each resource provider, it can ensure that the solution explored by the initial agent meets the requirements and expectations of each resource provider, which helps to optimize the resource allocation and management of each resource provider.
[0089] Furthermore, it is necessary to determine the first action configuration parameter of the initial agent. The first action configuration parameter indicates the action of the initial agent, and the action is to select a switch parameter combination sample from multiple switch parameter combination samples as the current switch parameter combination sample. Specifically, the first action configuration parameter can define the strategy and rules of the initial agent in selecting switch parameter combination samples. For example, a random selection method can be used to allow the initial agent to randomly select a switch parameter combination sample as the current switch parameter combination sample in each iteration; or, a predictive selection can be performed based on certain algorithms or models, so that the initial agent selects the switch parameter combination sample that is most likely to bring high rewards as the current switch parameter combination sample based on historical data and pattern recognition results.
[0090] Furthermore, it is necessary to determine the first reward mechanism parameter of the initial agent. Among them, the first reward mechanism parameter indicates that after the initial agent performs an action, the first feedback data returned by the first interactive environment that changes over time within the historical time range is obtained, and the reward data of the initial agent under the current switch parameter combination sample is determined according to the first feedback data. Specifically, the first reward mechanism parameter may include the definition of the reward function and the calculation method of the reward value. The reward function can be determined based on factors such as the impact effect data of the switch parameter combination sample on the target transaction within the historical time range, the resource asset status of the resource provider, and the planning and needs of the resource provider, so as to reflect the quality and degree of excellence of the initial agent's execution of the action. The calculation method of the reward value can be determined based on the reward function and the first feedback data. For example, the reward value can be calculated based on the quantitative data of the benefits, risks, and evaluation of the target transaction after the initial agent performs the action. By reasonably setting the first reward mechanism parameter, the initial agent can be encouraged to explore and try new switch parameter combination samples to obtain higher rewards and better performance.
[0091] After determining the first solution space, the first action configuration parameters, and the first reward mechanism parameters, the initial agent can be iteratively trained in the first interactive environment based on the first solution space, the first action configuration parameters, and the first reward mechanism parameters. Specifically, the iterative training process can be: first, the initial agent is placed in the first interactive environment, and it is asked to select a switch parameter combination sample as the current switch parameter combination sample according to the first action configuration parameters; then, the action is executed and the state change and the first feedback data of the first interactive environment are observed; then, the reward data of the initial agent under the current switch parameter combination sample is calculated according to the first reward mechanism parameters and the first feedback data; finally, the strategy and behavior of the initial agent are updated according to the reward data and the reinforcement learning algorithm. Through continuous iterative training, the strategy of the initial agent will gradually stabilize and show better performance.
[0092] If the preset convergence conditions are met during the iterative training process, it means that the initial intelligent agent has been fully trained and its strategies and behaviors have been optimized to a certain extent. At this time, the initial intelligent agent can be determined as the target intelligent agent.
[0093] Regarding the above-mentioned convergence conditions, in some possible implementations, the strategy of the initial agent no longer changes significantly in multiple consecutive iterations, or the change range of the strategy is lower than a preset threshold, and it can be determined that the convergence conditions are met. In some possible implementations, the reward data obtained by the initial agent under the current switch parameter combination sample reaches or exceeds a preset maximum value, and it can be determined that the convergence conditions are met. In some possible implementations, the number of iterative training reaches the preset upper limit value, and it can be determined that the convergence conditions are met. Although the strategy of the initial agent may not be completely stable or the reward value has not reached the maximum value at this time, the training can be terminated to avoid overtraining and waste of resources. In some possible implementations, the performance of the initial agent is evaluated by a series of performance evaluation indicators (such as accuracy, recall rate, output score, etc., selected according to the specific application scenario). When the evaluation indicator reaches or exceeds the preset standard, it is considered that the initial agent has been trained, that is, it is determined that the convergence conditions are met. In some possible implementations, the state of the first interactive environment no longer changes significantly in multiple consecutive iterations, or the amplitude of the state change is lower than a preset threshold, indicating that the behavior of the initial intelligent agent has had a stable impact on the first interactive environment, and the environmental state tends to be stable, and it can be determined that the convergence condition is met.
[0094] In this embodiment, the first solution space limits the exploration scope of the initial intelligent agent, so that it can efficiently find a better solution within a reasonable area. The first action configuration parameter clarifies the selection strategy of the initial intelligent agent, so that it can try different switch parameter combinations in a targeted manner. The first reward mechanism parameter motivates the initial intelligent agent to select a better switch parameter combination through feedback data, thereby maximizing the reward it obtains. By determining the first solution space, the first action configuration parameter and the first reward mechanism parameter of the initial intelligent agent, and performing iterative training in the first interactive environment, the decision-making ability of the target intelligent agent obtained by subsequent training can be effectively improved, which helps resource providers to achieve more reasonable resource allocation and management.
[0095] In one embodiment, Figure 3 Step S204 of the illustrated embodiment is further refined and may include the following steps:
[0096] Binary encoding is performed on each switch parameter combination sample based on a preset step length to obtain each binary switch parameter combination sample, wherein the deviation between any two adjacent switch parameter combination samples in the multiple binary switch parameter combination samples is equal to the preset step length;
[0097] The target action configuration parameters of the initial intelligent agent are determined based on a preset step size, the target action configuration parameters indicate the target action of the initial intelligent agent, the target action is to select a switch parameter combination sample from multiple binary switch parameter combination samples based on the preset step size as the current switch parameter combination sample, the target action configuration parameters are one of the first action configuration parameters of the initial intelligent agent, and the target action is one of the actions of the initial intelligent agent.
[0098] Specifically, this embodiment introduces binary coding and preset step size to more accurately control the behavior and strategy of the initial agent when selecting switch parameter combination samples. First, each switch parameter combination sample is binary-coded based on the preset step size, and the switch parameter combination sample that may originally exist in a complex form is converted into a binary form to facilitate subsequent processing and calculation. Binary coding is a way of representing information as binary numbers (0 and 1), which is used here to represent the state of the switch parameter combination sample. Through binary coding, each switch parameter combination sample can be mapped to a unique binary number, so that it is convenient to compare, select and operate.
[0099] In the binary encoding process, ensure that the deviation between any two adjacent switch parameter combination samples in the binary multiple switch parameter combination samples is equal to the preset step size. The preset step size is a pre-set value that determines the difference between two adjacent switch parameter combination samples in the binary encoding space. By setting the preset step size, the stride of the initial agent when exploring the switch parameter combination samples can be controlled.
[0100] Furthermore, the target action configuration parameters of the initial agent are determined based on the preset step size. The target action configuration parameters are the rules and strategies that the initial agent needs to follow when performing actions, indicating how the initial agent should select a switch parameter combination sample as the current switch parameter combination sample. The target action is to select a switch parameter combination sample from multiple binary switch parameter combination samples based on the preset step size as the current switch parameter combination sample. In other words, each time the initial agent selects, it will jump to select the next switch parameter combination sample in the binary coding space according to the preset step size.
[0101] In this embodiment, since the switch parameter combination samples (or switch parameter combinations) exist in a discrete form, by introducing binary coding and preset step length, the complex switch parameter combination samples are converted into concise binary numbers, which is convenient for subsequent processing and calculation by the initial intelligent agent, and improves the speed and accuracy of data processing. The setting of the preset step length provides a clear step length for the exploration of the initial intelligent agent in the binary coding space, effectively controls the step length of the exploration process, and avoids blind search and waste of resources.
[0102] In one embodiment, Figure 3 Step S202 of the illustrated embodiment is further refined and may include the following steps:
[0103] Obtain the first transaction constraint data of each resource provider in the historical time range;
[0104] Based on the switch effect information sample, the resource asset information sample of each resource provider, and the first transaction constraint data, determine the upper and lower limits of assets of each resource provider in the historical time range;
[0105] Based on the analysis of the upper and lower limits of assets of each resource provider in the historical time range, the first solution space of the initial intelligent agent is obtained.
[0106] Specifically, the first transaction constraint data of each resource provider in the historical time range involved in this embodiment refers to the relevant data on the constraints and restrictions on resource provision, resource conversion, risk management, etc. involved in the execution of the target transaction by each resource provider in the historical time range. For example, the first transaction constraint data may include the resource provider's requirements on the resource provision ratio, the setting of resource conversion conditions, the assessment of risk tolerance, and the expectation of future resource asset status, etc. They exist in quantitative or non-quantitative form and are the basic data for determining the upper and lower limits of assets of each resource provider in the historical time range.
[0107] Regarding the process of obtaining the first transaction constraint data of each resource provider in the historical time range, in some possible implementations, data related to transaction constraints can be extracted from the document data of each resource provider as the first transaction constraint data of each resource provider in the historical time range. In some possible implementations, the first transaction constraint data of each resource provider in the historical time range is predetermined and stored locally, and the first transaction constraint data of each resource provider in the historical time range can be directly read locally.
[0108] Furthermore, based on the switch effect information sample and the resource asset information sample and the first transaction constraint data of each resource provider, the upper and lower limits of the assets of each resource provider in the historical time range are determined. Specifically, there is a mapping relationship between the switch effect information sample, the resource asset information sample of each resource provider, the first transaction constraint data, and the upper and lower limits of the assets of each resource provider in the historical time range. On the basis of determining the switch effect information sample, the resource asset information sample of each resource provider, and the first transaction constraint data, the upper and lower limits of the assets of each resource provider in the historical time range can be determined based on the mapping relationship.
[0109] In some possible implementations, the mapping relationship between the switch effect information sample, the resource asset information sample of each resource provider, the first transaction constraint data, and the upper and lower limits of the assets of each resource provider in the historical time range can be recorded in the form of a function. Substituting the switch effect information sample, the resource asset information sample of each resource provider, and the first transaction constraint data into the corresponding function, the upper and lower limits of the assets of each resource provider in the historical time range can be solved.
[0110] In some possible implementations, the mapping relationship between the switch effect information sample, the resource asset information sample of each resource provider, the first transaction constraint data, and the upper and lower limits of the assets of each resource provider in the historical time range can be recorded in the form of a database. By constructing a query statement with the switch effect information sample, the resource asset information sample of each resource provider, and the first transaction constraint data, the upper and lower limits of the assets of each resource provider in the historical time range can be queried from the relevant database.
[0111] In some possible implementations, the mapping relationship between the switch effect information sample, the resource asset information sample of each resource provider, the first transaction constraint data, and the upper and lower limits of the assets of each resource provider in the historical time range can be recorded in the form of a machine learning model. The query statement constructed by the switch effect information sample, the resource asset information sample of each resource provider, and the first transaction constraint data is input into the corresponding machine learning model, that is, the upper and lower limits of the assets of each resource provider in the historical time range output by the machine learning model are obtained.
[0112] Then, the first solution space of the initial agent can be obtained based on the asset upper and lower limits of each resource provider in the historical time range. This process can be specifically manifested as: taking the asset upper and lower limits of each resource provider as constraints, combining the switch effect information samples and resource asset information samples, and constructing a virtual or simulated environment (i.e., the first solution space). In this environment, the initial agent can freely try different switch parameter combination samples, but the scope of its exploration is limited by the asset upper and lower limits of each resource provider in the historical time range.
[0113] In this embodiment, by obtaining the first transaction constraint data of each resource provider in the historical time range, and combining the switch effect information sample and the resource asset information sample, the upper and lower limits of the assets of each resource provider in the historical time range are determined, and then the first solution space of the initial intelligent agent is obtained through analysis. In this way, not only the rationality of the initial intelligent agent in exploring the switch parameter combination is ensured, blind search and waste of resources are avoided, but also the adaptability of the initial intelligent agent to the dynamic needs of the resource provider is improved.
[0114] See also Figure 4 , is a schematic diagram of a process for obtaining a switching effect information sample according to an embodiment of this specification. Figure 4 As shown, the method of the embodiment of this specification may include the following steps S302 to S308, and steps S302 to S308 may be used as Figure 2 The detailed steps of step S102 of the illustrated embodiment.
[0115] S302, determining a plurality of first switch parameter combination samples and a plurality of second switch parameter combination samples;
[0116] S304, obtaining the effect data of each first switch parameter combination sample in a plurality of first switch parameter combination samples on the target transaction within a historical time range;
[0117] S306, calling a preset linear regression model based on each first switch parameter combination sample, the impact effect data of each first switch parameter combination sample, and multiple second switch parameter combination samples, to obtain the impact effect data of each second switch parameter combination sample on the target transaction within a historical time range among the multiple second switch parameter combination samples output by the linear regression model;
[0118] S308, constructing a switch effect information sample based on each first switch parameter combination sample, the influence effect data of each first switch parameter combination sample, each second switch parameter combination sample, and the influence effect data of each second switch parameter combination sample.
[0119] Specifically, in some cases, a large amount of computing resources and time may be required to obtain the impact effect data of the switch parameter combination samples on the target transaction within the historical time range, and the control process of the resource provision ratio of each resource provider usually has time-limited requirements. In order to more quickly obtain the impact effect data of all switch parameter combination samples required to construct the switch effect information sample on the target transaction within the historical time range, this embodiment proposes to use a linear regression model to reduce the consumption of computing resources and time.
[0120] Among them, the linear regression model refers to a mathematical model or machine learning model based on statistics, which is used to establish a linear relationship between independent variables (here, it can be the first switch parameter combination sample, the effect data of the first switch parameter combination sample on the target transaction within the historical time range, and the second switch parameter combination sample) and the dependent variable (here, it can be the effect data of the second switch parameter combination sample on the target transaction within the historical time range). The linear regression model is pre-set and can be trained and optimized based on historical data or experience to ensure the accuracy and reliability of its prediction results.
[0121] First, among the multiple switch parameter combination samples required to construct the switch effect information sample, multiple first switch parameter combination samples and multiple second switch parameter combination samples are respectively determined, and the first switch parameter combination samples and the second switch parameter combination samples are not repeated.
[0122] Furthermore, the effect data of each first switch parameter combination sample in the plurality of first switch parameter combination samples on the target transaction within the historical time range is obtained. Regarding this process, in a possible implementation, the effect data of each first switch parameter combination sample on the target transaction within the historical time range can be recorded within the historical time range, and the effect data of each first switch parameter combination sample on the target transaction within the historical time range that is pre-recorded can be directly read.
[0123] Then, based on each first switch parameter combination sample, the impact effect data of each first switch parameter combination sample and multiple second switch parameter combination samples, the preset linear regression model can be called to obtain the impact effect data of each second switch parameter combination sample on the target transaction within the historical time range among the multiple second switch parameter combination samples output by the linear regression model.
[0124] Furthermore, a switch effect information sample can be constructed by using the previously determined first switch parameter combination samples, the influence effect data of the first switch parameter combination samples, the second switch parameter combination samples, and the influence effect data of the second switch parameter combination samples.
[0125] In this embodiment, a linear regression model is used to predict the impact effect data of the second switch parameter combination sample on the target transaction, and a switch effect information sample is constructed based on these predicted data and the actual impact effect data of the first switch parameter combination sample, thereby significantly improving the efficiency and accuracy of data acquisition.
[0126] See also Figure 5 , provides a flow chart of obtaining a resource asset information sample for the embodiment of this specification. Figure 5 As shown, the method of the embodiment of this specification may include the following steps S402 to S404, and steps S402 to S404 may be used as Figure 2 The detailed steps of step S104 of the illustrated embodiment.
[0127] S402, obtaining resource asset status prediction data and resource asset status observation data of each resource provider within a historical time range;
[0128] S404, determining a resource asset information sample of each resource provider within a historical time range based on the resource asset status prediction data and resource asset status observation data of each resource provider within a historical time range.
[0129] Specifically, regarding the process of obtaining resource asset information samples of each resource provider in a historical time range, on the one hand, it is necessary to obtain resource asset status prediction data of each resource provider in a historical time range. Among them, the resource asset status prediction data of each resource provider in a historical time range may include data predicted by an artificial intelligence model, a statistical model or a related mathematical model, and may also include some prior data based on manual input, which reflects the possible resource asset status of the resource provider, such as the expected increase or decrease of resources, the expected change of assets, and the expected conversion between resources and assets. In some possible implementations, machine learning algorithms, such as time series analysis, regression analysis, etc., can be used to predict the resource asset status of each resource provider based on historical data, thereby obtaining resource asset status prediction data.
[0130] On the one hand, it is necessary to obtain the resource asset status observation data of each resource provider in the historical time range. Among them, the resource asset status observation data of each resource provider in the historical time range may include data obtained through actual observation and recording, and may also include some prior data based on manual input. These data reflect the actual resource asset status of the resource provider in the past, such as the actual increase or decrease of resources, the actual change of assets, and the actual conversion between resources and assets. In some possible implementation methods, the historical resource asset status data of each resource provider can be obtained from the data platform of each resource provider through a data interface or data capture technology as the resource asset status observation data.
[0131] Furthermore, based on the resource asset status prediction data and resource asset status observation data of each resource provider in the historical time range, the resource asset information samples of each resource provider in the historical time range are determined. Specifically,
[0132] Then, for each resource provider, the resource asset status prediction data and resource asset status observation data of the resource provider in the historical time range are integrated to obtain the resource asset information sample of the resource provider in the historical time range. In this way, the resource asset information sample of each resource provider in the historical time range can be determined.
[0133] In this embodiment, the resource asset status prediction data and resource asset status observation data of each resource provider in the historical time range are obtained, and the resource asset information samples are determined based on these two types of data. Among them, the resource asset status prediction data reflects the possible resource asset status of the resource provider, and the resource asset status observation data reflects the actual resource asset status of the resource provider. The combination of the two can more comprehensively characterize the resource asset status of the resource provider, and provide comprehensive and accurate data support for the training of the initial intelligent agent. In this way, the initial intelligent agent can more accurately learn the characteristics and laws of different resource providers, so as to make more reasonable decisions.
[0134] See also Figure 6 , is a flow chart of a data configuration method provided in the embodiment of this specification. Figure 6 As shown, the method of the embodiment of this specification may include the following steps S502 to S510.
[0135] S502, obtaining switch effect information, the switch effect information including multiple switch parameter combinations and effect data of each of the multiple switch parameter combinations on the target transaction within a target time range, the switch parameter combination indicating a resource provision ratio of each of the multiple resource providers to convert resources into assets based on the target transaction;
[0136] S504, obtaining resource asset information of each resource provider within the target time range;
[0137] S506, constructing a second interactive environment based on the switch effect information and the resource asset information of each resource provider;
[0138] S508, calling the target agent based on the second interactive environment, so that the target agent determines a target switch parameter combination from a plurality of switch parameter combinations;
[0139] S510: configuring a target resource provision ratio of each resource provider based on a target switch parameter combination.
[0140] Specifically, the target time range involved in this embodiment refers to a certain time interval at the current moment, or a certain time interval in the future, such as one hour, one day, one month or several months in the future; the impact effect data of each switch parameter combination sample on the target transaction within the target time range refers to the quantitative data of the execution effect of each switch parameter combination sample on the target transaction within the target time range.
[0141] Specifically, on one hand, the switch effect information includes multiple switch parameter combinations, and the switch parameter combination refers to a set of data formed by combining the state parameters of each switch in the multiple switches. On the one hand, the switch effect information includes the effect data of each switch parameter combination in the multiple switch parameter combinations on the target transaction within the target time range, wherein the target transaction refers to one of the multiple transactions involved in financial activities, for example, the target transaction may be a credit transaction; the target time range refers to a certain time interval in the future, such as the next day, month or several months; the effect data of each switch parameter combination on the target transaction within the target time range refers to the quantitative data of the execution effect of each switch parameter combination on the target transaction within the target time range.
[0142] In some possible cases, the types of impact effect data may include impact effect data of resource providers and impact effect data of resource acquirers. Resource acquirers refer to the party that acquires resources through target transactions and converts them into assets in financial activities, such as borrowers in credit transactions. For example, the impact effect data of resource providers may include quantitative data on the income, risk, evaluation, resource liquidity, etc. of resource providers under the influence of corresponding switch parameter combinations; similarly, the impact effect data of resource acquirers may include quantitative data on the income, risk, evaluation, etc. of resource providers under the influence of corresponding switch parameter combinations.
[0143] A resource provider refers to a party that provides resources for a target transaction in a financial activity. A switch parameter combination indicates the resource provision ratio of each resource provider among multiple resource providers to convert resources into assets based on the target transaction, which means that there is a mapping relationship between the switch parameter combination and the resource provision ratio of each resource provider. A specific switch parameter combination indicates a specific resource provision ratio of each resource provider, and different switch parameter combinations indicate different specific resource provision ratios of each resource provider.
[0144] Regarding the process of obtaining switch effect information, on the one hand, it is necessary to obtain multiple switch parameter combinations. Among them, multiple switch parameter combinations can be obtained by analyzing all possible state parameters of the corresponding multiple switches. For example, assuming that there are 3 enumerated switches, and each enumerated switch has 3 state parameters, then by combining all possible state parameters of these 3 enumerated switches, 3 to the power of 3, that is, 27 different switch parameter combinations, can be obtained. Each switch parameter combination represents a specific switch state configuration, and the switch state configuration determines the resource provision ratio of each resource provider in the target transaction.
[0145] On the other hand, it is necessary to obtain the effect data of each switch parameter combination on the target transaction within the target time range. In some possible cases, the effect data of each switch parameter combination on the target transaction within the target time range can be determined based on historical data analysis. In some possible cases, the associated time range corresponding to the target time range can be determined from the historical time range, and then the effect data of each switch parameter combination sample on the target transaction within the associated time range can be determined from the effect data of each switch parameter combination sample on the target transaction within the historical time range. Then, the effect data of each switch parameter combination sample on the target transaction within the associated time range is determined as the effect data of each switch parameter combination on the target transaction within the target time range.
[0146] Based on multiple switch parameter combinations and the effect data of each switch parameter combination in the multiple switch parameter combinations on the target transaction within the target time range, the switch effect information can be determined.
[0147] The resource asset information of each resource provider within the target time range refers to the quantitative data on the resource status, asset status, and conversion between resources and assets of each resource provider within the target time range, which is used to reflect the resource asset status and operation status of the resource provider within the target time range.
[0148] It should be noted that the data related to the resource asset information of each resource provider within the target time range may include at least one of the associated prediction data and observation data. Among them, prediction data refers to data predicted by artificial, machine learning models or related mathematical models, reflecting the possible resource asset status of the resource provider; observation data refers to data obtained through actual observation and recording, reflecting the actual resource asset status of the resource provider.
[0149] Regarding the process of obtaining the resource asset information of each resource provider within the target time range, in a possible implementation, the relevant documents of each resource provider can be obtained, and then the resource asset information of each resource provider within the target time range can be extracted from the relevant documents of each resource provider. In a possible implementation, the resource asset information of each resource provider within the target time range can be obtained from the data platform of each resource provider using a data interface or data crawling technology. In a possible implementation, the resource asset information of each resource provider within the target time range is predetermined and stored locally, and the resource asset information of each resource provider within the target time range can be directly read locally.
[0150] In this embodiment, the second interactive environment is constructed based on the switch effect information and the resource asset information of each resource provider, which simulates the impact of the switch parameter combination on the target transaction in the financial activity, as well as the resource asset status of each resource provider. Regarding the process of constructing the second interactive environment based on the switch effect information and the resource asset information of each resource provider, in some possible implementations, the switch effect information and the resource asset information of each resource provider can be mapped to the pre-created initial interactive environment to obtain the second interactive environment. Specifically, on the one hand, the switch effect information is mapped to the action space of the initial interactive environment, and on the other hand, the resource asset information of each resource provider is mapped to the state space of the initial interactive environment, and then the initial interactive environment is determined as the second interactive environment, providing a basis for subsequent reinforcement learning training.
[0151] Regarding the process of calling the target agent based on the second interactive environment so that the target agent determines the target switch parameter combination from multiple switch parameter combinations, the target agent is first placed in the constructed second interactive environment so that the target agent can perceive the resource asset information of each resource provider within the target time range, the effect data of the switch parameter combination on the target transaction within the target time range and other information from the second interactive environment. Then, the target agent begins to explore the second interactive environment according to the strategy learned in the training process. During the exploration process, the target agent will evaluate the impact of the switch parameter combination on the target transaction one by one, and consider the resource asset status and planning of each resource provider. This evaluation process is based on the reward mechanism parameters and action configuration parameters learned by the target agent in the training process. Then, the target agent selects a better switch parameter combination from multiple switch parameter combinations as the target switch parameter combination according to the evaluation results. Finally, the target agent outputs the selected target switch parameter combination as the configuration basis for the resource provision ratio of each resource provider.
[0152] Regarding the process of configuring the target resource provision ratio of each resource provider based on the target switch parameter combination, firstly, the target resource provision ratio of each resource provider in the target transaction is determined according to the target switch parameter combination output by the target agent. Then, the target resource provision ratio of each resource provider in the target transaction is applied to each resource provider to guide its resource investment and configuration in the target transaction. In some possible cases, the configured resource provision ratio can also be monitored and evaluated to ensure that it meets the goals and requirements of financial activities.
[0153] For example, assuming that the target time range is from the 10th to the 15th of a certain month, the resource providers are resource provider 1, resource provider 2, and resource provider 3. Based on the target switch parameter combination, it is determined that the target resource provision ratio of resource provider 1 from the 10th to the 15th of the month is 10%, the target resource provision ratio of resource provider 2 from the 10th to the 15th of the month is 30%, and the target resource provision ratio of resource provider 3 from the 10th to the 15th of the month is 60%. Then, after configuring the target resource provision ratios of resource provider 1, resource provider 2, and resource provider 3 from the 10th to the 15th of the month, resource provider 1, resource provider 2, and resource provider 3 will provide resources for the target transaction at the ratios of 10%, 30%, and 60% respectively from the 10th to the 15th of the month.
[0154] In some possible cases, the target time range is divided into multiple time periods, and the target agent determines the target switch parameter combination from multiple switch parameter combinations, including the target switch parameter combination corresponding to each time period in the multiple time periods, then the target resource provision ratio of each resource provider in each time period can be configured based on the target switch parameter combination corresponding to each time period. For example, assuming that the target time range is from the 10th to the 15th of a certain month, the target time range is further divided into the 10th to the 12th (time period 1), and the 13th to the 15th (time period 2). The resource providers are resource provider 1, resource provider 2, and resource provider 3, wherein the target resource provision ratios of resource provider 1, resource provider 2, and resource provider 3 in time period 1 are 10%, 30%, and 60%, respectively, and the target resource provision ratios of resource provider 1, resource provider 2, and resource provider 3 in time period 2 are 20%, 40%, and 40%, respectively. Then, after configuring the target resource provision ratios of resource provider 1, resource provider 2, and resource provider 3 in time period 1, resource provider 1, resource provider 2, and resource provider 3 will provide resources for the target transaction in the ratios of 10%, 30%, and 60% respectively during time period 1; after configuring the target resource provision ratios of resource provider 1, resource provider 2, and resource provider 3 in time period 2, resource provider 1, resource provider 2, and resource provider 3 will provide resources for the target transaction in the ratios of 20%, 40%, and 40% respectively during time period 2.
[0155] In this embodiment, first, a switch effect information sample is obtained, and the switch effect information sample includes multiple switch parameter combination samples and the effect data of each switch parameter combination sample on the target transaction within the target time range. Secondly, a resource asset information sample of each resource provider within the target time range is obtained. A second interactive environment is constructed based on the switch effect information sample and the resource asset information sample of each resource provider, so that the second interactive environment can reflect the impact of the switch parameter combination sample on the target transaction, as well as the resource asset situation of each resource provider. By calling the target intelligent agent trained by reinforcement learning, a better target switch parameter combination can be determined quickly and accurately from multiple switch parameter combinations. Finally, based on the configuration of the target switch parameter combination, it can ensure that each resource provider provides resources in an appropriate proportion, which helps to optimize the resource allocation and management of each resource provider.
[0156] See also Figure 7 , which is a schematic diagram of a process for determining target switch parameters according to an embodiment of this specification. Figure 7 As shown, the method of the embodiment of this specification may include the following steps S602 to S608, and steps S602 to S608 may be used as Figure 6The detailed steps of step S508 of the illustrated embodiment.
[0157] S602, determining a second solution space of the target intelligent agent;
[0158] S604, determining a second action configuration parameter of the target agent, where the second action configuration parameter indicates an action of the target agent, where the action is selecting a switch parameter combination from a plurality of switch parameter combinations as a current switch parameter combination;
[0159] S606, determining a second reward mechanism parameter of the target agent, the second reward mechanism parameter indicating that after the target agent performs an action, second feedback data that changes over time within a target time range returned by the second interactive environment is obtained, and reward data of the target agent under the current switch parameter combination is determined according to the second feedback data;
[0160] S608, based on the second solution space, the second action configuration parameters and the second reward mechanism parameters, calling the target agent in the second interactive environment, so that the target agent determines a target switch parameter combination from a plurality of switch parameter combinations.
[0161] Specifically, resource providers may propose corresponding plans based on their own management requirements for resource assets. For example, some resource providers require that their resources cannot be lower than a specific threshold within the target time range, or some resource providers plan to achieve a specific target ratio of resource conversion into assets at a certain point in the future or time period. These plans reflect the personalized needs and goals of resource providers for resource management, and are factors that intelligent agents need to consider when making decisions.
[0162] To this end, this embodiment proposes to determine a second solution space of the target agent, and use the second solution space as a range limit for the target agent to explore the switch parameter combination. In the second solution space, the target agent can freely try different switch parameter combinations, so as to ensure that the exploration process of the target agent is carried out within a reasonable range, avoiding blind search and waste of resources.
[0163] Regarding the process of determining the second solution space of the target agent, in some possible implementations, the second solution space of the target agent can be obtained based on the analysis of the second transaction constraint data of each resource provider within the target time range. Among them, the second transaction constraint data reflects the constraints and restrictions of each resource provider on the resource provision ratio, resource conversion conditions, risk management requirements, etc. of the target transaction within the target time range. In the second solution space, the target agent can freely try different switch parameter combinations, observe their impact on the target transaction, and adjust its strategy and behavior according to the second feedback data of the second interactive environment. At the same time, since the second solution space is determined based on the second transaction constraint data of each resource provider, it can ensure that the solution explored by the target agent meets the requirements and expectations of each resource provider, which helps to optimize the resource allocation and management of each resource provider.
[0164] Furthermore, it is necessary to determine the second action configuration parameter of the target agent. The second action configuration parameter indicates the action of the target agent, and the action is to select a switch parameter combination from multiple switch parameter combinations as the current switch parameter combination. Specifically, the second action configuration parameter can define the strategy and rules of the target agent in selecting the switch parameter combination. For example, a random selection method can be used to allow the target agent to randomly select a switch parameter combination as the current switch parameter combination in each iteration; or, a predictive selection can be performed based on certain algorithms or models, so that the target agent selects the switch parameter combination that is most likely to bring high rewards as the current switch parameter combination based on historical data and pattern recognition results.
[0165] Furthermore, it is necessary to determine the second reward mechanism parameter of the target agent. Among them, the second reward mechanism parameter indicates that after the target agent performs an action, the second feedback data returned by the second interactive environment that changes over time within the target time range is obtained, and the reward data of the target agent under the current switch parameter combination is determined according to the second feedback data. Specifically, the second reward mechanism parameter may include the definition of the reward function and the calculation method of the reward value. The reward function can be determined based on factors such as the impact effect data of the switch parameter combination on the target transaction within the target time range, the resource asset status of the resource provider, and the planning and needs of the resource provider, so as to reflect the quality and degree of excellence of the target agent's execution of the action. The calculation method of the reward value can be determined based on the reward function and the second feedback data. For example, the reward value can be calculated based on the quantitative data of the benefits, risks, and evaluation of the target transaction after the target agent performs the action. By reasonably setting the second reward mechanism parameter, the target agent can be encouraged to explore and try new switch parameter combinations to obtain higher rewards and better performance.
[0166] The process of calling the target agent in the second interactive environment based on the second solution space, the second action configuration parameters and the second reward mechanism parameters so that the target agent determines the target switch parameter combination from multiple switch parameter combinations is specifically as follows: first, the target agent is placed in the constructed second interactive environment. The second interactive environment has been fully simulated and constructed based on the switch effect information and the resource asset information of each resource provider, and can accurately reflect the impact of the switch parameter combination on the target transaction in financial activities and the resource asset status of each resource provider. Then, according to the predetermined second solution space, the exploration scope of the target agent is restricted. The second solution space is obtained based on the analysis of the second transaction constraint data of each resource provider in the target time range. It limits the reasonable range of the target agent when selecting the switch parameter combination, ensuring that the exploration process of the target agent will neither be too conservative and miss the better solution, nor be too aggressive and waste resources. Then, the target agent starts to select one from multiple switch parameter combinations as the current switch parameter combination according to the strategy and rules defined by the second action configuration parameters. This selection process can be random or based on the prediction results of certain algorithms or models. Regardless of the method, the goal is to enable the target agent to efficiently explore the second solution space and find potential better solutions. After the target agent performs an action, that is, after selecting a switch parameter combination as the current switch parameter combination, it will obtain the second feedback data that changes over time within the target time range from the second interactive environment. The second feedback data reflects the impact of the current switch parameter combination on the target transaction and the changes in the resource asset status of each resource provider. Subsequently, according to the reward function and reward value calculation method defined by the second reward mechanism parameters, the target agent calculates the reward data under the current switch parameter combination. The reward function can comprehensively consider the impact effect data of the switch parameter combination on the target transaction, the resource asset status of the resource provider, and the planning and needs of the resource provider to reflect the quality and degree of the target agent's execution of the action. The calculation of the reward value is based on the reward function and the second feedback data, which provides direct feedback to the target agent and guides its subsequent exploration and decision-making. Finally, the target agent updates its strategy and behavior based on the reward data and the reinforcement learning algorithm, and continues to iterate to find a better switch parameter combination to maximize its reward. When the preset end condition is met, the target agent will stop iterating and output the final target switch parameter combination as the basis for configuring the resource provision ratio of each resource provider.
[0167] In this embodiment, the second solution space limits the exploration scope of the target agent, so that it can efficiently find a better solution within a reasonable area. The second action configuration parameters clarify the selection strategy of the target agent, so that it can try different switch parameter combinations in a targeted manner. The second reward mechanism parameters motivate the target agent to select a better switch parameter combination through feedback data, thereby maximizing the reward it obtains. Based on the second solution space, the second action configuration parameters and the second reward mechanism parameters, the target agent can efficiently and accurately determine the target switch parameter combination in the second interactive environment.
[0168] See also Figure 8 , provides a flow chart of determining a second solution space for the embodiment of this specification. Figure 8 As shown, the method of the embodiment of this specification may include the following steps S702 to S706, and steps S702 to S706 may be used as Figure 7 The detailed steps of step S602 of the illustrated embodiment.
[0169] S702, obtaining second transaction constraint data of each resource provider within a target time range;
[0170] S704, determining the upper and lower limits of assets of each resource provider within the target time range based on the switch effect information, the resource asset information of each resource provider, and the second transaction constraint data;
[0171] S706, obtaining a second solution space of the target intelligent agent based on the upper and lower limits of the assets of each resource provider within the target time range.
[0172] Specifically, the second transaction constraint data of each resource provider within the target time range involved in this embodiment refers to the relevant data on the constraints and restrictions on resource provision, resource conversion, risk management, etc. involved in the execution of the target transaction by each resource provider within the target time range. For example, the second transaction constraint data may include the resource provider's requirements on the resource provision ratio, the setting of resource conversion conditions, the assessment of risk tolerance, and the expectation of future resource asset status, etc. They exist in quantitative or non-quantitative form and are the basic data for determining the upper and lower limits of assets of each resource provider within the target time range.
[0173] Regarding the process of obtaining the second transaction constraint data of each resource provider in the target time range, in some possible implementations, data related to transaction constraints can be extracted from the document data of each resource provider as the second transaction constraint data of each resource provider in the target time range. In some possible implementations, the second transaction constraint data of each resource provider in the target time range is predetermined and stored locally, and the second transaction constraint data of each resource provider in the target time range can be directly read locally.
[0174] Furthermore, based on the switch effect information, the resource asset information of each resource provider, and the second transaction constraint data, the upper and lower limits of the assets of each resource provider in the target time range are determined. Specifically, there is a mapping relationship between the switch effect information, the resource asset information of each resource provider, and the second transaction constraint data, and the upper and lower limits of the assets of each resource provider in the target time range. On the basis of determining the switch effect information, the resource asset information of each resource provider, and the second transaction constraint data, the upper and lower limits of the assets of each resource provider in the target time range can be determined based on the mapping relationship.
[0175] In some possible implementations, the mapping relationship between the switch effect information, the resource asset information of each resource provider, the second transaction constraint data, and the upper and lower limits of the assets of each resource provider in the target time range can be recorded in the form of a function. Substituting the switch effect information, the resource asset information of each resource provider, and the second transaction constraint data into the corresponding function, the upper and lower limits of the assets of each resource provider in the target time range can be solved.
[0176] In some possible implementations, the mapping relationship between the switch effect information, the resource asset information of each resource provider, the second transaction constraint data, and the upper and lower limits of the assets of each resource provider in the target time range can be recorded in the form of a database. By constructing a query statement with the switch effect information, the resource asset information of each resource provider, and the second transaction constraint data, the upper and lower limits of the assets of each resource provider in the target time range can be queried from the relevant database.
[0177] In some possible implementations, the mapping relationship between the switch effect information, the resource asset information of each resource provider, the second transaction constraint data, and the upper and lower limits of the assets of each resource provider in the target time range can be recorded in the form of a machine learning model. The switch effect information, the resource asset information of each resource provider, and the second transaction constraint data are used to construct a query statement and input it into the corresponding machine learning model, that is, the upper and lower limits of the assets of each resource provider in the target time range output by the machine learning model are obtained.
[0178] Then, the second solution space of the target agent can be obtained based on the upper and lower limits of the assets of each resource provider in the target time range. This process can be specifically manifested as: taking the upper and lower limits of the assets of each resource provider as constraints, combining the switch effect information and resource asset information, and constructing a virtual or simulated environment (i.e., the second solution space). In this environment, the target agent can freely try different switch parameter combinations, but the scope of its exploration is limited by the upper and lower limits of the assets of each resource provider in the target time range.
[0179] In this embodiment, by obtaining the second transaction constraint data of each resource provider in the target time range, and combining the switch effect information sample and the resource asset information sample, the upper and lower limits of the assets of each resource provider in the target time range are determined, and then the second solution space of the target agent is obtained through analysis. In this way, not only the rationality of the target agent in exploring the switch parameter combination is ensured, blind search and waste of resources are avoided, but also the adaptability of the target agent to the dynamic needs of the resource provider is improved.
[0180] In one embodiment, based on Figure 8 The embodiment shown, Figure 7 Step S608 of the illustrated embodiment is further refined and may include the following steps:
[0181] Based on the switch effect information, the resource asset information of each resource provider, and the second transaction constraint data, determine a function curve of the asset upper limit and the asset lower limit of the target resource provider changing over time within the target time range;
[0182] Based on the function curve, determining at least one target time point of the target resource provider within the target time range;
[0183] Dividing the target time range based on at least one target time point to obtain multiple time periods;
[0184] Based on multiple time periods, a second solution space, second action configuration parameters, and second reward mechanism parameters, a target agent is called in a second interactive environment so that the target agent determines, from a plurality of switch parameter combinations, a target switch parameter combination corresponding to each of the multiple time periods.
[0185] Specifically, within the target time range, the resource asset status of different resource providers may change dynamically, and the resource asset planning of different resource providers may also change dynamically. If the target time range is a long time interval, then only determining one target switch parameter combination may be difficult to adapt to the dynamic needs of the resource provider. To this end, this embodiment proposes dividing the target time range into multiple time periods, determining the target switch parameter combination corresponding to each time period, so as to adapt to the dynamic needs of the resource provider as much as possible.
[0186] First, the target resource provider involved in this embodiment refers to one or more of the multiple resource providers. Figure 8 As can be seen from the illustrated embodiment, based on the switch effect information, resource asset information of each resource provider, and the second transaction constraint data, the upper and lower limits of assets of each resource provider in the target time range can be determined. Furthermore, the upper and lower limits of assets of the target resource provider in the target time range can be determined from the upper and lower limits of assets of each resource provider in the target time range.
[0187] It is understandable that within the target time range, the upper and lower limits of the target resource provider’s assets change over time. Therefore, the target resource provider’s assets can be determined by performing a time series analysis on the upper and lower limits of the target resource provider’s assets within the target time range.
[0188] In the expected effect, a certain time period should make the upper and lower limits of the assets of the target resource provider stable or meet certain requirements under the action of the corresponding target switch parameter combination. Therefore, the target time point involved in this embodiment should be a time point that can reflect the significant change in the upper and lower limits of the assets of the target resource provider within the target time range or reach a specific condition (such as reaching a preset threshold, a trend reversal, etc.). These target time points are crucial for dividing time periods and determining the target switch parameter combinations corresponding to each time period, because they mark important change nodes of the asset status of the target resource provider, which helps the target intelligent agent to more accurately adapt to the dynamic needs of the target resource provider.
[0189] Therefore, the process of determining at least one target time point of the target resource provider within the target time range based on the function curve can be specifically expressed as follows: according to the characteristics of the function curve (such as slope change, extreme points, etc.), identify the target time point where the upper and lower limits of the assets of the target resource provider within the target time range change significantly or meet specific conditions.
[0190] Further, the target time range is divided based on at least one target time point to obtain multiple time periods. For example, a set of time points to be divided can be constructed based on the start time point, the end time point and at least one target time point of the target time range, and any two adjacent time points in the set of time points to be divided can constitute a corresponding time period.
[0191] To facilitate understanding of the process of determining the target time point and dividing the time period in this embodiment, please refer to Fig. 9 , Fig. 9 It is an example schematic diagram of a function curve provided by an embodiment of this specification. It involves a plane coordinate system, whose vertical axis is assets and horizontal axis is time. The function curve in the plane coordinate system includes a first function curve in which the upper limit of the assets of the target resource provider changes with time within the target time range, and a second function curve in which the lower limit of the assets of the target resource provider changes with time within the target time range. In addition, the target time range is defined by a starting time point and an ending time point, and there is also an asset warning line in the plane coordinate system (representing a preset threshold value of the lower limit of the assets). The first function curve has two interval curves below the asset warning line, which indicates that the lower limit of the assets of the target resource provider is too low, so the target time point 1 and the target time point 2 can be determined based on these two interval curves. Further, the starting time point and the target time point 1 can constitute a corresponding time period, the target time point 1 and the target time point 2 can constitute a corresponding time period, and the target time point 2 and the ending time point can constitute a corresponding time period.
[0192] Furthermore, the target agent is placed in the constructed second interactive environment. For any target time period among the multiple time periods, the exploration scope of the target agent is limited according to the second solution space. Then, the target agent starts to select a switch parameter combination as the switch parameter combination of the target time period from multiple switch parameter combinations according to the strategy and rules defined by the second action configuration parameters. This selection process can be random or based on the prediction results of certain algorithms or models. The target agent will select the most appropriate switch parameter combination according to the characteristics of the target time period (such as the change trend of the upper and lower limits of assets, the planning of resource providers, etc.) and the strategy learned during the training process. After the target agent executes the action, that is, after selecting a switch parameter combination as the switch parameter combination of the target time period, it will obtain feedback data that changes over time in the time period from the second interactive environment. These data reflect the impact of the current switch parameter combination on the target transaction and the changes in the resource asset status of each resource provider. Subsequently, according to the reward function and reward value calculation method defined by the second reward mechanism parameters, the target agent calculates the reward data under the target time period and the current switch parameter combination. The target agent updates its strategy and behavior based on the reward data and reinforcement learning algorithm to adapt to the characteristics of the target time period, and then iteratively finds the target switch parameter combination corresponding to the target time period. The above process will be repeated in each time period until the target switch parameter combination corresponding to each time period is determined. The target agent will output the target switch parameter combination corresponding to each time period as the basis for configuring the resource provision ratio of each resource provider in different time periods.
[0193] In this embodiment, a function curve is first determined, a target time point is determined based on the function curve, and the target time range is reasonably divided into multiple time periods based on the target time point, so that the target intelligent agent can determine a better target switch parameter combination according to the characteristics of different time periods, so that the output of the target intelligent agent can effectively adapt to the dynamic needs of the resource provider.
[0194] To facilitate understanding of this embodiment Figures 2 to 9 For the interaction between the agent and the interactive environment in the illustrated embodiment, see Fig.10 , Fig.10 This is an example schematic diagram of an intelligent agent interacting with an interactive environment provided in an embodiment of this specification.
[0195] The switch effect information is determined based on multiple switch parameter combinations and the impact effect data of each switch parameter combination on the target transaction within a specific time range; the resource asset information of each resource provider within a specific time range is determined based on the resource asset status prediction data and resource asset status observation data of each resource provider within a specific time range; and the interaction environment is determined based on the switch effect information and the resource asset information of each resource provider within a specific time range.
[0196] In addition, the switch effect information provides the intelligent agent with a set of switch parameter combinations, from which the intelligent agent can find a switch parameter combination that meets the needs; the resource asset information of each resource provider within a specific time range combined with the transaction constraint data provides the intelligent agent with a solution space to limit the action of the intelligent agent in finding the switch parameter combination.
[0197] The agent is placed in an interactive environment, which has been fully simulated and constructed based on the switch effect information and the resource asset information of each resource provider in a specific time range. The agent can perceive various state information in the interactive environment and begin to explore the interactive environment. During the exploration process, the agent will execute the current action, select a switch parameter combination from the switch parameter combination set, and cause the state of the interactive environment to change, so that the interactive environment generates corresponding feedback data. The feedback data reflects the effect of the current switch parameter combination on the target transaction and the changes in the resource asset status of each resource provider in a specific time range. The agent calculates the reward value obtained by the current action based on these feedback data and the preset reward mechanism parameters (such as the definition of the reward function and the calculation method of the reward value). Then, the agent updates its strategy based on the reward value and the reinforcement learning algorithm. If the current action obtains a high reward value, it means that the action is effective, and the agent may tend to repeat the action or choose a similar action in future exploration. On the contrary, if the current action obtains a low reward value, the agent may adjust its strategy and try other different switch parameter combinations. Through continuous exploration, feedback, and strategy updates, the agent gradually learns how to find the best switch parameter combination in the interactive environment to maximize the reward value it obtains. This process not only enables the agent to adapt to the dynamic changes of each resource provider within a specific time frame, but also meets the personalized needs and goals of each resource provider for resource management.
[0198] Based on this, the intelligent agent can quickly and accurately find the optimal switch parameter combination according to the switch effect information and the resource asset information of each resource provider in a specific time range, so as to optimize the resource allocation and management of each resource provider.
[0199] It should be noted that the intelligent agent of this embodiment can be the above-mentioned initial intelligent agent or target intelligent agent, and various types of information involved in this embodiment can exist as sample data during the training process.
[0200] For the effects that can be achieved by this embodiment, please refer to the relevant embodiments of the above-mentioned intelligent agent training method and data configuration method, which will not be repeated here.
[0201] based on Figure 1 The following will combine Fig.11 , the intelligent agent training device provided in the embodiment of this specification is introduced in detail. It should be noted that, Fig.11 The intelligent agent training device 1 in the present invention is used to execute the present invention Figure 2 - Figure 5 For the convenience of explanation, only the part related to the embodiment of this specification is shown. For the specific technical details not disclosed, please refer to this specification. Figure 2 - Figure 5 The embodiment shown. Wherein, the intelligent agent training device 1 specifically comprises:
[0202] A first acquisition unit 11 is used to acquire a switch effect information sample, the switch effect information sample including a plurality of switch parameter combination samples and effect data of each of the plurality of switch parameter combination samples on a target transaction within a historical time range, the switch parameter combination sample indicating a resource provision ratio of each of the plurality of resource providers to convert resources into assets based on the target transaction;
[0203] The second acquisition unit 12 is used to acquire resource asset information samples of each resource provider within a historical time range;
[0204] A construction unit 13, configured to construct a first interactive environment based on the switch effect information sample and the resource asset information samples of each resource provider;
[0205] The training unit 14 is used to perform reinforcement learning training on the initial intelligent agent based on the first interactive environment to obtain a target intelligent agent.
[0206] Optionally, the training unit 14 is also used to: determine a first solution space for the initial intelligent agent; determine a first action configuration parameter of the initial intelligent agent, the first action configuration parameter indicates an action of the initial intelligent agent, the action is to select a switch parameter combination sample from multiple switch parameter combination samples as the current switch parameter combination sample; determine a first reward mechanism parameter for the initial intelligent agent, the first reward mechanism parameter indicates that after the initial intelligent agent performs the action, first feedback data that changes with time within a historical time range and is returned by the first interactive environment is obtained, and reward data for the initial intelligent agent under the current switch parameter combination sample is determined based on the first feedback data; based on the first solution space, the first action configuration parameter and the first reward mechanism parameter, the initial intelligent agent is iteratively trained in the first interactive environment; if the preset convergence conditions are met, the initial intelligent agent is determined as the target intelligent agent.
[0207] Optionally, the training unit 14 is also used to: obtain the first transaction constraint data of each resource provider in the historical time range; determine the upper and lower asset limits of each resource provider in the historical time range based on the switch effect information sample and the resource asset information sample and the first transaction constraint data of each resource provider; and obtain the first solution space of the initial intelligent agent based on the analysis of the upper and lower asset limits of each resource provider in the historical time range.
[0208] Optionally, the first acquisition unit 11 is also used to: determine multiple first switch parameter combination samples and multiple second switch parameter combination samples; obtain impact effect data of each first switch parameter combination sample in the multiple first switch parameter combination samples on the target transaction within a historical time range; call a preset linear regression model based on each first switch parameter combination sample, the impact effect data of each first switch parameter combination sample, and multiple second switch parameter combination samples to obtain the impact effect data of each second switch parameter combination sample in the multiple second switch parameter combination samples output by the linear regression model on the target transaction within the historical time range; construct a switch effect information sample based on each first switch parameter combination sample, the impact effect data of each first switch parameter combination sample, each second switch parameter combination sample, and the impact effect data of each second switch parameter combination sample.
[0209] Optionally, the second acquisition unit 12 is also used to: obtain the resource asset status prediction data and resource asset status observation data of each resource provider in the historical time range; and determine the resource asset information sample of each resource provider in the historical time range based on the resource asset status prediction data and resource asset status observation data of each resource provider in the historical time range.
[0210] For the effects that can be achieved by this embodiment, please refer to the relevant embodiments of the above-mentioned intelligent agent training method, which will not be repeated here.
[0211] based on Figure 1 The following will combine Fig.12 , the data configuration device provided in the embodiment of this specification is introduced in detail. It should be noted that, Fig.12 The data configuration device 2 is used to execute this instruction Figure 6 - Fig. 9 For the convenience of explanation, only the part related to the embodiment of this specification is shown. For the specific technical details not disclosed, please refer to this specification. Figure 6 - Fig. 9 The embodiment shown. Wherein, the data configuration device 2 specifically includes:
[0212] A first acquisition unit 21 is used to acquire switch effect information, where the switch effect information includes multiple switch parameter combinations and effect data of each of the multiple switch parameter combinations on the target transaction within a target time range, where the switch parameter combination indicates a resource provision ratio of each of the multiple resource providers to convert resources into assets based on the target transaction;
[0213] The second acquisition unit 22 is used to acquire resource asset information of each resource provider within a target time range;
[0214] A construction unit 23, configured to construct a second interactive environment based on the switch effect information and the resource asset information of each resource provider;
[0215] A calling unit 24 is used to call a target agent based on the second interactive environment, so that the target agent determines a target switch parameter combination from a plurality of switch parameter combinations, wherein the target agent is trained based on the relevant embodiment of the above-mentioned agent training method;
[0216] The configuration unit 25 is used to configure the target resource provision ratio of each resource provider based on the target switch parameter combination.
[0217] Optionally, the calling unit 24 is also used to: determine a second solution space for the target intelligent agent; determine a second action configuration parameter for the target intelligent agent, the second action configuration parameter indicates an action of the target intelligent agent, the action being to select a switch parameter combination from multiple switch parameter combinations as the current switch parameter combination; determine a second reward mechanism parameter for the target intelligent agent, the second reward mechanism parameter indicates that after the target intelligent agent performs the action, second feedback data that changes with time within a target time range and is returned by the second interactive environment is obtained, and reward data for the target intelligent agent under the current switch parameter combination is determined based on the second feedback data; based on the second solution space, the second action configuration parameter and the second reward mechanism parameter, the target intelligent agent is called in the second interactive environment to enable the target intelligent agent to determine the target switch parameter combination from multiple switch parameter combinations.
[0218] Optionally, the calling unit 24 is also used to: obtain the second transaction constraint data of each resource provider within the target time range; determine the upper and lower asset limits of each resource provider within the target time range based on the switch effect information and the resource asset information and the second transaction constraint data of each resource provider; and obtain the second solution space of the target intelligent entity based on the analysis of the upper and lower asset limits of each resource provider within the target time range.
[0219] Optionally, the calling unit 24 is also used to: determine a function curve of the asset upper limit and asset lower limit of the target resource provider changing with time within the target time range based on the switch effect information, the resource asset information of each resource provider, and the second transaction constraint data; determine at least one target time point of the target resource provider within the target time range based on the function curve; divide the target time range based on at least one target time point to obtain multiple time periods; call the target intelligent agent in the second interactive environment based on the multiple time periods, the second solution space, the second action configuration parameters, and the second reward mechanism parameters, so that the target intelligent agent determines the target switch parameter combination corresponding to each time period in the multiple time periods from the multiple switch parameter combinations.
[0220] For the effects that can be achieved by this embodiment, please refer to the relevant embodiments of the above-mentioned intelligent agent training method and data configuration method, which will not be repeated here.
[0221] See also Fig.13 , is a schematic diagram of the structure of an electronic device provided in the embodiment of this specification. Fig.13 As shown, the electronic device 1000 may include: at least one processor 1001, such as a CPU, at least one network interface 1004, an input / output interface 1003, a memory 1005, and at least one communication bus 1002. The communication bus 1002 is used to realize the connection and communication between these components. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a non-volatile memory (non-volatile memory), such as at least one disk storage. The memory 1005 may optionally be at least one storage device located away from the aforementioned processor 1001. As shown in FIG. Fig.13 As shown, the memory 1005 as a computer storage medium may include an operating system, a network communication module, an input and output interface module, an agent training application, and a data configuration application.
[0222] exist Fig.13 In the electronic device 1000 shown, the input-output interface 1003 is mainly used to provide an input interface for the user and obtain data input by the user.
[0223] In one embodiment, the processor 1001 may be used to call the agent training application stored in the memory 1005, and specifically perform the following operations:
[0224] Obtaining a switch effect information sample, the switch effect information sample including a plurality of switch parameter combination samples and effect data of each of the plurality of switch parameter combination samples on a target transaction within a historical time range, the switch parameter combination sample indicating a resource provision ratio of each of the plurality of resource providers to convert resources into assets based on the target transaction;
[0225] Obtain resource asset information samples of each resource provider within a historical time range;
[0226] Building a first interactive environment based on the switch effect information sample and the resource asset information sample of each resource provider;
[0227] Based on the first interactive environment, reinforcement learning training is performed on the initial intelligent agent to obtain a target intelligent agent.
[0228] Optionally, when the processor 1001 performs reinforcement learning training on the initial intelligent agent based on the first interactive environment to obtain the target intelligent agent, the following specific operations are performed: determining the first solution space of the initial intelligent agent; determining the first action configuration parameter of the initial intelligent agent, the first action configuration parameter indicates the action of the initial intelligent agent, and the action is to select a switch parameter combination sample from multiple switch parameter combination samples as the current switch parameter combination sample; determining the first reward mechanism parameter of the initial intelligent agent, the first reward mechanism parameter indicates that after the initial intelligent agent performs the action, the first feedback data that changes with time within the historical time range returned by the first interactive environment is obtained, and the reward data of the initial intelligent agent under the current switch parameter combination sample is determined according to the first feedback data; based on the first solution space, the first action configuration parameter and the first reward mechanism parameter, the initial intelligent agent is iteratively trained in the first interactive environment; if the preset convergence condition is met, the initial intelligent agent is determined as the target intelligent agent.
[0229] Optionally, when executing to determine the first solution space of the initial intelligent entity, the processor 1001 specifically performs the following operations: obtaining the first transaction constraint data of each resource provider in the historical time range; determining the upper and lower asset limits of each resource provider in the historical time range based on the switch effect information sample and the resource asset information sample and the first transaction constraint data of each resource provider; and obtaining the first solution space of the initial intelligent entity based on the analysis of the upper and lower asset limits of each resource provider in the historical time range.
[0230] Optionally, when executing the acquisition of switch effect information samples, the processor 1001 specifically performs the following operations: determining multiple first switch parameter combination samples and multiple second switch parameter combination samples; obtaining the impact effect data of each first switch parameter combination sample in the multiple first switch parameter combination samples on the target transaction within the historical time range; calling a preset linear regression model based on each first switch parameter combination sample, the impact effect data of each first switch parameter combination sample, and multiple second switch parameter combination samples to obtain the impact effect data of each second switch parameter combination sample in the multiple second switch parameter combination samples output by the linear regression model on the target transaction within the historical time range; constructing a switch effect information sample based on each first switch parameter combination sample, the impact effect data of each first switch parameter combination sample, each second switch parameter combination sample, and the impact effect data of each second switch parameter combination sample.
[0231] Optionally, when executing the process of obtaining resource asset information samples of each resource provider in a historical time range, the processor 1001 specifically performs the following operations: obtaining resource asset status prediction data and resource asset status observation data of each resource provider in a historical time range; and determining resource asset information samples of each resource provider in the historical time range based on the resource asset status prediction data and resource asset status observation data of each resource provider in the historical time range.
[0232] In one embodiment, the processor 1001 may be used to call the data configuration application stored in the memory 1005, and specifically perform the following operations:
[0233] Obtaining switch effect information, the switch effect information including multiple switch parameter combinations and effect data of each of the multiple switch parameter combinations on the target transaction within a target time range, the switch parameter combination indicating a resource provision ratio of each of the multiple resource providers to convert resources into assets based on the target transaction;
[0234] Obtain resource asset information of each resource provider within the target time range;
[0235] Building a second interactive environment based on the switch effect information and the resource asset information of each resource provider;
[0236] Invoking a target agent based on the second interactive environment so that the target agent determines a target switch parameter combination from a plurality of switch parameter combinations, wherein the target agent is trained based on an agent training program;
[0237] The target resource provision ratio of each resource provider is configured based on the target switch parameter combination.
[0238] Optionally, when executing the calling of the target intelligent agent based on the second interactive environment so that the target intelligent agent determines a target switch parameter combination from multiple switch parameter combinations, the processor 1001 specifically performs the following operations: determining a second solution space for the target intelligent agent; determining a second action configuration parameter for the target intelligent agent, the second action configuration parameter indicating an action of the target intelligent agent, the action being selecting a switch parameter combination from multiple switch parameter combinations as the current switch parameter combination; determining a second reward mechanism parameter for the target intelligent agent, the second reward mechanism parameter indicating obtaining second feedback data that changes with time within a target time range returned by the second interactive environment after the target intelligent agent performs the action, and determining reward data for the target intelligent agent under the current switch parameter combination based on the second feedback data; based on the second solution space, the second action configuration parameter and the second reward mechanism parameter, calling the target intelligent agent in the second interactive environment so that the target intelligent agent determines the target switch parameter combination from multiple switch parameter combinations.
[0239] Optionally, when executing to determine the second solution space of the target intelligent entity, the processor 1001 specifically performs the following operations: obtaining the second transaction constraint data of each resource provider in the target time range; determining the upper and lower asset limits of each resource provider in the target time range based on the switch effect information and the resource asset information and the second transaction constraint data of each resource provider; and obtaining the second solution space of the target intelligent entity based on the analysis of the upper and lower asset limits of each resource provider in the target time range.
[0240] Optionally, when the processor 1001 executes the calling of the target intelligent agent in the second interactive environment based on the second solution space, the second action configuration parameters and the second reward mechanism parameters so that the target intelligent agent determines the target switch parameter combination from multiple switch parameter combinations, the following specific operations are performed: based on the switch effect information and the resource asset information of each resource provider and the second transaction constraint data, a function curve of the asset upper limit and asset lower limit of the target resource provider changing with time within the target time range is determined; based on the function curve, at least one target time point of the target resource provider within the target time range is determined; based on the at least one target time point, the target time range is divided into multiple time periods; based on the multiple time periods, the second solution space, the second action configuration parameters and the second reward mechanism parameters, the target intelligent agent is called in the second interactive environment so that the target intelligent agent determines the target switch parameter combination corresponding to each of the multiple time periods from the multiple switch parameter combinations.
[0241] For the effects that can be achieved by this embodiment, please refer to the relevant embodiments of the above-mentioned intelligent agent training method and data configuration method, which will not be repeated here.
[0242] The embodiment of the present specification also provides a computer storage medium, which computer readable storage medium stores computer program code. When the computer program code is executed, the above Figure 2 - Fig.10 The method of the embodiment shown in the figure can be found in the specific implementation process. Figure 2 - Fig.10 The specific description of the illustrated embodiment will not be repeated here.
[0243] The embodiment of the present specification also provides a computer program product, which stores at least one instruction, and when the at least one instruction is executed by a processor, the above Figure 2 - Fig.10 The method of the embodiment shown in the figure can be found in the specific implementation process. Figure 2 - Fig.10 The specific description of the illustrated embodiment will not be repeated here.
[0244] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.
[0245] The above disclosure is only the preferred embodiment of this specification, which certainly cannot be used to limit the scope of rights of this specification. Therefore, equivalent changes made according to the claims of this specification are still within the scope covered by this specification.
Claims
1. A method for training an intelligent agent, the method comprising: Acquire a switch effect information sample, the switch effect information sample comprising a plurality of switch parameter combination samples and effect data of each of the plurality of switch parameter combination samples on a target transaction within a historical time range, the switch parameter combination sample indicating a resource provision ratio of each of the plurality of resource providers to convert resources into assets based on the target transaction; Obtaining a sample of resource asset information of each of the resource providers within the historical time range; Building a first interactive environment based on the switch effect information sample and the resource asset information samples of each resource provider; Based on the first interactive environment, reinforcement learning training is performed on the initial intelligent agent to obtain a target intelligent agent.
2. According to the method of claim 1, the step of performing reinforcement learning training on the initial agent based on the first interactive environment to obtain the target agent comprises: Determining a first solution space of the initial agent; Determine a first action configuration parameter of the initial agent, where the first action configuration parameter indicates an action of the initial agent, where the action is to select a switch parameter combination sample from a plurality of switch parameter combination samples as a current switch parameter combination sample; Determine a first reward mechanism parameter of the initial agent, wherein the first reward mechanism parameter indicates that after the initial agent performs the action, first feedback data returned by the first interactive environment that changes over time within the historical time range is obtained, and reward data of the initial agent under the current switch parameter combination sample is determined according to the first feedback data; Iteratively training the initial agent in the first interactive environment based on the first solution space, the first action configuration parameters, and the first reward mechanism parameters; If the preset convergence condition is met, the initial agent is determined as the target agent.
3. The method according to claim 2, wherein determining the first solution space of the initial agent comprises: Acquire the first transaction constraint data of each resource provider within the historical time range; Determine the upper and lower limits of assets of each resource provider within the historical time range based on the switch effect information sample, the resource asset information sample of each resource provider, and the first transaction constraint data; The first solution space of the initial intelligent agent is obtained based on the analysis of the upper and lower limits of assets of each resource provider in the historical time range.
4. The method according to claim 1, wherein obtaining a switching effect information sample comprises: Determine a plurality of first switch parameter combination samples and a plurality of second switch parameter combination samples; Acquire the effect data of each of the first switch parameter combination samples among the plurality of the first switch parameter combination samples on the target transaction within a historical time range; Based on each of the first switch parameter combination samples, the impact effect data of each of the first switch parameter combination samples, and a plurality of the second switch parameter combination samples, a preset linear regression model is called to obtain the impact effect data of each of the second switch parameter combination samples on the target transaction within the historical time range among the plurality of second switch parameter combination samples output by the linear regression model; A switch effect information sample is constructed based on each of the first switch parameter combination samples, the influence effect data of each of the first switch parameter combination samples, each of the second switch parameter combination samples, and the influence effect data of each of the second switch parameter combination samples.
5. According to the method of claim 1, the step of obtaining a sample of resource asset information of each resource provider within the historical time range comprises: Obtain resource asset status prediction data and resource asset status observation data of each resource provider within the historical time range; The resource asset information samples of each resource provider within the historical time range are determined based on the resource asset status prediction data and resource asset status observation data of each resource provider within the historical time range.
6. A data configuration method, the method comprising: Acquire switch effect information, the switch effect information including a plurality of switch parameter combinations and effect data of each of the plurality of switch parameter combinations on a target transaction within a target time range, the switch parameter combination indicating a resource provision ratio of each of the plurality of resource providers to convert resources into assets based on the target transaction; Acquire resource asset information of each resource provider within the target time range; Building a second interactive environment based on the switch effect information and the resource asset information of each resource provider; Invoking a target agent based on the second interactive environment, so that the target agent determines a target switch parameter combination from the plurality of switch parameter combinations, wherein the target agent is trained based on the method according to any one of claims 1 to 6; The target resource provision ratio of each of the resource providers is configured based on the target switch parameter combination.
7. The method according to claim 6, wherein the calling of the target agent based on the second interactive environment so that the target agent determines a target switch parameter combination from the plurality of switch parameter combinations comprises: Determining a second solution space of the target agent; Determine a second action configuration parameter of the target agent, where the second action configuration parameter indicates an action of the target agent, where the action is to select a switch parameter combination from a plurality of switch parameter combinations as a current switch parameter combination; Determine a second reward mechanism parameter of the target agent, wherein the second reward mechanism parameter indicates that after the target agent performs the action, second feedback data returned by the second interactive environment that changes over time within the target time range is obtained, and reward data of the target agent under the current switch parameter combination is determined according to the second feedback data; Based on the second solution space, the second action configuration parameters and the second reward mechanism parameters, the target agent is called in the second interactive environment so that the target agent determines a target switch parameter combination from the multiple switch parameter combinations.
8. The method according to claim 7, wherein determining the second solution space of the target agent comprises: Acquire second transaction constraint data of each resource provider within the target time range; Determine the upper and lower limits of assets of each resource provider within the target time range based on the switch effect information, resource asset information of each resource provider, and second transaction constraint data; The second solution space of the target intelligent entity is obtained based on the upper and lower asset limits of each resource provider within the target time range.
9. The method according to claim 8, wherein the step of calling the target agent in the second interactive environment based on the second solution space, the second action configuration parameter, and the second reward mechanism parameter, so that the target agent determines a target switch parameter combination from a plurality of switch parameter combinations, comprises: Based on the switch effect information, the resource asset information of each resource provider, and the second transaction constraint data, determine a function curve of the asset upper limit and the asset lower limit of the target resource provider changing with time within the target time range; Based on the function curve, determining at least one target time point of the target resource provider within the target time range; Dividing the target time range based on at least one target time point to obtain multiple time periods; Based on the multiple time periods, the second solution space, the second action configuration parameters and the second reward mechanism parameters, the target intelligent agent is called in the second interactive environment so that the target intelligent agent determines the target switch parameter combination corresponding to each of the multiple time periods from the multiple switch parameter combinations.
10. An intelligent agent training device, the device comprising: A first acquisition unit is used to acquire a switch effect information sample, wherein the switch effect information sample includes a plurality of switch parameter combination samples and effect data of each of the plurality of switch parameter combination samples on a target transaction within a historical time range, wherein the switch parameter combination sample indicates a resource provision ratio of each of the plurality of resource providers to convert resources into assets based on the target transaction; A second acquisition unit is used to acquire resource asset information samples of each resource provider within the historical time range; A construction unit, configured to construct a first interactive environment based on the switch effect information sample and the resource asset information samples of each resource provider; A training unit is used to perform reinforcement learning training on the initial intelligent agent based on the first interactive environment to obtain a target intelligent agent.
11. A data configuration device, comprising: A first acquisition unit is used to acquire switch effect information, wherein the switch effect information includes a plurality of switch parameter combinations and effect data of each of the plurality of switch parameter combinations on a target transaction within a target time range, wherein the switch parameter combination indicates a resource provision ratio of each of the plurality of resource providers to convert resources into assets based on the target transaction; A second acquisition unit is used to acquire resource asset information of each resource provider within the target time range; A construction unit, configured to construct a second interactive environment based on the switch effect information and the resource asset information of each resource provider; a calling unit, configured to call a target agent based on the second interactive environment, so that the target agent determines a target switch parameter combination from the plurality of switch parameter combinations, wherein the target agent is trained based on the method according to any one of claims 1 to 6; A configuration unit is used to configure a target resource provision ratio of each resource provider based on the target switch parameter combination.
12. A computer-readable storage medium storing a computer program code, wherein when the computer program code is executed, the method according to any one of claims 1 to 9 is implemented.
13. An electronic device, comprising: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the steps of the method as claimed in any one of claims 1 to 9.
14. A computer program product, wherein the computer program product stores at least one instruction, and when the at least one instruction is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.