Illegal capital collection suspicious subject identification method and system based on reinforcement learning
Through the method based on reinforcement learning, the characteristics of illegal fundraising funds data are analyzed and reinforcement learning models are constructed, which solves the limitations of dynamic changes and complex feature extraction in the existing technology, and achieves more efficient and accurate identification of illegal fundraising risks.
Patent Information
- Application Number
- CN202510030512.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing methods for identifying illegal fundraising risks have limitations in dealing with dynamic changes in fund data and complex feature extraction, and cannot fully mine the implicit information in the data, and lack accuracy and flexibility.
Using a method based on reinforcement learning, we analyze the transaction flows of historical suspected illegal fundraising entities, extract the characteristics of fund data, and build a reinforcement learning model, and dynamically adjust the recognition strategy to gradually improve the accuracy and robustness of the recognition.
Effectively identifying suspicious subjects of illegal fundraising improves the accuracy and efficiency of risk identification, can dynamically adapt to changes in illegal fundraising behavior, and identify abnormal patterns and risk points that are difficult to detect by traditional methods.
Smart Images

Figure CN120013229A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of identification of suspicious entities of illegal fund-raising, and in particular to a method and system for identifying suspicious entities of illegal fund-raising based on reinforcement learning. Background Art
[0002] Illegal fund-raising refers to the act of raising funds from unspecified persons by promising to repay principal and interest or provide other investment returns without the permission of the financial management department of the State Council or in violation of the national financial management regulations. With the rapid development of Internet finance in recent years, illegal fund-raising has become more hidden and complex, posing a huge challenge to regulators. First, the concealment and complexity of illegal fund-raising increase the difficulty of risk identification. Illegal fund-raising often disguises itself as legal investment and financial products, using high returns to attract investors, and has a variety of operating methods and involves complex capital flows, which are difficult to identify through simple rules and feature extraction. Secondly, the dynamic and changeable nature of illegal fund-raising makes it difficult for traditional static identification methods to cope with it. Illegal fund-raisers often change their methods and strategies to evade supervision and crackdowns, which requires the identification method to have the ability to dynamically learn and adapt to changes. Finally, the existing illegal fund-raising risk identification methods have limitations in data utilization and analysis, and it is difficult to fully explore the implicit information in the capital flow data, resulting in unsatisfactory identification results.
[0003] Traditional methods for identifying illegal fundraising risks mainly rely on rule models and supervised learning, but these two methods have obvious shortcomings when dealing with complex and changeable illegal fundraising behaviors. Rule models rely on expert experience to formulate rules, and there are problems such as incomplete rule coverage and poor flexibility, which makes it difficult to deal with complex and changeable illegal fundraising behaviors. Supervised learning methods require a large amount of labeled data for training, but the labeling cost of illegal fundraising data is high and time-consuming, and the labeled data is often incomplete, which affects the training effect and generalization ability of the model. In contrast, reinforcement learning does not require a large amount of labeled data, and can gradually improve the accuracy and robustness of recognition through interactive learning strategies with the environment. By setting a reasonable reward mechanism, the reinforcement learning method can effectively capture abnormal patterns and changing trends in the data, dynamically adapt to changes in illegal fundraising behaviors, and improve the recognition effect.
[0004] Fund data is a direct reflection of illegal fund-raising activities. Analyzing fund flows can reveal illegal fund-raising activities hidden behind legitimate businesses. By extracting and analyzing key features in fund flow data, such as the frequency of fund inflows and outflows, abnormal fluctuations in transaction amounts, and transaction behaviors of related accounts, suspicious entities of illegal fund-raising can be effectively identified.
[0005] Therefore, the technical problems that need to be solved urgently are: the existing methods have limitations in dealing with dynamic changes in financial data and extracting complex features, and are unable to fully explore the implicit information in the data. In addition, the existing methods for identifying risks of illegal fundraising lack accuracy and flexibility. Summary of the invention
[0006] The present invention is made to solve the above-mentioned problems, and aims to provide a method and system for identifying suspicious entities of illegal fund-raising based on reinforcement learning.
[0007] The present invention provides a method for identifying suspicious entities of illegal fund-raising based on reinforcement learning, which has the following characteristics and specifically includes the following steps: S1, analyzing the data characteristics of illegal fund-raising funds according to the transaction flow of the entities suspected of illegal fund-raising in history, and sorting out the types of fund risks and key characteristics; S2, sorting out the key data items of each type of risk based on the risk type and risk characteristics; S3, collecting the transaction flow data of the entities suspected of / not suspected of illegal fund-raising in history, pre-processing the original data by means of data cleaning, missing value processing, data standardization, data encoding, data statistics, etc., and extracting key data characteristics; S4, forming a time series illegal fund-raising data feature table corresponding to each risk entity with a time interval of one month; S5, constructing a reinforcement learning model, wherein the reinforcement learning mainly consists of an intelligent agent, an environment, a state, an action, and a reward, and the intelligent agent observes the state s at the current time t from the environment t Then an action a is selected according to the strategy π t , the agent executed a t After that, the environment will transition to a new state s t+1 , for the new state s t+1 The environment will give a reward signal r t (positive reward or negative reward), then the agent performs new actions according to a certain strategy based on the new state and the reward of environmental feedback. The goal of reinforcement learning is to find a strategy that maximizes the long-term discounted cumulative reward in a given environment, which can be used in state s t Take action t The cumulative expected discounted reward is calculated by the state-action value function Q π (s,a) means,
[0008] S6. Based on the reinforcement learning model, try to select a most suitable reinforcement learning algorithm for model training, including DQN, A3C, DDPG, PPO, TD3 and other algorithms. The algorithm selection principles include continuity / discreteness of state and action space, iteration method and learning efficiency; S7. Train the reinforcement learning model; S8. Comprehensively compare the performance of each algorithm based on indicators such as convergence speed, recognition accuracy, and training efficiency, and select the best algorithm; S9. After determining the reinforcement learning algorithm with the best overall performance, carry out reinforcement learning optimization steps such as environmental testing, hyperparameter testing, and network structure optimization to gradually improve the performance of the algorithm, and train the reinforcement learning model based on the optimal parameter settings; S10. Integrate the trained reinforcement learning model into the illegal fund-raising risk identification system, connect it with the existing financial supervision system and data analysis platform, and realize automated and intelligent risk identification; S11. Based on the identification results of the reinforcement learning agent, establish an illegal fund-raising early warning mechanism. When a suspicious subject is identified, promptly notify the regulatory agency for further investigation and disposal.
[0009] In the method for identifying suspicious entities of illegal fund-raising based on reinforcement learning provided by the present invention, it can also have the following characteristics: wherein, step S1 also includes the following steps: S1-1, the risks of illegal fund-raising funds include 8 risk types: daily transaction scale, dispersed transfer-in, concentrated transfer-out, abnormal time transactions, 24-hour uninterrupted transactions, risky transaction objects, abnormal transaction patterns and fuzzy transactions; S1-2, the daily transaction scale includes 2 characteristics: the number of daily transactions and the daily transaction amount; dispersed transfer-in includes 2 characteristics: a single transfer amount of ten thousand and the number of monthly transaction counterparties; concentrated transfer-out includes 1 characteristic: the number of daily transaction counterparties; abnormal time transactions include 1 characteristic: the number of transactions at abnormal times; 24-hour uninterrupted transactions include 1 characteristic: an account has transactions every hour for 24 consecutive hours; risky transaction objects include 1 characteristic: the number of transaction objects that are risky enterprises / individuals; abnormal transaction patterns include 3 characteristics: the number of repeated small transactions, the number of repeated transactions with a specific amount, and third-party account transfers; fuzzy transactions include 1 characteristic: the number of transactions in which characteristic words such as "investment return", "fundraising funds", and "loan repayment" appear in the monthly transaction remarks.
[0010] In the method for identifying suspicious entities of illegal fund-raising based on reinforcement learning provided by the present invention, it can also have the following characteristics: wherein, the key data items of daily transaction scale are transaction account, transaction ID, transaction date, and transaction amount; the key data items of decentralized transfer-in are transaction account, transaction date, transaction amount, and counterparty account; the key data items of centralized transfer-out are transaction account, transaction date, and counterparty account; the key data items of abnormal time transactions are transaction account, transaction date, transaction time, and transaction ID; the key data items of 24-hour uninterrupted transactions are transaction account, transaction date, transaction time, and transaction ID; the key data items of risky transaction objects are transaction account, counterparty account, and list of risky enterprises / individuals; the key data items of abnormal transaction patterns are transaction account, transaction amount, transaction time, and counterparty account; the key data items of fuzzy transactions are transaction account, transaction remark, and transaction ID.
[0011] The method for identifying suspicious entities of illegal fund-raising based on reinforcement learning provided by the present invention may also have the following features: wherein step S5 also includes the following steps: S5-1, discretizing the fund data corresponding to each entity into time series data with one month as the time interval for each moment; S5-2, constructing a reinforcement learning environment, and using the fund data and the extracted features as the state input of the reinforcement learning agent; S5-3, the state s t Composed of extracted fund flow characteristics; S5-4, action a t is the action that the agent can perform in each state, that is, marking a subject as a suspected illegal fundraising subject or a legal subject, with a value of 0 or 1; S5-5, reward r t It is defined as the sum of positive rewards and negative rewards. Positive rewards are given for correctly identifying the illegal fund-raising entity, while negative rewards are given for failure to identify or misjudgement. The positive reward is positively correlated with the amount involved in the case corresponding to the entity. In addition to being related to the amount involved, negative rewards also need to take into account the losses caused by misjudgment, regulatory penalties, and the human costs of correcting the error.
[0012] In the method for identifying suspicious entities of illegal fund-raising based on reinforcement learning provided by the present invention, it can also have the following characteristics: wherein, step S7 also includes the following steps: S7-1, initializing the parameters of the reinforcement learning agent, including the weights, learning rate, discount factor, etc. of the neural network; S7-2, the agent continuously interacts with the data in the environment, obtains feedback (rewards) by performing actions (marking entities), and adjusts strategies; S7-3, stores the agent's historical experience in a memory bank, and performs random sampling training to reduce the correlation between samples and stabilize the training process; S7-4, based on the rewards obtained and state transfers, uses optimization methods such as gradient descent to update the agent's strategy parameters to improve recognition accuracy.
[0013] The present invention also provides a system for identifying suspected illegal fund-raising entities based on reinforcement learning, which has the following characteristics: an information combing module, which analyzes the data characteristics of illegal fund-raising funds according to the transaction flow of the entities suspected of illegal fund-raising in history, and combs the fund risk types and key characteristics; a data analysis module, which combs the key data items of each type of risk based on the risk type and risk characteristics; a data feature extraction module, which collects the transaction flow data of the entities suspected of / not suspected of illegal fund-raising in history, pre-processes the original data by means of data cleaning, missing value processing, data standardization, data encoding, data statistics, etc., and extracts key data features; a data feature processing module, which forms a time series illegal fund-raising data feature table corresponding to each risk entity with a time interval of one month; a model construction module, which constructs a reinforcement learning model, wherein the reinforcement learning mainly consists of an intelligent agent, an environment, a state, an action, and a reward, and the intelligent agent observes the state s at the current time t from the environment t Then an action a is selected according to the strategy π t , the agent executed a t After that, the environment will transition to a new state s t+1 , for the new state s t+1 The environment will give a reward signal r t (positive reward or negative reward), then the agent performs new actions according to a certain strategy based on the new state and the reward of environmental feedback. The goal of reinforcement learning is to find a strategy that maximizes the long-term discounted cumulative reward in a given environment, which can be used in state s t Take action t The cumulative expected discounted reward is calculated by the state-action value function Q π (s,a) means,
[0014] The model selection module attempts to select a most suitable reinforcement learning algorithm for model training based on the reinforcement learning model, including DQN, A3C, DDPG, PPO, TD3 and other algorithms. The algorithm selection principles include the continuity / discreteness of the state and action space, the iteration method and the learning efficiency; the model training module trains the reinforcement learning model; the algorithm selection module compares the performance of each algorithm based on indicators such as convergence speed, recognition accuracy and training efficiency, and selects the best algorithm; the algorithm optimization module determines the reinforcement learning algorithm with the best overall performance, and then carries out reinforcement learning optimization steps such as environmental testing, hyperparameter testing, and network structure optimization to gradually improve the performance of the algorithm, and trains the reinforcement learning model based on the optimal parameter settings; the risk identification module integrates the trained reinforcement learning model into the illegal fund-raising risk identification system, connects with the existing financial regulatory system and data analysis platform, and realizes automated and intelligent risk identification; the early warning module establishes an illegal fund-raising early warning mechanism based on the identification results of the reinforcement learning agent, and when a suspicious subject is identified, promptly notifies the regulatory agency for further investigation and disposal.
[0015] Functions and Effects of the Invention
[0016] According to the reinforcement learning-based method and system for identifying suspicious entities of illegal fund-raising involved in the present invention, the reinforcement learning-based method for identifying suspicious entities of illegal fund-raising proposed in the present invention utilizes key fund characteristics and data to train reinforcement learning agents, so that the agents can accurately identify suspicious entities of illegal fund-raising, thereby improving the accuracy and efficiency of risk identification.
[0017] The present invention innovatively adopts a method based on fund data and reinforcement learning, effectively solving the deficiencies in existing illegal fund-raising risk identification methods, and provides a new idea and new means for illegal fund-raising risk identification. In addition, the reinforcement learning method can fully tap the implicit information in the fund data, identify abnormal patterns and risk points that are difficult to detect with traditional rule models and supervised learning methods, and provide regulators with an efficient and intelligent illegal fund-raising risk identification tool. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a recognition schematic diagram of a method for identifying suspected entities of illegal fund-raising based on reinforcement learning in an embodiment of the present invention. DETAILED DESCRIPTION
[0019] In the description of this application, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, an electrical connection, or mutual communication; it can be a direct connection, or an indirect connection through an intermediate medium, it can be the internal connection of two elements or the interaction relationship between two elements. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0020] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the following embodiments and the accompanying drawings specifically illustrate a method for identifying suspicious entities of illegal fund-raising based on reinforcement learning of the present invention.
[0021] Figure 1 It is a recognition schematic diagram of a method for identifying suspected entities of illegal fund-raising based on reinforcement learning in an embodiment of the present invention.
[0022] like Figure 1 As shown, the method for identifying suspicious entities of illegal fund-raising based on reinforcement learning in this embodiment specifically includes the following steps:
[0023] S1. Analyze the data characteristics of illegal fund-raising funds based on the historical transaction flows of entities suspected of illegal fund-raising, and sort out the types of fund risks and key characteristics.
[0024] S1-1, the risks of illegal fund-raising include eight risk types: daily transaction scale, dispersed transfer-in, concentrated transfer-out, abnormal time transaction, 24-hour uninterrupted transaction, risky transaction objects, abnormal transaction patterns and ambiguous transactions.
[0025] S1-2, daily transaction scale includes two features: daily transaction number and daily transaction amount; dispersed transfer-in includes two features: single transfer-in amount is a whole ten thousand and the number of monthly trading counterparties; concentrated transfer-out includes one feature: the number of daily trading counterparties; abnormal time transactions include one feature: the number of transactions at abnormal times; 24-hour uninterrupted transactions include one feature: the account has transactions every hour for 24 consecutive hours; risky transaction objects include one feature: the number of risky enterprises / individuals; abnormal transaction patterns include three features: the number of repeated small transactions, the number of repeated transactions with specific amounts, and third-party account transfers; fuzzy transactions include one feature: the number of transactions in which characteristic words such as "investment return", "fundraising funds", and "loan repayment" appear in the remarks of monthly transactions.
[0026] S2, sort out the key data items for each type of risk based on risk type and risk characteristics.
[0027] Table 1
[0028]
[0029]
[0030] Table 1 is a table of risk types, characteristics and key data items of illegal fund-raising funds.
[0031] S3, collects historical transaction flow data of entities suspected / not suspected of illegal fund-raising, pre-processes the original data through data cleaning, missing value processing, data standardization, data coding, data statistics and other means, and extracts key data features.
[0032] S4, with one month as a time interval, forms a time series illegal fund-raising data feature table corresponding to each risk subject.
[0033] Table 2
[0034]
[0035]
[0036] Table 2 is a table of illegal fund-raising data characteristics, where “yes” is represented by 1 and “no” is represented by 0.
[0037] S5, build a reinforcement learning model.
[0038] Reinforcement learning mainly consists of an agent, an environment, a state, an action, and a reward. The agent observes the state s at the current time t from the environment. t Then an action a is selected according to the strategy π t , the agent executed a t After that, the environment will transition to a new state s t+1 , for the new state s t+1 The environment will give a reward signal r t (positive reward or negative reward), then the agent performs new actions according to a certain strategy based on the new state and the reward of environmental feedback. The goal of reinforcement learning is to find a strategy that maximizes the long-term discounted cumulative reward in a given environment, which can be used in state s t Take action t The cumulative expected discounted reward is calculated by the state-action value function Q π (s,a) means,
[0039]
[0040] S5-1, discretize the capital data corresponding to each subject into time series data with one month as the time interval of each moment.
[0041] S5-2, build a reinforcement learning environment and use the funding data and extracted features as the state input of the reinforcement learning agent.
[0042] S5-3, state s t Consists of extracted fund flow features.
[0043] S5-4, Action a t It is the action that the agent can perform in each state, that is, marking a subject as a suspected illegal fund-raising subject or a legal subject, and the value is 0 or 1.
[0044] S5-5, Reward r t It is defined as the sum of positive rewards and negative rewards. Positive rewards are given for correctly identifying the illegal fund-raising entity, while negative rewards are given for failure to identify or misjudgement. The positive reward is positively correlated with the amount involved in the case corresponding to the entity. In addition to being related to the amount involved, negative rewards also need to take into account the losses caused by misjudgment, regulatory penalties, and the human costs of correcting the error.
[0045] S6, based on the reinforcement learning model, try to select a most suitable reinforcement learning algorithm for model training, including DQN, A3C, DDPG, PPO, TD3 and other algorithms. The algorithm selection principles include the continuity / discreteness of the state and action space, the iteration method and the learning efficiency.
[0046] S7, train the reinforcement learning model.
[0047] S7-1, initialize the parameters of the reinforcement learning agent, including the weights of the neural network, learning rate, discount factor, etc.
[0048] S7-2, the agent continuously interacts with the data in the environment, obtains feedback (rewards) by performing actions (marked subjects), and adjusts strategies.
[0049] S7-3, stores the agent’s historical experience in the memory bank and performs random sampling training to reduce the correlation between samples and stabilize the training process.
[0050] S7-4, based on the rewards and state transitions obtained, use optimization methods such as gradient descent to update the agent's policy parameters to improve recognition accuracy.
[0051] S8, comprehensively compare the performance of each algorithm based on indicators such as convergence speed, recognition accuracy, and training efficiency, and select the best algorithm.
[0052] S9, after determining the reinforcement learning algorithm with the best overall performance, carry out reinforcement learning optimization steps such as environmental testing, hyperparameter testing, and network structure optimization to gradually improve the performance of the algorithm and train the reinforcement learning model based on the optimal parameter settings.
[0053] S10, integrates the trained reinforcement learning model into the illegal fundraising risk identification system, connects it with the existing financial regulatory system and data analysis platform, and realizes automated and intelligent risk identification.
[0054] S11, based on the recognition results of the reinforcement learning agent, establish an early warning mechanism for illegal fund-raising. When a suspicious entity is identified, promptly notify the regulatory agency for further investigation and disposal.
[0055] Based on the above method, the present invention provides a system for identifying suspicious entities of illegal fundraising based on reinforcement learning, including an information combing module, a data analysis module, a data feature extraction module, a data feature processing module, a model building module, a model selection module, a model training module, an algorithm selection module, an algorithm optimization module, a risk identification module and an early warning module.
[0056] The information sorting module analyzes the data characteristics of illegal fund-raising funds based on the historical transaction flows of entities suspected of illegal fund-raising, and sorts out the types of fund risks and key characteristics.
[0057] The data analysis module sorts out the key data items of each type of risk based on risk type and risk characteristics.
[0058] The data feature extraction module collects historical transaction flow data of entities suspected / not suspected of illegal fund-raising, pre-processes the original data through data cleaning, missing value processing, data standardization, data coding, data statistics and other means, and extracts key data features.
[0059] The data feature processing module forms a time-series illegal fund-raising data feature table corresponding to each risk subject at a time interval of one month.
[0060] Model building module, builds reinforcement learning models.
[0061] Reinforcement learning mainly consists of an agent, an environment, a state, an action, and a reward. The agent observes the state s at the current time t from the environment. t Then an action a is selected according to the strategy π t , the agent executed a t After that, the environment will transition to a new state s t+1 , for the new state s t+1 The environment will give a reward signal r t (positive reward or negative reward), then the agent performs new actions according to a certain strategy based on the new state and the reward of environmental feedback. The goal of reinforcement learning is to find a strategy that maximizes the long-term discounted cumulative reward in a given environment, which can be used in state s t Take action t The cumulative expected discounted reward is calculated by the state-action value function Qπ (s,a) means,
[0062]
[0063] The model selection module attempts to select a most suitable reinforcement learning algorithm for model training based on the reinforcement learning model, including DQN, A3C, DDPG, PPO, TD3 and other algorithms. The algorithm selection principles include the continuity / discreteness of the state and action space, the iteration method and the learning efficiency.
[0064] Model training module, training reinforcement learning model.
[0065] The algorithm selection module compares the performance of various algorithms based on indicators such as convergence speed, recognition accuracy, and training efficiency, and selects the best algorithm;
[0066] The algorithm optimization module determines the reinforcement learning algorithm with the best overall performance and then carries out reinforcement learning optimization steps such as environmental testing, hyperparameter testing, and network structure optimization to gradually improve the performance of the algorithm and train the reinforcement learning model based on the optimal parameter settings.
[0067] The risk identification module integrates the trained reinforcement learning model into the illegal fundraising risk identification system, connects it with the existing financial regulatory system and data analysis platform, and realizes automated and intelligent risk identification.
[0068] The early warning module establishes an early warning mechanism for illegal fund-raising based on the recognition results of the reinforcement learning agent. When a suspicious entity is identified, the regulatory agency is notified in a timely manner for further investigation and disposal.
[0069] Functions and Effects of the Embodiments
[0070] According to the reinforcement learning-based method and system for identifying suspicious entities of illegal fund-raising involved in the present invention, the reinforcement learning-based method for identifying suspicious entities of illegal fund-raising proposed in the present invention utilizes key fund characteristics and data to train reinforcement learning agents, so that the agents can accurately identify suspicious entities of illegal fund-raising, thereby improving the accuracy and efficiency of risk identification.
[0071] The present invention innovatively adopts a method based on fund data and reinforcement learning, effectively solving the deficiencies in existing illegal fund-raising risk identification methods, and provides a new idea and new means for illegal fund-raising risk identification. In addition, the reinforcement learning method can fully tap the implicit information in the fund data, identify abnormal patterns and risk points that are difficult to detect with traditional rule models and supervised learning methods, and provide regulators with an efficient and intelligent illegal fund-raising risk identification tool.
[0072] The present invention improves the accuracy and efficiency of risk identification by effectively utilizing capital data and dynamically adjusting identification strategies, and has strong dynamic adaptability and real-time learning capabilities.
[0073] Those skilled in the art should understand that the present invention is not limited to the above embodiments, and the above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, and these changes and improvements fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.
Claims
1. A method for identifying suspicious entities of illegal fund-raising based on reinforcement learning, characterized in that: The specific steps include: S1, analyze the data characteristics of illegal fund-raising funds based on the historical transaction flow of entities suspected of illegal fund-raising, and sort out the types of fund risks and key characteristics; S2, sort out the key data items of each type of risk based on risk type and risk characteristics; S3, collect historical transaction flow data of entities suspected / not suspected of illegal fund-raising, pre-process the original data through data cleaning, missing value processing, data standardization, data coding, data statistics and other means, and extract key data features; S4, taking one month as a time interval, forming a time series illegal fund-raising data feature table corresponding to each risk subject; S5, build reinforcement learning model, Reinforcement learning mainly consists of an agent, an environment, a state, an action, and a reward. The agent observes the state s at the current time t from the environment. t Then an action a is selected according to the strategy π t , the agent executed a t After that, the environment will transition to a new state s t+1 , for the new state s t+1 The environment will give a reward signal r t (positive reward or negative reward), then the agent performs new actions according to a certain strategy based on the new state and the reward of environmental feedback. The goal of reinforcement learning is to find a strategy that maximizes the long-term discounted cumulative reward in a given environment, which can be used in state s t Take action t The cumulative expected discounted reward is calculated by the state-action value function Q π (s,a) means, S6, based on the reinforcement learning model, try to select a most suitable reinforcement learning algorithm for model training, including DQN, A3C, DDPG, PPO, TD3 and other algorithms. The selection principles of the algorithm include the continuity / discreteness of the state and action space, the iteration method and the learning efficiency; S7, training reinforcement learning model; S8, comprehensively compare the performance of each algorithm based on indicators such as convergence speed, recognition accuracy, and training efficiency, and select the best algorithm; S9, after determining the reinforcement learning algorithm with the best overall performance, carry out reinforcement learning optimization steps such as environmental testing, hyperparameter testing, and network structure optimization to gradually improve the performance of the algorithm and train the reinforcement learning model based on the optimal parameter settings; S10, integrate the trained reinforcement learning model into the illegal fund-raising risk identification system, connect it with the existing financial supervision system and data analysis platform, and realize automated and intelligent risk identification; S11, based on the recognition results of the reinforcement learning agent, establish an early warning mechanism for illegal fundraising. When a suspicious entity is identified, promptly notify the regulatory agency for further investigation and disposal.
2. According to claim 1, the method for identifying suspicious entities of illegal fund-raising based on reinforcement learning, Features: Wherein, the step S1 includes the following sub-steps: S1-1, the risk of illegal fund-raising includes eight risk types: daily transaction scale, dispersed transfer-in, concentrated transfer-out, abnormal time transaction, 24-hour non-stop transaction, risky transaction object, abnormal transaction mode and ambiguous transaction; S1-2, daily transaction scale includes two features: daily transaction number and daily transaction amount; dispersed transfer-in includes two features: single transfer-in amount is a whole ten thousand and the number of monthly trading counterparties; concentrated transfer-out includes one feature: the number of daily trading counterparties; abnormal time transactions include one feature: the number of abnormal time transactions; 24-hour uninterrupted transactions include one feature: the account has transactions every hour for 24 consecutive hours; risky transaction objects include one feature: the number of risky enterprises / individuals; abnormal transaction patterns include three features: the number of repeated small transactions, the number of repeated transactions with specific amounts, and third-party account transfers; fuzzy transactions include one feature: the number of monthly transactions with characteristic words such as "investment return", "fund raising", and "loan repayment" in the remarks.
3. The method for identifying suspicious entities of illegal fund-raising based on reinforcement learning according to claim 2 is characterized by: in, The key data items for daily transaction scale are transaction account, transaction ID, transaction date, and transaction amount; the key data items for decentralized transfer-in are transaction account, transaction date, transaction amount, and counterparty account; the key data items for centralized transfer-out are transaction account, transaction date, and counterparty account; the key data items for abnormal time transactions are transaction account, transaction date, transaction time, and transaction ID; the key data items for 24-hour uninterrupted transactions are transaction account, transaction date, transaction time, and transaction ID; the key data items for risky transaction objects are transaction account, counterparty account, and list of risky enterprises / individuals; the key data items for abnormal transaction patterns are transaction account, transaction amount, transaction time, and counterparty account; the key data items for ambiguous transactions are transaction account, transaction remark, and transaction ID.
4. The method for identifying suspicious entities of illegal fund-raising based on reinforcement learning according to claim 1 is characterized by: in, The step S5 includes the following sub-steps: S5-1, with one month as the time interval for each moment, discretize the capital data corresponding to each subject into time series data; S5-2, build a reinforcement learning environment, taking the funding data and extracted features as the state input of the reinforcement learning agent; S5-3, state s t It is composed of extracted fund flow characteristics; S5-4, Action a t is the action that the agent can perform in each state, i.e. marking a subject as a suspected illegal fund-raising subject or a legal subject, with a value of 0 or 1; S5-5, Reward r t It is defined as the sum of positive rewards and negative rewards. Positive rewards are given for correctly identifying the illegal fund-raising entity, while negative rewards are given for failure to identify or misjudgement. The positive reward is positively correlated with the amount involved in the case corresponding to the entity. In addition to being related to the amount involved, negative rewards also need to take into account the losses caused by misjudgment, regulatory penalties, and the human costs of correcting the error.
5. The method for identifying suspicious entities of illegal fund-raising based on reinforcement learning according to claim 1 is characterized by: in, The step S7 includes the following sub-steps: S7-1, initialize the parameters of the reinforcement learning agent, including the weights of the neural network, learning rate, discount factor, etc.; S7-2, the agent continuously interacts with data in the environment, obtains feedback (rewards) by performing actions (marking subjects), and adjusts strategies; S7-3, stores the agent's historical experience in the memory bank and performs random sampling training to reduce the correlation between samples and stabilize the training process; S7-4, based on the rewards and state transitions obtained, use optimization methods such as gradient descent to update the agent's policy parameters to improve recognition accuracy.
6. A system for identifying suspicious entities of illegal fund-raising based on reinforcement learning, characterized in that: include: The information sorting module analyzes the data characteristics of illegal fund-raising funds based on the historical transaction flows of entities suspected of illegal fund-raising, and sorts out the types of fund risks and key characteristics; The data analysis module sorts out the key data items of each type of risk based on risk type and risk characteristics; The data feature extraction module collects historical transaction flow data of entities suspected or not suspected of illegal fund-raising, pre-processes the original data through data cleaning, missing value processing, data standardization, data coding, data statistics, etc., and extracts key data features; The data feature processing module forms a time series illegal fund-raising data feature table corresponding to each risk subject at a time interval of one month; Model building module, building reinforcement learning models, Reinforcement learning mainly consists of an agent, an environment, a state, an action, and a reward. The agent observes the state s at the current time t from the environment. t Then an action a is selected according to the strategy π t , the agent executed a t After that, the environment will transition to a new state s t+1 , for the new state s t+1 The environment will give a reward signal r t (positive reward or negative reward), then the agent performs new actions according to a certain strategy based on the new state and the reward of environmental feedback. The goal of reinforcement learning is to find a strategy that maximizes the long-term discounted cumulative reward in a given environment, which can be used in state s t Take action t The cumulative expected discounted reward is calculated by the state-action value function Q π (s,a) means, The model selection module attempts to select a most suitable reinforcement learning algorithm for model training based on the reinforcement learning model, including DQN, A3C, DDPG, PPO, TD3 and other algorithms. The algorithm selection principles include the continuity / discreteness of the state and action space, the iteration method and the learning efficiency. Model training module, training reinforcement learning model; The algorithm selection module compares the performance of various algorithms based on indicators such as convergence speed, recognition accuracy, and training efficiency, and selects the best algorithm; The algorithm optimization module determines the reinforcement learning algorithm with the best overall performance and then conducts reinforcement learning optimization steps such as environmental testing, hyperparameter testing, and network structure optimization to gradually improve the performance of the algorithm and train the reinforcement learning model based on the optimal parameter settings; The risk identification module integrates the trained reinforcement learning model into the illegal fund-raising risk identification system, connects with the existing financial supervision system and data analysis platform, and realizes automated and intelligent risk identification; The early warning module establishes an early warning mechanism for illegal fund-raising based on the recognition results of the reinforcement learning agent. When a suspicious entity is identified, the regulatory agency is notified in a timely manner for further investigation and disposal.
Citation Information
Cited By
Mobile phone lease risk control optimization method and system based on multi-dimensional AI model analysis
CN120975537A
Mobile phone leasing risk control optimization method and system based on multi-dimensional AI model analysis
CN120975537B
Customer identification and dynamic supervision analysis system and method based on intelligent agent
CN121010380A
A Customer Identification and Dynamic Monitoring Analysis System and Method Based on Intelligent Agents
CN121010380B
Asset transaction risk monitoring system and method
CN121860748A