Recommended method and device for risk identification strategy
By periodically acquiring and testing data to generate risk identification strategies that meet the set identification quality, the problems of high cost, long time and low credibility of manual experience generation strategies are solved, and efficient and reliable risk identification is achieved.
Patent Information
- Application Number
- CN202210725695.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-24
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2042-06-24
AI Technical Summary
In the existing technology, relying on artificial historical experience to generate risk identification strategies results in high risk identification costs, long time and low credibility.
By periodically acquiring multiple training data and test data, the decision tree classification algorithm is used to generate candidate recognition strategies that meet the set recognition quality. The recognition quality is determined through test data, and efficient risk identification strategies are recommended.
It reduces the processing cost and time of risk identification, while improving the credibility of risk identification.
Smart Images

Figure CN114997704B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer information processing technology, and in particular to a method and device for recommending a risk identification strategy. Background Art
[0002] With the rapid development of internet technology, more and more merchants are able to conduct online transactions with users through online platforms. Generally, to ensure transaction security, online platforms can identify the risks of merchants before they conduct online transactions. If they determine that the merchants do not pose any risks, they will allow merchants to conduct online transactions with users.
[0003] In related technologies, risk identification strategies are mainly generated based on manual historical experience to perform risk identification. However, this manual judgment method increases labor costs and processing time, and reduces the credibility of risk identification. Summary of the Invention
[0004] The present disclosure aims to solve one of the technical problems in the related art at least to a certain extent.
[0005] To this end, the first purpose of the present disclosure is to propose a method for recommending risk identification strategies, so as to automatically generate at least one group of candidate identification strategies that meet the set identification quality based on multiple training data obtained periodically and risk labels annotated with the training data. At the same time, multiple test data annotated with risk labels obtained periodically are used to test the candidate identification strategies, and risk identification strategies are recommended based on the identification quality of each group of candidate identification strategies obtained from the test. Then, risk identification is performed according to the recommended risk identification strategy, which reduces the processing cost and time of risk identification and improves the credibility of risk identification.
[0006] The second objective of the present disclosure is to provide a device for recommending risk identification strategies.
[0007] A third objective of the present disclosure is to provide an electronic device.
[0008] A fourth object of the present disclosure is to provide a computer-readable storage medium.
[0009] A fifth object of the present disclosure is to provide a computer program product.
[0010] To achieve the above-mentioned purpose, the first embodiment of the present disclosure proposes a method for recommending a risk identification strategy, including: periodically obtaining multiple training data and multiple test data; wherein each training data and each test data is marked with a risk label; for any target period, based on the multiple training data obtained in the target period and the risk labels marked with each training data, generate at least one group of candidate identification strategies that meet the set identification quality; using the multiple test data marked with risk labels obtained in the target period, test each of the candidate identification strategies separately to determine the identification quality of each group of candidate identification strategies; and based on the identification quality of each group of candidate identification strategies, recommend a risk identification strategy for the target period.
[0011] The risk identification strategy recommendation method of the disclosed embodiment is to periodically obtain multiple training data and multiple test data; wherein each training data and each test data is labeled with a risk label; for any target period, based on the multiple training data obtained in the target period and the risk labels labeled on each training data, generate at least one set of candidate identification strategies that meet the set recognition quality; use the multiple test data labeled with risk labels obtained in the target period to test each candidate identification strategy to determine the recognition quality of each group of candidate identification strategies; and recommend risk identification strategies for the target period based on the recognition quality of each group of candidate identification strategies. The method automatically generates at least one set of candidate identification strategies that meet the set recognition quality based on the multiple training data obtained periodically and the risk labels labeled on the training data; at the same time, uses the multiple test data labeled with risk labels obtained periodically to test the candidate identification strategies, and recommends risk identification strategies based on the recognition quality of each group of candidate identification strategies obtained from the test. Then, risk identification is performed based on the recommended risk identification strategies, reducing the processing cost and time of risk identification and improving the credibility of risk identification.
[0012] To achieve the above-mentioned purpose, the second aspect embodiment of the present disclosure proposes a risk identification strategy recommendation device, including: an acquisition module, used to periodically acquire multiple training data and multiple test data; wherein, each of the training data and each of the test data is marked with a risk label; a generation module, used to generate at least one group of candidate identification strategies that meet the set identification quality for any target period based on the multiple training data acquired in the target period and the risk labels marked on each of the training data; a testing module, used to use the multiple test data marked with risk labels acquired in the target period to test each of the candidate identification strategies separately to determine the identification quality of each group of the candidate identification strategies; a recommendation module, used to recommend the risk identification strategy for the target period based on the identification quality of each group of the candidate identification strategies.
[0013] To achieve the above-mentioned purpose, the third aspect embodiment of the present disclosure proposes an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and runnable on the processor, characterized in that when the processor executes the program, the recommendation method of the risk identification strategy described in the first aspect embodiment of the present disclosure is implemented.
[0014] In order to achieve the above-mentioned objectives, the fourth embodiment of the present disclosure proposes a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, it implements the recommendation method of the risk identification strategy described in the first embodiment of the present disclosure.
[0015] In order to achieve the above-mentioned objectives, the fifth embodiment of the present disclosure proposes a computer program product, which, when the instruction processor in the computer program product is executed, executes the risk identification strategy recommendation method described in the first embodiment of the present disclosure.
[0016] Additional aspects and advantages of the present disclosure will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The above and / or additional aspects and advantages of the present disclosure will become apparent and readily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0018] Figure 1 A flowchart of a recommended method for risk identification strategy provided by one embodiment of the present disclosure;
[0019] Figure 2 A flowchart of a recommended method for risk identification strategy provided by another embodiment of the present disclosure;
[0020] Figure 3 A flowchart of a recommended method for risk identification strategy provided by another embodiment of the present disclosure;
[0021] Figure 4 A flowchart of a recommended method for risk identification strategy provided by another embodiment of the present disclosure;
[0022] Figure 5 A flowchart of a recommended method for risk identification strategy provided by another embodiment of the present disclosure;
[0023] Figure 6 A schematic diagram of a risk identification strategy recommendation system provided by one embodiment of the present disclosure;
[0024] Figure 7 A flowchart of a method for recommending a risk identification strategy according to another embodiment of the present disclosure;
[0025] Figure 8 A schematic diagram of the structure of a device for recommending risk identification strategies provided by one embodiment of the present disclosure;
[0026] Figure 9 The present invention is a block diagram of an electronic device showing a method for recommending a risk identification strategy according to an exemplary embodiment. DETAILED DESCRIPTION
[0027] The following describes in detail embodiments of the present disclosure, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present disclosure, and should not be construed as limiting the present disclosure.
[0028] In the technical solution disclosed herein, the acquisition, collection, storage, use, processing, transmission, provision, disclosure and application of data comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good morals.
[0029] Risk identification strategies are increasingly being used in various business scenarios (e.g., Internet transaction scenarios). Various risk identification strategies are obtained through processing business data, processing and selecting features, and adjusting models and parameters. These strategies are then run online and used to handle various business scenarios where risks may occur by adjusting risk identification strategy thresholds.
[0030] At present, risk identification strategy platforms generally run multiple risk identification strategies together in a certain combination to play a role in risk control; in addition, various risk identification strategies have their own threshold adjustment ranges and methods, but these threshold adjustments are mainly determined by the experience of algorithm developers and specific business scenarios.
[0031] However, the development of a specific risk identification strategy and its threshold adjustment relies on the experience of the algorithm developers behind that strategy, making it difficult to replicate in actual deployment and achieving refined control. Furthermore, risk identification strategies are often manually developed by analysts analyzing cases across various scenarios, making it difficult to identify the underlying strategies between risk characteristics and risk labels. Discovering combinations of strategies is even more challenging. Furthermore, in business scenarios where multiple risk identification strategies are combined, each strategy has its own threshold adjustment, making global optimization and refined control difficult. Furthermore, it's difficult to implement unified global risk strategy threshold adjustment and control for a specific business objective to achieve the ultimate business goal. To effectively identify risks, risk identification strategies must be frequently updated and iterated, but existing purely manual evaluation methods struggle to adapt to changes in a timely manner.
[0032] In response to the above problems, the present disclosure proposes a method and device for recommending risk identification strategies.
[0033] The following describes a method and apparatus for recommending risk identification strategies according to an embodiment of the present disclosure with reference to the accompanying drawings.
[0034] Figure 1 This is a flow chart of a method for recommending a risk identification strategy according to one embodiment of the present disclosure. It should be noted that this method for recommending a risk identification strategy can be applied to a device for recommending a risk identification strategy. This device can be configured in an electronic device. The electronic device can be a mobile terminal, such as a mobile phone, tablet computer, personal digital assistant, or other hardware device with various operating systems.
[0035] like Figure 1 As shown, the recommended approach for this risk identification strategy includes the following steps:
[0036] Step 101: Periodically obtain multiple pieces of training data and multiple pieces of test data; wherein each piece of training data and each piece of test data is marked with a risk label.
[0037] In an embodiment of the present disclosure, multiple target data records corresponding to the risk control attributes under the risk control scenario can be generated periodically (e.g., at 00:00 every day), and the multiple target data records can be labeled with risk labels, and the labeled target data records can be divided into training data and test data. For example, using "gender" and "shopping price" in a shopping scenario as risk control attributes, multiple shopping data records corresponding to the risk control attributes are generated, and the shopping data records are labeled as to whether they have risks, and the risk-labeled shopping data records are divided into training data and test data. For example, shopping data records from 1 to 4 months from the current time are used as training data, and shopping data records from 0 to 1 month from the current time are used as test data.
[0038] Step 102 : for any target period, based on multiple training data acquired in the target period and risk labels annotated on each training data, generate at least one set of candidate recognition strategies that meet the set recognition quality.
[0039] Optionally, for any target period, based on multiple training data obtained in the target period and risk labels annotated with each training data, a decision tree classification algorithm is used to generate at least one set of candidate recognition strategies that meet the set recognition quality.
[0040] As a possible implementation method of the embodiment of the present disclosure, for any target period, multiple training data obtained in the target period and the risk labels annotated with each training data can be input into a strategy generation model (e.g., a decision tree classification model). The strategy generation model can output at least one group of candidate recognition strategies that meet the set recognition quality, wherein the recognition quality of the candidate recognition strategies for risk recognition of multiple training data is greater than the corresponding set threshold, and the recognition quality includes accuracy and / or recall. It should be noted that each group of candidate recognition strategies can correspond to a decision classification tree, and each decision classification tree can contain less than or equal to a set number of decision sub-branches, and each decision sub-branch can be a judgment condition.
[0041] Step 103 : Using a plurality of test data pieces annotated with risk labels acquired during the target period, each candidate recognition strategy is tested respectively to determine the recognition quality of each group of candidate recognition strategies.
[0042] Furthermore, in order to determine the recognition strategies with higher interpretability and recognition quality among the candidate recognition strategies, multiple test data marked with risk labels obtained in the target period can be used to test each group of candidate recognition strategies separately to obtain the recognition quality of each group of candidate recognition strategies.
[0043] Step 104 : Recommend a risk identification strategy for the target period based on the identification quality of each group of candidate identification strategies.
[0044] As a possible implementation method of the embodiment of the present disclosure, each group of candidate identification strategies can be sorted according to the recognition quality of each group of candidate identification strategies, and the candidate identification strategies with recognition quality lower than the set quality threshold can be discarded, and the candidate identification strategies with recognition quality higher than the set quality threshold can be retained. Based on the retained candidate identification strategies, the risk identification strategy recommended for the target period can be selected.
[0045] As another possible implementation method of the embodiment of the present disclosure, each group of candidate identification strategies can be sorted from high to low according to their recognition quality, and a set number of candidate identification strategies ranked at the top can be retained. Based on the retained candidate identification strategies, the risk identification strategy recommended for the target period can be selected.
[0046] In summary, at least one group of candidate recognition strategies that meet the set recognition quality is automatically generated based on multiple training data obtained periodically and the risk labels annotated with the training data. At the same time, multiple test data annotated with risk labels obtained periodically are used to test the candidate recognition strategies, and risk recognition strategies are recommended based on the recognition quality of each group of candidate recognition strategies obtained from the test. Then, risk identification is performed based on the recommended risk identification strategy, which reduces the processing cost and time of risk identification and improves the credibility of risk identification.
[0047] In order to clearly illustrate how to recommend a risk identification strategy for a target period based on the identification quality of each group of candidate identification strategies, the embodiment of the present disclosure proposes another risk identification strategy recommendation method. Figure 2 A flowchart of a recommended method for risk identification strategy provided by another embodiment of the present disclosure.
[0048] like Figure 2 As shown, the recommended approach for the risk identification strategy may include the following steps:
[0049] Step 201 : Periodically obtain multiple pieces of training data and multiple pieces of test data; wherein each piece of training data and each piece of test data is marked with a risk label.
[0050] Step 202 : for any target period, based on multiple training data acquired in the target period and risk labels annotated on each training data, generate at least one set of candidate recognition strategies that meet the set recognition quality.
[0051] In step 203 , a plurality of test data pieces annotated with risk labels obtained in a target period are used to test each candidate recognition strategy respectively to determine the recognition quality of each group of candidate recognition strategies.
[0052] Step 204 : Screen at least one group of candidate recognition strategies based on the recognition quality of each group of candidate recognition strategies to obtain a retained target strategy.
[0053] As a possible implementation method of the embodiment of the present disclosure, the recognition quality of each group of candidate recognition strategies can be compared with a set quality threshold, and the recognition strategies in each group of candidate recognition strategies with a recognition quality lower than the set quality threshold can be discarded, and the candidate recognition strategies with a recognition quality greater than or equal to the set quality threshold can be retained, and the retained candidate recognition strategies can be used as target strategies.
[0054] Step 205: Recommend a risk identification strategy for the target period based on the target strategy.
[0055] Optionally, obtain the historical identification strategy recommended by the reference period before the target period; generate a strategy set based on the historical identification strategy and the target strategy; use multiple test data marked with risk labels obtained in the target period to test the historical identification strategy to determine the recognition quality of the historical identification strategy; based on the recognition quality of the historical identification strategy and the recognition quality of the candidate identification strategy, select the risk identification strategy recommended for the target period from the strategy set.
[0056] That is to say, in order to recommend a risk identification strategy with higher recognition quality in the target period, the recommended historical identification strategy and the target strategy can be merged to generate a strategy set, and the historical identification strategy can be tested using multiple test data marked with risk labels obtained in the target period to obtain the recognition quality of the historical identification strategy. Then, based on the recognition quality of the historical identification strategy and the recognition quality of the target strategy in the candidate identification strategy, the identification strategies with recognition quality lower than the set quality threshold are screened out, and the identification strategies with recognition quality greater than or equal to the set quality threshold are retained. From the retained identification strategies, the risk identification strategy recommended for the target period is determined. It should be noted that when there are multiple groups of retained identification strategies, the identification strategy with the highest recognition quality can be selected as the risk identification strategy recommended for the target period, or a group of identification strategies can be randomly selected from multiple groups of retained identification strategies as the risk identification strategy recommended for the target period. This is not specifically limited in the present disclosure.
[0057] Among them, when determining the recognition quality of the historical recognition strategy, the historical recognition strategy can be used to perform risk identification on multiple test data to obtain the first prediction label corresponding to each test data. The recognition quality of the historical recognition strategy can be determined based on the number of data pieces with the same risk label and / or different risk labels in each test data. That is to say, by adopting the historical identification strategy to perform risk identification on multiple test data, the first prediction label corresponding to each test data can be obtained. Then, the first prediction label corresponding to each test data is compared with the labeled risk label of the corresponding test data to determine the number of data items with the same labeled risk label and / or the number of data items with different labeled risk labels and the corresponding first prediction label in each test data. Further, according to the number of data items with the same labeled risk label and / or the number of data items with different labeled risk labels and the corresponding first prediction label in each test data, the recognition quality of the historical identification strategy is determined, wherein the number of data items with the same labeled risk label and the corresponding first prediction label in each test data is positively correlated with the recognition quality of the historical identification strategy, that is, the more data items with the same labeled risk label and the corresponding first prediction label in each test data, the higher the recognition quality of the historical identification strategy; the number of data items with different labeled risk labels and the corresponding first prediction label in each test data is negatively correlated with the recognition quality of the historical identification strategy, that is: the more data items with different labeled risk labels and the corresponding first prediction label in each test data, the lower the recognition quality of the historical identification strategy.
[0058] It should be noted that the execution process of steps 201 to 203 can be implemented in any of the embodiments of the present disclosure, and the embodiments of the present disclosure do not limit this and will not be described in detail.
[0059] In summary, at least one group of candidate identification strategies is screened according to the recognition quality of each group of candidate identification strategies to obtain the retained target strategy; the risk identification strategy for the target period is recommended based on the target strategy, thereby recommending a risk identification strategy with higher recognition quality in the target period.
[0060] In order to clearly illustrate how to periodically obtain multiple pieces of training data and multiple pieces of test data, the embodiment of the present disclosure proposes another method for recommending a risk identification strategy. Figure 3 A flowchart of a recommended method for risk identification strategy provided by another embodiment of the present disclosure.
[0061] like Figure 3 As shown, the recommended method for the risk identification strategy may include the following steps:
[0062] Step 301: Periodically collect risk control attributes of multiple user accounts and obtain risk tags annotated by multiple user accounts.
[0063] In the disclosed embodiments, attributes of multiple user accounts can be periodically collected and matched to corresponding risk control scenarios to determine at least one corresponding risk control attribute from the collected attributes. For example, in a shopping scenario, the risk control attributes may be "gender" and "price," and the risk tag may be whether there is a fraud risk.
[0064] In step 302, a corresponding target data record is generated according to the risk control attributes of the same user account, and is annotated with a corresponding risk tag.
[0065] As a possible implementation of the embodiment of the present disclosure, at least one risk control attribute of the same user account may be used as a corresponding target data record, and the target data record may be labeled with a corresponding risk label.
[0066] Step 303 : Divide the target data records corresponding to the multiple user accounts into multiple pieces of training data and multiple pieces of testing data according to a set ratio.
[0067] Furthermore, the target data records corresponding to multiple user accounts can be divided according to a set ratio to obtain multiple training data and multiple test data. For example, 75% of the target data records can be used as training data, and the remaining 25% of the target data records can be used as multiple test data.
[0068] Step 304 : for any target period, at least one set of candidate recognition strategies that meet the set recognition quality is generated based on the multiple training data obtained in the target period and the risk labels annotated on each training data.
[0069] In step 305 , a plurality of test data pieces annotated with risk labels acquired during the target period are used to test each candidate recognition strategy respectively to determine the recognition quality of each group of candidate recognition strategies.
[0070] Step 306 : Recommend a risk identification strategy for the target period based on the identification quality of each group of candidate identification strategies.
[0071] It should be noted that the execution process of steps 304 to 306 can be implemented in any of the embodiments of the present disclosure, and the embodiments of the present disclosure do not limit this and will not be described in detail.
[0072] In summary, by periodically collecting risk control attributes of multiple user accounts and obtaining risk labels for these accounts, generating a corresponding target data record based on the risk control attributes of the same user account and labeling it with the corresponding risk label, and then dividing the target data records corresponding to the multiple user accounts into multiple training data and multiple test data according to a set ratio, multiple training data and multiple test data can be periodically obtained.
[0073] In order to clearly illustrate how to use multiple test data marked with risk labels obtained in a target period to test each candidate identification strategy separately to determine the identification quality of each candidate identification strategy, the embodiment of the present disclosure proposes another risk identification strategy recommendation method. Figure 4 A flowchart of a recommended method for risk identification strategy provided by another embodiment of the present disclosure.
[0074] like Figure 4 As shown, the recommended method for the risk identification strategy may include the following steps:
[0075] Step 401 : Periodically obtain multiple pieces of training data and multiple pieces of test data; wherein each piece of training data and each piece of test data is marked with a risk label.
[0076] Step 402 : for any target period, based on multiple training data acquired in the target period and risk labels annotated on each training data, generate at least one set of candidate recognition strategies that meet the set recognition quality.
[0077] Step 403 : Using any candidate identification strategy, perform risk identification on multiple test data to obtain a second prediction label corresponding to each test data.
[0078] In the embodiment of the present disclosure, any candidate identification strategy can be used to perform risk identification on multiple test data, and a second prediction label corresponding to each test data can be obtained.
[0079] Step 404 : Determine the recognition quality of the corresponding candidate recognition strategy based on the number of data items in which the risk label annotated in each test data item is identical to and / or different from the corresponding second prediction label.
[0080] Next, the second prediction label corresponding to each test data is compared with the labeled risk label of the corresponding test data to determine the number of data items whose labeled risk label is the same as the corresponding second prediction label and / or the number of data items whose labeled risk label is different from the corresponding second prediction label in each test data. Further, the recognition quality of the candidate recognition strategy is determined based on the number of data items whose labeled risk label is the same as the corresponding second prediction label and / or the number of data items whose labeled risk label is different from the corresponding second prediction label in each test data, wherein the number of data items whose labeled risk label is the same as the corresponding second prediction label in each test data is positively correlated with the recognition quality of the candidate recognition strategy, that is, the more data items whose labeled risk label is the same as the corresponding second prediction label in each test data, the higher the recognition quality of the candidate recognition strategy; the number of data items whose labeled risk label is different from the corresponding second prediction label in each test data is negatively correlated with the recognition quality of the candidate recognition strategy, that is: the more data items whose labeled risk label is different from the corresponding second prediction label in each test data, the lower the recognition quality of the candidate recognition strategy.
[0081] Step 405 : Recommend a risk identification strategy for the target period based on the identification quality of each group of candidate identification strategies.
[0082] It should be noted that the execution process of steps 401 to 402 and step 405 can be implemented in any way in the embodiments of the present disclosure, and the embodiments of the present disclosure do not limit this and will not be described in detail.
[0083] In summary, by adopting any candidate identification strategy, risk identification is performed on multiple test data to obtain the second prediction label corresponding to each test data. The recognition quality of the corresponding candidate identification strategy is determined based on the number of data pieces with the same risk label and / or different risk label as the corresponding second prediction label in each test data. Therefore, by testing the candidate identification strategy with multiple test data, the recognition quality of each group of candidate identification strategies can be accurately determined.
[0084] In order to clearly illustrate how to generate at least one set of candidate recognition strategies that meet the set recognition quality based on multiple training data obtained in the target period and the risk labels annotated on each training data, Figure 5 A flowchart of a recommended method for risk identification strategy provided by another embodiment of the present disclosure.
[0085] like Figure 5 As shown, the recommended method for the risk identification strategy may include the following steps:
[0086] Step 501 : Periodically obtain multiple pieces of training data and multiple pieces of test data; wherein each piece of training data and each piece of test data is marked with a risk label.
[0087] Step 502: for any target period, multiple training data obtained in the target period and risk labels annotated on each training data are input into the strategy generation model to obtain candidate identification strategies generated by the strategy generation model according to set strategy generation conditions.
[0088] Among them, the setting strategy generation conditions include at least one of the following: the number of judgment conditions included in the candidate identification strategy is less than or equal to the set number; the recognition quality of the candidate strategy for risk identification of multiple training data is greater than the corresponding set threshold; wherein the recognition quality includes accuracy and / or recall rate.
[0089] As a possible implementation of the embodiment of the present disclosure, the strategy generation model can be a decision tree classification model, and the multiple training data obtained in the target period and the risk labels annotated with each training data can be input into the decision tree classification model, and the maximum depth of the decision tree classification model is set to a set number, and the accuracy and / or recall rate of risk identification for multiple training data is greater than the corresponding set threshold. The decision tree classification model can output the risk probability corresponding to the risk label annotated with each training data, and determine whether each training data has a risk based on the risk probability. It should be noted that the maximum depth of the decision tree classification model is a set number, and the number of judgment conditions (depth) contained in the candidate identification strategy generated by the decision tree classification model can be less than or equal to the set number. In other words, according to the required risk identification accuracy and / or recall rate, multiple combinations of candidate identification strategies can be generated. For example, the maximum depth of the decision tree classification model is 3, and a candidate identification strategy containing one judgment condition, a candidate identification strategy containing two judgment conditions, and a candidate identification strategy containing three judgment conditions can be generated. It should be noted that the candidate identification strategy contains multiple (at least two) judgment conditions. The candidate identification strategy can identify that data that meets multiple (at least two) judgment conditions is risky. For example, the three judgment conditions included in the candidate identification strategy are "A>1", "B>2" and "C>3", that is, the candidate identification strategy identifies that data that meets "A>1", "B>2" and "C>3" at the same time is risky.
[0090] Step 503 : Using a plurality of test data pieces annotated with risk labels acquired during the target period, each candidate recognition strategy is tested respectively to determine the recognition quality of each group of candidate recognition strategies.
[0091] Step 504 : Recommend a risk identification strategy for the target period based on the identification quality of each group of candidate identification strategies.
[0092] It should be noted that the execution process of step 501 and step 503 to step 504 can be implemented in any way in the embodiments of the present disclosure, and the embodiments of the present disclosure do not limit this and will not be described in detail.
[0093] In summary, for any target period, multiple training data obtained in the target period and the risk labels annotated with each training data are input into the strategy generation model to obtain the candidate recognition strategies generated by the strategy generation model according to the set strategy generation conditions. Thus, the strategy generation model can be used to generate candidate recognition strategies that meet the set recognition quality.
[0094] In order to explain the above embodiment more clearly, an example is given below.
[0095] For example, the recommended method of risk identification strategy of the embodiment of the present disclosure is applicable to Figure 6 Taking the risk identification strategy recommendation system shown in the figure as an example, the first step is to prepare the dependent libraries. You can declare the dependent libraries in advance (mainly traditional machine learning, rule learning, explainable machine learning algorithm libraries, and self-developed optimization algorithm libraries). These dependent Python libraries will be automatically installed during installation.
[0096] The second step is data preparation. This involves obtaining features (risk control attributes) and label data (risk labels) related to the risk identification strategy to be generated. The data is then divided into two parts: one for training the model, which can be saved in a file such as "train_data.csv", and the other for testing the model's effectiveness, which can be saved in a file such as "test_data.csv".
[0097] The third step is to configure the algorithm configuration file. The format is as follows:
[0098] "{"max_depth_duplication":6,
[0099] "max_depth":[1,2,3,4,5,6],
[0100] "max_features":1.0,
[0101] "max_samples_features":1.0,
[0102] "n_estimators":100,
[0103] "recall_min":0.04,
[0104] "precision_min": 0.1,
[0105] “features”:[],
[0106] “label”:“is_fraud”
[0107] }";
[0108] This file defines parameters such as the depth (max_depth), minimum precision (precision_min), and minimum recall (recall_min) of the algorithm generation strategy. It also defines feature data (risk control attributes, features) and label item names (risk control labels). For example, a risk control scenario uses "gender" and "price" as feature data. These feature data columns are also included in the training and test data. Add the following to the configuration file:
[0109] "features": ["gender", "price"]"; In addition, a label column needs to be defined. For example, if there is a column "is_fraud" in the training and test data files to indicate that the bank is risky, then add "label": "is_fraud" in the configuration file accordingly;
[0110] Step 4: Strategy generation, run: python pipeline.py -c config.json -i train_data.csv -o rules.txt
[0111] Among them, config.json is the configuration file, train_data.csv is the training data file, and the generation strategy is saved in the rules.txt file.
[0112] The fifth step is strategy scoring. We can re-score the generated strategy using the test data by running:
[0113] python rules_scorer.py -c config.json -d test_data.csv -r rules.txt
[0114] Among them, config.json is the configuration file, test_data.csv is the test data file, and rules.txt is the existing strategy file. The re-scored precision and recall will be saved in rules.txt.
[0115] Correspondingly, such as Figure 7 As shown, Figure 7This is a flow chart of a method for recommending risk identification strategies according to another embodiment of the present disclosure. First, data from 1 to 4 months away from the current time is periodically obtained as training data, and data from the most recent month is used as test data. A high-precision strategy (target strategy) is obtained through training using a strategy recommendation algorithm (e.g., a decision tree classification model). In the presence of a historical strategy (historical identification strategy), the high-precision strategy is merged with the historical strategy to obtain a strategy set. The strategy set is then scored and screened using the test data, and the retained strategies with strong explanatory power and good performance are recommended. It should be noted that each time a new strategy is generated, the existing strategies are merged and scored, strategies that do not meet the performance requirements are eliminated, and strategies that meet the performance requirements are screened out for reference by algorithms and business personnel for selection and launch.
[0116] The risk identification strategy recommendation method of the disclosed embodiment periodically obtains multiple training data and multiple test data, wherein each training data and each test data is labeled with a risk label. For any target period, based on the multiple training data obtained in the target period and the risk labels labeled with each training data, at least one group of candidate identification strategies that meet a set recognition quality is generated. The multiple test data labeled with risk labels obtained in the target period are used to test each candidate identification strategy to determine the recognition quality of each group of candidate identification strategies. Based on the recognition quality of each group of candidate identification strategies, a risk identification strategy for the target period is recommended. The method automatically generates at least one group of candidate identification strategies that meet the set recognition quality based on the multiple training data obtained periodically and the risk labels labeled with the training data. At the same time, the candidate identification strategies are tested using the multiple test data labeled with risk labels obtained periodically. Risk identification strategies are recommended based on the recognition quality of each group of candidate identification strategies obtained from the test. Risk identification is then performed based on the recommended risk identification strategies, reducing the processing cost and time of risk identification while improving the credibility of risk identification.
[0117] To implement the above embodiment, the present disclosure further proposes a device for recommending risk identification strategies.
[0118] Figure 8 A schematic diagram of the structure of a device for recommending risk identification strategies provided by one embodiment of the present disclosure.
[0119] like Figure 8 As shown, the risk identification strategy recommendation device 800 includes: an acquisition module 810 , a generation module 820 , a testing module 830 and a recommendation module 840 .
[0120] Among them, the acquisition module 810 is used to periodically acquire multiple training data and multiple test data; wherein, each training data and each test data is marked with a risk label; the generation module 820 is used to generate at least one group of candidate recognition strategies that meet the set recognition quality for any target period based on the multiple training data acquired in the target period and the risk labels marked on each training data; the testing module 830 is used to use the multiple test data marked with risk labels acquired in the target period to test each candidate recognition strategy separately to determine the recognition quality of each group of candidate recognition strategies; the recommendation module 840 is used to recommend the risk recognition strategy for the target period based on the recognition quality of each group of candidate recognition strategies.
[0121] As a possible implementation method of the embodiment of the present disclosure, the recommendation module 840 is also used to: screen at least one group of candidate identification strategies based on the identification quality of each group of candidate identification strategies to obtain a retained target strategy; and recommend a risk identification strategy for a target period based on the target strategy.
[0122] As a possible implementation method of an embodiment of the present disclosure, the recommendation module 840 is also used to: obtain the historical identification strategy recommended by the reference period before the target period; generate a strategy set based on the historical identification strategy and the target strategy; use multiple test data marked with risk tags obtained in the target period to test the historical identification strategy to determine the recognition quality of the historical identification strategy; select the risk identification strategy recommended for the target period from the strategy set based on the recognition quality of the historical identification strategy and the recognition quality of the candidate identification strategy.
[0123] As a possible implementation method of the embodiment of the present disclosure, the recommendation module 840 is also used to: adopt a historical identification strategy to perform risk identification on multiple test data to obtain a first prediction label corresponding to each test data; determine the recognition quality of the historical identification strategy based on the number of data items in each test data whose marked risk label is the same as the corresponding first prediction label and / or the number of data items that are different.
[0124] As a possible implementation method of the embodiment of the present disclosure, the acquisition module 810 is also used to: periodically collect the risk control attributes of multiple user accounts, and obtain risk labels annotated by multiple user accounts; generate a corresponding target data record based on the risk control attributes of the same user account, and annotate it with the corresponding risk label; divide the target data records corresponding to multiple user accounts into multiple training data and multiple test data according to a set ratio.
[0125] As a possible implementation method of the embodiment of the present disclosure, the test module 830 is also used to: adopt any candidate identification strategy to perform risk identification on multiple test data to obtain a second prediction label corresponding to each test data; determine the recognition quality of the corresponding candidate identification strategy based on the number of data items with the same risk label and / or different risk label as the corresponding second prediction label in each test data.
[0126] As a possible implementation method of the embodiment of the present disclosure, the generation module 820 is also used to: for any target period, input multiple training data obtained in the target period and the risk labels annotated with each training data into the strategy generation model to obtain a candidate identification strategy generated by the strategy generation model according to the set strategy generation conditions; wherein the set strategy generation conditions include at least one of the following: the number of judgment conditions included in the candidate identification strategy is less than or equal to the set number; the recognition quality of the candidate strategy for risk identification of multiple training data is greater than the corresponding set threshold; wherein the recognition quality includes accuracy and / or recall rate.
[0127] The risk identification strategy recommendation device of the disclosed embodiment periodically obtains multiple training data and multiple test data, wherein each training data and each test data is labeled with a risk label. For any target period, based on the multiple training data obtained in the target period and the risk labels labeled with each training data, at least one group of candidate identification strategies that meet a set recognition quality is generated. The multiple test data labeled with risk labels obtained in the target period are used to test each candidate identification strategy to determine the recognition quality of each group of candidate identification strategies. Based on the recognition quality of each group of candidate identification strategies, a risk identification strategy for the target period is recommended. The device can automatically generate at least one group of candidate identification strategies that meet the set recognition quality based on the multiple training data obtained periodically and the risk labels labeled with the training data. At the same time, the candidate identification strategies are tested using the multiple test data labeled with risk labels obtained periodically, and a risk identification strategy is recommended based on the recognition quality of each group of candidate identification strategies obtained from the test. Risk identification is then performed based on the recommended risk identification strategy, thereby reducing the processing cost and time of risk identification and improving the credibility of risk identification.
[0128] It should be noted that the above explanations of the embodiment of the method for recommending a risk identification strategy are also applicable to the device for recommending a risk identification strategy in this embodiment, and will not be repeated here.
[0129] In order to implement the above embodiments, the present disclosure also proposes an electronic device, including: a memory, a processor, and a computer program stored in the memory and runnable on the processor, characterized in that when the processor executes the program, the recommendation method of the risk identification strategy described in the above embodiments is implemented.
[0130] In order to implement the above embodiments, the present disclosure further proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the risk identification strategy recommendation method described in the above embodiments.
[0131] In order to implement the above embodiments, the present disclosure further proposes a computer program product, which, when an instruction processor in the computer program product is executed, implements the risk identification strategy recommendation method described in the above embodiments.
[0132] In order to implement the above embodiment, the present disclosure also proposes an electronic device, such as Figure 9 As shown, Figure 9 The present invention is a block diagram of an electronic device showing a method for recommending a risk identification strategy according to an exemplary embodiment.
[0133] like Figure 9 As shown, the electronic device 900 includes:
[0134] The memory 910 and the processor 920, a bus 930 connecting different components (including the memory 910 and the processor 920), the memory 910 stores a computer program, and when the processor 920 executes the program, the recommendation method of the risk identification strategy described in the embodiment of the present disclosure is implemented.
[0135] Bus 930 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor or a local bus using any of a variety of bus architectures. Examples of these architectures include, but are not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MAC) bus, an Enhanced ISA bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.
[0136] The electronic device 900 typically includes a variety of electronic device-readable media, which can be any available media that can be accessed by the electronic device 900, including volatile and non-volatile media, removable and non-removable media.
[0137] The memory 910 may also include computer system readable media in the form of volatile memory, such as random access memory (RAM) 940 and / or cache memory 950. The electronic device 900 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 960 may be used to read and write non-removable, non-volatile magnetic media ( Figure 9 Not shown, often called a "hard drive"). Although Figure 9Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to the bus 930 via one or more data medium interfaces. The memory 910 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the various embodiments of the present disclosure.
[0138] A program / utility 980 having a set (at least one) of program modules 970 may be stored, for example, in memory 910. Such program modules 970 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which, or some combination thereof, may include an implementation of a network environment. Program modules 970 generally implement the functions and / or methods of the embodiments described herein.
[0139] The electronic device 900 may also communicate with one or more external devices 990 (e.g., a keyboard, a pointing device, a display, etc.), one or more devices that enable a user to interact with the electronic device 900, and / or any device that enables the electronic device 900 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication may be performed through an input / output (I / O) interface 992. Furthermore, the electronic device 900 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 993. Figure 9 As shown, the network adapter 993 communicates with other modules of the electronic device 900 via the bus 930. Figure 9 Not shown, other hardware and / or software modules may be used in conjunction with electronic device 900, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0140] The processor 920 executes various functional applications and data processing by running programs stored in the memory 910 .
[0141] It should be noted that the implementation process and technical principles of the electronic device of this embodiment can be found in Figures 1 to 7 The explanation of the recommended method for the risk identification strategy of the embodiment of the present disclosure will not be repeated here.
[0142] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification and features of different embodiments or examples, unless they are mutually inconsistent.
[0143] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the present disclosure, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0144] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present disclosure includes additional implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present disclosure belong.
[0145] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing it in another suitable manner if necessary, and then storing it in a computer memory.
[0146] It should be understood that various parts of the present disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0147] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0148] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a single processing module, or each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium.
[0149] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present disclosure have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. A person of ordinary skill in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present disclosure.
Claims
1. A method for recommending a risk identification strategy, characterized in that: include: Periodically acquiring multiple pieces of training data and multiple pieces of test data; wherein each piece of training data and each piece of test data is labeled with a risk label; For any target period, based on multiple training data obtained in the target period and the risk labels annotated on each of the training data, generate at least one set of candidate recognition strategies that meet the set recognition quality; Using a plurality of test data marked with risk labels obtained during the target period, each of the candidate recognition strategies is tested to determine the recognition quality of each group of the candidate recognition strategies; Recommend a risk identification strategy for the target period based on the identification quality of each group of candidate identification strategies.
2. The recommendation method according to claim 1, characterized in that The recommending of a risk identification strategy for the target period based on the identification quality of each group of candidate identification strategies includes: screening the at least one group of candidate recognition strategies according to the recognition quality of each group of candidate recognition strategies to obtain a retained target strategy; Recommend a risk identification strategy for the target period based on the target strategy.
3. The recommendation method according to claim 2, characterized in that: The recommending of a risk identification strategy for the target period according to the target strategy includes: Obtaining a historical identification strategy recommended by a reference period prior to the target period; generating a policy set according to the historical identification policy and the target policy; Using a plurality of test data pieces marked with risk labels acquired during the target period, the historical identification strategy is tested to determine the identification quality of the historical identification strategy; The risk identification strategy recommended for the target period is selected from the strategy set according to the identification quality of the historical identification strategy and the identification quality of the candidate identification strategy.
4. The recommendation method according to claim 3, characterized in that: The testing of the historical identification strategy using the plurality of test data pieces marked with risk labels obtained during the target period to determine the identification quality of the historical identification strategy includes: Using the historical identification strategy, risk identification is performed on multiple test data to obtain a first prediction label corresponding to each test data; The recognition quality of the historical recognition strategy is determined based on the number of data items in each of the test data whose risk labels are identical to and / or different from the corresponding first prediction labels.
5. The recommendation method according to any one of claims 1 to 4, characterized in that: The periodic acquisition of multiple pieces of training data and multiple pieces of test data includes: Periodically collect risk control attributes of multiple user accounts and obtain risk tags annotated by the multiple user accounts; Generate a corresponding target data record according to the risk control attribute of the same user account, and mark it with a corresponding risk tag; The target data records corresponding to the multiple user accounts are divided into the multiple training data and the multiple test data according to a set ratio.
6. The recommendation method according to any one of claims 1 to 4, characterized in that: The plurality of test data pieces marked with risk labels obtained in the target period are used to test each candidate identification strategy to determine the identification quality of each candidate identification strategy, including: Using any candidate identification strategy to perform risk identification on the plurality of test data to obtain a second prediction label corresponding to each of the test data; The recognition quality of the corresponding candidate recognition strategy is determined based on the number of data items in which the risk label annotated in each of the test data is identical to and / or the number of data items in which the risk label annotated in the test data is different from the corresponding second prediction label.
7. The recommendation method according to any one of claims 1 to 4, characterized in that: For any target period, generating at least one set of candidate recognition strategies that meet the set recognition quality based on the multiple training data obtained in the target period and the risk labels annotated on each of the training data, includes: For any target period, multiple training data obtained in the target period and the risk labels annotated on each training data are input into the strategy generation model to obtain candidate identification strategies generated by the strategy generation model according to the set strategy generation conditions; The setting policy generation condition includes at least one of the following: The number of judgment conditions included in the candidate identification strategy is less than or equal to a set number; The recognition quality of the risk identification performed by the candidate identification strategy on the plurality of training data is greater than a corresponding set threshold; wherein the recognition quality includes accuracy and / or recall rate.
8. A device for recommending risk identification strategies, characterized in that: include: An acquisition module, configured to periodically acquire a plurality of training data and a plurality of test data; wherein each of the training data and each of the test data is marked with a risk label; A generation module is configured to generate, for any target period, at least one set of candidate recognition strategies that meet a set recognition quality based on multiple training data obtained during the target period and risk labels annotated with each training data; A testing module, configured to test each of the candidate identification strategies using a plurality of risk-labeled test data obtained during the target period to determine the recognition quality of each group of the candidate identification strategies; The recommendation module is used to recommend a risk identification strategy for the target period based on the identification quality of each group of candidate identification strategies.
9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 7 when executing the program.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Rule recommendation method, apparatus and device
CN108665142A
Hybrid Techniques for Quality Estimation of a Decision-Making Policy in a Computer System
US20230122472A1