Method and device for identifying shell company, storage medium and electronic device

CN115496364BActive Publication Date: 2026-09-29CHINA CONSTRUCTION BANK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211156462.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-22
Publication Date
2026-09-29
Estimated Expiration
2042-09-22

AI Technical Summary

Technical Problem

[0005]有鉴于此,本发明实施例提供了一种幌子企业识别方法,以解决现有的幌子企业识别方式中,依赖单一数据来源,需人工解读风险场景,识别准确率较低,且需耗费大量人力资源的问题

Benefits of technology

[0048]基于上述本发明实施例提供的一种幌子企业识别方法,包括:当需要对目标企业进行识别时,确定目标企业对应的企业信息,其中包括企业基础数据、交易流水数据、渠道登陆数据、工商数据和征信数据;在企业信息中,确定场景特征数据和每个预设特征维度对应的特征数据;将各个预设特征维度对应的特征数据输入已构建的评分模型,经评分模型处理后,获得目标企业对应的风险评分;将场景特征数据和风险评分输入已构建的场景解释模型,经场景解释模型处理后,获得风险评分对应的场景映射规则;依据风险评分和场景映射规则,确定目标企业对应的识别结果;若识别结果表征目标企业为幌子企业,则确定目标企业对应的欺诈场景,完成目标企业的识别过程。应用本发明实施例提供的方法,可结合企业多维度的特征数据,通过预先构建的评分模型对企业是否为幌子企业进行风险评估,并可通过预先构建的场景解释模型,对风险评分进行规则映射,得到场景映射规则,可根据风险评分和场景映射规则识别企业是否可能为幌子企业,若为幌子企业,可继而识别其可能涉及的欺诈场景是什么。在识别过程中可从多维度的数据中挖掘风险特征,有利于提高识别准确率,其次,可实现欺诈场景的自动化识别,无需依赖于人工处理,可节省人力资源,避免出现人为理解偏差。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115496364B_ABST
    Figure CN115496364B_ABST
Patent Text Reader

Abstract

The application provides a curtain enterprise identification method and device, a storage medium and an electronic device. The method comprises the following steps: when it is necessary to identify a target enterprise, determining enterprise information corresponding to the target enterprise, wherein the enterprise information comprises enterprise basic data, transaction flow data, channel login data, business data and credit data; in the enterprise information, determining scene characteristic data and characteristic data corresponding to each preset characteristic dimension; inputting the characteristic data of each dimension into a scoring model, and obtaining a risk score after processing; inputting the scene characteristic data and the risk score into a scene interpretation model, and obtaining a scene mapping rule after processing; determining an identification result according to the risk score and the scene mapping rule; if the identification result indicates that the target enterprise is a curtain enterprise, determining a fraud scene corresponding to the target enterprise. By applying the method, multi-dimensional data is combined for identification, the identification accuracy can be improved, and automatic identification of the fraud scene can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of risk control technology, and in particular to a method and apparatus, storage medium and electronic device for identifying shell companies. Background Technology

[0002] With the development of financial services, banks and other financial institutions are facing an increasing number of illegal and fraudulent activities. "Fake companies" refer to enterprises that do not actually operate. In financial transactions, criminals often use fake companies to commit illegal and fraudulent acts. Therefore, identifying fake companies is one of the primary tasks in the risk control work of financial institutions.

[0003] Currently, the usual approach is to extract corresponding features based on the company's business registration information, and identify whether the company is a shell company by whether it has features indicating that it has no actual business operations or that it was not established with actual business operations as its starting point.

[0004] Existing methods for identifying shell companies rely solely on characteristics found in business registration information, resulting in limited data sources and low accuracy. Furthermore, these methods only determine whether a company is a shell company; the identification results require manual interpretation of the business context and risk scenarios, consuming significant human resources and being prone to misunderstandings. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide a method for identifying covert enterprises, in order to solve the problems of existing covert enterprise identification methods that rely on a single data source, require manual interpretation of risk scenarios, have low identification accuracy, and require a large amount of human resources.

[0006] This invention also provides a device for identifying fraudulent enterprises, to ensure the practical implementation and application of the above method.

[0007] To achieve the above objectives, the embodiments of the present invention provide the following technical solutions:

[0008] A method for identifying shell companies includes:

[0009] When it is necessary to identify a target enterprise, the enterprise information corresponding to the target enterprise is determined. The enterprise information includes basic enterprise data, transaction flow data, channel login data, business registration data, and credit data.

[0010] In the enterprise information, scene feature data and feature data corresponding to each preset feature dimension are determined;

[0011] The feature data corresponding to each of the preset feature dimensions is input into the constructed scoring model. After processing by the scoring model, the risk score corresponding to the target enterprise is obtained.

[0012] The scene feature data and the risk score are input into the constructed scene interpretation model. After processing by the scene interpretation model, the scene mapping rule corresponding to the risk score is obtained.

[0013] Based on the risk score and the scenario mapping rules, the identification result corresponding to the target enterprise is determined;

[0014] If the identification result indicates that the target company is a front company, then the fraud scenario corresponding to the target company is determined, and the identification process of the target company is completed.

[0015] Optionally, the process of constructing the scoring model in the above method includes:

[0016] Determine the initial sample set corresponding to each preset feature dimension; the initial sample set corresponding to each preset feature dimension includes multiple sample data corresponding to that preset feature dimension.

[0017] For each preset feature dimension, the training sample set corresponding to the preset feature dimension is determined based on the initial sample set corresponding to the preset feature dimension.

[0018] For each preset feature dimension, a sub-model corresponding to the preset feature dimension is constructed based on the training sample set corresponding to the preset feature dimension;

[0019] The sub-models are fused to obtain a fused model, which is then used as the scoring model.

[0020] Optionally, in the above method, determining the training sample set corresponding to the preset feature dimension based on the initial sample set corresponding to the preset feature dimension includes:

[0021] Based on the preset oversampling strategy, the initial sample set corresponding to the preset feature dimension is oversampled to obtain the first sample set;

[0022] Based on a preset data cleaning strategy, the first sample set is cleaned to obtain a second sample set corresponding to the first sample set.

[0023] Based on the preset community detection algorithm, the second sample set is subjected to bad sample diffusion processing to obtain the third sample set corresponding to the second sample set;

[0024] Based on a preset variable derivation strategy, the third sample set is subjected to variable derivation processing to obtain the fourth sample set corresponding to the third sample set;

[0025] Based on a preset variable selection strategy, the fourth sample set is subjected to variable selection processing to obtain a fifth sample set corresponding to the fourth sample set, and the fifth sample set is used as the training sample set corresponding to the preset feature dimension.

[0026] Optionally, in the above method, constructing a sub-model corresponding to the preset feature dimension based on the training sample set corresponding to the preset feature dimension includes:

[0027] Based on a number of preset ensemble tree algorithms, construct an ensemble tree model corresponding to each ensemble tree algorithm;

[0028] For each ensemble tree model, the ensemble tree model is trained based on the training sample set corresponding to the preset feature dimension, and the ensemble tree model that has been trained is determined as a candidate model.

[0029] For each candidate model, the candidate model is verified based on the preset reserved sample set and the cross-time sample set to obtain the verification result corresponding to the candidate model;

[0030] Based on the verification results corresponding to each of the candidate models, a target candidate model is determined among the candidate models, and the target candidate model is determined as the sub-model corresponding to the preset feature dimension.

[0031] Optionally, in the above method, the step of fusing the various sub-models to obtain a fused model includes:

[0032] Determine the weights corresponding to each of the sub-models;

[0033] According to the weights corresponding to each sub-model, the sub-models are weighted and fused, and the fusion result is used as the fusion model.

[0034] Optionally, the construction process of the scenario interpretation model in the above method includes:

[0035] Based on a pre-defined set of scene mapping rules and a pre-defined gradient boosting decision tree algorithm, a gradient boosting decision tree model is constructed.

[0036] Based on the preset scenario explanatory variables and bad sample set, the gradient boosting decision tree model is trained to obtain the trained gradient boosting decision tree model.

[0037] Determine whether the trained gradient boosting decision tree model meets the preset test and verification conditions. If the trained gradient boosting decision tree model meets the test and verification conditions, then the trained gradient boosting decision tree model is determined as the scenario explanation model.

[0038] The above method may optionally include the following preset feature dimensions: basic information dimension, enterprise cash flow dimension, actual controller cash flow dimension, business registration and enterprise credit dimension, and actual controller credit dimension.

[0039] A decoy company identification device, comprising:

[0040] The first determining unit is used to determine the enterprise information corresponding to the target enterprise when it is necessary to identify the target enterprise. The enterprise information includes basic enterprise data, transaction flow data, channel login data, business registration data and credit data.

[0041] The second determining unit is used to determine the scene feature data and the feature data corresponding to each preset feature dimension from the enterprise information.

[0042] The first processing unit is used to input the feature data corresponding to each of the preset feature dimensions into the constructed scoring model, and after processing by the scoring model, obtain the risk score corresponding to the target enterprise.

[0043] The second processing unit is used to input the scene feature data and the risk score into the constructed scene interpretation model, and after processing by the scene interpretation model, obtain the scene mapping rule corresponding to the risk score;

[0044] The third determining unit is used to determine the identification result corresponding to the target enterprise based on the risk score and the scenario mapping rule;

[0045] The fourth determining unit is used to determine the fraud scenario corresponding to the target enterprise if the identification result indicates that the target enterprise is a front enterprise, and to complete the identification process of the target enterprise.

[0046] A storage medium comprising stored instructions, wherein, when the instructions are executed, the device in which the storage medium resides executes the aforementioned method for identifying fraudulent enterprises.

[0047] An electronic device includes a memory and one or more instructions, wherein one or more instructions are stored in the memory and configured to be executed by one or more processors as described above for identifying a fraudulent enterprise.

[0048] A method for identifying fraudulent enterprises based on the above embodiments of the present invention includes: when it is necessary to identify a target enterprise, determining the enterprise information corresponding to the target enterprise, including basic enterprise data, transaction flow data, channel login data, business registration data, and credit data; determining scenario feature data and feature data corresponding to each preset feature dimension in the enterprise information; inputting the feature data corresponding to each preset feature dimension into a pre-constructed scoring model, and obtaining a risk score corresponding to the target enterprise after processing by the scoring model; inputting the scenario feature data and risk score into a pre-constructed scenario interpretation model, and obtaining a scenario mapping rule corresponding to the risk score after processing by the scenario interpretation model; determining the identification result corresponding to the target enterprise based on the risk score and the scenario mapping rule; if the identification result indicates that the target enterprise is a fraudulent enterprise, then determining the fraud scenario corresponding to the target enterprise, and completing the identification process of the target enterprise. The method provided in this invention can combine multi-dimensional characteristic data of an enterprise to conduct a risk assessment on whether the enterprise is a fraudulent enterprise through a pre-built scoring model. Furthermore, a pre-built scenario interpretation model can be used to map the risk score into rules, resulting in scenario mapping rules. Based on the risk score and scenario mapping rules, it can be identified whether the enterprise is likely to be a fraudulent enterprise. If it is a fraudulent enterprise, the potential fraud scenarios it may be involved in can then be identified. During the identification process, risk characteristics can be mined from multi-dimensional data, which helps improve the accuracy of identification. Secondly, it enables automated identification of fraud scenarios without relying on manual processing, saving human resources and avoiding human error. Attached Figure Description

[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0050] Figure 1 A flowchart illustrating a method for identifying a front company, as provided in an embodiment of the present invention;

[0051] Figure 2 Another flowchart of a method for identifying a front company provided in an embodiment of the present invention;

[0052] Figure 3 Another flowchart of a method for identifying a front company provided in an embodiment of the present invention;

[0053] Figure 4 This is a schematic diagram of the structure of a decoy company identification device provided in an embodiment of the present invention;

[0054] Figure 5This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0055] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0056] In this application, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0057] As the background technology shows, the existing process of identifying shell companies usually relies on business registration data for risk identification. For example, machine learning models are built using publicly available business registration data to identify shell companies. However, in real life, the risk characteristics of shell companies are not limited to business registration data. Identification based solely on business registration data has a low accuracy rate, and the business scenario interpretability of the existing identification process is insufficient.

[0058] Therefore, this invention provides a method for identifying covert enterprises, which combines multi-dimensional data to identify covert enterprises and further identifies scene mapping rules, thereby improving the accuracy of identification and the business interpretability of the identification results.

[0059] This invention provides a method for identifying fraudulent enterprises. This method can be applied to a risk identification system, and its execution entity can be the system's server. The method flowchart is shown below. Figure 1 As shown, it includes:

[0060] S101: When it is necessary to identify a target enterprise, determine the enterprise information corresponding to the target enterprise, the enterprise information including basic enterprise data, transaction flow data, channel login data, business registration data and credit data;

[0061] In the method provided by this invention, when a user needs to identify whether a certain enterprise may be a front-end enterprise, they can input the key identifier of the target enterprise (i.e., the enterprise to be identified) through the system front-end and send an instruction to the system to identify the target enterprise. When the system receives the instruction to identify the target enterprise, it can obtain multi-dimensional data corresponding to the target enterprise from the database based on the key identifier of the target enterprise, including basic enterprise data, transaction data, channel login data, business registration data, and credit data, to obtain the enterprise information corresponding to the target enterprise.

[0062] It should be noted that, in the specific implementation process, enterprise information may also include data from other dimensions.

[0063] S102: In the enterprise information, determine the scene feature data and the feature data corresponding to each preset feature dimension;

[0064] In the method provided by this invention, scenario feature data attributes and data attributes corresponding to each preset feature dimension can be pre-set. Data corresponding to the scenario feature data attributes is obtained from enterprise information and used as scenario feature data. Data corresponding to the data attributes for each preset feature dimension is obtained from enterprise information and used as feature data for each preset feature dimension. Scenario feature data attributes refer to data attributes that are highly interpretable for risk scenarios. For example, transaction data provides good interpretability for money laundering risk scenarios, and bill business data provides good interpretability for corporate bill fraud risk scenarios. Each preset feature dimension is a data dimension referenced for identifying the risk level of an enterprise as a shell company, such as basic information dimension, enterprise cash flow dimension, etc.

[0065] S103: Input the feature data corresponding to each of the preset feature dimensions into the constructed scoring model, and after processing by the scoring model, obtain the risk score corresponding to the target enterprise;

[0066] In the method provided by the embodiments of the present invention, a scoring model can be pre-built. This scoring model is a model built based on machine learning algorithms. It can mine the risk level of an enterprise being a shell company through data of various preset feature dimensions and output a risk score.

[0067] In the method provided by this invention, feature data corresponding to each preset feature dimension is loaded into the input layer of the scoring model. The scoring model can process the input multi-dimensional feature data. After the scoring model completes the processing, the risk score corresponding to the target enterprise can be obtained from the output layer of the scoring model.

[0068] S104: Input the scene feature data and the risk score into the constructed scene interpretation model, and after processing by the scene interpretation model, obtain the scene mapping rule corresponding to the risk score;

[0069] In the method provided by the embodiments of the present invention, a scenario interpretation model can be pre-built. This scenario interpretation model is also a model built based on machine learning algorithms. It can mine scenario mapping rules for interpreting the risk score of an enterprise through scenario feature data and risk scores. That is, it explains the identification results indicated by the risk score (whether it belongs to a fraudulent enterprise) and the rules that may be involved in fraud scenarios.

[0070] In the method provided by the embodiments of the present invention, the scene feature data and risk score corresponding to the target enterprise can be loaded into the input layer of the scene interpretation model. After processing by the scene interpretation model, the scene mapping rules corresponding to the risk score can be obtained from its output layer.

[0071] It should be noted that in the specific implementation process, the scoring model and the scenario interpretation model can be integrated into a single model. After inputting the relevant data into the overall model, the risk score and the corresponding scenario mapping rules can be obtained directly.

[0072] S105: Based on the risk score and the scenario mapping rules, determine the identification result corresponding to the target enterprise;

[0073] In the method provided by this invention, the risk score can be interpreted according to the rules of the scene mapping rules to obtain the identification result corresponding to the target enterprise. For example, the scene mapping rules state that when the risk score is greater than or equal to a first preset threshold, the identification object is more likely to be a decoy enterprise, and when the risk score is less than the first preset threshold, the identification object is less likely to be a decoy enterprise. Based on the comparison result between the current risk score and the first preset threshold, the identification result corresponding to the target enterprise is determined.

[0074] S106: If the identification result indicates that the target company is a front company, then the fraud scenario corresponding to the target company is determined, and the identification process of the target company is completed.

[0075] In the method provided by this invention, the scenario mapping rules include potential fraud scenarios (if the target company is a front company). If the identification result indicates that the target company is a front company, meaning the risk level of a front company is high, then the fraud scenario corresponding to the target company can be obtained from the scenario mapping rules. A fraud scenario refers to a risky scenario where criminals use a front company to commit illegal or fraudulent acts, such as the following scenarios:

[0076] Money laundering: Underground banks use shell companies to open multiple bank accounts to transfer funds for money laundering.

[0077] Retail loan fraud and credit card fraud: Using sham companies to "package" borrowers as company executives, and forging social security, housing provident fund, bank statements and other documents to apply for personal loans and large credit cards from banks;

[0078] Corporate inclusive finance loan fraud: Borrowers purchase shell companies and forge transaction contracts and other false materials to fraudulently obtain loans; they manipulate shell companies to provide guarantees for loan companies to fraudulently obtain loans.

[0079] Corporate bill fraud: Bill brokers register a large number of sham companies to apply for bill discounting at banks and fraudulently obtain bank acceptance bills based on false trade backgrounds.

[0080] If the identification results indicate that the target company is not a front company, the identification process for the target company can be terminated directly.

[0081] Based on the method provided in this invention, when it is necessary to identify a target enterprise, the enterprise information corresponding to the target enterprise is determined, including basic enterprise data, transaction flow data, channel login data, business registration data, and credit data. Within the enterprise information, scenario feature data and feature data corresponding to each preset feature dimension are determined. The feature data corresponding to each preset feature dimension are input into a pre-constructed scoring model. After processing by the scoring model, a risk score corresponding to the target enterprise is obtained. The scenario feature data and risk score are input into a pre-constructed scenario interpretation model. After processing by the scenario interpretation model, a scenario mapping rule corresponding to the risk score is obtained. Based on the risk score and scenario mapping rule, the identification result corresponding to the target enterprise is determined. If the identification result indicates that the target enterprise is a fraudulent enterprise, the fraud scenario corresponding to the target enterprise is determined, thus completing the identification process of the target enterprise. The method provided in this invention can combine multi-dimensional characteristic data of an enterprise to conduct a risk assessment on whether the enterprise is a fraudulent enterprise through a pre-built scoring model. Furthermore, a pre-built scenario interpretation model can be used to map the risk score into rules, resulting in scenario mapping rules. Based on the risk score and scenario mapping rules, it can be identified whether the enterprise is likely to be a fraudulent enterprise. If it is a fraudulent enterprise, the potential fraud scenarios it may be involved in can then be identified. During the identification process, risk characteristics can be mined from multi-dimensional data, which helps improve the accuracy of identification. Secondly, it enables automated identification of fraud scenarios without relying on manual processing, saving human resources and avoiding human error.

[0082] exist Figure 1 Based on the method shown, refer to Figure 2 The flowchart shown illustrates the construction process of the scoring model mentioned in step S103 of the method provided in this embodiment of the invention, which includes:

[0083] S201: Determine the initial sample set corresponding to each preset feature dimension; the initial sample set corresponding to each preset feature dimension includes multiple sample data corresponding to that preset feature dimension;

[0084] In the method provided by this invention, sample data corresponding to each preset feature dimension can be pre-configured to obtain an initial sample set. The sample data includes sample input and sample output. The sample input is the data of the preset feature dimension of the sample enterprise, and the sample output is the risk level of the sample enterprise being a shell company.

[0085] S202: For each preset feature dimension, determine the training sample set corresponding to the preset feature dimension based on the initial sample set corresponding to the preset feature dimension;

[0086] In the method provided by this invention, data processing is performed on the initial sample set corresponding to each preset feature dimension according to a pre-configured data processing strategy to obtain the training sample set corresponding to each preset feature dimension. Specific data processing operations may include data sampling, bad sample diffusion, variable derivation, etc.

[0087] It should be noted that in the specific implementation process, the data processing operations performed on the initial sample sets corresponding to each preset feature dimension can be different. The specific processing operation can be selected according to the actual needs, with the processing effect as the standard.

[0088] S203: For each preset feature dimension, construct a sub-model corresponding to the preset feature dimension based on the training sample set corresponding to the preset feature dimension;

[0089] In the method provided by the embodiments of the present invention, a sub-model corresponding to each preset feature dimension can be constructed based on the training sample set corresponding to each preset feature dimension. The sub-model is a model constructed based on machine learning algorithms and can be used to predict the risk level of an enterprise being a shell company based on the data of the corresponding preset feature dimension.

[0090] S204: Perform fusion processing on each of the sub-models to obtain a fusion model, and use the fusion model as the scoring model.

[0091] In the method provided by the embodiments of the present invention, each sub-model can be fused according to a preset model fusion strategy, and the fused model can be used as a scoring model.

[0092] Based on the methods provided in the above embodiments, refer to Figure 3 The flowchart shown illustrates the process in step S202 of the method provided by this embodiment of the invention, which involves determining the training sample set corresponding to the preset feature dimension based on the initial sample set corresponding to the preset feature dimension. This process includes:

[0093] S301: Based on the preset oversampling strategy, the initial sample set corresponding to the preset feature dimension is oversampled to obtain the first sample set;

[0094] In the method provided by this invention, considering that the purity of black samples in the model is often on the order of one in a thousand or even one in ten thousand, the class imbalance problem is particularly prominent. Therefore, a sampling method is used to process the initial sample set. By changing the original imbalanced sample set, a balanced sample distribution is obtained, thereby training a better model. Sampling methods can be broadly divided into oversampling and undersampling. The method provided by this invention first processes the initial sample set using an oversampling method, and uses the processed sample set as the first sample set. The oversampling strategy can adopt SMOTE (Synthetic Minority Oversampling Technique).

[0095] The core idea of ​​SMOTE is to generate additional samples by interpolating between minority class samples. Specifically, for a minority class sample x... i Using the K-nearest neighbor method (the value of k needs to be specified in advance), calculate the distance from x. i The k nearest minority class samples are identified, where the distance is defined as the Euclidean distance between samples in the n-dimensional feature space. Then, a new sample is generated by randomly selecting one of these k nearest neighbors using the following formula:

[0096]

[0097] in, For the selected k nearest neighbors, δ∈[0,1] is a random number. The samples generated by SMOTE are generally found in... and x i On the connected straight lines, x new This is a new sample.

[0098] SMOTE randomly selects minority class samples to synthesize new samples without considering the surrounding samples. This can lead to two problems: First, if the selected minority class sample is surrounded by minority class samples, the newly synthesized sample will not provide much useful information. Second, if the selected minority class sample is surrounded by majority class samples, which may be noise, the newly synthesized sample will have a large overlap with the surrounding majority class samples, making classification difficult.

[0099] S302: According to the preset data cleaning strategy, the first sample set is cleaned to obtain the second sample set corresponding to the first sample set;

[0100] In the method provided by this invention, a preset data cleaning strategy can be used to process the data of a first sample set, and the processed sample set can be used as a second sample set. The data cleaning strategy mainly uses certain rules to clean overlapping data, thereby achieving the purpose of undersampling. These rules are often heuristic, such as Tomek Link and ENN.

[0101] Tomek links represent the closest pair of samples between different classes; that is, the two samples are each other's nearest neighbors and belong to different classes. If two samples form a Tomek link, either one of them is noise, or both samples are near the boundary. Removing Tomek links can "clean up" overlapping samples between classes, ensuring that the nearest neighbors all belong to the same class, thus improving classification accuracy.

[0102] ENN (Edited Nearest Neighbors) removes a sample if more than half of its K nearest neighbors do not belong to the majority class. A variation of this method removes a sample if none of its K nearest neighbors belong to the majority class.

[0103] The method provided in this embodiment of the invention adopts a method of oversampling followed by data cleaning, namely SMOTE+ENN or SMOTE+Tomek, which uses data cleaning technology to remove overlapping samples, thus overcoming the disadvantage of the SMOTE algorithm that the generated minority class samples are easy to overlap with the surrounding majority class samples and are difficult to classify.

[0104] S303: Based on the preset community detection algorithm, perform bad sample diffusion processing on the second sample set to obtain the third sample set corresponding to the second sample set;

[0105] In the method provided by this invention embodiment, a community detection algorithm, namely the Louvain graph clustering algorithm, is used to perform bad sample diffusion on the second sample set in order to find a subset of nodes similar to known black samples and group them together. The sample set that has undergone bad sample diffusion is used as the third sample set.

[0106] The main processing steps of the Louvain algorithm include: treating each node in the graph as a community, attempting to have a node join a neighboring community, calculating the modularity exponential increment ΔQ of the graph, and finally selecting a neighboring community with the largest ΔQ to join; treating the communities partitioned in the previous step as supernodes and calculating the relevant feature values ​​of the supernodes; when the algorithm has reached its objective (e.g., the maximum increment ΔQ is less than a certain value), the algorithm terminates and outputs the result. Otherwise, the supernode is treated as a regular node, and the algorithm returns to the first step; the result obtained by the Louvain algorithm is a graph composed of nodes partitioned into communities.

[0107] For example, in the method provided in this embodiment of the invention, an organizational structure relationship diagram between enterprises is established based on business registration data. The organizational structure specifically refers to the legal representative, directors, supervisors, senior executives, and shareholders of micro and small enterprises. First, a relationship diagram of the organizational structure between micro and small enterprises is constructed, with each sample debt item as a node. When a node of a sample debt item corresponds to a person in the organizational structure of a micro and small enterprise that is the same person in the organizational structure of another debt item (e.g., the legal representative of one company is a shareholder of another company), then an edge can be formed between these two nodes. The more overlapping nodes, the higher the weight of this edge. The graph relationships are divided into communities based on the Louvain community detection algorithm. After dividing the communities, the proportion of bad samples in each community is calculated. In communities where the proportion of bad samples is >= 20%, other good samples are diffused into bad samples.

[0108] S304: Based on a preset variable derivation strategy, perform variable derivation processing on the third sample set to obtain a fourth sample set corresponding to the third sample set;

[0109] In the method provided by this invention, to efficiently utilize data from various dimensions, the sample set can be processed by variable derivation according to a preset variable derivation strategy, and the sample set after variable derivation can be used as a fourth sample set. For example, NLP (Natural Language Processing) technology can be used to vectorize text for categorical variables such as business scope; anomaly detection technology can be used to capture abnormal data in continuous variables; and consistency cross-validation can be performed using cross-validation between multi-dimensional data. Ultimately, multi-dimensional variables such as basic variables, statistical variables, consistency cross-validation variables, time series variables, and anomaly detection variables are formed.

[0110] S305: Based on the preset variable selection strategy, perform variable selection processing on the fourth sample set to obtain the fifth sample set corresponding to the fourth sample set, and use the fifth sample set as the training sample set corresponding to the preset feature dimension.

[0111] In the method provided by the embodiments of the present invention, variables in the sample set can be filtered through a preset variable filtering strategy, and the sample set after variable filtering can be used as the final training sample set for model training.

[0112] Specific variable selection strategies can employ filtering, embedding, integration, or other variable selection methods to comprehensively select variables for inclusion in the model. The specific strategies and overviews of various variable selection methods are shown in the table below:

[0113] Table 1

[0114]

[0115]

[0116] To better illustrate the method provided in the embodiments of the present invention, the following section will further explain the various variable screening strategies in conjunction with the table above.

[0117] The main points regarding filtration methods are as follows:

[0118] Each feature is scored based on its divergence or correlation, and a threshold or the number of features to be selected is set. The filtering method is independent of the machine learning classification algorithm, which is not conducive to optimizing classification performance. However, this method uses a simple singular value decomposition method to reduce dimensionality, resulting in high computational efficiency and clear results.

[0119] IV (Information Value): The purpose of IV is to measure the overall predictive power of a variable. The advantage is that the IV values ​​of each variable are comparable. The IV value indicates the information contribution of a variable in determining the category to which a sample belongs; the greater the contribution, the larger the IV value. During initial variable screening, variables with IV values ​​less than 0.02 can be directly removed and not included in the subsequent algorithm fitting process.

[0120] PSI: PSI is a univariate stability measure. The specific calculation logic involves dividing the data into equal-frequency and equal-interval bins based on all values ​​of the feature in the current dimension. The proportion of people in each bin is calculated separately, resulting in the proportion of each bin in the sample set and the proportion in the out-of-time validation set. The PSI of each bin is then calculated separately, and the PSIs of all bins are summed. If the univariate PSI is greater than 0.01, the variable is considered unstable.

[0121] Mutual information: Mutual information is used to evaluate the amount of information that the occurrence of one event contributes to the occurrence of another event. For events A and B occurring simultaneously, mutual information is a method described in information theory, and it is calculated as follows:

[0122]

[0123] The significance of mutual information lies in the amount of information provided by the correlation between the occurrence of event A and the occurrence of event B. When extracting features for classification problems, mutual information can be used to measure the correlation between a feature and a specific category. The greater the information content, the stronger the correlation between the feature and the category, and vice versa. The range of mutual information values ​​is [0,1]. A larger value indicates a stronger correlation, while a value of 0 indicates that the feature is not related to the target.

[0124] Variance: In mathematical statistics, variance is the most important and commonly used indicator for measuring the dispersion of random variables. Variance is the average of the squared deviations of each variable value from the mean, and it is the most important method for measuring the dispersion of numerical data. When the data distribution is relatively concentrated, the sum of the squared differences between each data point and the mean is small. When the data distribution is relatively dispersed, i.e., the data fluctuates significantly around the mean, the variance is large. Therefore, the larger the variance, the greater the data fluctuation; the smaller the variance, the smaller the data fluctuation. Thus, it is necessary to prioritize eliminating features with zero or small variance.

[0125] For explanations of univariate AUC, KS_2SAMP, and Rank-sum tests, please refer to Table 1.

[0126] The main points regarding the embedding method are as follows:

[0127] Embedded Method: This method uses a machine learning model to train and obtain the weight coefficients of each feature. Features are then selected based on the size of the coefficients.

[0128] Table 1 explains the importance of Lasso regression, LightGBM model, and zero importance.

[0129] The main points regarding the integration method are as follows:

[0130] Wrapper Method: Select several features each time and retain or remove them based on the result of the objective function.

[0131] For an explanation of the genetic algorithm and recursive feature elimination, please refer to Table 1.

[0132] The main points regarding other methods are as follows:

[0133] Univariate Head-of-the-Line Fraud Rate Screening Method: Since anti-fraud focuses on the number of fraudulent samples hit by the debt items with the highest predicted scores, this embodiment of the invention proposes a univariate head-of-the-line fraud rate evaluation metric. A univariate learner is constructed, which can be LightGBM, XGBoost, etc. On the test set, the number of customers with the highest predicted scores by the learner is m, of which n are fraudulent. The head-of-the-line fraud rate metric n / m is used as the variable evaluation standard. At the same time, the AUC of the learner on the test set is calculated. If the AUC does not exceed 0.5 or n / m does not exceed the fraud sample concentration, then the variable is directly excluded from the model.

[0134] It should be noted that in the specific implementation process, different variable screening strategies can be adopted for each sample set that needs to be screened. The selection can be made according to the actual screening effect and will not affect the function of the method provided in the embodiments of the present invention.

[0135] exist Figure 2 Based on the aforementioned method, the process of constructing a sub-model corresponding to the preset feature dimension based on the training sample set corresponding to the preset feature dimension, mentioned in step S203 of the method provided in this embodiment of the invention, includes:

[0136] Based on a number of preset ensemble tree algorithms, construct an ensemble tree model corresponding to each ensemble tree algorithm;

[0137] In the method provided by this invention, various ensemble tree algorithms can be pre-set, such as LightGBM, GBDT, XGBoost, and Random Forest. An ensemble tree model can be constructed based on each type of ensemble tree algorithm.

[0138] For each ensemble tree model, the ensemble tree model is trained based on the training sample set corresponding to the preset feature dimension, and the ensemble tree model that has been trained is determined as a candidate model.

[0139] In the method provided by the embodiments of the present invention, each ensemble tree model is trained based on the training sample set, and the trained ensemble tree model is used as a candidate model.

[0140] For each candidate model, the candidate model is verified based on the preset reserved sample set and the cross-time sample set to obtain the verification result corresponding to the candidate model;

[0141] The method provided in this embodiment of the invention performs two verifications based on the feasibility of the data to ensure the robustness of the model. These verification methods mainly include the following:

[0142] Validation on Reserved Samples: Validation on reserved samples is part of the scoring model development process. At the modeling point, 70% of the samples are randomly selected as development samples for the scoring model, and the model results are applied to the remaining 30% of reserved samples to test the model's stability and effectiveness. The purpose of validating the model on reserved samples is to assess the predictive power of the scoring model using independent samples not used in the modeling process. If there is a significant difference in the model's predictive power between the reserved and development samples, it indicates overfitting during development, and the model will not be able to distinguish between good and bad performance in real-world applications.

[0143] Cross-time sample validation: With supporting data, the scoring model is validated across time samples. Cross-time sample validation applies the developed model to samples at different points in time to test whether the model is stable and effective.

[0144] In the method provided by the embodiments of the present invention, each candidate model can be verified based on a preset reserved sample set and a cross-time sample set to obtain the verification result of each candidate model.

[0145] Based on the verification results corresponding to each of the candidate models, a target candidate model is determined among the candidate models, and the target candidate model is determined as the sub-model corresponding to the preset feature dimension.

[0146] In the method provided by the embodiments of the present invention, the model with the best discrimination ability and stability can be selected from various candidate models as the target candidate model, and the target candidate model can be used as a sub-model corresponding to the preset feature dimension.

[0147] exist Figure 2 Based on the method shown, the process of fusing the various sub-models to obtain a fused model, mentioned in step S204 of the method provided in this embodiment of the invention, includes:

[0148] Determine the weights corresponding to each of the sub-models;

[0149] According to the weights corresponding to each sub-model, the sub-models are weighted and fused, and the fusion result is used as the fusion model.

[0150] In the method provided by this invention, the various sub-models are fused using a weighted fusion approach. Based on the validation results of each sub-model, the weight corresponding to each sub-model is determined, and sub-models with more accurate predictions are assigned higher weights to improve the prediction accuracy of the fused model. Then, the sub-models are weighted and fused together to obtain the fused model.

[0151] In the method provided by this invention, different weights are assigned to each voter (sub-model) to change their influence on the final result. Models with poor performance are given lower weights, while models with better performance are given higher weights. Since fraud detection focuses on the precision of the head samples, the assigned weight is the ratio of the number of head hits of the sub-model to the number of head hits of all available data blocks for that sample. Assuming there are n available data blocks for the sample, the predicted probability Prob after fusion is:

[0152]

[0153] It should be noted that the fusion method provided in the embodiments of the present invention is only for better illustrating the specific embodiments provided by the embodiments of the present invention. In the specific implementation process, other fusion methods can also be used for model fusion, which will not affect the implementation function of the method provided by the embodiments of the present invention.

[0154] exist Figure 1 Based on the method shown, the construction process of the scene interpretation model mentioned in step S104 of the method provided in this embodiment of the invention includes:

[0155] Based on a pre-defined set of scene mapping rules and a pre-defined gradient boosting decision tree algorithm, a gradient boosting decision tree model is constructed.

[0156] In the method provided by this invention, a Gradient Boosting Decision Tree (GBDT) algorithm is used for modeling, and the paths traversed by all leaf nodes of the GBDT model are used as the rule set for scene mapping. Therefore, an initial GBDT model is constructed based on the preset scene mapping rule set and the GBDT algorithm.

[0157] Based on the preset scenario explanatory variables and bad sample set, the gradient boosting decision tree model is trained to obtain the trained gradient boosting decision tree model.

[0158] In the method provided by this embodiment of the invention, variables that are good at explaining the decoy enterprise can be set as scenario explanatory variables, scenario explanatory variables are taken as X, and real bad samples are taken as Y, and the GBDT model is trained.

[0159] Determine whether the trained gradient boosting decision tree model meets the preset test and verification conditions. If the trained gradient boosting decision tree model meets the test and verification conditions, then the trained gradient boosting decision tree model is determined as the scenario explanation model.

[0160] In the method provided by the embodiments of the present invention, the trained GBDT model is tested and verified, that is, it is determined whether it meets the preset test and verification conditions. If it does not meet the conditions, the model parameters are adjusted and training continues until the trained GBDT model meets the test and verification conditions, and then it is used as the scene interpretation model.

[0161] exist Figure 1 Based on the method shown, the preset feature dimensions provided in the embodiments of the present invention include: basic information dimension, enterprise cash flow dimension, actual controller cash flow dimension, business and enterprise credit dimension, and actual controller credit dimension.

[0162] In the method provided by this invention, each preset feature dimension includes a basic information dimension, a corporate transaction volume dimension, a controlling shareholder transaction volume dimension, a business registration and corporate credit dimension, and a controlling shareholder credit dimension. Correspondingly, the scoring model can be formed by fusing the sub-models corresponding to each preset feature dimension. Therefore, the sub-models in the scoring model include: a basic information consistency sub-model, a corporate transaction volume sub-model, a controlling shareholder transaction volume sub-model, a business registration and corporate credit sub-model, and a controlling shareholder credit sub-model.

[0163] To better illustrate the method provided in the embodiments of the present invention, the model development process in the method provided in the embodiments of the present invention will be briefly described below in conjunction with actual application scenarios.

[0164] The development process of decision-making models based on heuristic machine learning mainly includes:

[0165] Determine key definitions: Define the time window for the sample and the definition of good or bad. The definition of good or bad is often based on existing lists or rules.

[0166] Data processing: Divide the modeling samples into training and test sets, and perform feature engineering and derivation on the variables;

[0167] Variable selection: Scoring models often have thousands of candidate variables. To improve model development efficiency, pre-screening of these variables is typically performed. The best practice for logistic regression scorecard models is to consider the stability and discriminative power of variables during pre-screening, and to incorporate business experience to complete the variable selection.

[0168] Model development: Constructing a supervised machine learning model using the selected variables. Commonly used algorithms include GBDT, XGBoost, and Random Forest. By training hundreds or thousands of decision trees, the results of the constructed decision trees are integrated (weighted, voted, etc.) to output the final classification probability value, and the model probability value is mapped to the model score using a linear or non-linear formula.

[0169] Model evaluation and application: Evaluate the model's results on out-of-time samples to ensure its robustness across different time windows and customer groups; observe the distribution of customer numbers in each score range and whether fraudulent behavior is involved to assess the model's ranking ability.

[0170] The method provided in this embodiment of the invention mainly includes the following model development process:

[0171] Model subdivision stage;

[0172] Front-up companies may exhibit different risk characteristics and trends across various data types. This invention's embodiments subdivide the model based on data block sources, categorizing data blocks into basic information, business registration, credit information, and transaction records.

[0173] Sampling process;

[0174] Please refer to the previous text based on Figure 3 In the provided embodiments, the descriptions of steps S301 and S302 will not be repeated here.

[0175] The diffusion process of bad samples;

[0176] Please refer to the previous text based on Figure 3 In the provided embodiments, the description of step S303 will not be repeated here.

[0177] Variable derivation process;

[0178] Please refer to the previous text based on Figure 3 In the provided embodiments, the description of step S304 will not be repeated here.

[0179] Variable selection process;

[0180] Please refer to the previous text based on Figure 3 In the provided embodiments, the description of step S305 will not be repeated here.

[0181] Imbalanced sample modeling stage;

[0182] LightGBM provides interfaces for custom objective and evaluation functions. The objective function is used to optimize the training data, while the evaluation function is used to assess the performance of the trained model on the validation set, primarily for optimizing hyperparameters.

[0183] To address data with excessively low bad sample concentrations, this invention employs a custom objective function to asymmetrically penalize incorrect predictions during modeling iterations. This means assigning a larger penalty to misclassified bad samples, thereby improving the capture rate of top-ranked bad samples. Specifically, the training loss is optimized during model training. Custom objective functions are highly effective for gradient-based Boosting models but not for Bagging models like Random Forests. Therefore, they can be used on the LightGBM model based on GBDT.

[0184] The objective function used in this embodiment of the invention can be selected from the following 6 objective functions. The objective function with the highest improvement in the application process of this embodiment of the invention is Cost-sensitive Logloss. In the actual training process, the optimal objective function can be selected to train the model.

[0185] The six main objective functions used include:

[0186] ① Interval-weighted loss function:

[0187]

[0188] ②Cost-sensitive Logloss:

[0189]

[0190]

[0191]

[0192] ③Fair Loss:

[0193]

[0194]

[0195]

[0196] ④Log Cosh Loss (Regression Loss Function):

[0197]

[0198] g i (x)=log(e -x +e x )

[0199]

[0200] ⑤Pseudo Huber Loss (an approximation of the Huber loss function):

[0201]

[0202]

[0203]

[0204] ⑥Focal Loss:

[0205]

[0206] The loss functions mentioned above are existing loss functions. This embodiment of the invention only provides illustrative representations of each loss function and will not be described in detail here.

[0207] Sub-model development phase;

[0208] To a large extent, model development is an interactive process that requires repeatedly following the steps until satisfactory results are obtained. Because different sub-models have varying data quality and variable distributions, a uniform modeling process and methodology may not yield optimal results. Therefore, based on the above process, the five data block sub-models use different data block variables, the same bad samples, and different methods for optimal modeling.

[0209] Each sub-model utilizes various machine learning algorithms for development; the modeling algorithm employs the LightGBM ensemble tree model, with local optimization using GridSearchCV, and parameter tuning incorporating expert experience and data feedback during the modeling process. The specific workflow involves using a small sample set (sampling samples) offline to iteratively compare and obtain model parameters using OPTUNA.lightGBM.

[0210] Model Result Evaluation: AUC, KS, and TOP N were used to evaluate the model's discriminative power, and PSI was used to evaluate its stability. Through multiple attempts at variable selection, modeling, and optimization, the model with the best discriminative power and stability was finally selected from the following sub-models: basic information consistency, enterprise cash flow, actual controller cash flow, business registration + enterprise credit, and actual controller credit. This model will serve as the basis for the next step of fusion.

[0211] Sub-model validation phase;

[0212] The description of "verifying the candidate model based on the preset reserved sample set and the cross-time sample set" in the previous embodiment for step S203 can be referred to, and will not be repeated here.

[0213] Sub-model fusion stage;

[0214] Different sub-models have different expressive capabilities on different data. Multiple machine learning models can often improve the overall predictive ability. After trying to fuse models based on data missing conditions, such as linear weighting and logistic regression weighting, considering that there are many combinations of missing data blocks, the logistic regression weighted fusion method needs to consider dozens of cases and the prediction results under different logistic regression coefficients need to be normalized before they can be compared. Therefore, considering the ease of model development and implementation, this embodiment of the invention adopts a model fusion method based on the performance of sub-models.

[0215] Rating conversion;

[0216] Scoring conversion is the process of mapping the results of a scoring model to a specific bad-to-good ratio. For ease of daily business management, conversion is typically used to establish a functional relationship between the scoring results and a specific risk level. For example, using 1000 points as the basis for score conversion (the higher the score, the greater the likelihood of fraud), the bad-to-good ratio is 60:1 at 1000 points, and halved for every 30 points decrease.

[0217] Mapping ratings to fraud scenarios;

[0218] To meet the model's need to interpret business scenarios, the GBDT algorithm is used to map the model's prediction results to combinations of business scenario variables, thus completing the mapping from fraud prediction results to business scenarios. The main process is as follows:

[0219] ① Select explanatory variables for the scenario: Based on human experience, select variables with good explanatory power for the front company from the processed variables as X variables, and model the Y variables using real bad samples.

[0220] ② Set the CUTOFF point: Currently determined based on the maximum F1 score of the model in the TOP N training set.

[0221] ③ Divide the samples according to the cutoff: Divide the training set into two parts, above and below the cutoff. This is used for the next step of modeling using the GBDT algorithm.

[0222] ④ Extracting rule combinations through GBDT tree model: The modeling algorithm adopts GBDT tree model algorithm, and the path traversed by all leaf nodes of GBDT is used as the rule set for scene mapping. The rule extraction of fraud scene is based on the training set to complete the modeling, and the effect is evaluated on the test set and OOT. Specifically, the three types of parameters Max_depth, Min_samples_leaf, and N_estimators were adjusted.

[0223] ⑤ Statistical screening: By setting conditions such as coverage, coverage rate, and precision in the training set, test set, and out-of-time samples, we initially select scene mapping rules with "low touch rate and high accuracy" for training and testing.

[0224] ⑥ Business logic filtering: The variable symbols on a single combined rule path may not conform to the business interpretation. It is necessary to remove some conditions containing symbols or those that do not conform to the threshold or business experience from the path.

[0225] ⑦ Recombining Rules: Based on the screening rules in the previous step, re-screen according to step ⑤ based on the performance of training, testing, and out-of-time samples.

[0226] The method provided in this invention expands data sources and fully mines risk characteristics. It incorporates internal enterprise basic information, transaction flow data, and channel login data, as well as external data from business registration, credit reporting, and judicial litigation. Utilizing machine learning, deep learning, natural language processing, anomaly detection, and graph algorithms, it fully mines fraud characteristics. It establishes machine learning modeling steps suitable for extremely imbalanced samples. It provides detailed methods for each stage, including sample sampling, black sample diffusion, sub-model partitioning, variable selection, imbalanced sample modeling, model result evaluation, and sub-model fusion, improving the overall accuracy and precision of the model. It establishes a mapping method from model to business interpretation. By combining model scoring results with business explanatory variables, it improves the business interpretability of the model results.

[0227] and Figure 1 Corresponding to the method for identifying decoy companies shown, this embodiment of the invention also provides a device for identifying decoy companies, used for... Figure 1 The specific implementation of the method shown is illustrated in the following diagram. Figure 4 As shown, it includes:

[0228] The first determining unit 401 is used to determine the enterprise information corresponding to the target enterprise when it is necessary to identify the target enterprise. The enterprise information includes basic enterprise data, transaction flow data, channel login data, business registration data and credit data.

[0229] The second determining unit 402 is used to determine scene feature data and feature data corresponding to each preset feature dimension in the enterprise information;

[0230] The first processing unit 403 is used to input the feature data corresponding to each of the preset feature dimensions into the constructed scoring model, and after processing by the scoring model, obtain the risk score corresponding to the target enterprise.

[0231] The second processing unit 404 is used to input the scene feature data and the risk score into the constructed scene interpretation model, and after processing by the scene interpretation model, obtain the scene mapping rule corresponding to the risk score.

[0232] The third determining unit 405 is used to determine the identification result corresponding to the target enterprise based on the risk score and the scene mapping rule;

[0233] The fourth determining unit 406 is used to determine the fraud scenario corresponding to the target enterprise if the identification result indicates that the target enterprise is a front enterprise, and to complete the identification process of the target enterprise.

[0234] Based on the apparatus provided in this embodiment of the invention, when it is necessary to identify a target enterprise, the enterprise information corresponding to the target enterprise is determined, including basic enterprise data, transaction flow data, channel login data, business registration data, and credit data; within the enterprise information, scenario feature data and feature data corresponding to each preset feature dimension are determined; the feature data corresponding to each preset feature dimension are input into a pre-constructed scoring model, and after processing by the scoring model, a risk score corresponding to the target enterprise is obtained; the scenario feature data and risk score are input into a pre-constructed scenario interpretation model, and after processing by the scenario interpretation model, a scenario mapping rule corresponding to the risk score is obtained; based on the risk score and the scenario mapping rule, the identification result corresponding to the target enterprise is determined; if the identification result indicates that the target enterprise is a fraudulent enterprise, the fraud scenario corresponding to the target enterprise is determined, and the identification process of the target enterprise is completed. The apparatus provided in this invention can combine multi-dimensional characteristic data of an enterprise to conduct a risk assessment on whether the enterprise is a fraudulent enterprise through a pre-built scoring model. Furthermore, a pre-built scenario interpretation model can be used to map the risk score into rules, resulting in scenario mapping rules. Based on the risk score and scenario mapping rules, it can be identified whether the enterprise is likely to be a fraudulent enterprise. If it is a fraudulent enterprise, the potential fraud scenarios it may be involved in can then be identified. During the identification process, risk characteristics can be mined from multi-dimensional data, which helps improve the accuracy of identification. Secondly, it enables automated identification of fraud scenarios without relying on manual processing, saving human resources and avoiding human error.

[0235] exist Figure 4 Based on the device shown, the device provided in this embodiment of the invention can be further extended to include multiple units. The functions of each unit can be found in the descriptions of the various embodiments of the method for identifying decoy companies provided above, and will not be further illustrated here.

[0236] A storage medium comprising stored instructions, wherein, when the instructions are executed, the device in which the storage medium resides executes the aforementioned method for identifying fraudulent enterprises.

[0237] This invention also provides an electronic device, the structural schematic of which is shown below. Figure 5 As shown, it specifically includes a memory 501 and one or more instructions 502, wherein one or more instructions 502 are stored in the memory 501 and configured to be executed by one or more processors 503 to perform the following operations:

[0238] When it is necessary to identify a target enterprise, the enterprise information corresponding to the target enterprise is determined. The enterprise information includes basic enterprise data, transaction flow data, channel login data, business registration data, and credit data.

[0239] In the enterprise information, scene feature data and feature data corresponding to each preset feature dimension are determined;

[0240] The feature data corresponding to each of the preset feature dimensions is input into the constructed scoring model. After processing by the scoring model, the risk score corresponding to the target enterprise is obtained.

[0241] The scene feature data and the risk score are input into the constructed scene interpretation model. After processing by the scene interpretation model, the scene mapping rule corresponding to the risk score is obtained.

[0242] Based on the risk score and the scenario mapping rules, the identification result corresponding to the target enterprise is determined;

[0243] If the identification result indicates that the target company is a front company, then the fraud scenario corresponding to the target company is determined, and the identification process of the target company is completed.

[0244] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and relevant parts can be referred to the descriptions in the method embodiments. The systems and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0245] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0246] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for identifying decoy companies, characterized in that, include: When it is necessary to identify a target enterprise, the enterprise information corresponding to the target enterprise is determined. The enterprise information includes basic enterprise data, transaction flow data, channel login data, business registration data, and credit data. In the enterprise information, scene feature data and feature data corresponding to each preset feature dimension are determined; The feature data corresponding to each of the preset feature dimensions is input into the constructed scoring model. After processing by the scoring model, the risk score corresponding to the target enterprise is obtained. The scene feature data and the risk score are input into the constructed scene interpretation model. After processing by the scene interpretation model, the scene mapping rule corresponding to the risk score is obtained. The scene interpretation model is constructed based on the gradient boosting decision tree algorithm, and the path traversed by all leaf nodes of the gradient boosting decision tree is used as the set of scene mapping rules. Based on the risk score and the scenario mapping rules, the identification result corresponding to the target enterprise is determined; If the identification result indicates that the target company is a front company, then the fraud scenario corresponding to the target company is determined, and the identification process of the target company is completed.

2. The method according to claim 1, characterized in that, The process of constructing the scoring model includes: Determine the initial sample set corresponding to each preset feature dimension; the initial sample set corresponding to each preset feature dimension includes multiple sample data corresponding to that preset feature dimension. For each preset feature dimension, the training sample set corresponding to the preset feature dimension is determined based on the initial sample set corresponding to the preset feature dimension. For each preset feature dimension, a sub-model corresponding to the preset feature dimension is constructed based on the training sample set corresponding to the preset feature dimension; The sub-models are fused to obtain a fused model, which is then used as the scoring model.

3. The method according to claim 2, characterized in that, The step of determining the training sample set corresponding to the preset feature dimension based on the initial sample set corresponding to the preset feature dimension includes: Based on the preset oversampling strategy, the initial sample set corresponding to the preset feature dimension is oversampled to obtain the first sample set; Based on a preset data cleaning strategy, the first sample set is cleaned to obtain a second sample set corresponding to the first sample set. Based on the preset community detection algorithm, the second sample set is subjected to bad sample diffusion processing to obtain the third sample set corresponding to the second sample set; Based on a preset variable derivation strategy, the third sample set is subjected to variable derivation processing to obtain the fourth sample set corresponding to the third sample set; Based on a preset variable selection strategy, the fourth sample set is subjected to variable selection processing to obtain a fifth sample set corresponding to the fourth sample set, and the fifth sample set is used as the training sample set corresponding to the preset feature dimension.

4. The method according to claim 2, characterized in that, The step of constructing a sub-model corresponding to the preset feature dimension based on the training sample set corresponding to the preset feature dimension includes: Based on a number of preset ensemble tree algorithms, construct an ensemble tree model corresponding to each ensemble tree algorithm; For each ensemble tree model, the ensemble tree model is trained based on the training sample set corresponding to the preset feature dimension, and the ensemble tree model that has been trained is determined as a candidate model. For each candidate model, the candidate model is verified based on the preset reserved sample set and the cross-time sample set to obtain the verification result corresponding to the candidate model; Based on the verification results corresponding to each of the candidate models, a target candidate model is determined among the candidate models, and the target candidate model is determined as the sub-model corresponding to the preset feature dimension.

5. The method according to claim 2, characterized in that, The process of fusing the various sub-models to obtain a fused model includes: Determine the weights corresponding to each of the sub-models; According to the weights corresponding to each sub-model, the sub-models are weighted and fused, and the fusion result is used as the fusion model.

6. The method according to claim 1, characterized in that, The construction process of the scenario interpretation model includes: Based on a pre-defined set of scene mapping rules and a pre-defined gradient boosting decision tree algorithm, a gradient boosting decision tree model is constructed. Based on the preset scenario explanatory variables and bad sample set, the gradient boosting decision tree model is trained to obtain the trained gradient boosting decision tree model. Determine whether the trained gradient boosting decision tree model meets the preset test and verification conditions. If the trained gradient boosting decision tree model meets the test and verification conditions, then the trained gradient boosting decision tree model is determined as the scenario explanation model.

7. The method according to claim 1, characterized in that, The preset feature dimensions include: basic information dimension, enterprise cash flow dimension, actual controller cash flow dimension, business registration and enterprise credit dimension, and actual controller credit dimension.

8. A device for identifying a decoy company, characterized in that, include: The first determining unit is used to determine the enterprise information corresponding to the target enterprise when it is necessary to identify the target enterprise. The enterprise information includes basic enterprise data, transaction flow data, channel login data, business registration data and credit data. The second determining unit is used to determine the scene feature data and the feature data corresponding to each preset feature dimension from the enterprise information. The first processing unit is used to input the feature data corresponding to each of the preset feature dimensions into the constructed scoring model, and after processing by the scoring model, obtain the risk score corresponding to the target enterprise. The second processing unit is used to input the scene feature data and the risk score into the constructed scene interpretation model. After processing by the scene interpretation model, the scene mapping rule corresponding to the risk score is obtained. The scene interpretation model is constructed based on the gradient boosting decision tree algorithm, and the path traversed by all leaf nodes of the gradient boosting decision tree is used as the set of scene mapping rules. The third determining unit is used to determine the identification result corresponding to the target enterprise based on the risk score and the scenario mapping rule; The fourth determining unit is used to determine the fraud scenario corresponding to the target enterprise if the identification result indicates that the target enterprise is a front enterprise, and to complete the identification process of the target enterprise.

9. A storage medium, characterized in that, The storage medium includes stored instructions, wherein, when the instructions are executed, the device containing the storage medium is controlled to perform the fake enterprise identification method as described in any one of claims 1 to 7.

10. An electronic device, characterized in that, It includes a memory, and one or more instructions, wherein one or more instructions are stored in the memory and configured to be executed by one or more processors as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Enterprise risk identification monitoring method, device and equipment and storage medium

    CN109492945A