Customer risk analysis method and system, electronic equipment and storage medium
By constructing a bill transaction map and using a logistic regression model to analyze the company's external operating and transaction data, the problem of accuracy in identifying fake companies was solved, and efficient identification and risk assessment of shell companies was achieved.
Patent Information
- Application Number
- CN202510908937.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-10-03
AI Technical Summary
Existing technologies cannot effectively guarantee the accuracy of customer risk analysis, especially the identification of companies registered with fictitious or false identities, resulting in inaccurate risk assessment.
By obtaining the company's external operating data and bill transaction data, constructing bill transaction map data, using modularity algorithm to divide it, selecting sample data to train the logistic regression model, and analyzing the company's shell risk.
It achieves efficient and accurate identification of shell companies, improves the accuracy of risk analysis, and reduces the impact of falsified data.
Smart Images

Figure CN120746718A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of risk analysis, and in particular to a customer risk analysis method and system, electronic equipment, and storage medium. Background Art
[0002] Risk control and anti-fraud systems are of vital strategic importance to banks. Their role is not only reflected in risk prevention and control and compliance management, but also directly impacts a bank's operational efficiency, customer trust, reputation, and long-term competitiveness. Therefore, it is crucial that risk control and anti-fraud systems can accurately analyze customer data for risk and fraud prevention.
[0003] Currently, risk and fraud analysis for customers primarily involves analyzing bank information, including company account information and identity information, as well as basic company data, to achieve a credit score and risk assessment for the company. Subsequent transactions will also be monitored.
[0004] However, because current data analysis primarily relies on basic company data, some companies with no actual business or capital, registered under fictitious or false identities, exist primarily for illegal purposes. Lacking a genuine business background, these companies fabricate data through fabricated documents and fictitious transactions to evade analysis. Therefore, these methods cannot effectively guarantee the accuracy of analysis. Summary of the Invention
[0005] Based on the above-mentioned deficiencies of the existing technology, the present application provides a customer risk analysis method and system, electronic equipment, and storage medium to solve the problem that the existing technology cannot guarantee the accuracy of risk analysis results.
[0006] In order to achieve the above objectives, this application provides the following technical solutions:
[0007] The first aspect of the present application provides a customer risk analysis method, comprising:
[0008] Obtaining external operating data of each company and bill transaction data of each of said companies;
[0009] Using the external operating data of each of the companies and the bill transaction data of each of the companies, construct bill transaction graph data;
[0010] Based on the bill transaction graph data, the number of regions to which the counterparties of each company's node belong and the number of counterparties are counted, and the numbers are added to the attributes of the nodes. Furthermore, modularity is divided for each company in the bill transaction graph data using a modularity algorithm, and the modularity value of each company's node is extracted and added to the attributes of the nodes.
[0011] Selecting attributes of nodes of multiple companies from the bill transaction graph data as sample data; wherein the sample data includes positive sample data of attributes of nodes of multiple selected core companies and negative sample data of attributes of nodes of multiple shell companies;
[0012] Using the sample data to train a logistic regression model;
[0013] Analyzing the attributes of the nodes of each of the companies in the current bill transaction graph data using the logistic regression model to obtain shell company analysis results for each of the companies;
[0014] Based on the shell analysis results of each of the companies, all the shell companies in the current bill transaction map data are determined.
[0015] Optionally, in the above-mentioned customer risk analysis method, after determining all shell companies in the current bill transaction graph data based on the shell company analysis results of each of the companies, the method further includes:
[0016] Adding the shell index in the shell analysis results of each of the companies to the attributes of the nodes of each of the companies;
[0017] The nodes of each of the companies in the bill transaction graph data are traversed in sequence, and the nodes of each of the companies that are shell companies are filtered out according to the shell index in the attributes of the traversed nodes of the companies.
[0018] Optionally, in the above-mentioned customer risk analysis method, after sequentially traversing the nodes of each company in the bill transaction graph data and filtering the nodes of each company that are shell companies according to the shell index in the attributes of the traversed company nodes, the method further includes:
[0019] The corresponding graph algorithm is used to calculate the bill transaction graph data after filtering out the nodes of each company that is a shell company, and the web page ranking, degree centrality value and betweenness centrality score of each company are obtained and fed back.
[0020] Optionally, in the above-mentioned customer risk analysis method, the use of the external operating data of each of the companies and the bill transaction data of each of the companies to construct the bill transaction graph data includes:
[0021] Storing the acquired external operating data of each of the companies and the bill transaction data of each of the companies in a data warehouse;
[0022] Writing the external operating data of each of the companies and the bill transaction data of each of the companies in the data warehouse into the target graph calculation library according to the preset spectrum modeling specification;
[0023] Through the target graph computing library, the external operating data of each company is used to construct the node of each company, and the bill transaction data of each company is used to establish the edges between the nodes of each company to obtain the bill transaction graph data.
[0024] Optionally, in the above-mentioned customer risk analysis method, using the external operating data of each company to construct a node of each company, and using the bill transaction data of each company to establish edges between the nodes of each company, to obtain bill transaction graph data, includes:
[0025] Using the external operating data of multiple companies as node attributes, constructing nodes for each of the companies;
[0026] For each of the bill transaction data of each transaction of each of the companies, a directed edge is established between the nodes of the two companies conducting the transaction, and the bill transaction data of the transaction is used as an attribute of the directed edge.
[0027] A second aspect of the present application provides a customer risk analysis system, comprising:
[0028] A data acquisition unit, configured to acquire external operating data of each company and bill transaction data of each of the companies;
[0029] A graph construction unit, configured to construct bill transaction graph data using the external operating data of each of the companies and the bill transaction data of each of the companies;
[0030] a feature analysis unit, configured to calculate, based on the bill transaction graph data, the number of regions to which the counterparties of each company's node belong and the number of counterparties, and add these to the node's attributes; and to perform modularity classification on each company in the bill transaction graph data using a modularity algorithm, extract the modularity value of each company's node, and add this to the node's attributes;
[0031] A sample selection unit, configured to select attributes of nodes of a plurality of the companies from the bill transaction graph data as sample data; wherein the sample data includes positive sample data of attributes of nodes of the selected plurality of core companies and negative sample data of attributes of nodes of the selected plurality of shell companies;
[0032] A model training unit, configured to train a logistic regression model using the sample data;
[0033] an analysis unit, configured to analyze the attributes of the nodes of each of the companies in the current bill transaction graph data using the logistic regression model to obtain shell company analysis results for each of the companies;
[0034] A result determination unit is configured to determine all shell companies in the current bill transaction graph data based on the shell analysis results of each of the companies.
[0035] Optionally, the above-mentioned customer risk analysis system further includes:
[0036] An index adding unit, configured to add the shell index in the shell analysis results of each of the companies to the attributes of the node of each of the companies;
[0037] The node filtering unit is used to traverse the nodes of each of the companies in the bill transaction graph data in sequence, and filter the nodes of each of the companies that are shell companies according to the shell index in the attributes of the traversed nodes of the companies.
[0038] Optionally, the above-mentioned customer risk analysis system further includes:
[0039] The indicator calculation unit is used to use the corresponding graph algorithm to calculate the bill transaction graph data after filtering the nodes of each company belonging to the shell company, obtain the web page ranking, degree centrality value and betweenness centrality score of each company and feedback them.
[0040] Optionally, in the above-mentioned customer risk analysis system, the graph construction unit includes:
[0041] A data storage unit, configured to store the acquired external operating data of each of the companies and the bill transaction data of each of the companies into a data warehouse;
[0042] A writing unit, configured to write the external operating data of each of the companies and the bill transaction data of each of the companies in the data warehouse into a target graph calculation library according to a preset spectrum modeling specification;
[0043] A construction subunit is used to construct the nodes of each of the companies using the external operating data of each of the companies through the target graph computing library, and to establish edges between the nodes of each of the companies using the bill transaction data of each of the companies to obtain bill transaction graph data.
[0044] Optionally, in the above-mentioned customer risk analysis system, the construction subunit includes:
[0045] A node construction unit, configured to use the external operating data of multiple companies as node attributes to construct nodes for each of the companies;
[0046] The edge establishment unit is used to establish a directed edge between the nodes of the two companies conducting the transaction for the bill transaction data of each transaction of each company, and use the bill transaction data of the transaction as an attribute of the directed edge.
[0047] A third aspect of the present application provides an electronic device, including:
[0048] memory and processor;
[0049] Wherein, the memory is used to store programs;
[0050] The processor is used to execute the program, and when the program is executed, it is specifically used to implement the customer risk analysis method as described in any one of the above.
[0051] In a fourth aspect, the present application provides a computer storage medium for storing a computer program, which, when executed by a processor, is used to implement the customer risk analysis method as described in any one of the above.
[0052] This application provides a customer risk analysis method that obtains the external operating data of each company and the bill transaction data of each company, so that the accuracy of the analysis results can be guaranteed by analyzing data that is not easy to forge, and the basic information data of the company is no longer used. Then, the external operating data of each company and the bill transaction data of each company are used to construct bill transaction map data, so that not only the information of each company itself can be reflected through the picture, but also the relationship between each company in the transaction can be reflected, so as to facilitate analysis based on the relationship between each company. Then, based on the bill transaction map data, the number of regions to which the counterparties of each company's nodes belong and the number of counterparties are counted, and added to the attributes of the node, and the modularity of each company in the bill transaction map data is divided into modules by a modularity algorithm, and the modularity value of each company's node is extracted and added to the attributes of the node, so that it is easy to accurately analyze whether the company is a shell company by extracting these features based on the relationship between the companies. Then, the attributes of the nodes of multiple companies are selected from the bill transaction map data as sample data. The sample data includes positive sample data for the attributes of nodes of multiple selected core companies and negative sample data for the attributes of nodes of multiple shell companies. This sample data is used to train a logistic regression model. The logistic regression model is then used to analyze the attributes of the nodes of each company in the current bill transaction graph data, obtaining shell analysis results for each company. This allows for efficient and accurate company analysis using the model. Finally, based on the shell analysis results for each company, all shell companies in the current bill transaction graph data are identified, thereby implementing a method for efficiently and accurately analyzing company risks. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.
[0054] Figure 1 A flowchart of a customer risk analysis method provided in an embodiment of the present application;
[0055] Figure 2 A flowchart of a method for constructing bill transaction graph data provided in an embodiment of the present application;
[0056] Figure 3 A flowchart of a method for constructing nodes and edges between nodes provided in an embodiment of the present application;
[0057] Figure 4 A flowchart of a method for risk indicator analysis based on bill transaction graph data provided in an embodiment of the present application;
[0058] Figure 5 A schematic diagram of the architecture of a customer risk analysis system provided in an embodiment of the present application;
[0059] Figure 6 A schematic diagram of the architecture of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0060] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0061] In this application, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.
[0062] This application embodiment provides a customer risk analysis method, such as Figure 1 As shown, the following steps are included:
[0063] S101. Obtain external operating data and bill transaction data of each company.
[0064] Among them, the company's external operating data refers to the data generated externally by the company due to business interactions. Because this data is generated by operations and stored externally, it is not internal company data and cannot be falsified by outsiders at will. Therefore, this data can reflect whether the company is operating normally. Therefore, optionally, in another embodiment of the present application, the company's external operating data may include but is not limited to the company's social security payment amount, enterprise bidding information, intellectual property rights, and other data.
[0065] A company's bill transaction data is the bill data of its transactions. Bills cannot be forged at will, and transactions can reveal a company's actual operating conditions and its relationships with other companies. This allows analysis of its position and importance within the company, enabling accurate identification of companies at risk of being shell companies.
[0066] Optionally, the bill transaction data may include but is not limited to the transaction amount, bill number, bill type, etc.
[0067] S102. Utilize the external operating data of each company and the bill transaction data of each company to construct bill transaction graph data.
[0068] It should be noted that the graph can not only reflect the information of the nodes, but also reflect the relationship between the nodes. The external operating data of each company is the information of each company itself, and the bill transaction data of each company includes the transaction relationship between each company. Therefore, in the embodiment of the present application, the external operating data of each company and the bill transaction data of each company are used to construct a graph reflecting the transactions between each company, so that the transaction relationship between each company can be reflected by constructing the bill transaction graph data. Through the transaction relationship between each company, not only the actual operating situation of each company can be reflected, but also the actual meaning, status and importance of each company can be reflected, thereby facilitating more accurate identification of shell companies.
[0069] Optionally, in another embodiment of the present application, a specific implementation of step S102 is as follows: Figure 2 As shown, the following steps are included:
[0070] S201: Store the acquired external operating data of each company and the bill transaction data of each company into a data warehouse.
[0071] Specifically, after obtaining the external operating data and bill transaction data of each company, in order to facilitate subsequent reading and construct bill transaction graph data, the external operating data and bill transaction data of each company are first stored in the data warehouse Hive.
[0072] S202: Write the external operating data of each company and the bill transaction data of each company in the data warehouse into the target graph calculation library according to the preset spectrum modeling specification.
[0073] The target graph database can be Spark GraphX. Because current knowledge graph data is larger than that stored in graph databases like Neo4j, these graph databases lack support for custom algorithm functionality and exhibit varying algorithm execution performance. Therefore, in this embodiment, Spark GraphX is employed on a big data cluster to facilitate customization and subsequent high-performance big data algorithm computation.
[0074] S203. Through the target graph computing library, each company's external operating data is used to construct the node of each company, and the bill transaction data of each company is used to establish the edge between the nodes of each company to obtain the bill transaction graph data.
[0075] Specifically, we use the external operating data of each company to construct a node corresponding to that company. Then, based on the bill transaction data of each company, we connect the nodes of the companies with transactions through directed edges to obtain the bill transaction graph data.
[0076] Optionally, in another embodiment of the present application, a specific implementation of step S203 is as follows: Figure 3 As shown, the following steps are included:
[0077] S301. Utilize the external operating data of multiple companies as node attributes to construct nodes for each company.
[0078] That is, a node is created for each company, and the business data of each company is set as the attribute of the node of the company.
[0079] S302 : For each company's bill transaction data of each transaction, establish a directed edge between the nodes of the two companies that conduct the transaction, and use the bill transaction data of the transaction as an attribute of the directed edge.
[0080] Specifically, based on each company's bill transaction data, each transaction is identified. A corresponding directed edge is then created for each transaction, and the bill transaction data for that transaction is set as an attribute of the directed edge. Therefore, if two companies have multiple transactions, a corresponding directed edge will be generated for each transaction. This means that multiple directed edges exist between the two companies, rather than just one, indicating a transactional relationship. This allows for more precise representation and analysis of the specific transactions between the companies, furthermore accurately reflecting a company's operating performance and its position in the industry, leading to more accurate analysis results.
[0081] S103. Based on the bill transaction map data, the number of regions to which the counterparties of the nodes of each company belong and the number of counterparties are counted, and added to the attributes of the nodes. In addition, the modularity of each company in the bill transaction map data is divided by the modularity algorithm, and the modularity value of each company's node is extracted and added to the attributes of the nodes.
[0082] It should be noted that after obtaining the bill transaction map data, relevant features can be extracted from the transaction relationships of various companies reflected in the bill transaction map data for subsequent analysis of the risk situation of each company.
[0083] In the embodiment of the present application, the specific features analyzed are the number of regions to which the counterparties of the company's nodes belong, the number of counterparties, and the modularity value.
[0084] The number of regions to which a company's node's counterparties belong is the number of regions to which the counterparties in each transaction with the company belong. This means the number of regions to which each node connected to the company's node belongs, for example, the number of provinces to which each company belongs. Therefore, for each company, we can analyze the nodes connected to the company's node based on the bill transaction graph data, determine the regions to which these nodes belong, and finally count the number of these regions.
[0085] The number of counterparties for a company's node is the total number of transactions it has conducted with all companies, not the number of companies it has traded with. Therefore, based on the bill transaction graph data, we can calculate the number of counterparties for that company's node by counting the attributes of the directed edges connecting the company's nodes.
[0086] Modularity is a core indicator for measuring the quality of community structure within a network. Therefore, it can reflect the division of a node and its relationship with other companies. Furthermore, modular bridge nodes are more likely to be shell companies, making them a fundamental feature for identifying shell companies.
[0087] Optionally, the modularity partitioning Louvain algorithm may be used to perform modularity partitioning on all companies and extract the modularity value of each node.
[0088] In order to facilitate the subsequent analysis of the node's features, these three features are also analyzed at the same time, so this feature is added to the node's attributes.
[0089] S104. Select attributes of nodes of multiple companies from the bill transaction graph data as sample data.
[0090] Because the bill transaction graph data contains a large number of companies, in order to efficiently and accurately analyze each company and to facilitate subsequent analysis of new companies, we selected some data as sample data to train the model. Subsequently, the model can be used to directly analyze all companies in the bill transaction graph data.
[0091] Among them, the sample data includes positive sample data of the attributes of the nodes of multiple selected core companies and negative sample data of the attributes of the nodes of multiple shell companies. That is, the sample data includes positive sample data and negative sample data. Specifically, the core companies in the industry are obviously the most reliable and the companies with the lowest risks, so the nodes of multiple core companies belonging to the industry are selected and their attribute data are used as positive sample data. Optionally, the nodes of the core companies can be analyzed based on the attributes in the nodes. For example, nodes with larger attributes such as the amount of social security paid, the number of successful bids for the company, intellectual property rights, the number of regions to which they belong, and the number of counterparties are selected as the nodes of the core companies. Or, directly based on the name of the company, companies that are obviously core companies, such as state-owned enterprises, well-known companies in the industry, etc., can be selected.
[0092] Furthermore, multiple nodes belonging to shell companies are selected and their attribute data is used as negative sample data. Similarly, shell companies can be selected based on node attributes. For example, nodes with small or zero attributes such as social security contributions, corporate bids, intellectual property rights, number of regions, and number of trading counterparties can be selected as shell company nodes.
[0093] S105. Use sample data to train the logistic regression model.
[0094] Optionally, after a series of operations such as standardization of the sample data and processing of outliers, the model may be iteratively trained multiple times using a weighted logistic regression algorithm to obtain a trained logistic regression model.
[0095] S106. Analyze the attributes of the nodes of each company in the current bill transaction graph data using a logistic regression model to obtain shell analysis results for each company.
[0096] The current bill transaction map data is the latest bill transaction map data, so if it has not been updated, it is the initial bill transaction map data. If it has been updated, it is the latest bill transaction map data after the update.
[0097] Specifically, the attributes of the nodes of each company in the current bill transaction graph data are input into the logistic regression model, and the logistic regression model is used to analyze whether the company is a shell company or the probability of it being a shell company based on the attributes of the node of each company.
[0098] S107. Based on the shell analysis results of each company, all shell companies in the current bill transaction map data are determined.
[0099] Alternatively, if the shell company analysis results in whether a company is a shell company, then companies with a "yes" result can be directly identified as shell companies. If the shell company analysis results in the probability of being a shell company, then companies with a probability below a threshold can be identified as shell companies. This ultimately yields a list of shell companies, which can be labeled in the attributes of each shell company node in the bill transaction graph data.
[0100] It should be noted that the above steps only analyze high-risk shell companies. In order to analyze the risk of other companies, in another embodiment of the present application, after executing step S107, it further includes risk indicator analysis based on bill transaction map data. Figure 4 As shown, an embodiment of a method for risk indicator analysis based on bill transaction graph data includes:
[0101] S401. Add the shell index in the shell analysis results of each company to the attributes of the node of each company.
[0102] In order to quickly identify the nodes of shell companies based on the shell index, the shell index in the shell analysis results of each company is added to the attributes of the node of each company.
[0103] S402: traverse the nodes of each company in the bill transaction graph data in sequence, and filter the nodes of each company that is a shell company according to the shell index in the attributes of the traversed company nodes.
[0104] Specifically, the nodes of each company in the bill transaction graph data are traversed in turn. If the shell index in the attribute of the traversed company node is greater than a threshold, it is deleted from the graph.
[0105] S403. Use the corresponding graph algorithm to calculate the bill transaction graph data after filtering out the nodes of each company that is a shell company, obtain the web page ranking, degree centrality value and betweenness centrality score of each company and feedback them.
[0106] Specifically, we use weighted Page Rank, centrality, and betweenness centrality algorithms to calculate each company's page ranking (rank ranking), degree centrality value, and betweenness centrality score. The betweenness centrality score refers to the betweenness centrality score relative to the core enterprise.
[0107] It's important to note that an investigation of companies with high rank, degree centrality, and betweenness centrality scores revealed that these companies are typically at the upstream end of the industry chain. While their market capitalizations may not be high, their industry standing is strong, making them highly resilient to risk and characterized by relative technological leadership or resource monopoly. Furthermore, according to research, many of these companies have been recognized as specialized, specialized, and innovative "Little Giants." Therefore, the eigenvalues calculated by these algorithms have significant practical significance for banks' credit and bill operations.
[0108] The embodiment of the present application provides a customer risk analysis method, which obtains the external operating data of each company and the bill transaction data of each company, so that the accuracy of the analysis results can be guaranteed by analyzing the data that is not easy to forge, and the basic information data of the company is no longer used. Then, the external operating data of each company and the bill transaction data of each company are used to construct the bill transaction map data, so that not only the information of each company itself can be reflected through the picture, but also the relationship between each company in the transaction can be reflected, so as to facilitate analysis based on the relationship between each company. Then, based on the bill transaction map data, the number of regions to which the counterparties of the nodes of each company belong and the number of counterparties are counted, and added to the attributes of the node, and the modularity of each company in the bill transaction map data is divided into modules by a modularity algorithm, and the modularity value of the node of each company is extracted and added to the attributes of the node, so that it is easy to accurately analyze whether the company is a shell company by extracting these features based on the relationship between the companies. Then, the attributes of the nodes of multiple companies are selected from the bill transaction map data as sample data. The sample data includes positive sample data for the attributes of nodes of multiple selected core companies and negative sample data for the attributes of nodes of multiple shell companies. This sample data is used to train a logistic regression model. The logistic regression model is then used to analyze the attributes of the nodes of each company in the current bill transaction graph data, obtaining shell analysis results for each company. This allows for efficient and accurate company analysis using the model. Finally, based on the shell analysis results for each company, all shell companies in the current bill transaction graph data are identified, thereby implementing a method for efficiently and accurately analyzing company risks.
[0109] Another embodiment of the present application provides a customer risk analysis system, such as Figure 5 As shown, including:
[0110] The data acquisition unit 501 is used to acquire the external operating data of each company and the bill transaction data of each company.
[0111] The graph construction unit 502 is used to construct bill transaction graph data using the external operating data of each company and the bill transaction data of each company.
[0112] The feature analysis unit 503 is used to count the number of regions to which the counterparties of the nodes of each company belong and the number of counterparties based on the bill transaction map data, and add them to the attributes of the nodes, and to divide the modularity of each company in the bill transaction map data through a modularity algorithm, extract the modularity value of the node of each company, and add it to the attributes of the node.
[0113] The sample selection unit 504 is configured to select attributes of nodes of multiple companies from the bill transaction graph data as sample data, wherein the sample data includes positive sample data of the attributes of nodes of multiple selected core companies and negative sample data of the attributes of nodes of multiple shell companies.
[0114] The model training unit 505 is used to train the logistic regression model using sample data.
[0115] The analysis unit 506 is used to analyze the attributes of the nodes of each company in the current bill transaction graph data using a logistic regression model to obtain the shell analysis results of each company.
[0116] The result determination unit 507 is used to determine all shell companies in the current bill transaction map data based on the shell analysis results of each company.
[0117] Optionally, in another embodiment of the present application, the customer risk analysis system further includes:
[0118] The index adding unit is used to add the shell index in the shell analysis results of each company to the attributes of the node of each company.
[0119] The node filtering unit is used to traverse the nodes of each company in the bill transaction graph data in sequence, and filter the nodes of each company that is a shell company according to the shell index in the attributes of the traversed company node.
[0120] Optionally, in another embodiment of the present application, the customer risk analysis system further includes:
[0121] The indicator calculation unit is used to use the corresponding graph algorithm to calculate the bill transaction graph data after filtering the nodes of each company belonging to the shell company, obtain the web page ranking, degree centrality value and betweenness centrality score of each company and feedback them.
[0122] Optionally, in the customer risk analysis system provided in another embodiment of the present application, the graph construction unit includes:
[0123] The data storage unit is used to store the acquired external operating data of each company and the bill transaction data of each company into the data warehouse.
[0124] The writing unit is used to write the external operating data of each company and the bill transaction data of each company in the data warehouse into the target graph calculation library according to the preset spectrum modeling specifications.
[0125] The sub-unit is constructed to construct the nodes of each company using the external operating data of each company through the target graph computing library, and to establish the edges between the nodes of each company using the bill transaction data of each company to obtain the bill transaction graph data.
[0126] Optionally, in a customer risk analysis system provided in another embodiment of the present application, constructing a sub-unit includes:
[0127] The node construction unit is used to use the external operating data of multiple companies as node attributes to construct nodes for each company.
[0128] The edge establishment unit is used to establish a directed edge between the nodes of two companies conducting transactions for the bill transaction data of each transaction of each company, and use the bill transaction data of the transaction as the attribute of the directed edge.
[0129] It should be noted that the specific working process of each unit provided in the above embodiments of the present application can refer to the implementation process of the corresponding steps in the above method embodiments, and will not be repeated here.
[0130] Another embodiment of the present application provides an electronic device, such as Figure 6 As shown, including:
[0131] Memory 601 and processor 602 .
[0132] The memory 601 is used to store programs.
[0133] The processor 602 is used to execute the program stored in the memory 601. When the program is executed, it is specifically used to implement the customer risk analysis method provided by any one of the above embodiments.
[0134] Another embodiment of the present application provides a computer storage medium for storing a computer program. When the computer program is executed by a processor, it is used to implement the customer risk analysis method provided in any of the above embodiments.
[0135] Computer storage media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0136] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0137] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A customer risk analysis method, characterized in that: include: Obtaining external operating data of each company and bill transaction data of each of said companies; Using the external operating data of each of the companies and the bill transaction data of each of the companies, construct bill transaction graph data; Based on the bill transaction graph data, the number of regions to which the counterparties of each company's node belong and the number of counterparties are counted, and the numbers are added to the attributes of the nodes. Furthermore, modularity is divided for each company in the bill transaction graph data using a modularity algorithm, and the modularity value of each company's node is extracted and added to the attributes of the nodes. Selecting attributes of nodes of multiple companies from the bill transaction graph data as sample data; wherein the sample data includes positive sample data of attributes of nodes of multiple selected core companies and negative sample data of attributes of nodes of multiple shell companies; Using the sample data to train a logistic regression model; Analyzing the attributes of the nodes of each of the companies in the current bill transaction graph data using the logistic regression model to obtain shell company analysis results for each of the companies; Based on the shell analysis results of each of the companies, all the shell companies in the current bill transaction map data are determined.
2. The method according to claim 1, characterized in that After determining all the shell companies in the current bill transaction graph data based on the shell analysis results of each of the companies, the method further includes: Adding the shell index in the shell analysis results of each of the companies to the attributes of the nodes of each of the companies; The nodes of each of the companies in the bill transaction graph data are traversed in sequence, and the nodes of each of the companies that are shell companies are filtered out according to the shell index in the attributes of the traversed nodes of the companies.
3. The method according to claim 2, characterized in that After sequentially traversing the nodes of each company in the bill transaction graph data and filtering the nodes of each company that are shell companies according to the shell index in the attributes of the traversed company nodes, the method further includes: The corresponding graph algorithm is used to calculate the bill transaction graph data after filtering out the nodes of each company that is a shell company, and the web page ranking, degree centrality value and betweenness centrality score of each company are obtained and fed back.
4. The method according to claim 1, wherein The method of constructing bill transaction graph data by utilizing the external operating data of each of the companies and the bill transaction data of each of the companies includes: Storing the acquired external operating data of each of the companies and the bill transaction data of each of the companies in a data warehouse; Writing the external operating data of each company and the bill transaction data of each company in the data warehouse into the target graph calculation library according to the preset spectrum modeling specification; Through the target graph computing library, the external operating data of each company is used to construct the node of each company, and the bill transaction data of each company is used to establish the edges between the nodes of each company to obtain the bill transaction graph data.
5. The method according to claim 1, wherein The nodes of each company are constructed by using the external operating data of each company, and the edges between the nodes of each company are established by using the bill transaction data of each company to obtain bill transaction graph data, including: Using the external operating data of multiple companies as node attributes, constructing nodes for each of the companies; For each of the bill transaction data of each transaction of each of the companies, a directed edge is established between the nodes of the two companies conducting the transaction, and the bill transaction data of the transaction is used as an attribute of the directed edge.
6. A customer risk analysis system, characterized in that: include: A data acquisition unit, configured to acquire external operating data of each company and bill transaction data of each of the companies; A graph construction unit, configured to construct bill transaction graph data using the external operating data of each of the companies and the bill transaction data of each of the companies; a feature analysis unit, configured to calculate, based on the bill transaction graph data, the number of regions to which the counterparties of each company's node belong and the number of counterparties, and add these to the node's attributes; and to perform modularity classification on each company in the bill transaction graph data using a modularity algorithm, extract the modularity value of each company's node, and add this to the node's attributes; A sample selection unit, configured to select attributes of nodes of a plurality of the companies from the bill transaction graph data as sample data; wherein the sample data includes positive sample data of attributes of nodes of the selected plurality of core companies and negative sample data of attributes of nodes of the selected plurality of shell companies; A model training unit, configured to train a logistic regression model using the sample data; an analysis unit, configured to analyze the attributes of the nodes of each of the companies in the current bill transaction graph data using the logistic regression model to obtain shell company analysis results for each of the companies; A result determination unit is configured to determine all shell companies in the current bill transaction graph data based on the shell analysis results of each of the companies.
7. The system according to claim 6, characterized in that Also includes: An index adding unit, configured to add the shell index in the shell analysis results of each of the companies to the attributes of the node of each of the companies; The node filtering unit is used to traverse the nodes of each of the companies in the bill transaction graph data in sequence, and filter the nodes of each of the companies that are shell companies according to the shell index in the attributes of the traversed nodes of the companies.
8. The system according to claim 7, characterized in that Also includes: The indicator calculation unit is used to use the corresponding graph algorithm to calculate the bill transaction graph data after filtering the nodes of each company belonging to the shell company, obtain the web page ranking, degree centrality value and betweenness centrality score of each company and feedback them.
9. An electronic device, characterized in that: include: memory and processor; Wherein, the memory is used to store programs; The processor is used to execute the program, and when the program is executed, it is specifically used to implement the customer risk analysis method according to any one of claims 1 to 5.
10. A computer storage medium, characterized in that Used to store a computer program, which, when executed by a processor, is used to implement the customer risk analysis method according to any one of claims 1 to 5.