Enterprise relevance analysis method, device, computer equipment and storage medium
By obtaining the financial flows and industrial and commercial information of the enterprise, generating association tables and generating graphs and grouping, the limitations of a single data source in the existing technology are solved, and more efficient enterprise association analysis is achieved, and accuracy and coverage are improved.
Patent Information
- Application Number
- CN202411369646.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-29
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2044-09-29
AI Technical Summary
Existing enterprise correlation analysis technology relies on a single data source, is difficult to identify complex multi-level relationships, is inefficient in processing, and has shortcomings in multi-media comprehensive analysis, which affects accuracy and coverage.
The company's financial flow data and industrial and commercial information are obtained, and the company grouping is generated through graph generation functions and connected branch clustering algorithms. The noise data is filtered in combination with the preset association medium filtering algorithm to identify hidden association relationships.
It improves the accuracy and coverage of enterprise correlation analysis, can identify multi-level enterprise group relationships, eliminate media that do not have actual correlation, and provide more accurate correlation reflections between enterprises.
Smart Images

Figure CN119357710B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of graph model technology, and in particular to a method, apparatus, computer equipment, and storage medium for analyzing enterprise relevance. Background Art
[0002] Financial institutions and industrial and commercial administration departments accumulate vast amounts of enterprise data during their operations, making enterprise correlation analysis crucial for credit assessment and risk control. Existing enterprise correlation analysis technologies typically rely on a single data source (such as shareholder information or financial data). This approach typically only identifies direct correlations and struggles to uncover complex, multi-layered relationships. Furthermore, existing analysis technologies primarily rely on table field associations (such as SQL Joins), resulting in low processing efficiency for large-scale data and difficulty coping with complex correlations across heterogeneous data sources. Furthermore, existing analysis technologies also have shortcomings in the comprehensive analysis of multiple media (such as IP addresses, phone numbers, and MAC addresses), making it difficult to fully tap into the potential correlation value of this data. These factors all hinder the accuracy and coverage of enterprise correlation analysis. Summary of the Invention
[0003] The purpose of the embodiments of the present application is to propose a method, apparatus, computer device and storage medium for enterprise relevance analysis to solve the problems of low accuracy and recognition coverage of enterprise relevance analysis.
[0004] In order to solve the above technical problems, the present application provides an enterprise relevance analysis method, which adopts the following technical solutions:
[0005] Acquire joint information of each enterprise, the joint information including financial transaction data and business information of each enterprise;
[0006] Determining the industrial and commercial association relationship and the alternative media association relationship between the enterprises based on the joint information, and generating a first association table between the enterprises based on the determined industrial and commercial association relationship and the alternative media association relationship;
[0007] Filter the candidate media in the first association table according to a preset association medium screening algorithm to obtain a second association table;
[0008] Based on the second association table, grouping the enterprises using a graph generation function and a connected branch clustering algorithm to obtain enterprise grouping information;
[0009] An enterprise relevance analysis result is generated according to the enterprise grouping information.
[0010] In order to solve the above technical problems, the present application also provides an enterprise relevance analysis device, which adopts the following technical solution:
[0011] An information acquisition module is used to acquire joint information of each enterprise, wherein the joint information includes financial flow data and business information of each enterprise;
[0012] an association generating module, configured to determine the industrial and commercial association relationship and the alternative medium association relationship between the enterprises based on the joint information, and generate a first association table between the enterprises based on the determined industrial and commercial association relationship and the alternative medium association relationship;
[0013] a medium screening module, configured to screen the candidate media in the first association table according to a preset association medium screening algorithm to obtain a second association table;
[0014] An enterprise grouping module, configured to group enterprises based on the second association table using a graph generation function and a connected branch clustering algorithm to obtain enterprise grouping information;
[0015] The result generating module is used to generate enterprise relevance analysis results according to the enterprise grouping information.
[0016] In order to solve the above technical problems, the embodiment of the present application further provides a computer device, which adopts the following technical solution:
[0017] Acquire joint information of each enterprise, the joint information including financial transaction data and business information of each enterprise;
[0018] Determining the industrial and commercial association relationship and the alternative media association relationship between the enterprises based on the joint information, and generating a first association table between the enterprises based on the determined industrial and commercial association relationship and the alternative media association relationship;
[0019] Filter the candidate media in the first association table according to a preset association medium screening algorithm to obtain a second association table;
[0020] Based on the second association table, grouping the enterprises using a graph generation function and a connected branch clustering algorithm to obtain enterprise grouping information;
[0021] An enterprise relevance analysis result is generated according to the enterprise grouping information.
[0022] In order to solve the above technical problems, the embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solution:
[0023] Acquire joint information of each enterprise, the joint information including financial transaction data and business information of each enterprise;
[0024] Determining the industrial and commercial association relationship and the alternative media association relationship between the enterprises based on the joint information, and generating a first association table between the enterprises based on the determined industrial and commercial association relationship and the alternative media association relationship;
[0025] Filtering the candidate media in the first association table according to a preset association medium screening algorithm to obtain a second association table;
[0026] Based on the second association table, grouping the enterprises using a graph generation function and a connected branch clustering algorithm to obtain enterprise grouping information;
[0027] An enterprise relevance analysis result is generated according to the enterprise grouping information.
[0028] Compared with the prior art, the embodiments of the present application mainly have the following beneficial effects: obtaining joint information of each enterprise, including the financial flow data and business information of each enterprise; determining the business association relationship and alternative medium association relationship between enterprises based on the joint information, and generating a first association table between enterprises based on the determined business association relationship and alternative medium association relationship. The analysis based on multiple media can effectively identify hidden association relationships, thereby improving the coverage of the analysis; according to the preset association medium screening algorithm, the alternative media in the first association table are screened to obtain a second association table, filtering out media that do not have actual association, eliminating public media and noise data, and retaining media that are highly correlated with the actual association between enterprises, so that the second association table can more accurately reflect the actual association relationship between enterprises; based on the second association table, the enterprises are grouped by using a graph generation function and a connected branch clustering algorithm to obtain enterprise grouping information, which can not only identify direct enterprise associations, but also mine potential enterprise group relationships through a multi-level and complex network structure, thereby improving the accuracy and coverage of the final enterprise association analysis results. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the solutions in this application, a brief introduction will be given below to the drawings required for use in the description of the embodiments of this application. Obviously, the drawings described below are some embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0030] Figure 1 is an exemplary system architecture diagram to which the present application may be applied;
[0031] Figure 2 is a flow chart of an embodiment of the enterprise relevance analysis method according to the present application;
[0032] Figure 3 is a structural diagram of an embodiment of an enterprise relevance analysis device according to the present application;
[0033] Figure 4 It is a structural diagram of an embodiment of a computer device according to the present application. DETAILED DESCRIPTION
[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of the application are for the purpose of describing specific embodiments only and are not intended to limit this application. The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.
[0035] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0036] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.
[0037] like Figure 1 As shown, system architecture 100 may include a terminal device 101, a network 102, and a server 103. Terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. Network 102 is a medium for providing a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0038] The user can use the terminal device 101 to interact with the server 103 via the network 102 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0039] The terminal device 101 can be various electronic devices with a display screen and supporting web browsing. In addition to the laptop computer 1011, the tablet computer 1012 or the mobile phone 1013, the terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer and a desktop computer, etc.
[0040] The server 103 may be a server that provides various services, such as a background server that provides support for web pages displayed on the terminal device 101 .
[0041] It should be noted that the enterprise relevance analysis method provided in the embodiment of the present application is generally executed by a server, and accordingly, the enterprise relevance analysis device is generally set in the server.
[0042] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0043] Continue to refer Figure 2 , shows a flow chart of an embodiment of the enterprise relevance analysis method according to the present application. The enterprise relevance analysis method includes the following steps:
[0044] Step S201: Acquire joint information of each enterprise, where the joint information includes financial flow data and business information of each enterprise.
[0045] In this embodiment, the enterprise relevance analysis method is executed on the electronic device (eg Figure 1 The server shown in the figure) can communicate with the terminal device through a wired connection or a wireless connection. It should be noted that the above-mentioned wireless connection method may include but is not limited to 3G / 4G / 5G connection, Wi-Fi connection, Bluetooth connection, Wi MAX connection, Zigbee connection, UWB (Ultra Wide Band) connection, and other wireless connection methods currently known or to be developed in the future.
[0046] Specifically, joint information is obtained for each enterprise. This information consists of two parts: financial transaction data and business information. Financial transaction data can be obtained from the Financial Regulatory Bureau using the EAST5 financial transaction data, which reflects the transaction behavior of the enterprise. EAST5 financial transaction data refers to transaction data from financial institutions submitted in accordance with specific data reporting requirements (EAST, which stands for Examination and Analysis System Technology), with the "5" representing the version number or update stage. Business information can be obtained from the business administration system, which provides information such as the legal structure, shareholder relationships, and legal representatives of the enterprise.
[0047] Step S202 : determining the industrial and commercial association relationship and the candidate media association relationship between the enterprises based on the joint information, and generating a first association table between the enterprises based on the determined industrial and commercial association relationship and the candidate media association relationship.
[0048] Specifically, based on the joint information, the industrial and commercial relationships between enterprises (such as direct legal relationships such as shareholders and legal persons) and alternative media relationships based on associated media (such as corporate managers, corporate bank accounts, transaction IP addresses, MAC addresses, contact numbers, email addresses, etc.) are determined.
[0049] Through these two types of associations, a first association table between enterprises can be generated. The first association table contains possible association information between enterprises. Each row in the first association table records the connection between two enterprises through the industrial and commercial association relationship and the alternative medium association relationship.
[0050] Step S203 : Screening the candidate media in the first association table according to a preset association medium screening algorithm to obtain a second association table.
[0051] Specifically, the association information recorded in the first association table may contain a large amount of noise data or erroneous association information, for example, multiple enterprises share common media such as IP addresses and MAC addresses. Directly using this information may lead to misjudgment of the association between enterprises.
[0052] A preset correlation medium screening algorithm can be used to filter candidate media in the first correlation table. This algorithm screens candidate media based on multiple preset dimensions, including correlation, clustering, and outlier determination. This filter eliminates media with no real correlation, removes common media and noise data, and retains media that are highly correlated with the actual relationships between enterprises. The retained media are highly relevant and representative, more accurately reflecting the actual relationships between enterprises. This not only provides a more reliable data foundation for enterprise correlation analysis, but also reduces data size, reduces computational complexity during subsequent processing, and improves overall data processing efficiency.
[0053] Step S204 : Based on the second association table, the enterprises are grouped by using a graph generation function and a connected branch clustering algorithm to obtain enterprise grouping information.
[0054] Specifically, based on the second association table, a graph generation function (graph) can be used to generate a graph structure between enterprises. In this graph structure, nodes are enterprises, and edges are relationships between enterprises established through industrial and commercial relationships and alternative media relationships. A graph generation function (graph) is a function used in graph theory to create graph structures.
[0055] The connected components clustering algorithm (connected_components algorithm) is then used to identify connected subgraphs between enterprises, that is, each enterprise group. A connected component is a subset of the graph in which all nodes are connected by edges. All nodes in a connected component can always be connected by a path. The connected component clustering algorithm identifies all connected subgraphs in the graph and divides them into different groups. Each connected component can be considered an independent group, with nodes within it connected to each other but not connected to nodes in other groups. Different connected components represent different enterprise groups, that is, different enterprise grouping information.
[0056] In addition, you can also create a first-in-first-out queue (FIFO) through the queue function in the Queue queue generation library, and export and store the nodes and media in the generated graph data as a two-dimensional data set to adapt to the existing analysis system.
[0057] NetworkX is a Python graph library that provides a variety of functions for creating and manipulating graphs, such as the graph() function and the connected_components() function. This application can use NetworkX to build a graph model of complex relationships between enterprises and identify enterprise groups.
[0058] Step S205: generating enterprise relevance analysis results based on the enterprise grouping information.
[0059] Specifically, enterprise correlation analysis results can be generated based on enterprise grouping information. These results can be used to assess potential enterprise risks, reveal complex inter-enterprise correlation networks, and provide financial institutions with a basis for risk control of multiple credit grants and over-credit grants. This helps financial institutions more comprehensively understand the inter-enterprise correlations and thus optimize credit management and risk control processes.
[0060] In this embodiment, joint information of each enterprise is obtained, including each enterprise's financial flow data and enterprise business information; the business association relationships and alternative medium association relationships between the enterprises are determined based on the joint information, and a first association table between the enterprises is generated based on the determined business association relationships and alternative medium association relationships. Analysis based on multiple media can effectively identify hidden association relationships, thereby improving the coverage of the analysis; according to a preset association medium screening algorithm, the alternative media in the first association table are screened to obtain a second association table, filtering out media with no actual association, eliminating common media and noise data, and retaining media that are highly correlated with the actual associations between the enterprises, so that the second association table can more accurately reflect the actual associations between the enterprises; based on the second association table, enterprises are grouped using a graph generation function and a connected branch clustering algorithm to obtain enterprise grouping information. This not only identifies direct enterprise associations, but also mines potential enterprise group relationships through a multi-level, complex network structure, thereby improving the accuracy and coverage of the final enterprise association analysis results.
[0061] Furthermore, the above-mentioned step S202 may include: determining the industrial and commercial association relationship between enterprises in the joint information according to the preset industrial and commercial association rules, and generating an industrial and commercial association relationship table based on the industrial and commercial association relationship, wherein the joint primary key in the industrial and commercial association relationship table is the enterprise identification of the two enterprises with the industrial and commercial association relationship; determining the alternative media association relationship between enterprises in the joint information according to the preset various types of alternative media, and generating an alternative media association relationship table based on the alternative media association relationship, wherein the joint primary key in the alternative media association relationship table is the enterprise identification of the two enterprises with the alternative media association relationship; taking the alternative media association relationship table as the left table, left-linking the industrial and commercial association relationship table to the alternative media association relationship table through the joint primary key to obtain the first association table between enterprises.
[0062] Specifically, first, the industrial and commercial association relationships between enterprises are determined in the joint information according to the preset industrial and commercial association rules, thereby obtaining a set of available association media. The industrial and commercial association rules can be determined according to the Company Law.
[0063] Article 216 of the Company Law stipulates that "related relationships" refer to the relationship between a company's controlling shareholder, actual controller, directors, supervisors, and senior management and the enterprises they directly or indirectly control, as well as other relationships that may result in the transfer of corporate interests. Therefore, the "relationships between a company's controlling shareholder, actual controller, directors, supervisors, and senior management and the enterprises they directly or indirectly control" mentioned in the joint information can be considered confirmed related relationships and used as criteria for media mining. EAST5 and industrial and commercial data contain a total of 151 types of industrial and commercial related relationships. These include 39 detailed industrial and commercial related relationship data items, including those involving actual controllers, corporate legal persons, controlling shareholders, and headquarters enterprises, which are defined as confirmed related enterprise relationships. These include superior legal persons, corporate owners, sole investors, actual controlling investors, sole investors, absolute controlling parties, beneficiaries, actual operators, parent-subsidiary companies, and branches.
[0064] Based on the industrial and commercial relationships, a table of industrial and commercial relationships can be generated. Each row in the table represents a certain industrial and commercial relationship between two enterprises. The corporate identifiers of the two enterprises form the joint primary key for each row in the table, uniquely identifying the relationship between the enterprises. Digits 9 to 17 of the enterprise's unified social credit code can be extracted to restore the 9-digit enterprise organization code, which serves as the enterprise identifier to better organize the associated media information under the 9-digit and 18-digit enterprise identifiers. In the industrial and commercial relationship table, a value of 1 in the association field indicates an association, and a value of 0 indicates no association.
[0065] This application also pre-sets multiple types of alternative media (including discrete character data, basic information of corporate entities and information related to active media, such as IP addresses, telephone numbers, MAC addresses, etc.).
[0066] Based on multiple types of alternative media, indirect relationships between enterprises are identified in the joint information. Each alternative medium represents a possible connection between enterprises through equipment, networks, or other means. Based on the alternative media relationships, an alternative media relationship table can be generated. Each row in this table records the relationship between two enterprises established through a certain alternative medium, and the joint primary key is still the corporate identity of the two enterprises. Unlike the industrial and commercial relationship table, the alternative media relationship table captures the connection between enterprises through indirect channels such as shared equipment or networks. For example, if Company A and Company B share an IP address, the relationship between the two companies will be recorded in the alternative media relationship table as being related through "IP address."
[0067] Using the SQL Left Join operation, the alternative media association table is used as the left table, and the industrial and commercial association table is used as the right table. The two tables are combined using the joint primary key to generate the first association table. This allows even pairs of companies without industrial and commercial connections to be associated through alternative media, ensuring that all potential association information is included in the first association table, providing a complete view for subsequent analysis.
[0068] The left join operation retains the data of all alternative media association relationships, and at the same time, when there is a business association, it supplements the relevant data of the business association, so that the two associations (direct legal association and indirect association through the medium) are presented in the same table.
[0069] In this embodiment, the industrial and commercial association relationship table and the alternative medium association relationship table are merged in a left-join manner, thereby improving the coverage and accuracy of enterprise association analysis. Not only can enterprises establish associations through direct industrial and commercial information, but hidden connections can also be mined through indirect alternative media, avoiding the loss of association information; the use of joint primary keys provides a unified identification standard for the integration of different data tables, improving the efficiency and consistency of data processing; the generation of the first association table lays the data foundation for subsequent data screening and graph model analysis, thereby improving the accuracy of enterprise association analysis.
[0070] Furthermore, the above-mentioned step of screening the alternative media in the first association table according to the preset association medium screening algorithm may include: for each pair of enterprises in the first association table, calculating the association strength between the two enterprises under each alternative medium; based on the association strength between the two enterprises under each alternative medium and the industrial and commercial association relationship between the two enterprises, calculating the Pearson correlation coefficient between the industrial and commercial association of the two enterprises and the association of each alternative medium; and screening the alternative media in the first association table according to the obtained Pearson correlation coefficients.
[0071] Specifically, for each pair of companies in the first association table, their association strength is calculated for each candidate medium (e.g., IP address, phone number, MAC address, etc.). Association strength can be measured by the frequency with which two companies share a particular medium or the time span over which they use the same medium. The association strength for each candidate medium quantitatively reflects the closeness of the connection between the two companies on that particular candidate medium.
[0072] Based on the calculated association strength, the Pearson correlation coefficient is further calculated between the two companies' industrial and commercial connections and their alternative media connections. In this application, this coefficient is used to measure the linear correlation between the two companies' industrial and commercial connections and their alternative media connections. The correlation coefficient ranges from -1 to 1, with values closer to 1 indicating a stronger association, closer to 0 indicating no association, and closer to -1 indicating a negative correlation.
[0073] Calculating the Pearson correlation coefficient helps determine whether a candidate medium effectively reflects the actual business connections between companies. A candidate medium with a high correlation coefficient is considered a more valuable connection medium.
[0074] Based on the Pearson correlation coefficient of each candidate medium, candidate media with a high correlation to the industrial and commercial relationship are screened. In one embodiment, a correlation threshold can be set; only candidate media with a correlation coefficient above the threshold are retained. This screening step eliminates candidate media with low correlation, ensuring that the second correlation table only contains candidate media with high correlation, thereby improving the accuracy of the correlation analysis.
[0075] In this embodiment, the association strength between two enterprises under each alternative medium is calculated, and the Pearson correlation coefficient is used to measure the correlation between these association strengths and industrial and commercial associations, thereby screening out the alternative medium that best reflects the true relationship between the enterprises, reducing inefficient noise data, filtering out low-correlation alternative media, improving the coverage and accuracy of inter-enterprise association analysis, reducing the interference of irrelevant data on the analysis results, and improving the overall calculation efficiency, providing a more reliable data basis for subsequent enterprise grouping.
[0076] Furthermore, the step of screening the candidate media in the first association table according to the preset association medium screening algorithm may further include: obtaining aggregation information of each candidate medium in the first association table; and eliminating candidate media whose aggregation information does not meet the aggregation conditions.
[0077] Specifically, based on the first association table, an aggregation analysis is performed on each candidate medium (such as IP addresses, phone numbers, MAC addresses, etc.) to calculate the distribution of each medium across different enterprises. Aggregation information refers to the degree to which a candidate medium is shared and used by multiple enterprises. A higher aggregation indicates that the medium is used by more enterprises. This aggregation information can be used to identify which candidate media are widely used public media (such as public Wi-Fi IP addresses) and have high aggregation.
[0078] In one embodiment, cluster analysis may include ratio analysis and density cluster analysis, wherein:
[0079] Ratio analysis: This is based on the ratio of a candidate medium to multiple companies. For example, for a single candidate medium with three or more entities, all entities are counted, eliminating duplicates. The resulting ratio is then divided by the number of entities. If the ratio of a candidate medium to multiple companies exceeds a preset threshold (e.g., 0.1), it indicates a certain degree of multi-agent aggregation and should be eliminated.
[0080] Density Cluster Analysis: If the media values of candidate media exhibit a certain order, density clustering (such as DBSCAN) can be performed on the ordered data to identify the degree of media aggregation and eliminate those that may lead to false associations. For example, for IP addresses, the IP field is encoded in numerical order, normalized using extreme value normalization, and clustered based on whether the enterprise has risk data.
[0081] Based on pre-set aggregation criteria, candidate media that do not meet the aggregation criteria are screened and removed. Aggregation criteria can be set based on the characteristics of the associations between enterprises. The purpose of removal is to prevent common media from causing erroneous associations between enterprises.
[0082] In this embodiment, the aggregation of candidate media is analyzed, and aggregation conditions are set to eliminate media that do not meet the conditions, thereby improving the accuracy of enterprise association analysis; by calculating the aggregation information of each candidate medium, widely used public media can be identified to avoid erroneous associations caused by these media; after eliminating highly aggregated media, the association information between enterprises is more accurate, and media that can truly reflect the relationship between enterprises are retained, reducing the impact of noise data, optimizing data quality, and providing a more accurate data basis for processing.
[0083] Furthermore, the step of screening the alternative media in the first association table according to a preset associated medium screening algorithm may also include: calculating, for each alternative medium in the first association table, the mean and standard deviation of each enterprise on the alternative medium; calculating, based on the mean and standard deviation, the outlier threshold corresponding to the alternative medium; comparing the usage frequency of the alternative medium with the outlier threshold, and retaining or deleting the alternative medium based on the comparison result.
[0084] Specifically, for each candidate medium in the first association table, the mean and standard deviation of the candidate medium in each enterprise are calculated. The mean reflects the average use of a certain candidate medium by each enterprise, and the standard deviation measures the differences in the use of the medium by enterprises.
[0085] Based on the calculated mean and standard deviation, the system determines the outlier threshold for each candidate medium. There are two types of outlier thresholds: the first outlier threshold, which is used frequently and can be defined as "mean + 2 standard deviation". If the frequency of use of a candidate medium by each enterprise exceeds this threshold, it is considered an outlier, and the system can then determine whether the candidate medium is used too frequently among enterprises (for example, a public IP address may be shared by many enterprises) and may be a public medium. The second outlier threshold, which is used less frequently and can be defined as "mean - 2 standard deviation", is used. If the frequency of use of a candidate medium by each enterprise is lower than this threshold, it is considered an outlier, and the system can then determine whether the use of a candidate medium among enterprises is accidental (for example, accidentally logging into a certain IP address). The medium can be retained for subsequent enterprise association analysis.
[0086] In this embodiment, the mean and standard deviation of each candidate medium across enterprises are calculated, and then the outlier threshold is determined, which can effectively screen out common media with abnormally high or low frequency of use. Outlier determination is based on statistical calculations of the mean and standard deviation, which can dynamically identify media that are frequently shared by a large number of enterprises or media that are used occasionally. Comparing the usage frequency of the candidate medium with the outlier threshold can retain media that can reflect the actual relationship between enterprises, thereby improving the accuracy of the relationship analysis.
[0087] Furthermore, the above-mentioned step S204 may include: determining a seed node in the second association table based on the second association table and a preset input medium dictionary; taking the seed node as the starting point, performing node traversal based on the second association table, and generating a graph model through a graph generation function during the traversal; running a connected branch clustering algorithm on the graph model to perform enterprise grouping and obtain enterprise grouping information.
[0088] Specifically, based on the second association table and the input medium dictionary, a seed node is first determined in the second association table. The second association table contains association information between enterprises, and the input medium dictionary defines the candidate media types that need to be added to the graph model.
[0089] Starting from the determined seed node, the second association table is used to traverse the nodes. During the traversal process, a graph model is constructed through a graph generation function (such as the graph function in the network kx library). In the graph model, nodes represent enterprises and edges represent the relationships between enterprises. The specific steps can be as follows: (1): Create a graph g and define the media dictionary contained in the graph, including business information, bank cards, contact numbers, IP addresses, MAC addresses and their associated result dictionaries. (2): Define data indexes and media data sets. (3): Construct a seed node C and update the associated result dictionary. (4): Traverse other nodes, create edges of the graph model through the add_edge function, and update the associated result dictionary.
[0090] Run a connected component clustering algorithm on the generated graph model to group companies. This step uses a connected component clustering algorithm (such as the connected_components function in the network kx library) to identify connected subgraphs in the graph as distinct company groups, thereby obtaining company grouping information. In implementation, use the nx.connected_components() function to retrieve the connected subgraphs in the graph; define queues and traversal nodes, group companies, and record the graph seed and subgraph information.
[0091] In this embodiment, the seed node is used as the starting point, and various media are combined in the traversal process. The graph model can comprehensively consider the various edge media between enterprises and the parent-child node relationships in the graph model, thereby improving the coverage and comprehensiveness of the enterprise network; the connected branch clustering algorithm can efficiently identify the connected subgraphs of the enterprise, accurately group the enterprises according to the actual association relationships, and can handle complex enterprise relationship networks, avoiding information loss and misassociation in traditional methods. For example, during the graph node and edge traversal aggregation process, it is easy for two or more non-related enterprise subgraphs to be mistakenly associated and aggregated into a large graph due to certain nodes. In the above embodiment, the graph can be split and reorganized according to the graph seed node and the parent nodes at all levels to achieve the accuracy of the logic of identifying the "same person" or "group" of enterprises in each associated subgraph; the construction and grouping results of the graph model provide strong support for subsequent data analysis, can clearly reveal the relationship structure between enterprises, and help to carry out further risk assessment and other operations.
[0092] Furthermore, after the above-mentioned step S205, it may also include: obtaining the enterprise correlation analysis results corresponding to the target enterprise; determining the credit information associated with the target enterprise based on the enterprise correlation analysis results, so as to conduct risk assessment on the target enterprise based on the credit information; or, identifying the false transactions associated with the target enterprise based on the enterprise correlation analysis results; and performing feature engineering optimization on the risk assessment model based on the identified false transactions.
[0093] Specifically, the target enterprise to be evaluated is obtained, and based on the target enterprise's enterprise identification, the target enterprise's related information is extracted from the enterprise relatedness analysis results. This information includes various related relationships between the target enterprise and other enterprises.
[0094] Utilize the target company's corporate relevance analysis results, combined with its credit information, to conduct a risk assessment. This credit information can include the company's loan history and credit limits at other financial institutions, as well as the credit history of its affiliated companies. Credit information can also include the credit information of other companies associated with the target company. This allows for the identification of potential risks such as multiple credit lines and over-credit lines, allowing for a comprehensive assessment of these risks.
[0095] The results of enterprise correlation analysis can identify unusual transactions between a target enterprise and its affiliates, particularly fraudulent transactions (e.g., fabricating cash flows by circulating funds between affiliated companies to circumvent risk assessment models and obtain loans). Once these fraudulent transactions are identified, feature engineering within the risk assessment model can be optimized, incorporating the features of these fraudulent transactions into the model to improve its ability to identify similar risks.
[0096] In this embodiment, obtaining the enterprise correlation analysis results of the target enterprise can discover hidden credit risks, especially in the case of multiple credit and over-credit, and can provide accurate risk assessment; through the enterprise correlation analysis results, false transaction behaviors can also be identified, and false transaction features can be incorporated into the risk assessment model through feature engineering to improve the model's detection ability for similar behaviors, which is conducive to enhancing the risk control capabilities of financial institutions and effectively preventing potential credit risks and fraud.
[0097] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware via computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes in the above-described method embodiments. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0098] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0099] Further references Figure 3 , as a response to the above Figure 2 In order to realize the method shown in the figure, the present application provides an embodiment of an enterprise correlation analysis device. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0100] like Figure 3As shown, the enterprise relevance analysis device 300 of this embodiment includes: an information acquisition module 301, a relevance generation module 302, a medium screening module 303, an enterprise grouping module 304, and a result generation module 305, wherein:
[0101] The information acquisition module 301 is used to acquire the joint information of each enterprise, which includes the financial flow data and business information of each enterprise.
[0102] The association generating module 302 is configured to determine the industrial and commercial association relationships and the candidate media association relationships between enterprises based on the joint information, and generate a first association table between enterprises based on the determined industrial and commercial association relationships and the candidate media association relationships.
[0103] The medium screening module 303 is configured to screen the candidate media in the first association table according to a preset association medium screening algorithm to obtain a second association table.
[0104] The enterprise grouping module 304 is configured to group enterprises based on the second association table using a graph generation function and a connected component clustering algorithm to obtain enterprise grouping information.
[0105] The result generating module 305 is used to generate enterprise relevance analysis results according to the enterprise grouping information.
[0106] In this embodiment, joint information of each enterprise is obtained, including each enterprise's financial flow data and enterprise business information; the business association relationships and alternative medium association relationships between the enterprises are determined based on the joint information, and a first association table between the enterprises is generated based on the determined business association relationships and alternative medium association relationships. Analysis based on multiple media can effectively identify hidden association relationships, thereby improving the coverage of the analysis; according to a preset association medium screening algorithm, the alternative media in the first association table are screened to obtain a second association table, filtering out media with no actual association, eliminating common media and noise data, and retaining media that are highly correlated with the actual associations between the enterprises, so that the second association table can more accurately reflect the actual associations between the enterprises; based on the second association table, enterprises are grouped using a graph generation function and a connected branch clustering algorithm to obtain enterprise grouping information. This not only identifies direct enterprise associations, but also mines potential enterprise group relationships through a multi-level, complex network structure, thereby improving the accuracy and coverage of the final enterprise association analysis results.
[0107] In some optional implementations of this embodiment, the association generation module 302 may include: an industrial and commercial generation submodule, an alternative generation submodule, and an association generation submodule, wherein:
[0108] The industrial and commercial generation submodule is used to determine the industrial and commercial association relationship between enterprises in the joint information according to the preset industrial and commercial association rules, and generate an industrial and commercial association relationship table based on the industrial and commercial association relationship. The joint primary key in the industrial and commercial association relationship table is the enterprise identifier of the two enterprises with industrial and commercial association relationship.
[0109] The alternative generation submodule is used to determine the alternative media association relationship between enterprises in the joint information according to the preset various alternative media, and generate an alternative media association relationship table based on the alternative media association relationship. The joint primary key in the alternative media association relationship table is the enterprise ID of the two enterprises with the alternative media association relationship.
[0110] The association generation submodule is used to use the alternative media association relationship table as the left table, and left-link the industrial and commercial association relationship table to the alternative media association relationship table through the joint primary key to obtain the first association table between enterprises.
[0111] In this embodiment, the industrial and commercial association relationship table and the alternative medium association relationship table are merged in a left-join manner, thereby improving the coverage and accuracy of enterprise association analysis. Not only can enterprises establish associations through direct industrial and commercial information, but hidden connections can also be mined through indirect alternative media, avoiding the loss of association information; the use of joint primary keys provides a unified identification standard for the integration of different data tables, improving the efficiency and consistency of data processing; the generation of the first association table lays the data foundation for subsequent data screening and graph model analysis, thereby improving the accuracy of enterprise association analysis.
[0112] In some optional implementations of this embodiment, the medium screening module 303 may include: a strength calculation submodule, a correlation calculation submodule, and a screening submodule, wherein:
[0113] The strength calculation submodule is used to calculate the association strength between the two enterprises under each candidate medium for each pair of enterprises in the first association table.
[0114] The correlation calculation submodule is used to calculate the Pearson correlation coefficient between the industrial and commercial association of the two enterprises and the association of each alternative medium based on the association strength of the two enterprises under each alternative medium and the industrial and commercial association relationship between the two enterprises.
[0115] The screening submodule is used to screen the candidate media in the first association table according to the obtained Pearson correlation coefficients.
[0116] In this embodiment, the association strength between two enterprises under each alternative medium is calculated, and the Pearson correlation coefficient is used to measure the correlation between these association strengths and industrial and commercial associations, thereby screening out the alternative medium that best reflects the true relationship between the enterprises, reducing inefficient noise data, filtering out low-correlation alternative media, improving the coverage and accuracy of inter-enterprise association analysis, reducing the interference of irrelevant data on the analysis results, and improving the overall calculation efficiency, providing a more reliable data basis for subsequent enterprise grouping.
[0117] In some optional implementations of this embodiment, the media screening module 303 may further include: an aggregation acquisition submodule and a candidate rejection submodule, wherein:
[0118] The aggregation acquisition submodule is used to obtain the aggregation information of each candidate medium in the first association table.
[0119] The candidate elimination submodule is used to eliminate candidate media whose aggregation information does not meet the aggregation conditions.
[0120] In this embodiment, the aggregation of candidate media is analyzed, and aggregation conditions are set to eliminate media that do not meet the conditions, thereby improving the accuracy of enterprise association analysis; by calculating the aggregation information of each candidate medium, widely used public media can be identified to avoid erroneous associations caused by these media; after eliminating highly aggregated media, the association information between enterprises is more accurate, and media that can truly reflect the relationship between enterprises are retained, reducing the impact of noise data, optimizing data quality, and providing a more accurate data basis for processing.
[0121] In some optional implementations of this embodiment, the medium screening module 303 may further include: a calculation submodule, a threshold calculation submodule, and a comparison submodule, wherein:
[0122] The calculation submodule is used to calculate the mean and standard deviation of each enterprise on the alternative medium for each alternative medium in the first association table.
[0123] The threshold calculation submodule is used to calculate the outlier threshold corresponding to the candidate medium based on the mean and standard deviation.
[0124] The comparison submodule is used to compare the usage frequency of the candidate medium with the outlier threshold, and retain or delete the candidate medium according to the comparison result.
[0125] In this embodiment, the mean and standard deviation of each candidate medium across enterprises are calculated, and then the outlier threshold is determined, which can effectively screen out common media with abnormally high or low frequency of use. Outlier determination is based on statistical calculations of the mean and standard deviation, which can dynamically identify media that are frequently shared by a large number of enterprises or media that are used occasionally. Comparing the usage frequency of the candidate medium with the outlier threshold can retain media that can reflect the actual relationship between enterprises, thereby improving the accuracy of the relationship analysis.
[0126] In some optional implementations of this embodiment, the enterprise grouping module 304 may include: a node determination submodule, a traversal submodule, and an enterprise grouping submodule, wherein:
[0127] The node determination submodule is used to determine the seed node in the second association table based on the second association table and a preset input medium dictionary.
[0128] The traversal submodule is used to take the seed node as the starting point, perform node traversal based on the second association table, and generate a graph model through a graph generation function during the traversal.
[0129] The enterprise grouping submodule is used to run the connected branch clustering algorithm on the graph model to group enterprises and obtain enterprise grouping information.
[0130] In this embodiment, the seed node is used as the starting point, and various media are combined in the traversal process. The graph model can comprehensively consider various relationships between enterprises, thereby improving the coverage and comprehensiveness of the enterprise network; the connected branch clustering algorithm can efficiently identify the connected subgraphs of the enterprise, accurately group the enterprises according to the actual association relationships, and can handle complex enterprise relationship networks, avoiding information loss and misassociation in traditional methods; the construction and grouping results of the graph model provide strong support for subsequent data analysis, can clearly reveal the relationship structure between enterprises, and facilitate further risk assessment and other operations.
[0131] In some optional implementations of this embodiment, the enterprise relevance analysis device 300 may further include: a result acquisition module, a risk assessment module, a transaction identification module, and a feature optimization module, wherein:
[0132] The result acquisition module is used to obtain the enterprise relevance analysis results corresponding to the target enterprise.
[0133] The risk assessment module is used to determine the credit information associated with the target enterprise based on the results of the enterprise correlation analysis, so as to conduct risk assessment on the target enterprise based on the credit information.
[0134] The transaction identification module is used to identify false transactions associated with the target enterprise based on the results of enterprise association analysis.
[0135] The feature optimization module is used to perform feature engineering optimization on the risk assessment model based on the identified fraudulent transactions.
[0136] In this embodiment, obtaining the enterprise correlation analysis results of the target enterprise can discover hidden credit risks, especially in the case of multiple credit and over-credit, and can provide accurate risk assessment; through the enterprise correlation analysis results, false transaction behaviors can also be identified, and false transaction features can be incorporated into the risk assessment model through feature engineering to improve the model's detection ability for similar behaviors, which is conducive to enhancing the risk control capabilities of financial institutions and effectively preventing potential credit risks and fraud.
[0137] To solve the above technical problems, the present application also provides a computer device. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.
[0138] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected through a system bus. It should be noted that the figure only shows a computer device 4 having a memory 41, a processor 42, and a network interface 43, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art will understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to a microprocessor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.
[0139] The computer device may be a desktop computer, notebook computer, PDA, cloud server, etc. The computer device may interact with the user via a keyboard, mouse, remote control, touchpad, or voice control device.
[0140] The memory 41 includes at least one type of readable storage medium, including flash memory, a hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic storage, a magnetic disk, an optical disk, etc. In some embodiments, the memory 41 may be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 may also be an external storage device of the computer device 4, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash memory card, etc. equipped on the computer device 4. Of course, the memory 41 may also include both the internal storage unit of the computer device 4 and its external storage device. In this embodiment, the memory 41 is generally used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions for the enterprise relevance analysis method. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or are to be output.
[0141] In some embodiments, the processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 42 is generally used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions stored in the memory 41 or process data, such as computer-readable instructions for executing the enterprise relevance analysis method.
[0142] The network interface 43 may include a wireless network interface or a wired network interface. The network interface 43 is generally used to establish a communication connection between the computer device 4 and other electronic devices.
[0143] The computer device provided in this embodiment can execute the above-mentioned enterprise relevance analysis method, which can be the enterprise relevance analysis method of each of the above-mentioned embodiments.
[0144] In this embodiment, joint information of each enterprise is obtained, including each enterprise's financial flow data and enterprise business information; the business association relationships and alternative medium association relationships between the enterprises are determined based on the joint information, and a first association table between the enterprises is generated based on the determined business association relationships and alternative medium association relationships. Analysis based on multiple media can effectively identify hidden association relationships, thereby improving the coverage of the analysis; according to a preset association medium screening algorithm, the alternative media in the first association table are screened to obtain a second association table, filtering out media with no actual association, eliminating common media and noise data, and retaining media that are highly correlated with the actual associations between the enterprises, so that the second association table can more accurately reflect the actual associations between the enterprises; based on the second association table, enterprises are grouped using a graph generation function and a connected branch clustering algorithm to obtain enterprise grouping information. This not only identifies direct enterprise associations, but also mines potential enterprise group relationships through a multi-level, complex network structure, thereby improving the accuracy and coverage of the final enterprise association analysis results.
[0145] The present application also provides another embodiment, namely, providing a computer-readable storage medium, wherein the computer-readable storage medium stores computer-readable instructions, and the computer-readable instructions can be executed by at least one processor to enable the at least one processor to perform the steps of the enterprise relevance analysis method as described above.
[0146] In this embodiment, joint information of each enterprise is obtained, including each enterprise's financial flow data and enterprise business information; the business association relationships and alternative medium association relationships between the enterprises are determined based on the joint information, and a first association table between the enterprises is generated based on the determined business association relationships and alternative medium association relationships. Analysis based on multiple media can effectively identify hidden association relationships, thereby improving the coverage of the analysis; according to a preset association medium screening algorithm, the alternative media in the first association table are screened to obtain a second association table, filtering out media with no actual association, eliminating common media and noise data, and retaining media that are highly correlated with the actual associations between the enterprises, so that the second association table can more accurately reflect the actual associations between the enterprises; based on the second association table, enterprises are grouped using a graph generation function and a connected branch clustering algorithm to obtain enterprise grouping information. This not only identifies direct enterprise associations, but also mines potential enterprise group relationships through a multi-level, complex network structure, thereby improving the accuracy and coverage of the final enterprise association analysis results.
[0147] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0148] Obviously, the embodiments described above are only some of the embodiments of the present application, rather than all of the embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions described in the aforementioned specific embodiments, or to make equivalent replacements for some of the technical features therein. Any equivalent structure made using the contents of the present application specification and the accompanying drawings, directly or indirectly used in other related technical fields, is also within the scope of patent protection of the present application.
Claims
1. A method for analyzing enterprise relevance, characterized in that: The steps include: Acquire joint information of each enterprise, the joint information including financial transaction data and business information of each enterprise; Determining the industrial and commercial association relationships and the alternative media association relationships between the enterprises based on the joint information, and generating a first association table between the enterprises based on the determined industrial and commercial association relationships and the alternative media association relationships, wherein the steps of determining the industrial and commercial association relationships and the alternative media association relationships between the enterprises based on the joint information, and generating the first association table between the enterprises based on the determined industrial and commercial association relationships and the alternative media association relationships include: Determining the industrial and commercial association relationship between the enterprises in the joint information according to the preset industrial and commercial association rules, and generating an industrial and commercial association relationship table based on the industrial and commercial association relationship, wherein the joint primary key in the industrial and commercial association relationship table is the enterprise identifiers of the two enterprises with the industrial and commercial association relationship; Determining candidate media association relationships between enterprises in the joint information according to various preset candidate media types, and generating a candidate media association table based on the candidate media association relationships, wherein the joint primary key in the candidate media association table is the enterprise IDs of the two enterprises with the candidate media association relationship; Using the candidate media association table as the left table, the industrial and commercial association table is left-linked to the candidate media association table through a joint primary key to obtain a first association table between enterprises; The candidate media in the first association table are screened according to a preset association medium screening algorithm to obtain a second association table, wherein the step of screening the candidate media in the first association table according to the preset association medium screening algorithm includes: For each pair of enterprises in the first association table, calculating the association strength between the two enterprises under each alternative medium; Calculate the Pearson correlation coefficient between the industrial and commercial relationship and the relationship between the two enterprises in each alternative medium based on the strength of the relationship between the two enterprises under each alternative medium and the industrial and commercial relationship between the two enterprises; Screening the candidate media in the first association table according to the obtained Pearson correlation coefficients; The step of screening the candidate media in the first association table according to a preset association medium screening algorithm further includes: For each candidate medium in the first association table, calculating the mean and standard deviation of each enterprise on the candidate medium; Calculating an outlier threshold corresponding to the candidate medium based on the mean and the standard deviation; Comparing the usage frequency of the candidate medium with the outlier threshold, and retaining or deleting the candidate medium based on the comparison result; Based on the second association table, grouping the enterprises using a graph generation function and a connected branch clustering algorithm to obtain enterprise grouping information; An enterprise relevance analysis result is generated according to the enterprise grouping information.
2. The enterprise relevance analysis method according to claim 1, characterized in that: The step of screening the candidate media in the first association table according to a preset association medium screening algorithm further includes: Obtaining aggregation information of each candidate medium in the first association table; Eliminate alternative media whose aggregation information does not meet the aggregation conditions.
3. The enterprise relevance analysis method according to claim 1, characterized in that: The step of grouping enterprises based on the second association table by using a graph generation function and a connected component clustering algorithm to obtain enterprise grouping information includes: Determining a seed node in the second association table based on the second association table and a preset input medium dictionary; Taking the seed node as a starting point, performing node traversal based on the second association table, and generating a graph model by a graph generation function during the traversal; A connected component clustering algorithm is run on the graph model to group enterprises and obtain enterprise grouping information.
4. The enterprise relevance analysis method according to claim 1, characterized in that: After the step of generating the enterprise relevance analysis result according to the enterprise grouping information, the method further includes: Obtain the enterprise relevance analysis results corresponding to the target enterprise; Determining the credit information associated with the target enterprise based on the enterprise association analysis results, so as to conduct a risk assessment on the target enterprise based on the credit information; or Identifying fraudulent transactions associated with the target enterprise based on the enterprise association analysis results; Perform feature engineering optimization on the risk assessment model based on the identified fraudulent transactions.
5. An enterprise relevance analysis device, characterized in that: The enterprise relevance analysis device is used to implement the steps of the enterprise relevance analysis method according to any one of claims 1 to 4, and the enterprise relevance analysis device includes: An information acquisition module is used to acquire joint information of each enterprise, wherein the joint information includes financial flow data and business information of each enterprise; an association generating module, configured to determine the industrial and commercial association relationship and the alternative medium association relationship between the enterprises based on the joint information, and generate a first association table between the enterprises based on the determined industrial and commercial association relationship and the alternative medium association relationship; a medium screening module, configured to screen the candidate media in the first association table according to a preset association medium screening algorithm to obtain a second association table; An enterprise grouping module, configured to group enterprises based on the second association table using a graph generation function and a connected branch clustering algorithm to obtain enterprise grouping information; The result generating module is used to generate enterprise relevance analysis results according to the enterprise grouping information.
6. A computer device, characterized in that: The system comprises a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the enterprise relevance analysis method according to any one of claims 1 to 4 when executing the computer-readable instructions.
7. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the enterprise relevance analysis method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Enterprise association relationship analysis method, equipment and medium
CN117592652A