Enterprise risk assessment method, device, equipment, storage medium and program product
By constructing an enterprise relationship graph and calculating node embedding vectors, the problem of insufficient accuracy in enterprise risk assessment is solved, and more efficient risk association identification and assessment are achieved.
Patent Information
- Application Number
- CN202510936101.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-07-08
AI Technical Summary
The accuracy of enterprise risk assessment in existing technologies is not high, especially when determining the risk correlation between two enterprises to be assessed.
By constructing an enterprise relationship graph, based on the feature vectors of enterprise nodes and the feature vectors of neighboring nodes, the node embedding vectors between enterprises are calculated. The risk association assessment results between enterprises are determined by similarity calculation, and the accuracy of the assessment is improved by combining enterprise registration information and verification process.
It improves the accuracy of enterprise risk assessment, enabling more accurate identification of risk-related enterprises and enhancing assessment efficiency and accuracy.
Smart Images

Figure CN120430637B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of risk assessment, and in particular to an enterprise risk assessment method, device, equipment, storage medium and program product. BACKGROUND
[0002] Enterprise risk assessment is to assess the risk of an enterprise to be assessed. The traditional enterprise risk assessment method is to directly analyze data based on enterprise data of the enterprise to be assessed to determine whether it is a risk enterprise; this method requires data analysis for each enterprise to be assessed. To improve the efficiency of enterprise risk assessment, the risk correlation of two associated enterprises to be assessed can be assessed first, so that only one of the two enterprises to be assessed needs to be analyzed. If the enterprise to be assessed is a risk enterprise, other enterprises associated with it can be quickly identified as risk enterprises without the need for separate data analysis. Therefore, how to determine the risk correlation assessment result of two enterprises to be assessed is a pressing need to be addressed.
[0003] At present, the risk correlation assessment result of two enterprises to be assessed is determined based on the similarity calculation result of the enterprise data of the two enterprises to be assessed. However, the risk correlation assessment result determined by the prior art is not accurate, for example, although the similarity of the enterprise data of the two enterprises to be assessed is high, the correlation between the two enterprises to be assessed is not strong in reality. Therefore, how to improve the accuracy of enterprise risk correlation assessment to improve the accuracy of enterprise risk assessment is a pressing technical problem to be solved. SUMMARY
[0004] The present application provides an enterprise risk assessment method, device, equipment, storage medium and program product to solve the defect of low accuracy of enterprise risk assessment in the prior art and achieve an enterprise risk assessment scheme with high accuracy.
[0005] The present application provides an enterprise risk assessment method, comprising:
[0006] determining two enterprises to be assessed that are associated with the risk to be assessed; the two enterprises to be assessed include a first enterprise to be assessed and a second enterprise to be assessed;
[0007] determining a first node embedding vector of the first enterprise to be assessed and a second node embedding vector of the second enterprise to be assessed based on an enterprise relationship graph; the nodes in the enterprise relationship graph are enterprise nodes, the edges in the enterprise relationship graph represent the correlation between enterprises, and the node feature vector of any enterprise node in the enterprise relationship graph is determined based on the enterprise data of the enterprise node;
[0008] determine a risk association assessment result of the first to-be-evaluated enterprise and the second to-be-evaluated enterprise based on the similarity calculation result of the first node embedding vector and the second node embedding vector;
[0009] The first node embedding vector is determined based on a node feature vector of the first to-be-evaluated enterprise and node feature vectors of each first neighbor node in a first neighbor node set, and each first neighbor node in the first neighbor node set is an enterprise node in the enterprise relationship graph that has an association relationship with the first to-be-evaluated enterprise. The second node embedding vector is determined based on a node feature vector of the second to-be-evaluated enterprise and node feature vectors of each second neighbor node in a second neighbor node set, and each second neighbor node in the second neighbor node set is an enterprise node in the enterprise relationship graph that has an association relationship with the second to-be-evaluated enterprise.
[0010] According to the enterprise risk assessment method provided in the application, the first node embedding vector of the first to-be-evaluated enterprise and the second node embedding vector of the second to-be-evaluated enterprise are determined based on the enterprise relationship graph, and the method comprises the following steps:
[0011] The association layer number of the first to-be-evaluated enterprise and the second to-be-evaluated enterprise is determined based on the enterprise relationship graph. The association layer number is the total edge number of the shortest association path of the first to-be-evaluated enterprise and the second to-be-evaluated enterprise in the enterprise relationship graph.
[0012] The first node embedding vector of the first to-be-evaluated enterprise and the second node embedding vector of the second to-be-evaluated enterprise are determined based on the association layer number.
[0013] If the association layer number is equal to 1, the first node embedding vector is the node feature vector of the first to-be-evaluated enterprise, and the second node embedding vector is the node feature vector of the second to-be-evaluated enterprise.
[0014] If the association layer number is greater than 1, the first node embedding vector is determined based on the node feature vector of the first to-be-evaluated enterprise and a neighbor aggregation vector of a first neighbor node set, the neighbor aggregation vector of the first neighbor node set is determined based on node feature vectors of each first neighbor node in the first neighbor node set, the second node embedding vector is determined based on the node feature vector of the second to-be-evaluated enterprise and a neighbor aggregation vector of a second neighbor node set, and the neighbor aggregation vector of the second neighbor node set is determined based on node feature vectors of each second neighbor node in the second neighbor node set.
[0015] According to the enterprise risk assessment method provided by the application, if the number of the associated layers is greater than 1, the number of the associated layers is K, and the first node embedding vector is determined based on the following manner:
[0016] The node embedding vectors of each first neighbor node in the K-1 layer are aggregated to obtain a neighbor aggregation vector of the first neighbor node set.
[0017] The neighbor aggregation vector of the first neighbor node set and the node embedding vector of the first enterprise to be evaluated in the K-1 layer are aggregated to obtain the first node embedding vector of the first enterprise to be evaluated in the K layer.
[0018] The node embedding vector of any first neighbor node in the first layer is the node feature vector of the first neighbor node, and the node embedding vector of the first enterprise to be evaluated in the first layer is the node feature vector of the first enterprise to be evaluated.
[0019] According to the enterprise risk assessment method provided by the application, the aggregation processing of the neighbor aggregation vector of the first neighbor node set and the node embedding vector of the first enterprise to be evaluated in the K-1 layer to obtain the first node embedding vector of the first enterprise to be evaluated in the K layer comprises:
[0020] The aggregation processing of the neighbor aggregation vector of the first neighbor node set and the node embedding vector of the first enterprise to be evaluated in the K-1 layer to obtain the first node embedding vector of the first enterprise to be evaluated in the K layer comprises:
[0021] The node embedding vector is input into a feature extraction layer to obtain a node extraction vector output by the feature extraction layer.
[0022] The node extraction vector is input into a nonlinear activation function layer to obtain the first node embedding vector output by the nonlinear activation function layer.
[0023] According to the enterprise risk assessment method provided by the application, the first neighbor node set is determined based on the following manner:
[0024] The neighbor node set of the first enterprise to be evaluated in the K layer is sampled to obtain the first neighbor node set of the first enterprise to be evaluated in the K layer.
[0025] The neighbor node set of the first enterprise to be evaluated in the K layer comprises all direct neighbor nodes of each neighbor node in the neighbor node set of the first enterprise to be evaluated in the K-1 layer, and the neighbor node set of the first enterprise to be evaluated in the first layer comprises all direct neighbor nodes of the first enterprise to be evaluated.
[0026] According to the enterprise risk assessment method provided by the application, the risk correlation assessment result of the first to-be-evaluated enterprise and the second to-be-evaluated enterprise is determined based on the similarity calculation result of the first node embedding vector and the second node embedding vector, and the method comprises the following steps:
[0027] The risk correlation assessment result is determined based on a comparison result of the similarity calculation result and a preset similarity threshold value; the risk correlation assessment result comprises a first risk correlation result and a second risk correlation result, and a risk correlation degree of the first risk correlation result is greater than a risk correlation degree of the second risk correlation result;
[0028] After the risk correlation assessment result of the first to-be-evaluated enterprise and the second to-be-evaluated enterprise is determined based on the similarity calculation result of the first node embedding vector and the second node embedding vector, the method further comprises the following steps:
[0029] In a case where the risk correlation assessment result is the first risk correlation result, enterprise registration information is acquired;
[0030] In a case where it is determined based on the enterprise registration information that the first to-be-evaluated enterprise and the second to-be-evaluated enterprise have associated records, it is determined that the first to-be-evaluated enterprise and the second to-be-evaluated enterprise have risk correlation;
[0031] In a case where it is determined based on the enterprise registration information that the first to-be-evaluated enterprise and the second to-be-evaluated enterprise have no associated records, a verification process is triggered; the verification process is used to verify whether the first to-be-evaluated enterprise and the second to-be-evaluated enterprise have risk correlation.
[0032] According to the enterprise risk assessment method provided by the application, the enterprise relationship graph is determined based on the following manner:
[0033] Enterprise data of a plurality of enterprises is acquired; the enterprise data comprises multidimensional data, and sources of the enterprise data comprise a plurality of different databases;
[0034] An enterprise relationship graph is constructed based on the enterprise data of the plurality of enterprises;
[0035] In the enterprise relationship graph, an enterprise node is represented by a node feature vector, the node feature vector is a multidimensional feature vector, and the multidimensional feature vector is determined based on real-time enterprise data.
[0036] According to the enterprise risk assessment method provided by the application, the enterprise risk assessment method further comprises the following steps:
[0037] In a case where it is determined that there is a risk enterprise, a to-be-risk-evaluated enterprise having an associated relationship with the risk enterprise is determined based on the enterprise relationship graph.
[0038] determine whether the enterprise to be risk evaluated is a risk enterprise based on the correlation degree coefficient of the risk enterprise and the enterprise to be risk evaluated.
[0039] The correlation degree coefficient is obtained by multiplying the relationship strength coefficients of each side of the shortest correlation path between the risk enterprise and the enterprise to be risk evaluated, and the relationship strength coefficient is used to represent the correlation strength between enterprises and is greater than 0 and less than or equal to 1.
[0040] According to the enterprise risk evaluation method provided by the application, the relationship strength coefficient of any side is determined based on the following method:
[0041] determine the relationship strength sub-coefficients of each correlation type of the side, and the relationship strength sub-coefficient of any correlation type is used to represent the correlation strength between enterprises in the correlation type;
[0042] weight and aggregate each relationship strength sub-coefficient based on the weight of each correlation type to obtain the relationship strength coefficient of the side.
[0043] The application further provides an enterprise risk evaluation device, which comprises:
[0044] an enterprise determination module configured to determine two enterprises to be evaluated for risk correlation, wherein the two enterprises to be evaluated comprise a first enterprise to be evaluated and a second enterprise to be evaluated;
[0045] a vector determination module configured to determine a first node embedding vector of the first enterprise to be evaluated and a second node embedding vector of the second enterprise to be evaluated based on an enterprise relationship graph, wherein the nodes in the enterprise relationship graph are enterprise nodes, the sides in the enterprise relationship graph are used to represent the correlation between enterprises, and the node feature vector of any enterprise node in the enterprise relationship graph is determined based on the enterprise data of the enterprise node;
[0046] a correlation evaluation module configured to determine a risk correlation evaluation result of the first enterprise to be evaluated and the second enterprise to be evaluated based on the similarity calculation result of the first node embedding vector and the second node embedding vector.
[0047] The first node embedding vector is determined based on a node feature vector of the first to-be-evaluated enterprise and node feature vectors of each first neighbor node in a first neighbor node set, each first neighbor node in the first neighbor node set being an enterprise node in the enterprise relationship graph that has a correlation with the first to-be-evaluated enterprise.
[0048] The application further provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the enterprise risk evaluation method according to any one of the above when executing the program.
[0049] The application further provides a non-transitory computer readable storage medium, which stores a computer program, and the computer program is executable on a processor to implement the enterprise risk evaluation method according to any one of the above.
[0050] The application further provides a computer program product, which includes a computer program, and the computer program is executable on a processor to implement the enterprise risk evaluation method according to any one of the above.
[0051] The enterprise risk assessment method, device, equipment, storage medium and program product provided by the application determine two to-be-evaluated enterprises related to a to-be-evaluated risk, the two to-be-evaluated enterprises include a first to-be-evaluated enterprise and a second to-be-evaluated enterprise, a first node embedding vector of the first to-be-evaluated enterprise and a second node embedding vector of the second to-be-evaluated enterprise are determined based on an enterprise relationship graph, the nodes in the enterprise relationship graph are enterprise nodes, the edges in the enterprise relationship graph are used to represent the association relationship between enterprises, the node feature vector of any enterprise node in the enterprise relationship graph is determined based on the enterprise data of the enterprise node, the first node embedding vector is determined based on the node feature vector of the first to-be-evaluated enterprise and the node feature vectors of each first neighbor node in a first neighbor node set, each first neighbor node in the first neighbor node set is an enterprise node in the enterprise relationship graph that has an association relationship with the first to-be-evaluated enterprise, the second node embedding vector is determined based on the node feature vector of the second to-be-evaluated enterprise and the node feature vectors of each second neighbor node in a second neighbor node set, each second neighbor node in the second neighbor node set is an enterprise node in the enterprise relationship graph that has an association relationship with the second to-be-evaluated enterprise, based on this, the first node embedding vector and the second node embedding vector are both determined based on not only the node feature vector of itself but also the feature vectors of its neighbor nodes, so that not only the features of the to-be-evaluated enterprises themselves are considered, but also the features of the enterprises associated with the to-be-evaluated enterprises are considered, and then the similarity calculation result determined based on the first node embedding vector and the second node embedding vector can also represent whether the features of the enterprises associated with the two to-be-evaluated enterprises are similar, so that it is avoided that the enterprise data of the two to-be-evaluated enterprises is high in similarity but the association relationship between the two to-be-evaluated enterprises is not strong in reality, and the accuracy of the determined risk association evaluation result is improved, that is, the accuracy of the enterprise risk association evaluation is improved, and finally the accuracy of the enterprise risk evaluation is improved. BRIEF DESCRIPTION OF DRAWINGS
[0052] In order to more clearly illustrate the technical solutions in the application or the prior art, the following will briefly introduce the drawings needed in the embodiments or the prior art description. Obviously, the drawings in the following description are some embodiments of the application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.
[0053] Figure 1 is one of the flowcharts of the enterprise risk assessment method provided by the application.
[0054] Figure 2 is the second flowchart of the enterprise risk assessment method provided by the application.
[0055] Figure 3 is the third flowchart of the enterprise risk assessment method provided by the application.
[0056] Figure 4 is a structural schematic diagram of an enterprise risk assessment device provided by the present application.
[0057] Figure 5 is a structural schematic diagram of an electronic device provided by the present application. DETAILED DESCRIPTION
[0058] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the protection scope of the present application.
[0059] The present application proposes the following embodiments. The present application will be described below with reference to the drawings. Figures 1-3 An enterprise risk assessment method is described.
[0060] Figure 1 is one of the flowcharts of the enterprise risk assessment method provided by the present application, as shown in the figure, the enterprise risk assessment method comprises the following steps 110, 120 and 130. Figure 1
[0061] Step 110, two to-be-evaluated enterprises associated with a to-be-evaluated risk are determined.
[0062] Among them, the two to-be-evaluated enterprises include a first to-be-evaluated enterprise and a second to-be-evaluated enterprise.
[0063] Here, the two to-be-evaluated enterprises are two enterprises to be evaluated for whether they have risk correlation.
[0064] Exemplarily, assuming that the first to-be-evaluated enterprise is a risk enterprise, if the first to-be-evaluated enterprise has risk correlation with the second to-be-evaluated enterprise, the second to-be-evaluated enterprise is also a risk enterprise, of course, the number of the second to-be-evaluated enterprise can be multiple, based on this, only one to-be-evaluated enterprise is determined as a risk enterprise, multiple other enterprises can be quickly determined as risk enterprises, thereby improving the enterprise risk assessment efficiency.
[0065] Step 120, based on an enterprise relationship graph, a first node embedding vector of the first to-be-evaluated enterprise and a second node embedding vector of the second to-be-evaluated enterprise are respectively determined.
[0066] Among them, the nodes in the enterprise relationship graph are enterprise nodes, the edges in the enterprise relationship graph are used to represent the correlation between enterprises, and the node feature vector of any enterprise node in the enterprise relationship graph is determined based on the enterprise data of the enterprise node.
[0067] Here, the enterprise relationship graph is constructed based on enterprise data of a plurality of enterprises, specifically, based on enterprise data of a plurality of enterprises, graph computing is performed to obtain the enterprise relationship graph. The plurality of enterprises should include the first to-be-evaluated enterprise and the second to-be-evaluated enterprise. Further, the enterprise data includes multi-dimensional data, thereby improving the construction accuracy of the enterprise relationship graph, and further improving the evaluation accuracy of the enterprise risk association, and ultimately improving the evaluation accuracy of the enterprise risk. Further, the sources of enterprise data include a plurality of different databases, so as to expand the sources of enterprise data, thereby improving the construction accuracy of the enterprise relationship graph, and further improving the evaluation accuracy of the enterprise risk association, and ultimately improving the evaluation accuracy of the enterprise risk. Further, the enterprise data is real-time enterprise data, thereby ensuring that the enterprise relationship graph is also real-time, i.e., improving the construction accuracy of the enterprise relationship graph, and further improving the evaluation accuracy of the enterprise risk association, and ultimately improving the evaluation accuracy of the enterprise risk.
[0068] Here, the enterprise node is represented by a node feature vector, and if the enterprise data includes multi-dimensional data, the node feature vector is a multi-dimensional feature vector. The edges in the enterprise relationship graph are represented by heterogeneous edges. Further, based on the heterogeneous graph-based enterprise relationship network modeling scheme, the enterprise relationship graph is constructed; specifically, node modeling is performed first, and then edge relationship modeling is performed.
[0069] Here, the association relationship can include but is not limited to at least one of the following: equity control relationship, supply chain relationship, cross-employment relationship of senior managers, common investment relationship, intellectual property association relationship, guarantee mutual guarantee relationship, bank-enterprise lending relationship, and judicial association relationship, etc. The association relationship between two enterprises can include multiple, i.e., multiple association relationships exist at the same time. For example, if the direct shareholding ratio is ≥30% or the indirect holding level is ≤3, there is an equity control relationship between the two enterprises; if the annual order transaction amount is > 1000 million yuan and the continuous cooperation cycle is ≥2 years, there is an equity control relationship between the two enterprises; if there is an overlap in the current director / supervisor / manager position and the overlap period is ≥6 months, there is a cross-employment relationship of senior managers between the two enterprises; if the total investment amount of joint investment in the same entity enterprise is > 500 million yuan, there is a common investment relationship between the two enterprises; if the number of joint patent applications within three years is ≥5 or the technical standard is jointly signed, there is an intellectual property association relationship between the two enterprises; if the single guarantee amount exceeds 10% of the net assets, there is a mutual guarantee agreement between the two enterprises.
[0070] The first node embedding vector is determined based on a node feature vector of the first to-be-evaluated enterprise and node feature vectors of each first neighbor node in a first neighbor node set, each first neighbor node in the first neighbor node set being an enterprise node in the enterprise relationship graph that has a correlation relationship with the first to-be-evaluated enterprise. The second node embedding vector is determined based on a node feature vector of the second to-be-evaluated enterprise and node feature vectors of each second neighbor node in a second neighbor node set, each second neighbor node in the second neighbor node set being an enterprise node in the enterprise relationship graph that has a correlation relationship with the second to-be-evaluated enterprise.
[0071] Specifically, the node feature vector of the first to-be-evaluated enterprise and the node feature vectors of each first neighbor node in the first neighbor node set can be aggregated to obtain the first node embedding vector. The node feature vector of the second to-be-evaluated enterprise and the node feature vectors of each second neighbor node in the second neighbor node set can be aggregated to obtain the second node embedding vector.
[0072] Since the enterprise nodes in the enterprise relationship graph are represented by node feature vectors, the node feature vector of the first to-be-evaluated enterprise and the node feature vector of the second to-be-evaluated enterprise can be determined based on the enterprise relationship graph.
[0073] Here, each first neighbor node in the first neighbor node set can have a direct correlation relationship with the node corresponding to the first to-be-evaluated enterprise, or can have an indirect correlation relationship. Each second neighbor node in the second neighbor node set can have a direct correlation relationship with the node corresponding to the second to-be-evaluated enterprise, or can have an indirect correlation relationship. For example, a direct correlation relationship means that two nodes are directly connected by one edge, and an indirect correlation relationship means that two nodes are not directly connected by one edge, i.e., there are other nodes between them. The first neighbor node set can include all neighbor nodes associated with the first to-be-evaluated enterprise, or can only include part of the neighbor nodes associated with the first to-be-evaluated enterprise, such as setting a correlation layer number, neighbor nodes with a correlation layer number greater than a preset correlation layer number are not included in the first neighbor node set. Similarly, the second neighbor node set can include all neighbor nodes associated with the second to-be-evaluated enterprise, or can only include part of the neighbor nodes associated with the second to-be-evaluated enterprise, such as setting a correlation layer number, neighbor nodes with a correlation layer number greater than a preset correlation layer number are not included in the second neighbor node set.
[0074] It should be understood that the first node embedding vector and the second node embedding vector are determined not only based on the node feature vector of the node itself, but also based on the feature vectors of the neighbor nodes, so as to consider not only the characteristics of the to-be-evaluated enterprise itself, but also the characteristics of the enterprises associated with the to-be-evaluated enterprise, and then the similarity calculation result determined based on the first node embedding vector and the second node embedding vector can also represent whether the characteristics of the enterprises associated with the two to-be-evaluated enterprises are similar, thereby avoiding that the enterprise data of the two to-be-evaluated enterprises is high in similarity, but in fact the association between the two to-be-evaluated enterprises is not strong, and then the accuracy of the subsequently determined risk association evaluation result is improved, that is, the accuracy of the enterprise risk association evaluation is improved, and finally the accuracy of the enterprise risk evaluation is improved.
[0075] Further, all neighbor nodes of the first to-be-evaluated enterprise can be sampled to obtain a first neighbor node set, so as to reduce the number of neighbor nodes in the first neighbor node set, and then reduce the calculation amount, improve the enterprise risk evaluation efficiency, and ensure that the first node embedding vector is accurately obtained, and then the accuracy of the subsequently determined risk association evaluation result is further improved, that is, the accuracy of the enterprise risk association evaluation is further improved, and finally the accuracy of the enterprise risk evaluation is further improved.
[0076] Further, all neighbor nodes of the second to-be-evaluated enterprise can be sampled to obtain a second neighbor node set, so as to reduce the number of neighbor nodes in the second neighbor node set, and then reduce the calculation amount, improve the enterprise risk evaluation efficiency, and ensure that the second node embedding vector is accurately obtained, and then the accuracy of the subsequently determined risk association evaluation result is further improved, that is, the accuracy of the enterprise risk association evaluation is further improved, and finally the accuracy of the enterprise risk evaluation is further improved.
[0077] Step 130, determining a risk association evaluation result of the first to-be-evaluated enterprise and the second to-be-evaluated enterprise based on the similarity calculation result of the first node embedding vector and the second node embedding vector.
[0078] Here, the calculation method of the similarity calculation result can be set according to actual needs, for example, a cosine similarity calculation method.
[0079] In a specific embodiment, the risk association evaluation result is determined based on a comparison result of the similarity calculation result and a preset similarity threshold. The risk association evaluation result includes a first risk association result and a second risk association result, and the risk association degree of the first risk association result is greater than the risk association degree of the second risk association result. Further, the first risk association result indicates that the first to-be-evaluated enterprise and the second to-be-evaluated enterprise have risk association, and the second risk association result indicates that the first to-be-evaluated enterprise and the second to-be-evaluated enterprise do not have risk association.
[0080] The enterprise risk assessment method provided by the embodiment of the present application determines two to-be-evaluated enterprises associated with a to-be-evaluated risk, the two to-be-evaluated enterprises including a first to-be-evaluated enterprise and a second to-be-evaluated enterprise, determines a first node embedding vector of the first to-be-evaluated enterprise and a second node embedding vector of the second to-be-evaluated enterprise based on an enterprise relationship graph, the nodes in the enterprise relationship graph are enterprise nodes, the edges in the enterprise relationship graph are used to represent the association relationship between enterprises, the node feature vector of any enterprise node in the enterprise relationship graph is determined based on the enterprise data of the enterprise node, the first node embedding vector is determined based on the node feature vector of the first to-be-evaluated enterprise and the node feature vectors of each first neighbor node in a first neighbor node set, each first neighbor node in the first neighbor node set is an enterprise node in the enterprise relationship graph that has an association relationship with the first to-be-evaluated enterprise, the second node embedding vector is determined based on the node feature vector of the second to-be-evaluated enterprise and the node feature vectors of each second neighbor node in a second neighbor node set, each second neighbor node in the second neighbor node set is an enterprise node in the enterprise relationship graph that has an association relationship with the second to-be-evaluated enterprise, based on this, the first node embedding vector and the second node embedding vector are both determined based on not only the node feature vector of itself but also the feature vectors of its neighbor nodes, so that not only the features of the to-be-evaluated enterprises themselves are considered, but also the features of the enterprises associated with the to-be-evaluated enterprises are considered, and then the similarity calculation result based on the first node embedding vector and the second node embedding vector can also represent whether the features of the enterprises associated with the two to-be-evaluated enterprises are similar, so that it is avoided that the enterprise data of the two to-be-evaluated enterprises is high in similarity but the association relationship between the two to-be-evaluated enterprises is not strong in fact, and the accuracy of the determined risk association evaluation result is improved, that is, the accuracy of the enterprise risk association evaluation is improved, and finally the accuracy of the enterprise risk assessment is improved.
[0081] According to any one of the above embodiments, in the method, the step 120 includes steps 121 and 122.
[0082] The step 121 determines the association layer number of the first to-be-evaluated enterprise and the second to-be-evaluated enterprise based on the enterprise relationship graph.
[0083] The association layer number is the total number of edges of the shortest association path between the first to-be-evaluated enterprise and the second to-be-evaluated enterprise in the enterprise relationship graph. Since the nodes in the enterprise relationship graph are enterprise nodes and the edges in the enterprise relationship graph are used to represent the association relationship between enterprises, the association layer number of the two to-be-evaluated enterprises can be determined based on the enterprise relationship graph.
[0084] Here, the shortest association path is the association path with the shortest number of edges among the multiple association paths between the node of the first to-be-evaluated enterprise and the node of the second to-be-evaluated enterprise. For example, the first to-be-evaluated enterprise is enterprise A, the second to-be-evaluated enterprise is enterprise B, there are two association paths between enterprise A and enterprise B, which are enterprise A-enterprise C-enterprise B and enterprise A-enterprise C-enterprise D-enterprise B, and the shortest association path is enterprise A-enterprise C-enterprise B, and the association path number is 2.
[0085] In step 122, based on the association path number, a first node embedding vector of the first to-be-evaluated enterprise and a second node embedding vector of the second to-be-evaluated enterprise are respectively determined.
[0086] If the association path number is equal to 1, the first node embedding vector is a node feature vector of the first to-be-evaluated enterprise, and the second node embedding vector is a node feature vector of the second to-be-evaluated enterprise. That is, the first to-be-evaluated enterprise and the second to-be-evaluated enterprise are in a direct association relationship, and there is no need to consider the features of the neighbor nodes.
[0087] If the association path number is greater than 1, the first node embedding vector is determined based on the node feature vector of the first to-be-evaluated enterprise and a neighbor aggregation vector of a first neighbor node set, the neighbor aggregation vector of the first neighbor node set is determined based on the node feature vectors of each first neighbor node in the first neighbor node set, the second node embedding vector is determined based on the node feature vector of the second to-be-evaluated enterprise and a neighbor aggregation vector of a second neighbor node set, and the neighbor aggregation vector of the second neighbor node set is determined based on the node feature vectors of each second neighbor node in the second neighbor node set. That is, the first to-be-evaluated enterprise and the second to-be-evaluated enterprise are in an indirect association relationship, and therefore the features of the neighbor nodes need to be considered.
[0088] Specifically, the node feature vectors of each first neighbor node can be aggregated to obtain the neighbor aggregation vector of the first neighbor node set. The node feature vectors of each second neighbor node can be aggregated to obtain the neighbor aggregation vector of the second neighbor node set.
[0089] The enterprise risk assessment method provided by the embodiment of the present application determines the correlation layer number of a first to-be-evaluated enterprise and a second to-be-evaluated enterprise based on an enterprise relationship graph, and the correlation layer number is the total number of edges of the shortest correlation path of the first to-be-evaluated enterprise and the second to-be-evaluated enterprise in the enterprise relationship graph, so that the correlation layer number is determined based on the shortest correlation path, and the risk correlation between the two to-be-evaluated enterprises can be more accurately determined. At the same time, based on the correlation layer number, a first node embedding vector of the first to-be-evaluated enterprise and a second node embedding vector of the second to-be-evaluated enterprise are respectively determined, and if the correlation layer number is equal to 1, the first node embedding vector is a node feature vector of the first to-be-evaluated enterprise, and the second node embedding vector is a node feature vector of the second to-be-evaluated enterprise, if the correlation layer number is greater than 1, the first node embedding vector is determined based on the node feature vector of the first to-be-evaluated enterprise and a neighbor aggregation vector of a first neighbor node set, the neighbor aggregation vector of the first neighbor node set is determined based on the node feature vectors of each first neighbor node in the first neighbor node set, the second node embedding vector is determined based on the node feature vector of the second to-be-evaluated enterprise and a neighbor aggregation vector of a second neighbor node set, and the neighbor aggregation vector of the second neighbor node set is determined based on the node feature vectors of each second neighbor node in the second neighbor node set. Based on this, the first node embedding vector and the second node embedding vector can be more accurately determined based on the correlation layer number, so as to further improve the determination accuracy of the risk correlation evaluation result, that is, to further improve the accuracy of the enterprise risk correlation evaluation, and finally to further improve the accuracy of the enterprise risk evaluation. Moreover, the first node embedding vector and the second node embedding vector are determined not only based on the node feature vectors of themselves, but also based on the feature vectors of their neighbor nodes, so as to consider not only the features of the to-be-evaluated enterprises themselves, but also the features of the enterprises associated with the to-be-evaluated enterprises, thereby improving the determination accuracy of the risk correlation evaluation result, that is, improving the accuracy of the enterprise risk correlation evaluation, and finally improving the accuracy of the enterprise risk evaluation.
[0090] Based on any of the above embodiments, in the method, if the correlation layer number is greater than 1, the correlation layer number is K, and the first node embedding vector is determined based on the following manner: steps 1221 and 1222.
[0091] Step 1221: The node embedding vectors of each first neighbor node in the first neighbor node set at the K-1 layer are aggregated to obtain a neighbor aggregation vector of the first neighbor node set.
[0092] Here, the determination manner of the node embedding vector of the first neighbor node at the K-1 layer is basically the same as that of the first node embedding vector, which will not be described here. Based on this, the node embedding vector of any first neighbor node at the first layer is the node feature vector of the first neighbor node.
[0093] The aggregation manner can be set according to actual conditions, and is not limited herein, for example, mean pooling, maximum pooling, weighted summation, LSTM (Long Short-Term Memory) aggregation, etc. The LSTM aggregation manner is specifically: randomly arranging each neighbor node to generate a sequence, capturing long-range dependencies through a gating mechanism, and finally the hidden state is taken as the aggregation result.
[0094] It should be understood that the node embedding vector of each first neighbor node at the K-1 layer is obtained by layer-by-layer aggregation based on the node feature vector of each first neighbor node. Therefore, the first node embedding vector is also determined based on the feature vectors of the neighbor nodes, so that the features of the enterprises associated with the enterprise to be evaluated are also considered, thereby improving the accuracy of the subsequent determined risk association evaluation result. The node embedding vector of the first neighbor node at the K-1 layer is determined based on not only the node feature vector of the first neighbor node itself, but also the feature vectors of the neighbor nodes, so that not only the features of the first neighbor node itself are considered, but also the features of the enterprises associated with the first neighbor node are considered, thereby improving the accuracy of the subsequent determined risk association evaluation result, i.e., improving the accuracy of the enterprise risk association evaluation, and finally improving the accuracy of the enterprise risk evaluation.
[0095] Further, all neighbor nodes of the first enterprise to be evaluated can be sampled to obtain a first neighbor node set, thereby reducing the number of neighbor nodes in the first neighbor node set, further reducing the calculation amount, improving the efficiency of enterprise risk evaluation, and ensuring that the first node embedding vector is accurately obtained, thereby further improving the accuracy of the subsequent determined risk association evaluation result, i.e., further improving the accuracy of the enterprise risk association evaluation, and finally further improving the accuracy of the enterprise risk evaluation.
[0096] In an embodiment, in order to balance the calculation efficiency and information integrity, a hierarchical sampling strategy is performed on the center node corresponding to the first enterprise to be evaluated. Specifically, equal-probability random sampling is performed on each layer of neighbor sets according to a predetermined number, so that the dimension of the adjacency matrix can be effectively controlled, and the problem of memory explosion caused by full neighborhood calculation can be avoided.
[0097] In another embodiment, the neighbor node set of the first enterprise to be evaluated at the K layer can be sampled to obtain a first neighbor node set of the first enterprise to be evaluated at the K layer, and all direct neighbor nodes of each neighbor node in the neighbor node set of the first enterprise to be evaluated at the K-1 layer, i.e., only the neighbor nodes at the K layer are selected, thereby reducing the number of neighbor nodes in the first neighbor node set, further reducing the calculation amount, improving the efficiency of enterprise risk evaluation, and ensuring that the first node embedding vector is accurately obtained, thereby further improving the accuracy of the subsequent determined risk association evaluation result, i.e., further improving the accuracy of the enterprise risk association evaluation, and finally further improving the accuracy of the enterprise risk evaluation.
[0098] Step 1222, aggregate the neighbor aggregation vectors of the first neighbor node set and the node embedding vector of the first to-be-evaluated enterprise at the K-1 layer to obtain the first node embedding vector of the first to-be-evaluated enterprise at the K layer.
[0099] Here, the determination manner of the node embedding vector of the first to-be-evaluated enterprise at the K-1 layer and the first node embedding vector of the first to-be-evaluated enterprise at the K layer is basically the same. Based on this, the node embedding vector of the first to-be-evaluated enterprise at the first layer is the node feature vector of the first to-be-evaluated enterprise.
[0100] Here, the aggregation manner can be set according to actual conditions, which is not limited here, for example, concat (concatenation).
[0101] It should be understood that the node embedding vector of the first to-be-evaluated enterprise at the K-1 layer is obtained by aggregating the node feature vector of the first to-be-evaluated enterprise layer by layer, and therefore, the first node embedding vector is determined not only based on the node feature vector of itself, but also based on the feature vectors of its neighbor nodes, so that not only the features of the to-be-evaluated enterprise itself are considered, but also the features of the enterprises associated with the to-be-evaluated enterprise are considered, and then the similarity calculation result determined based on the first node embedding vector and the second node embedding vector can also represent whether the features of the enterprises associated with the two to-be-evaluated enterprises are similar, thereby avoiding that the enterprise data similarity of the two to-be-evaluated enterprises is high, but actually the association relationship of the two to-be-evaluated enterprises is not strong, and then improving the accuracy of the determined risk association evaluation result, i.e., improving the accuracy of the enterprise risk association evaluation, and finally improving the accuracy of the enterprise risk evaluation. And the node embedding vector of the first to-be-evaluated enterprise at the K-1 layer is determined not only based on the node feature vector of itself, but also based on the feature vectors of its neighbor nodes, so that not only the features of itself are considered, but also the features of the enterprises associated with it are considered, and then the accuracy of the subsequent determined risk association evaluation result is improved, i.e., the accuracy of the enterprise risk association evaluation is improved, and finally the accuracy of the enterprise risk evaluation is improved. And the node embedding vector of the first to-be-evaluated enterprise at the K-1 layer is aggregated, so as to retain the node feature evolution track, i.e., fuse the node historical state, and then improve the accuracy of the subsequent determined risk association evaluation result, i.e., improve the accuracy of the enterprise risk association evaluation, and finally improve the accuracy of the enterprise risk evaluation.
[0102] The enterprise risk assessment method provided by the embodiment of the application aggregates the node embedding vectors of each first neighbor node at the K-1 layer of the first neighbor node set to obtain a neighbor aggregation vector of the first neighbor node set, so that the node embedding vector of the first neighbor node at the K-1 layer is determined not only based on the node feature vector of the first neighbor node itself but also based on the feature vectors of the neighbor nodes, so that the feature of the first neighbor node itself and the features of the enterprises associated with the first neighbor node are considered, and the accuracy of the risk correlation assessment result determined subsequently is further improved, that is, the accuracy of the enterprise risk correlation assessment is further improved, and finally the accuracy of the enterprise risk assessment is further improved. The node embedding vector of the first neighbor node at the K-1 layer of the first neighbor node set and the node embedding vector of the first enterprise to be evaluated at the K-1 layer are aggregated to obtain the first node embedding vector of the first enterprise to be evaluated at the K layer, so that the node embedding vector of the first enterprise to be evaluated at the K-1 layer is determined not only based on the node feature vector of the first enterprise to be evaluated itself but also based on the feature vectors of the neighbor nodes, so that the feature of the first enterprise to be evaluated itself and the features of the enterprises associated with the first enterprise to be evaluated are considered, and the accuracy of the risk correlation assessment result determined subsequently is further improved, that is, the accuracy of the enterprise risk correlation assessment is further improved, and finally the accuracy of the enterprise risk assessment is further improved. The node embedding vector of the first enterprise to be evaluated at the K-1 layer is aggregated to retain the node self-feature evolution track, that is, to fuse the node historical state, and the accuracy of the risk correlation assessment result determined subsequently is further improved, that is, the accuracy of the enterprise risk correlation assessment is further improved, and finally the accuracy of the enterprise risk assessment is further improved.
[0103] Based on any of the above embodiments, in the method, the step 1222 includes a step 12221, a step 12222 and a step 12223.
[0104] In the step 12221, the node embedding vector of the first enterprise to be evaluated at the K-1 layer and the neighbor aggregation vector of the first neighbor node set are aggregated to obtain the node embedding vector of the first enterprise to be evaluated at the K layer.
[0105] In the step 12222, the node embedding vector is input into a feature extraction layer to obtain a node extraction vector output by the feature extraction layer.
[0106] Here, the feature extraction layer is used for further feature extraction of the node embedding vector, that is, cross-layer feature dimension transformation is realized; the feature extraction layer is a trainable network layer. Further, the feature extraction layers corresponding to different correlation layers are different, that is, the node embedding vector is input into the feature extraction layer corresponding to the correlation layer. For example, the feature extraction layer can use a trainable parameter matrix Representation.
[0107] Step 12223, input the node extraction vector into a nonlinear activation function layer to obtain the first node embedding vector output by the nonlinear activation function layer.
[0108] Here, the nonlinear activation function layer can be constructed by a nonlinear activation function, which can be an activation function such as ReLU or LeakyReLU, to enhance the representation ability of the node embedding vector.
[0109] In a specific embodiment, the node embedding vector is iteratively updated by a multi-layer aggregation function to realize hierarchical extraction of graph structure information. It adopts a message passing mechanism, and each layer of network performs two core operations: neighbor information aggregation and feature nonlinear transformation.
[0110] Exemplarily, for the first node embedding vector of the first to-be-evaluated enterprise at the Kth layer, the aggregation formula is as follows, that is, the multi-layer aggregation function is as follows:
[0111] ;
[0112] In the formula, denotes the first to-be-evaluated enterprise the first node embedding vector at the Kth layer, denotes a nonlinear activation function, denotes a trainable parameter matrix corresponding to the Kth layer, denotes a concatenation function, denotes the first to-be-evaluated enterprise the node embedding vector at the K-1th layer, denotes an aggregation function, denotes the first neighbor node the node embedding vector at the K-1th layer, denotes the first to-be-evaluated enterprise the first neighbor node set at the Kth layer. Wherein, by cross-layer feature dimension transformation can be realized; the nonlinear activation function may be an activation function such as ReLU or LeakyReLU to enhance the representation ability of the node embedding vector.
[0113] It should be understood that the network model including the feature extraction layer and the nonlinear activation function layer is trainable, so as to improve the representation ability of the first node embedding vector by training the model, and further improve the accuracy of the subsequent determined risk correlation evaluation result, that is, further improve the accuracy of enterprise risk correlation evaluation, and finally further improve the accuracy of enterprise risk evaluation.
[0114] The enterprise risk assessment method provided by the embodiment of the application fuses nodes into a vector input into a feature extraction layer, obtains a node extraction vector output by the feature extraction layer, and inputs the node extraction vector into a nonlinear activation function layer to obtain a first node embedding vector output by the nonlinear activation function layer, so that the representation capability of the first node embedding vector is further improved through feature extraction and a nonlinear activation function, and the accuracy of a risk correlation assessment result determined subsequently is further improved, that is, the accuracy of enterprise risk correlation assessment is further improved, and finally the accuracy of enterprise risk assessment is further improved.
[0115] Based on any of the above embodiments, in the method, the first neighbor node set is determined based on the following manner:
[0116] The neighbor node set of the first to-be-evaluated enterprise at the Kth layer is sampled to obtain a first neighbor node set of the first to-be-evaluated enterprise at the Kth layer.
[0117] The neighbor node set of the first to-be-evaluated enterprise at the Kth layer includes all direct neighbor nodes of each neighbor node in the neighbor node set of the first to-be-evaluated enterprise at the (K-1)th layer, and the neighbor node set of the first to-be-evaluated enterprise at the first layer includes all direct neighbor nodes of the first to-be-evaluated enterprise.
[0118] Here, the direct neighbor node indicates that the correlation between two nodes is a direct correlation, and the direct correlation indicates that the two nodes are directly connected by an edge.
[0119] In an embodiment, the number of samples of each correlation type can be determined according to the correlation between each neighbor node in the neighbor node set of the first to-be-evaluated enterprise at the (K-1)th layer and its direct neighbor node, and the neighbor node set of the first to-be-evaluated enterprise at the Kth layer is sampled based on the number of samples. More specifically, a relationship-aware sampling function can be used for sampling, that is, truncation control is performed based on the number of samples to control the number of neighbor nodes of each correlation. For heterogeneous graph data, different sampling strategies can be configured for different correlation types, such as down-sampling for high-frequency correlations and up-sampling for low-frequency correlations, so as to improve the modeling capability for long-tail distribution, further improve the representation capability of the first node embedding vector, further improve the accuracy of the risk correlation assessment result determined subsequently, that is, further improve the accuracy of enterprise risk correlation assessment, and finally further improve the accuracy of enterprise risk assessment.
[0120] The enterprise risk assessment method provided by the embodiment of the present application samples the neighbor node set of the first to-be-assessed enterprise at the Kth layer, reduces the number of neighbor nodes in the first neighbor node set by sampling and selecting only the neighbor nodes at the Kth layer, thereby reducing the calculation amount, improving the enterprise risk assessment efficiency, and ensuring that the first node embedding vector is accurately obtained, thereby further improving the accuracy of the subsequent determined risk correlation assessment result, i.e., further improving the accuracy of the enterprise risk correlation assessment, and finally further improving the accuracy of the enterprise risk assessment.
[0121] Based on any of the above embodiments, in the method, if the number of correlation layers is greater than 1 and the number of correlation layers is K, the second node embedding vector is determined based on the following manner: step 1223 and step 1224.
[0122] In step 1223, the node embedding vectors of each second neighbor node at the K-1th layer in the second neighbor node set are aggregated to obtain a neighbor aggregation vector of the second neighbor node set.
[0123] Here, the determination manner of the node embedding vector of the second neighbor node at the K-1th layer and the second node embedding vector is basically the same, which will not be described here. Based on this, the node embedding vector of any of the second neighbor nodes at the first layer is the node feature vector of the second neighbor node.
[0124] Here, the aggregation manner can be set according to actual conditions, which will not be specifically limited here, for example, mean pooling, maximum pooling, weighted summation, LSTM aggregation, etc.
[0125] It should be understood that the node embedding vector of each second neighbor node at the K-1th layer is obtained by layer-by-layer aggregation based on the node feature vector of each second neighbor node, and therefore, the second node embedding vector is also determined based on the feature vectors of its neighbor nodes, so that the features of the enterprises associated with the to-be-assessed enterprise are also considered, thereby improving the accuracy of the subsequently determined risk correlation assessment result. In addition, the node embedding vector of the second neighbor node at the K-1th layer is determined based on not only the node feature vector of itself but also the feature vectors of its neighbor nodes, so that not only the features of itself but also the features of the enterprises associated with it are considered, thereby improving the accuracy of the subsequently determined risk correlation assessment result, i.e., improving the accuracy of the enterprise risk correlation assessment, and finally improving the accuracy of the enterprise risk assessment.
[0126] Further, all neighbor nodes of the second to-be-evaluated enterprise can be sampled to obtain a second neighbor node set, so as to reduce the number of neighbor nodes in the second neighbor node set, further reduce the calculation amount, improve the enterprise risk evaluation efficiency, and ensure that the second node embedding vector is accurately obtained, and further improve the accuracy of the subsequent determined risk correlation evaluation result, that is, further improve the accuracy of the enterprise risk correlation evaluation, and finally further improve the accuracy of the enterprise risk evaluation.
[0127] In an embodiment, in order to balance the calculation efficiency and information integrity, a hierarchical sampling strategy is performed on the center node corresponding to the second to-be-evaluated enterprise. Specifically, equal-probability random sampling is performed on a predetermined number of neighbor sets in each layer, so that the dimension of the adjacency matrix can be effectively controlled, and the problem of video memory explosion caused by full neighborhood calculation can be avoided.
[0128] In another embodiment, the neighbor node set of the second to-be-evaluated enterprise at the Kth layer can be sampled to obtain a second neighbor node set of the second to-be-evaluated enterprise at the Kth layer, and all direct neighbor nodes of each neighbor node in the neighbor node set of the second to-be-evaluated enterprise at the (K-1)th layer, that is, only the neighbor nodes at the Kth layer are selected, so as to reduce the number of neighbor nodes in the second neighbor node set, further reduce the calculation amount, improve the enterprise risk evaluation efficiency, and ensure that the second node embedding vector is accurately obtained, and further improve the accuracy of the subsequent determined risk correlation evaluation result, that is, further improve the accuracy of the enterprise risk correlation evaluation, and finally further improve the accuracy of the enterprise risk evaluation.
[0129] In step 1224, the neighbor aggregation vector of the second neighbor node set and the node embedding vector of the second to-be-evaluated enterprise at the (K-1)th layer are aggregated to obtain a second node embedding vector of the second to-be-evaluated enterprise at the Kth layer.
[0130] Here, the determination manner of the node embedding vector of the second to-be-evaluated enterprise at the (K-1)th layer and the second node embedding vector of the second to-be-evaluated enterprise at the Kth layer is basically the same. Based on this, the node embedding vector of the second to-be-evaluated enterprise at the first layer is the node feature vector of the second to-be-evaluated enterprise.
[0131] Here, the aggregation manner can be set according to actual conditions, which is not limited here, for example, concat (concatenation).
[0132] It should be understood that the node embedding vector of the second to-be-evaluated enterprise at the K-1 layer is obtained by aggregating the node feature vectors of the second to-be-evaluated enterprise layer by layer, and therefore, the second node embedding vector is determined not only based on the node feature vector of the second to-be-evaluated enterprise, but also based on the feature vectors of the neighbor nodes, so that the features of the to-be-evaluated enterprise and the features of the enterprises associated with the to-be-evaluated enterprise are considered, and then the similarity calculation result determined based on the first node embedding vector and the second node embedding vector can also represent whether the features of the enterprises associated with the two to-be-evaluated enterprises are similar, thereby avoiding the case that the enterprise data of the two to-be-evaluated enterprises is similar, but the association between the two to-be-evaluated enterprises is not strong in fact, and thus the accuracy of the determined risk association evaluation result is improved, that is, the accuracy of the enterprise risk association evaluation is improved, and finally the accuracy of the enterprise risk evaluation is improved. Moreover, the node embedding vector of the second to-be-evaluated enterprise at the K-1 layer is determined not only based on the node feature vector of the second to-be-evaluated enterprise, but also based on the feature vectors of the neighbor nodes, so that the features of the to-be-evaluated enterprise and the features of the enterprises associated with the to-be-evaluated enterprise are considered, and thus the accuracy of the subsequently determined risk association evaluation result is improved, that is, the accuracy of the enterprise risk association evaluation is improved, and finally the accuracy of the enterprise risk evaluation is improved. Moreover, the node embedding vector of the second to-be-evaluated enterprise at the K-1 layer is aggregated, so that the evolution track of the node itself is retained, that is, the historical state of the node is fused, and thus the accuracy of the subsequently determined risk association evaluation result is improved, that is, the accuracy of the enterprise risk association evaluation is improved, and finally the accuracy of the enterprise risk evaluation is improved.
[0133] The enterprise risk assessment method provided by the embodiment of the application aggregates the node embedding vectors of each second neighbor node at the K-1 layer of the second neighbor node set to obtain a neighbor aggregation vector of the second neighbor node set, so that the node embedding vector of the second neighbor node at the K-1 layer is determined not only based on the node feature vector of the second neighbor node itself but also based on the feature vectors of the neighbor nodes of the second neighbor node, so that the feature of the second neighbor node itself and the features of the enterprises associated with the second neighbor node are considered, and the accuracy of the risk correlation assessment result determined subsequently is further improved, that is, the accuracy of the enterprise risk correlation assessment is further improved, and finally the accuracy of the enterprise risk assessment is further improved. The node embedding vector of the second neighbor node at the K-1 layer of the second neighbor node set and the node embedding vector of the second enterprise to be evaluated at the K-1 layer are aggregated to obtain a second node embedding vector of the second enterprise to be evaluated at the K layer, so that the node embedding vector of the second enterprise to be evaluated at the K-1 layer is determined not only based on the node feature vector of the second enterprise to be evaluated itself but also based on the feature vectors of the neighbor nodes of the second enterprise to be evaluated, so that the feature of the second enterprise to be evaluated itself and the features of the enterprises associated with the second enterprise to be evaluated are considered, and the accuracy of the risk correlation assessment result determined subsequently is further improved, that is, the accuracy of the enterprise risk correlation assessment is further improved, and finally the accuracy of the enterprise risk assessment is further improved. The node embedding vector of the second enterprise to be evaluated at the K-1 layer is aggregated to retain the node self-feature evolution track, that is, to fuse the node historical state, and the accuracy of the risk correlation assessment result determined subsequently is further improved, that is, the accuracy of the enterprise risk correlation assessment is further improved, and finally the accuracy of the enterprise risk assessment is further improved.
[0134] According to any one of the above embodiments, in the method, the step 1224 comprises: a step 12241, a step 12242 and a step 12243.
[0135] In the step 12241, the node embedding vector of the second neighbor node at the K-1 layer of the second neighbor node set and the node embedding vector of the second enterprise to be evaluated at the K-1 layer are aggregated to obtain a node embedding vector of the second enterprise to be evaluated at the K layer.
[0136] In the step 12242, the node embedding vector is input into a feature extraction layer to obtain a node extraction vector output by the feature extraction layer.
[0137] Here, the feature extraction layer is used for further feature extraction of the node embedding vector, that is, to realize cross-layer feature dimension transformation; the feature extraction layer is a trainable network layer. Further, the feature extraction layers corresponding to different correlation layers are different, that is, the node embedding vector is input into the feature extraction layer corresponding to the correlation layer. For example, the feature extraction layer can use a trainable parameter matrix Representation.
[0138] Step 12243, input the node extraction vector into a nonlinear activation function layer to obtain the second node embedding vector output by the nonlinear activation function layer.
[0139] Here, the nonlinear activation function layer can be constructed by a nonlinear activation function, which can be an activation function such as ReLU or LeakyReLU, to enhance the representation ability of the node embedding vector.
[0140] In a specific embodiment, the node embedding vector is iteratively updated by a multi-layer aggregation function to realize hierarchical extraction of graph structure information. It adopts a message passing mechanism, and each layer of network performs two core operations: neighbor information aggregation and feature nonlinear transformation.
[0141] Exemplarily, for the second node embedding vector of the second to-be-evaluated enterprise at the Kth layer, the aggregation formula is as follows, that is, the multi-layer aggregation function is as follows:
[0142] ;
[0143] In the formula, denotes the second to-be-evaluated enterprise the node embedding vector of the second to-be-evaluated enterprise at the Kth layer, denotes a nonlinear activation function, denotes a trainable parameter matrix corresponding to the Kth layer, denotes a concatenation function, denotes the second to-be-evaluated enterprise the node embedding vector of the second to-be-evaluated enterprise at the K-1th layer, denotes an aggregation function, denotes the second neighbor node the node embedding vector of the second neighbor node at the K-1th layer, denotes the second to-be-evaluated enterprise the second neighbor node set of the second to-be-evaluated enterprise at the Kth layer. Wherein, the cross-layer feature dimension transformation can be realized by The nonlinear activation function may be an activation function such as ReLU or LeakyReLU to enhance the representation ability of the node embedding vector.
[0144] It should be understood that the network model including the feature extraction layer and the nonlinear activation function layer is trainable, so as to improve the representation ability of the second node embedding vector by training the model, and further improve the accuracy of the subsequent determined risk correlation evaluation result, that is, further improve the accuracy of enterprise risk correlation evaluation, and finally further improve the accuracy of enterprise risk evaluation.
[0145] The enterprise risk assessment method provided by the embodiment of the application fuses nodes into a vector input into a feature extraction layer, obtains a node extraction vector output by the feature extraction layer, and inputs the node extraction vector into a nonlinear activation function layer to obtain a second node embedding vector output by the nonlinear activation function layer, so that the representation capability of the second node embedding vector is further improved through feature extraction and a nonlinear activation function, and the accuracy of a risk correlation assessment result determined subsequently is further improved, that is, the accuracy of enterprise risk correlation assessment is further improved, and finally the accuracy of enterprise risk assessment is further improved.
[0146] Based on any of the above embodiments, in the method, the second neighbor node set is determined based on the following manner:
[0147] The neighbor node set of the second to-be-evaluated enterprise at the Kth layer is sampled to obtain a second neighbor node set of the second to-be-evaluated enterprise at the Kth layer.
[0148] The neighbor node set of the second to-be-evaluated enterprise at the Kth layer includes all direct neighbor nodes of each neighbor node in the neighbor node set of the second to-be-evaluated enterprise at the (K-1)th layer, and the neighbor node set of the second to-be-evaluated enterprise at the first layer includes all direct neighbor nodes of the second to-be-evaluated enterprise.
[0149] Here, the direct neighbor node indicates that the correlation between two nodes is a direct correlation, and the direct correlation indicates that the two nodes are directly connected by an edge.
[0150] In an embodiment, the number of samples of each correlation type can be determined according to the correlation between each neighbor node in the neighbor node set of the second to-be-evaluated enterprise at the (K-1)th layer and its direct neighbor node, and the neighbor node set of the second to-be-evaluated enterprise at the Kth layer is sampled based on the number of samples. More specifically, a relationship-aware sampling function can be used for sampling, that is, truncation control is performed based on the number of samples to control the number of neighbor nodes of each correlation. For heterogeneous graph data, different sampling strategies can be configured for different correlation types, such as down-sampling for high-frequency correlations and up-sampling for low-frequency correlations, so as to improve the modeling capability for long-tail distribution, further improve the representation capability of the second node embedding vector, further improve the accuracy of the risk correlation assessment result determined subsequently, that is, further improve the accuracy of enterprise risk correlation assessment, and finally further improve the accuracy of enterprise risk assessment.
[0151] The enterprise risk assessment method provided by the embodiment of the present application samples the neighbor node set of the second to-be-evaluated enterprise at the Kth layer, reduces the number of neighbor nodes in the second neighbor node set by sampling and selecting only the neighbor nodes at the Kth layer, thereby reducing the calculation amount, improving the enterprise risk assessment efficiency, and ensuring that the second node embedding vector is accurately obtained, thereby further improving the accuracy of the subsequent determined risk correlation assessment result, i.e., further improving the accuracy of the enterprise risk correlation assessment, and finally further improving the accuracy of the enterprise risk assessment.
[0152] Based on any of the above embodiments, Figure 2 is a flowchart of the enterprise risk assessment method provided by the present application, as shown in Figure 2 The enterprise risk assessment method comprises steps 110, 120, 131, 140, 150 and 160. Step 140 is after step 131.
[0153] Step 131, based on the comparison result of the similarity calculation result and the preset similarity threshold, determine the risk correlation assessment result.
[0154] Here, the preset similarity threshold can be set according to the actual situation, for example, the value range of the preset similarity threshold is [0.6, 0.8], and correspondingly, the value range of the similarity calculation result is [0, 1].
[0155] Further, the preset similarity threshold is dynamically changed, thereby realizing dynamic threshold adjustment.
[0156] Further, the preset similarity threshold can be calibrated by Bayesian optimization, and the preset similarity threshold is determined by loading 1432 cases in the past three years, and the recall rate reaches 89.6% when the preset similarity threshold is determined to be 0.72 by grid search.
[0157] The risk correlation assessment result comprises a first risk correlation result and a second risk correlation result, and the risk correlation degree of the first risk correlation result is greater than the risk correlation degree of the second risk correlation result.
[0158] Further, the first risk correlation result indicates that the first to-be-evaluated enterprise and the second to-be-evaluated enterprise have risk correlation, and the second risk correlation result indicates that the first to-be-evaluated enterprise and the second to-be-evaluated enterprise do not have risk correlation.
[0159] In an embodiment, the greater the similarity calculation result, the more relevant the first to-be-evaluated enterprise is to the second to-be-evaluated enterprise, and based on this, in a case where the similarity calculation result is greater than a preset similarity threshold, the risk correlation evaluation result is determined to be the first risk correlation result, and in a case where the similarity calculation result is less than or equal to the preset similarity threshold, the risk correlation evaluation result is determined to be the second risk correlation result.
[0160] In another embodiment, the smaller the similarity calculation result, the more relevant the first to-be-evaluated enterprise is to the second to-be-evaluated enterprise, and based on this, in a case where the similarity calculation result is greater than a preset similarity threshold, the risk correlation evaluation result is determined to be the second risk correlation result, and in a case where the similarity calculation result is less than or equal to the preset similarity threshold, the risk correlation evaluation result is determined to be the first risk correlation result.
[0161] Step 140, in a case where the risk correlation evaluation result is the first risk correlation result, obtaining enterprise registration information.
[0162] Specifically, in a case where the risk correlation degree of the first to-be-evaluated enterprise and the second to-be-evaluated enterprise is greater, further obtaining enterprise registration information to verify whether the first to-be-evaluated enterprise and the second to-be-evaluated enterprise are risk correlated, that is, further determining whether the two to-be-evaluated enterprises are high risk correlated, so as to further improve the accuracy of enterprise risk correlation evaluation through a further verification mechanism, and finally further improve the accuracy of enterprise risk evaluation.
[0163] Here, the enterprise registration information includes association record information between multiple enterprises, and the associated enterprises of the enterprise can be associated when the enterprise is registered. Further, the association record is a direct association record.
[0164] Step 150, in a case where it is determined based on the enterprise registration information that the first to-be-evaluated enterprise and the second to-be-evaluated enterprise have an association record, determining that the first to-be-evaluated enterprise and the second to-be-evaluated enterprise are risk correlated.
[0165] If the first to-be-evaluated enterprise and the second to-be-evaluated enterprise have an association record, there is no need to trigger a verification process, and the first to-be-evaluated enterprise and the second to-be-evaluated enterprise can be directly determined to be risk correlated.
[0166] Further, if the first to-be-evaluated enterprise is a risk enterprise, the second to-be-evaluated enterprise is also a risk enterprise. Of course, the number of second to-be-evaluated enterprises can be multiple, and based on this, only one to-be-evaluated enterprise is determined to be a risk enterprise, and multiple other enterprises can be quickly determined to be risk enterprises, thereby improving the efficiency of enterprise risk evaluation.
[0167] Step 160, in the case that it is determined based on the enterprise registration information that the first to-be-evaluated enterprise and the second to-be-evaluated enterprise have no association record, triggering a verification process.
[0168] The verification process is used to verify whether the first to-be-evaluated enterprise and the second to-be-evaluated enterprise have a risk association.
[0169] If the first to-be-evaluated enterprise and the second to-be-evaluated enterprise have no association record, the verification process needs to be triggered to further verify whether the first to-be-evaluated enterprise and the second to-be-evaluated enterprise have a risk association.
[0170] Exemplarily, assuming that the first to-be-evaluated enterprise is a risk enterprise, if the first to-be-evaluated enterprise and the second to-be-evaluated enterprise have a risk association, the second to-be-evaluated enterprise is also a risk enterprise, of course, the number of the second to-be-evaluated enterprise can be multiple, based on this, only one to-be-evaluated enterprise is determined to be a risk enterprise, a plurality of other enterprises can be quickly determined to be risk enterprises, thereby improving the enterprise risk evaluation efficiency.
[0171] The enterprise risk evaluation method provided by the embodiment of the application determines that the first to-be-evaluated enterprise and the second to-be-evaluated enterprise have a risk association in the case that it is determined based on the enterprise registration information that the first to-be-evaluated enterprise and the second to-be-evaluated enterprise have an association record, and triggers a verification process in the case that it is determined based on the enterprise registration information that the first to-be-evaluated enterprise and the second to-be-evaluated enterprise have no association record, thereby further verifying whether the first to-be-evaluated enterprise and the second to-be-evaluated enterprise have a risk association, that is, further improving the accuracy of the enterprise risk association evaluation through a further verification mechanism, and finally further improving the accuracy of the enterprise risk evaluation.
[0172] Based on any of the above embodiments, in the method, the enterprise relationship graph is determined based on the following manner:
[0173] Obtaining enterprise data of a plurality of enterprises;
[0174] Based on the enterprise data of the plurality of enterprises, constructing an enterprise relationship graph.
[0175] Here, the enterprise data can include but is not limited to at least one of the following: enterprise operation data, enterprise operation data, public opinion transmission data, Internet of Things sensing data, etc. For example, the enterprise data includes financial statements, equity structure, operating indicators, audit conclusions, compliance archives, etc. The enterprise data can include structured data and unstructured data.
[0176] Specifically, based on the enterprise data of a plurality of enterprises, a graph calculation is performed to obtain an enterprise relationship graph. The plurality of enterprises should include a first to-be-evaluated enterprise and a second to-be-evaluated enterprise. Further, based on a heterogeneous graph enterprise relationship network modeling scheme, the enterprise relationship graph is constructed; specifically, node modeling is performed first, and then edge relationship modeling is performed.
[0177] In a specific embodiment, the enterprise data is digitized and converted into quantitative data, facilitating the extraction of node feature vectors.
[0178] In a specific embodiment, natural language processing and data analysis techniques in deep learning are used to analyze and extract key information from enterprise data, such as enterprise size, operating status, risk level, etc., and then construct an enterprise relationship graph based on enterprise data, and use graph computing techniques to discover the association between risk enterprises, such as industry chain relationship, equity relationship, etc.
[0179] Further, data preprocessing is performed on the enterprise data, including data cleaning and data standardization. In terms of data cleaning, duplicate data is removed to ensure the uniqueness of the data set; missing values are handled using interpolation, regression prediction or machine learning-based prediction models; and outliers are detected and handled. In terms of data standardization, for continuous variables, z-score standardization or min-max normalization methods are used; for categorical variables, one-hot encoding or label encoding is performed.
[0180] Among them, the enterprise data includes multi-dimensional data, thereby improving the construction accuracy of the enterprise relationship graph, and further improving the evaluation accuracy of the enterprise risk association, and finally improving the evaluation accuracy of the enterprise risk. Correspondingly, the enterprise nodes in the enterprise relationship graph are represented by node feature vectors, and the node feature vectors are multi-dimensional feature vectors.
[0181] It should be noted that the existing technology generally has the defects of single data dimension and weak association analysis, for example, the regulatory system only integrates financial data to implement risk early warning, and lacks the ability to process unstructured data.
[0182] Among them, the sources of the enterprise data include a plurality of different databases, so as to expand the sources of the enterprise data, thereby improving the construction accuracy of the enterprise relationship graph, and further improving the evaluation accuracy of the enterprise risk association, and finally improving the evaluation accuracy of the enterprise risk. For example, data from different regulatory systems, such as tax, business, etc., through data fusion technology, cross-validation and supplement of enterprise data, improve the accuracy and integrity of enterprise data.
[0183] The multi-dimensional feature vector is determined based on real-time enterprise data, so that the enterprise relationship graph is also real-time, that is, the construction accuracy of the enterprise relationship graph is improved, and then the evaluation accuracy of the enterprise risk association is improved, and finally the evaluation accuracy of the enterprise risk is improved.
[0184] Considering the current situation that the existing technology enterprise data acquisition channel is single and the update cycle is more than 30 days, the multi-dimensional behavior characteristics generated in the enterprise operation process in real time cannot be effectively captured, and more than 60% of the early warning signals are false or missed; based on this, the source of the enterprise data includes multiple different databases, and the enterprise data includes multi-dimensional data, and the multi-dimensional feature vector is determined based on real-time enterprise data, so as to realize real-time supervision data flow to build an enterprise relationship graph, and then improve the evaluation accuracy of the enterprise risk association, and finally improve the evaluation accuracy of the enterprise risk.
[0185] It should be understood that by constructing the enterprise relationship graph, a panoramic enterprise portrait can be provided, and then an intelligent early warning mechanism is provided, and the timeliness and accuracy of the enterprise risk evaluation are significantly improved.
[0186] The enterprise risk evaluation method provided by the embodiment of the application, the enterprise data includes multi-dimensional data, and the source of the enterprise data includes multiple different databases, so as to expand the source of the enterprise data, and the multi-dimensional feature vector is determined based on real-time enterprise data, so as to ensure that the enterprise relationship graph is also real-time; in the above manner, the construction accuracy of the enterprise relationship graph can be improved, and then the evaluation accuracy of the enterprise risk association is improved, and finally the evaluation accuracy of the enterprise risk is improved.
[0187] Based on any one of the above embodiments, Figure 3 is a third flowchart of the enterprise risk evaluation method provided by the application, as shown in Figure 3 The enterprise risk evaluation method further includes steps 310 and 320.
[0188] Step 310, in the case of determining that there is a risk enterprise, determining a risk evaluation enterprise associated with the risk enterprise based on the enterprise relationship graph.
[0189] Here, the association relationship between the risk enterprise and the risk evaluation enterprise can be a direct association relationship or an indirect association relationship; the direct association relationship means that two nodes are directly connected by one edge, and the indirect association relationship means that two nodes are not directly connected by one edge, that is, there are other nodes between them.
[0190] Since the edge in the enterprise relationship graph is used to represent the association relationship between enterprises, the risk evaluation enterprise associated with the risk enterprise can be determined based on the enterprise relationship graph.
[0191] In step 320, whether the enterprise to be risk evaluated is a risk enterprise is determined based on the correlation degree coefficient of the risk enterprise and the enterprise to be risk evaluated.
[0192] The correlation degree coefficient is obtained by multiplying the relationship strength coefficients of the edges of the shortest correlation path between the risk enterprise and the enterprise to be risk evaluated, the relationship strength coefficient being used to represent the correlation strength between enterprises and being greater than 0 and less than or equal to 1.
[0193] Here, the shortest correlation path is the correlation path with the shortest number of edges between the node of the risk enterprise and the node of the enterprise to be risk evaluated in the multiple correlation paths between the risk enterprise and the enterprise to be risk evaluated. For example, the risk enterprise is enterprise A, the enterprise to be risk evaluated is enterprise B, there are two correlation paths between enterprise A and enterprise B, which are enterprise A-enterprise C-enterprise B and enterprise A-enterprise C-enterprise D-enterprise B, and the shortest correlation path is enterprise A-enterprise C-enterprise B.
[0194] Here, the relationship strength coefficient can be calculated according to a relevant calculation method, for example, a holding ratio normalization, a transaction amount logarithmic conversion, a tenure coefficient and the like.
[0195] The enterprise risk evaluation method provided by the embodiment of the present application can quickly determine whether the enterprise to be risk evaluated is a risk enterprise, thereby improving the efficiency of enterprise risk evaluation.
[0196] Based on any of the above embodiments, in the method, the relationship strength coefficient of any edge is determined based on the following manner:
[0197] The relationship strength sub-coefficients of the correlation relationship categories of the edge are determined, the relationship strength sub-coefficient of any correlation relationship category being used to represent the correlation strength between enterprises in the correlation relationship category;
[0198] The relationship strength sub-coefficients are weighted and aggregated based on the weights of the correlation relationship categories, to obtain the relationship strength coefficient of the edge.
[0199] Here, the relationship strength sub-coefficient can be calculated according to a relevant calculation method, for example, a holding ratio normalization, a transaction amount logarithmic conversion, a tenure coefficient and the like.
[0200] Since the correlation relationship between two enterprises can include multiple correlation relationships, the relationship strength sub-coefficients of the correlation relationship categories of the edge can be determined.
[0201] Here, the weights of the correlation relationship categories can be set in advance, for example, the weight of the equity control relationship is 0.38. Further, the weights of the correlation relationship categories are added to 1.
[0202] The enterprise risk assessment method provided by the embodiment of the application can accurately obtain the relationship strength coefficient of the edge by weighting and aggregating each relationship strength sub-coefficient based on the weight of each correlation relationship category, that is, considering that different correlation relationship categories have different correlation strengths, thereby improving the accuracy of enterprise risk assessment.
[0203] The enterprise risk assessment device provided by the application is described below, and the enterprise risk assessment device described below can be correspondingly referred to the enterprise risk assessment method described above.
[0204] Figure 4 The enterprise risk assessment device provided by the application is described below, and the enterprise risk assessment device described below can be correspondingly referred to the enterprise risk assessment method described above. Figure 4 As shown in the figure, the enterprise risk assessment device comprises an enterprise determination module 410, a vector determination module 420 and a correlation assessment module 430.
[0205] The enterprise determination module 410 is configured to determine two enterprises to be evaluated for risk correlation; the two enterprises to be evaluated comprise a first enterprise to be evaluated and a second enterprise to be evaluated.
[0206] The vector determination module 420 is configured to determine a first node embedding vector of the first enterprise to be evaluated and a second node embedding vector of the second enterprise to be evaluated based on an enterprise relationship graph; the nodes in the enterprise relationship graph are enterprise nodes, the edges in the enterprise relationship graph are used to represent the correlation between enterprises, and the node feature vector of any enterprise node in the enterprise relationship graph is determined based on the enterprise data of the enterprise node.
[0207] The correlation assessment module 430 is configured to determine a risk correlation assessment result of the first enterprise to be evaluated and the second enterprise to be evaluated based on the similarity calculation result of the first node embedding vector and the second node embedding vector.
[0208] The first node embedding vector is determined based on the node feature vector of the first enterprise to be evaluated and the node feature vector of each first neighbor node in the first neighbor node set, and each first neighbor node in the first neighbor node set is an enterprise node in the enterprise relationship graph that has a correlation with the first enterprise to be evaluated; the second node embedding vector is determined based on the node feature vector of the second enterprise to be evaluated and the node feature vector of each second neighbor node in the second neighbor node set, and each second neighbor node in the second neighbor node set is an enterprise node in the enterprise relationship graph that has a correlation with the second enterprise to be evaluated.
[0209] Figure 5 An example of an entity structure diagram of an electronic device is shown in the figure, Figure 5As shown, the electronic device can include a processor 510, a communications interface 520, a memory 530, and a communications bus 540, wherein the processor 510, the communications interface 520, and the memory 530 complete mutual communication through the communications bus 540. The processor 510 can invoke a logical instruction in the memory 530 to execute an enterprise risk assessment method, which includes: determining two to-be-evaluated enterprises associated with a to-be-evaluated risk; the two to-be-evaluated enterprises include a first to-be-evaluated enterprise and a second to-be-evaluated enterprise; determining a first node embedding vector of the first to-be-evaluated enterprise and a second node embedding vector of the second to-be-evaluated enterprise based on an enterprise relationship graph; the nodes in the enterprise relationship graph are enterprise nodes, the edges in the enterprise relationship graph are used to represent the association relationship between enterprises, and the node feature vector of any enterprise node in the enterprise relationship graph is determined based on the enterprise data of the enterprise node; determining a risk association evaluation result of the first to-be-evaluated enterprise and the second to-be-evaluated enterprise based on a similarity calculation result of the first node embedding vector and the second node embedding vector; wherein the first node embedding vector is determined based on the node feature vector of the first to-be-evaluated enterprise and the node feature vectors of each first neighbor node in a first neighbor node set, each first neighbor node in the first neighbor node set is an enterprise node in the enterprise relationship graph that has an association relationship with the first to-be-evaluated enterprise; and the second node embedding vector is determined based on the node feature vector of the second to-be-evaluated enterprise and the node feature vectors of each second neighbor node in a second neighbor node set, each second neighbor node in the second neighbor node set is an enterprise node in the enterprise relationship graph that has an association relationship with the second to-be-evaluated enterprise.
[0210] In addition, the logical instructions in the memory 530 described above can be implemented in the form of a software function unit and sold or used as an independent product, which can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0211] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program being stored in a non-transitory computer readable storage medium, and the computer program being executable by a processor to cause a computer to perform the enterprise risk assessment method provided by any of the above methods, the method comprising: determining two to-be-evaluated enterprises associated with a to-be-evaluated risk; the two to-be-evaluated enterprises comprising a first to-be-evaluated enterprise and a second to-be-evaluated enterprise; determining a first node embedding vector of the first to-be-evaluated enterprise and a second node embedding vector of the second to-be-evaluated enterprise based on an enterprise relationship graph; the nodes in the enterprise relationship graph are enterprise nodes, the edges in the enterprise relationship graph are used to represent the association relationship between enterprises, and the node feature vector of any enterprise node in the enterprise relationship graph is determined based on the enterprise data of the enterprise node; determining a risk association evaluation result of the first to-be-evaluated enterprise and the second to-be-evaluated enterprise based on a similarity calculation result of the first node embedding vector and the second node embedding vector; wherein the first node embedding vector is determined based on the node feature vector of the first to-be-evaluated enterprise and the node feature vectors of each first neighbor node in a first neighbor node set, each first neighbor node in the first neighbor node set being an enterprise node in the enterprise relationship graph that has an association relationship with the first to-be-evaluated enterprise; and the second node embedding vector is determined based on the node feature vector of the second to-be-evaluated enterprise and the node feature vectors of each second neighbor node in a second neighbor node set, each second neighbor node in the second neighbor node set being an enterprise node in the enterprise relationship graph that has an association relationship with the second to-be-evaluated enterprise.
[0212] In yet another aspect, the present application also provides a non-transitory computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the enterprise risk assessment method provided by any of the above methods, and the method comprises: determining two to-be-evaluated enterprises associated with a to-be-evaluated risk; the two to-be-evaluated enterprises comprise a first to-be-evaluated enterprise and a second to-be-evaluated enterprise; determining a first node embedding vector of the first to-be-evaluated enterprise and a second node embedding vector of the second to-be-evaluated enterprise based on an enterprise relationship graph; the nodes in the enterprise relationship graph are enterprise nodes, the edges in the enterprise relationship graph are used to represent the association relationship between enterprises, and the node feature vector of any enterprise node in the enterprise relationship graph is determined based on the enterprise data of the enterprise node; determining a risk association evaluation result of the first to-be-evaluated enterprise and the second to-be-evaluated enterprise based on the similarity calculation result of the first node embedding vector and the second node embedding vector; wherein the first node embedding vector is determined based on the node feature vector of the first to-be-evaluated enterprise and the node feature vectors of each first neighbor node in a first neighbor node set, and each first neighbor node in the first neighbor node set is an enterprise node in the enterprise relationship graph that has an association relationship with the first to-be-evaluated enterprise; the second node embedding vector is determined based on the node feature vector of the second to-be-evaluated enterprise and the node feature vectors of each second neighbor node in a second neighbor node set, and each second neighbor node in the second neighbor node set is an enterprise node in the enterprise relationship graph that has an association relationship with the second to-be-evaluated enterprise.
[0213] The device embodiments described above are only schematic, wherein the units illustrated as separate components can or can not be physically separate, and the components illustrated as units can or can not be physical units, i.e., can be located in one place or distributed on a plurality of network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment. Those skilled in the art can understand and implement without creative labor.
[0214] From the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be realized by means of software plus a necessary general hardware platform, and of course can also be realized by hardware. Based on such understanding, the above technical solutions, essentially or in terms of the contribution to the prior art, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.
[0215] It should be pointed out finally that the above embodiments are only used to illustrate the technical solutions of the present application, but not to limit the same; and although the present application has been described in detail with reference to the foregoing embodiments, it should be appreciated by those skilled in the art that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features thereof can be replaced equivalently; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method of enterprise risk assessment, characterized by, The method comprises the following steps: determining two to-be-evaluated enterprises associated with a risk to be evaluated; the two to-be-evaluated enterprises comprise a first to-be-evaluated enterprise and a second to-be-evaluated enterprise; the two to-be-evaluated enterprises are two enterprises to be evaluated for whether they have a risk association; based on an enterprise relationship graph, a first node embedding vector of the first to-be-evaluated enterprise and a second node embedding vector of the second to-be-evaluated enterprise are determined respectively; the nodes in the enterprise relationship graph are enterprise nodes, the edges in the enterprise relationship graph represent the association relationship between enterprises, and the node feature vector of any enterprise node in the enterprise relationship graph is determined based on the enterprise data of the enterprise node; based on the similarity calculation result of the first node embedding vector and the second node embedding vector, a risk association evaluation result of the first to-be-evaluated enterprise and the second to-be-evaluated enterprise is determined; wherein the first node embedding vector is determined based on the node feature vector of the first to-be-evaluated enterprise and the node feature vectors of each first neighbor node in the first neighbor node set, each first neighbor node in the first neighbor node set is an enterprise node in the enterprise relationship graph that has an association relationship with the first to-be-evaluated enterprise; the second node embedding vector is determined based on the node feature vector of the second to-be-evaluated enterprise and the node feature vectors of each second neighbor node in the second neighbor node set, each second neighbor node in the second neighbor node set is an enterprise node in the enterprise relationship graph that has an association relationship with the second to-be-evaluated enterprise; the first node embedding vector of the first to-be-evaluated enterprise and the second node embedding vector of the second to-be-evaluated enterprise are determined based on the enterprise relationship graph, comprising: based on the enterprise relationship graph, the number of association layers of the first to-be-evaluated enterprise and the second to-be-evaluated enterprise is determined; the number of association layers is the total number of edges of the shortest association path between the first to-be-evaluated enterprise and the second to-be-evaluated enterprise in the enterprise relationship graph; the shortest association path is the association path with the shortest number of edges in the multiple association paths between the node of the first to-be-evaluated enterprise and the node of the second to-be-evaluated enterprise; based on the number of association layers, the first node embedding vector of the first to-be-evaluated enterprise and the second node embedding vector of the second to-be-evaluated enterprise are determined respectively; wherein, if the number of association layers is equal to 1, the first node embedding vector is the node feature vector of the first to-be-evaluated enterprise, and the second node embedding vector is the node feature vector of the second to-be-evaluated enterprise. If the number of associated layers is greater than 1, the first node embedding vector is determined based on a node feature vector of the first to-be-evaluated enterprise and a neighbor aggregation vector of a first neighbor node set, the neighbor aggregation vector of the first neighbor node set being determined based on node feature vectors of first neighbor nodes in the first neighbor node set, and the second node embedding vector is determined based on a node feature vector of the second to-be-evaluated enterprise and a neighbor aggregation vector of a second neighbor node set, the neighbor aggregation vector of the second neighbor node set being determined based on node feature vectors of second neighbor nodes in the second neighbor node set; If the number of associated layers is greater than 1, the number of associated layers is K, and the first node embedding vector is determined in the following manner: aggregating node embedding vectors of the first neighbor nodes in the first neighbor node set at the K-1th layer to obtain a neighbor aggregation vector of the first neighbor node set; aggregating the neighbor aggregation vector of the first neighbor node set and the node embedding vector of the first to-be-evaluated enterprise at the K-1th layer to obtain the first node embedding vector of the first to-be-evaluated enterprise at the Kth layer; wherein the node embedding vector of any first neighbor node at the first layer is a node feature vector of the first neighbor node, and the node embedding vector of the first to-be-evaluated enterprise at the first layer is a node feature vector of the first to-be-evaluated enterprise.
2. The method of enterprise risk assessment of claim 1, wherein, The aggregating the neighbor aggregation vector of the first neighbor node set and the node embedding vector of the first to-be-evaluated enterprise at the K-1th layer to obtain the first node embedding vector of the first to-be-evaluated enterprise at the Kth layer comprises: aggregating the neighbor aggregation vector of the first neighbor node set and the node embedding vector of the first to-be-evaluated enterprise at the K-1th layer to obtain a node fusion vector of the first to-be-evaluated enterprise at the Kth layer; inputting the node fusion vector into a feature extraction layer to obtain a node extraction vector output by the feature extraction layer; inputting the node extraction vector into a nonlinear activation function layer to obtain the first node embedding vector output by the nonlinear activation function layer.
3. The method of enterprise risk assessment of claim 1, wherein, The first neighbor node set is determined in the following manner: sampling a neighbor node set of the first to-be-evaluated enterprise at the Kth layer to obtain the first neighbor node set of the first to-be-evaluated enterprise at the Kth layer; wherein the neighbor node set of the first to-be-evaluated enterprise at the Kth layer comprises all direct neighbor nodes of each neighbor node in a neighbor node set of the first to-be-evaluated enterprise at the K-1th layer, and the neighbor node set of the first to-be-evaluated enterprise at the first layer comprises all direct neighbor nodes of the first to-be-evaluated enterprise.
4. The business risk assessment method according to any one of claims 1 to 3, characterized in that, The determining the risk association evaluation result of the first to-be-evaluated enterprise and the second to-be-evaluated enterprise based on the similarity calculation result of the first node embedding vector and the second node embedding vector comprises: determine the risk association assessment result based on a comparison result of the similarity calculation result and a preset similarity threshold; the risk association assessment result includes a first risk association result and a second risk association result, and a risk association degree of the first risk association result is greater than a risk association degree of the second risk association result; after determining the risk association assessment result of the first to-be-evaluated enterprise and the second to-be-evaluated enterprise based on the similarity calculation result of the first node embedding vector and the second node embedding vector, the method further includes: in a case where the risk association assessment result is the first risk association result, obtaining enterprise registration information; in a case where it is determined based on the enterprise registration information that the first to-be-evaluated enterprise and the second to-be-evaluated enterprise have associated records, determining that the first to-be-evaluated enterprise and the second to-be-evaluated enterprise have risk association; in a case where it is determined based on the enterprise registration information that the first to-be-evaluated enterprise and the second to-be-evaluated enterprise have no associated records, triggering a verification process; the verification process is used to verify whether the first to-be-evaluated enterprise and the second to-be-evaluated enterprise have risk association.
5. The business risk assessment method according to any one of claims 1 to 3, characterized in that, The enterprise relationship graph is determined based on the following manner: obtaining enterprise data of a plurality of enterprises; the enterprise data includes multidimensional data, and sources of the enterprise data include a plurality of different databases; constructing an enterprise relationship graph based on the enterprise data of the plurality of enterprises; wherein an enterprise node in the enterprise relationship graph is represented by a node feature vector, the node feature vector is a multidimensional feature vector, and the multidimensional feature vector is determined based on real-time enterprise data.
6. The business risk assessment method according to any one of claims 1 to 3, wherein, The enterprise risk assessment method further includes: in a case where it is determined that there is a risk enterprise, determining a to-be-risk-evaluated enterprise having an associated relationship with the risk enterprise based on the enterprise relationship graph; determining whether the to-be-risk-evaluated enterprise is a risk enterprise based on an association degree coefficient of the risk enterprise and the to-be-risk-evaluated enterprise; wherein the association degree coefficient is obtained by multiplying relationship strength coefficients of edges of a shortest association path between the risk enterprise and the to-be-risk-evaluated enterprise, the relationship strength coefficient is used to represent an association relationship strength between enterprises, and the relationship strength coefficient is greater than 0 and less than or equal to 1.
7. The method of enterprise risk assessment of claim 6, wherein, The relationship strength coefficient of any edge is determined based on the following manner: determining relationship strength sub-coefficients of each association relationship category of the edge; the relationship strength sub-coefficient of any association relationship category is used to represent an association relationship strength between enterprises in the association relationship category; weighting and aggregating each relationship strength sub-coefficient based on weights of the association relationship categories to obtain the relationship strength coefficient of the edge.
8. An enterprise risk assessment apparatus, characterized by comprising: includes: an enterprise determination module configured to determine two to-be-evaluated enterprises for risk association; the two to-be-evaluated enterprises include a first to-be-evaluated enterprise and a second to-be-evaluated enterprise; the two to-be-evaluated enterprises are two enterprises to be evaluated for risk association determine a first node embedding vector of the first to-be-evaluated enterprise and a second node embedding vector of the second to-be-evaluated enterprise based on the enterprise relationship graph; the nodes in the enterprise relationship graph are enterprise nodes, the edges in the enterprise relationship graph are used to represent the association relationship between enterprises, and the node feature vector of any enterprise node in the enterprise relationship graph is determined based on enterprise data of the enterprise node; an association evaluation module is configured to determine a risk association evaluation result of the first to-be-evaluated enterprise and the second to-be-evaluated enterprise based on a similarity calculation result of the first node embedding vector and the second node embedding vector; wherein the first node embedding vector is determined based on the node feature vector of the first to-be-evaluated enterprise and the node feature vectors of each first neighbor node in a first neighbor node set, each first neighbor node in the first neighbor node set is an enterprise node in the enterprise relationship graph that has an association relationship with the first to-be-evaluated enterprise; and the second node embedding vector is determined based on the node feature vector of the second to-be-evaluated enterprise and the node feature vectors of each second neighbor node in a second neighbor node set, each second neighbor node in the second neighbor node set is an enterprise node in the enterprise relationship graph that has an association relationship with the second to-be-evaluated enterprise; the first node embedding vector of the first to-be-evaluated enterprise and the second node embedding vector of the second to-be-evaluated enterprise are determined based on the enterprise relationship graph, including: determining an association layer number of the first to-be-evaluated enterprise and the second to-be-evaluated enterprise based on the enterprise relationship graph; the association layer number is the total number of edges of the shortest association path of the first to-be-evaluated enterprise and the second to-be-evaluated enterprise in the enterprise relationship graph; the shortest association path is an association path with the shortest number of edges in the association paths between the node of the first to-be-evaluated enterprise and the node of the second to-be-evaluated enterprise; determining the first node embedding vector of the first to-be-evaluated enterprise and the second node embedding vector of the second to-be-evaluated enterprise based on the association layer number; wherein if the association layer number is equal to 1, the first node embedding vector is the node feature vector of the first to-be-evaluated enterprise, and the second node embedding vector is the node feature vector of the second to-be-evaluated enterprise; if the association layer number is greater than 1, the first node embedding vector is determined based on the node feature vector of the first to-be-evaluated enterprise and the neighbor aggregation vector of the first neighbor node set, the neighbor aggregation vector of the first neighbor node set is determined based on the node feature vectors of each first neighbor node in the first neighbor node set, the second node embedding vector is determined based on the node feature vector of the second to-be-evaluated enterprise and the neighbor aggregation vector of the second neighbor node set, and the neighbor aggregation vector of the second neighbor node set is determined based on the node feature vectors of each second neighbor node in the second neighbor node set; if the association layer number is greater than 1 and the association layer number is K, the first node embedding vector is determined based on the following manner: aggregating the node embedding vectors of each first neighbor node in the first neighbor node set at the K-1th layer to obtain a neighbor aggregation vector of the first neighbor node set; aggregating the neighbor aggregation vector of the first neighbor node set and the node embedding vector of the first to-be-evaluated enterprise at the K-1th layer to obtain a first node embedding vector of the first to-be-evaluated enterprise at the Kth layer; wherein the node embedding vector of any first neighbor node at the first layer is a node feature vector of the first neighbor node, and the node embedding vector of the first to-be-evaluated enterprise at the first layer is a node feature vector of the first to-be-evaluated enterprise.
9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, The processor executes the computer program to implement the enterprise risk evaluation method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the enterprise risk evaluation method according to any one of claims 1 to 7.
11. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the enterprise risk evaluation method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Enterprise evaluation method, device and equipment
CN113112186A
Risk identification method and device, computer equipment and storage medium
CN116503156A