Enterprise risk assessment method and device, equipment, storage medium and program product
By building an enterprise relationship map and embedding vector calculation based on node characteristics and neighbor node characteristics, the problem of low accuracy of enterprise risk assessment is solved, and more accurate risk correlation assessment is achieved.
Patent Information
- Application Number
- CN202510936101.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-07-08
AI Technical Summary
In the prior art, the accuracy of enterprise risk assessment is not high, and the risk correlation between two companies to be evaluated cannot be effectively determined, resulting in inaccurate evaluation results.
By constructing an enterprise relationship map, the node embedding vectors between enterprises are calculated based on the node feature vectors of enterprise nodes and the feature vectors of neighbor nodes, the risk association evaluation results are determined using the similarity calculation results, and the characteristics of the enterprise itself and the associated enterprises are considered.
It improves the accuracy of corporate risk assessment, avoids misjudgment when corporate data is similar but the relationship between the relationship is not strong, and improves the accuracy of risk correlation assessment.
Smart Images

Figure CN120430637A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of risk assessment technology, and in particular to an enterprise risk assessment method, apparatus, device, storage medium and program product. Background Art
[0002] Enterprise risk assessment involves evaluating the risks of the enterprise under assessment. Traditional enterprise risk assessment methods directly analyze the enterprise's data to determine whether it is a risky enterprise. This method requires data analysis for each enterprise under assessment. To improve the efficiency of enterprise risk assessment, the risk correlation between two related enterprises under assessment can be evaluated first. This allows only data analysis of one of the enterprises under assessment to be performed separately. If the enterprise under assessment is a risky enterprise, the other risk-related enterprises can be quickly identified as risky enterprises without the need for separate data analysis. Therefore, determining the risk correlation assessment results of two enterprises under assessment is an urgent need.
[0003] Currently, risk correlation assessments of two companies are determined based on the similarity of their data. However, the accuracy of risk correlation assessments determined by existing technologies is low. For example, while the data of two companies under evaluation may be highly similar, the actual correlation between the two companies is not strong. Therefore, improving the accuracy of risk correlation assessments and, therefore, the accuracy of risk assessments, is a pressing technical issue. Summary of the Invention
[0004] The present invention provides an enterprise risk assessment method, device, equipment, storage medium and program product to solve the defect of low accuracy of enterprise risk assessment in the prior art and realize a high-accuracy enterprise risk assessment solution.
[0005] The present invention provides an enterprise risk assessment method, comprising: Determine two enterprises to be assessed that are associated with the risks to be assessed; the two enterprises to be assessed include a first enterprise to be assessed and a second enterprise to be assessed; Based on the enterprise relationship graph, determining a first node embedding vector of the first enterprise to be evaluated and a second node embedding vector of the second enterprise to be evaluated; the nodes in the enterprise relationship graph are enterprise nodes, the edges in the enterprise relationship graph are used to represent the association relationship between enterprises, and the node feature vector of any enterprise node in the enterprise relationship graph is determined based on the enterprise data of the enterprise node; Determining a risk association assessment result between the first enterprise to be assessed and the second enterprise to be assessed based on a similarity calculation result between the first node embedding vector and the second node embedding vector; Among them, the first node embedding vector is determined based on the node feature vector of the first enterprise to be evaluated and the node feature vectors of each first neighbor node in the first neighbor node set, and each first neighbor node in the first neighbor node set is an enterprise node that has an association relationship with the first enterprise to be evaluated in the enterprise relationship graph; the second node embedding vector is determined based on the node feature vector of the second enterprise to be evaluated and the node feature vectors of each second neighbor node in the second neighbor node set, and each second neighbor node in the second neighbor node set is an enterprise node that has an association relationship with the second enterprise to be evaluated in the enterprise relationship graph.
[0006] According to an enterprise risk assessment method provided by the present invention, determining a first node embedding vector of the first enterprise to be assessed and a second node embedding vector of the second enterprise to be assessed based on the enterprise relationship graph respectively includes: Determine, based on the enterprise relationship graph, the number of association layers between the first enterprise to be evaluated and the second enterprise to be evaluated; the number of association layers is the total number of edges in the shortest association path between the first enterprise to be evaluated and the second enterprise to be evaluated in the enterprise relationship graph; Determining a first node embedding vector of the first enterprise to be evaluated and a second node embedding vector of the second enterprise to be evaluated based on the number of association layers; Wherein, if the number of association layers is equal to 1, the first node embedding vector is the node feature vector of the first enterprise to be evaluated, and the second node embedding vector is the node feature vector of the second enterprise to be evaluated; If the number of association layers is greater than 1, the first node embedding vector is determined based on the node feature vector of the first enterprise to be evaluated and the neighbor aggregation vector of the first neighbor node set, the neighbor aggregation vector of the first neighbor node set is determined based on the node feature vector of each first neighbor node in the first neighbor node set, the second node embedding vector is determined based on the node feature vector of the second enterprise to be evaluated and the neighbor aggregation vector of the second neighbor node set, and the neighbor aggregation vector of the second neighbor node set is determined based on the node feature vector of each second neighbor node in the second neighbor node set.
[0007] According to an enterprise risk assessment method provided by the present invention, if the number of association layers is greater than 1 and the number of association layers is K, the first node embedding vector is determined based on the following method: Aggregating the node embedding vectors of each first neighbor node in the first neighbor node set at the K-1 layer to obtain a neighbor aggregation vector of the first neighbor node set; Aggregating the neighbor aggregation vectors of the first neighbor node set and the node embedding vector of the first enterprise to be evaluated at the K-1 layer to obtain a first node embedding vector of the first enterprise to be evaluated at the K layer; Among them, the node embedding vector of any first neighbor node in the first layer is the node feature vector of the first neighbor node; the node embedding vector of the first enterprise to be evaluated in the first layer is the node feature vector of the first enterprise to be evaluated.
[0008] According to an enterprise risk assessment method provided by the present invention, aggregating the neighbor aggregation vector of the first neighbor node set and the node embedding vector of the first enterprise to be assessed at the K-1 layer to obtain the first node embedding vector of the first enterprise to be assessed at the K layer includes: Aggregating the neighbor aggregation vectors of the first neighbor node set and the node embedding vector of the first enterprise to be evaluated at the K-1 layer to obtain the node embedding vector of the first enterprise to be evaluated at the K layer; Integrating the node into a vector input to a feature extraction layer to obtain a node extraction vector output by the feature extraction layer; The node extraction vector is input into a nonlinear activation function layer to obtain the first node embedding vector output by the nonlinear activation function layer.
[0009] According to an enterprise risk assessment method provided by the present invention, the first neighbor node set is determined based on the following method: Sampling the neighbor node set of the first enterprise to be evaluated at the Kth layer to obtain the first neighbor node set of the first enterprise to be evaluated at the Kth layer; Among them, the neighbor node set of the first enterprise to be evaluated at the Kth layer includes: all direct neighbor nodes of each neighbor node in the neighbor node set of the first enterprise to be evaluated at the K-1th layer; the neighbor node set of the first enterprise to be evaluated at the first layer includes: all direct neighbor nodes of the first enterprise to be evaluated.
[0010] According to an enterprise risk assessment method provided by the present invention, determining a risk association assessment result between the first enterprise to be assessed and the second enterprise to be assessed based on a similarity calculation result between the first node embedding vector and the second node embedding vector includes: Determining the risk association assessment result based on a comparison result of the similarity calculation result and a preset similarity threshold; the risk association assessment result includes a first risk association result and a second risk association result, and the risk association degree of the first risk association result is greater than the risk association degree of the second risk association result; After determining the risk association assessment result between the first to-be-assessed enterprise and the second to-be-assessed enterprise based on the similarity calculation result between the first node embedding vector and the second node embedding vector, the method further includes: When the risk association assessment result is the first risk association result, obtaining enterprise registration information; If it is determined based on the enterprise registration information that the first enterprise to be assessed has an association record with the second enterprise to be assessed, determining that there is a risk association between the first enterprise to be assessed and the second enterprise to be assessed; If it is determined based on the enterprise registration information that the first enterprise to be evaluated has no association record with the second enterprise to be evaluated, a verification process is triggered; the verification process is used to verify whether there is a risk association between the first enterprise to be evaluated and the second enterprise to be evaluated.
[0011] According to an enterprise risk assessment method provided by the present invention, the enterprise relationship map is determined based on the following method: Acquire enterprise data of multiple enterprises; the enterprise data includes multi-dimensional data, and the sources of the enterprise data include multiple different databases; Building an enterprise relationship map based on the enterprise data of the plurality of enterprises; Among them, the enterprise nodes in the enterprise relationship map are represented by node feature vectors, and the node feature vectors are multidimensional feature vectors, which are determined based on real-time enterprise data.
[0012] According to an enterprise risk assessment method provided by the present invention, the enterprise risk assessment method further includes: In the case where it is determined that there is a risk enterprise, determining enterprises to be risk-assessed that have a relationship with the risk enterprise based on the enterprise relationship map; Determining whether the enterprise to be risk-assessed is a risky enterprise based on a correlation coefficient between the risk enterprise and the enterprise to be risk-assessed; Among them, the correlation degree coefficient is obtained by multiplying the relationship strength coefficients of each edge of the shortest correlation path between the risk enterprise and the enterprise to be risk assessed. The relationship strength coefficient is used to characterize the strength of the correlation relationship between enterprises. The relationship strength coefficient is greater than 0 and less than or equal to 1.
[0013] According to an enterprise risk assessment method provided by the present invention, the relationship strength coefficient on any side is determined based on the following method: Determine the relationship strength sub-coefficient of each association relationship category of the edge; the relationship strength sub-coefficient of any association relationship category is used to represent the strength of the association relationship between enterprises related to the association relationship category; Based on the weights of the association relationship categories, each of the relationship strength sub-coefficients is weightedly aggregated to obtain the relationship strength coefficient of the edge.
[0014] The present invention also provides an enterprise risk assessment device, comprising: An enterprise determination module is used to determine two enterprises to be assessed that are associated with the risks to be assessed; the two enterprises to be assessed include a first enterprise to be assessed and a second enterprise to be assessed; a vector determination module, configured to determine, based on an enterprise relationship graph, a first node embedding vector for the first enterprise to be evaluated and a second node embedding vector for the second enterprise to be evaluated; wherein the nodes in the enterprise relationship graph are enterprise nodes, the edges in the enterprise relationship graph are used to represent the association relationships between enterprises, and the node feature vector of any enterprise node in the enterprise relationship graph is determined based on the enterprise data of the enterprise node; an association assessment module, configured to determine a risk association assessment result between the first enterprise to be assessed and the second enterprise to be assessed based on a similarity calculation result between the first node embedding vector and the second node embedding vector; Among them, the first node embedding vector is determined based on the node feature vector of the first enterprise to be evaluated and the node feature vectors of each first neighbor node in the first neighbor node set, and each first neighbor node in the first neighbor node set is an enterprise node that has an association relationship with the first enterprise to be evaluated in the enterprise relationship graph; the second node embedding vector is determined based on the node feature vector of the second enterprise to be evaluated and the node feature vectors of each second neighbor node in the second neighbor node set, and each second neighbor node in the second neighbor node set is an enterprise node that has an association relationship with the second enterprise to be evaluated in the enterprise relationship graph.
[0015] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, any of the enterprise risk assessment methods described above is implemented.
[0016] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described enterprise risk assessment methods.
[0017] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned enterprise risk assessment methods.
[0018] The enterprise risk assessment method, apparatus, equipment, storage medium and program product provided by the present invention determine two enterprises to be assessed that are associated with the risks to be assessed, the two enterprises to be assessed including a first enterprise to be assessed and a second enterprise to be assessed, and respectively determine a first node embedding vector of the first enterprise to be assessed and a second node embedding vector of the second enterprise to be assessed based on an enterprise relationship graph, wherein the nodes in the enterprise relationship graph are enterprise nodes, the edges in the enterprise relationship graph are used to characterize the association relationship between enterprises, the node feature vector of any enterprise node in the enterprise relationship graph is determined based on the enterprise data of the enterprise node, and the first node embedding vector is determined based on the node feature vector of the first enterprise to be assessed and the node feature vectors of each first neighbor node in the first neighbor node set, wherein each first neighbor node in the first neighbor node set is an enterprise node in the enterprise relationship graph that has an association relationship with the first enterprise to be assessed, and the second node embedding vector is determined based on the node feature vector of the second enterprise to be assessed. The first node embedding vector and the second node embedding vector are determined by the node feature vector of the first node and the node feature vector of each second neighbor node in the second neighbor node set. Each second neighbor node in the second neighbor node set is an enterprise node that has an association relationship with the second enterprise to be evaluated in the enterprise relationship graph. Based on this, the first node embedding vector and the second node embedding vector are determined not only based on their own node feature vectors, but also based on the feature vectors of their neighbor nodes. Therefore, not only the characteristics of the enterprise to be evaluated itself are considered, but also the characteristics of the enterprises associated with the enterprise to be evaluated. Furthermore, based on the similarity calculation results determined by the first node embedding vector and the second node embedding vector, it can also be characterized whether the characteristics of the enterprises associated with the two enterprises to be evaluated are similar, thereby avoiding the situation where the enterprise data of the two enterprises to be evaluated have a high similarity, but in fact the association relationship between the two enterprises to be evaluated is not strong, thereby improving the accuracy of the determined risk association assessment results, that is, improving the accuracy of enterprise risk association assessment, and ultimately improving the accuracy of enterprise risk assessment. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0020] Figure 1 This is one of the flow charts of the enterprise risk assessment method provided by the present invention.
[0021] Figure 2 This is the second flow chart of the enterprise risk assessment method provided by the present invention.
[0022] Figure 3 This is the third flow chart of the enterprise risk assessment method provided by the present invention.
[0023] Figure 4 It is a structural diagram of the enterprise risk assessment device provided by the present invention.
[0024] Figure 5 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0025] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0026] The present invention proposes the following embodiments. Figure 1-Figure 3 The enterprise risk assessment method of the present invention is described.
[0027] Figure 1 This is one of the flow charts of the enterprise risk assessment method provided by the present invention, such as Figure 1 As shown, the enterprise risk assessment method includes the following steps 110, 120 and 130.
[0028] Step 110: Determine two companies to be assessed that are associated with the risks to be assessed.
[0029] The two enterprises to be evaluated include a first enterprise to be evaluated and a second enterprise to be evaluated.
[0030] Here, the two companies to be assessed are two companies to be assessed for risk correlation.
[0031] For example, assuming that the first enterprise to be evaluated is a risk enterprise, if the first enterprise to be evaluated has a risk correlation with the second enterprise to be evaluated, then the second enterprise to be evaluated is also a risk enterprise. Of course, the number of second enterprises to be evaluated can be multiple. Based on this, only one enterprise to be evaluated needs to be determined as a risk enterprise, and multiple other enterprises can be quickly determined as risk enterprises, thereby improving the efficiency of enterprise risk assessment.
[0032] Step 120: Based on the enterprise relationship graph, determine a first node embedding vector of the first enterprise to be evaluated and a second node embedding vector of the second enterprise to be evaluated.
[0033] Among them, the nodes in the enterprise relationship graph are enterprise nodes, the edges in the enterprise relationship graph are used to represent the association relationship between enterprises, and the node feature vector of any enterprise node in the enterprise relationship graph is determined based on the enterprise data of the enterprise node.
[0034] Here, the enterprise relationship map is constructed based on the enterprise data of multiple enterprises. Specifically, the enterprise relationship map is obtained by performing graph calculation based on the enterprise data of multiple enterprises. The multiple enterprises should include a first enterprise to be evaluated and a second enterprise to be evaluated. Furthermore, the enterprise data includes multi-dimensional data, thereby improving the accuracy of constructing the enterprise relationship map, thereby improving the accuracy of assessing enterprise risk associations, and ultimately improving the accuracy of assessing enterprise risks. Furthermore, the source of enterprise data includes multiple different databases to expand the source of enterprise data, thereby improving the accuracy of constructing the enterprise relationship map, thereby improving the accuracy of assessing enterprise risk associations, and ultimately improving the accuracy of assessing enterprise risks. Furthermore, the enterprise data is real-time enterprise data, thereby ensuring that the enterprise relationship map is also real-time, that is, improving the accuracy of constructing the enterprise relationship map, thereby improving the accuracy of assessing enterprise risk associations, and ultimately improving the accuracy of assessing enterprise risks.
[0035] Here, enterprise nodes are represented by node feature vectors. If enterprise data includes multidimensional data, the node feature vectors are multidimensional feature vectors. Edges in the enterprise relationship graph are represented by heterogeneous edges. Furthermore, based on the enterprise relationship network modeling solution for heterogeneous graphs, the enterprise relationship graph is constructed; specifically, node modeling is performed first, followed by edge relationship modeling.
[0036] Here, the term "affiliated relationship" may include, but is not limited to, at least one of the following: equity control, supply chain, cross-executive appointments, joint investment, intellectual property, mutual guarantee, bank-enterprise loan, and judicial relationships. The relationship between two companies can include multiple relationships, meaning multiple relationships exist simultaneously. For example, if the direct shareholding ratio is ≥30% or the indirect holding level is ≤3, then the two companies have an equity control relationship; if the annual order transaction volume is greater than RMB 10 million and the duration of the cooperation is ≥2 years, then the two companies have an equity control relationship; if there is overlap in the positions of current directors / supervisors / management personnel and the tenure overlap is ≥6 months, then the two companies have a cross-executive appointment relationship; if the total capital contribution to the joint equity investment in the same entity is greater than RMB 5 million, then the two companies have a joint investment relationship; if the number of joint patent applications within three years is ≥5 or the number of joint authorships on technical standards is ≥5, then the two companies have an intellectual property relationship; if there is a mutual guarantee agreement with a single guarantee amount exceeding 10% of the net assets, then the two companies have a mutual guarantee relationship.
[0037] Among them, the first node embedding vector is determined based on the node feature vector of the first enterprise to be evaluated and the node feature vectors of each first neighbor node in the first neighbor node set, and each first neighbor node in the first neighbor node set is an enterprise node that has an association relationship with the first enterprise to be evaluated in the enterprise relationship graph; the second node embedding vector is determined based on the node feature vector of the second enterprise to be evaluated and the node feature vectors of each second neighbor node in the second neighbor node set, and each second neighbor node in the second neighbor node set is an enterprise node that has an association relationship with the second enterprise to be evaluated in the enterprise relationship graph.
[0038] Specifically, feature aggregation can be performed on the node feature vector of the first enterprise to be evaluated and the node feature vectors of each first neighbor node in the first neighbor node set to obtain a first node embedding vector. Feature aggregation can be performed on the node feature vector of the second enterprise to be evaluated and the node feature vectors of each second neighbor node in the second neighbor node set to obtain a second node embedding vector.
[0039] Since enterprise nodes in the enterprise relationship graph are represented by node feature vectors, the node feature vector of the first enterprise to be evaluated and the node feature vector of the second enterprise to be evaluated can be determined based on the enterprise relationship graph.
[0040] Here, each first neighbor node in the first neighbor node set may have a direct or indirect association with the node corresponding to the first enterprise to be evaluated; each second neighbor node in the second neighbor node set may have a direct or indirect association with the node corresponding to the second enterprise to be evaluated. For example, a direct association indicates that the two nodes are directly connected by an edge, while an indirect association indicates that the two nodes are not directly connected by an edge, that is, there are other nodes between them. The first neighbor node set may include all neighbor nodes associated with the first enterprise to be evaluated, or only some of them. If a set number of association levels is set, neighbor nodes with association levels exceeding the preset number of association levels are not included in the first neighbor node set. Similarly, the second neighbor node set may include all neighbor nodes associated with the second enterprise to be evaluated, or only some of them. If a set number of association levels is set, neighbor nodes with association levels exceeding the preset number of association levels are not included in the second neighbor node set.
[0041] It should be understood that the first node embedding vector and the second node embedding vector are determined not only based on their own node feature vectors, but also based on the feature vectors of their neighboring nodes, so that not only the characteristics of the enterprise to be evaluated itself are taken into account, but also the characteristics of the enterprises associated with the enterprise to be evaluated. Furthermore, the similarity calculation results determined based on the first node embedding vector and the second node embedding vector can also characterize whether the characteristics of the enterprises associated with the two enterprises to be evaluated are similar, thereby avoiding the situation where the enterprise data of the two enterprises to be evaluated have high similarity, but in fact the correlation between the two enterprises to be evaluated is not strong, thereby improving the accuracy of the subsequently determined risk association assessment results, that is, improving the accuracy of enterprise risk association assessment, and ultimately improving the accuracy of enterprise risk assessment.
[0042] Furthermore, all neighbor nodes of the first enterprise to be evaluated can be sampled to obtain a first neighbor node set, thereby reducing the number of neighbor nodes in the first neighbor node set, thereby reducing the amount of calculation, improving the efficiency of enterprise risk assessment, and ensuring that the first node embedding vector is accurately obtained, thereby further improving the accuracy of the subsequent risk association assessment results, that is, further improving the accuracy of enterprise risk association assessment, and ultimately further improving the accuracy of enterprise risk assessment.
[0043] Furthermore, all neighbor nodes of the second enterprise to be evaluated can be sampled to obtain a second neighbor node set, thereby reducing the number of neighbor nodes in the second neighbor node set, thereby reducing the amount of calculation, improving the efficiency of enterprise risk assessment, and ensuring that the second node embedding vector is accurately obtained, thereby further improving the accuracy of the subsequent risk association assessment results, that is, further improving the accuracy of enterprise risk association assessment, and ultimately further improving the accuracy of enterprise risk assessment.
[0044] Step 130: Determine a risk association assessment result of the first enterprise to be assessed and the second enterprise to be assessed based on a similarity calculation result between the first node embedding vector and the second node embedding vector.
[0045] Here, the calculation method of the similarity calculation result can be set according to actual needs, for example, a cosine similarity calculation method.
[0046] In one specific embodiment, a risk association assessment result is determined based on a comparison between the similarity calculation result and a preset similarity threshold. The risk association assessment result includes a first risk association result and a second risk association result, wherein the risk association degree of the first risk association result is greater than the risk association degree of the second risk association result. Furthermore, the first risk association result indicates that the first enterprise to be assessed and the second enterprise to be assessed have a risk association, while the second risk association result indicates that the first enterprise to be assessed and the second enterprise to be assessed do not have a risk association.
[0047] The enterprise risk assessment method provided by an embodiment of the present invention determines two enterprises to be assessed that are associated with the risks to be assessed, the two enterprises to be assessed including a first enterprise to be assessed and a second enterprise to be assessed, and determines a first node embedding vector of the first enterprise to be assessed and a second node embedding vector of the second enterprise to be assessed based on an enterprise relationship graph, respectively. The nodes in the enterprise relationship graph are enterprise nodes, and the edges in the enterprise relationship graph are used to characterize the association relationship between enterprises. The node feature vector of any enterprise node in the enterprise relationship graph is determined based on the enterprise data of the enterprise node, and the first node embedding vector is determined based on the node feature vector of the first enterprise to be assessed and the node feature vectors of each first neighbor node in the first neighbor node set, and each first neighbor node in the first neighbor node set is an enterprise node in the enterprise relationship graph that has an association relationship with the first enterprise to be assessed, and the second node embedding vector is determined based on the node feature vector of the second enterprise to be assessed. and the node feature vector of each second neighbor node in the second neighbor node set. Each second neighbor node in the second neighbor node set is an enterprise node that has an association relationship with the second enterprise to be evaluated in the enterprise relationship graph. Based on this, the first node embedding vector and the second node embedding vector are determined not only based on their own node feature vectors, but also based on the feature vectors of their neighbor nodes. Therefore, not only the characteristics of the enterprise to be evaluated itself are considered, but also the characteristics of the enterprises associated with the enterprise to be evaluated. Furthermore, based on the similarity calculation results determined by the first node embedding vector and the second node embedding vector, it can also be characterized whether the characteristics of the enterprises associated with the two enterprises to be evaluated are similar, thereby avoiding the situation where the enterprise data of the two enterprises to be evaluated have a high similarity, but in fact the association relationship between the two enterprises to be evaluated is not strong, thereby improving the accuracy of the determined risk association assessment results, that is, improving the accuracy of the enterprise risk association assessment, and ultimately improving the accuracy of the enterprise risk assessment.
[0048] Based on any of the above embodiments, in the method, step 120 includes step 121 and step 122 .
[0049] Step 121: Determine the number of association layers between the first enterprise to be evaluated and the second enterprise to be evaluated based on the enterprise relationship map.
[0050] The number of association levels is the total number of edges in the shortest association path between the first enterprise to be evaluated and the second enterprise to be evaluated in the enterprise relationship graph. Since the nodes in the enterprise relationship graph are enterprise nodes, and the edges in the enterprise relationship graph represent the associations between enterprises, the number of association levels between the two enterprises to be evaluated can be determined based on the enterprise relationship graph.
[0051] Here, the shortest connection path is the connection path with the shortest number of edges among the multiple connection paths between the node of the first enterprise to be evaluated and the node of the second enterprise to be evaluated. For example, if the first enterprise to be evaluated is Enterprise A and the second enterprise to be evaluated is Enterprise B, and there are two connection paths between Enterprise A and Enterprise B: Enterprise A-Enterprise C-Enterprise B and Enterprise A-Enterprise C-Enterprise D-Enterprise B. In this case, the shortest connection path is Enterprise A-Enterprise C-Enterprise B, and the number of connection levels is 2.
[0052] Step 122: Based on the number of association layers, determine a first node embedding vector of the first enterprise to be evaluated and a second node embedding vector of the second enterprise to be evaluated.
[0053] If the number of association levels is 1, the first node embedding vector is the node feature vector of the first enterprise to be evaluated, and the second node embedding vector is the node feature vector of the second enterprise to be evaluated. This means that the first and second enterprises to be evaluated are directly associated, and the characteristics of neighboring nodes do not need to be considered.
[0054] If the number of association levels is greater than 1, the first node embedding vector is determined based on the node feature vector of the first enterprise to be evaluated and the neighbor aggregation vector of the first neighbor node set. The neighbor aggregation vector of the first neighbor node set is determined based on the node feature vectors of each first neighbor node in the first neighbor node set. The second node embedding vector is determined based on the node feature vector of the second enterprise to be evaluated and the neighbor aggregation vector of the second neighbor node set. The neighbor aggregation vector of the second neighbor node set is determined based on the node feature vectors of each second neighbor node in the second neighbor node set. In other words, the first enterprise to be evaluated and the second enterprise to be evaluated are indirectly associated, so the characteristics of the neighbor nodes need to be considered.
[0055] Specifically, feature aggregation can be performed on the node feature vectors of each first neighbor node to obtain a neighbor aggregation vector of the first neighbor node set. Feature aggregation can be performed on the node feature vectors of each second neighbor node to obtain a neighbor aggregation vector of the second neighbor node set.
[0056] The enterprise risk assessment method provided by the embodiment of the present invention determines the number of association layers between the first enterprise to be assessed and the second enterprise to be assessed based on the enterprise relationship map, and the number of association layers is the total number of edges of the shortest association path between the first enterprise to be assessed and the second enterprise to be assessed in the enterprise relationship map, so as to determine the number of association layers based on the shortest association path, and can more accurately determine the risk correlation between the two enterprises to be assessed; at the same time, based on the number of association layers, the first node embedding vector of the first enterprise to be assessed and the second node embedding vector of the second enterprise to be assessed are determined respectively, and if the number of association layers is equal to 1, the first node embedding vector is the node feature vector of the first enterprise to be assessed, and the second node embedding vector is the node feature vector of the second enterprise to be assessed; if the number of association layers is greater than 1, the first node embedding vector is determined based on the node feature vector of the first enterprise to be assessed and the neighbor aggregation vector of the first neighbor node set, and the neighbor aggregation vector of the first neighbor node set is based on the node feature vector of each first neighbor node in the first neighbor node set. The first node embedding vector and the second node embedding vector are determined based on the node feature vector of the second enterprise to be evaluated and the neighbor aggregation vector of the second neighbor node set. The neighbor aggregation vector of the second neighbor node set is determined based on the node feature vector of each second neighbor node in the second neighbor node set. Based on this, the first node embedding vector and the second node embedding vector can be determined more accurately based on the number of association layers, thereby further improving the accuracy of the risk association assessment results, that is, further improving the accuracy of the enterprise risk association assessment, and ultimately further improving the accuracy of the enterprise risk assessment; and the first node embedding vector and the second node embedding vector are determined not only based on their own node feature vectors, but also based on the feature vectors of their neighbor nodes, thereby considering not only the characteristics of the enterprise to be evaluated itself, but also the characteristics of enterprises associated with the enterprise to be evaluated, thereby improving the accuracy of the risk association assessment results, that is, improving the accuracy of the enterprise risk association assessment, and ultimately improving the accuracy of the enterprise risk assessment.
[0057] Based on any of the above embodiments, in this method, if the number of association layers is greater than 1 and the number of association layers is K, the first node embedding vector is determined based on the following method: step 1221 and step 1222.
[0058] Step 1221 : Aggregate the node embedding vectors of the first neighbor nodes in the first neighbor node set at the K-1 layer to obtain a neighbor aggregation vector of the first neighbor node set.
[0059] Here, the node embedding vector of the first neighbor node at the K-1 layer is determined in a manner substantially the same as the first node embedding vector, and is not further described here. Based on this, the node embedding vector of any of the first neighbor nodes at the first layer is the node feature vector of the first neighbor node.
[0060] The aggregation method can be set based on actual conditions and is not specifically limited here. Examples include mean pooling, max pooling, weighted summation, and LSTM (Long Short-Term Memory) aggregation. Specifically, the LSTM aggregation method randomly arranges neighbor nodes to generate a sequence, captures long-range dependencies through a gating mechanism, and uses the final hidden state as the aggregation result.
[0061] It should be understood that the node embedding vector of each first neighbor node at the K-1 layer is obtained by aggregating the node feature vectors of each first neighbor node layer by layer. Therefore, the first node embedding vector is also determined based on the feature vectors of its neighbor nodes, thereby also considering the characteristics of enterprises associated with the enterprise to be assessed, thereby improving the accuracy of the subsequently determined risk association assessment results. Moreover, the node embedding vector of the first neighbor node at the K-1 layer is determined not only based on its own node feature vector, but also based on the feature vectors of its neighbor nodes, thereby considering not only its own characteristics but also the characteristics of the enterprises associated with it, thereby improving the accuracy of the subsequently determined risk association assessment results, that is, improving the accuracy of the enterprise risk association assessment, and ultimately improving the accuracy of the enterprise risk assessment.
[0062] Furthermore, all neighbor nodes of the first enterprise to be evaluated can be sampled to obtain a first neighbor node set, thereby reducing the number of neighbor nodes in the first neighbor node set, thereby reducing the amount of calculation, improving the efficiency of enterprise risk assessment, and ensuring that the first node embedding vector is accurately obtained, thereby further improving the accuracy of the subsequent risk association assessment results, that is, further improving the accuracy of enterprise risk association assessment, and ultimately further improving the accuracy of enterprise risk assessment.
[0063] In one embodiment, to balance computational efficiency and information integrity, a stratified sampling strategy is implemented for the central node corresponding to the first enterprise to be evaluated. Specifically, a predetermined number of nodes are randomly sampled with equal probability from each layer's neighbor set. This effectively controls the adjacency matrix dimension and avoids the memory usage issue caused by full neighborhood computation.
[0064] In another embodiment, the neighbor node set of the first enterprise to be evaluated in the K layer can be sampled to obtain the first neighbor node set of the first enterprise to be evaluated in the K layer, and all direct neighbor nodes of each neighbor node in the neighbor node set of the first enterprise to be evaluated in the K-1 layer, that is, only the neighbor nodes of the K layer are selected, thereby reducing the number of neighbor nodes in the first neighbor node set, thereby reducing the amount of calculation, improving the efficiency of enterprise risk assessment, and ensuring that the first node embedding vector is accurately obtained, thereby further improving the accuracy of the subsequently determined risk association assessment results, that is, further improving the accuracy of enterprise risk association assessment, and ultimately further improving the accuracy of enterprise risk assessment.
[0065] Step 1222: Aggregate the neighbor aggregation vectors of the first neighbor node set and the node embedding vector of the first enterprise to be evaluated at the K-1 layer to obtain the first node embedding vector of the first enterprise to be evaluated at the K layer.
[0066] Here, the node embedding vector of the first enterprise to be evaluated at layer K-1 is determined in substantially the same manner as the first node embedding vector of the first enterprise to be evaluated at layer K. Based on this, the node embedding vector of the first enterprise to be evaluated at the first layer is the node feature vector of the first enterprise to be evaluated.
[0067] Here, the aggregation processing method can be set according to actual conditions and is not specifically limited here, for example, concat (splicing).
[0068] It should be understood that the node embedding vector of the first enterprise to be evaluated at the K-1 layer is obtained by aggregating the node feature vectors of the first enterprise to be evaluated layer by layer. Therefore, the first node embedding vector is determined not only based on its own node feature vector, but also based on the feature vectors of its neighboring nodes. This not only considers the characteristics of the enterprise to be evaluated itself, but also considers the characteristics of enterprises associated with the enterprise to be evaluated. Furthermore, the similarity calculation result determined based on the first node embedding vector and the second node embedding vector can also indicate whether the characteristics of the enterprises associated with the two enterprises to be evaluated are similar. This avoids the situation where the enterprise data of the two enterprises to be evaluated are highly similar, but in fact the relationship between the two enterprises to be evaluated is not strong, thereby improving the accuracy of the determined risk association assessment results, that is, improving the accuracy of the enterprise risk association assessment, and ultimately improving the accuracy of the enterprise risk assessment. Moreover, the node embedding vector of the first enterprise to be evaluated at the K-1 layer is determined not only based on its own node feature vector, but also based on the feature vectors of its neighboring nodes. This not only considers its own characteristics, but also considers the characteristics of the enterprises associated with it. This further improves the accuracy of the subsequently determined risk association assessment results, that is, improving the accuracy of the enterprise risk association assessment, and ultimately improving the accuracy of the enterprise risk assessment. The node embedding vectors of the first enterprise to be evaluated at the K-1 layer are aggregated to retain the evolution trajectory of the node's own characteristics, that is, to fuse the historical status of the node, thereby improving the accuracy of the subsequent risk association assessment results, that is, improving the accuracy of the enterprise's risk association assessment, and ultimately improving the accuracy of the enterprise's risk assessment.
[0069] The enterprise risk assessment method provided by the embodiment of the present invention aggregates the node embedding vectors of each first neighbor node in the first neighbor node set at the K-1 layer to obtain the neighbor aggregation vector of the first neighbor node set, so that the node embedding vector of the first neighbor node at the K-1 layer is determined not only based on its own node feature vector, but also based on the feature vectors of its neighbor nodes, thereby considering not only its own characteristics but also the characteristics of the enterprises associated with it, thereby further improving the accuracy of the subsequently determined risk association assessment results, that is, further improving the accuracy of the enterprise risk association assessment, and ultimately further improving the accuracy of the enterprise risk assessment; aggregate the neighbor aggregation vector of the first neighbor node set and the node embedding vector of the first enterprise to be assessed at the K-1 layer to obtain the node embedding vector of the first enterprise to be assessed at the K-1 layer. The first node embedding vector of the layer is obtained, so that the node embedding vector of the first enterprise to be evaluated in the K-1 layer is determined not only based on its own node feature vector, but also based on the feature vectors of its neighboring nodes, so that not only its own characteristics but also the characteristics of the enterprises associated with it are considered, thereby further improving the accuracy of the risk association assessment results determined subsequently, that is, further improving the accuracy of the enterprise risk association assessment, and ultimately further improving the accuracy of the enterprise risk assessment; and the node embedding vector of the first enterprise to be evaluated in the K-1 layer is aggregated, so as to retain the node's own feature evolution trajectory, that is, to fuse the node's historical state, thereby further improving the accuracy of the risk association assessment results determined subsequently, that is, further improving the accuracy of the enterprise risk association assessment, and ultimately further improving the accuracy of the enterprise risk assessment.
[0070] Based on any of the above embodiments, in the method, the above step 1222 includes: step 12221, step 12222 and step 12223.
[0071] Step 12221: Aggregate the neighbor aggregation vector of the first neighbor node set and the node embedding vector of the first enterprise to be evaluated at the K-1 layer to obtain the node integration vector of the first enterprise to be evaluated at the K layer.
[0072] Step 12222: Input the node integration vector into the feature extraction layer to obtain the node extraction vector output by the feature extraction layer.
[0073] Here, the feature extraction layer is used to further extract features from the node integration vector, that is, to achieve cross-layer feature dimension transformation; the feature extraction layer is a trainable network layer. Furthermore, different numbers of associated layers correspond to different feature extraction layers, that is, the node integration vector is input to the feature extraction layer corresponding to the number of associated layers. For example, the feature extraction layer can use a trainable parameter matrix representation.
[0074] Step 12223: Input the node extraction vector to a nonlinear activation function layer to obtain the first node embedding vector output by the nonlinear activation function layer.
[0075] Here, the nonlinear activation function layer can be constructed by a nonlinear activation function, and the nonlinear activation function can be an activation function such as ReLU or LeakyReLU to enhance the representation ability of the node embedding vector.
[0076] In one specific embodiment, a multi-layer aggregation function is used to iteratively update node embedding vectors, achieving hierarchical extraction of graph structure information. This uses a message passing mechanism, with each layer performing two core operations: neighbor information aggregation and feature nonlinear transformation.
[0077] For example, for the first node embedding vector of the first enterprise to be evaluated at the Kth layer, its aggregation formula is as follows, that is, the multi-layer aggregation function is as follows: ; Where, Indicates the first enterprise to be evaluated The first node embedding vector at layer K is, represents a nonlinear activation function, represents the trainable parameter matrix corresponding to the K-th layer, represents the concatenation function, Indicates the first enterprise to be evaluated The node embedding vector at the K-1th layer, Represents an aggregate function, Indicates the first neighbor node The node embedding vector at the K-1th layer, Indicates the first enterprise to be evaluated The first neighbor node set in the Kth layer. Can realize cross-layer feature dimension transformation; nonlinear activation function The activation function can be ReLU or LeakyReLU to enhance the representation ability of the node embedding vector.
[0078] It should be understood that the network model including the feature extraction layer and the nonlinear activation function layer is trainable, thereby improving the representation ability of the first node embedding vector by training the model, and further improving the accuracy of the subsequently determined risk association assessment results, that is, further improving the accuracy of the enterprise risk association assessment, and ultimately further improving the accuracy of the enterprise risk assessment.
[0079] The enterprise risk assessment method provided by an embodiment of the present invention inputs a node integration vector into a feature extraction layer to obtain a node extraction vector output by the feature extraction layer, and inputs the node extraction vector into a nonlinear activation function layer to obtain a first node embedding vector output by the nonlinear activation function layer, thereby further improving the representation ability of the first node embedding vector through feature extraction and nonlinear activation function, and further improving the accuracy of the subsequently determined risk association assessment results, that is, further improving the accuracy of the enterprise risk association assessment, and ultimately further improving the accuracy of the enterprise risk assessment.
[0080] Based on any of the above embodiments, in this method, the first set of neighbor nodes is determined based on the following method: The neighbor node set of the first enterprise to be evaluated at the Kth layer is sampled to obtain a first neighbor node set of the first enterprise to be evaluated at the Kth layer.
[0081] Among them, the neighbor node set of the first enterprise to be evaluated at the Kth layer includes: all direct neighbor nodes of each neighbor node in the neighbor node set of the first enterprise to be evaluated at the K-1th layer; the neighbor node set of the first enterprise to be evaluated at the first layer includes: all direct neighbor nodes of the first enterprise to be evaluated.
[0082] Here, a direct neighbor node indicates that the relationship between two nodes is a direct relationship, and a direct relationship indicates that the two nodes are directly connected by an edge.
[0083] In one embodiment, the sampling number of each type of association relationship can be determined based on the association relationship between each neighbor node in the neighbor node set of the first enterprise to be evaluated at the K-1 layer and its direct neighbor node, and based on each sampling number, the neighbor node set of the first enterprise to be evaluated at the K layer is sampled. More specifically, a relationship-aware sampling function can be used for sampling, that is, truncation control is performed based on the sampling number to control the number of neighbor nodes of each association relationship. For heterogeneous graph data, different relationship types can be configured with differentiated sampling strategies, such as downsampling high-frequency relationships and oversampling low-frequency relationships, so as to improve the modeling ability of long-tail distributions, thereby further improving the representation ability of the first node embedding vector, and further improving the accuracy of the subsequently determined risk association assessment results, that is, further improving the accuracy of the enterprise risk association assessment, and ultimately further improving the accuracy of the enterprise risk assessment.
[0084] The enterprise risk assessment method provided by an embodiment of the present invention samples the neighbor node set of the first enterprise to be assessed in the Kth layer, and reduces the number of neighbor nodes in the first neighbor node set through sampling and selecting only the neighbor nodes in the Kth layer, thereby reducing the amount of calculation, improving the efficiency of enterprise risk assessment, and ensuring that the first node embedding vector is accurately obtained, thereby further improving the accuracy of the subsequently determined risk association assessment results, that is, further improving the accuracy of the enterprise risk association assessment, and ultimately further improving the accuracy of the enterprise risk assessment.
[0085] Based on any of the above embodiments, in this method, if the number of association layers is greater than 1 and the number of association layers is K, the second node embedding vector is determined based on the following method: step 1223 and step 1224.
[0086] Step 1223: Aggregate the node embedding vectors of the second neighbor nodes in the second neighbor node set at the K-1 layer to obtain a neighbor aggregation vector of the second neighbor node set.
[0087] Here, the node embedding vector of the second neighbor node at the K-1 layer is determined in a manner substantially the same as the second node embedding vector, and is not further described here. Based on this, the node embedding vector of any second neighbor node at the first layer is the node feature vector of the second neighbor node.
[0088] Here, the aggregation processing method can be set according to the actual situation and is not specifically limited here, for example, mean pooling, maximum pooling, weighted summation, and LSTM aggregation, etc.
[0089] It should be understood that the node embedding vector of each second neighbor node at the K-1 layer is obtained by aggregating the node feature vectors of each second neighbor node layer by layer. Therefore, the second node embedding vector is also determined based on the feature vectors of its neighbor nodes, thereby also considering the characteristics of enterprises associated with the enterprise to be assessed, thereby improving the accuracy of the subsequently determined risk association assessment results. Moreover, the node embedding vector of the second neighbor node at the K-1 layer is determined not only based on its own node feature vector, but also based on the feature vectors of its neighbor nodes, thereby considering not only its own characteristics but also the characteristics of the enterprises associated with it, thereby improving the accuracy of the subsequently determined risk association assessment results, that is, improving the accuracy of the enterprise risk association assessment, and ultimately improving the accuracy of the enterprise risk assessment.
[0090] Furthermore, all neighbor nodes of the second enterprise to be evaluated can be sampled to obtain a second neighbor node set, thereby reducing the number of neighbor nodes in the second neighbor node set, thereby reducing the amount of calculation, improving the efficiency of enterprise risk assessment, and ensuring that the second node embedding vector is accurately obtained, thereby further improving the accuracy of the subsequent risk association assessment results, that is, further improving the accuracy of enterprise risk association assessment, and ultimately further improving the accuracy of enterprise risk assessment.
[0091] In one embodiment, to balance computational efficiency and information integrity, a stratified sampling strategy is implemented for the central node corresponding to the second enterprise to be evaluated. Specifically, a predetermined number of nodes are randomly sampled with equal probability from each layer's neighbor set. This effectively controls the adjacency matrix dimension and avoids the memory usage issue caused by full neighborhood computation.
[0092] In another embodiment, the neighbor node set of the second enterprise to be evaluated in the K layer can be sampled to obtain the second neighbor node set of the second enterprise to be evaluated in the K layer, and all direct neighbor nodes of each neighbor node in the neighbor node set of the second enterprise to be evaluated in the K-1 layer, that is, only the neighbor nodes of the K layer are selected, thereby reducing the number of neighbor nodes in the second neighbor node set, thereby reducing the amount of calculation, improving the efficiency of enterprise risk assessment, and ensuring that the second node embedding vector is accurately obtained, thereby further improving the accuracy of the subsequently determined risk association assessment results, that is, further improving the accuracy of the enterprise risk association assessment, and ultimately further improving the accuracy of the enterprise risk assessment.
[0093] Step 1224 , aggregate the neighbor aggregation vector of the second neighbor node set and the node embedding vector of the second enterprise to be evaluated at the K-1 layer to obtain the second node embedding vector of the second enterprise to be evaluated at the K layer.
[0094] Here, the node embedding vector of the second enterprise to be evaluated at layer K-1 is determined in substantially the same manner as the second node embedding vector of the second enterprise to be evaluated at layer K. Based on this, the node embedding vector of the second enterprise to be evaluated at the first layer is the node feature vector of the second enterprise to be evaluated.
[0095] Here, the aggregation processing method can be set according to actual conditions and is not specifically limited here, for example, concat (splicing).
[0096] It should be understood that the node embedding vector of the second enterprise to be evaluated at the K-1 layer is obtained by aggregating the node feature vectors of the second enterprise to be evaluated layer by layer. Therefore, the second node embedding vector is determined not only based on its own node feature vector, but also based on the feature vectors of its neighboring nodes. This not only considers the characteristics of the enterprise to be evaluated itself, but also considers the characteristics of enterprises associated with the enterprise to be evaluated. Furthermore, the similarity calculation result determined based on the first node embedding vector and the second node embedding vector can also indicate whether the characteristics of the enterprises associated with the two enterprises to be evaluated are similar. This avoids the situation where the enterprise data of the two enterprises to be evaluated are highly similar, but in fact the relationship between the two enterprises to be evaluated is not strong, thereby improving the accuracy of the determined risk association assessment results, that is, improving the accuracy of the enterprise risk association assessment, and ultimately improving the accuracy of the enterprise risk assessment. Moreover, the node embedding vector of the second enterprise to be evaluated at the K-1 layer is determined not only based on its own node feature vector, but also based on the feature vectors of its neighboring nodes. This not only considers its own characteristics, but also considers the characteristics of the enterprises associated with it. This further improves the accuracy of the subsequently determined risk association assessment results, that is, improving the accuracy of the enterprise risk association assessment, and ultimately improving the accuracy of the enterprise risk assessment. The node embedding vectors of the second enterprise to be evaluated at the K-1 layer are aggregated to retain the evolution trajectory of the node's own characteristics, that is, to fuse the historical status of the node, thereby improving the accuracy of the subsequent risk association assessment results, that is, improving the accuracy of the enterprise risk association assessment, and ultimately improving the accuracy of the enterprise risk assessment.
[0097] The enterprise risk assessment method provided by the embodiment of the present invention aggregates the node embedding vectors of each second neighbor node in the second neighbor node set at the K-1 layer to obtain the neighbor aggregation vector of the second neighbor node set, so that the node embedding vector of the second neighbor node at the K-1 layer is determined not only based on its own node feature vector, but also based on the feature vector of its neighbor node, thereby considering not only its own characteristics but also the characteristics of the enterprise associated with it, thereby further improving the accuracy of the subsequently determined risk association assessment results, that is, further improving the accuracy of the enterprise risk association assessment, and ultimately further improving the accuracy of the enterprise risk assessment; aggregate the neighbor aggregation vector of the second neighbor node set and the node embedding vector of the second enterprise to be assessed at the K-1 layer to obtain the node embedding vector of the second enterprise to be assessed at the K-1 layer. The second node embedding vector of the layer is obtained, so that the node embedding vector of the second enterprise to be evaluated in the K-1 layer is determined not only based on its own node feature vector, but also based on the feature vector of its neighboring nodes, so that not only its own characteristics but also the characteristics of the enterprises associated with it are considered, thereby further improving the accuracy of the risk association assessment results determined subsequently, that is, further improving the accuracy of the enterprise risk association assessment, and ultimately further improving the accuracy of the enterprise risk assessment; and the node embedding vector of the second enterprise to be evaluated in the K-1 layer is aggregated, so as to retain the node's own feature evolution trajectory, that is, to fuse the node's historical state, thereby further improving the accuracy of the risk association assessment results determined subsequently, that is, further improving the accuracy of the enterprise risk association assessment, and ultimately further improving the accuracy of the enterprise risk assessment.
[0098] Based on any of the above embodiments, in this method, step 1224 includes: step 12241 , step 12242 and step 12243 .
[0099] Step 12241: Aggregate the neighbor aggregation vector of the second neighbor node set and the node embedding vector of the second enterprise to be evaluated at the K-1 layer to obtain the node integration vector of the second enterprise to be evaluated at the K layer.
[0100] Step 12242: Input the node integration vector into the feature extraction layer to obtain the node extraction vector output by the feature extraction layer.
[0101] Here, the feature extraction layer is used to further extract features from the node integration vector, that is, to achieve cross-layer feature dimension transformation; the feature extraction layer is a trainable network layer. Furthermore, different numbers of associated layers correspond to different feature extraction layers, that is, the node integration vector is input to the feature extraction layer corresponding to the number of associated layers. For example, the feature extraction layer can use a trainable parameter matrix representation.
[0102] Step 12243: input the node extraction vector into a nonlinear activation function layer to obtain the second node embedding vector output by the nonlinear activation function layer.
[0103] Here, the nonlinear activation function layer can be constructed by a nonlinear activation function, and the nonlinear activation function can be an activation function such as ReLU or LeakyReLU to enhance the representation ability of the node embedding vector.
[0104] In one specific embodiment, a multi-layer aggregation function is used to iteratively update node embedding vectors, achieving hierarchical extraction of graph structure information. This uses a message passing mechanism, with each layer performing two core operations: neighbor information aggregation and feature nonlinear transformation.
[0105] For example, for the second node embedding vector of the second enterprise to be evaluated at the Kth layer, its aggregation formula is as follows, that is, the multi-layer aggregation function is as follows: ; Where, Indicates the second enterprise to be evaluated The second node embedding vector at layer K is, represents a nonlinear activation function, represents the trainable parameter matrix corresponding to the K-th layer, represents the concatenation function, Indicates the second enterprise to be evaluated The node embedding vector at the K-1th layer, Represents an aggregate function, Indicates the second neighbor node The node embedding vector at the K-1th layer, Indicates the second enterprise to be evaluated The second neighbor node set in the Kth layer. Can realize cross-layer feature dimension transformation; nonlinear activation function The activation function can be ReLU or LeakyReLU to enhance the representation ability of the node embedding vector.
[0106] It should be understood that the network model including the feature extraction layer and the nonlinear activation function layer is trainable, thereby improving the representation ability of the second node embedding vector by training the model, thereby further improving the accuracy of the subsequently determined risk association assessment results, that is, further improving the accuracy of the enterprise risk association assessment, and ultimately further improving the accuracy of the enterprise risk assessment.
[0107] The enterprise risk assessment method provided by an embodiment of the present invention inputs a node embedding vector into a feature extraction layer to obtain a node extraction vector output by the feature extraction layer, and inputs the node extraction vector into a nonlinear activation function layer to obtain a second node embedding vector output by the nonlinear activation function layer, thereby further improving the representation ability of the second node embedding vector through feature extraction and nonlinear activation function, and further improving the accuracy of the subsequently determined risk association assessment results, that is, further improving the accuracy of the enterprise risk association assessment, and ultimately further improving the accuracy of the enterprise risk assessment.
[0108] Based on any of the above embodiments, in this method, the second set of neighbor nodes is determined based on the following method: The neighbor node set of the second enterprise to be evaluated in the Kth layer is sampled to obtain a second neighbor node set of the second enterprise to be evaluated in the Kth layer.
[0109] Among them, the neighbor node set of the second enterprise to be evaluated at the Kth layer includes: all direct neighbor nodes of each neighbor node in the neighbor node set of the second enterprise to be evaluated at the K-1th layer; the neighbor node set of the second enterprise to be evaluated at the first layer includes: all direct neighbor nodes of the second enterprise to be evaluated.
[0110] Here, a direct neighbor node indicates that the relationship between two nodes is a direct relationship, and a direct relationship indicates that the two nodes are directly connected by an edge.
[0111] In one embodiment, the sampling number of each type of association relationship can be determined based on the association relationship between each neighbor node in the neighbor node set of the second enterprise to be evaluated at the K-1 layer and its direct neighbor node, and the neighbor node set of the second enterprise to be evaluated at the K layer can be sampled based on each sampling number. More specifically, a relationship-aware sampling function can be used for sampling, that is, truncation control is performed based on the sampling number to control the number of neighbor nodes of each association relationship. For heterogeneous graph data, different relationship types can be configured with differentiated sampling strategies, such as downsampling high-frequency relationships and oversampling low-frequency relationships, so as to improve the modeling ability of long-tail distributions, thereby further improving the representation ability of the second node embedding vector, and further improving the accuracy of the subsequently determined risk association assessment results, that is, further improving the accuracy of the enterprise risk association assessment, and ultimately further improving the accuracy of the enterprise risk assessment.
[0112] The enterprise risk assessment method provided by an embodiment of the present invention samples the neighbor node set of the second enterprise to be assessed in the Kth layer, and reduces the number of neighbor nodes in the second neighbor node set through sampling and selecting only the neighbor nodes in the Kth layer, thereby reducing the amount of calculation, improving the efficiency of enterprise risk assessment, and ensuring that the second node embedding vector is accurately obtained, thereby further improving the accuracy of the subsequently determined risk association assessment results, that is, further improving the accuracy of the enterprise risk association assessment, and ultimately further improving the accuracy of the enterprise risk assessment.
[0113] Based on any of the above embodiments, Figure 2 This is the second flow chart of the enterprise risk assessment method provided by the present invention, such as Figure 2 As shown, the enterprise risk assessment method includes: step 110, step 120, step 131, step 140, step 150 and step 160. Step 140 is after step 131.
[0114] Step 131 : determining the risk association assessment result based on a comparison result between the similarity calculation result and a preset similarity threshold.
[0115] Here, the preset similarity threshold can be set according to actual conditions. For example, the value range of the preset similarity threshold is [0.6, 0.8]. Correspondingly, the value range of the similarity calculation result is [0, 1].
[0116] Furthermore, the preset similarity threshold is dynamically changed, thereby achieving dynamic threshold adjustment.
[0117] Furthermore, the preset similarity threshold can be calibrated through Bayesian optimization. 1,432 cases from the past three years are loaded to determine the preset similarity threshold. Grid search determines that when the preset similarity threshold is 0.72, the recall rate reaches 89.6%.
[0118] The risk association assessment result includes a first risk association result and a second risk association result, and the risk association degree of the first risk association result is greater than the risk association degree of the second risk association result.
[0119] Furthermore, the first risk association result indicates that the first enterprise to be evaluated and the second enterprise to be evaluated have risk association, and the second risk association result indicates that the first enterprise to be evaluated and the second enterprise to be evaluated do not have risk association.
[0120] In one embodiment, a larger similarity calculation result indicates a greater correlation between the first enterprise to be evaluated and the second enterprise to be evaluated. Based on this, when the similarity calculation result is greater than a preset similarity threshold, the risk association assessment result is determined to be a first risk association result; when the similarity calculation result is less than or equal to the preset similarity threshold, the risk association assessment result is determined to be a second risk association result.
[0121] In another embodiment, the smaller the similarity calculation result, the more related the first enterprise to be evaluated is to the second enterprise to be evaluated. Based on this, when the similarity calculation result is greater than the preset similarity threshold, the risk association assessment result is determined to be the second risk association result; when the similarity calculation result is less than or equal to the preset similarity threshold, the risk association assessment result is determined to be the first risk association result.
[0122] Step 140: When the risk association assessment result is the first risk association result, obtain enterprise registration information.
[0123] Specifically, when the risk correlation between the first enterprise to be assessed and the second enterprise to be assessed is relatively high, further enterprise registration information is obtained to verify whether there is a risk correlation between the first enterprise to be assessed and the second enterprise to be assessed, that is, to further determine whether the two enterprises to be assessed are highly risk-related, thereby further improving the accuracy of enterprise risk correlation assessment through further verification mechanism, and ultimately further improving the accuracy of enterprise risk assessment.
[0124] Here, the enterprise registration information includes association record information between multiple enterprises, and when an enterprise is registered, an association record can be made for the enterprise's associated enterprises. Furthermore, the association record is a direct association record.
[0125] Step 150: When it is determined based on the enterprise registration information that the first enterprise to be evaluated and the second enterprise to be evaluated have an association record, it is determined that there is a risk association between the first enterprise to be evaluated and the second enterprise to be evaluated.
[0126] If the first enterprise to be assessed has associated records with the second enterprise to be assessed, there is no need to trigger the verification process again, and it can be directly determined that the first enterprise to be assessed and the second enterprise to be assessed have a risk association.
[0127] Furthermore, if the first enterprise to be evaluated is a risk enterprise, then the second enterprise to be evaluated is also a risk enterprise. Of course, there can be multiple second enterprises to be evaluated. Based on this, only one enterprise to be evaluated needs to be determined as a risk enterprise, and multiple other enterprises can be quickly determined as risk enterprises, thereby improving the efficiency of enterprise risk assessment.
[0128] Step 160 : triggering a verification process when it is determined based on the enterprise registration information that there is no association record between the first enterprise to be evaluated and the second enterprise to be evaluated.
[0129] The verification process is used to verify whether there is a risk association between the first enterprise to be evaluated and the second enterprise to be evaluated.
[0130] If there is no association record between the first enterprise to be assessed and the second enterprise to be assessed, a verification process needs to be triggered to further verify whether there is a risk association between the first enterprise to be assessed and the second enterprise to be assessed.
[0131] For example, assuming that the first enterprise to be evaluated is a risk enterprise, if there is a risk association between the first enterprise to be evaluated and the second enterprise to be evaluated, then the second enterprise to be evaluated is also a risk enterprise. Of course, the number of second enterprises to be evaluated can be multiple. Based on this, only one enterprise to be evaluated needs to be determined as a risk enterprise, and multiple other enterprises can be quickly determined as risk enterprises, thereby improving the efficiency of enterprise risk assessment.
[0132] The enterprise risk assessment method provided by an embodiment of the present invention determines that there is a risk association between the first enterprise to be assessed and the second enterprise to be assessed when it is determined based on the enterprise registration information that there is an association record between the first enterprise to be assessed and the second enterprise to be assessed; and triggers a verification process when it is determined based on the enterprise registration information that there is no association record between the first enterprise to be assessed and the second enterprise to be assessed, thereby further verifying whether there is a risk association between the first enterprise to be assessed and the second enterprise to be assessed. That is, through the further verification mechanism, the accuracy of the enterprise risk association assessment is further improved, and ultimately the accuracy of the enterprise risk assessment is further improved.
[0133] Based on any of the above embodiments, in this method, the enterprise relationship map is determined based on the following method: Obtain enterprise data for multiple enterprises; Based on the enterprise data of the multiple enterprises, an enterprise relationship map is constructed.
[0134] Here, enterprise data may include, but is not limited to, at least one of the following: business data, operational data, public opinion data, and IoT sensor data. For example, enterprise data includes financial statements, equity structure, operating indicators, audit conclusions, and compliance records. This enterprise data can include both structured and unstructured data.
[0135] Specifically, a graph computation is performed based on the enterprise data of multiple enterprises to obtain an enterprise relationship graph. The multiple enterprises should include a first enterprise to be evaluated and a second enterprise to be evaluated. Furthermore, based on an enterprise relationship network modeling solution for heterogeneous graphs, the enterprise relationship graph is constructed. Specifically, node modeling is performed first, followed by edge relationship modeling.
[0136] In a specific embodiment, enterprise data is digitized and converted into quantitative data to facilitate extraction of node feature vectors.
[0137] In a specific embodiment, natural language processing and data analysis techniques in deep learning are used to parse enterprise data and extract key information, such as enterprise size, operating conditions, risk level, etc., and then a corporate relationship map based on enterprise data is constructed, and graph computing technology is used to discover the associations between risky enterprises, such as industrial chain relationships, equity relationships, etc.
[0138] Furthermore, enterprise data is preprocessed, which includes data cleaning and standardization. Data cleaning involves removing duplicate data to ensure dataset uniqueness; addressing missing values by filling them with interpolation, regression prediction, or machine learning-based predictive models; and detecting and addressing outliers. For data standardization, continuous variables are normalized using z-score or min-max normalization, while categorical variables are one-hot encoded or label encoded.
[0139] The enterprise data includes multi-dimensional data, thereby improving the accuracy of constructing the enterprise relationship map, further improving the accuracy of assessing enterprise risk associations, and ultimately improving the accuracy of assessing enterprise risks. Accordingly, enterprise nodes in the enterprise relationship map are represented by node feature vectors, which are multi-dimensional feature vectors.
[0140] It should be noted that existing technologies generally have defects such as single data dimension and weak correlation analysis. For example, the regulatory system only integrates financial data to implement risk warnings and lacks the ability to process unstructured data.
[0141] The sources of enterprise data include multiple different databases, expanding the sources of enterprise data, thereby improving the accuracy of enterprise relationship map construction, further improving the accuracy of enterprise risk association assessment, and ultimately improving the accuracy of enterprise risk assessment. For example, data from different regulatory systems, such as tax, industry and commerce, can be cross-validated and supplemented with enterprise data through data fusion technology, improving the accuracy and completeness of enterprise data.
[0142] Among them, the multidimensional feature vector is determined based on real-time enterprise data, thereby ensuring that the enterprise relationship map is also real-time, that is, improving the accuracy of constructing the enterprise relationship map, thereby improving the accuracy of assessing enterprise risk associations, and ultimately improving the accuracy of assessing enterprise risks.
[0143] Taking into account the current situation that the existing technology has a single enterprise data collection channel and an update cycle of more than 30 days, it is impossible to effectively capture the multi-dimensional behavioral characteristics generated in real time during the operation of the enterprise, resulting in more than 60% of the early warning signals being false alarms or missed reports; based on this, the sources of enterprise data include multiple different databases, and enterprise data includes multi-dimensional data, and the multi-dimensional feature vector is determined based on real-time enterprise data, thereby realizing real-time supervision of data streams to build an enterprise relationship map, thereby improving the assessment accuracy of enterprise risk associations, and ultimately improving the assessment accuracy of enterprise risks.
[0144] It should be understood that by constructing a corporate relationship map, a panoramic corporate portrait can be provided, and then an intelligent early warning mechanism can be provided, significantly improving the timeliness and accuracy of corporate risk assessment.
[0145] In the enterprise risk assessment method provided by an embodiment of the present invention, enterprise data includes multi-dimensional data, and the source of the enterprise data includes multiple different databases to expand the source of enterprise data, and the multi-dimensional feature vector is determined based on real-time enterprise data, thereby ensuring that the enterprise relationship map is also real-time; through the above method, the accuracy of constructing the enterprise relationship map can be improved, thereby improving the assessment accuracy of enterprise risk associations, and ultimately improving the assessment accuracy of enterprise risks.
[0146] Based on any of the above embodiments, Figure 3 This is the third flow chart of the enterprise risk assessment method provided by the present invention, such as Figure 3 As shown, the enterprise risk assessment method further includes: step 310 and step 320.
[0147] Step 310: When it is determined that there is a risk enterprise, determine the enterprises to be risk-assessed that have a relationship with the risk enterprise based on the enterprise relationship map.
[0148] Here, the association relationship between the risk enterprise and the enterprise to be risk assessed can be a direct association relationship or an indirect association relationship; a direct association relationship means that the two nodes are directly connected by an edge, and an indirect association relationship means that the two nodes are not directly connected by an edge, that is, there are other nodes between the two nodes.
[0149] Since the edges in the enterprise relationship graph are used to represent the association relationships between enterprises, the enterprises to be risk-assessed that have an association relationship with the risky enterprises can be determined based on the enterprise relationship graph.
[0150] Step 320: Determine whether the enterprise to be risk-assessed is a risky enterprise based on the correlation coefficient between the risk enterprise and the enterprise to be risk-assessed.
[0151] Among them, the correlation degree coefficient is obtained by multiplying the relationship strength coefficients of each edge of the shortest correlation path between the risk enterprise and the enterprise to be risk assessed. The relationship strength coefficient is used to characterize the strength of the correlation relationship between enterprises. The relationship strength coefficient is greater than 0 and less than or equal to 1.
[0152] Here, the shortest path is the path with the shortest number of edges among the multiple paths between the venture enterprise's node and the enterprise to be risk assessed. For example, if the venture enterprise is Enterprise A and the enterprise to be risk assessed is Enterprise B, and there are two paths between Enterprise A and Enterprise B: Enterprise A-Enterprise C-Enterprise B and Enterprise A-Enterprise C-Enterprise D-Enterprise B. In this case, the shortest path is Enterprise A-Enterprise C-Enterprise B.
[0153] Here, the relationship strength coefficient can be calculated according to relevant calculation methods, such as normalization of shareholding ratio, logarithmic transformation of transaction volume, and tenure length coefficient.
[0154] The enterprise risk assessment method provided by the embodiment of the present invention can quickly determine whether the enterprise to be risk-assessed is a risky enterprise through the above-mentioned method, thereby improving the efficiency of enterprise risk assessment.
[0155] Based on any of the above embodiments, in this method, the relationship strength coefficient of any side is determined based on the following method: Determine the relationship strength sub-coefficient of each association relationship category of the edge; the relationship strength sub-coefficient of any association relationship category is used to represent the strength of the association relationship between enterprises related to the association relationship category; Based on the weights of the association relationship categories, each of the relationship strength sub-coefficients is weightedly aggregated to obtain the relationship strength coefficient of the edge.
[0156] Here, the relationship strength sub-coefficient can be calculated according to relevant calculation methods, such as normalization of shareholding ratio, logarithmic transformation of transaction volume, and tenure length coefficient.
[0157] Since the association relationship between two enterprises may include multiple ones, that is, multiple association relationships may exist at the same time, the relationship strength sub-coefficient of each association relationship category of the edge may be determined.
[0158] Here, the weight of each relationship category can be set in advance. For example, the weight of equity control relationship is 0.38. Furthermore, the weight of each relationship category is added up to 1.
[0159] The enterprise risk assessment method provided by the embodiment of the present invention performs weighted aggregation on the relationship strength sub-coefficients based on the weights of each association relationship category, and can accurately obtain the relationship strength coefficient of the edge, that is, taking into account that different association relationship categories have different association strengths, thereby improving the accuracy of enterprise risk assessment.
[0160] The enterprise risk assessment device provided by the present invention is described below. The enterprise risk assessment device described below and the enterprise risk assessment method described above can be referenced to each other.
[0161] Figure 4 This is a schematic diagram of the structure of the enterprise risk assessment device provided by the present invention. Figure 4 As shown, the enterprise risk assessment device includes: an enterprise determination module 410 , a vector determination module 420 and an association assessment module 430 .
[0162] The enterprise determination module 410 is configured to determine two enterprises to be assessed that are associated with the risks to be assessed; the two enterprises to be assessed include a first enterprise to be assessed and a second enterprise to be assessed.
[0163] Vector determination module 420 is used to determine the first node embedding vector of the first enterprise to be evaluated and the second node embedding vector of the second enterprise to be evaluated based on the enterprise relationship graph; the nodes in the enterprise relationship graph are enterprise nodes, the edges in the enterprise relationship graph are used to represent the association relationship between enterprises, and the node feature vector of any enterprise node in the enterprise relationship graph is determined based on the enterprise data of the enterprise node.
[0164] The association assessment module 430 is configured to determine a risk association assessment result of the first enterprise to be assessed and the second enterprise to be assessed based on a similarity calculation result between the first node embedding vector and the second node embedding vector.
[0165] Among them, the first node embedding vector is determined based on the node feature vector of the first enterprise to be evaluated and the node feature vectors of each first neighbor node in the first neighbor node set, and each first neighbor node in the first neighbor node set is an enterprise node that has an association relationship with the first enterprise to be evaluated in the enterprise relationship graph; the second node embedding vector is determined based on the node feature vector of the second enterprise to be evaluated and the node feature vectors of each second neighbor node in the second neighbor node set, and each second neighbor node in the second neighbor node set is an enterprise node that has an association relationship with the second enterprise to be evaluated in the enterprise relationship graph.
[0166] Figure 5 An example of a physical structure diagram of an electronic device is shown below. Figure 5As shown, the electronic device may include: a processor 510, a communications interface 520, a memory 530 and a communication bus 540, wherein the processor 510, the communications interface 520 and the memory 530 communicate with each other via the communication bus 540. The processor 510 may call the logic instructions in the memory 530 to execute the enterprise risk assessment method, which includes: determining two enterprises to be assessed that are associated with the risks to be assessed; the two enterprises to be assessed include a first enterprise to be assessed and a second enterprise to be assessed; based on the enterprise relationship graph, determining the first node embedding vector of the first enterprise to be assessed and the second node embedding vector of the second enterprise to be assessed respectively; the nodes in the enterprise relationship graph are enterprise nodes, the edges in the enterprise relationship graph are used to represent the association relationship between enterprises, and the node feature vector of any enterprise node in the enterprise relationship graph is determined based on the enterprise data of the enterprise node; based on the similarity calculation between the first node embedding vector and the second node embedding vector The calculation result is used to determine the risk association assessment result of the first enterprise to be assessed and the second enterprise to be assessed; wherein, the first node embedding vector is determined based on the node feature vector of the first enterprise to be assessed and the node feature vectors of each first neighbor node in the first neighbor node set, and each first neighbor node in the first neighbor node set is an enterprise node that has an association relationship with the first enterprise to be assessed in the enterprise relationship graph; the second node embedding vector is determined based on the node feature vector of the second enterprise to be assessed and the node feature vectors of each second neighbor node in the second neighbor node set, and each second neighbor node in the second neighbor node set is an enterprise node that has an association relationship with the second enterprise to be assessed in the enterprise relationship graph.
[0167] Furthermore, the logic instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0168] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the enterprise risk assessment method provided by the above methods, which includes: determining two enterprises to be assessed that are associated with the risks to be assessed; the two enterprises to be assessed include a first enterprise to be assessed and a second enterprise to be assessed; based on the enterprise relationship graph, determining the first node embedding vector of the first enterprise to be assessed and the second node embedding vector of the second enterprise to be assessed respectively; the nodes in the enterprise relationship graph are enterprise nodes, and the edges in the enterprise relationship graph are used to represent the association relationship between enterprises. The node feature vector of any enterprise node in the enterprise relationship graph is based on the enterprise data of the enterprise node. According to; based on the similarity calculation result of the first node embedding vector and the second node embedding vector, the risk association assessment result of the first to-be-assessed enterprise and the second to-be-assessed enterprise is determined; wherein, the first node embedding vector is determined based on the node feature vector of the first to-be-assessed enterprise and the node feature vectors of each first neighbor node in the first neighbor node set, and each first neighbor node in the first neighbor node set is an enterprise node that has an association relationship with the first to-be-assessed enterprise in the enterprise relationship graph; the second node embedding vector is determined based on the node feature vector of the second to-be-assessed enterprise and the node feature vectors of each second neighbor node in the second neighbor node set, and each second neighbor node in the second neighbor node set is an enterprise node that has an association relationship with the second to-be-assessed enterprise in the enterprise relationship graph.
[0169] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented by a processor to execute the enterprise risk assessment method provided by the above-mentioned methods, the method comprising: determining two enterprises to be assessed that are associated with the risks to be assessed; the two enterprises to be assessed include a first enterprise to be assessed and a second enterprise to be assessed; based on an enterprise relationship graph, determining a first node embedding vector of the first enterprise to be assessed and a second node embedding vector of the second enterprise to be assessed respectively; the nodes in the enterprise relationship graph are enterprise nodes, the edges in the enterprise relationship graph are used to represent the association relationship between enterprises, and the node feature vector of any enterprise node in the enterprise relationship graph is determined based on the enterprise data of the enterprise node; based on the first node embedding vector The similarity calculation result of the vector and the second node embedding vector is used to determine the risk association assessment result of the first enterprise to be assessed and the second enterprise to be assessed; wherein, the first node embedding vector is determined based on the node feature vector of the first enterprise to be assessed and the node feature vectors of each first neighbor node in the first neighbor node set, and each first neighbor node in the first neighbor node set is an enterprise node that has an association relationship with the first enterprise to be assessed in the enterprise relationship graph; the second node embedding vector is determined based on the node feature vector of the second enterprise to be assessed and the node feature vectors of each second neighbor node in the second neighbor node set, and each second neighbor node in the second neighbor node set is an enterprise node that has an association relationship with the second enterprise to be assessed in the enterprise relationship graph.
[0170] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0171] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0172] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for enterprise risk assessment, characterized in that: include: Identify two companies to be assessed that have associated risks; The two enterprises to be assessed include a first enterprise to be assessed and a second enterprise to be assessed; Based on the enterprise relationship graph, determining a first node embedding vector of the first enterprise to be evaluated and a second node embedding vector of the second enterprise to be evaluated; The nodes in the enterprise relationship graph are enterprise nodes, the edges in the enterprise relationship graph are used to represent the association relationship between enterprises, and the node feature vector of any enterprise node in the enterprise relationship graph is determined based on the enterprise data of the enterprise node; Determining a risk association assessment result between the first enterprise to be assessed and the second enterprise to be assessed based on a similarity calculation result between the first node embedding vector and the second node embedding vector; The first node embedding vector is determined based on the node feature vector of the first enterprise to be evaluated and the node feature vectors of each first neighbor node in the first neighbor node set, and each first neighbor node in the first neighbor node set is an enterprise node in the enterprise relationship graph that has an association relationship with the first enterprise to be evaluated; The second node embedding vector is determined based on the node feature vector of the second enterprise to be evaluated and the node feature vector of each second neighbor node in the second neighbor node set, and each second neighbor node in the second neighbor node set is an enterprise node that has an association relationship with the second enterprise to be evaluated in the enterprise relationship graph.
2. The enterprise risk assessment method according to claim 1, characterized in that: The determining, based on the enterprise relationship graph, a first node embedding vector of the first enterprise to be evaluated and a second node embedding vector of the second enterprise to be evaluated respectively includes: Determine, based on the enterprise relationship graph, the number of association layers between the first enterprise to be evaluated and the second enterprise to be evaluated; the number of association layers is the total number of edges in the shortest association path between the first enterprise to be evaluated and the second enterprise to be evaluated in the enterprise relationship graph; Determining a first node embedding vector of the first enterprise to be evaluated and a second node embedding vector of the second enterprise to be evaluated based on the number of association layers; Wherein, if the number of association layers is equal to 1, the first node embedding vector is the node feature vector of the first enterprise to be evaluated, and the second node embedding vector is the node feature vector of the second enterprise to be evaluated; If the number of association layers is greater than 1, the first node embedding vector is determined based on the node feature vector of the first enterprise to be evaluated and the neighbor aggregation vector of the first neighbor node set, the neighbor aggregation vector of the first neighbor node set is determined based on the node feature vector of each first neighbor node in the first neighbor node set, the second node embedding vector is determined based on the node feature vector of the second enterprise to be evaluated and the neighbor aggregation vector of the second neighbor node set, and the neighbor aggregation vector of the second neighbor node set is determined based on the node feature vector of each second neighbor node in the second neighbor node set.
3. The enterprise risk assessment method according to claim 2, characterized in that: If the number of association levels is greater than 1 and is K, the first node embedding vector is determined based on the following method: Aggregating the node embedding vectors of each first neighbor node in the first neighbor node set at the K-1 layer to obtain a neighbor aggregation vector of the first neighbor node set; Aggregating the neighbor aggregation vectors of the first neighbor node set and the node embedding vector of the first enterprise to be evaluated at the K-1 layer to obtain a first node embedding vector of the first enterprise to be evaluated at the K layer; The node embedding vector of any of the first neighbor nodes in the first layer is the node feature vector of the first neighbor node; The node embedding vector of the first enterprise to be evaluated in the first layer is the node feature vector of the first enterprise to be evaluated.
4. The enterprise risk assessment method according to claim 3, characterized in that: The aggregating the neighbor aggregation vector of the first neighbor node set and the node embedding vector of the first enterprise to be evaluated at the K-1 layer to obtain the first node embedding vector of the first enterprise to be evaluated at the K layer includes: Aggregating the neighbor aggregation vectors of the first neighbor node set and the node embedding vector of the first enterprise to be evaluated at the K-1 layer to obtain the node embedding vector of the first enterprise to be evaluated at the K layer; Integrating the node into a vector input to a feature extraction layer to obtain a node extraction vector output by the feature extraction layer; The node extraction vector is input into a nonlinear activation function layer to obtain the first node embedding vector output by the nonlinear activation function layer.
5. The enterprise risk assessment method according to claim 3, characterized in that: The first set of neighbor nodes is determined based on the following method: Sampling the neighbor node set of the first enterprise to be evaluated at the Kth layer to obtain the first neighbor node set of the first enterprise to be evaluated at the Kth layer; Among them, the neighbor node set of the first enterprise to be evaluated at the Kth layer includes: all direct neighbor nodes of each neighbor node in the neighbor node set of the first enterprise to be evaluated at the K-1th layer; the neighbor node set of the first enterprise to be evaluated at the first layer includes: all direct neighbor nodes of the first enterprise to be evaluated.
6. The enterprise risk assessment method according to any one of claims 1 to 5, characterized in that: The determining, based on the similarity calculation result of the first node embedding vector and the second node embedding vector, the risk association assessment result between the first to-be-assessed enterprise and the second to-be-assessed enterprise includes: Determining the risk association assessment result based on a comparison result of the similarity calculation result and a preset similarity threshold; the risk association assessment result includes a first risk association result and a second risk association result, and the risk association degree of the first risk association result is greater than the risk association degree of the second risk association result; After determining the risk association assessment result between the first to-be-assessed enterprise and the second to-be-assessed enterprise based on the similarity calculation result between the first node embedding vector and the second node embedding vector, the method further includes: When the risk association assessment result is the first risk association result, obtaining enterprise registration information; If it is determined based on the enterprise registration information that the first enterprise to be assessed has an association record with the second enterprise to be assessed, determining that there is a risk association between the first enterprise to be assessed and the second enterprise to be assessed; If it is determined based on the enterprise registration information that the first enterprise to be evaluated has no association record with the second enterprise to be evaluated, a verification process is triggered; the verification process is used to verify whether there is a risk association between the first enterprise to be evaluated and the second enterprise to be evaluated.
7. The enterprise risk assessment method according to any one of claims 1 to 5, characterized in that: The enterprise relationship map is determined based on the following method: Acquire enterprise data of multiple enterprises; the enterprise data includes multi-dimensional data, and the sources of the enterprise data include multiple different databases; Building an enterprise relationship map based on the enterprise data of the plurality of enterprises; Among them, the enterprise nodes in the enterprise relationship map are represented by node feature vectors, and the node feature vectors are multidimensional feature vectors, which are determined based on real-time enterprise data.
8. The enterprise risk assessment method according to any one of claims 1 to 5, characterized in that: The enterprise risk assessment method further includes: In the case where it is determined that there is a risk enterprise, determining enterprises to be risk-assessed that have a relationship with the risk enterprise based on the enterprise relationship map; Determining whether the enterprise to be risk-assessed is a risky enterprise based on a correlation coefficient between the risk enterprise and the enterprise to be risk-assessed; Among them, the correlation degree coefficient is obtained by multiplying the relationship strength coefficients of each edge of the shortest correlation path between the risk enterprise and the enterprise to be risk assessed. The relationship strength coefficient is used to characterize the strength of the correlation relationship between enterprises. The relationship strength coefficient is greater than 0 and less than or equal to 1.
9. The enterprise risk assessment method according to claim 8, characterized in that: The relationship strength coefficient for either side is determined based on the following: Determine the relationship strength sub-coefficient of each association relationship category of the edge; the relationship strength sub-coefficient of any association relationship category is used to represent the strength of the association relationship between enterprises related to the association relationship category; Based on the weights of the association relationship categories, each of the relationship strength sub-coefficients is weightedly aggregated to obtain the relationship strength coefficient of the edge.
10. An enterprise risk assessment device, characterized in that: include: An enterprise identification module is used to identify two enterprises to be assessed that are associated with the risks to be assessed; The two enterprises to be assessed include a first enterprise to be assessed and a second enterprise to be assessed; a vector determination module, configured to determine, based on the enterprise relationship graph, a first node embedding vector of the first enterprise to be evaluated and a second node embedding vector of the second enterprise to be evaluated; The nodes in the enterprise relationship graph are enterprise nodes, the edges in the enterprise relationship graph are used to represent the association relationship between enterprises, and the node feature vector of any enterprise node in the enterprise relationship graph is determined based on the enterprise data of the enterprise node; an association assessment module, configured to determine a risk association assessment result between the first enterprise to be assessed and the second enterprise to be assessed based on a similarity calculation result between the first node embedding vector and the second node embedding vector; The first node embedding vector is determined based on the node feature vector of the first enterprise to be evaluated and the node feature vectors of each first neighbor node in the first neighbor node set, and each first neighbor node in the first neighbor node set is an enterprise node in the enterprise relationship graph that has an association relationship with the first enterprise to be evaluated; The second node embedding vector is determined based on the node feature vector of the second enterprise to be evaluated and the node feature vector of each second neighbor node in the second neighbor node set, and each second neighbor node in the second neighbor node set is an enterprise node that has an association relationship with the second enterprise to be evaluated in the enterprise relationship graph.
11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the enterprise risk assessment method according to any one of claims 1 to 9 is implemented.
12. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the enterprise risk assessment method according to any one of claims 1 to 9 is implemented.
13. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the enterprise risk assessment method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Enterprise evaluation method, device and equipment
CN113112186A
Enterprise risk infection path analysis method, device and equipment and storage medium
CN113688287A
Internet service determination method and device based on enterprise relationship network
CN114331142A
User data processing method and device, computer equipment and storage medium
CN115630973A
Risk enterprise identification method and device, equipment and medium
CN115796572A