Comprehensive risk early warning system based on market big data and credit risks of two transaction parties
Through the integrated risk warning system, the multi-source data and entity relationship network are integrated, and the problems of data monitoring of credit risk analysis in the existing technology are solved, achieving more efficient and accurate credit risk monitoring and early warning.
Patent Information
- Application Number
- CN202510449500.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-09
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing credit risk analysis technologies rely too much on single-dimensional data monitoring, making it difficult to timely capture and correlate external environmental factors, real-time market information and disaster events, resulting in risk warning lag and prediction deviation.
The comprehensive risk warning system based on market big data is adopted, and through the market data aggregation module, signal quantization analysis module, associated risk penetration module and comprehensive risk warning module, ESG rating, news and emotion, geographical activity, disaster warning and counterparty industrial and commercial financial data are integrated to generate structured market factor flow, and through entity relationship network and risk conduction calculation, weighted risk signal vector and counterparty risk index scores are obtained.
It improves the real-time and sensitivity of credit risk monitoring, enhances the accuracy and sensitivity of risk factor correlation, achieves the accuracy and timeliness of risk prediction, and improves the speed and effectiveness of risk incident response.
Smart Images

Figure CN119963203A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of credit risk analysis, and in particular to a comprehensive risk early warning system based on market big data and the credit risks of both parties to a transaction. Background Art
[0002] The technical field of credit risk analysis focuses on identifying, assessing and predicting the risk that a company or individual will be unable to fulfill its financial obligations, specifically through continuous monitoring and analysis of the financial data, market behavior, credit history and external environmental factors (such as disaster events or economic fluctuations) of both parties to the transaction.
[0003] Existing technologies are overly dependent on single-dimensional data monitoring in credit risk analysis. In actual operation, they often only focus on the evaluation of financial data or historical credit records, lack timely capture and correlation analysis of external environmental factors, real-time market information and disaster events, resulting in delayed risk warnings and difficulty in accurately assessing the risks brought about by sudden or emerging events. For example, when sudden disasters or regulatory adjustments occur, existing technologies often lead to risk prediction deviations due to the difficulty in quickly integrating multi-source data such as market sentiment and geographic information changes. Therefore, improvements are needed. Summary of the invention
[0004] The purpose of the present invention is to solve the shortcomings of the prior art and to propose a comprehensive risk early warning system based on market big data and the credit risks of both parties to the transaction.
[0005] In order to achieve the above-mentioned purpose, the present invention adopts the following technical scheme: A comprehensive risk early warning system based on market big data and credit risk of both parties to a transaction includes: The market data aggregation module collects ESG ratings, news sentiment, geographic activity and disaster warnings, as well as counterparty business financial summary data, performs format alignment and timestamp marking, and screens out incomplete records to obtain structured market factor flows; The signal quantitative analysis module extracts credit risk text items based on the structured market factor flow according to ESG deviation values, sentiment polarity and geographic activity change rate, generates an initial signal measurement matrix, matches geographic changes with disaster areas based on the initial signal measurement matrix, associates penalty records with counterparties, sets impact weights, and obtains a weighted risk signal vector; The associated risk penetration module identifies counterparties, associated parties and relationship edges, builds a basic entity relationship network, assigns the weighted risk signal vector to nodes based on the basic entity relationship network, calculates the credit risk transmission impact along the ownership and supply chain edges, updates node metrics, and establishes a risk transmission association view; The comprehensive risk warning module, based on the risk transmission association view, summarizes the direct and transmission credit risk metrics of each counterparty node, and calculates the counterparty risk index score by weighted calculation; uses the counterparty risk index score to compare with the preset classification threshold, determines the risk range, and generates a classification warning notification.
[0006] Preferably, the steps of obtaining the structured market factor flow are: Collect ESG rating data, news sentiment data, geographic activity data, and disaster warning data, as well as counterparty business financial summary data to obtain a preliminary market data set; Based on the preliminary market data set, the field integrity in each record is scanned one by one, and each record is marked according to the position of the missing field. The marked results are called to remove the marked records one by one to obtain a structured market factor flow.
[0007] Preferably, the steps of obtaining the initial signal metric matrix are: Based on the structured market factor flow, extract the ESG rating data, news sentiment data and geographic activity data in each record, calculate the change rate and baseline deviation of each data, and generate a basic risk feature matrix for each record; Based on the basic risk characteristic matrix, the credit risk identification index of a single record is calculated using the following formula: ; in, is the credit risk identification index of a single record, is the ESG rating value of a single record, is the overall mean of ESG rating data, is the news sentiment polarity value of a single record, is the geographic activity value of a single record, is the average value of geographical activity, For the press release hour of this record, The smallest press release hour on record; Based on the credit risk identification index of each record, the corresponding original text items are evaluated, and the key text items are selected according to the value of the credit risk identification index. The key text items are combined to construct an initial signal measurement matrix.
[0008] Preferably, the step of obtaining the weighted risk signal vector is: Based on the initial signal metric matrix, signal items affected by geographic changes and disaster areas are extracted, and the risk areas associated with disasters are identified in combination with real-time geographic monitoring data to obtain an initial geographic risk data set; Based on the initial geographic risk data set, a comprehensive risk metric for each data point relative to the disaster area is calculated using the following formula: ; in, is the comprehensive risk measure of data point p, is the distance from data point p to disaster area r, is the time from the disaster occurrence to the data point p being recorded, is the total number of disaster areas considered, is the angle difference between the geographical location of data point p and the disaster center; Based on the comprehensive risk metric and combined with the penalty record of each counterparty, a weighted risk signal vector is formed.
[0009] Preferably, the steps of acquiring the basic entity relationship network are: Based on the business financial summary data of the counterparty, retrieve the subject identification information in the business financial summary data one by one, identify the entity identities of all the counterparties one by one through the subject identification information, and generate a counterparty entity list; According to the counterparty entity list, query the information of shareholders, subsidiaries, senior executives, suppliers and customers of each counterparty entity in the industrial and commercial financial summary data, verify the authenticity of each item one by one, and generate a set of related party entities of the counterparty; Based on the counterparty entity list and the counterparty's associated entity set, the ownership relationship and supply chain relationship between each counterparty entity and the associated entity are established one by one, the ownership relationship and supply chain relationship are converted into relationship edges one by one, and the entity nodes are connected to generate a basic entity relationship network.
[0010] Preferably, the steps of obtaining the risk transmission association view are: Based on the basic entity relationship network, a corresponding weighted risk signal vector is assigned to each enterprise node to generate an enterprise node network; Based on the enterprise node network, the total credit risk impact on each node is calculated using the following formula: ; in, represents the total credit risk impact on node q, is the risk level of node s, is the set of neighbor nodes directly connected to node q, is the connection strength between node q and node s, is the logarithm of the transaction amount between node q and node s, is the number of connected nodes of node q; Based on the total credit risk impact of each node, the risk level of the node is re-evaluated and a risk transmission association view is generated.
[0011] Preferably, the steps for obtaining the counterparty risk index score are: Based on the risk transmission association view, extract the direct credit risk level of each counterparty node and the total credit risk impact on each node to form a node credit risk value set; According to the node credit risk value set, the risk index score of the counterparty node is calculated, and the calculation formula is: ; in, is the risk index score of the counterparty node, is the risk level of node k, is the total credit risk impact on node k, is the total number of relationship edges associated with counterparty node m, is the total number of nodes directly associated with the counterparty node m; Based on the risk index score of the counterparty node, the risk mark of each counterparty node in the risk transmission association view is updated one by one to generate a counterparty risk index score.
[0012] Preferably, the steps for obtaining the graded warning notification are: Using the counterparty risk index score, the counterparty risk index score of each transaction counterparty node is called one by one, and the numerical comparison and analysis are performed one by one with the pre-set grading threshold value to obtain a preliminary result of risk level determination; Based on the preliminary results of the risk level determination, the counterparty node is divided into risk intervals according to the degree of deviation between the counterparty risk index score of each counterparty and the classification threshold, and a counterparty node risk interval set is generated; Based on the counterparty node risk interval set, the warning notification templates corresponding to each risk interval are matched one by one to form a graded warning notification.
[0013] Compared with the prior art, the advantages and positive effects of the present invention are: In the present invention, by comprehensively collecting ESG ratings, news sentiment, geographic activity, disaster warning and industrial and commercial financial summary data, format alignment and unified timestamp marking are performed on multi-source heterogeneous data to form a structured market factor flow, and multi-dimensional fusion and efficient integration of market information are realized in processing logic, thereby improving the real-time and sensitivity of credit risk monitoring; quantitative analysis of credit risk text items is performed based on ESG deviation values, sentiment polarity and geographic activity change rates, and accurate risk matching is performed in combination with geographic location and disaster information, directly improving the accuracy and sensitivity of risk factor associations; a basic entity relationship network between counterparties and related parties is constructed, and risk signal vectors are assigned, and risk transmission calculations are performed along ownership and supply chain relationships to achieve cross-entity risk transmission path tracking, thereby enhancing the accuracy and timeliness of credit risk prediction; risk index scores are calculated based on the fusion of node direct risk and transmission risk metrics, risk intervals are dynamically updated, and graded warning notifications are automatically triggered to achieve automation and intelligence in risk prediction and warning, thereby improving the speed and effectiveness of risk event response. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 It is a system flow chart of the present invention. DETAILED DESCRIPTION
[0015] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0016] See also Figure 1 The present invention provides a technical solution: a comprehensive risk early warning system based on market big data and credit risks of both parties to a transaction includes: The market data aggregation module collects ESG ratings, news sentiment, geographic activity and disaster warnings, as well as counterparty business financial summary data, performs format alignment and timestamp marking, and screens out incomplete records to obtain structured market factor flows; The signal quantitative analysis module, based on the structured market factor flow, extracts credit risk text items according to ESG deviation values, sentiment polarity and geographic activity change rate, generates an initial signal measurement matrix, matches geographic changes with disaster areas, associates penalty records with counterparties, sets impact weights, and obtains weighted risk signal vectors based on the initial signal measurement matrix; The associated risk penetration module identifies counterparties, associated parties, and relationship edges, builds a basic entity relationship network, assigns weighted risk signal vectors to nodes based on the basic entity relationship network, calculates the credit risk transmission impact along the ownership and supply chain edges, updates node metrics, and establishes a risk transmission association view; The comprehensive risk warning module, based on the risk transmission association view, summarizes the direct and transmission credit risk metrics of each counterparty node, and calculates the counterparty risk index score through weighted calculation; uses the counterparty risk index score to compare with the preset classification threshold, determines the risk range, and generates a classification warning notification.
[0017] The steps to obtain structured market factor flows are: Collect ESG rating data, news sentiment data, geographic activity data, and disaster warning data, as well as counterparty business financial summary data to obtain a preliminary market data set; Based on the preliminary market data set, the field integrity in each record is scanned one by one, and each record is marked according to the position of the missing field. The marking results are called to remove the marked records one by one to obtain the structured market factor flow.
[0018] Specifically, preparation work is carried out with reference to the ESG rating data, news sentiment data, geographic activity data, disaster warning data, and the counterparty's industrial and commercial financial summary data that have been obtained. First, different types of information are matched through unified timestamps at the multi-source collection end to ensure time zone consistency. Then, an empirical range is set for the ESG rating value for verification, for example, the minimum is set to 0 and the maximum is set to 100. If there are records exceeding this range, they are marked as suspicious data and compared again based on the judgment threshold constructed based on the historical average and standard deviation. The judgment threshold can be set by adding or subtracting two times the standard deviation of the average value of the historical data and some extreme cases are manually verified. If the threshold still cannot contain data that is too large or too small, the data will be eliminated. Taking news sentiment data as an example, a numerical range of -1 to 1 can be used. If it exceeds this range, the same detection and elimination method is used. The geographic activity data and disaster warning information are divided according to the frequency of occurrence in the actual monitoring results. For example, the activity frequency is divided into an integer range of 0 to 30 and the prediction threshold is set to 15 to determine whether it belongs to a high activity record. A similar statistical method is used to determine whether to eliminate extreme outliers. Next, a multi-layer model based on the combination of convolutional neural network and recurrent neural network can be introduced to automatically identify suspicious data. The model can include an input layer, four convolutional layers, a bidirectional recurrent layer and an output layer. Six types of data, including ESG ratings, news sentiment, geographic location coordinates, activity values, disaster feature labels and counterparty financial statements, are used as multi-channel inputs. All inputs are filtered layer by layer in the convolutional layer to extract temporal and spatial features and enter the bidirectional recurrent layer for short-term and long-term association identification. The stochastic gradient descent method with momentum factor is used for training and the cross entropy loss is calculated in each iteration cycle. The validation set is used to monitor the overfitting situation and the learning rate and regularization coefficient are adjusted to ensure the gradual convergence of the error. For example, the initial learning rate is set to 0.001 and the regularization coefficient is set to 0.0005. After thirty rounds of iterations, a relatively stable convergence result is obtained. After the training is completed, the suspicious data is inferred to determine whether it has reached the credibility level. If it is not credible, it is regarded as an invalid record and excluded in subsequent processing. Finally, all field information that can be recognized after multi-layer verification is integrated to form a preliminary market data set suitable for subsequent analysis.
[0019] Based on the preliminary market data set obtained in the previous paragraph, each record is scanned in turn according to the field completeness and the missing positions are compared one by one. For example, check whether there is a valid value between 0 and 100 in the ESG rating field, check whether it is between -1 and 1 in the news sentiment field, check whether it is lower than the activity frequency threshold of 15 mentioned above or higher than the extreme upper limit of 30 in the geographic activity field, and identify whether the corresponding disaster type in the disaster warning field is consistent with the data within the actual detection range. If it is found that some key fields of a record are empty or in an unreasonable range, mark the record as missing or abnormal. In order to accurately determine the records that can be retained, a simplified logical judgment method can be introduced. For example, when the same record ESG ratings that are blank and unreasonable news sentiment values are considered unavailable data. If only one field is missing, it is initially filled in based on the historical statistical mean and the threshold is compared again to see if it is reasonable. If the sentiment value after filling is still between -1 and 1, it is retained; otherwise, it is eliminated. The judgment process will also be roughly measured in combination with the asset or liability fields in the counterparty's financial summary. For example, it can be set that assets below RMB 1 million or above RMB 1 billion need to be checked to see if they match the values of other fields to exclude records that are obviously distorted. The above-mentioned scanning and marking steps are used to eliminate records marked as invalid and retain valid information that meets multiple judgment criteria. After aggregation, the structured market factor flow for subsequent quantitative analysis is obtained.
[0020] The steps to obtain the initial signal metric matrix are: Based on the structured market factor flow, extract the ESG rating data, news sentiment data and geographic activity data in each record, calculate the change rate and baseline deviation of each data, and generate a basic risk feature matrix for each record; Based on the basic risk characteristic matrix, the credit risk identification index of a single record is calculated using the following formula: ; in, is the credit risk identification index of a single record, is the ESG rating value of a single record, is the overall mean of ESG rating data, is the news sentiment polarity value of a single record, is the geographic activity value of a single record, is the average value of geographical activity, For the press release hour of this record, The smallest press release hour on record; Based on the credit risk identification index of each record, the corresponding original text items are evaluated, and the key text items are selected according to the value of the credit risk identification index. The key text items are combined to construct an initial signal measurement matrix.
[0021] Specifically, based on the structured market factor flow, the ESG rating data, news sentiment data and geographic activity data in each record are first read, and the ESG rating data is compared with a determined valid range when reading. For example, the ESG rating value is limited to between 0 and 100, the news sentiment value is limited to between -1 and 1, and the geographic activity value is limited to between 0 and 150. If any field exceeds the range, it is compared and judged with the ESG rating mean, news sentiment mean or geographic activity mean obtained in the previous step, and its validity is confirmed by combining the actual observation results and the previous accumulated data. If it is confirmed that the data is not If it is invalid, the record will be directly eliminated. After verifying the valid data, the corresponding change rate is calculated for ESG rating, news sentiment and geographical activity according to the time series information of each record, and the overall mean calculated previously is used as a benchmark to measure the degree of deviation of each record from the baseline. This is used to construct a numerical vector of basic risk characteristics. Subsequently, the numerical vectors of each record are summarized and sorted one by one, and the basic indicators such as the difference between ESG rating and baseline, news sentiment difference, and geographical activity difference in each vector are marked respectively. Finally, these indicators are combined in the same row of data to form a single matrix row, and all record matrix rows are summarized into a basic risk characteristic matrix.
[0022] formula: The formula is beneficial in that it introduces the multiple impacts of ESG ratings, news sentiment polarity, geographic activity, and news release time on credit risk, and generates a comprehensive measurement value for each record by performing logarithmic amplification and deviation measurement on these values, so as to finely screen records that may have credit risk hazards in subsequent steps; The parameter acquisition step is to retrieve the first The ESG rating value corresponding to the record is given by a professional evaluation agency in the range of 0 to 100. To determine the effectiveness of the rating, it is necessary to compare the mean values of similar rating samples from at least 300 companies within one year and set an industry benchmark range of 20 to 90. If the rating value of a company is lower than 20 or higher than 90, it is necessary to conduct a secondary review of its financing scale, industry reputation and previous ESG project scores. During the secondary review, it is necessary to collect information such as equity changes and administrative penalties that occurred in the company during the same period, and assign quantitative scores to it. The calculation process can be expressed as follows: ,in Score the enterprise in different evaluation dimensions. is the weight of each evaluation dimension (calculated by the importance of the dimension after the actual survey), and after completing the weighting, we get For example, when collecting statistics on a certain enterprise, its environmental compliance score is 70, social responsibility score is 75, and governance level score is 68, and weights of 0.3, 0.4, and 0.3 are assigned respectively. , and finally Set to 71.6; The steps to obtain the parameters are to count the ESG ratings of all companies in the same time period, divide their sum by the total number of companies, and set up a smoothing coefficient for weighted synthesis based on the ESG average of the previous year, denoted as ,in is the number of enterprises participating in the statistics, For each company's ESG rating value, is the comprehensive average of the previous year. is the smoothing coefficient, which can be set by comparing the deviation between the current year and the previous year. It can be set to 0.7. If the number of participating enterprises in the current year is large, it will be increased. If the number of enterprises is small this year, the proportion will be reduced. For example, if there are 500 companies this year, the average ESG score of the previous year is 65, and the total ESG score of the companies this year is 33,500. ,but ; The parameter acquisition step is to parse the quantitative value of sentiment tendency from the news text, with a value range of -1 to 1, where -1 represents highly negative, 1 represents highly positive, and 0 represents neutral. To obtain the value of , we need to count the frequency of adjectives, positive and negative words, and emotional symbols in the news sentences, and assign coefficients to the emotional intensity keywords. The calculation process can be expressed as ,in For the The intensity factor of the emotional keyword (range -3 to 3), is the frequency of occurrence of the keyword. The numerical result between -3 and 3 needs to be normalized and mapped to the range of -1 to 1. For example, when a news text contains a high proportion of negative keywords and the keyword intensity is high, For example, a news article contains ten sentiment keywords, including two strongly negative keywords (k=-3, f=3), one moderately negative keyword (k=-2, f=1), five neutral keywords (k=0, f=5), and two positive keywords (k=2, f=2). The weighted sum is , the total frequency is 3+1+5+2=11, and the result -7 / 11 is approximately -0.636, ; The steps to obtain the parameters are as follows: based on the multiple monitoring statistics of the geographical activity of enterprises in the previous structured market factor flow, the frequency of geographical activities actually surveyed by the enterprise within a period of time is quantified into specific values, and the frequency of logistics, registration, investment promotion or meetings of enterprises in a certain area is obtained by counting the number of events and multiplying it by the corresponding event weight. , where the event weight can be set according to the scale of the event. For example, the weight of logistics times can be set to 1, the weight of investment in establishing a subsidiary can be set to 5, and the weight of large-scale investment promotion activities can be set to 8. If an enterprise has 40 logistics events, 1 subsidiary establishment, and 1 investment promotion activity during the inspection period, then ; The steps to obtain the parameters are to summarize the geographic activity values of all enterprises in the same period and take the average value. If the number of sampled enterprises is , the total activity is ,but In the statistical process, the benchmark weights can be set in combination with the economic scale of different regions. For example, the activity of coastal areas can be magnified by 1.2 times and then summed, while the mountainous areas can be kept at 1 times. The magnification factor can be set by comprehensively evaluating the economic indicators of each region in the past five years and the actual activity records of enterprises. For example, when there are 100 enterprises in a certain region, the total activity value after weighted sum is 4900, then ; The parameter acquisition step is to record the release time hours of each news from the news data. For example, if the news is released at 10 o'clock on the same day, ; The step of obtaining the parameters is to find the minimum release hour value of all news records in the same statistical interval. For example, if the release time of all news is distributed between 1 and 23 hours, then ; Calculation process: The first step is to substitute the example values and set , , , , , , ; The second step is to calculate ,for , squared to 9.61; Step 3: Calculate , first ask , , so ; Step 4: Calculate ; Step 5: Multiply the numerators to get ; Step 6: Calculate the denominator ; Step 7: Take the logarithm. ; This result shows that when When it is close to or greater than 0.364, it means that the record has a certain cumulative impact on ESG rating deviation, negative news level or geographical activity. A larger value often indicates that the credit risk contained in the record is higher. If it is too low, it can be used as a record of relatively low risk in subsequent screening.
[0023] When evaluating the associated original text items based on the credit risk identification index of each record, first compare the CRI value of each record with a pre-set multi-level interval. For example, the CRI value can be divided into four levels: 0 to 0.3, 0.3 to 0.6, 0.6 to 1.0, and greater than 1.0. The proportion of each record in the corresponding level is recorded with reference to the standard comparison table set in the previous link. When the CRI value corresponding to a record falls into a higher level, sentences with obvious negative or extreme expressions are extracted from the original news text, and the keywords and intensity factors mentioned in the sentiment analysis are used to compare the CRI value of each record. The sub-statistical results confirm whether there is major negative information, and compare the previous ESG rating fluctuations with the geographical activity of the record to determine whether there are special circumstances in the record. If the ESG rating is significantly lower than 60 or the geographical activity is significantly higher than 80, focus on the core nouns appearing in the emotional text of the news, such as administrative penalties and insolvency, so as to screen out a number of key text items and correspond to the CRI index. After completing the same steps for all records, the corresponding relationship between the set of key text items and their index can be obtained, and finally the initial signal measurement matrix is formed based on the summary.
[0024] The steps to obtain the weighted risk signal vector are: Based on the initial signal metric matrix, the signal items affected by geographical changes and disaster areas are extracted, and combined with real-time geographical monitoring data, the risk areas associated with disasters are identified to obtain the initial geographical risk data set; Based on the initial geographic risk dataset, a comprehensive risk metric is calculated for each data point relative to the disaster area using the following formula: ; in, is the comprehensive risk measure of data point p, is the distance from data point p to disaster area r, is the time from the disaster occurrence to the data point p being recorded, is the total number of disaster areas considered, is the angle difference between the geographical location of data point p and the disaster center; Based on the comprehensive risk metric and combined with the penalty record of each counterparty, a weighted risk signal vector is formed.
[0025] Specifically, based on the initial signal measurement matrix, firstly, the signal items with high correlation with geographical change information and disaster areas are retrieved and arranged in chronological order, and then the corresponding comparison is carried out in combination with the geographical distribution information obtained from the previous monitoring. The collected signal items are located in the spatial coordinates and distinguished and marked according to the geographical coordinates and recording time of each record. The coordinates of all records are verified one by one with reference to the previously obtained geographical monitoring data, with particular attention paid to the records within the range of 0 km to 50 km from the disaster center. If there are records indicating that a certain area has suffered from disasters such as floods or earthquakes in the previous monitoring cycle, the area is marked as a high-risk area for disasters, and a weight score is set according to the disaster type and disaster scale of the area. For example, when the earthquake intensity is greater than level 5, the weight is set to 3, and when the flood-inundated area exceeds 5,000 square meters, the weight is set to 4. The process of obtaining the weight needs to refer to the historical data of the same type of disasters in the area in previous years, and the casualties and losses are quantified and scored year by year. The average impact level of this type of disaster in the past ten years is taken as the benchmark. Then, the final value is obtained by comparison based on the specific situation this time. The signals appearing in these high-weight areas are marked one by one in all records, and routine checks are further carried out on records in non-disaster areas. For example, based on the existing geographic monitoring data, one type of label is made for coastal areas with an altitude between 0 and 100 meters, and another type of label is made for mountainous terrain or areas with an altitude of more than 1,000 meters. The on-site meteorological information, population mobility data, and historical disaster records of each area in the past 30 days are compared. Areas that meet the characteristics of proneness to geological disasters will be identified as potential risk areas. Subsequently, the signal items of the identified high-risk and potential risk areas are aggregated, and the timestamps, geographic coordinates, disaster types, signal strength and other attributes are placed in the same data table. The geographic monitoring information obtained previously is associated through key fields, and the records with geographic coordinates within 5 kilometers and timestamps with a difference of no more than 24 hours are merged and marked so that the relevant information of the same time period and the same location can be processed uniformly in subsequent steps. When all data are checked, the initial geographic risk data set can be obtained.
[0026] formula: The benefit of the formula is that it combines multiple disaster areas and the distance differences between each data point and these disaster areas in spatial and temporal dimensions, comprehensively considers the differences in geographical locations, and thus subdivides the potential disaster threat level faced by each data point.
[0027] The steps to obtain the parameters are to collect the distance from the data point p to the disaster area r from the surveying and mapping data or GIS geographic information records, and complete the numerical calculation under the map projection or geographic coordinate system in combination with the distance measurement standard. The earthquake zone or flood range needs to be divided into several boundaries. If point p falls within the r area, the distance can be set to 0. Otherwise, the straight-line distance between point p and the nearest point of the r area contour is calculated. In the example, the horizontal distance between point p and the r area is measured using a 1:50,000 topographic map in a mountainous environment and is 8.2 kilometers. The altitude difference of 100 meters can be converted to approximately 0.1 kilometers, so .
[0028] The steps to obtain the parameters are to record the length of time from the time the disaster occurs to the time when the data point p is observed. The value needs to be obtained by subtracting the time when the disaster event occurs from the observation time of point p. If point p is observed multiple times during this period, the earliest valid recording time of the point is selected. If point p appears 5 hours after the disaster occurs, then , which is obtained by comparing the official disaster occurrence time of the central meteorological department or earthquake center with the observation timestamp of point p. In this example, the earthquake occurred at 07:00 and the timestamp of point p was 12:00. For 5 hours.
[0029] The steps for obtaining the parameters are to determine the total number of disaster areas of concern. These disaster areas are derived from statistical records of disasters of different types or locations that actually occurred in the same period. When there are multiple disasters such as earthquakes, floods, fires, etc. within a certain period of time, or the same type breaks out in multiple locations, a number must be assigned to each disaster area, and finally the number range is summarized to obtain the n value. For example, if 3 earthquake disasters and 2 flood disasters are registered in a certain monitoring period, then n=5.
[0030] The steps for obtaining the parameters are as follows: according to the geographical measurement results, the directional angle difference between the data point p and the disaster center in longitude and latitude is obtained. p and the disaster center are regarded as vectors, and the north direction is selected as the reference zero degree on the map. If the coordinates of the disaster center are ( ), the coordinates of point p are ( ), you can first convert them into the x and y coordinates of the plane projection, and then use atan2 and other methods to find the angle, or directly use the spherical geometry formula to calculate the azimuth, and finally get It is recorded in degrees. In this example, if the disaster center is located at longitude 114.00° and latitude 30.50°, and p is located at longitude 114.31° and latitude 30.52°, it can be converted to .
[0031] Calculation process: The first step is to substitute the example values and set , , , degrees, and calculate in sequence and Then multiply and sum; The second step is to set up all disaster areas. The distances are 8.3, 10.2, 3.6, 15.0, and 20.5 kilometers respectively, and the time record is unified as 5 hours. , first find ln(6) , and substitute the distances into , , , , , multiply by 1.791759 and sum to get ; Step 3: Take the cube root of 0.9392 , calculate the denominator , first , , , the sum of the two is approximately 1.3782; Step 4, Final The results show that for data point p1, the comprehensive risk metric is about 0.711. If it is greater than 1, it means that the coupling degree between the distance to the disaster area and the time point of the disaster is relatively close. If it is less than 0.5, it means that the data point is relatively less affected by the disaster.
[0032] Based on the comprehensive risk measurement, when matching the penalty records of each counterparty, first combine the penalty information table of each counterparty with the time-space corresponding information of the comprehensive risk measurement one by one, and set a benchmark range for the counterparty. For example, if the number of penalties in the past two years is less than 5 times, it is considered that the penalties are small, and if it is more than 5 times, it is considered that the penalties are large. If the same counterparty is located in a high disaster risk area in the near future and has a large number of penalty records, it will be marked and a weight coefficient will be set for it. The weight coefficient can be obtained by comparing the penalty records of previous years with the average value of risk measurement, accumulating scores based on information such as the frequency of penalties and the amount involved in each penalty, and simplifying the calculation in combination with the comprehensive risk measurement obtained previously. For example, when When the risk measure is greater than 0.8 and the number of penalties exceeds 5 times, the weight coefficient can be set to 2, otherwise the coefficient is set to 1. These weight coefficients are multiplied with the corresponding comprehensive risk measures to obtain weighted signal items. This process is performed for all counterparties to accumulate and record the comprehensive results. During the recording process, the geographic location labels and historical penalty details of the counterparties are also retained. If there are large fluctuations in the penalty amount, further proofreading is required with reference to the penalty amount range division. For example, 0 to 10,000 yuan corresponds to the lowest score range, 10,000 to 100,000 yuan corresponds to the middle score range, and 100,000 yuan and above corresponds to the higher score range. Finally, the comprehensive risk measure of each counterparty is combined with the weighted results of the penalty record to form a complete weighted risk signal vector.
[0033] The steps to obtain the basic entity relationship network are: Based on the business financial summary data of the counterparty, retrieve the subject identification information in the business financial summary data one by one, identify the entity identities of all the counterparties one by one through the subject identification information, and generate a counterparty entity list; Based on the counterparty entity list, query the information of shareholders, subsidiaries, senior executives, suppliers and customers in the industrial and commercial financial summary data of each counterparty entity, verify the authenticity of each item, and generate the counterparty's related party entity set; Based on the counterparty entity list and the counterparty's associated entity set, the ownership relationship and supply chain relationship between each counterparty entity and the associated entity are established one by one, the ownership relationship and supply chain relationship are converted into relationship edges one by one, and the entity nodes are connected to generate a basic entity relationship network.
[0034] Specifically, based on the counterparty's industrial and commercial financial summary data, first extract all subject identification information from the data and compare it with the corresponding company name or industrial and commercial registration number in a two-way manner. If it is found that any field information in the company name and the industrial and commercial registration number is incomplete, the entry is recorded and retained for a short period of time. Through data comparison from other sources, it is confirmed whether the information can be supplemented by data from companies with the same name or similar registration numbers. After that, a batch search process is performed for the subject identification information that can be matched normally. The search process requires reading the identification information one by one and searching for the corresponding company entry in the unified database. If the company has been recorded in the system, its existing profile is directly indexed. If the company has not yet appeared, it is searched according to the industrial and commercial financial A basic entry is created for the summary data and a unique code is assigned. Then the company name, industrial and commercial registration time, business scope code and main financial fields are compared one by one. For the business scope code, a four-digit or six-digit code interval can be set according to the national industry classification standard. For example, a range of 1001 to 1999 is set for the manufacturing industry, and a range of 2001 to 2999 is set for the wholesale and retail industry. When searching, if the business scope code of the company is inconsistent with the system setting, it will be included in the additional comparison list, and then combined with the annual operating income and net profit data retrieved on the spot for numerical verification. If the operating income is much greater than the average level of the same industry, it will be recorded in the list of large-value transaction companies in a graded manner. If the business income is obviously lower than the industry lower limit, it will be recorded in the list of doubtful business capabilities. In this way, the main information comparison and financial confirmation of all enterprises are completed. After the confirmation is completed, a complete counterparty entity entry table is formed inside the system. Each entry contains attributes such as enterprise name, business registration number, business scope code, financial indicators, etc., and the entry table is centrally stored as a counterparty entity list. In the final confirmation stage, a global duplicate item check will be carried out to identify the situation where the same business registration number may correspond to multiple names or the same name may correspond to multiple business registration numbers. All duplicate items will be submitted to manual verification or reference to the internal files of the enterprise for verification. If it is confirmed that they are indeed duplicates after verification, only the confirmed ones will be retained. Entries are checked and invalid entries are discarded. If new controversial entries need to be reconfirmed, a threshold for manual intervention is set. For example, if the dispute degree score exceeds 10, it will be subject to key review. The dispute score can be calculated based on the name similarity, registration number matching degree, and business scope matching degree. The dispute score can be calculated by giving 2 points for a complete name match, 5 points for a complete registration number match, and 3 points for a consistent business scope code. If the total score is less than 5, it means a high degree of difference. If it exceeds 10, it means that it is close to complete overlap but there are a few doubts. Finally, the data after the comparison is de-duplicated and integrated to obtain the entity identity information of each counterparty and generate a counterparty entity list.
[0035] According to the counterparty entity list, first obtain the unique code and company name of each counterparty, and then retrieve the additional information items corresponding to the unique code in the industrial and commercial financial summary data, such as shareholder name and equity ratio, subsidiary list, senior management position details, and supplier and customer records. These field information comes from the company's annual reports and industrial and commercial public documents at different time points. They need to be aligned before batch query. The query process can read the counterparty unique code one by one and associate it with the shareholder relationship table of the industrial and commercial financial summary data. Set a comparison threshold for the equity ratio, such as not less than 0.01, which is calculated based on 1% of the shareholding. If the shareholding ratio is less than 1%, it is recorded in the low equity details table, and at the same time check whether it is repeated with the name in other records. If it is repeated, it means that there may be multiple shareholders with the same name or the same shareholder is registered multiple times. In this case, the system can judge by the certificate number or the same contact information, and then When searching the subsidiary field downward, it is necessary to check the registration number of the subsidiary, the correspondence between the business scope and the parent company one by one. If the span between the parent company's main business scope and the subsidiary's business scope is too large, a high difference mark can be added in the system and included in the subsequent re-examination list. For the senior executives' employment information, it is necessary to check the company name, position name, and the start and end time of the employment. If the employment time span is too long or the number of positions exceeds 3, it may trigger a system warning with additional comments. For supplier and customer information, the transaction cycle, contract line number and payment time can be checked in the record to verify its authenticity. If a conflict is found with the payment time or amount of other records, it will be listed for review. After completing the cross-check of all information, a set of related party entities of the counterparty can be generated, which contains a group of shareholders, senior executives, subsidiaries and upstream and downstream partners corresponding to each counterparty. Records with inconsistent information sources or large discrepancies will also be included in this set to prevent omissions.
[0036] Based on the counterparty entity list and the related party entity set obtained above, the relationship field between each counterparty and its related party entity is first read. Taking the ownership relationship as an example, the parent-subsidiary relationship can be mapped based on the shareholder and subsidiary information to logically be a parent company shareholding or subsidiary ownership relationship. The ownership quantity can be recorded by referring to the shareholder shareholding ratio or the parent company's shareholding ratio of the subsidiary. If the share ratio is greater than 20, it is classified as a significant control category, if it is between 5 and 20, it is classified as a minor control category, and if it is less than 5, it is recorded as a trace control category. For supply chain relationships, the purchaser or seller role played by the enterprise in the same transaction event can be marked, and a judgment range can be set according to the number and amount of transactions appearing in the product or service transaction list. For example, if more than 3 orders are supplied each year and the total amount exceeds RMB 300,000, it is marked as an active supply chain relationship, otherwise it is marked as an ordinary supply chain relationship. After completing the identification of these two types of relationships, each enterprise can be regarded as an entity node and the enterprise and its shareholders can be summarized one by one in the system. , subsidiaries, suppliers, customers and executives, and record the shareholding ratio or transaction amount as part of the edge attribute. The association with executives can be stored separately in the annotation field according to the management attribute to distinguish the specific roles such as legal persons or agents for executive affairs. When establishing these relationship edges, it is necessary to avoid the generation of duplicate edges. If it is recognized that the same pair of entities have overlapping relationships in time, they are merged into one edge and the attributes are stored in parallel. If it is found that a certain enterprise is both a shareholder and a customer, multiple attributes are noted in the unified relationship edge information. If the same executive appears repeatedly in different enterprises, they are compared and classified according to their identity documents or valid contact information, and multiple nodes of the executive’s position are woven into an intersection relationship. A manual review threshold is set for possible fuzzy records. For example, entries with a shareholding ratio or transaction amount greater than RMB 1 million are set as the reexamination line. After exceeding this line, a secondary query is performed on the business license information of the enterprise or related party. Finally, all valid relationships and nodes are uniformly constructed in a basic entity relationship network.
[0037] The steps to obtain the risk transmission association view are as follows: Based on the basic entity relationship network, a corresponding weighted risk signal vector is assigned to each enterprise node to generate an enterprise node network; Based on the enterprise node network, the total credit risk impact on each node is calculated using the following formula: ; in, represents the total credit risk impact on node q, is the risk level of node s, is the set of neighbor nodes directly connected to node q, is the connection strength between node q and node s, is the logarithm of the transaction amount between node q and node s, is the number of connected nodes of node q; Based on the total credit risk impact of each node, the risk level of the node is re-evaluated and a risk transmission association view is generated.
[0038] Specifically, based on the basic entity relationship network, first select the financial and credit risk information of each enterprise node that has been obtained in the previous steps and uniformly identify the corresponding enterprise numbers, compare these enterprise numbers with the weighted risk signal vector generated previously, and check whether there are inconsistent data formats or missing fields. For example, if a node does not contain an ESG deviation value or a news sentiment fluctuation value in each dimension of the weighted risk signal vector, then record the situation and check for omissions and supplement it against the original risk data. Then, put each complete record into an enterprise node index table, set up a space that can accommodate multi-dimensional vectors for each enterprise node to save the weighted risk signal component corresponding to the node, and merge key identification fields such as node name, node ID and vector information. Then, according to a unified timestamp, the weighted risk signal value of each node is compared with the basic entity relationship network. The nodes in the data are related to the node positions in the table. For enterprises with more than 10 node connections, their ultra-high connectivity attributes can be additionally marked in the index table for subsequent identification in the visualization link. If the number of connections is less than or equal to 10, it is classified as a common node attribute. If some nodes recorded in the historical data have withdrawn from operation or been cancelled, they can be marked as frozen nodes. The corresponding weighted risk signal vectors are still retained but distinguished by different colors or symbols in the network. After completing the above processing, the node number and vector information are mapped and summarized within the system to form a set of matrix tables. The matrix table is combined with the existing entity relationship information in the same data set. Finally, the nodes that have completed the vector allocation are displayed in a network-style link, that is, each enterprise node is connected to its adjacent nodes through relationship edges, and the weighted risk signal vector of the node is attached to the network to generate an enterprise node network.
[0039] formula: The benefit of the formula is that it comprehensively considers the multiple relationship attributes between the enterprise node and the neighboring nodes, including the risk level of the neighboring nodes, as well as factors such as the connection strength and the logarithm of the transaction amount, and introduces the weighting factor caused by the number of connections of the node itself, so that the multi-dimensional risk transmission characteristics can be reflected in one formula.
[0040] The steps to obtain the parameters are as follows: The risk level values calculated in the previous stage are retrieved into a unified table in the system. The risk level is set to range from 0 to 10, corresponding to different degrees of potential credit risk. It is obtained by combining ESG scores, geographic activity, and sentiment polarity. In order to clearly quantify, ESG scores and sentiment values can be converted into a scoring system of 0 to 100, and then multiple scores can be superimposed and mapped to the range of 0 to 10 through a calculation formula. For example, for ESG scores and news sentiment comprehensive score Add and normalize to get the middle score , and then use Converted to the final range of 0 to 10 After collecting the values of each node, they are saved as a column of data. Each node has a corresponding index. In the example, a node scored 72 points in ESG score and 28 points after sentiment analysis. Add them together to 100 points and multiply by 0.1 to get 10. That is, it is marked as the highest risk level.
[0041] The steps to obtain the parameters are as follows: With Node The connection strength in the relationship network can be calculated based on the equity relationship ratio between the two parties or the frequency of cooperation in supply chain transaction events. Taking equity relationship as an example, if hold The equity ratio of the enterprise is 30% and Also holds If the two parties have a certain share, the proportion can be recorded in both directions. If the transaction frequency between the two parties in the supply chain reaches 10 times within a year, a certain score can be added. The total strength obtained by superimposing the two dimensions is included in the , for precise quantification, an additive formula can be defined ,in All need to be set in the early stage according to the industry average. In the example, if , , total equity 40%, annual transaction 10 times .
[0042] The steps to obtain the parameters are as follows: With Node The amount incurred during the transaction is calculated by taking the logarithm of the amount, which is denominated in RMB. The total amount of all transactions is collected and recorded by retrieving the previous contracts or transaction vouchers of both parties, and then get If the total transaction amount is zero, it means that there is no real capital flow, which can be listed as a zero transaction situation. During the storage process, a security check range will be set for the transaction amount. For example, the more common transaction amount range in the financial industry (100,000 to 100 million yuan) will be compared first. If an abnormally large or extremely small amount appears, historical bills need to be retrieved to confirm the validity of the data. For example, if and If the total transaction amount in the previous year is RMB 3 million, then .
[0043] The steps to obtain the parameters are as follows: The number of directly connected neighbor nodes in the network, that is, the number of nodes connected to each other in the previously built relationship network diagram. All the relationship edges are used to count how many independent nodes are connected connected, thus obtaining ,like Connected to 10 neighbors, then .
[0044] Calculation process: The first step is to select the sample node , its neighbor node set There are three nodes in , respectively obtained in the previous , , ; The second step is to bring in the logarithm of connection strength and transaction amount, such as , , ; Corresponding transaction amount logarithm (corresponding to approximately RMB 1 million), (corresponding to about RMB 100,000), (corresponding to approximately RMB 5 million); The third step is to calculate the contribution of each neighbor: , first calculate the denominator 15+1=16, ,get ; , the denominator is 5+1=6, ,get ; , denominator 25+1=26, ,get ; Step 4: Add these three items together , and then calculate , if the example node Number of connected nodes ,but , ,thus ; Step 5: Multiply 0.649 by 0.321 to get .
[0045] The results show that the node The total credit risk impact value is approximately 0.208. If it is greater than 1, it means that the overall risk of its neighboring nodes is high and the connection strength or transaction amount is relatively large. If it is less than 0.1, it means that its risk transmission is weak.
[0046] According to the total credit risk impact of each node, the values of each node are first retrieved one by one and compared with a rating cutoff standard. The rating cutoff standard can be set as follows: If the risk is less than 0.3, it is marked as low risk; if it is between 0.3 and 0.7, it is marked as medium risk; if it is greater than 0.7, it is marked as high risk. In order to determine these demarcation standards, we can refer to the past bankruptcy and default probability of enterprises in the same industry for multiple statistics, obtain the real risk events of dozens of enterprises and verify them. The corresponding reference interval is then determined. For example, if the majority of bankrupt companies in the past data The values are concentrated between 0.8 and 1.2, so 0.7 can be set as the high risk dividing line. If the values are concentrated between 0 and 0.3, 0.3 can be set as the low risk dividing line. After the dividing standard is confirmed, the calculation results of each node are classified according to the interval and its risk level is updated. If the value falls in the low risk range, you can consider re-examining the risk signal source of the node in the previous step. If the value is too high, it will be readjusted to a high-risk mark. After completing the re-evaluation of all nodes, these nodes and their interconnections will be summarized in a network diagram, and their risk level will be reflected by changing the color or mark of each node in the diagram. If there are many interconnected high-risk nodes, the local network can be reviewed again. Ultimately, all risk level information can be combined into an overall graphical perspective, and a new risk transmission association view can be formed in the system.
[0047] The steps to obtain the counterparty risk index score are as follows: Based on the risk transmission association view, the direct credit risk level of each counterparty node and the total credit risk impact on each node are extracted to form a node credit risk value set; According to the node credit risk value set, the risk index score of the counterparty node is calculated using the following formula: ; in, is the risk index score of the counterparty node, is the risk level of node k, is the total credit risk impact on node k, is the total number of relationship edges associated with counterparty node m, is the total number of nodes directly associated with the counterparty node m; Based on the risk index score of the counterparty node, the risk mark of each counterparty node in the risk transmission association view is updated one by one to generate a counterparty risk index score.
[0048] Specifically, based on the risk transmission association view, first read the direct credit risk level of each counterparty node and match the data with the total credit risk impact of each node obtained previously. Then create a unified index for each counterparty node in the internal data table, and record the credit risk level, total credit risk impact and the unique identifier of the counterparty node accordingly, so that subsequent operations can accurately trace back to the specific node. When retrieving the credit risk level, you can refer to a pre-set valid range. For example, the credit risk level is divided between 0 and 10. If the level registered for a node is higher than 10, it is necessary to check whether the node has repeated superposition or data import duplication during the previous processing. At the same time, a reference interval is also set for the total credit risk impact. For example, 0 to 2 is regarded as a normal impact, 2 to 5 is regarded as a high impact, and more than 5 is classified as an extreme impact. If the risk impact value corresponding to a node is abnormally large, it is necessary to verify its actual business volume and the number of associated nodes to confirm the integrity of the data source. When the direct credit risk level and total credit risk impact of all nodes are proofread and recorded, a set of node credit risk data is obtained in the system. Next, these data are merged into a node credit risk value set. The specific approach is to package the node ID, credit risk level and total credit risk impact in the same record, and sort each record by node ID. If the node ID is not registered in the previous index table, it is necessary to list it as a new node and retrieve its possible related corporate information or financial situation. If a duplicate node ID is found, it is combined with the above conflict troubleshooting process to uniquely merge. In order to ensure the integrity of the collection data, each node will also be checked for data completeness. For example, when a node lacks a direct credit risk level, it is necessary to query the previous credit score record again or contact the counterparty's business transaction general ledger. If the total credit risk impact is missing, it is necessary to check in the transmission association view whether the node has not participated in any related transactions or has never generated risk transmission. Once confirmed to be missing, it will be recorded with the default value of 0, and it will be marked in the record set that this node has no active transactions. At the same time, the nodes will be verified one by one in different dimensions, such as the number of related parties and whether it is identified as a geographical high-risk area in the previous step, etc., to ensure that each node can be successfully included in the credit risk value set during this aggregation process. When all verification steps are completed, a node credit risk value set that can be used in subsequent calculation links is formally formed.
[0049] formula: The benefit of the formula is that it combines the risk levels of multiple neighboring nodes with the total credit risk impact, and also considers the number of relationship edges owned by the counterparty node itself. In this way, the scale of the associated edges can be suppressed or amplified, thereby integrating the multi-dimensional elements of the counterparty node and its surrounding environment in the same expression.
[0050] The steps to obtain the parameters are to retrieve the node The risk level value stored in the previously accumulated risk control database is calculated in the same way as same.
[0051] The step of obtaining parameters is to call the node from the risk transmission association view established previously The total credit risk exposure.
[0052] The steps to obtain the parameters are as follows: Count all the relationship edges, treat all equity, supply chain or other actual business connections between enterprises as one edge, remove duplicates and add them up, the total number is , when counting, it is necessary to record different types of relationships in layers. For example, equity relationships are marked as type A edges, and supply chain cooperation is marked as type B edges. For the same enterprise, if both type A and type B exist, they are counted as two relationship edges. If there are multiple duplicate records, they are merged into one. In the example, if the node If there are 5 equity relationships and 3 supply chain relationships at the same time, then .
[0053] The steps to obtain the parameters are as follows: count the counterparty nodes The number of all directly connected neighbor nodes is calculated, and a unique index is assigned based on the node ID in the system to prevent data duplication. For example, if If there is a direct relationship with 8 entities in the network .
[0054] Calculation process: The first step is to select a counterparty node , count all directly associated nodes quantity In this example, , assuming that its neighbor nodes , query in the system separately and ; The second step is to set , , , execute on these three nodes : ; ; ; Step 3: Take the square root , and then query the node The number of associated edges In this example, ,but , in the denominator ,thus ; The results show that when The higher the value, the greater the risk index of the counterparty node. If it reaches 8 or 9 or above, it can be regarded as a relatively extreme high risk. If it is between 0 and 2, it tends to be low risk. After the values are aggregated, they can be compared vertically and the risk markers in the risk transmission association view can be updated.
[0055] Based on the risk index score of the counterparty node, first read all the completed calculations The value is compared with the existing risk information of the current node in the system. If a node was originally listed as low risk but the risk of this calculation is If the value significantly rises to above 5, it will be marked in the internal data table and the numerical comparison of the last evaluation will be noted in the list. For example, the last evaluation was 2, and now it has risen to 5, indicating that the risk performance of the node has fluctuated significantly, and it is necessary to check its business data source again in the future. If the node was originally in the medium risk range of 3 to 5 but this time it only has a risk index of 1 to 2, it will be downgraded to low risk. After completing the risk index comparison of all nodes, the updated scores will be matched with the nodes in the associated view. In the graphical interface, the different intervals can be The values are displayed through different colors or icons. Nodes with a score of 8 or above can be highlighted, and multiple transaction data can be deeply searched. When all nodes are re-scored and marked, the final counterparty risk index score of this iteration is formed.
[0056] The steps to obtain the graded warning notification are as follows: Using the counterparty risk index score, call the counterparty risk index score of each counterparty node one by one, and compare and analyze the values one by one with the pre-set grading threshold to obtain the preliminary results of risk level determination; Based on the preliminary results of risk level determination, the counterparty node is divided into risk intervals according to the degree of deviation between the counterparty risk index score of each counterparty and the classification threshold, and a counterparty node risk interval set is generated; Based on the counterparty node risk range set, the warning notification templates corresponding to each risk range are matched one by one to form a graded warning notification.
[0057] Specifically, when calling the counterparty risk index score of each transaction counterparty node one by one and comparing and analyzing the values one by one with the pre-set grading thresholds, first retrieve all the counterparty risk index scores that have been calculated in the internal database and correspond them according to the node unique identifier, and then extract the grading threshold from the pre-established threshold definition table. These thresholds can be set with reference to the historical statistical process. For example, when the counterparty risk index score is between 0 and 3, it is low, when it is between 3 and 6, it is medium, and when it is higher than 6, it is high. If the threshold setting comes from experience, the source of the experience needs to be explained. It can be based on 10 to 30 representative enterprise sample records, and the enterprises that have experienced risk events are tracked longitudinally, and the distribution range of their counterparty risk index scores before and after the incident is counted. The corresponding critical point is determined by the control data of the enterprise's real default cases. If the range of the counterparty risk index score is large, several floating coefficients can be added when setting the threshold and verified by examples. For example, enterprises whose scores have exceeded 5 for three consecutive times and have subsequently encountered credit problems are marked, so as to deduce the reasonable boundaries of the medium and high grade intervals. Next, the actual counterparty risk index score of each node is compared with this pre-grading threshold one by one. If it exceeds a certain upper limit, it is classified into the corresponding grade. If it is between two dividing values, it enters the corresponding segment. During the comparison, the difference between the score of the current node and the threshold is recorded one by one, so as to prepare for the subsequent further division of the degree of deviation. If the score of the node exceeds the highest grade interval, for example, it is greater than 8, it will be additionally marked as an extreme grade and added to the special attention list. If the score of the node is lower than the lowest grade interval, it is marked as a low grade and no further investigation will be carried out subsequently. After completing this round of matching, the counterparty risk index score of each node and its grade can be presented in the list. Those high-scoring or over-standard nodes can be quickly located by data sorting or filtering. At this point, it is summarized into a complete numerical comparison and analysis process to obtain the preliminary results of risk grade determination.
[0058] When dividing the risk interval to which the counterparty node belongs according to the degree of deviation between the counterparty risk index score of each counterparty and the grading threshold, first read the node value comparison result generated in the previous link in the data table, and select the field of the difference between the counterparty risk index score and the threshold for key processing. If the difference is greater than a certain upper and lower limit, it can be classified into a higher risk interval. If the difference is lower than another reference value, it can be classified into a lower risk interval. The specific interval division can be based on 3 to 5 range segments, for example, 0 to 2 is low risk, 2 to 5 is medium risk, 5 to 8 is high risk, and 8 or more is extreme risk. If precise judgment is required, the threshold difference can be further subdivided into multiple micro-levels. By comparing the counterparty risk index score of each node minus the corresponding threshold, the distance from the grading threshold can be quantified. For example, a difference of more than 3 can be regarded as a large deviation from the threshold. When setting the difference, the specific segmentation value can be calculated by referring to the average fluctuation of the enterprise's past failure cases and the counterparty risk index score. For example, for enterprises with long-term overdue records, if the difference between their counterparty risk index score and the threshold reaches 2, they can be included in the warning category. When all nodes are compared, interval labels are added to each node according to the above standards, such as "low risk", "medium risk", "high risk" or "extreme risk", etc., and the node ID, counterparty risk index score, and interval label are combined and recorded in a result list. If the value of a node is abnormal and exceeds the established range by a lot, continue to make notes in this step. Finally, all nodes with interval labels are aggregated into a set and retained in the system for subsequent risk notification matching to generate a counterparty node risk interval set.
[0059] When matching the warning notification templates corresponding to each risk interval one by one, the set of counterparty node risk intervals that have been delineated will be loaded into the internal table first, and the notification style with the same interval label will be retrieved in the corresponding template library. For example, a general prompt is set for the low-risk interval, a regular warning is set for the medium-risk interval, and more rigorous or intensive warning content is set for the high-risk and extreme-risk intervals. The specific template may include fields such as title, main points of the text and subsequent follow-up requirements. In order to achieve targeted processing, it is necessary to read the node interval labels one by one and match the corresponding templates. For example, when a node is found to be marked as extreme risk, the corresponding template is extracted and matched in Insert the node's identification information and risk index value into the main text, and then load its main related transaction data and possible overdue records and other details. If there is no corresponding template, a general warning format can be called. After completing the matching of the interval labels of all nodes with the notification template, a matching prompt list will be generated, which lists the template name, node ID and risk description used for each node. If the number of medium-risk nodes exceeds 20, centralized screening is required, and nodes with high-value transactions or abnormal capital flows are given priority attention. Finally, all successfully matched nodes and their corresponding warning contents are packaged and pushed to form a graded warning notification.
Claims
1. A comprehensive risk early warning system based on market big data and the credit risk of both parties to the transaction, characterized by: The system comprises: The market data aggregation module collects ESG ratings, news sentiment, geographic activity and disaster warnings, as well as counterparty business financial summary data, performs format alignment and timestamp marking, and screens out incomplete records to obtain structured market factor flows; The signal quantitative analysis module extracts credit risk text items based on the structured market factor flow according to ESG deviation values, sentiment polarity and geographic activity change rate, generates an initial signal measurement matrix, matches geographic changes with disaster areas based on the initial signal measurement matrix, associates penalty records with counterparties, sets impact weights, and obtains a weighted risk signal vector; The associated risk penetration module identifies counterparties, associated parties and relationship edges, builds a basic entity relationship network, assigns the weighted risk signal vector to nodes based on the basic entity relationship network, calculates the credit risk transmission impact along the ownership and supply chain edges, updates node metrics, and establishes a risk transmission association view; The comprehensive risk warning module, based on the risk transmission association view, summarizes the direct and transmission credit risk metrics of each counterparty node, and calculates the counterparty risk index score by weighted calculation; uses the counterparty risk index score to compare with the preset classification threshold, determines the risk range, and generates a classification warning notification.
2. The comprehensive risk early warning system based on market big data and credit risk of both parties to a transaction according to claim 1 is characterized in that: The steps for obtaining the structured market factor flow are: Collect ESG rating data, news sentiment data, geographic activity data, and disaster warning data, as well as counterparty business financial summary data to obtain a preliminary market data set; Based on the preliminary market data set, the field integrity in each record is scanned one by one, and each record is marked according to the position of the missing field. The marked results are called to remove the marked records one by one to obtain a structured market factor flow.
3. The comprehensive risk early warning system based on market big data and credit risk of both parties to a transaction according to claim 1 is characterized in that: The steps of obtaining the initial signal metric matrix are: Based on the structured market factor flow, extract the ESG rating data, news sentiment data and geographic activity data in each record, calculate the change rate and baseline deviation of each data, and generate a basic risk feature matrix for each record; Based on the basic risk characteristic matrix, the credit risk identification index of a single record is calculated using the following formula: ; in, is the credit risk identification index of a single record, is the ESG rating value of a single record, is the overall mean of ESG rating data, is the news sentiment polarity value of a single record, is the geographic activity value of a single record, is the average value of geographical activity, For the press release hour of this record, The smallest press release hour on record; Based on the credit risk identification index of each record, the corresponding original text items are evaluated, and the key text items are selected according to the value of the credit risk identification index. The key text items are combined to construct an initial signal measurement matrix.
4. The comprehensive risk early warning system based on market big data and credit risk of both parties to a transaction according to claim 1 is characterized in that: The steps of obtaining the weighted risk signal vector are: Based on the initial signal metric matrix, signal items affected by geographic changes and disaster areas are extracted, and the risk areas associated with disasters are identified in combination with real-time geographic monitoring data to obtain an initial geographic risk data set; Based on the initial geographic risk data set, a comprehensive risk metric for each data point relative to the disaster area is calculated using the following formula: ; in, is the comprehensive risk measure of data point p, is the distance from data point p to disaster area r, is the time from the disaster occurrence to the data point p being recorded, is the total number of disaster areas considered, is the angle difference between the geographical location of data point p and the disaster center; Based on the comprehensive risk metric and combined with the penalty record of each counterparty, a weighted risk signal vector is formed.
5. The comprehensive risk early warning system based on market big data and credit risk of both parties to a transaction according to claim 1 is characterized in that: The steps for obtaining the basic entity relationship network are: Based on the business financial summary data of the counterparty, retrieve the subject identification information in the business financial summary data one by one, identify the entity identities of all the counterparties one by one through the subject identification information, and generate a counterparty entity list; According to the counterparty entity list, query the information of shareholders, subsidiaries, senior executives, suppliers and customers of each counterparty entity in the industrial and commercial financial summary data, verify the authenticity of each item one by one, and generate a set of related party entities of the counterparty; Based on the counterparty entity list and the counterparty's associated entity set, the ownership relationship and supply chain relationship between each counterparty entity and the associated entity are established one by one, the ownership relationship and supply chain relationship are converted into relationship edges one by one, and the entity nodes are connected to generate a basic entity relationship network.
6. The comprehensive risk early warning system based on market big data and credit risk of both parties to a transaction according to claim 1 is characterized in that: The steps for obtaining the risk transmission association view are as follows: Based on the basic entity relationship network, a corresponding weighted risk signal vector is assigned to each enterprise node to generate an enterprise node network; Based on the enterprise node network, the total credit risk impact on each node is calculated using the following formula: ; in, represents the total credit risk impact on node q, is the risk level of node s, is the set of neighbor nodes directly connected to node q, is the connection strength between node q and node s, is the logarithm of the transaction amount between node q and node s, is the number of connected nodes of node q; Based on the total credit risk impact of each node, the risk level of the node is re-evaluated and a risk transmission association view is generated.
7. The comprehensive risk early warning system based on market big data and credit risk of both parties to a transaction according to claim 1 is characterized in that: The steps for obtaining the counterparty risk index score are as follows: Based on the risk transmission association view, extract the direct credit risk level of each counterparty node and the total credit risk impact on each node to form a node credit risk value set; According to the node credit risk value set, the risk index score of the counterparty node is calculated, and the calculation formula is: ; in, is the risk index score of the counterparty node, is the risk level of node k, is the total credit risk impact on node k, is the total number of relationship edges associated with counterparty node m, is the total number of nodes directly associated with the counterparty node m; Based on the risk index score of the counterparty node, the risk mark of each counterparty node in the risk transmission association view is updated one by one to generate a counterparty risk index score.
8. The comprehensive risk early warning system based on market big data and credit risk of both parties to a transaction according to claim 1 is characterized in that: The steps for obtaining the graded warning notification are as follows: Using the counterparty risk index score, the counterparty risk index score of each transaction counterparty node is called one by one, and the numerical comparison and analysis are performed one by one with the pre-set grading threshold value to obtain a preliminary result of risk level determination; Based on the preliminary results of the risk level determination, the counterparty node is divided into risk intervals according to the degree of deviation between the counterparty risk index score of each counterparty and the classification threshold, and a counterparty node risk interval set is generated; Based on the counterparty node risk interval set, the warning notification templates corresponding to each risk interval are matched one by one to form a graded warning notification.
Citation Information
Cited By
Bid inviting and tendering risk early warning system and method based on large model
CN120634281A