Enterprise credit supervision information verification method and system based on multi-source data

By employing dual-coefficient risk quantification and graph neural network propagation, a multi-level adaptive enterprise credit supervision information verification system is constructed. This system addresses the shortcomings of existing technologies in static assessment and implicit correlation identification, enabling full-time, full-chain enterprise credit supervision and improving the accuracy and interpretability of verification.

CN122022848APending Publication Date: 2026-05-12CHINA NAT INST OF STANDARDIZATION
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-05
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing enterprise credit supervision information verification technologies suffer from static data assessment that ignores dynamic factors, lack of implicit correlation mining, inability to locate contradictory time windows and correct data sources, resulting in delayed risk warnings and insufficient reliability of verification conclusions.

Method used

By employing dual-coefficient risk quantification, dynamic credibility fusion, graph neural network propagation, and attention attribution correction, a multi-level adaptive intelligent verification system for enterprise credit supervision information is constructed. Through the synergy of explicit and implicit verification, risk transmission and sources of responsibility are identified.

Benefits of technology

It has achieved full-domain, all-time, and full-chain enterprise credit supervision, improved the accuracy and robustness of verification results, accurately identified credit risks and sources of responsibility, and enhanced the interpretability and correctability of verification conclusions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122022848A_ABST
    Figure CN122022848A_ABST
Patent Text Reader

Abstract

The invention discloses an enterprise credit supervision information verification method and system based on multi-source data, and the method comprises the steps: carrying out the preliminary screening of data, calculating a first risk coefficient of each piece of multi-source credit supervision information, calculating a second risk coefficient according to the metadata of each piece of multi-source credit supervision information, and carrying out the dominant verification, calculating the dynamic credibility of each piece of multi-source credit supervision information, carrying out adaptive weighted fusion to obtain a multi-source data fusion decision value, constructing an enterprise credit association graph, carrying out implicit association risk propagation, calculating the association consistency index of an enterprise corresponding to each graph node, carrying out implicit verification according to a time period, and determining a contradictory time period of the multi-source credit supervision information. And calculating a contradictory attention score of each piece of multi-source credit supervision information, determining a responsibility source in a contradictory time period, and carrying out information correction. The method not only can improve the efficiency and accuracy of enterprise credit supervision information verification, but also has good interpretability, and can be directly applied to an enterprise credit supervision information verification system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of enterprise credit supervision technology, and in particular to a method and system for verifying enterprise credit supervision information based on multi-source data. Background Technology

[0002] With the deepening of the "streamlining administration, delegating power, and improving services" reform and the advancement of digital government construction, corporate credit supervision is undergoing a paradigm shift from traditional manual spot checks to data-driven and intelligent early warning. The rapid development of new-generation information technologies such as big data and artificial intelligence has provided a technical foundation for the integration of multi-source heterogeneous data for corporate credit supervision, making it possible to integrate credit information across departments and levels. Currently, the convergence and cross-verification of multi-dimensional credit supervision information from industry and commerce, taxation, judiciary, and public opinion have become key means to improve the accuracy of supervision and prevent systemic risks.

[0003] However, existing technologies for verifying corporate credit regulatory information still have significant limitations: First, traditional verification methods often employ static data quality assessments, neglecting dynamic contextual factors such as market fluctuations and policy adjustments at the time of data generation, leading to delayed risk warnings; second, existing technologies focus on isolated assessments of a single company's credit status, lacking in-depth analysis of implicit relationships between companies, making it difficult to identify the chain reaction of risk transmission across entities; finally, the detection of contradictions between multi-source data often stops at the immediate resolution of field-level conflicts, failing to pinpoint the time window in which the contradictions arose, let alone trace back to specific data sources for precise correction, resulting in insufficient reliability of verification conclusions. Therefore, this invention proposes a method and system for verifying corporate credit regulatory information based on multi-source data. Through dual-coefficient risk quantification, dynamic credibility fusion, graph neural network propagation, and attention attribution correction, a multi-level, adaptive, and interpretable intelligent verification system for corporate credit regulatory information is constructed. This system enables full-domain, all-time, and full-chain verification of corporate credit regulatory information, providing a practical and effective technical solution for addressing the complex and ever-changing business environment and increasingly hidden credit risks. Summary of the Invention

[0004] The purpose of this invention is to provide a method and system for verifying enterprise credit supervision information based on multi-source data.

[0005] To achieve the above objectives, the present invention is implemented according to the following technical solution: This invention includes the following steps: Obtain multi-source credit supervision information and metadata of enterprises, perform initial data screening to calculate the first risk coefficient of each multi-source credit supervision information, calculate the second risk coefficient based on the metadata of each multi-source credit supervision information, and conduct explicit verification according to the first and second risk coefficients by time period. Calculate the dynamic credibility of each multi-source credit supervision information, and adaptively weight and fuse each multi-source credit supervision information according to the dynamic credibility to obtain the multi-source data fusion decision value; Based on the multi-source data fusion decision values ​​and basic attributes of each enterprise, an enterprise credit association graph is constructed. Multi-source credit supervision information is continuously acquired to spread implicit association risks. The association consistency index of each graph node corresponding to the enterprise is calculated and implicitly verified according to time period. Based on the explicit and implicit verification results, the time period of conflict in multi-source credit supervision information is determined. The conflict attention score of each multi-source credit supervision information is calculated based on the first risk coefficient, the second risk coefficient, and the dynamic credibility. The source of responsibility within the conflict time period is determined, and the information is corrected. The multi-source credit supervision information includes industrial and commercial supervision information, tax supervision information, judicial supervision information, and public opinion supervision information; the metadata includes the market environment and data operation information at the time of generation of the credit supervision information; the first risk coefficient is negatively correlated with the data quality of the multi-source credit supervision information; the second risk coefficient is negatively correlated with the generation environment of the multi-source credit supervision information.

[0006] Furthermore, the method for calculating the first risk coefficient of each multi-source credit regulatory information includes: The types and quantities of blank data in various multi-source credit supervision information are statistically analyzed; the blank data types include critical blank data and non-critical blank data. Potentially abnormal data is identified by filtering various multi-source credit supervision information based on standard data value ranges, business rules, and standard data formats; the potentially abnormal data includes statistically abnormal data, logically abnormal data, and format-abnormal data. The first risk coefficient for each source of credit regulatory information is calculated based on statistical indicators of blank data and potentially abnormal data, expressed as follows: ; ; in for The primary risk factor for data source credit regulatory information is... For scale parameters, For translation parameters, As the first risk score, , , As the first risk weight, for Data source Weights for blank data for Data source Blank data blank quantity, for Data source Total number of data types for Total number of potentially anomalous data from the data source for Data source Potentially abnormal data Abnormality score, This is the anomaly score mapping function. for The difference between the timestamp of the data source corresponding to the time period and the current verification time. To update the delay risk threshold.

[0007] Furthermore, the method for calculating the second risk coefficient includes: Extract market environment and data operation information from various multi-source credit supervision information; the market environment includes industry credit deviation, macroeconomic volatility, and enterprise relative size risk; the data operation information includes collection duration, process complexity, abnormal access density, correction frequency, and operator trust level. The environmental risk of corporate credit regulatory information is calculated based on the market environment, expressed as follows: ; in Environmental risks related to corporate credit supervision information , , , , As a weight of the market environment, for Corporate credit score during the assessment period , The long-term historical average score and standard deviation for the entire industry. for Macroeconomic indicator values ​​for the assessment period Macroeconomic indicator values The long-term trend value, The standard deviation of macroeconomic indicator values. This refers to the company's size and position within the industry. As an index of industry public opinion, As a regional judicial activity index; The operational risk of each multi-source credit supervision information is calculated based on the data operation information, expressed as follows: ; in for Operational risks associated with data source credit regulatory information. , , , , Weights for data operations, for Data collection time per data session This is the historical average time for similar tasks. for The number of system interfaces and manual approval steps involved in the data source collection process. for Data source abnormal access density, for Number of data source version changes Operator trust level; The second risk coefficient for each multi-source credit supervision information is calculated based on the environmental risk of enterprise credit supervision information and the operational risk of each multi-source credit supervision information. The expression is as follows: ; ; in for The second risk factor of data source credit regulatory information, For the second risk score, For scaling parameters, For curvature parameters, For curvature parameters, It is a very small constant.

[0008] Furthermore, the method for explicit verification by time period includes: Calculate the average of the first risk coefficient and the average of the second risk coefficient for all multi-source credit supervision information within the same time period. Take the product of the average of the first risk coefficient and the average of the second risk coefficient as the explicit risk score. When the explicit risk score is greater than the explicit risk threshold, it is determined that the multi-source credit supervision information for the corresponding time period has failed explicit verification.

[0009] Furthermore, the method for obtaining multi-source data fusion decision values ​​includes: The dynamic credibility of each multi-source credit supervision information is calculated based on its generation time, basic credibility rating, and historical consistency. Then, the multi-source credit supervision information is adaptively weighted and fused according to this dynamic credibility to obtain the multi-source data fusion decision value, expressed as: ; ; ; in for Time period Enterprises' multi-source data fusion decision-making value Total number of data sources for Dynamic reliability of the data source for Business weight of the data source for Time period enterprise Data source: Enterprise credit supervision information For dynamic weights, for Basic credibility rating of the data source for The time sensitivity of the data source For the verification period, for The generation time of data source credit supervision information The historical consistency correction factor is calculated using an exponentially weighted moving average. To adjust the weights, for Time period Weighted median of multi-source credit supervision information for enterprises It is a very small constant.

[0010] Furthermore, the method of implicit verification by time period includes: Graph nodes are set up according to each enterprise, and the association edges of each graph node are set according to the enterprise relationships. The basic attributes of each enterprise and the multi-source data fusion decision value are used as the node representation of the corresponding graph node to construct an enterprise credit association graph. The enterprise relationships include equity relationships, guarantee relationships, transaction relationships, and personnel relationships. The attention coefficients of the graph nodes corresponding to the target enterprise and the graph nodes corresponding to the neighboring enterprises of related enterprises are calculated and normalized. Multi-source credit supervision information of each enterprise is continuously acquired, and a graph attention network is used to propagate implicit association risks. The expression is as follows: ; ; in For graph nodes and Attention coefficient For graph nodes and Prior weights corresponding to firm relationships, For attention parameter vectors, For learnable parameter matrix, , For node representation vectors, For relational projection matrices, Embed vectors for relation types. For graph nodes updated after the attention layer exist The layer's representation vector, For graph nodes The set of neighboring nodes, For graph nodes and Normalized attention coefficient exist Layer representation, for The learnable parameter matrix of the layer, It is a non-linear activation function; Based on the updated enterprise credit association graph, calculate the association consistency index of the enterprises corresponding to each graph node, as expressed in the following expression: ; in For graph nodes The correlation consistency index, For graph nodes The set of second-order neighbors, For graph nodes and Value deviation, , For graph nodes and Multi-source data fusion decision values For graph nodes and The closeness of related-party transactions between the corresponding enterprises; When a company's correlation consistency index is greater than the implicit risk threshold, it is determined that the company's multi-source credit supervision information for the corresponding time period has failed the implicit verification.

[0011] Furthermore, the method for determining the source of responsibility within the conflicting time period includes: The time periods during which explicit or implicit verification fails are defined as conflicting periods for multi-source credit regulatory information. The conflict attention score for each multi-source credit regulatory information within a conflicting period is calculated based on the first risk coefficient, the second risk coefficient, and dynamic credibility. The expression is as follows: ; in Contradictory time period Inside Contradictory attention scores of the data source , Contradictory time period Inside The first and second risk coefficients of the data source To ensure the dynamic credibility of corresponding credit supervision information, In order to correspond credit supervision information Query vector of contradictory state The correlation score A collection of data sources; When the contradiction attention score is greater than the contradiction attention threshold, the corresponding data source is determined to be the source of responsibility for the enterprise's credit risk within the target contradiction time period, and manual verification is introduced to correct the information.

[0012] Secondly, the enterprise credit supervision information verification system based on multi-source data includes: Explicit verification module: used to perform initial data screening, calculate the first risk coefficient of each multi-source credit supervision information, calculate the second risk coefficient based on the metadata of each multi-source credit supervision information, and perform explicit verification according to the first and second risk coefficients by time period. Fusion module: used to calculate the dynamic credibility of each multi-source credit supervision information, and adaptively weight and fuse each multi-source credit supervision information according to the dynamic credibility to obtain the multi-source data fusion decision value; Implicit Verification Module: Used to construct enterprise credit association graphs, continuously acquire multi-source credit supervision information to spread implicit association risks, calculate the association consistency index of enterprises corresponding to each graph node, and perform implicit verification over time periods; The responsibility source tracing module is used to determine the time period of conflict between multi-source credit supervision information. Based on the first risk coefficient, the second risk coefficient, and the dynamic credibility, it calculates the conflict attention score of each multi-source credit supervision information, determines the source of responsibility within the conflict time period, and corrects the information.

[0013] The beneficial effects of this invention are: This invention relates to a method and system for verifying enterprise credit supervision information based on multi-source data. Compared with existing technologies, this invention has the following technical advantages: This invention introduces a first risk coefficient and a second risk coefficient, and combines the two to calculate the dynamic credibility of the data source through adaptive weighted fusion. This enables the verification system to perceive and respond to fluctuations in the reliability of the data source in real time, thereby improving the accuracy and robustness of the verification results. This invention constructs a dual-mode collaborative verification mechanism by combining explicit verification (consistency analysis based on time slices of multi-source data from a single enterprise) and implicit verification (risk propagation simulation based on enterprise credit association diagrams). This mechanism breaks through the dimensional limitations of single-subject assessment and achieves a leap from point-based supervision to network-based supervision. This invention proposes a responsibility source tracing method based on contradiction time period location and contradiction attention score calculation. It breaks through the limitation of traditional contradiction resolution that "only corrects errors but does not determine responsibility". It can accurately identify the source of responsibility that leads to verification conflict, enhance the interpretability and correctability of verification conclusions, and build a multi-level, adaptive, and interpretable intelligent verification system for enterprise credit supervision information. It provides a practical and effective technical solution for dealing with the complex and ever-changing business environment and increasingly hidden credit risks. Attached Figure Description

[0014] Figure 1 This is a flowchart illustrating the steps of the enterprise credit supervision information verification method based on multi-source data according to the present invention. Detailed Implementation

[0015] The present invention will be further described below through specific embodiments. The illustrative embodiments and descriptions herein are used to explain the present invention, but are not intended to limit the present invention.

[0016] The present invention provides a method and system for verifying enterprise credit supervision information based on multi-source data, comprising the following steps: like Figure 1 As shown, this embodiment includes the following steps: Obtain multi-source credit supervision information and metadata of enterprises, perform initial data screening to calculate the first risk coefficient of each multi-source credit supervision information, calculate the second risk coefficient based on the metadata of each multi-source credit supervision information, and conduct explicit verification according to the first and second risk coefficients by time period. Calculate the dynamic credibility of each multi-source credit supervision information, and adaptively weight and fuse each multi-source credit supervision information according to the dynamic credibility to obtain the multi-source data fusion decision value; Based on the multi-source data fusion decision values ​​and basic attributes of each enterprise, an enterprise credit association graph is constructed. Multi-source credit supervision information is continuously acquired to spread implicit association risks. The association consistency index of each graph node corresponding to the enterprise is calculated and implicitly verified according to time period. Based on the explicit and implicit verification results, the time period of conflict in multi-source credit supervision information is determined. The conflict attention score of each multi-source credit supervision information is calculated based on the first risk coefficient, the second risk coefficient, and the dynamic credibility. The source of responsibility within the conflict time period is determined, and the information is corrected. The multi-source credit supervision information includes industrial and commercial supervision information, tax supervision information, judicial supervision information, and public opinion supervision information; the metadata includes the market environment and data operation information at the time of generation of the credit supervision information; the first risk coefficient is negatively correlated with the data quality of the multi-source credit supervision information; the second risk coefficient is negatively correlated with the generation environment of the multi-source credit supervision information.

[0017] In this embodiment, the method for calculating the first risk coefficient of each multi-source credit supervision information includes: The types and quantities of blank data in various multi-source credit supervision information are statistically analyzed; the blank data types include critical blank data and non-critical blank data. Potentially abnormal data is identified by filtering various multi-source credit supervision information based on standard data value ranges, business rules, and standard data formats; the potentially abnormal data includes statistically abnormal data, logically abnormal data, and format-abnormal data. The first risk coefficient for each source of credit regulatory information is calculated based on statistical indicators of blank data and potentially abnormal data, expressed as follows: ; ; in for The primary risk factor for data source credit regulatory information is... For scale parameters, For translation parameters, As the first risk score, , , As the first risk weight, for Data source Weights for blank data for Data source Blank data blank quantity, for Data source Total number of data types for Total number of potentially anomalous data from the data source for Data source Potentially abnormal data Abnormality score, This is the anomaly score mapping function. for The difference between the timestamp of the data source corresponding to the time period and the current verification time. To update the delay risk threshold; In actual assessments, blank data refers to missing or empty data in fields that should be collected. Based on their nature and impact, blank data can be divided into: critical blank data (missing fields that have a decisive impact on credit judgment, such as a company's registered capital, legal representative / industrial and commercial regulatory information, taxpayer status / tax regulatory information, amount of enforcement target / judicial regulatory information, entity identity confirmation / public opinion monitoring information) and non-critical blank data (missing auxiliary or descriptive fields, such as business scope / industrial and commercial regulatory information, industry details / tax regulatory information, case details / judicial regulatory information, etc.). Potentially abnormal data refers to data whose field values ​​exist but clearly exceed the reasonable range, violate business logic, or exhibit extreme anomalies. It can be divided into: statistically abnormal data (numerical data deviates from the normal range, such as a negative number of years of establishment or a revenue growth rate exceeding 1000%), logically abnormal data (contradictory relationships between data, such as total profit being much greater than operating revenue), and format-abnormal data (text data does not conform to established standards, such as an incorrect number of digits in the unified social credit code). When calculating the first risk coefficient, credit regulatory information from different data sources is processed in batches according to the data generation time period, and anomaly scoring mapping function is used. Based on rules and statistics, the following scoring methods are used: 1) Score the potential abnormal data based on the degree to which it deviates from the normal range (e.g., 0.5 for deviations of 3-5 standard deviations, 1.0 for deviations of more than 5 standard deviations); 2) Score the potential abnormal data based on the importance of the business rules it violates (e.g., 1.0 for errors in core financial reconciliation relationships, 0.5 for minor relationship contradictions); 3) Score the potential abnormal data based on the criticality of the field it is located in (e.g., 1.0 for incorrect format of key ID, 0.2 for poor format of description field).

[0018] In this embodiment, the method for calculating the second risk coefficient includes: Extract market environment and data operation information from various multi-source credit supervision information; the market environment includes industry credit deviation, macroeconomic volatility, and enterprise relative size risk; the data operation information includes collection duration, process complexity, abnormal access density, correction frequency, and operator trust level. The environmental risk of corporate credit regulatory information is calculated based on the market environment, expressed as follows: ; in Environmental risks related to corporate credit supervision information , , , , As a weight of the market environment, for Corporate credit score during the assessment period , The long-term historical average score and standard deviation for the entire industry. for Macroeconomic indicator values ​​for the assessment period Macroeconomic indicator values The long-term trend value, The standard deviation of macroeconomic indicator values. This refers to the company's size and position within the industry. As an index of industry public opinion, As a regional judicial activity index; The operational risk of each multi-source credit supervision information is calculated based on the data operation information, expressed as follows: ; in for Operational risks associated with data source credit regulatory information. , , , , Weights for data operations, for Data collection time per data session This is the historical average time for similar tasks. for The number of system interfaces and manual approval steps involved in the data source collection process. for Data source abnormal access density, for Number of data source version changes Operator trust level; The second risk coefficient for each multi-source credit supervision information is calculated based on the environmental risk of enterprise credit supervision information and the operational risk of each multi-source credit supervision information. The expression is as follows: ; ; in for The second risk factor of data source credit regulatory information, For the second risk score, For scaling parameters, For curvature parameters, For curvature parameters, It is a very small constant; In actual assessments, when calculating the second risk coefficient, credit regulatory information from different data sources is processed in batches according to the data generation time period. First, calculate the environmental risk of corporate credit regulatory information based on the market environment. Industry credit deviation is the standardized deviation of the target company's credit score from the long-term historical average credit score of the entire industry (when the overall credit of the industry declines significantly, the negative data of a single company may be amplified). Macroeconomic volatility, macroeconomic indicator values long-term trend value The moving average reflects the deviation of the key indicator of overall economic activity (GDP growth rate) from its long-term trend (during periods of severe economic fluctuations, abnormal business data may be more due to systemic shocks than individual credit issues). Risk related to the relative size of the enterprise, based on the enterprise's size position within the industry (the enterprise's size position within the industry). Assess its data resilience by ranking it by percentile of total assets (SMEs' data volatility is typically higher than that of large enterprises, and their data is more easily distorted in volatile environments); Industry public opinion heat index The regional judicial activity index is determined by calculating the ratio of the industry's public opinion volume (news, social media platforms) to the base period during the corresponding time period (when the popularity is abnormally high, the signal-to-noise ratio of a single public opinion message may decrease); The deviation of the number of enterprise-related litigation cases in the region at time t from the historical average is determined (a surge in cases may reflect regional systemic disputes). The operational risks of each data source of credit regulatory information were then calculated separately. To standardize the collection duration, To standardize process complexity, For standardization correction frequency; Data source abnormal access density Operator trust level is obtained by standardizing the number of non-routine queries within a time window before and after data generation (using historical baseline). Operator history violation records Sure.

[0019] In this embodiment, the method for explicit verification by time period includes: Calculate the average of the first risk coefficient and the average of the second risk coefficient for all multi-source credit supervision information within the same time period. Take the product of the average of the first risk coefficient and the average of the second risk coefficient as the explicit risk score. When the explicit risk score is greater than the explicit risk threshold, it is determined that the multi-source credit supervision information for the corresponding time period has failed explicit verification.

[0020] In this embodiment, the method for obtaining multi-source data fusion decision values ​​includes: The dynamic credibility of each multi-source credit supervision information is calculated based on its generation time, basic credibility rating, and historical consistency. Then, the multi-source credit supervision information is adaptively weighted and fused according to this dynamic credibility to obtain the multi-source data fusion decision value, expressed as: ; ; ; in for Time period Enterprises' multi-source data fusion decision-making value Total number of data sources for Dynamic reliability of the data source for Business weight of the data source for Time period enterprise Data source: Enterprise credit supervision information For dynamic weights, for Basic credibility rating of the data source for The time sensitivity of the data source For the verification period, for The generation time of data source credit supervision information The historical consistency correction factor is calculated using an exponentially weighted moving average. To adjust the weights, for Time period Weighted median of multi-source credit supervision information for enterprises It is a very small constant; In actual assessments, when calculating the dynamic credibility of various multi-source credit supervision information, the basic credibility rating of each data source is determined by the data source generating unit (e.g., for tax data corresponding to the tax bureau, the basic credibility rating is 0.95; for public opinion data corresponding to media platforms, the basic credibility rating is 0.6); the timeliness sensitivity coefficient... Related to data attributes (e.g., tax data is generated regularly and quantitatively by professional institutions, so it is less affected by time and has a lower time sensitivity coefficient; public opinion data is generated randomly and must have high timeliness, so it has a higher time sensitivity coefficient); the business weight of the data source is determined based on the degree of influence of the data source on credit (e.g., the weight of tax abnormalities on credit is higher than that of water and electricity payments).

[0021] In this embodiment, the method for implicit verification by time period includes: Graph nodes are set up according to each enterprise, and the association edges of each graph node are set according to the enterprise relationships. The basic attributes of each enterprise and the multi-source data fusion decision value are used as the node representation of the corresponding graph node to construct an enterprise credit association graph. The enterprise relationships include equity relationships, guarantee relationships, transaction relationships, and personnel relationships. The attention coefficients of the graph nodes corresponding to the target enterprise and the graph nodes corresponding to the neighboring enterprises of related enterprises are calculated and normalized. Multi-source credit supervision information of each enterprise is continuously acquired, and a graph attention network is used to propagate implicit association risks. The expression is as follows: ; ; in For graph nodes and Attention coefficient For graph nodes and Prior weights corresponding to firm relationships, For attention parameter vectors, For learnable parameter matrix, , For node representation vectors, For relational projection matrices, Embed vectors for relation types. For graph nodes updated after the attention layer exist The layer's representation vector, For graph nodes The set of neighboring nodes, For graph nodes and Normalized attention coefficient exist Layer representation, for The learnable parameter matrix of the layer, It is a non-linear activation function; Based on the updated enterprise credit association graph, calculate the association consistency index of the enterprises corresponding to each graph node, as expressed in the following expression: ; in For graph nodes The correlation consistency index, For graph nodes The set of second-order neighbors, For graph nodes and Value deviation, , For graph nodes and Multi-source data fusion decision values For graph nodes and The closeness of related-party transactions between the corresponding enterprises; When a company’s correlation consistency index is greater than the implicit risk threshold, it is determined that the company’s multi-source credit supervision information for the corresponding time period has failed the implicit verification. In actual evaluation, when calculating the attention coefficient between nodes in each graph, the relationship type embedding vector is used to distinguish different semantic associations such as equity / guarantee / transaction. The prior weights are preset values ​​determined by the enterprise based on the enterprise relationship (e.g., risk transmission coefficient of 0.9 for guarantee relationship and 0.3 for ordinary transaction relationship). The calculated attention coefficients are normalized into a probability distribution through softmax. The association consistency index of each node in the graph is used to measure the consistency of the credit status of the enterprise with other enterprises in its association network. The larger the value, the more suspicious it is (when the value is close to 0, the credit status of the enterprise and its association network is homogeneous and there is no abnormal transmission; when the value is close to 1, it is completely separated from the association network, and there may be data fraud or malicious isolation). Among them, the closeness of related transactions is determined by the proportion of related transaction amount to the enterprise's revenue (reflecting the strength of association).

[0022] In this embodiment, the method for determining the source of responsibility within the conflicting time period includes: The time periods during which explicit or implicit verification fails are defined as conflicting periods for multi-source credit regulatory information. The conflict attention score for each multi-source credit regulatory information within a conflicting period is calculated based on the first risk coefficient, the second risk coefficient, and dynamic credibility. The expression is as follows: ; in Contradictory time period Inside Contradictory attention scores of the data source , Contradictory time period Inside The first and second risk coefficients of the data source To ensure the dynamic credibility of corresponding credit supervision information, In order to correspond credit supervision information Query vector of contradictory state The correlation score A collection of data sources; When the contradiction attention score is greater than the contradiction attention threshold, the corresponding data source is determined to be the source of responsibility for the enterprise's credit risk within the target contradiction time period, and manual verification is introduced to correct the information. In actual assessment, taking the credit supervision information verification of X Technology Co., Ltd. in the first quarter of 2023 (January 1, 2023 to March 31, 2023) as an example, the specific implementation process of this invention is illustrated as follows: With a scaling parameter of 2, a translation parameter of 0.5, and a first risk weight of 0.4 / 0.4 / 0.2, for tax supervision information, the types and quantities of blank data are statistically analyzed to identify potential abnormal data. Considering the data update delay, the first risk coefficient of this data source in the first quarter of 2023 is calculated to be 0.75. With market environment weights of 0.3 / 0.25 / 0.2 / 0.15 / 0.1, data operation weights of 0.3 / 0.25 / 0.2 / 0.15 / 0.1, scaling parameter of 1.2, and curvature parameter of 0.8 / 0.6, the environmental risk of enterprise credit supervision information for this period is calculated to be 0.7, and the operational risk of tax supervision information is calculated to be 0.65. Further calculation yields a second risk coefficient of 0.68 for this data source. The explicit risk threshold is set at 0.3. The product of the average first risk coefficient (0.72) and the average second risk coefficient (0.65) of all data sources (industrial and commercial, tax, judicial, and public opinion) within this time period is calculated to be 0.468. It is determined that the enterprise credit supervision information within this time period has failed explicit verification. Set a dynamic weight of 0.7 and a correction weight of 0.8, calculate the dynamic credibility of each multi-source credit supervision information, and adaptively weight and fuse the information from each data source according to the dynamic credibility to obtain the multi-source data fusion decision value (vector form) of the enterprise in the first quarter of 2023. The basic attributes of each enterprise (such as industry, registered capital, and years of establishment) and the decision values ​​of multi-source data fusion are used as the node representations of the corresponding graph nodes to construct an enterprise credit association graph. Graph Attention Network (GAT) is used to propagate implicit association risks. Based on the updated enterprise credit association graph, the association consistency index of each graph node is calculated. The association consistency index of the target enterprise (XX Technology Co., Ltd.) is calculated to be 0.78 (the implicit risk threshold is 0.6), indicating that the enterprise credit supervision information has not passed the implicit verification during this period. Since this time period failed both explicit and implicit verification, it was determined to be a contradictory time period. For the tax supervision information, based on the first risk coefficient of 0.75, the second risk coefficient of 0.68, the dynamic credibility of 0.62, and the correlation score between the tax supervision information and the query vector of the current contradictory state of 2.3, its contradiction attention score was calculated to be 1.87 (greater than the contradiction attention threshold of 0.8). The tax supervision information was determined to be the source of responsibility for the enterprise's credit risk within this contradictory time period, and the system will prompt the introduction of manual verification to correct the information.

[0023] Secondly, the enterprise credit supervision information verification system based on multi-source data includes: Explicit verification module: used to perform initial data screening, calculate the first risk coefficient of each multi-source credit supervision information, calculate the second risk coefficient based on the metadata of each multi-source credit supervision information, and perform explicit verification according to the first and second risk coefficients by time period. Fusion module: used to calculate the dynamic credibility of each multi-source credit supervision information, and adaptively weight and fuse each multi-source credit supervision information according to the dynamic credibility to obtain the multi-source data fusion decision value; Implicit Verification Module: Used to construct enterprise credit association graphs, continuously acquire multi-source credit supervision information to spread implicit association risks, calculate the association consistency index of enterprises corresponding to each graph node, and perform implicit verification over time periods; The responsibility source tracing module is used to determine the time period of conflict between multi-source credit supervision information. Based on the first risk coefficient, the second risk coefficient, and the dynamic credibility, it calculates the conflict attention score of each multi-source credit supervision information, determines the source of responsibility within the conflict time period, and corrects the information.

[0024] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for verifying enterprise credit supervision information based on multi-source data, characterized in that, Includes the following steps: S1. Obtain multi-source credit supervision information and metadata of enterprises, perform preliminary data screening and calculate the first risk coefficient of each multi-source credit supervision information, calculate the second risk coefficient based on the metadata of each multi-source credit supervision information, and perform explicit verification according to the first risk coefficient and the second risk coefficient by time period. S2. Calculate the dynamic credibility of each multi-source credit supervision information, and adaptively weight and fuse each multi-source credit supervision information according to the dynamic credibility to obtain the multi-source data fusion decision value. S3. Construct an enterprise credit association graph based on the multi-source data fusion decision value and basic attributes of each enterprise, continuously acquire multi-source credit supervision information to spread implicit association risks, and calculate the association consistency index of each graph node corresponding to the enterprise for implicit verification over time periods. S4. Determine the time period of conflict between multi-source credit supervision information based on explicit and implicit verification results. Calculate the conflict attention score of each multi-source credit supervision information based on the first risk coefficient, the second risk coefficient, and dynamic credibility. Determine the source of responsibility within the conflict time period and correct the information accordingly. The multi-source credit supervision information includes industrial and commercial supervision information, tax supervision information, judicial supervision information, and public opinion supervision information; the metadata includes the market environment and data operation information at the time of generation of the credit supervision information; the first risk coefficient is negatively correlated with the data quality of the multi-source credit supervision information; The second risk coefficient is negatively correlated with the multi-source credit regulatory information generation environment.

2. The method for verifying enterprise credit supervision information based on multi-source data according to claim 1, characterized in that, The method for calculating the first risk coefficient of each multi-source credit supervision information includes: The types and quantities of blank data in various multi-source credit supervision information are statistically analyzed; the blank data types include critical blank data and non-critical blank data. Potentially abnormal data is identified by filtering various multi-source credit supervision information based on standard data value ranges, business rules, and standard data formats; the potentially abnormal data includes statistically abnormal data, logically abnormal data, and format-abnormal data. The first risk coefficient for each source of credit regulatory information is calculated based on statistical indicators of blank data and potentially abnormal data, expressed as follows: ; ; in for The primary risk factor for data source credit regulatory information is... For scale parameters, For translation parameters, As the first risk score, , , As the first risk weight, for Data source Weights for blank data for Data source Blank data blank quantity, for Data source Total number of data types for Total number of potentially anomalous data from the data source for Data source Potentially abnormal data Abnormality score, This is the anomaly score mapping function. for The difference between the timestamp of the data source corresponding to the time period and the current verification time. To update the delay risk threshold.

3. The method for verifying enterprise credit supervision information based on multi-source data according to claim 1, characterized in that, The method for calculating the second risk coefficient includes: Extract market environment and data operation information from various multi-source credit supervision information; the market environment includes industry credit deviation, macroeconomic volatility, and enterprise relative size risk; the data operation information includes collection duration, process complexity, abnormal access density, correction frequency, and operator trust level. The environmental risk of corporate credit regulatory information is calculated based on the market environment, expressed as follows: ; in Environmental risks related to corporate credit supervision information , , , , As a weight of the market environment, for Corporate credit score during the assessment period , The long-term historical average score and standard deviation for the entire industry. for Macroeconomic indicator values ​​for the assessment period Macroeconomic indicator values The long-term trend value, The standard deviation of macroeconomic indicator values. This refers to the company's size and position within the industry. As an index of industry public opinion, As a regional judicial activity index; The operational risk of each multi-source credit supervision information is calculated based on the data operation information, expressed as follows: ; in for Operational risks associated with data source credit regulatory information. , , , , Weights for data operations, for Data collection time per data session This is the historical average time for similar tasks. for The number of system interfaces and manual approval steps involved in the data source collection process. for Data source abnormal access density, for Number of data source version changes Operator trust level; The second risk coefficient for each multi-source credit supervision information is calculated based on the environmental risk of enterprise credit supervision information and the operational risk of each multi-source credit supervision information. The expression is as follows: ; ; in for The second risk factor of data source credit regulatory information, For the second risk score, For scaling parameters, For curvature parameters, For curvature parameters, It is a very small constant.

4. The method for verifying enterprise credit supervision information based on multi-source data according to claim 1, characterized in that, The method for explicit verification by time period includes: Calculate the average of the first risk coefficient and the average of the second risk coefficient for all multi-source credit supervision information within the same time period. Take the product of the average of the first risk coefficient and the average of the second risk coefficient as the explicit risk score. When the explicit risk score is greater than the explicit risk threshold, it is determined that the multi-source credit supervision information for the corresponding time period has failed explicit verification.

5. The method for verifying enterprise credit supervision information based on multi-source data according to claim 1, characterized in that, The method for obtaining multi-source data fusion decision values ​​includes: The dynamic credibility of each multi-source credit supervision information is calculated based on its generation time, basic credibility rating, and historical consistency. Then, the multi-source credit supervision information is adaptively weighted and fused according to this dynamic credibility to obtain the multi-source data fusion decision value, expressed as: ; ; ; in for Time period Enterprises' multi-source data fusion decision-making value Total number of data sources for Dynamic reliability of the data source for Business weight of the data source for Time period enterprise Data source: Enterprise credit supervision information For dynamic weights, for Basic credibility rating of the data source for The time sensitivity of the data source For the verification period, for The generation time of data source credit supervision information The historical consistency correction factor is calculated using an exponentially weighted moving average. To adjust the weights, for Time period Weighted median of multi-source credit supervision information for enterprises It is a very small constant.

6. The method for verifying enterprise credit supervision information based on multi-source data according to claim 1, characterized in that, The method for implicit verification based on time periods includes: Graph nodes are set up according to each enterprise, and the association edges of each graph node are set according to the enterprise relationships. The basic attributes of each enterprise and the multi-source data fusion decision value are used as the node representation of the corresponding graph node to construct an enterprise credit association graph. The enterprise relationships include equity relationships, guarantee relationships, transaction relationships, and personnel relationships. The attention coefficients of the graph nodes corresponding to the target enterprise and the graph nodes corresponding to the neighboring enterprises of related enterprises are calculated and normalized. Multi-source credit supervision information of each enterprise is continuously acquired, and a graph attention network is used to propagate implicit association risks. The expression is as follows: ; ; in For graph nodes and Attention coefficient For graph nodes and Prior weights corresponding to firm relationships, For attention parameter vectors, For learnable parameter matrix, , For node representation vectors, For relational projection matrices, Embed vectors for relation types. For graph nodes updated after the attention layer exist The layer's representation vector, For graph nodes The set of neighboring nodes, For graph nodes and Normalized attention coefficient exist Layer representation, for The learnable parameter matrix of the layer, It is a non-linear activation function; Based on the updated enterprise credit association graph, calculate the association consistency index of the enterprises corresponding to each graph node, as expressed in the following expression: ; in For graph nodes The correlation consistency index, For graph nodes The set of second-order neighbors, For graph nodes and Value deviation, , For graph nodes and Multi-source data fusion decision values For graph nodes and The closeness of related-party transactions between the corresponding enterprises; When a company's correlation consistency index is greater than the implicit risk threshold, it is determined that the company's multi-source credit supervision information for the corresponding time period has failed the implicit verification.

7. The method for verifying enterprise credit supervision information based on multi-source data according to claim 1, characterized in that, The method for determining the source of responsibility within the time period of the conflict includes: The time periods during which explicit or implicit verification fails are defined as conflicting periods for multi-source credit regulatory information. The conflict attention score for each multi-source credit regulatory information within a conflicting period is calculated based on the first risk coefficient, the second risk coefficient, and dynamic credibility. The expression is as follows: ; in Contradictory time period Inside Contradictory attention scores of the data source , Contradictory time period Inside The first and second risk coefficients of the data source To ensure the dynamic credibility of corresponding credit supervision information, In order to correspond credit supervision information Query vector of contradictory state The correlation score A collection of data sources; When the contradiction attention score is greater than the contradiction attention threshold, the corresponding data source is determined to be the source of responsibility for the enterprise's credit risk within the target contradiction time period, and manual verification is introduced to correct the information.

8. A corporate credit supervision information verification system based on multi-source data, used to execute the method described in any one of claims 1-7, characterized in that, include: Explicit verification module: used to perform initial data screening, calculate the first risk coefficient of each multi-source credit supervision information, calculate the second risk coefficient based on the metadata of each multi-source credit supervision information, and perform explicit verification according to the first and second risk coefficients by time period. Fusion module: used to calculate the dynamic credibility of each multi-source credit supervision information, and adaptively weight and fuse each multi-source credit supervision information according to the dynamic credibility to obtain the multi-source data fusion decision value; Implicit Verification Module: Used to construct enterprise credit association graphs, continuously acquire multi-source credit supervision information to spread implicit association risks, calculate the association consistency index of enterprises corresponding to each graph node, and perform implicit verification over time periods; The responsibility source tracing module is used to determine the time period of conflict between multi-source credit supervision information. Based on the first risk coefficient, the second risk coefficient, and the dynamic credibility, it calculates the conflict attention score of each multi-source credit supervision information, determines the source of responsibility within the conflict time period, and corrects the information.