Credit investigation rating system based on multi-source data fusion and enterprise knowledge graph
By integrating multi-source data and enterprise knowledge graph technology, a credit rating system is constructed, which solves the problems of data integration and knowledge graph fragmentation in existing technologies. It enables a comprehensive and dynamic assessment of enterprise credit, improves the accuracy and robustness of the rating, and outputs rating results in multiple forms, making it easier for financial institutions to manage credit risk.
Patent Information
- Application Number
- CN202511467170.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2026-01-09
AI Technical Summary
The existing credit reporting system is fragmented in terms of data integration and graph construction, failing to form a synergistic effect. It is difficult to comprehensively and dynamically assess corporate credit, and traditional rating methods are inadequate in the face of rapidly changing markets.
A credit rating system is constructed by employing multi-source data fusion and enterprise knowledge graph technology. The system includes modules for data collection, preprocessing, fusion, knowledge graph construction, and credit rating. A comprehensive credit score is generated through entity parsing, relation extraction, and graph analysis algorithms.
It enables a comprehensive and dynamic assessment of corporate credit, improves the accuracy and robustness of ratings, outputs rating results in multiple forms, facilitates financial institutions in managing credit risk, and promotes financial market stability.
Smart Images

Figure CN121304292A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of credit rating systems, in particular to a credit rating system based on multi-source data fusion and enterprise knowledge graph. BACKGROUND
[0002] With the vigorous development of the economy, the importance of enterprise credit rating in the field of financial risk control is increasingly prominent. Traditional credit systems mainly focus on the historical financial data and credit records of enterprises, and the data sources are relatively limited and have static characteristics, which makes it difficult to fully and truly reflect the credit of enterprises. Especially in the current wave of big data, enterprise data is showing a trend of multi-source convergence, heterogeneous interweaving and massive growth, while the existing credit system is not competent in integrating such data, and it lacks consideration of the complex relationship network between enterprises, such as supply chain association and investment link, which makes the accuracy of the rating results greatly discounted and makes it difficult to capture potential risks. In addition, traditional rating methods have shortcomings in dynamic updating and depth correlation analysis, which are difficult to adapt to the changing market rhythm. Although there are many attempts in the prior art to use multi-source data or knowledge graph to build a credit system, these attempts often separate data fusion and graph construction, and do not form a synergistic effect. Data fusion is more than surface aggregation, and the application of knowledge graph is only focused on relationship retrieval, and cannot be deeply integrated into the rating model. SUMMARY
[0003] Therefore, the present application provides a credit rating system based on multi-source data fusion and enterprise knowledge graph to solve the problem that data fusion and graph construction are separated in the prior art and do not form a synergistic effect.
[0004] In order to achieve the above purpose, the present application provides the following technical solutions:
[0005] The credit rating system based on multi-source data fusion and enterprise knowledge graph comprises the following modules: a data acquisition module, a data preprocessing module, a data fusion module, a knowledge graph construction module and a credit rating module.
[0006] The data acquisition module is configured to acquire enterprise-related data from a plurality of heterogeneous data sources, including but not limited to public financial statements, tax data, business registration information, judicial litigation records, news public opinion data, industry reports and supply chain data,
[0007] The data preprocessing module is connected to the data acquisition module and is used to perform data cleaning, data transformation and data enrichment on the acquired raw data. Data cleaning includes deduplication, error correction, outlier detection and processing, and uses rule-based and statistical methods to identify invalid data. Data transformation involves format standardization, unit unification and encoding conversion. Data enrichment extracts key information through natural language processing technology.
[0008] The data fusion module is connected to the data preprocessing module and adopts multi-source data fusion technology, including rule-based methods and machine learning algorithms, to perform entity parsing, record linking and conflict resolution. Entity parsing uses fuzzy matching and semantic analysis to identify different representations of the same enterprise. Record linking establishes a data association network through graph algorithms. Conflict resolution integrates contradictory data based on data source credibility weighting or voting mechanisms to output a consistent enterprise data view.
[0009] The knowledge graph construction module is connected to the data fusion module and is used to construct an enterprise knowledge graph. Enterprise entities serve as nodes, and node attributes include enterprise name, registered address, industry classification, financial indicators, and operating status. Edges represent relationships between enterprises, including equity relationships, transaction relationships, guarantee relationships, and competitive relationships. The construction process includes entity extraction, relationship extraction, and graph storage. Entity extraction uses named entity recognition technology to extract enterprise information from text data. Relationship extraction uses a relationship classification model to identify relationship types. Graph storage uses a graph database for storage.
[0010] The credit rating module is connected to the knowledge graph construction module and is used to conduct credit rating based on the enterprise knowledge graph. The rating method integrates graph analysis algorithms and traditional credit scoring models. The graph analysis algorithms include calculating node degree centrality, eigenvector centrality, community detection, and shortest path analysis to assess enterprise influence, risk contagion, and clustering effects. The traditional credit scoring model uses logistic regression, decision trees, or neural networks, combined with financial ratios and macroeconomic indicators, to generate a comprehensive credit score.
[0011] Preferably, the data acquisition module also includes a web crawler submodule and an API interface submodule. The web crawler submodule adopts a distributed crawler architecture and supports timed or real-time data crawling from public websites and databases. The API interface submodule connects with external data providers through RESTful API or GraphQL protocol to achieve automatic access to structured data and has a built-in data caching mechanism to improve acquisition efficiency.
[0012] The data acquisition module also includes a data source management submodule, which can dynamically adjust the acquisition strategy;
[0013] It also includes an output module, which is connected to the credit rating module and is used to output the rating results in various forms, including generating credit report documents, providing them to third-party systems through API interfaces, or displaying knowledge graphs and rating details in a visualization interface that supports interactive queries and dynamic filtering.
[0014] Preferably, the data acquisition module further includes a data source management submodule, used to configure and manage the connection parameters, acquisition frequency and authentication information of different data sources.
[0015] Preferably, the data source management submodule further includes a data quality monitoring component. The data quality monitoring component monitors the availability, data integrity, and consistency of the data source in real time. By setting thresholds and alarm rules, when data anomalies are detected, such as excessively high missing rates or format errors, a retry mechanism is automatically triggered or a backup data source is switched. The data quality monitoring component also integrates a log recording function to record the collection history and quality indicators.
[0016] Preferably, the data preprocessing module further includes a data cleaning submodule and a data transformation submodule. The data cleaning submodule uses a machine learning-based data anomaly detection algorithm to identify and process outliers, and the data transformation submodule uses a template engine to map heterogeneous data to a standard schema.
[0017] Preferably, the data cleaning submodule also integrates data verification rules, which include format checking, range verification, and logical consistency checking; the data conversion submodule supports custom mapping rules, allowing users to define conversion logic through a graphical interface.
[0018] Preferably, the data fusion module further includes an entity parsing submodule and a conflict resolution submodule. The entity parsing submodule uses graph embedding technology to enhance entity matching accuracy, while the conflict resolution submodule uses a context-based weighted average method to integrate conflict data.
[0019] Preferably, the entity parsing submodule is configured to use a deep learning model to calculate entity similarity and improve matching accuracy by training labeled data; the conflict resolution submodule introduces a credibility assessment mechanism, assigns credibility scores to each data source, dynamically adjusts the fusion weights based on the scores, and supports a manual intervention interface to allow experts to correct the fusion results; in addition, the data fusion module also includes a fusion effect evaluation component, which periodically tests the fusion performance using a benchmark dataset and optimizes the algorithm parameters based on feedback to ensure the long-term stability of the system.
[0020] Preferably, the credit rating module further includes a risk warning submodule, which is used to trigger risk alerts based on real-time changes in the knowledge graph, such as automatically notifying users when an anomaly in the guarantee chain is detected.
[0021] Compared with the prior art, this application has at least the following beneficial effects:
[0022] This invention achieves a comprehensive and dynamic assessment of corporate credit by integrating multi-source data fusion and enterprise knowledge graph technology. The system can collect data from multiple channels, and after preprocessing and fusion, construct a unified enterprise data view. The enterprise knowledge graph constructed on this basis not only includes the attributes of the enterprise itself, but also reveals various relationships between enterprises, so that credit rating can take into account associated risks, such as guarantee chain risk or supply chain risk.
[0023] By combining graph analysis algorithms and traditional scoring methods into a rating model, the accuracy and robustness of ratings are greatly improved. Moreover, the system achieves automated processing, reduces manual intervention, and improves efficiency. The output module provides rating results in multiple formats to facilitate use by different users. Overall, it provides a more scientific and reliable solution for enterprise credit investigation, helps financial institutions better manage credit risk, and promotes the stable development of the financial market. Attached Figure Description
[0024] To more intuitively illustrate the prior art and this application, exemplary drawings are provided below. It should be understood that the specific shapes and structures shown in the drawings should not generally be regarded as limiting conditions for implementing this application; for example, based on the technical concept disclosed in this application and the exemplary drawings, those skilled in the art are able to easily make conventional adjustments or further optimizations to the addition / reduction / classification, specific shapes, positional relationships, connection methods, size ratios, etc. of certain units (components).
[0025] Figure 1 This is a module diagram of the credit rating system based on multi-source data fusion and enterprise knowledge graph proposed in this application. Detailed Implementation
[0026] The present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0027] like Figure 1 As shown, this application discloses a credit rating system based on multi-source data fusion and enterprise knowledge graph, including the following modules: data acquisition module, data preprocessing module, data fusion module, knowledge graph construction module, credit rating module, and output module;
[0028] The data acquisition module is configured to collect enterprise-related data from multiple heterogeneous data sources, including but not limited to publicly available financial statements, tax data, business registration information, legal records, news and public opinion data, industry reports, and supply chain data.
[0029] The data acquisition module further includes a web crawler submodule and an API interface submodule. The web crawler submodule adopts a distributed crawler architecture and supports timed or real-time data crawling from public websites and databases. The API interface submodule connects with external data providers through RESTful API or GraphQL protocol to achieve automatic access to structured data and has a built-in data caching mechanism to improve acquisition efficiency.
[0030] The data acquisition module also includes a data source management submodule, which can dynamically adjust the acquisition strategy, such as optimizing the acquisition order based on data source priority and network conditions, to ensure the comprehensiveness and timeliness of the data.
[0031] The data preprocessing module is connected to the data acquisition module and is used to perform data cleaning, data transformation, and data enrichment on the acquired raw data. Data cleaning includes deduplication, error correction, outlier detection and processing, and uses rule-based and statistical methods to identify invalid data. Data transformation involves format standardization, unit unification, and encoding conversion, such as unifying different date formats to the ISO standard and converting currency units to base currencies. Data enrichment uses natural language processing technology to extract key information, such as extracting corporate events from news texts and adding tags to enhance data dimensions, generating high-quality standardized datasets.
[0032] The data fusion module is connected to the data preprocessing module and adopts multi-source data fusion technology, including rule-based methods and machine learning algorithms, to perform entity parsing, record linking and conflict resolution. Entity parsing uses fuzzy matching and semantic analysis to identify different representations of the same enterprise. Record linking establishes a data association network through graph algorithms. Conflict resolution integrates contradictory data based on data source credibility weighting or voting mechanisms to output a consistent enterprise data view.
[0033] The knowledge graph construction module is connected to the data fusion module and is used to construct an enterprise knowledge graph. Enterprise entities serve as nodes, and node attributes include enterprise name, registered address, industry classification, financial indicators, and operating status. Edges represent relationships between enterprises, including equity relationships, transaction relationships, guarantee relationships, and competitive relationships. The construction process includes entity extraction, relationship extraction, and graph storage. Entity extraction uses named entity recognition technology to extract enterprise information from text data. Relationship extraction uses a relationship classification model to identify relationship types. Graph storage uses a graph database such as Neo4j or JanusGraph, which supports SPARQL queries and real-time updates.
[0034] The credit rating module is connected to the knowledge graph construction module and is used to conduct credit rating based on the enterprise knowledge graph. The rating method integrates graph analysis algorithms and traditional credit scoring models. The graph analysis algorithms include calculating node degree centrality, eigenvector centrality, community detection and shortest path analysis to assess enterprise influence, risk contagion and cluster effect. The traditional credit scoring model uses logistic regression, decision tree or neural network, combined with financial ratios and macroeconomic indicators to generate a comprehensive credit score.
[0035] The output module is connected to the credit rating module and is used to output the rating results in various forms, including generating credit report documents, providing them to third-party systems through API interfaces, or displaying knowledge graphs and rating details in a visualization interface. The visualization interface supports interactive queries and dynamic filtering, which facilitates in-depth analysis by users.
[0036] In implementation, this invention involves a data acquisition module that automatically obtains raw enterprise data from diverse and heterogeneous data sources (such as financial, business registration, and public opinion data). Next, a data preprocessing module cleans, transforms, and enriches this data, eliminating noise and standardizing the format to create a high-quality dataset. Then, the core data fusion module integrates fragmented information from different channels pointing to the same enterprise using entity resolution and conflict resolution technologies, generating a unified and accurate enterprise data view. Based on this, a knowledge graph construction module uses information extraction technology to construct a visualized knowledge graph of enterprise entities and their complex relationships (such as equity, guarantees, and supply chains), storing it in a graph database. Subsequently, a credit rating module comprehensively utilizes traditional scoring models (such as logistic regression) and graph analysis algorithms (such as centrality calculation and community discovery) to not only assess the enterprise's own qualifications but also deeply explore potential risks within its interconnected networks, generating a more comprehensive and dynamic credit score. Finally, the output module delivers the rating results and analysis process to users in the form of reports, APIs, or visual interfaces to assist in decision-making. The entire process achieves a closed loop from automatic integration of multi-source data to intelligent risk insight.
[0037] For example, when evaluating Company A, the following situations may occur:
[0038] Step 1: Data Collection
[0039] System Operation: The data acquisition module begins automatic operation. It uses web crawlers to retrieve "Company A's" business registration information, annual financial report summaries, and related press releases (such as "Company A wins an innovation award") from public channels. Simultaneously, it obtains its tax payment records, legal proceedings information (discovering a sales contract dispute in which it is the defendant), and a list of its major suppliers and customers (e.g., its core supplier is "Company B") from third-party data service providers via API interfaces.
[0040] Step 2: Data Preprocessing
[0041] System Operation: The collected raw data was in a disorganized format. The data preprocessing module cleaned and standardized it. For example, it converted the date "October 1, 2023" in the financial report to "2023-10-01", and converted "Revenue: One Hundred Million Yuan" to the number "100,000,000". It also extracted key events (such as "awards") from news texts and labeled them with "positive public opinion".
[0042] Step 3: Data Fusion
[0043] System Operation: Currently, the system's data on "Company A" may come from the Administration for Industry and Commerce (registered capital of 100 million), the Tax Bureau (annual tax payment of 5 million), and financial reports (revenue of 100 million). The core task of the data fusion module is to confirm that all this data belongs to the same company (entity resolution) and resolve any potential conflicts. For example, if another industry report shows its revenue is only 80 million, the module will perform weighted fusion based on the credibility of the data source (e.g., official tax bureau data has higher weight), ultimately generating the most reliable and unified data profile for "Company A".
[0044] Step 4: Knowledge Graph Construction
[0045] System Operation: The knowledge graph construction module uses "Company A" as a node, filling in its attributes (registered capital, revenue, etc.). Then, it creates other related nodes, such as the supplier "Company B," the customer "Company C," and the opposing party in the lawsuit, "Company D." Next, it establishes relationship edges between these nodes:
[0046] Company A - [Supplier] -> Company B
[0047] Company A - [Client] - > Company C
[0048] "Company D" - [Defendant in the lawsuit] -> "Company A"
[0049] In this way, a local corporate relationship network is formed, which intuitively shows the position of "Company A" in the business ecosystem.
[0050] Step 5: Credit Rating
[0051] System Operation: The credit rating module has begun a comprehensive assessment.
[0052] Traditional model section: It analyzes Company A's financial ratios (such as debt-to-equity ratio and profit margin) and gives a base score.
[0053] The graph analysis section analyzes the knowledge graph. For example, by calculating node degree centrality, it discovers that "Company B" is a core supplier for many companies (i.e., high degree centrality). This means that if this supplier encounters problems, the risk will spread to "Company A" through the supply chain. Simultaneously, the litigation relationship will also be identified as a negative signal. The module quantifies these graph analysis results (association risk, public opinion risk) and weights them together with scores from traditional models to generate a comprehensive credit score that considers both the company's individual strength and its external associated risks. This score is more comprehensive and accurate than simply looking at financial statements.
[0054] Step 6: Output the results
[0055] System Operation: The output module generates a credit report from the final credit rating (e.g., "BBB+") and detailed analysis, and pushes it to the bank's credit approval system via API. Bank managers can also directly view the knowledge graph of "Company A" on the system's visual interface, clearly seeing its supply chain relationships and risk points, providing strong data support for loan decisions.
[0056] The data acquisition module further includes a data source management submodule, which is used to configure and manage the connection parameters, acquisition frequency and authentication information of different data sources. The data source management submodule supports multiple protocols, such as HTTP, FTP and JDBC, and has a load balancing function, which automatically allocates acquisition tasks according to the data source response time.
[0057] By adding a data source management submodule, connection parameters and collection frequencies for various data sources can be centrally configured and managed. This acts like a "central dispatch center," enabling the system to flexibly and uniformly handle interactions with different data providers, improving the manageability and efficiency of the data collection process. For example, this submodule can be configured to collect data from data source A hourly, data source B daily, and access keys can be managed centrally.
[0058] The data source management submodule also includes a data quality monitoring component. This component monitors the availability, integrity, and consistency of the data source in real time. By setting thresholds and alarm rules, it automatically triggers a retry mechanism or switches to a backup data source when data anomalies are detected, such as excessively high missing rates or format errors, ensuring the reliability of the data collection process. The data quality monitoring component also integrates a logging function to record collection history and quality indicators, facilitating auditing and optimization. In addition, this submodule supports dynamic data source registration, allowing users to add new data sources through the configuration interface without modifying the code, thus improving the system's scalability and flexibility.
[0059] The data quality monitoring component acts like a "quality inspector," monitoring the health and quality of data sources in real time. Upon detecting problems (such as data source downtime or incorrect data format), it automatically triggers countermeasures (such as switching to a backup source), ensuring the stability and reliability of data collection. The dynamic registration function allows users to easily add new data sources through the configuration interface without modifying program code, greatly enhancing the system's scalability.
[0060] The data preprocessing module further includes a data cleaning submodule and a data transformation submodule. The data cleaning submodule uses machine learning-based data anomaly detection algorithms, such as isolated forest or autoencoder, to identify and process outliers. The data transformation submodule uses a template engine to map heterogeneous data to a standard schema. The transformation submodule uses a template engine to efficiently map diverse raw data into a unified standardized format within the system, providing a foundation for subsequent fusion.
[0061] The data cleaning submodule also integrates data verification rules, including format checking, range verification, and logical consistency checking, to ensure data quality; the data transformation submodule supports custom mapping rules, allowing users to define transformation logic through a graphical interface, reducing coding workload.
[0062] Data validation rules (such as checking whether values are within a reasonable range) ensure the correctness of the data logic. Meanwhile, the data transformation submodule supports custom mapping rules through a graphical interface, lowering the technical barrier and allowing business experts to participate in the development of data standards. This reduces reliance on developers and improves the system's usability and adaptability.
[0063] The data fusion module further includes an entity parsing submodule and a conflict resolution submodule. The entity parsing submodule uses graph embedding technology to enhance entity matching accuracy, while the conflict resolution submodule uses a context-based weighted average method to integrate conflict data.
[0064] The entity parsing submodule transforms enterprise entities into mathematical vectors and calculates vector similarity to more accurately determine whether two records point to the same enterprise, which is especially useful when names are similar or there are spelling errors. The conflict resolution submodule uses a context-based weighted average method, which can more intelligently integrate contradictory information rather than simply discarding it.
[0065] The entity parsing submodule is configured to use deep learning models, such as Siamese networks or Transformers, to calculate entity similarity and improve matching accuracy by training on a large amount of labeled data. The conflict resolution submodule introduces a credibility assessment mechanism, assigns credibility scores to each data source, dynamically adjusts the fusion weights based on the scores, and supports a manual intervention interface, allowing experts to correct the fusion results. In addition, the data fusion module also includes a fusion effect evaluation component, which regularly tests the fusion performance using benchmark datasets and optimizes algorithm parameters based on feedback to ensure the long-term stability of the system.
[0066] The credit rating module further includes a risk warning submodule, which triggers risk alerts based on real-time changes in the knowledge graph. For example, it automatically notifies users when an anomaly in the guarantee chain is detected, thus enabling the system to not only perform static ratings but also dynamically monitor based on real-time changes in the knowledge graph. Once an abnormal pattern appears in the graph (such as a rapid expansion of the guarantee circle of a core enterprise), the risk warning submodule will immediately trigger an alarm to notify users of potential risks, achieving a leap from "post-event evaluation" to "pre-event warning," which is of great value for financial risk control.
[0067] The technical features of the above embodiments can be combined in any way (as long as there is no contradiction in the combination of these technical features). For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described; these embodiments not explicitly written should also be considered to be within the scope of this specification.
Claims
1. A credit rating system based on multi-source data fusion and enterprise knowledge graph, characterized in that, It includes the following modules: data acquisition module, data preprocessing module, data fusion module, knowledge graph construction module, and credit rating module; The data acquisition module is configured to collect enterprise-related data from multiple heterogeneous data sources, including but not limited to publicly available financial statements, tax data, business registration information, legal records, news and public opinion data, industry reports, and supply chain data. The data preprocessing module is connected to the data acquisition module and is used to perform data cleaning, data transformation and data enrichment on the acquired raw data. Data cleaning includes deduplication, error correction, outlier detection and processing, and uses rule-based and statistical methods to identify invalid data. Data transformation involves format standardization, unit unification and encoding conversion. Data enrichment extracts key information through natural language processing technology. The data fusion module is connected to the data preprocessing module and adopts multi-source data fusion technology, including rule-based methods and machine learning algorithms, to perform entity parsing, record linking and conflict resolution. Entity parsing uses fuzzy matching and semantic analysis to identify different representations of the same enterprise. Record linking establishes a data association network through graph algorithms. Conflict resolution integrates contradictory data based on data source credibility weighting or voting mechanisms to output a consistent enterprise data view. The knowledge graph construction module is connected to the data fusion module and is used to construct an enterprise knowledge graph. Enterprise entities serve as nodes, and node attributes include enterprise name, registered address, industry classification, financial indicators, and operating status. Edges represent relationships between enterprises, including equity relationships, transaction relationships, guarantee relationships, and competitive relationships. The construction process includes entity extraction, relationship extraction, and graph storage. Entity extraction uses named entity recognition technology to extract enterprise information from text data. Relationship extraction uses a relationship classification model to identify relationship types. Graph storage uses a graph database for storage. The credit rating module is connected to the knowledge graph construction module and is used to conduct credit rating based on the enterprise knowledge graph. The rating method integrates graph analysis algorithms and traditional credit scoring models. The graph analysis algorithms include calculating node degree centrality, eigenvector centrality, community detection, and shortest path analysis to assess enterprise influence, risk contagion, and clustering effects. The traditional credit scoring model uses logistic regression, decision trees, or neural networks, combined with financial ratios and macroeconomic indicators, to generate a comprehensive credit score.
2. The credit rating system based on multi-source data fusion and enterprise knowledge graph as described in claim 1, characterized in that, The data acquisition module also includes a web crawler submodule and an API interface submodule. The web crawler submodule adopts a distributed crawler architecture and supports timed or real-time data crawling from public websites and databases. The API interface submodule connects with external data providers through RESTful API or GraphQL protocol to achieve automatic access to structured data and has a built-in data caching mechanism to improve acquisition efficiency. The data acquisition module also includes a data source management submodule, which can dynamically adjust the acquisition strategy; It also includes an output module, which is connected to the credit rating module and is used to output the rating results in various forms, including generating credit report documents, providing them to third-party systems through API interfaces, or displaying knowledge graphs and rating details in a visualization interface that supports interactive queries and dynamic filtering.
3. The credit rating system based on multi-source data fusion and enterprise knowledge graph as described in claim 2, characterized in that, The data acquisition module also includes a data source management submodule, which is used to configure and manage the connection parameters, acquisition frequency and authentication information of different data sources.
4. The credit rating system based on multi-source data fusion and enterprise knowledge graph as described in claim 2, characterized in that, The data source management submodule also includes a data quality monitoring component. The data quality monitoring component monitors the availability, data integrity, and consistency of the data source in real time. By setting thresholds and alarm rules, when data anomalies are detected, such as excessively high missing rates or format errors, a retry mechanism is automatically triggered or a backup data source is switched. The data quality monitoring component also integrates a log recording function to record the collection history and quality indicators.
5. The credit rating system based on multi-source data fusion and enterprise knowledge graph as described in claim 1, characterized in that, The data preprocessing module also includes a data cleaning submodule and a data transformation submodule. The data cleaning submodule uses a machine learning-based data anomaly detection algorithm to identify and process outliers. The data transformation submodule uses a template engine to map heterogeneous data to a standard schema.
6. The credit rating system based on multi-source data fusion and enterprise knowledge graph as described in claim 5, characterized in that, The data cleaning submodule also integrates data verification rules, including format checking, range verification, and logical consistency checking; the data transformation submodule supports custom mapping rules, allowing users to define transformation logic through a graphical interface.
7. The credit rating system based on multi-source data fusion and enterprise knowledge graph as described in claim 1, characterized in that, The data fusion module also includes an entity parsing submodule and a conflict resolution submodule. The entity parsing submodule uses graph embedding technology to enhance entity matching accuracy, while the conflict resolution submodule uses a context-based weighted average method to integrate conflict data.
8. The credit rating system based on multi-source data fusion and enterprise knowledge graph as described in claim 7, characterized in that, The entity parsing submodule is configured to use a deep learning model to calculate entity similarity and improve matching accuracy by training labeled data. The conflict resolution submodule introduces a credibility assessment mechanism, assigns credibility scores to each data source, dynamically adjusts the fusion weights based on the scores, and supports a manual intervention interface, allowing experts to correct the fusion results. In addition, the data fusion module also includes a fusion effect evaluation component, which regularly tests the fusion performance using a benchmark dataset and optimizes the algorithm parameters based on feedback to ensure the long-term stability of the system.
9. The credit rating system based on multi-source data fusion and enterprise knowledge graph as described in claim 1, characterized in that, The credit rating module also includes a risk warning sub-module, which is used to trigger risk alerts based on real-time changes in the knowledge graph, such as automatically notifying users when an anomaly in the guarantee chain is detected.
Citation Information
Cited By
Asset risk analysis method and system for multi-source heterogeneous data fusion
CN122264944A