Multi-source data fusion credit rating evaluation method and system
By unifying the processing and quantitative evaluation of multi-source data, and combining real-time monitoring and dynamic weight adjustment, the problem of difficulty in quantifying credibility in multi-source data fusion credit rating has been solved, thus achieving the accuracy and real-time nature of credit rating.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- IND & COMMERCIAL BANK OF CHINA CO LTD KAIFENG BRANCH
- Filing Date
- 2025-12-02
- Publication Date
- 2026-05-12
AI Technical Summary
In existing credit rating assessments using multi-source data fusion by financial institutions, the credibility of data sources is difficult to quantify. In particular, the accuracy, completeness, timeliness, and stability of third-party data change over time and are difficult to monitor in real time, leading to a lag in weight adjustments and affecting the accuracy of the rating.
By unifying timestamps, field naming, and data types, missing value filling, outlier removal, and duplicate data deduplication are performed. The accuracy, completeness, timeliness, and stability indicators are quantified. The comprehensive credibility score is calculated by weighting and summing according to preset weights. The fusion weights are monitored in real time and adjusted according to score changes. Key risk signals of low credibility data are extracted, and the credit rating results are dynamically adjusted.
It achieves accurate measurement of the credibility of multi-source data, avoids noise interference from low-credibility data, retains important risk information, responds quickly to changes in the credibility of data sources, solves the problem of lagging dynamic adjustment of weights, and improves the accuracy and real-time performance of credit rating.
Smart Images

Figure CN122022979A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data management technology, specifically to a credit rating assessment method and system that integrates multi-source data. Background Technology
[0002] Credit rank refers to the use of symbols, based on rigorous analysis, to provide users with easily understandable information reflecting the creditworthiness of an entity being rated. Credit rank is a method of expressing and transmitting assessment information; if the symbols are complex and difficult to decipher, and the explanations are obscure, then such assessment information will be difficult for investors to understand and accept. In financial activities, the use of credit rank and the application of different credit rank ratings are widespread.
[0003] Patent document CN120182006A discloses a big data-based financial compliance risk assessment method, comprising: collecting and standardizing financial transaction data, customer information data, regulatory requirement data, and market data to obtain a structured dataset; constructing a multi-level financial compliance risk assessment indicator system based on the structured dataset, the indicator system being a three-level structure including risk domains, risk subcategories, and risk indicators; training a layered fusion network including decision trees, neural networks, and support vector machines using the structured dataset and the multi-level financial compliance risk assessment indicator system to obtain a financial compliance risk assessment model; identifying potential compliance risk points based on the compliance risk assessment report, and quantitatively assessing the compliance risk points by assigning weights to them to obtain quantitative risk point assessment results; setting graded early warning thresholds based on the quantitative risk point assessment results, and generating an early warning signal containing a risk point description, risk level, and handling suggestions when the risk quantitative indicator exceeds the preset threshold, and sending the early warning signal to the corresponding risk management department.
[0004] In existing credit rating assessment technologies for financial institutions, including the aforementioned technologies, the credibility of multi-source data varies significantly. For example, the accuracy of internal data from financial institutions is better than that of third-party data. When integrating data, the weights need to be dynamically adjusted according to credibility. However, quantifying credibility is a technical challenge: the credibility of data sources needs to consider "accuracy, completeness, timeliness, and stability," but these indicators are difficult to quantify. The credibility of data sources may change over time (e.g., a third-party credit reporting agency may experience a decline in accuracy due to data leakage). Real-time monitoring and weight adjustments are necessary. Some low-credibility data (e.g., social media sentiment) may contain key risk signals (e.g., negative news about companies). Directly discarding such data would result in information loss, but assigning it low weights could lead to it being overwhelmed by noise from high-credibility data.
[0005] Therefore, optimizing existing credit rating assessment methods is a problem worth studying. Summary of the Invention
[0006] To address the shortcomings of the existing technology, the present invention aims to provide a credit rating assessment method based on multi-source data fusion, and a credit rating assessment system based on multi-source data fusion, in order to solve the problems mentioned in the background art.
[0007] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0008] Firstly, the credit rating assessment method based on multi-source data fusion includes the following steps:
[0009] It integrates three types of data: internal data from financial institutions, third-party data, and low-reliability supplementary data. It unifies timestamps, field naming, and data types, and cleans the data by filling missing values, removing outliers, and deduplicating duplicate data, thus laying the foundation for credibility assessment.
[0010] The accuracy, completeness, timeliness, and stability of the four major indicators are converted into quantitative values by calculating the data matching rate, field completeness rate, data freshness, and data quality fluctuation coefficient.
[0011] The overall credibility score of each data source is obtained by weighted summation according to preset weights, a mapping library is established, and the real-time monitoring process is started.
[0012] The data is divided into three levels—high, medium, and low—based on its overall credibility score, and corresponding fusion weights are assigned accordingly. For low-credibility data, key risk signals are extracted through keyword matching and sentiment analysis.
[0013] By combining the data features and key risk signals from each layer, the raw score is calculated by weighting the data according to the weights, and the score is mapped to the corresponding credit rating level.
[0014] The system updates the data source credibility score in real time, adjusts the fusion weights based on score changes, and adjusts supplementary weights accordingly if high-priority risk signals are present, while simultaneously updating the credit rating results.
[0015] Furthermore, the access to three types of data—internal data from financial institutions, third-party data, and low-reliability supplementary data—is standardized in terms of timestamps, field naming, and data types. Data cleaning is achieved through missing value filling, outlier removal, and duplicate data deduplication, laying the foundation for credibility assessment. Specifically, rule-based mapping is performed on fields with the same or different names across data sources. Data cleaning is completed by specifically filling missing values, removing outliers, and removing duplicate data according to data source priority and timestamp rules. This ensures that the accessed data has a unified format and that invalid data is thoroughly removed, providing a high-quality data foundation for subsequent credibility assessment.
[0016] Furthermore, the four indicators of accuracy, completeness, timeliness, and stability are converted into quantitative values through data matching rate, field completeness rate, data freshness, and data quality fluctuation coefficient. Specifically, the following steps are taken: using core fields that have passed internal verification by financial institutions as a benchmark, the matching rate of third-party data is calculated by associating it with the benchmark data, and low-reliability data is verified by random sampling, thus quantifying accuracy; determining the list of core fields for each data source and quantifying completeness based on the proportion of non-missing core fields; quantifying timeliness based on the longest valid period of the data and the time difference between generation and storage; and quantifying stability by calculating the fluctuation coefficient based on the median reliability value within the statistical period, excluding data from abnormal periods, and adjusting for insufficient periods according to the corresponding coefficient.
[0017] Furthermore, the process of obtaining the comprehensive credibility score of each data source by weighted summation according to preset weights, establishing a mapping library, and starting a real-time monitoring process specifically involves: setting the weights of the four major credibility indicators for filing, which can be adjusted according to the process; verifying the validity of the indicators and calculating the comprehensive credibility score of each data source according to the preset formula; assigning a unique identifier to the data source; establishing a mapping library to store key information of the data source and adopting an adaptive storage method to ensure query and update efficiency; and starting a real-time monitoring process to periodically collect and verify data and store it in the warehouse.
[0018] Furthermore, the division into high, medium, and low layers based on the comprehensive credibility score and the allocation of corresponding fusion weights are as follows: the comprehensive credibility score is divided into high, medium, and low layers according to preset rules, and the weights are allocated according to the score gradient. Data from within financial institutions is assigned the highest weight by default in the high credibility layer, and the weights of other data sources are positively matched with the scores. Only the comprehensive credibility score is associated without additional adjustments based on type. The weight allocation rules are linked with the mapping library in real time, and the layering is automatically verified and the weights are matched as the credibility score is updated.
[0019] Furthermore, for low-reliability data, key risk signals are extracted through keyword matching and sentiment analysis. Specifically, taking financial credit risk scenarios as the core, a multi-dimensional risk keyword library containing synonym variant word mappings is constructed and updated regularly. A combination of exact matching and semantic similarity matching is used, along with a pre-trained model for sentiment analysis. Valid risk signals are filtered based on the relevance of the publishing entity, authority, validity of the publishing time, and deduplication conditions. The signals are then structured to generate feature sets, invalid information is removed, and suspected high-risk records awaiting review are marked and included in manual review. The review results are used to optimize the model and keyword library, and the extraction process is monitored in real time.
[0020] Furthermore, the process of splicing data features and key risk signals across different layers specifically involves: using the customer's unique identifier as the core association key, extracting effective structured features from the high, medium, and low confidence layers and eliminating redundant content; treating the key risk feature set of the low confidence data as an independent group and horizontally splicing it with the three layers of structured features to form a credit rating feature matrix; executing field deduplication rules; verifying the integrity of the matrix and standardizing the labeling of missing non-core fields; using a unified timestamp to record the splicing time; and linking it with the real-time monitoring process for scheduled execution, retaining the update trajectory to support traceability.
[0021] Furthermore, the step of calculating the original score by weight and mapping the score to the corresponding credit rating level specifically involves: standardizing the structured features at each layer, calculating the weighted sum of the structured features at each layer according to the corresponding fusion weight, multiplying the score of key risk features by supplementary weights after calculating the score according to the rules, and then summing the results to obtain the original score. The original score is then mapped to the credit rating level according to preset rules. The threshold value can be fine-tuned by combining risk signals and the proportion of high-confidence data. The calculation and mapping are performed synchronously with the feature matrix update every hour.
[0022] Furthermore, the real-time update of the data source credibility score and the adjustment of the fusion weights based on score changes are specifically as follows: the latest quality data of the data source is collected based on the real-time monitoring process, the credibility score is recalculated after validity verification, the weight adjustment is triggered according to preset conditions by comparing with the historical score, and the new weights are matched according to the hierarchical gradient rules.
[0023] Secondly, a credit rating and assessment system based on multi-source data fusion includes:
[0024] The data access module is used to access three types of data: internal data from financial institutions, third-party data, and low-reliability supplementary data.
[0025] The data standardization module is used to unify the timestamps, field names, and data types of various types of data.
[0026] The data cleaning module is used to clean data by filling in missing values, removing outliers, and deduplicating duplicate data, laying the foundation for credibility assessment.
[0027] The credibility index quantification module is used to convert four major indicators—accuracy, completeness, timeliness, and stability—into quantitative values through calculations of data matching rate, field completeness rate, data freshness, and data quality fluctuation coefficient.
[0028] The comprehensive credibility calculation module is used to calculate the comprehensive credibility score of each data source by weighted summation according to preset weights, and to establish a mapping library;
[0029] The real-time monitoring module is used to start the real-time monitoring process and update the data source credibility score in real time.
[0030] The credibility stratification module is used to divide the credibility into three layers: high, medium and low, based on the comprehensive credibility score, and to assign corresponding fusion weights to each layer.
[0031] The risk signal extraction module is used to extract key risk signals from low-reliability data through keyword matching and sentiment analysis.
[0032] The credit rating calculation module is used to combine data features and key risk signals from each layer, calculate the raw score by weighting, and map the score to the corresponding credit rating level.
[0033] The weighting and rating adjustment module is used to adjust the fusion weights based on changes in the credibility score. If there are high-priority risk signals, the supplementary weights will be adjusted accordingly, and the credit rating results will be updated synchronously.
[0034] Compared with existing technologies, this invention has the following advantages: By transforming abstract core credibility indicators into calculable quantitative values, it effectively solves the technical pain point of difficulty in accurately measuring the credibility of multi-source data; by using a model of allocating fusion weights in a hierarchical manner according to credibility, combined with the targeted extraction and independent weight configuration of key risk signals in low-credibility data, it avoids noise interference from low-credibility data while fully preserving important risk information, thus eliminating the problem of information loss; by monitoring changes in data source quality in real time, dynamically updating credibility scores and fusion weights, and synchronously adjusting credit rating results, it achieves rapid response to changes in data source credibility, solving the drawback of lagging dynamic weight adjustment. Attached Figure Description
[0035] Figure 1 This is a flowchart of the credit rating assessment method based on multi-source data fusion according to the present invention. Detailed Implementation
[0036] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments and accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0037] This application discloses a credit rating assessment method based on multi-source data fusion, such as... Figure 1 The steps include:
[0038] S1. Access three types of data: internal data from financial institutions, third-party data, and low-reliability supplementary data. Unify timestamps, field names, and data types. Clean the data by filling missing values, removing outliers, and deduplicating duplicate data to lay the foundation for credibility assessment.
[0039] S2. The four indicators of accuracy, completeness, timeliness, and stability are converted into quantitative values by calculating the data matching rate, field completeness rate, data freshness, and data quality fluctuation coefficient respectively.
[0040] S3. Calculate the overall credibility score of each data source by weighted summation according to preset weights, establish a mapping library, and start the real-time monitoring process;
[0041] S4. Divide the data into three layers—high, medium, and low—based on the overall credibility score and assign corresponding fusion weights; for low credibility data, extract key risk signals through keyword matching and sentiment analysis.
[0042] S5. Combine the data features and key risk signals from each layer, calculate the raw score by weighting, and map the score to the corresponding credit rating level;
[0043] S6. Update the data source credibility score in real time, adjust the fusion weight according to the score changes, and adjust the supplementary weight accordingly if there are high-priority risk signals, and update the credit rating results synchronously.
[0044] S1 integrates three types of data: internal financial institution data, third-party data, and low-reliability supplementary data, unifying timestamps, field naming, and data types. Data cleaning is achieved through missing value imputation, outlier removal, and duplicate data deduplication, laying the foundation for reliability assessment. Specifically, this involves integrating structured data from within financial institutions, such as transaction logs, credit records, and basic customer information; compliant data sources from third-party credit reports, business registration information, and tax declarations; and low-reliability supplementary data, such as social media sentiment, industry forum comments, and publicly available online information. Internal financial institution data is synchronized in real-time via core business system interfaces, while third-party data is synchronized via API interfaces on a contractual basis. Data is retrieved at a fixed frequency, while low-reliability data is collected through web crawlers combined with authorized data sources; the unified timestamp format is YYYY-MM-DDHH:MM:SS (accurate to the second); field names use a lowercase letter standard separated by underscores (e.g., overdue_times, debt_amount); data types are uniformly mapped to numeric (including integer and floating-point), character (UTF-8 encoding), and boolean (0 for no and 1 for yes); and fields with the same name but different meanings across data sources are mapped in a regularized manner (e.g., "number of overdue days" and "number of overdue records" are uniformly standardized to "overdue_times").
[0045] When performing data cleaning, missing numerical data is filled with the mean of the same customer dimension or the same business scenario, missing character data is uniformly marked as "none", and missing date data is marked as "no valid date". The mean and standard deviation of each numerical field are calculated based on the 3σ principle, and abnormal records that deviate from the mean by more than 3 times the standard deviation are removed. Extreme data (such as transaction amounts that far exceed the industry norm) are additionally marked for manual review.
[0046] Data deduplication follows the rule of "internal data of financial institutions has the highest priority, followed by third-party data, and supplementary data with low credibility has the lowest priority." For data sources with the same priority, the record with the latest timestamp is retained. Before deduplication, the consistency of core key fields (such as customer unique identifier and business number) is verified to avoid the accidental deletion of valid data due to field differences. Through the above standardization and cleaning operations, it is ensured that the access data format is uniform and invalid data is thoroughly removed, providing a high-quality data foundation for subsequent credibility assessment.
[0047] In S2, the four indicators of accuracy, completeness, timeliness, and stability are converted into quantitative values through data matching rate, field completeness rate, data freshness, and data quality fluctuation coefficient. In practice, the accuracy quantification uses the core fields (including customer unique identifier, transaction amount, overdue status, and debt balance) in the internal data of financial institutions that have been verified by the business system as the benchmark dataset, and the benchmark accuracy rate is fixed at 100%. When calculating the accuracy of third-party data, the benchmark data of financial institutions is first associated with the customer unique identifier, and the number of records with completely consistent core fields between the two is counted. Then, the number of records is divided by the total number of records in the batch of third-party data, and the result is multiplied by 100% to obtain the data matching rate.
[0048] The accuracy of supplementary data with low credibility is verified by daily random sampling. 5%-10% of the total number of records of this type are sampled each day and their authenticity is verified by personnel with credit rating business qualifications. The number of records that pass the verification is divided by the total number of sampled records, and the result is multiplied by 100% to obtain the data matching rate. The sample must cover different source platforms and different release time periods to ensure the representativeness of the verification.
[0049] Before quantifying completeness, the credit rating business committee of financial institutions determines the list of core fields for each data source based on regulatory requirements and rating model needs. The core fields for third-party credit reporting data include 12 items such as number of overdue payments, debt amount, and repayment records. The core fields for low-reliability supplementary data include 5 items such as the publishing entity, content keywords, and publication time. The list of core fields is reviewed and updated quarterly. When calculating the field completeness rate, the number of non-missing core fields in the data source is counted, divided by the total number of core fields in the data source, and the result is multiplied by 100%. If a core field is optional for business purposes and has no actual business significance, it is not included in the missing field statistics.
[0050] In the timeliness quantification, the longest validity period for each data source is set according to regulatory requirements, data update frequency and business scenarios. The longest validity period for internal transaction data of financial institutions is 24 hours, for business registration information it is 30 days, for tax declaration data it is 90 days, and for social media public opinion it is 7 days. The longest validity period can be adjusted by the business administrator according to actual business changes.
[0051] When calculating data freshness, the time difference (accurate to the second) is calculated between the data generation timestamp and the system timestamp when the data is fully entered into the database and passes the format verification. This time difference is then divided by the longest valid period of the corresponding data source. The result is obtained by subtracting this ratio from 1 and multiplying it by 100%. If the data generation timestamp is missing, the average update period of the data source is calculated backward from the access system timestamp as a reference for the generation time. In this case, the timeliness score needs to be reduced by 30%.
[0052] Stability quantification uses a 24-hour statistical period, starting from the moment the data source is first connected to the system and the credibility calculation is completed. The median credibility value of the last three consecutive periods is calculated (the arithmetic mean of the three indicators of accuracy, completeness, and timeliness in each period). When calculating the data quality fluctuation coefficient, the standard deviation and mean of these three median values are first calculated. The fluctuation ratio is obtained by dividing the standard deviation by the mean. Then, the fluctuation ratio is subtracted from 1 and multiplied by 100% to obtain the stability score.
[0053] If the indicator score drops by more than 50% in a certain period due to objective factors such as data source system failure or network interruption, the data for that period will not be included in the statistics and will be supplemented in the next period. If the data source access time is less than 3 statistical periods, it will be calculated based on the actual number of periods. For 1 period, the stability score will be multiplied by a period coefficient of 0.6; for 2 periods, it will be multiplied by a period coefficient of 0.8; and for 3 or more periods, the coefficient will be 1.0.
[0054] In step S3, the comprehensive credibility scores of each data source are obtained by weighted summation according to preset weights, a mapping library is established, and the real-time monitoring process is started. Specifically, the weights of the four credibility indicators are preset and confirmed by the credit rating business committee of the financial institution in combination with regulatory requirements and business practices, and are fixed as accuracy 0.4, completeness 0.2, timeliness 0.2, and stability 0.2. The weight values are stored in the system configuration center and filed. If adjustments are needed, they must be reviewed and approved by the committee and updated synchronously.
[0055] Furthermore, the current method uses fixed weights (accuracy 0.4, completeness 0.2, timeliness 0.2, stability 0.2) to calculate the overall credibility score. This fails to consider the differentiated requirements of data source type (e.g., internal data from financial institutions vs. supplementary data with low credibility) and business scenario (personal credit vs. corporate credit) regarding the importance of indicators. This leads to a disconnect between weight allocation and actual risk assessment needs, affecting the accuracy of the credibility score. Therefore, this application proposes using a "multi-dimensional constraint dynamic weight allocation algorithm" to calculate the weights of the four credibility indicators (accuracy A, completeness C, timeliness T, stability S), instead of fixed weights, as detailed below:
[0056] Input data includes basic input and the quantified values of the four confidence metrics (A∈[0,100], C ...
[0057] [0,100], T∈[0,100], S∈[0,100]); Constraint input, data source type coefficient λ (λ∈
[0058] {0.8, 0.6, 0.4}): Internal data from financial institutions λ = 0.8, third-party data λ = 0.6, and supplementary data with low reliability λ = 0.4, reflecting the impact of the inherent reliability of the data source on the indicator weights; Business relevance coefficient μ (μ∈[0,1]), set by the credit rating business committee of the financial institution based on the degree of relevance between the indicator and the rating objective (such as the "accuracy" relevance μ in personal credit scenarios). A =0.9, in corporate lending scenarios
[0059] "Stability" correlation μ S =0.9); quality fluctuation sensitivity coefficient σ (σ∈[1,2]), default σ
[0060] =1.5, when the data source quality fluctuation coefficient in the last 3 periods is >0.3, σ = 1.8, which enhances the sensitivity of the weights to changes in data quality.
[0061] Then, the objective weight ω of the index is calculated based on the entropy weight method. obj (Reflecting the inherent fluctuation characteristics of the data), the calculation formula includes:
[0062]
[0063] Information entropy H i This is used to measure the degree of dispersion of the values of the i-th indicator. The higher the degree of dispersion, the greater the weight.
[0064] In the formula, i∈{1,2,3,4} corresponds to the four major indicators: accuracy A, completeness C, timeliness T, and stability S.
[0065] k∈[1,n], where n is the quantified value of the indicator for the data source over the last n statistical periods (n≥3, calculated based on the actual number of periods if less than 3 periods); x ik p is the quantized value of the i-th indicator in the k-th period; ik H is the normalized probability of the value of the i-th indicator in the k-th period; i The information entropy (H) of the i-th index i ∈[0,1]), H i The smaller the value, the greater the difference in the indicator values, and the higher its contribution to the credibility assessment; ω i obj The objective weight of the i-th indicator (∑ω) i obj =1).
[0066] Then, the subjective weights ω of the indicators are calculated based on the analytic hierarchy process. subj (Reflecting business needs) Construct a business need judgment matrix M, and calculate the weights after passing a consistency check. The calculation formula includes:
[0067]
[0068] In the formula, M ij The importance of the i-th indicator relative to the j-th indicator (set based on the business relevance coefficient μ, such as μ) A >μ C At that time, a 12 =2, indicating that accuracy is twice as important as completeness; λ max The largest eigenvalue of the judgment matrix M is determined; CI is the consistency index, and RI is the random consistency index (RI = 0.90 for a 4th-order matrix); CR < 0.1 is required, otherwise the judgment matrix M is adjusted until the consistency requirement is met; ω i subj The subjective weight of the i-th indicator (∑ω) i subj =1).
[0069] Then, the subjective and objective weights are dynamically coupled to obtain the final weight ω. i Introducing the data source type coefficient λ and the quality fluctuation sensitivity coefficient σ, a coupling function is constructed:
[0070]
[0071] The larger λ is (the higher the reliability of the data source), the greater the objective weight ω. i obj The higher the weight ratio, the larger the σ (the more sensitive the data quality fluctuations), the stronger the adjustment effect of the objective weight, ensuring that the weight not only matches the actual characteristics of the data, but also does not deviate from business needs.
[0072] Then, calculate the overall credibility score:
[0073] Score = ∑ i=1 4 ω i ·x i ;
[0074] In the formula, x i Let be the current quantified value (A / C / T / S) of the i-th indicator, and Score be the comprehensive credibility score (rounded to two decimal places, Score∈[0,100]).
[0075] Finally, the output includes: the dynamic weights (ω) of the four credibility indicators. A ω C ω T ω S ),
[0076] Satisfying ∑ω i =1;
[0077] The system generates a comprehensive credibility score for each data source. It establishes a mapping library of "unique data source identifier - dynamic weight - comprehensive credibility score - λ / μ / σ", uses a distributed database for storage, and initiates a real-time monitoring process. The system collects the latest indicator data every hour, repeats the above steps to update the dynamic weight and comprehensive credibility score, and synchronously records the weight adjustment trajectory (including weights before and after adjustment, changes in parameters triggering the adjustment, and calculation results).
[0078] By using three constraint parameters, λ, μ, and σ, the weights are dynamically adapted to data source types, business scenarios, and data quality fluctuations, thus solving the problem of poor adaptability of fixed weights. The constraint parameters can be flexibly adjusted according to changes in regulatory policies and business scenarios (e.g., when adding a "green credit" scenario, the μ value can be adjusted to strengthen the weight of environmental compliance-related indicators), thereby expanding the applicability of the technical solution.
[0079] When calculating the overall credibility score, the validity of the four quantitative indicators—accuracy, completeness, timeliness, and stability—is first verified to ensure that the values of each indicator are within the range of 0-100. If an indicator exceeds the range due to data anomalies (such as collection failure or invalid verification), the average of the three most recent valid scores for that indicator is used as a substitute. Then, the overall credibility score is calculated according to the formula: "Overall Credibility Score = Accuracy × 0.4 + Completeness × 0.2 + Timeliness × 0.2 + Stability × 0.2". The result is rounded to two decimal places. Each data source is assigned a unique identifier, with the identifier format being "data source type_organization code_data category" (e.g., "internal_bank001_transaction", "thirdparty_credit002_report", "supplement_social003_opinion"). In addition to the unique identifier and comprehensive credibility score associated with the data source, the mapping database also synchronously stores key information such as data source type, access method (interface synchronization / API pull / web crawling), core field list, longest validity period, and sampling verification ratio. The mapping database uses distributed database storage, partitioned by data source type and timestamp-based, supporting millisecond-level queries and updates.
[0080] After establishing the mapping library, an independent real-time monitoring process is started. This process is associated with the monitoring frequency configured in the system (once every hour) and periodically collects raw quality data from each data source (including the number of matching records with the benchmark data of financial institutions, missing core fields, data generation and storage timestamps, and the median reliability value of the last 3 periods, etc.). After the collected data is format-validated, it is stored in the monitoring data warehouse. At the same time, the monitoring process has built-in anomaly detection rules. If data collection fails, a certain indicator value drops by more than 50%, or no valid data is obtained for two consecutive periods, a tiered alarm is immediately triggered (SMS + system pop-up notification to maintenance personnel and business managers), and an alarm log is recorded (including the time of the anomaly, data source identifier, anomaly type, and processing status).
[0081] The mapping library is updated synchronously every hour. Each update retains the historical credibility score and calculation basis, forming a data source credibility change trajectory archive. It supports traceability and query by time range and data source identifier, ensuring that the score is traceable and verifiable.
[0082] In S4, the system is divided into three layers: high, medium, and low, based on the overall credibility score, and corresponding fusion weights are assigned. In practice, the overall credibility score stratification standard strictly follows the preset rules, namely, the high credibility layer is defined as an overall score ≥ 80 points, the medium credibility layer is defined as an overall score ≤ 60 points and < 80 points, and the low credibility layer is defined as an overall score < 60 points. The stratification thresholds are stored in the system configuration center and are linked to the credibility score calculation module.
[0083] The weight allocation adopts a gradient rule. In the high credibility layer, the internal data of financial institutions are assigned the highest fusion weight of 0.8 by default, regardless of whether the score is in the range of 80-100. Only when the comprehensive score of the internal data of financial institutions drops below 80 will the weight be reallocated according to the corresponding layer.
[0084] Third-party data is weighted according to a score gradient: 0.6 for 80-84 points, 0.7 for 85-89 points, and 0.8 for 90-100 points, ensuring a positive match between scores and weights. The medium-confidence layer uses a more refined weighting gradient: 0.3 for 60-64 points, 0.4 for 65-69 points, and 0.5 for 70-79 points, achieving precise weight adjustment based on confidence scores. The low-confidence layer also uses a gradient allocation: 0.1 for 1-39 points and 0.2 for 40-59 points. These base weights are only linked to the overall confidence score and are not adjusted based on the data source type.
[0085] All weight allocation rules are linked in real time with the "data source-credibility score" mapping library. After each credibility score update, the system automatically verifies the stratification and matches the corresponding weight. The weight takes effect at the same time as the credibility score update (synchronized every hour).
[0086] After weight allocation, the system automatically records adjustment logs, including the unique identifier of the data source, the weight before adjustment, the weight after adjustment, the credibility score that triggered the adjustment, and the adjustment time. The logs are stored in association with the data source credibility change trajectory archive, with a retention period of no less than 1 year, and support traceability and verification according to business needs.
[0087] If the data source undergoes a cross-layer weight adjustment (such as being downgraded from a high-confidence layer to a medium-confidence layer), the system will trigger an additional lightweight alarm and simultaneously notify the credit rating business operations and maintenance personnel to ensure that the weight adjustment process is monitorable and traceable, fully meeting the core requirements of multi-source data layered fusion.
[0088] In S4, key risk signals are extracted from low-reliability data through keyword matching and sentiment analysis. Specifically, the extraction of key risk signals from low-reliability data focuses on financial credit risk scenarios. First, a multi-dimensional risk keyword library is constructed. Keywords are divided into four categories according to risk type: debt risk (such as "debt default", "inability to repay", "overdue payment"), compliance risk (such as "tax evasion", "illegal operation", "regulatory penalty"), operational risk (such as "broken capital chain", "shutdown", "negative exposure"), and related risk (such as "related enterprise default", "guarantee compensation"). The keyword library is reviewed by the credit rating business committee of financial institutions in combination with regulatory policies, historical default cases and industry risk characteristics, and is updated monthly. New risk event-related terms are added simultaneously. At the same time, a synonym and variant word mapping table is established (such as "tax evasion" and "tax fraud" are uniformly mapped to the keyword "tax violation") to ensure comprehensive matching.
[0089] Keyword matching adopts a combination of "exact matching + semantic similarity matching". First, the records that hit the core keywords are screened through exact matching. Then, semantic analysis algorithms in the financial field are used to identify semantically similar expressions to avoid risk omissions due to differences in expression. During the matching process, the completeness of the core fields of the text is checked at the same time. Only records with complete core fields of the publishing subject, content text and publishing time are analyzed later.
[0090] Sentiment analysis employs a pre-trained model based on the financial sector. The model is optimized and trained using over 100,000 pieces of financial sentiment data internally labeled by financial institutions, adapting to the risk assessment needs of credit scenarios. During analysis, the text is first segmented and stop words are removed. Then, the sentiment intensity calculation results are adjusted by combining the weights of risk keywords to ensure the accuracy of sentiment judgment in the financial sector. A negative sentiment intensity threshold of ≥0.8 is set. When the text sentiment intensity reaches this threshold and hits risk keywords, it is initially determined to be a valid risk signal.
[0091] Further refine the signal screening criteria: The publishing entity must be directly related to the target assessment object (such as the target company itself, its holding subsidiaries, or core related companies) or be an authoritative source of information (such as authoritative financial media, official accounts of regulatory agencies, or industry association publishing platforms). Anonymous content published by unrelated entities or without substantial information support must have an emotional intensity threshold of ≥0.9 to be included.
[0092] The publication time must be within the longest valid period (7 days) of this type of data. Expired data will only retain major risk signals with an emotional intensity ≥ 0.95. For content that is repeatedly published on multiple platforms for the same risk event, duplicates will be removed by text similarity algorithm (similarity ≥ 90% is judged as duplicates), and only the record with the highest emotional intensity and the highest source authority will be retained.
[0093] The selected valid risk signals are structured to generate a set of key risk features, including risk category, core hit keywords, sentiment intensity value (retaining two decimal places), name and type of the publishing entity, publishing time (accurate to the second, format consistent with standardized data), name and authority score of the source platform (set based on the platform's historical information authenticity verification results, ranging from 0 to 1.0), and risk description summary (extracting core information from the text, controlled within 50 words). At the same time, records with vague descriptions or incomplete information but suspected high risk are marked with "pending review" and included in the manual review queue (review cycle not exceeding 24 hours). The review results are simultaneously used to optimize the keyword library and sentiment analysis model parameters.
[0094] For information that is maliciously spammed, lacks substantial content, has a text length of less than 10 characters, or contains invalid characters, invalid information is removed through text length filtering and content relevance verification (based on semantic similarity <30%) to avoid interfering with credit rating results.
[0095] The extraction process is linked with the real-time monitoring process. The latest low-confidence data is collected and the extraction process is executed every hour. New risk signals are added to the feature matrix in real time, and signals that have expired (exceeded the validity period or have been verified as false information) are automatically removed from the feature set to ensure the timeliness and accuracy of key risk signals.
[0096] In S5, the data features and key risk signals of each layer are concatenated. Specifically, the concatenation operation uses the customer's unique identifier as the core association key to ensure that the data features of each layer correspond accurately to the key risk signals. First, the structured features of high, medium, and low credibility layers are extracted: the structured features of the high credibility layer include standardized numerical and character fields such as the total amount of transaction flow, the number of times of credit delinquency, and the debt balance (all following unified field naming and data type specifications); the structured features of the medium credibility layer cover core fields such as repayment records, years of business registration, and tax declaration amounts from third-party credit reports; the structured features of the low credibility layer only retain the cleaned and effective fields such as the type of publishing entity and the time period of information release, excluding redundant content without substantial credit assessment value.
[0097] Subsequently, the key risk feature set extracted from the low-credibility data (including risk category, core hit keywords, sentiment intensity value, name and type of the publishing entity, publishing time, authority score of the source platform, and risk description summary) is used as an independent feature group and horizontally spliced with the above three-layer structured features to form a complete credit rating feature matrix.
[0098] During the concatenation process, the following field deduplication rules are applied: If fields with the same name or meaning exist at different levels (e.g., "overdue_times" in the high-confidence layer and "overdue times" in the third-party data have been standardized to the same field), only the field value corresponding to the high-confidence layer is retained; if fields with the same name but different meanings exist (verified by the field mapping library), a data source identifier suffix (e.g., "profit_internal" or "profit_thirdparty") is added to the field name to distinguish them.
[0099] At the same time, the integrity of the feature matrix is verified to ensure that there are no missing customer unique identifiers and core rating fields (such as number of overdue payments, amount of debt, intensity of key risk sentiment, etc.). Missing non-core fields are marked with "none" or "0" (adapted according to data type). The spliced feature matrix adopts the same timestamp format as the "data source-credibility score" mapping library to record the splicing completion time and ensure the consistency with the subsequent weighted calculation and rating update process.
[0100] The splicing operation is linked with the real-time monitoring process. It is executed synchronously every hour after the data source credibility score is updated and key risk signals are added or removed. The newly generated feature matrix covers historical versions and retains the update trajectory. It supports tracing the feature composition of different stages by timestamp, providing data support for the traceability of credit rating results.
[0101] In step S5, the original score is calculated by weighting and mapped to the corresponding credit rating level. Specifically, before calculating the original score, the structured features of each layer are standardized. The min-max standardization method is used to map all numerical features to the 0-10 score range (formula: standardized score = (original feature value - minimum value of the feature) / (maximum value of the feature - minimum value of the feature) × 10. If the maximum value and minimum value of the feature are equal, they are uniformly assigned a value of 5 points). Character features are converted into quantitative scores according to preset rules (e.g., in "repayment status", "normal" is assigned 10 points, "overdue for 1-30 days" is assigned 5 points, and "overdue for more than 30 days" is assigned 0 points), ensuring that features of different types and scales are comparable. The weighted sum of structured features in the high-confidence layer is the sum of all valid standardized feature scores in that layer multiplied by the corresponding fusion weight (e.g., internal data from financial institutions is calculated with a weight of 0.8). The weighted sums of structured features in the medium-confidence and low-confidence layers are calculated according to their respective assigned fusion weights. The low-confidence layer only includes cleaned valid structured features; redundant fields are not included in the calculation. When calculating the score of key risk features, first, a score for each valid risk signal is calculated based on "sentiment intensity × source authority" (e.g., a signal with sentiment intensity of 0.85 and authority of 0.9 scores 0.85 × 0.9 = 0.765). Then, the scores of all valid signals are summed, and finally, the sum is multiplied by a supplementary weight (default 0.1, 0.2 for high-priority risk signals). If there is no valid risk signal, the key risk feature score is 0.
[0102] The original credit rating score is calculated using the formula "Original Score = Weighted Sum of High Credibility Layer + Weighted Sum of Medium Credibility Layer + Weighted Sum of Structured Features of Low Credibility Layer + Key Risk Feature Score". The result is rounded to two decimal places. If the total score exceeds the range of 0-100, it is truncated to the threshold (scores below 0 are counted as 0, and scores above 100 are counted as 100).
[0103] The rating mapping strictly follows preset rules: 90-100 points correspond to AAA, 80-89 points to AA, 70-79 points to A, 60-69 points to BBB, 50-59 points to BB, and below 50 points to B and below. These mapping rules are stored in the system configuration center and filed after review and approval by the financial institution's credit rating business committee. Adjustments require an approval process. During the mapping process, if the original score is at a critical level (e.g., 79.5, 69.8, etc.) and there is a significant risk signal with a sentiment intensity ≥0.95, the rating can be adjusted down one level. If the high-confidence layer data accounts for ≥80% and there are no risk signals, the critical score can be slightly adjusted up one level. The adjustment rules must be clearly indicated in the rating results.
[0104] The calculation of the original score and the rating mapping are performed synchronously with the feature matrix update every hour. During the calculation process, information such as the weight values of each layer, the details of the standardized feature scores, and the composition of the key risk signal scores are recorded and stored in the rating calculation log. This log is associated with the data source credibility change trajectory and the feature matrix update archive and is stored for a period of no less than one year. It supports tracing the basis of the original score calculation and the rating mapping process by customer unique identifier and time range to ensure that the rating results are verifiable and traceable.
[0105] In S6, the data source credibility score is updated in real time, and the fusion weight is adjusted according to the score changes. Specifically, the real-time update of the data source credibility score and the weight adjustment are carried out in conjunction with the real-time monitoring process every hour. The monitoring process collects the latest quality data of each data source at a preset frequency (including matching records with the benchmark data of financial institutions, missing core fields, data generation and entry timestamps, and the median credibility value of the last three statistical periods). After collection, the data validity is first checked to remove abnormal data caused by system failure or network interruption. If the data of a certain indicator is missing or exceeds the range of 0-100, the average of the last three valid scores of the indicator is used as a substitute. Then, the comprehensive credibility score is recalculated according to the formula "Comprehensive Credibility Score = Accuracy × 0.4 + Completeness × 0.2 + Timeliness × 0.2 + Stability × 0.2". The result is rounded to two decimal places.
[0106] After the new score is calculated, it is compared with the previous valid score of the data source in the mapping library. If the score change (absolute value) is ≥5 points, or if there is a score fluctuation across the layer threshold (such as from 82 points to 78 points, or from 61 points to 59 points), the fusion weight adjustment process is automatically triggered; if the score change is <5 points and does not cross the layer, the current weight remains unchanged.
[0107] When adjusting weights, the system matches the corresponding tiered weight rules based on the updated comprehensive credibility score (high credibility tier: 80-84 points → 0.6, 85-89 points → 0.7, 90-100 points → 0.8; medium credibility tier: 60-64 points → 0.3, 65-69 points → 0.4, 70-79 points → 0.5; low credibility tier: 1-39 points → 0.1, 40-59 points → 0.2; internal data of financial institutions remains fixed at 0.8 in the high credibility tier). The system generates adjusted weight values. If the data source stability index drops below 60 points after the update, the weight is further reduced by 0.1 on top of the matched tiered weights, and a manual review process is automatically triggered. The weights officially take effect after the review is passed; otherwise, they are corrected according to the review comments.
[0108] After the weight adjustment is completed, the system updates the weight field in the "Data Source-Confidence Score" mapping library in real time and records the adjustment log synchronously. The log includes information such as the unique identifier of the data source, the previous score and weight, the current score and weight, the score change range, the adjustment trigger conditions (range meets the standard / cross-layer / stability does not meet the standard), the adjustment time, and the review status (if required). It is stored in association with the data source credibility change trajectory archive and the retention period is no less than 1 year.
[0109] Meanwhile, the weight update signal is synchronized to the feature fusion module and credit rating calculation module in real time to ensure that the latest weights are used in subsequent fusion calculations and rating result updates. The entire update and adjustment process is completed within 3 minutes after the monitoring data is collected, ensuring the real-time nature of weight adjustment and the timeliness of rating results.
[0110] If a high-priority risk signal exists in S6, the weights will be adjusted accordingly, and the credit rating results will be updated synchronously. Specifically, the determination of high-priority risk signals strictly relies on the key risk feature set and must meet two core conditions simultaneously: first, the emotional intensity of the risk signal is ≥0.9; second, the authority score of the source platform is ≥0.8, and the risk category belongs to a major risk type in debt risk (such as "debt default" or "inability to repay") or compliance risk (such as "regulatory penalties" or "tax evasion"). The determination criteria are consistent with the key risk signal extraction rules and are automatically verified by the system without manual intervention.
[0111] When the system detects a high-priority risk signal that meets the criteria, it immediately triggers a supplementary weight adjustment process, adjusting the original default supplementary weight of 0.1 for key risk features to 0.2. The adjustment instruction is synchronized in real time to the feature fusion module and the credit rating calculation module to ensure that subsequent calculations use the updated weight. The supplementary weight adjustment is effective for 24 hours. If a new high-priority risk signal of the same type is added during this period, the duration is recalculated from the trigger time of the new signal for 24 hours. If multiple high-priority risk signals of different types are superimposed on the same assessment object (such as simultaneously hitting debt default and regulatory penalty signals), the supplementary weight can be adjusted to a maximum of 0.3. When superimposed, the maximum adjustment range is determined by the sum of the products of "emotional intensity × source authority" of each signal, and the maximum cannot exceed 0.3.
[0112] After the supplementary weight adjustment is completed, the system will initiate the credit rating result synchronization update process within 3 minutes, re-execute feature standardization, weighted summation and rating mapping operations, and prioritize ensuring that the scores of high-priority risk signals are fully included during the recalculation. If the original score crosses the rating level threshold due to the supplementary weight adjustment (e.g., from 79.8 points to 80.2 points, or from 60.3 points to 59.7 points), it will be directly mapped to the corresponding level according to the new score, without the need to execute additional threshold fine-tuning rules. If the high-priority risk signal is subsequently confirmed by manual review as false information (e.g., malicious rumors, misreporting), the current supplementary weight adjustment will be immediately terminated, the default supplementary weight of 0.1 will be restored, and the credit rating result will be recalculated retrospectively. The corrected result will overwrite the original update record, and the reason for the retrospective and the basis for the correction will be noted in the log.
[0113] The updated credit rating results must be pushed to the financial institution's credit rating system in real time to ensure that the business side receives the latest rating information. At the same time, the system triggers a lightweight alarm, notifying the credit rating business manager and the corresponding account manager through system pop-ups and SMS. The alarm content includes the unique identifier of the assessed object, the core information of the high-priority risk signal (risk category, sentiment intensity, source platform), the adjustment of supplementary weights, and the magnitude of the change in the rating result. The adjustment and update process must be logged in its entirety. The log content includes the unique identifier of the high-priority risk signal, the trigger time, the supplementary weight before the adjustment, the supplementary weight after the adjustment, the duration, the score and level of the rating result before and after the update, the review status (if necessary), the personnel handling the process, and other information. This log is stored in association with the key risk feature set file, credit rating calculation log, and data source credibility change trajectory, with a retention period of no less than one year. It supports traceability queries by assessed object, time range, and signal type, ensuring the traceability and verifiability of supplementary weight adjustments and rating result updates, fully meeting the core requirements of real-time adjustment and risk control.
[0114] This application also protects a credit rating assessment system that integrates multi-source data, including: a data access module for accessing three types of data: internal data from financial institutions, third-party data, and low-reliability supplementary data;
[0115] The data standardization module is used to unify the timestamps, field names, and data types of various types of data.
[0116] The data cleaning module is used to clean data by filling in missing values, removing outliers, and deduplicating duplicate data, laying the foundation for credibility assessment.
[0117] The credibility index quantification module is used to convert four major indicators—accuracy, completeness, timeliness, and stability—into quantitative values through calculations of data matching rate, field completeness rate, data freshness, and data quality fluctuation coefficient.
[0118] The comprehensive credibility calculation module is used to calculate the comprehensive credibility score of each data source by weighted summation according to preset weights, and to establish a mapping library;
[0119] The real-time monitoring module is used to start the real-time monitoring process and update the data source credibility score in real time.
[0120] The credibility stratification module is used to divide the credibility into three layers: high, medium and low, based on the comprehensive credibility score, and to assign corresponding fusion weights to each layer.
[0121] The risk signal extraction module is used to extract key risk signals from low-reliability data through keyword matching and sentiment analysis.
[0122] The credit rating calculation module is used to combine data features and key risk signals from each layer, calculate the raw score by weighting, and map the score to the corresponding credit rating level.
[0123] The weighting and rating adjustment module is used to adjust the fusion weights based on changes in the credibility score. If there are high-priority risk signals, the supplementary weights will be adjusted accordingly, and the credit rating results will be updated synchronously.
[0124] This application also protects a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the credit rating assessment method for multi-source data fusion as provided in the embodiments of the present invention.
[0125] The computer storage medium of this invention can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, system, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system, or device.
[0126] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit programs for use by or with an instruction execution system, system, or device.
[0127] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including—but not limited to—wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0128] The above description is merely one specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A credit rating assessment method based on multi-source data fusion, characterized in that, Including the following steps: It integrates three types of data: internal data from financial institutions, third-party data, and low-reliability supplementary data. It unifies timestamps, field names, and data types, and cleans the data by filling missing values, removing outliers, and deduplicating duplicate data, thus laying the foundation for credibility assessment. The accuracy, completeness, timeliness, and stability of the four major indicators are converted into quantitative values by calculating the data matching rate, field completeness rate, data freshness, and data quality fluctuation coefficient. The overall credibility score of each data source is obtained by weighted summation according to preset weights, a mapping library is established, and the real-time monitoring process is started. The layers are divided into high, medium, and low based on the overall credibility score, and corresponding fusion weights are assigned accordingly. For low-reliability data, key risk signals are extracted through keyword matching and sentiment analysis; By combining the data features and key risk signals from each layer, the raw score is calculated by weighting the data according to the weights, and the score is mapped to the corresponding credit rating level. The system updates the data source credibility score in real time, adjusts the fusion weights based on score changes, and adjusts supplementary weights accordingly if high-priority risk signals are present, while simultaneously updating the credit rating results.
2. The credit rating assessment method based on multi-source data fusion according to claim 1, characterized in that, The system integrates three types of data: data from within financial institutions, third parties, and data supplements with low credibility. It unifies timestamps, field names, and data types, and cleans the data by filling missing values, removing outliers, and deduplicating duplicate data. This lays the foundation for credibility assessment. Specifically, it performs rule-based mapping on fields with the same or different names across data sources, and cleans the data by filling missing values, removing outliers, and removing duplicate data according to data source priority and timestamp rules. This ensures that the integrated data has a unified format and that invalid data is thoroughly removed, providing a high-quality data foundation for subsequent credibility assessment.
3. The credit rating assessment method based on multi-source data fusion according to claim 1, characterized in that, The accuracy, completeness, timeliness, and stability of the four indicators are converted into quantitative values by calculating the data matching rate, field completeness rate, data freshness, and data quality fluctuation coefficient. Specifically, the core fields that have passed internal verification by financial institutions are used as the benchmark, the matching rate of third-party data is calculated by associating it with the benchmark data, and low-reliability data is verified by random sampling, and the accuracy is quantified respectively. Determine the list of core fields for each data source and quantify completeness by the proportion of non-missing core fields; quantify timeliness by the longest valid period of data and the time difference between generation and storage; calculate the fluctuation coefficient based on the median reliability value within the statistical period, exclude abnormal period data, and adjust for insufficient periods according to the corresponding coefficient to quantify stability.
4. The credit rating assessment method based on multi-source data fusion according to claim 1, characterized in that, The process involves obtaining a comprehensive credibility score for each data source by weighted summation according to preset weights, establishing a mapping library, and initiating a real-time monitoring process. Specifically, this includes: setting the weights of the four major credibility indicators, which can be adjusted according to the process after filing; verifying the validity of the indicators; calculating the comprehensive credibility score of each data source according to a preset formula; assigning a unique identifier to each data source; establishing a mapping library to store key information of the data sources and using an adaptive storage method to ensure query and update efficiency; and initiating a real-time monitoring process to periodically collect and verify data and store it in the warehouse.
5. The credit rating assessment method based on multi-source data fusion according to claim 1, characterized in that, The process of dividing data into high, medium, and low layers based on comprehensive credibility scores and assigning corresponding fusion weights is as follows: the comprehensive credibility scores are divided into high, medium, and low layers according to preset rules, and the weights are allocated according to the score gradient. Data from financial institutions is assigned the highest weight by default in the high credibility layer, and the weights of other data sources are positively matched with their scores. Only the comprehensive credibility scores are associated without additional adjustments based on type. The weight allocation rules are linked with the mapping library in real time, and the layering is automatically verified and the weights are matched as the credibility scores are updated.
6. The credit rating assessment method based on multi-source data fusion according to claim 1, characterized in that, For low-reliability data, key risk signals are extracted through keyword matching and sentiment analysis. Specifically, taking financial credit risk scenarios as the core, a multi-dimensional risk keyword library containing synonym variant word mappings is constructed and updated regularly. A combination of exact matching and semantic similarity matching is used, along with a pre-trained model for sentiment analysis. Valid risk signals are filtered based on the relevance of the publishing entity, authority, validity of the publishing time, and deduplication conditions. The signals are structured to generate feature sets, invalid information is removed, and suspected high-risk records awaiting review are marked and included in manual review. The review results are used to optimize the model and keyword library, and the extraction process is monitored in real time.
7. The credit rating assessment method based on multi-source data fusion according to claim 1, characterized in that, The process of splicing data features and key risk signals across different layers is as follows: using the customer's unique identifier as the core association key, effective structured features of high, medium, and low confidence layers are extracted and redundant content is eliminated. The key risk feature set of low confidence data is treated as an independent group and horizontally spliced with the three layers of structured features to form a credit rating feature matrix. Field deduplication rules are executed, the integrity of the matrix is verified, and missing non-core fields are labeled in a standardized manner. A unified timestamp is used to record the splicing time, and the process is linked with the real-time monitoring process for scheduled execution, retaining the update trajectory to support traceability.
8. The credit rating assessment method based on multi-source data fusion according to claim 1, characterized in that, The process of calculating the original score by weight and mapping the score to the corresponding credit rating level involves: standardizing the structured features of each layer, calculating the weighted sum of the structured features of each layer according to the corresponding fusion weight, multiplying the score of the key risk features by the supplementary weight according to the rules, and superimposing them to obtain the original score. The credit rating level is then mapped according to the preset rules. The threshold value can be fine-tuned by combining the risk signal and the proportion of high-confidence data. The calculation and mapping are performed synchronously with the feature matrix update every hour.
9. The credit rating assessment method based on multi-source data fusion according to claim 1, characterized in that, The real-time update of the data source credibility score and the adjustment of the fusion weights based on score changes are specifically as follows: the latest quality data of the data source is collected based on the real-time monitoring process, the credibility score is recalculated after validity verification, the weight adjustment is triggered according to preset conditions by comparing with the historical score, and the new weights are matched according to the hierarchical gradient rules.
10. A credit rating assessment system based on multi-source data fusion, used to implement the method according to any one of claims 1-9, characterized in that, include: The data access module is used to access three types of data: internal data from financial institutions, third-party data, and low-reliability supplementary data. The data standardization module is used to unify the timestamps, field names, and data types of various types of data. The data cleaning module is used to clean data by filling in missing values, removing outliers, and deduplicating duplicate data, laying the foundation for credibility assessment. The credibility index quantification module is used to convert four major indicators—accuracy, completeness, timeliness, and stability—into quantitative values through calculations of data matching rate, field completeness rate, data freshness, and data quality fluctuation coefficient. The comprehensive credibility calculation module is used to calculate the comprehensive credibility score of each data source by weighted summation according to preset weights, and to establish a mapping library; The real-time monitoring module is used to start the real-time monitoring process and update the data source credibility score in real time. The credibility stratification module is used to divide the credibility into three layers: high, medium and low, based on the comprehensive credibility score, and to assign corresponding fusion weights to each layer. The risk signal extraction module is used to extract key risk signals from low-reliability data through keyword matching and sentiment analysis. The credit rating calculation module is used to combine data features and key risk signals from each layer, calculate the raw score by weight, and map the score to the corresponding credit rating level. The weighting and rating adjustment module is used to adjust the fusion weights based on changes in the credibility score. If there are high-priority risk signals, the supplementary weights will be adjusted accordingly, and the credit rating results will be updated synchronously.