Enterprise credit assessment method and system based on multi-source heterogeneous data fusion
By acquiring and cleaning multi-source heterogeneous data, and using large models for in-depth analysis and calibration, the problem of low efficiency and accuracy in enterprise credit assessment has been solved, thereby improving enterprise management and business optimization.
Patent Information
- Application Number
- CN202510700354.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-10-28
AI Technical Summary
The diversity and complexity of data sources lead to low efficiency and accuracy in corporate credit assessment, and existing technologies struggle to effectively manage and organize multi-source heterogeneous data.
By acquiring data characteristics from multi-source heterogeneous data, we perform data cleaning, transformation, and fusion, utilize large models for in-depth analysis and calibration, and combine them with enterprise credit rating models for evaluation.
It improves the efficiency and accuracy of corporate credit assessment, reveals potential problems in business management, optimizes business layout, and enhances management level.
Smart Images

Figure CN120852028A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence assessment technology, specifically relating to a method and system for enterprise credit assessment based on the fusion of multi-source heterogeneous data. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] With the development of the digital economy, data has become a crucial production factor, making market-oriented allocation reform of data elements key. Trusted Data Space, as a new type of data infrastructure, can meet the consensus rules of multiple stakeholders, including data providers, users, and service providers, regarding business, management, and trust, enabling trusted data circulation and sharing, and providing richer data sources for corporate credit assessment. Data security and privacy protection are paramount in the process of data sharing and circulation. Trusted Data Space, through privacy computing technology, virtual sandbox technology, identity authentication, and authorization management, ensures the security, privacy, and trustworthiness of data during sharing and use, providing a reliable data environment for corporate credit assessment. Trusted Data Space can encapsulate data resources, products, and services from different sources into a unified data model, enabling unified data release, efficient querying, and cross-entity mutual recognition, improving the overall efficiency of data sharing and providing a more efficient data management and organization method for corporate credit assessment.
[0004] However, the diversity and complexity of data sources pose serious challenges to data management and organization, which will directly affect the efficiency and accuracy of corporate credit assessment. Summary of the Invention
[0005] To address the aforementioned issues, this invention proposes a method and system for enterprise credit assessment based on multi-source heterogeneous data fusion. The method involves acquiring data features from multi-source heterogeneous data, fusing the multi-source heterogeneous data based on these features, verifying and adjusting the fused multi-source heterogeneous data, and then completing the enterprise credit assessment within a trusted data space based on the verified multi-source heterogeneous fused data and the enterprise credit evaluation model.
[0006] According to some embodiments, the first aspect of the present invention provides a method for enterprise credit assessment based on multi-source heterogeneous data fusion, employing the following technical solution: A corporate credit assessment method based on multi-source heterogeneous data fusion includes: Obtain data from different business sources within the enterprise; Extract data features from different business data sources; Considering the extracted data features, determine the entities of data sources for different businesses, and perform multi-source heterogeneous data fusion of data sources from different businesses under the determined entities; The fused multi-source heterogeneous data is verified and adjusted to obtain multi-source heterogeneous verification data. Based on the obtained multi-source heterogeneous verification data and enterprise credit evaluation model, an enterprise credit assessment was completed by fusing multi-source heterogeneous data in a trusted data space.
[0007] As a further technical constraint, data extraction and transformation are performed on the obtained data from different data sources to build a unified data warehouse, identify the data source entities in the data warehouse, determine the same entity records described by different data sources, and establish the association relationship between the same entity; under the determined entity, combined with the data characteristics of different business data sources and the data fusion goal, a data fusion method is selected to perform multi-source heterogeneous data fusion.
[0008] As a further technical limitation, in the process of enterprise credit assessment, the obtained multi-source heterogeneous verification data is input into the enterprise credit assessment model to obtain the enterprise credit estimate prediction value. The obtained enterprise credit estimate prediction value is evaluated and analyzed by evaluation indicators including at least accuracy, recall, and mean square error. Based on the evaluation and analysis results, the enterprise credit estimate prediction value is calibrated and adjusted to complete the enterprise credit assessment in the trusted data space through multi-source heterogeneous data fusion.
[0009] As a further technical constraint, keywords are extracted from the data source for preliminary screening of data features, redundant features are removed, the preliminary screened data features are evaluated, and cross-validation is used to complete the verification and optimization of data features, thereby obtaining data features and completing the extraction of data features from different business data sources.
[0010] Furthermore, after data feature extraction, the system combines related features and time series features based on business logic, combines different data features to extract deeper information, and iterates continuously to obtain the optimal combination of data features and derived feature set.
[0011] As a further technical limitation, before extracting the data features of the different business data sources, the acquired data source data is cleaned to remove noise, duplicate data, missing values and outliers. The cleaned data source data is then normalized.
[0012] According to some embodiments, the second aspect of the present invention provides an enterprise credit assessment system based on multi-source heterogeneous data fusion, employing the following technical solution: A corporate credit assessment system based on multi-source heterogeneous data fusion includes: The acquisition module is configured to acquire data from different business data sources within the enterprise. The extraction module is configured to extract data features from different business data sources. The fusion module is configured to consider the extracted data features, determine the entities of the data sources of different businesses, and perform multi-source heterogeneous data fusion of data from different business data sources under the determined entities. The verification module is configured to verify and adjust the fused multi-source heterogeneous data to obtain multi-source heterogeneous verification data. The evaluation module is configured to perform enterprise credit evaluation by fusing multi-source heterogeneous data in a trusted data space, based on the obtained multi-source heterogeneous verification data and enterprise credit evaluation model.
[0013] According to some embodiments, a third aspect of the present invention provides a computer-readable storage medium, employing the following technical solution: A computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps of an enterprise credit assessment method based on multi-source heterogeneous data fusion as described in the first aspect of the present invention.
[0014] According to some embodiments, the fourth aspect of the present invention provides an electronic device, which adopts the following technical solution: An electronic device includes a memory, a processor, and a program stored in the memory and running on the processor, wherein the processor executes the program to implement the steps in the enterprise credit assessment method based on multi-source heterogeneous data fusion as described in the first aspect of the present invention.
[0015] According to some embodiments, the fifth aspect of the present invention provides a computer program product, which adopts the following technical solution: A computer program product includes software code, wherein the program in the software code performs the steps of an enterprise credit assessment method based on multi-source heterogeneous data fusion as described in the first aspect of the present invention.
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention acquires data features from multi-source heterogeneous data, fuses the multi-source heterogeneous data based on the acquired data features, verifies and adjusts the fused multi-source heterogeneous data, and completes enterprise credit assessment in a trusted data space based on the verified multi-source heterogeneous fused data and an enterprise credit evaluation model.
[0017] In the process of conducting corporate credit assessment, this invention reveals potential problems and deficiencies in corporate management and risk control through in-depth analysis of multi-dimensional data polarity, prompting companies to take targeted improvement measures, thereby enhancing their overall management level and operational efficiency; and optimizes the company's business layout through refined analysis of data from different business segments. Attached Figure Description
[0018] The accompanying drawings, which form part of this embodiment, are used to provide a further understanding of this embodiment. The illustrative embodiments and their descriptions are used to explain this embodiment and do not constitute an improper limitation of this embodiment.
[0019] Figure 1 This is a flowchart of a corporate credit assessment method based on multi-source heterogeneous data fusion according to Embodiment 1 of the present invention; Figure 2 This is a schematic diagram of data acquisition and preprocessing in Embodiment 1 of the present invention; Figure 3 This is a feature engineering business logic diagram in Embodiment 1 of the present invention; Figure 4 This is a business logic diagram of multi-source heterogeneous data fusion in Embodiment 1 of the present invention; Figure 5 This is a flowchart of the large model training and optimization process in Embodiment 1 of the present invention; Figure 6 This is a schematic diagram of the credit calculation and evaluation of the multi-source heterogeneous large model in Embodiment 1 of the present invention; Figure 7 This is a schematic diagram illustrating the data trust space guarantee in Embodiment 1 of the present invention; Figure 8 This is a structural block diagram of an enterprise credit assessment system based on multi-source heterogeneous data fusion, as shown in Embodiment 2 of the present invention. Detailed Implementation
[0020] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0021] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0022] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0023] In this invention, terms such as "upper," "lower," "left," "right," "front," "back," "vertical," "horizontal," "side," and "bottom" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are used only to facilitate the description of the structural relationships of the various components or elements of this invention and do not specifically refer to any component or element in this invention. They should not be construed as limiting the invention.
[0024] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0025] Example 1 Embodiment 1 of this invention introduces a method for enterprise credit assessment based on the fusion of multi-source heterogeneous data.
[0026] like Figure 1 The enterprise credit assessment method shown includes: Obtain data from different business sources within the enterprise; Extract data features from different business data sources; Considering the extracted data features, determine the entities of data sources for different businesses, and perform multi-source heterogeneous data fusion of data sources from different businesses under the determined entities; The fused multi-source heterogeneous data is verified and adjusted to obtain multi-source heterogeneous verification data. Based on the obtained multi-source heterogeneous verification data and enterprise credit evaluation model, an enterprise credit assessment was completed by fusing multi-source heterogeneous data in a trusted data space.
[0027] Regarding multi-source heterogeneous data, it's important to note that enterprise data comes from a wide range of sources, including structured, semi-structured, and unstructured data. These data vary significantly in structure, format, and semantics, creating data silos that are difficult to utilize effectively. Multi-source heterogeneous data processing technology can break down these data silos, integrating information from multiple heterogeneous data sources to form a more comprehensive, accurate, and consistent knowledge system, providing more comprehensive data support for enterprise credit assessment. To improve data quality and usability, multi-source heterogeneous data needs to be fused and preprocessed. Data cleaning, transformation, and integration technologies can handle issues such as errors, missing values, and duplicate data, converting data of different formats and types into a unified format and data type, and establishing relationships between data points, laying the foundation for subsequent data analysis and mining. Extracting key information reflecting enterprise credit characteristics from raw data and uncovering potential relationships between data points is crucial for enterprise credit assessment. Feature engineering and data mining techniques can extract, select, and reduce features to identify key variables that significantly impact corporate credit assessment, while also uncovering hidden relationships within the data, providing deeper insights for corporate credit evaluation.
[0028] Regarding the large-scale model, it's important to note that based on the Transformer architecture and self-attention mechanism, it possesses powerful language understanding and generation capabilities, enabling it to handle long-sequence data and capture long-term dependencies within the data. In corporate credit assessment, the large-scale model can conduct in-depth analysis of relevant textual data about enterprises, such as corporate announcements, news reports, and user reviews, thereby gaining a more comprehensive understanding of the enterprise's operating status and market reputation, providing richer information for credit assessment. It employs a pre-training and fine-tuning strategy, first pre-training on large-scale unsupervised or weakly supervised data to learn general language knowledge and semantic representations, and then fine-tuning for specific corporate credit assessment tasks. This strategy fully utilizes large amounts of unlabeled data, reducing data labeling costs while improving the model's performance and generalization ability on specific tasks, allowing the model to better adapt to the needs of corporate credit assessment. As the amount of data increases and computing power improves, the scale of the large-scale model continues to expand, and its representational and learning capabilities also continuously enhance. Meanwhile, to improve the performance and efficiency of large models, researchers are continuously exploring and applying various model expansion and optimization techniques, such as distributed training, model parallelism, mixed precision training, layer normalization, and adaptive learning rate adjustment. These techniques provide strong support for the application of large models in corporate credit assessment. In addition to textual data, corporate credit assessment can also incorporate multimodal data such as images and audio. Multimodal large models can fuse and collaboratively analyze data from different modalities, fully leveraging the advantages of each modality to provide a more comprehensive and multi-dimensional perspective for corporate credit assessment, thereby more accurately evaluating the overall creditworthiness of enterprises.
[0029] This embodiment assesses corporate credit through a multi-source heterogeneous large-scale model in a trusted data space. In terms of data fusion and governance, it efficiently integrates data from different sources and formats, improving data quality and refining the data governance framework. Regarding trusted computing and security, it applies privacy computing to build a comprehensive data security protection system. At the model training and optimization level, it designs an adaptive architecture, develops efficient algorithms, and conducts thorough optimization evaluations. In terms of intelligent decision-making and application, it can perform intelligent data analysis to assist decision-making, provide scenario-based customized services, and promote cross-domain collaborative innovation. These innovative aspects collectively drive its important role in multiple fields.
[0030] This embodiment uses, as follows: Figure 2 The process shown involves data collection and preprocessing, mainly consisting of three parts: (1) Collect data from multiple sources Data is collected from various data sources, including but not limited to transaction records from financial institutions, consumer behavior data from e-commerce platforms, public service data from government departments, energy consumption data from manufacturing, transaction data from logistics and transportation, equipment operation data, and product or semi-finished product procurement data. For example, bank loan repayment records, purchase frequency, product types, and consumption amounts on e-commerce platforms, as well as enterprise hot data such as water, electricity, gas, and energy consumption, logistics costs and frequency, equipment utilization rate, and costs of purchasing raw materials from third parties.
[0031] (2) Data cleaning The collected data undergoes data cleaning, including removing noise, duplicate data, missing values, and outliers. For example, for missing consumption amount data, methods such as filling in based on the company's historical consumption data or using the mean can be used. For obviously abnormal high-value consumption records, their authenticity needs to be further verified.
[0032] (3) Data standardization and normalization Data of different formats and ranges are uniformly converted into a standard format and range suitable for model processing; for example, consumption amount data is normalized so that its value range is between 0 and 1, so as to facilitate comparison and comprehensive calculation between different data.
[0033] This embodiment uses, as follows: Figure 3 The feature engineering shown specifically includes: (1) Feature extraction Extract representative and discriminative features from the raw data. For text data, natural language processing techniques can be used to extract features such as keywords; for image data, features such as image texture and color can be extracted; for structured data, relevant fields can be directly selected as features.
[0034] (2) Feature selection This embodiment first performs preliminary feature screening. Through business understanding and feature listing, along with correlation analysis and information gain methods, it identifies features that significantly impact credit assessment and removes redundant or obviously irrelevant features. Next, it uses statistical analysis and model-based methods to assess feature importance. Then, it implements feature selection strategies such as filtering, wrapping, and embedding. Finally, it uses cross-validation and other methods to validate and optimize feature subsets, iterating continuously to obtain an ideal feature subset, thereby aiding subsequent data analysis and modeling. For example, when analyzing a company's consumer behavior characteristics, it was found that the frequency of purchasing high-end equipment or products has a high correlation with the company's credit level, while the frequency of purchasing low-end equipment or products has a low correlation. Therefore, the frequency of purchasing high-end equipment or products can be selected as an important feature.
[0035] (3) Feature combination and derivation Understanding data and objectives, and combining related features and time-series features based on business logic, involves mathematical operations on numerical features and encoding categorical features. Feature selection methods are used to eliminate features of low value, and validation and optimization techniques are employed to combine different features or derive new features to uncover deeper information within the data. Continuous iteration yields ideal feature combinations and derived feature sets that effectively improve data analysis and modeling results. For example, combining a company's revenue and debt levels to calculate the debt-to-income ratio, a derived feature, provides a more comprehensive reflection of the company's solvency.
[0036] This embodiment adopts Figure 4 The multi-source heterogeneous data fusion illustrated involves extracting and transforming data from different data sources (such as data from different internal business systems and external partners) and loading it into a data warehouse. During the extraction process, ETL (Extract, Transform, Load) tools are used to aggregate the scattered data according to pre-defined rules, performing operations such as format unification and semantic alignment. For example, date data in different formats is uniformly converted into a specific standard format, and fields with the same meaning but different expressions are standardized. This allows various types of data to be integrated within a unified storage environment of the data warehouse, facilitating subsequent analysis and application.
[0037] As one or more implementation methods, this embodiment establishes a comprehensive metadata management mechanism to provide detailed descriptions of the data in the data warehouse, including data sources, data definitions, and relationships between data. This helps to better understand the characteristics of data from different sources, coordinate the data fusion process, ensure the quality of the fused data, and facilitate accurate querying and use of the fused data by users. It identifies records describing the same entity in different data sources and establishes relationships between them. For example, using unique identifiers such as a company's unified social credit code, ID card number, or mobile phone number, it matches and associates companies in financial institution data with companies in e-commerce platform data, enabling comprehensive utilization of data from different sources to assess a company's creditworthiness. Based on the characteristics of the data and the fusion objectives, this embodiment selects appropriate data fusion algorithms, such as weighted average, Bayesian networks, and neural networks. If the data quality and reliability of different data sources differ, a weighted average can be used to assign greater weights to higher-quality data. For complex nonlinear relational data, neural networks can be used for fusion processing.
[0038] This embodiment verifies the fused data to ensure its consistency and accuracy; if any unreasonable aspects are found in the fused data, timely adjustments and corrections are necessary.
[0039] The process of multi-source heterogeneous data fusion in this embodiment includes: 1. Based on business rules, conduct checks according to existing rules in the business areas involved in the integrated data; 2. Entity relationship logic check: When the merged data involves multiple related entities, verify whether the relationships between them are logical. 3. Compare with authoritative data sources: If there are recognized authoritative data sources, compare the merged data with them; 4. Check for duplicate data: Look for duplicate records in the merged dataset, as duplicate data may cause bias in statistical analysis and other results; 5. Field Integrity Check: Check whether each field of each data record after merging has the corresponding value and whether there are any missing values; 6. Verify the integrity of the data volume to confirm whether the overall merged data volume meets expectations; 7. Timestamp check: For data with timestamps, check whether the time order and time range of the merged data are reasonable; 8. Statistical analysis methods: Use statistical tools, such as calculating the mean and standard deviation, and then set a reasonable interval based on the distribution of the data to determine the data points that exceed the interval as outliers. 9. Based on model detection, machine learning and other related models are used to train the model on historical normal data, allowing it to learn normal data patterns. This model is then used to determine which data in the fused dataset are outliers. For example, clustering models can be used; normal data points tend to cluster within specific clusters, while data points isolated outside these clusters may be outliers and need to be identified and processed. The above fusion process requires further examination of the data source and the fusion process to identify and correct any problems.
[0040] This embodiment uses, as follows: Figure 5 The large model shown is trained and optimized. Specifically, the Transformer large model architecture is selected based on factors such as the scale of the data, the complexity of the features, and computing resources.
[0041] This embodiment utilizes a multi-layered neural network structure to fully extract features and patterns from the data, including: 1. Multi-head attention mechanism: The model simultaneously focuses on input information from different representation subspaces, enabling it to capture the relationships between different positions in the input sequence, which is very effective for processing long sequence data. For example, in natural language processing, when analyzing the semantic relationships between different words in a sentence, the multi-head attention mechanism allows the model to consider the mutual influence between words from multiple perspectives.
[0042] 2. Feedforward neural networks perform further nonlinear transformations on the information processed by the attention mechanism, enhancing the model's expressive power and enabling more complex mapping of input features to extract higher-level features.
[0043] 3. Layer normalization helps stabilize the training process, making gradient propagation smoother during training, reducing problems such as gradient vanishing or exploding, improving the efficiency and effectiveness of model training, and ensuring that the distribution of input data in different layers is relatively stable.
[0044] 4. Residual connections make it easier to train deep models by directly adding the input of the previous layer to the output of the next layer, allowing gradients to propagate back more smoothly. This avoids performance degradation in deep networks and allows the model to continuously stack layers to improve its ability to process complex information.
[0045] This embodiment utilizes large-scale unsupervised data to pre-train a large model, learning the general features and patterns of the data. Then, supervised labeled data is used to fine-tune the pre-trained model to adapt it to the specific task of credit score calculation. For example, a language model is first pre-trained using massive amounts of text data to learn the grammar and semantics of the language, and then labeled credit data is used to fine-tune the model so that it can accurately predict credit scores.
[0046] This embodiment improves the model's performance and generalization ability by adjusting its hyperparameters and optimizing algorithms. Optimization algorithms such as stochastic gradient descent, Adagrad, and Adadelta can be used to continuously adjust the model's weight parameters until it converges to the optimal solution. Simultaneously, regularization techniques, such as L1 and L2 regularization and Dropout, can be employed to prevent overfitting. The specific process is as follows: (1) Optimization stage Data optimization and expansion involve using data augmentation techniques to increase data diversity, such as rotation and flipping operations in image processing and synonym replacement in natural language processing. This allows the model to learn data features from more diverse situations, enhancing its generalization ability. During data sampling, to address class imbalance, undersampling or oversampling methods are employed to balance the amount of data in each class, enabling the model to better learn the features of each data type and avoid overfitting or underlearning to a particular class. Model structure optimization involves adjusting network depth and width, appropriately increasing or decreasing the number of network layers and neurons per layer to enhance model expressive power, avoid overfitting, and reduce computational resource consumption. Pruning can also be used to remove unimportant connections and neurons, optimizing the model structure. Adopting different architectures or module combinations, such as replacing a simple architecture with a more advanced one suitable for the task, or introducing advanced modules like attention mechanisms, can improve the model's ability to capture key information and enhance overall performance. Algorithm optimization involves adjusting relevant parameters based on convergence during model training. This includes changing the learning rate and adjusting coefficients in adaptive learning rate algorithms to enable the model to converge to a better solution more stably and quickly. Hyperparameter tuning utilizes methods such as grid search, random search, and Bayesian optimization to systematically explore different combinations of hyperparameter values, finding the optimal hyperparameter settings to further improve model performance.
[0047] (2) Adjustment phase Early stopping strategy involves continuously monitoring the model's performance on the validation set (e.g., accuracy, loss value) during training. When performance stops improving or even begins to decline, training is stopped early to prevent overfitting of the training data and preserve the optimal model parameters at that point. Ensemble learning combines multiple trained models (which can be the same model structure with different initializations or hyperparameter settings, or models with different structures). Methods such as averaging prediction results (regression tasks) or voting (classification tasks) can improve the overall predictive performance and stability of the model, reducing variance and bias. Continuous learning and fine-tuning allows the model to continuously learn as new data is generated. Based on the existing model, fine-tuning is performed using new data. This can involve fixing the parameters of some model layers and updating only the parameters of specific layers, allowing the model to retain previously learned knowledge while adapting to new data features and task requirements.
[0048] This embodiment uses, as follows: Figure 6 The multi-source heterogeneous large model shown is used for calculating and evaluating corporate credit: As one or more implementation methods, preprocessed and fused multi-source heterogeneous data is input into a trained large model to obtain the user's credit score prediction result.
[0049] Data preprocessing and fusion review: Before inputting the data into the trained large model, preprocessing and fusion were completed. Preprocessing encompassed operations such as missing value handling (filling in or estimating missing parts of the data appropriately), outlier handling (identifying and reasonably correcting or removing abnormal data points), data standardization or normalization (unifying the data scale of different features), and feature encoding for categorical variables (converting categorical features into numerical forms that can be recognized by the model using one-hot encoding). Data fusion integrated data from different data sources, removing duplicate and redundant information and extracting information valuable for credit score prediction. 2. Data format adjustment: ensuring that the format of the preprocessed and fused data matches the input format required by the trained large model. If the large model requires the input to be in tensor form with a specific dimension, the data is transformed accordingly, involving data reshaping and dimensional adjustment to ensure that the data is successfully received by the model.
[0050] In this embodiment, the formatted multi-source heterogeneous data is accurately input into the model according to the input interface requirements of the large model. This process follows the model's established input specifications, such as the order and data type of the input data, and strictly adheres to these requirements to ensure that the model correctly interprets the input data.
[0051] In this embodiment, the feature extraction and representation learning utilizes a well-trained large model with powerful feature extraction capabilities, automatically mining deep-level feature information from multi-source heterogeneous input data. For data features from different sources, the large model uses a multi-head attention mechanism within its internal multi-layer architecture to fuse and encode these features, transforming them into more abstract feature representations that better reflect the essence of user credit. 2. Inference based on learned patterns: Based on the relationship patterns between user features and credit scores learned during the training phase, the large model infers from the current input data. If, during training, features such as stable high income, rational consumption behavior, and positive social credit feedback are found to be associated with higher credit scores, then when new input data exhibits similar features, the model will perform corresponding credit score inference and calculation according to the learned patterns. The model will output a specific credit assessment score based on the features of the input data and the learned patterns. The credit score ranges from 30 to 100 points and corresponds to a certain credit rating interval (such as good, medium, poor, etc.), depending on the output format set during model training and the specific application scenario requirements.
[0052] As one or more implementation methods, this embodiment uses evaluation metrics such as accuracy, recall, F1 score, and mean squared error to evaluate and analyze the model's prediction results. A high accuracy indicates that the model can accurately predict the user's credit score; a low mean squared error indicates that the deviation between the model's prediction results and the true value is small, and the model's performance is good.
[0053] As one or more implementation methods, this embodiment performs credit score calibration and adjustment. Based on actual conditions, business needs, and model rules, the credit score is calibrated and adjusted. For example, the credit score of users in certain high-risk industries may be appropriately lowered; for long-term users with good credit records, a certain credit bonus may be given. Simultaneously, the credit score calculation model can be periodically retrained and adjusted according to changes in the market environment and data updates to ensure the accuracy and effectiveness of the credit score. Specifically: (1) Calibration of mean and standard deviation A certain scale of sample data with labeled real credit scores is collected and input into a large model to obtain the corresponding predicted credit scores. The mean deviation and standard deviation difference between the predicted credit scores and the real credit scores are calculated. If the mean of the predicted credit scores is found to be higher than the mean of the real credit scores by a certain value, it indicates that there is an overall overestimation, which is corrected by subtracting the corresponding mean deviation. At the same time, if the standard deviation difference is large, it means that the dispersion of the predicted results does not match the actual situation. The predicted credit scores are scaled and adjusted accordingly based on the standard deviation ratio to make their dispersion closer to the actual data.
[0054] (2) Quantile calibration The actual credit score and projected credit score are arranged in ascending order, dividing them into different quantile intervals (such as quartiles, decimals, etc.). The actual and projected credit scores are compared within each quantile interval. If, within a certain quantile interval, the projected credit score is generally higher than the actual credit score, then corresponding rules can be established to appropriately lower the projected credit score within that interval; conversely, if the projected credit score is lower than the actual credit score, it is adjusted upwards, so that the projected results more closely match the actual situation within each quantile interval.
[0055] (3) Business rule calibration The model is calibrated based on credit assessment rules specific to the business scenario. For example, in financial lending, there may be rules for adding or deducting credit scores for users in certain industries under specific circumstances. After the large model outputs the predicted credit score, the predicted credit score is adjusted accordingly based on the user's industry and other attributes, following the business rules. The model also considers the user's historical credit record. If a user has a history of timely and good performance, their predicted credit score can be increased according to established rules; conversely, if there are negative records such as overdue payments or defaults, their predicted credit score will be appropriately decreased.
[0056] (4) Adjustment of rules based on external factors Adjustments should be made in conjunction with the macroeconomic situation. During periods of economic prosperity, when the overall credit environment is relatively favorable, the predicted credit scores for users can be appropriately increased (within a reasonable range). Conversely, during economic downturns, considering increased potential risks, the predicted credit scores for some users can be moderately lowered. Credit scores should also be adjusted in accordance with industry development trends. If an industry is experiencing rapid growth and has a promising future, the predicted credit scores for users within that industry can be appropriately increased according to relevant rules; conversely, if an industry faces difficulties or stricter regulations, its predicted credit scores should be lowered accordingly.
[0057] (5) Multi-model fusion calibration In addition to the currently used large model, other different types of credit assessment models (such as traditional statistical models, machine learning decision tree models, etc.) can be selected. The same preprocessed and fused multi-source heterogeneous data can be input into these models to obtain their respective credit score predictions. The predicted credit scores from multiple models can be merged using fusion strategies (such as simple weighted averaging, weighted averaging based on model performance, etc.). Then, the original predicted credit scores from the large model are compared, the differences are analyzed, and the credit score output by the large model is calibrated and adjusted based on the fusion results to make it closer to the accurate result after integrating the advantages of multiple models.
[0058] (6) Dynamic adjustment strategy Based on monitoring results, once a systematic bias in the predicted credit score is detected (such as an overall bias towards higher or lower scores) or significant inaccuracies are observed under certain specific circumstances (such as specific regions or user types), corresponding adjustment strategies will be immediately initiated. These strategies will be applied flexibly according to actual conditions to ensure that the predicted credit score consistently and accurately reflects the user's true creditworthiness. The user credit score predictions output by the large model will be continuously optimized to make them more realistic and reliable, providing accurate decision-making support for various credit-related businesses.
[0059] like Figure 7 As shown, the protection of the trusted data space in this embodiment includes at least the following: (1) Data security and privacy protection: Ensure data security and privacy throughout the credit score calculation process. Employ technologies such as data encryption, access control, and anonymization to prevent data leakage and misuse. For example, sensitive user information is encrypted during storage and transmission, and only authorized personnel can access and use this data.
[0060] (2) Data Quality Monitoring and Auditing: Establish a data quality monitoring mechanism to regularly check and evaluate data quality. Simultaneously, conduct data audits, recording data access, usage, and modification to trace and identify the root causes of data quality problems. If a decline in data quality or any anomalies are detected, take timely measures to repair and improve the data.
[0061] (3) Model Interpretation and Interpretability: Improve the interpretability of large models so that the calculation results of credit scores can be understood and accepted by users and relevant institutions. Model interpretation techniques, such as feature importance analysis and local interpretability models, can be used to explain how the model derives credit scores from input data, thereby enhancing the transparency and credibility of the model.
[0062] This embodiment utilizes a multi-source heterogeneous large-scale model to assess corporate credit in a trusted data space. It integrates massive, multi-source, and heterogeneous data resources, providing richer and more accurate data support for the construction of a social credit system. This further improves credit infrastructure and promotes the digital and intelligent development of the social credit system. It breaks down information silos between different entities, such as enterprises, financial institutions, and the government, enabling efficient sharing and circulation of credit information. This increases the transparency of credit across society, allowing all entities to gain a more comprehensive and accurate understanding of their counterparts' creditworthiness in economic activities, reducing the risks associated with information asymmetry. By accurately assessing corporate credit, it incentivizes trustworthy enterprises and constrains and punishes dishonest ones, guiding enterprises to establish a sense of integrity and regulate market behavior. This fosters a positive atmosphere of honesty and trustworthiness throughout society, continuously optimizing the social credit environment and promoting the healthy and sustainable development of the social economy.
[0063] This embodiment enables enterprises to store data according to unified standards and specifications, improving data consistency and availability. In the manufacturing supply chain, storing production and quality data from different enterprises according to unified standards facilitates data sharing and interaction between upstream and downstream enterprises, reduces communication costs and errors caused by inconsistent data formats, improves supply chain collaboration efficiency, and enhances the resilience of the supply chain to cope with complex business processes. Real-time and accurate data transmission allows enterprises in the supply chain to understand the production, inventory, and logistics information of upstream and downstream enterprises in a timely manner, realizing visualized and collaborative management of the supply chain. Taking the automotive supply chain as an example, parts suppliers can adjust their production schedules and inventory levels according to the production plans of vehicle manufacturers, and logistics companies can arrange transportation plans in advance, improving the responsiveness and flexibility of the entire supply chain, reducing the risks of inventory backlog and production interruptions caused by poor information flow, and enhancing the resilience of the supply chain to changes in demand and unexpected events.
[0064] Example 2 Embodiment 2 of the present invention introduces an enterprise credit assessment system based on multi-source heterogeneous data fusion.
[0065] like Figure 8 The enterprise credit assessment system shown includes: The acquisition module is configured to acquire data from different business data sources within the enterprise. The extraction module is configured to extract data features from different business data sources. The fusion module is configured to consider the extracted data features, determine the entities of the data sources of different businesses, and perform multi-source heterogeneous data fusion of data from different business data sources under the determined entities. The verification module is configured to verify and adjust the fused multi-source heterogeneous data to obtain multi-source heterogeneous verification data. The evaluation module is configured to perform enterprise credit evaluation by fusing multi-source heterogeneous data in a trusted data space, based on the obtained multi-source heterogeneous verification data and enterprise credit evaluation model.
[0066] The detailed steps are the same as those provided in Example 1 for a corporate credit assessment method based on multi-source heterogeneous data fusion, and will not be repeated here.
[0067] Example 3 Embodiment 3 of the present invention provides a computer-readable storage medium.
[0068] A computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps of an enterprise credit assessment method based on multi-source heterogeneous data fusion as described in Embodiment 1 of the present invention.
[0069] The detailed steps are the same as those provided in Example 1 for a corporate credit assessment method based on multi-source heterogeneous data fusion, and will not be repeated here.
[0070] Example 4 Embodiment 4 of the present invention provides an electronic device.
[0071] An electronic device includes a memory, a processor, and a program stored in the memory and running on the processor. When the processor executes the program, it implements the steps in the enterprise credit assessment method based on multi-source heterogeneous data fusion as described in Embodiment 1 of the present invention.
[0072] The detailed steps are the same as those provided in Example 1 for a corporate credit assessment method based on multi-source heterogeneous data fusion, and will not be repeated here.
[0073] Example 5 Embodiment 5 of the present invention provides a computer program product.
[0074] A computer program product includes software code, wherein the program in the software code executes the steps of an enterprise credit assessment method based on multi-source heterogeneous data fusion as described in Embodiment 1 of the present invention.
[0075] The detailed steps are the same as those provided in Example 1 for a corporate credit assessment method based on multi-source heterogeneous data fusion, and will not be repeated here.
[0076] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present invention can be implemented using various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.
[0077] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0078] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0079] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0080] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0081] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
[0082] The above description is merely a preferred embodiment of this practice and is not intended to limit the scope of this practice. Various modifications and variations can be made to this practice by those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this practice should be included within the protection scope of this practice.
Claims
1. A method for enterprise credit assessment based on multi-source heterogeneous data fusion, characterized in that, include: Obtain data from different business sources within the enterprise; Extract data features from different business data sources; Considering the extracted data features, determine the entities of data sources for different businesses, and perform multi-source heterogeneous data fusion of data sources from different businesses under the determined entities; The fused multi-source heterogeneous data is verified and adjusted to obtain multi-source heterogeneous verification data. Based on the obtained multi-source heterogeneous verification data and enterprise credit evaluation model, an enterprise credit assessment was completed by fusing multi-source heterogeneous data in a trusted data space.
2. The enterprise credit assessment method based on multi-source heterogeneous data fusion as described in claim 1, characterized in that, Data extraction and transformation are performed on the obtained data from different data sources to build a unified data warehouse. Data source entities in the data warehouse are identified, and records of the same entity described by different data sources are determined. Relationships between the same entities are established. Under the determined entities, data fusion methods are selected to perform multi-source heterogeneous data fusion based on the data characteristics of different business data sources and the data fusion objectives.
3. The enterprise credit assessment method based on multi-source heterogeneous data fusion as described in claim 1, characterized in that, In the process of corporate credit assessment, the obtained multi-source heterogeneous validation data is input into the corporate credit assessment model to obtain the corporate credit estimate prediction value. The obtained corporate credit estimate prediction value is evaluated and analyzed by evaluation indicators including at least accuracy, recall, and mean square error. Based on the evaluation and analysis results, the corporate credit estimate prediction value is calibrated and adjusted to complete the corporate credit assessment in the trusted data space through multi-source heterogeneous data fusion.
4. The enterprise credit assessment method based on multi-source heterogeneous data fusion as described in claim 1, characterized in that, Keywords are extracted from the data source for preliminary data feature screening. Redundant features are removed, and the preliminary data features are evaluated. Cross-validation is then used to verify and optimize the data features, thus obtaining the data features and completing the extraction of data features from different business data sources.
5. The enterprise credit assessment method based on multi-source heterogeneous data fusion as described in claim 4, characterized in that, After data feature extraction, the data features are combined with related features and time series features based on business logic. Different data features are combined and derived to mine deeper information. The process is iterated continuously to obtain the optimal combination of data features and the derived feature set.
6. The enterprise credit assessment method based on multi-source heterogeneous data fusion as described in claim 1, characterized in that, Before extracting the data features of the different business data sources, the acquired data source data is cleaned to remove noise, duplicate data, missing values and outliers. The cleaned data source data is then normalized.
7. A corporate credit assessment system based on multi-source heterogeneous data fusion, characterized in that, include: The acquisition module is configured to acquire data from different business data sources within the enterprise. The extraction module is configured to extract data features from different business data sources. The fusion module is configured to consider the extracted data features, determine the entities of the data sources of different businesses, and perform multi-source heterogeneous data fusion of data from different business data sources under the determined entities. The verification module is configured to verify and adjust the fused multi-source heterogeneous data to obtain multi-source heterogeneous verification data. The evaluation module is configured to perform enterprise credit evaluation by fusing multi-source heterogeneous data in a trusted data space, based on the obtained multi-source heterogeneous verification data and enterprise credit evaluation model.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of an enterprise credit assessment method based on multi-source heterogeneous data fusion as described in any one of claims 1-6.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the program, it implements the steps of the enterprise credit assessment method based on multi-source heterogeneous data fusion as described in any one of claims 1-6.
10. A computer program product, comprising software code, characterized in that, The program in the software code executes the steps of a corporate credit assessment method based on multi-source heterogeneous data fusion as described in any one of claims 1-6.
Citation Information
Cited By
Credit investigation data processing method and device
CN121120237A
Credit investigation data processing method and device
CN121437131A
Credit investigation data processing method and device
CN121437131B