A data quality assessment method and system for enterprise credit information fusion
By preprocessing, feature selection, reliability analysis and model optimization of enterprise credit information, the problems of data inconsistency and conflict in traditional evaluation methods are solved, and efficient and accurate credit information evaluation is achieved, which is suitable for data quality evaluation of different enterprises.
Patent Information
- Application Number
- CN202510089696.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-01-21
AI Technical Summary
In the evaluation of enterprise credit information, traditional methods rely on preset data, which poses risks of untimely updates, incomplete information or artificial tampering. The introduction of market data leads to wide data sources, diverse formats and conflict handling problems, affecting the accuracy and reliability of the assessment.
By collecting enterprise presets and market data, performing preprocessing, evaluating feature selection, reliability analysis, and bias analysis, constructing quality evaluation functions and optimizing the credit fusion data quality evaluation model, and using time series comparison, isolated forest and machine learning algorithms for data fusion and adjustment.
It improves the accuracy and reliability of enterprise credit information fusion data quality evaluation, realizes automated and real-time data fusion and adjustment, adapts to the evaluation needs of different standards and enterprises, saves resources and improves efficiency.
Smart Images

Figure CN119904148B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data evaluation, and in particular to a data quality evaluation method and system in enterprise credit information fusion. Background Art
[0002] In today's business environment, the accuracy and completeness of corporate credit information is crucial for financial institutions, investors, and various business partners. Corporate credit information not only reflects a company's ability to fulfill its obligations and operating conditions but also directly impacts its financing, partnerships, and market competitiveness. Therefore, effectively integrating and evaluating the quality of corporate credit information has become a major challenge in the current field of business information analysis.
[0003] Traditional corporate credit assessment methods rely primarily on pre-defined data such as financial statements and audit reports provided by companies. However, this data often faces risks of delayed updates, incomplete information, or tampering, raising questions about the accuracy and reliability of assessment results. To address this shortcoming, some institutions are experimenting with incorporating market data, such as publicly available corporate information, industry reports, and media reports, to enrich and deepen credit assessments.
[0004] However, the introduction of market data also presents new challenges. On the one hand, market data comes from a wide range of sources and formats, requiring effective preprocessing and integration to ensure accuracy and consistency. On the other hand, market data can differ or conflict with the company's pre-set data. Managing these discrepancies effectively is a key factor affecting the quality of credit assessments. Therefore, Kou Dai developed a new data quality assessment method to improve the accuracy, reliability, and comprehensiveness of enterprise data quality assessments. Summary of the Invention
[0005] The purpose of the present invention is to provide a data quality assessment method in enterprise credit information fusion.
[0006] To achieve the above object, the present invention is implemented according to the following technical solutions:
[0007] The present invention comprises the following steps:
[0008] Collecting preset data and market data of the enterprise and pre-processing the preset data and the market data; the preset data includes enterprise credit data and related data;
[0009] Performing feature selection on the relevant data to obtain first data, and performing reliability analysis on the market data based on the first data to obtain reliable data;
[0010] Performing deviation analysis on the reliable data to obtain a difference coefficient, constructing a quality assessment function based on the difference coefficient, the reliable data, and the enterprise credit data, and building a credit fusion data quality assessment model based on the quality assessment function;
[0011] The credit fusion data quality assessment model is optimized according to the assessment error, the data to be assessed is input into the credit fusion data quality assessment model, and an assessment result is output.
[0012] Furthermore, the method of performing feature selection on the relevant data to obtain first data includes:
[0013] Calculate the contribution of relevant data:
[0014]
[0015] The zth related data is b z , the contribution of the zth related data is The difference between the zth related data and the ath related data is F z,a , the average value of the difference value of the relevant data is The variance explanation rate of the zth related data is ξ z , related data b z The weight is λ z , the number of relevant data is
[0016] Calculate the evaluation value of the relevant data:
[0017]
[0018] The relevant data b z The evaluation value is S(b z ), the a-th related data is b a , the number of relevant data is Related datab z The marginal probability distribution of is p(b z ), related data b a The marginal probability distribution of is p(b a ), related data b z and related data b a The joint probability distribution of z ,b a ), related data b a The contribution of The zth adjustment coefficient is τ z , related data b z and related data b a The minimum function is min(b z ,b a);
[0019] The correlation data having an evaluation value greater than 0.4071 is output as the first data.
[0020] Furthermore, the method of performing reliability analysis on the market data based on the first data to obtain reliable data includes:
[0021] Compare and screen market data with market data, and consider data with a difference greater than 0.307 as abnormal data. Calculate the reliability of market data:
[0022]
[0023] The reliability of the jth market data at the sth moment is g j (s), the number of market data at time s is N s , the number of abnormal data at time s is NC s , the jth market data at time s is u j (s), the jth first data at the sth moment is u1 j (s);
[0024] When the reliability is greater than 0.6931, the market data is reliable data; otherwise, it is unreliable data and is eliminated.
[0025] Furthermore, the method of performing deviation analysis on the reliable data to obtain the coefficient of variation includes:
[0026] Sort the reliable data according to the time series and input the sorted reliable data into the deviation analysis model;
[0027] Calculate the change in reliable data:
[0028]
[0029] The cth reliable data at the sth moment is f c (s), the cth reliable data at the s+1th moment is f c (s+1), the bandwidth parameter is The suppression coefficient is γ, and the number of reliable data is The kernel density estimation function is The norm function is ||·||, and the change in the cth reliable data between the sth moment and the s+1th moment is h c (s,s+1);
[0030] Calculate the sensitive correlation of reliable data:
[0031]
[0032] The sensitive correlation of the cth reliable data is The upper limit of the timing is T s , the average value of the cth reliable data is The control constant is ε;
[0033] Take the reliable data with a sensitivity correlation greater than 0 as the key data and calculate the difference coefficient of the key data:
[0034]
[0035] The cth key data is f c , the coefficient of difference for the dth evaluation is
[0036] Furthermore, a method for constructing a quality assessment function based on the variance coefficient, reliable data, and corporate credit data includes:
[0037] Obtain enterprise credit through enterprise credit information and give the quality evaluation function expression:
[0038]
[0039] The coefficient of variation for the dth evaluation is The enterprise credit assessed in the dth time is y d , the quality evaluation function of the d-th evaluation is The jth reliable data is q j , the number of reliable data is M1, and the number of evaluation data is M.
[0040] Furthermore, a method for constructing a credit fusion data quality assessment model based on the quality assessment function includes:
[0041] According to the quality assessment function, the objective function of the credit fusion data quality assessment model is given as follows:
[0042]
[0043] The objective function of the d-th evaluation is The actual quality assessment value of the d-th evaluation is The loss function is L(·,·), and the prediction quality evaluation value of the d-th evaluation is
[0044] Credit fusion data quality assessment models include time series comparison algorithm, isolation forest algorithm, and machine learning algorithm;
[0045] The time series comparison algorithm analyzes and compares the trends of different time series data and identifies the differences between input data to obtain time-varying features;
[0046] The isolation forest algorithm constructs multiple isolated trees, uses randomly selected features to isolate time-varying features, calculates the average path length in all isolated trees, and marks time-varying features with a small number of isolations and a path length less than the average path length as abnormal to obtain labeled data;
[0047] The machine learning algorithm learns the objective function and the quality assessment pattern of historical data by building a model, and automatically evaluates and labels the input data using the quality assessment pattern and labeled data.
[0048] Furthermore, the method for optimizing the credit fusion data quality assessment model according to the assessment error includes:
[0049] Introducing the intelligent group optimization algorithm, the evaluation error is used as the fitness function, and the minimum evaluation error is used as the search goal;
[0050] Perform chaotic mapping on the particle population, and the expression is:
[0051]
[0052] The i-th original sequence is x i , the i+1th chaotic mapping sequence is x i+1 , the control parameter is η, and the random number between 0 and 1 is r1;
[0053] Take the particle position with the smallest fitness as the optimal position and calculate the particle position:
[0054]
[0055] The initial optimal position of the particle is The control coefficient is v, and the position of the wth particle in the t+1th iteration is
[0056] Use the inhibition factor to update the particle position and obtain the inhibition position. The expression is:
[0057]
[0058] The suppression factor is χ, and the random number between -1 and 1 is The optimal position of the particle at the tth iteration is The position of the wth particle at the tth iteration is The suppression position of the wth particle in the t+1th iteration is The current number of iterations is t, and the maximum number of iterations is t max ;
[0059] Adaptive weights are used to update the particle positions to obtain the adapted positions. The expression is:
[0060]
[0061] The adaptive weight of the tth iteration is The balance factor is σ, the attenuation coefficient is ρ, and the suppression position of the wth particle in the tth iteration is The normally distributed random number is β, and the optimal position of the t-th iteration is The adapted position of the wth particle in the t+1th iteration is
[0062] Iterate continuously until the evaluation error is minimized, otherwise adjust the particle population and update the adaptive weight.
[0063] Second, a data quality assessment system for enterprise credit information integration includes:
[0064] Data collection module: used to collect the enterprise's preset data and market data, and pre-process the preset data and the market data; the preset data includes enterprise credit data and related data;
[0065] Selection and analysis module: used for performing feature selection on the relevant data to obtain first data, and performing reliability analysis on the market data based on the first data to obtain reliable data;
[0066] Deviation modeling module: used to perform deviation analysis on the reliable data to obtain a difference coefficient, construct a quality assessment function based on the difference coefficient, the reliable data and the enterprise credit data, and build a credit fusion data quality assessment model based on the quality assessment function;
[0067] Optimization output module: used to optimize the credit fusion data quality assessment model according to the assessment error, input the data to be assessed into the credit fusion data quality assessment model, and output the assessment result.
[0068] The beneficial effects of the present invention are:
[0069] The present invention is a data quality assessment method and system for enterprise credit information fusion. Compared with the prior art, the present invention has the following technical effects:
[0070] The present invention can improve the accuracy of data quality assessment in enterprise credit information fusion through preprocessing, evaluation feature selection, reliability analysis, deviation analysis, construction of quality assessment function, model building and model optimization steps, thereby improving the precision of data quality assessment in enterprise credit information fusion, optimizing data quality assessment in enterprise credit information fusion, greatly saving resources and improving work efficiency, and can realize automatic evaluation of data quality in enterprise credit information fusion, and perform data fusion and data adjustment on data quality assessment in enterprise credit information fusion in real time, which is of great significance to data quality assessment in enterprise credit information fusion, can adapt to data quality assessment in enterprise credit information fusion with different standards and data quality assessment requirements in different enterprise credit information fusions, and has a certain universality. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 This is a flow chart of the steps of a data quality assessment method in enterprise credit information fusion according to the present invention. DETAILED DESCRIPTION
[0072] The present invention will be further described below through specific examples. The illustrative examples and descriptions of the present invention are used to explain the present invention but are not intended to limit the present invention.
[0073] The present invention provides a data quality assessment method and system for enterprise credit information fusion, comprising the following steps:
[0074] like Figure 1 As shown, in this embodiment, the following steps are included:
[0075] Collecting preset data and market data of the enterprise and pre-processing the preset data and the market data; the preset data includes enterprise credit data and related data;
[0076] In actual assessments, corporate credit data includes financial statements and audit reports; related data includes industry ratings, historical default records, and legal cases; market data includes corporate product sales data, public information, product industry reports, media opinions, customer feedback, and maintenance and repairs;
[0077] Take XXX medium-sized enterprise as the research object and obtain the preset data and market data for 2023, which includes product A and product B;
[0078] Performing feature selection on the relevant data to obtain first data, and performing reliability analysis on the market data based on the first data to obtain reliable data;
[0079] In actual evaluation, the first data is the relevant data of Product A; reliable data includes enterprise product sales data, enterprise public information, product industry reports, and customer feedback;
[0080] Performing deviation analysis on the reliable data to obtain a difference coefficient, constructing a quality assessment function based on the difference coefficient, the reliable data, and the enterprise credit data, and building a credit fusion data quality assessment model based on the quality assessment function;
[0081] In the actual evaluation, the coefficient of variation was 0.057;
[0082] Optimizing the credit fusion data quality assessment model according to the assessment error, inputting the data to be assessed into the credit fusion data quality assessment model, and outputting the assessment result;
[0083] In the actual assessment, the quality assessment score of the preset data for 2023 was 0.6954.
[0084] In this embodiment, the method of performing evaluation feature selection on the relevant data to obtain first data includes:
[0085] Calculate the contribution of relevant data:
[0086]
[0087] The zth related data is b z , the contribution of the zth related data is The difference between the zth related data and the ath related data is F z,a , the average value of the difference value of the relevant data is The variance explanation rate of the zth related data is ξ z , related data b z The weight is λ z , the number of relevant data is
[0088] Calculate the evaluation value of the relevant data:
[0089]
[0090] The relevant data b z The evaluation value is S(b z ), the a-th related data is b a , the number of relevant data is Related datab z The marginal probability distribution of is p(b z ), related data b a The marginal probability distribution of is p(b a ), related data b z and related data ba The joint probability distribution of z ,b a ), related data b a The contribution of The zth adjustment coefficient is τ z , related data b z and related data b a The minimum function is min(b z ,b a );
[0091] Output the relevant data with an evaluation value greater than 0.4071 as the first data;
[0092] In the actual evaluation, the contribution is 0.794 and the evaluation value is 0.481.
[0093] In this embodiment, the method for performing reliability analysis on the market data based on the first data to obtain reliable data includes:
[0094] Compare and screen market data with market data, and consider data with a difference greater than 0.307 as abnormal data. Calculate the reliability of market data:
[0095]
[0096] The reliability of the jth market data at the sth moment is g j (s), the number of market data at time s is N s , the number of abnormal data at time s is NC s , the jth market data at time s is u j (s), the jth first data at the sth moment is u1 j (s);
[0097] When the reliability is greater than 0.6931, the market data is reliable data; otherwise, it is unreliable data and is eliminated.
[0098] In this embodiment, the method of performing deviation analysis on the reliable data to obtain the coefficient of variation includes:
[0099] Sort the reliable data according to the time series and input the sorted reliable data into the deviation analysis model;
[0100] Calculate the change in reliable data:
[0101]
[0102] The cth reliable data at the sth moment is f c (s), the cth reliable data at the s+1th moment is f c(s+1), the bandwidth parameter is The suppression coefficient is γ, and the number of reliable data is The kernel density estimation function is The norm function is ||·||, and the change in the cth reliable data between the sth moment and the s+1th moment is h c (s,s+1);
[0103] Calculate the sensitive correlation of reliable data:
[0104]
[0105] The sensitive correlation of the cth reliable data is The upper limit of the timing is T s , the average value of the cth reliable data is The control constant is ε;
[0106] Take the reliable data with a sensitivity correlation greater than 0 as the key data and calculate the difference coefficient of the key data:
[0107]
[0108] The cth key data is f c , the coefficient of difference for the dth evaluation is
[0109] In this embodiment, the method for constructing a quality assessment function based on the variance coefficient, reliability data, and enterprise credit data includes:
[0110] Obtain enterprise credit through enterprise credit information and give the quality evaluation function expression:
[0111]
[0112] The coefficient of variation for the dth evaluation is The enterprise credit assessed in the dth time is y d , the quality evaluation function of the d-th evaluation is The jth reliable data is q j , the number of reliable data is M1, and the number of evaluation data is M.
[0113] In this embodiment, the method for constructing a credit fusion data quality assessment model based on the quality assessment function includes:
[0114] According to the quality assessment function, the objective function of the credit fusion data quality assessment model is given as follows:
[0115]
[0116] The objective function of the d-th evaluation is The actual quality assessment value of the d-th evaluation is The loss function is L(·,·), and the prediction quality evaluation value of the d-th evaluation is
[0117] Credit fusion data quality assessment models include time series comparison algorithm, isolation forest algorithm, and machine learning algorithm;
[0118] The time series comparison algorithm analyzes and compares the trends of different time series data and identifies the differences between input data to obtain time-varying features;
[0119] The isolation forest algorithm constructs multiple isolated trees, uses randomly selected features to isolate time-varying features, calculates the average path length in all isolated trees, and marks time-varying features with a small number of isolations and a path length less than the average path length as abnormal to obtain labeled data;
[0120] The machine learning algorithm learns the objective function and the quality assessment pattern of historical data by building a model, and automatically evaluates and labels the input data using the quality assessment pattern and labeled data.
[0121] In this embodiment, the method for optimizing the credit fusion data quality assessment model based on the assessment error includes:
[0122] Introducing the intelligent group optimization algorithm, the evaluation error is used as the fitness function, and the minimum evaluation error is used as the search goal;
[0123] Perform chaotic mapping on the particle population, and the expression is:
[0124]
[0125] The i-th original sequence is x i , the i+1th chaotic mapping sequence is x i+1 , the control parameter is η, and the random number between 0 and 1 is r1;
[0126] Take the particle position with the smallest fitness as the optimal position and calculate the particle position:
[0127]
[0128] The initial optimal position of the particle is The control coefficient is v, and the position of the wth particle in the t+1th iteration is
[0129] Use the inhibition factor to update the particle position and obtain the inhibition position. The expression is:
[0130]
[0131] The suppression factor is χ, and the random number between -1 and 1 is The optimal position of the particle at the tth iteration is The position of the wth particle at the tth iteration is The suppression position of the wth particle in the t+1th iteration is The current number of iterations is t, and the maximum number of iterations is t max ;
[0132] Adaptive weights are used to update the particle positions to obtain the adapted positions. The expression is:
[0133]
[0134] The adaptive weight of the tth iteration is The balance factor is σ, the attenuation coefficient is ρ, and the suppression position of the wth particle in the tth iteration is The normally distributed random number is β, and the optimal position of the t-th iteration is The adapted position of the wth particle in the t+1th iteration is
[0135] Iterate continuously until the evaluation error is minimized, otherwise adjust the particle population and update the adaptive weight.
[0136] Second, a data quality assessment system for enterprise credit information integration includes:
[0137] Data collection module: used to collect the enterprise's preset data and market data, and pre-process the preset data and the market data; the preset data includes enterprise credit data and related data;
[0138] Selection and analysis module: used for performing feature selection on the relevant data to obtain first data, and performing reliability analysis on the market data based on the first data to obtain reliable data;
[0139] Deviation modeling module: used to perform deviation analysis on the reliable data to obtain a difference coefficient, construct a quality assessment function based on the difference coefficient, the reliable data and the enterprise credit data, and build a credit fusion data quality assessment model based on the quality assessment function;
[0140] Optimization output module: used to optimize the credit fusion data quality assessment model according to the assessment error, input the data to be assessed into the credit fusion data quality assessment model, and output the assessment result.
[0141] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A data quality assessment method for enterprise credit information fusion, characterized in that: The following steps are involved: Collecting preset data and market data of the enterprise and pre-processing the preset data and the market data; the preset data includes enterprise credit data and related data; the related data includes industry ratings, historical default records, and legal caseloads; Performing feature selection on the relevant data to obtain first data, and performing reliability analysis on the market data based on the first data to obtain reliable data; Performing deviation analysis on the reliable data to obtain a difference coefficient, constructing a quality assessment function based on the difference coefficient, the reliable data, and the enterprise credit data, and building a credit fusion data quality assessment model based on the quality assessment function; Obtain enterprise credit through enterprise credit information and give the quality evaluation function expression: , The coefficient of variation for the dth evaluation is , the enterprise credit of the dth evaluation is , the quality evaluation function of the d-th evaluation is , the jth reliable data is , the amount of reliable data is , the number of evaluation data is M; According to the quality assessment function, the objective function of the credit fusion data quality assessment model is given as follows: , The objective function of the d-th evaluation is , the actual quality assessment value of the d-th evaluation is , the loss function is , the prediction quality assessment value of the d-th evaluation is ; Credit fusion data quality assessment models include time series comparison algorithm, isolation forest algorithm, and machine learning algorithm; The time series comparison algorithm analyzes and compares the trends of different time series data and identifies the differences between input data to obtain time-varying features; The isolation forest algorithm constructs multiple isolated trees, uses randomly selected features to isolate time-varying features, calculates the average path length in all isolated trees, and marks time-varying features with a small number of isolations and a path length less than the average path length as abnormal to obtain labeled data; The machine learning algorithm learns the objective function and the quality assessment pattern of historical data by building a model, and automatically evaluates and labels the input data using the quality assessment pattern and labeled data; The credit fusion data quality assessment model is optimized according to the assessment error, the data to be assessed is input into the credit fusion data quality assessment model, and an assessment result is output.
2. The data quality assessment method in enterprise credit information fusion according to claim 1 is characterized in that: The method of performing evaluation feature selection on the relevant data to obtain first data includes: Calculate the contribution of relevant data: , The zth related data is , the contribution of the zth related data is , the difference between the zth related data and the ath related data is , the average value of the difference value of the relevant data is , the variance explanation rate of the zth related data is , relevant data The weight is , the number of relevant data is ; Calculate the evaluation value of the relevant data: , The relevant data The evaluation value is , the ath related data is , relevant data The marginal probability distribution of , relevant data The marginal probability distribution of , relevant data and related data The joint probability distribution of , relevant data The contribution of , the zth adjustment coefficient is , relevant data and related data The minimum function of ; The correlation data having an evaluation value greater than 0.4071 is output as the first data.
3. The data quality assessment method in enterprise credit information fusion according to claim 1 is characterized in that: The method for performing reliability analysis on the market data based on the first data to obtain reliable data includes: Compare and screen market data with market data, and consider data with a difference greater than 0.307 as abnormal data. Calculate the reliability of market data: , The reliability of the jth market data at the sth moment is , the amount of market data at time s is , the number of abnormal data at time s is , the jth market data at time s is , the jth first data at the sth moment is ; When the reliability is greater than 0.6931, the market data is reliable data; otherwise, it is unreliable data and is eliminated.
4. The data quality assessment method in enterprise credit information fusion according to claim 1 is characterized in that: The method of performing deviation analysis on the reliable data to obtain a coefficient of variation comprises: Sort the reliable data according to the time series and input the sorted reliable data into the deviation analysis model; Calculate the change in reliable data: , The cth reliable data at the sth moment is , the cth reliable data at time s+1 is , the bandwidth parameter is , the suppression coefficient is , the amount of reliable data is , the kernel density estimation function is , the norm function is , the change in the cth reliable data between the sth moment and the s+1th moment is ; Calculate the sensitive correlation of reliable data: , The sensitive correlation of the cth reliable data is , the upper limit of the timing is , the average value of the cth reliable data is , the control constant is ; Take the reliable data with a sensitivity correlation greater than 0 as the key data and calculate the difference coefficient of the key data: , The cth key data is , the coefficient of difference for the dth evaluation is .
5. The data quality assessment method in enterprise credit information fusion according to claim 1 is characterized in that: The method for optimizing the credit fusion data quality assessment model according to the assessment error includes: Introducing the intelligent group optimization algorithm, the evaluation error is used as the fitness function, and the minimum evaluation error is used as the search goal; Perform chaotic mapping on the particle population, and the expression is: , The i-th original sequence is , the i+1th chaotic mapping sequence is , the control parameters are , a random number from 0 to 1 is ; Take the particle position with the smallest fitness as the optimal position and calculate the particle position: , The initial optimal position of the particle is , the control coefficient is , the position of the wth particle in the t+1th iteration is ; Use the inhibition factor to update the particle position and obtain the inhibition position. The expression is: , , The inhibitory factor is , a random number between -1 and 1 is , the optimal position of the particle in the tth iteration is , the position of the wth particle in the tth iteration is , the suppression position of the wth particle in the t+1th iteration is , the current number of iterations is t, and the maximum number of iterations is ; Adaptive weights are used to update the particle positions to obtain the adapted positions. The expression is: , , The adaptive weight of the tth iteration is , the balance factor is , the attenuation coefficient is , the suppression position of the wth particle in the tth iteration is , the normal distribution random number is , the optimal position of the tth iteration is , the adaptation position of the wth particle in the t+1th iteration is ; Iterate continuously until the evaluation error is minimized, otherwise adjust the particle population and update the adaptive weight.
6. A data quality assessment system for enterprise credit information fusion, used to execute the method according to any one of claims 1 to 5, characterized in that: include: Data collection module: used to collect the enterprise's preset data and market data, and pre-process the preset data and the market data; The preset data includes enterprise credit data and related data; Selection and analysis module: used for performing feature selection on the relevant data to obtain first data, and performing reliability analysis on the market data based on the first data to obtain reliable data; Deviation modeling module: used to perform deviation analysis on the reliable data to obtain a variance coefficient, construct a quality assessment function based on the variance coefficient, the reliable data, and the enterprise credit data, and build a credit fusion data quality assessment model based on the quality assessment function; including: Obtain enterprise credit through enterprise credit information and give the quality evaluation function expression: , The coefficient of variation for the dth evaluation is , the enterprise credit of the dth evaluation is , the quality evaluation function of the d-th evaluation is , the jth reliable data is , the amount of reliable data is , the number of evaluation data is M; According to the quality assessment function, the objective function of the credit fusion data quality assessment model is given as follows: , The objective function of the d-th evaluation is , the actual quality assessment value of the d-th evaluation is , the loss function is , the prediction quality assessment value of the d-th evaluation is ; Credit fusion data quality assessment models include time series comparison algorithm, isolation forest algorithm, and machine learning algorithm; The time series comparison algorithm analyzes and compares the trends of different time series data and identifies the differences between input data to obtain time-varying features; The isolation forest algorithm constructs multiple isolated trees, uses randomly selected features to isolate time-varying features, calculates the average path length in all isolated trees, and marks time-varying features with a small number of isolations and a path length less than the average path length as abnormal to obtain labeled data; The machine learning algorithm learns the objective function and the quality assessment pattern of historical data by building a model, and automatically evaluates and labels the input data using the quality assessment pattern and labeled data; Optimization output module: used to optimize the credit fusion data quality assessment model according to the assessment error, input the data to be assessed into the credit fusion data quality assessment model, and output the assessment result.
Citation Information
Patent Citations
Enterprise credit rating method and system based on financial big data
CN117557281A