Enterprise credit analysis method and system based on big data
Through the enterprise credit analysis method based on big data, the problem of insufficient data integration and deep mining capabilities of traditional systems is solved, and comprehensive and accurate analysis of enterprise credit status and credit evaluation are achieved, and evaluation efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202510343138.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-27
AI Technical Summary
Traditional enterprise credit reporting systems have shortcomings in terms of dispersed data sources, inconsistent formats, difficulty in data integration, and lack of in-depth data mining capabilities, which leads to the inability to accurately identify key corporate credit indicators and potential trends, affecting the accuracy and efficiency of credit assessments.
The enterprise credit analysis method based on big data is adopted to achieve comprehensive analysis and credit evaluation of enterprise credit data through data collection, cleaning, analysis, model construction and visual display. The specific steps include building a large database, using kettle tools to collect and clean data, identifying risk factors, building a credit evaluation model, and visually outputting the results.
It realizes a comprehensive and accurate analysis of the company's credit status, can quickly respond to user credit assessment requests, provide reliable decision-making basis, reduce credit risks, and improve the efficiency and quality of credit assessment.
Smart Images

Figure CN120219065A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of enterprise credit data processing, and specifically relates to a method and system for enterprise credit investigation and analysis based on big data. Background Art
[0002] With the development of enterprises, enterprise credit information has become increasingly important. There are many deficiencies in traditional enterprise credit investigation systems, making it difficult for users to quickly and accurately obtain comprehensive enterprise credit status.
[0003] Firstly, the data sources of current enterprise credit investigation systems are relatively scattered, and the data formats and standards of different data sources vary, making it very difficult to aggregate and integrate data. When comprehensively and deeply analyzing the enterprise credit status, the integrity and relevance of the data are insufficient, and the key indicators and potential trends of enterprise credit cannot be accurately identified. Eventually, it is difficult for users to quickly obtain comprehensive and accurate enterprise credit status. For example, when analyzing the correlation between enterprise business operations and credit scores, due to the inability to effectively integrate enterprise financial data and industry competition data, it is difficult to extract valuable information and provide strong data support for credit assessment.
[0004] Secondly, most existing enterprise credit investigation systems rely on traditional statistical analysis methods and lack the ability to deeply mine massive multi-source data. For example, traditional credit assessment uses the Naive Bayes model, which analyzes independent characteristics, while the characteristics of the data involved in actual enterprise credit investigation are often correlated, resulting in poor classification effects. In addition, existing enterprise credit investigation systems often cannot effectively identify key risk factors when dealing with complex data relationships, thus affecting the accuracy of credit assessment. Summary of the Invention
[0005] In a first aspect, an embodiment of this application provides a method for enterprise business analysis based on privacy computing, including the following steps: S1. Construct a large database according to the data sources of enterprise credit investigation analysis, use the kettle tool to collect the required enterprise credit investigation data from the large database, and perform cleaning and sorting; S2. Analyze the collected enterprise credit investigation data, identify risk factors, evaluate the risk level, and give risk warnings to enterprises with credit risks; S3. Determine the business requirements of enterprise credit investigation analysis, select algorithms from the algorithm library to construct a credit assessment model according to the business requirements, and use the collected enterprise credit investigation data and data analysis results to train the credit assessment model; S4. Respond to the user's credit assessment request for the target enterprise, input the enterprise credit investigation data of the target enterprise into the trained credit assessment model to obtain the credit assessment result; S5. Determine the data display method according to the data types of business requirements and credit assessment results, and visually output the credit assessment results according to the determined data display method.
[0006] Furthermore, the specific steps of step S1 are as follows: S11. Determine the data sources for enterprise credit investigation analysis and apply for access rights to the corresponding data in the data sources; S12. Use the data for which the access right application in the data sources is approved to build a large database; S13. Build a relational target database, use the kettle tool to create a scheduled data update task, and regularly export the required enterprise credit investigation data from the large database and import it into the relational target database; S14. Clean the enterprise credit investigation data to remove duplicate, invalid, and incorrect data; S15. Classify and organize the enterprise credit investigation data according to the target fields and then fill it into the database tables of the corresponding categories, and establish associations between different database tables.
[0007] Furthermore, the specific steps of step S15 are as follows: S151. Determine the categories of enterprise credit investigation data; S152. Determine the target fields of the enterprise credit investigation data in each category; S153. Unify the attributes of the enterprise credit investigation data in the relational target database according to the field names, data types, and lengths of the target fields, and process the missing or abnormal data; S154. Fill the processed enterprise credit investigation data into the database tables of the corresponding categories; S155. Establish associations between different database tables through the primary and foreign key constraints of the database tables; S156. Determine the fields that need to be format-converted in each database table and perform the conversion according to the preset rules.
[0008] Furthermore, the specific steps of step S2 are as follows: S21. Summarize the basic characteristics of the enterprise credit investigation data that has been cleaned and sorted, and complete the descriptive statistical analysis; S22. Select the independent variable fields and dependent variable fields from the target fields, fit the corresponding data of the selected fields through a pre-constructed regression model to obtain inferential characteristics, and analyze the relationship between the two fields through the inferential characteristic values; S23. Compare the enterprise credit investigation data with the preset index thresholds, and cluster according to financial information, business information, and credit history to obtain different risk levels; S24. Select decision fields from the enterprise credit investigation data, screen out field combinations from the target fields according to support and confidence, and extract association rules.
[0009] Further, the specific steps of step S23 are as follows: S231. Compare the financial information and business information in the enterprise credit investigation data with the preset index thresholds, and conduct risk warnings when the thresholds are exceeded. S232. Cluster the enterprise credit investigation data according to financial information and credit history to obtain different risk levels, and fill the risk levels into the target fields of the corresponding database tables.
[0010] Further, the specific steps of step S24 are as follows: S241. Use the risk level as the decision field, and screen out the fields relevant to the decision field from the enterprise credit investigation data according to data types and business logics to construct a candidate item set. S242. Count the number of records containing each candidate item set in the enterprise credit investigation data as the number of transactions of the candidate item set. S243. Calculate the support of each candidate item set according to the number of transactions of each candidate item set. S244. Determine the minimum support threshold according to the business requirements of enterprise credit investigation analysis, and take the candidate item sets with support greater than the minimum support threshold as frequent item sets. S245. Determine the association rule set according to the frequent item sets and calculate the confidence of each association rule. S246. Determine the minimum confidence threshold according to the business requirements of enterprise credit investigation analysis, and take the management rules with confidence greater than the minimum confidence threshold as effective association rules.
[0011] Further, the specific steps of step S3 are as follows: S31. Construct a credit assessment algorithm library, which includes a logistic regression algorithm, a decision tree algorithm, a random forest algorithm, and a neural network algorithm. S32. Determine the business requirements of enterprise credit investigation analysis. When the business requirement is to explain the credit assessment results and the data dimension is lower than the dimension threshold, select the logistic regression algorithm or the decision tree algorithm from the credit assessment algorithm library. When the business requirement has an accuracy requirement for the model, the data dimension is higher than the dimension threshold, and there is a non-linearity between the data, select the random forest algorithm or the neural network algorithm from the credit assessment algorithm library. S33. Use the selected algorithm to construct a credit assessment model. S34. Use the enterprise credit investigation data that has been cleaned, sorted, and analyzed as training data to train the constructed credit assessment model.
[0012] Furthermore, the specific steps of step S4 are as follows: S41. Receive a credit assessment request from the user for the target enterprise, parse the user identity and relevant identifiers of the target enterprise, and verify the credit assessment permission according to the user identity; Exemplarily, for financial institution users, it is necessary to verify their institutional qualifications and user permissions; for ordinary investor users, it is necessary to confirm their registration information and scope of usage permissions; S42. Query and obtain the corresponding enterprise credit investigation data from the relational target database according to the relevant identifiers of the target enterprise input by the user; S43. Organize and preprocess the obtained enterprise credit investigation data of the target enterprise according to the format and feature order required by the model, and then input it into the trained credit assessment model.
[0013] Furthermore, the specific steps of step S5 are as follows: S51. Determine the content displayed by the credit assessment result according to the business requirements; S52. Judge the data type of the credit assessment result, and determine the data display method according to the business requirements and the data type; S53. On the user interface of the enterprise credit investigation visualization platform, visually present the credit assessment result according to the determined data display method.
[0014] In a second aspect, an embodiment of the present application further provides a big data-based enterprise credit investigation analysis system, including: A data collection module, which is used to construct a big database according to the data sources of enterprise credit investigation analysis, collect the required enterprise credit investigation data from the big database using the kettle tool, and perform cleaning and sorting; A data analysis module, which is used to analyze the collected enterprise credit investigation data, identify risk factors, evaluate the risk level, and give risk warnings to enterprises with credit risks; A credit assessment model construction module, which is used to determine the business requirements of enterprise credit investigation analysis, select algorithms from the algorithm library according to the business requirements to construct a credit assessment model, and train the credit assessment model using the collected enterprise credit investigation data and data analysis results; A credit assessment module, which is used to respond to the user's credit assessment request for the target enterprise, and input the enterprise credit investigation data of the target enterprise into the trained credit assessment model to obtain a credit assessment result; An enterprise credit investigation analysis display module, which is used to determine the data display method according to the business requirements and the data type of the credit assessment result, and visually output the credit assessment result according to the determined data display method.
[0015] It can be seen from the above technical solutions that the present application has the following advantages: The enterprise credit investigation analysis method and system based on big data provided by this application can achieve a comprehensive analysis of enterprise credit investigation data through data collection, cleaning, analysis, model construction, and visual display. It can accurately evaluate the credit status of enterprises, provide reliable decision-making basis for financial institutions, investors, etc., and effectively reduce credit risks. Data collection and regular updates are carried out through the kettle tool, combined with data cleaning, sorting, and classified storage to ensure the accuracy, integrity, and timeliness of data, and improve the efficiency and quality of enterprise credit investigation analysis. By constructing a credit evaluation algorithm library containing a variety of classic algorithms and flexibly selecting algorithms to construct models according to business needs, it can adapt to the enterprise credit investigation analysis needs in different scenarios and provide appropriate solutions for both the requirements of model interpretability and accuracy. Determine the data display method according to business needs and data types, and intuitively present the credit evaluation results through a visualization platform, which is convenient for users to quickly understand the enterprise credit status, improve decision-making efficiency, and at the same time support real-time updates to ensure that users obtain the latest information. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of this application, the drawings required for description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0017] Figure 1 FIG. is a schematic flow chart of an enterprise credit investigation analysis method based on big data.
[0018] Figure 2 FIG. is a schematic diagram of an enterprise credit investigation analysis system based on big data. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] In the following, the specific steps of the enterprise credit investigation analysis method based on big data will be described in detail, and various embodiments of the present disclosure will be described more comprehensively. The present disclosure can have various embodiments, and adjustments and changes can be made therein. However, it should be understood that there is no intention to limit the various embodiments of the present disclosure to the specific embodiments disclosed herein, but the present disclosure should be understood to cover all adjustments, equivalents, and / or alternative solutions that fall within the spirit and scope of the various embodiments of the present disclosure.
[0020] Exemplarily, with the development of enterprises, the importance of enterprise credit information has become increasingly prominent. However, traditional enterprise credit investigation systems have limitations in many aspects, making it difficult for users to quickly and accurately grasp the comprehensive credit status of enterprises. The primary problem is that the data sources of current enterprise credit investigation systems are highly dispersed, and the data formats and standards among different data sources lack uniformity. This dispersion and difference pose great difficulties in the process of data aggregation and integration, resulting in the inability to guarantee the integrity and relevance of data when conducting in-depth and comprehensive analysis of enterprise credit status. Therefore, key credit indicators and potential trends are difficult to be accurately identified, and users are thus unable to quickly obtain comprehensive and accurate enterprise credit status. For example, when attempting to analyze the correlation between enterprise operating conditions and credit scores, valuable information is difficult to be mined due to the inability to effectively integrate the financial data and industry competition data of the enterprise, and thus a solid data foundation cannot be provided for credit assessment.
[0021] In addition, most existing enterprise credit investigation systems rely on traditional statistical analysis methods and lack the ability to deeply mine massive and diverse data sources. Taking traditional credit assessment methods as an example, such as the Naive Bayes model, this method mainly conducts analysis based on independent characteristics. However, in actual enterprise credit investigation scenarios, there are often complex correlations between data characteristics, which leads to unsatisfactory classification effects of the Naive Bayes model. Additionally, when facing complex data relationships, existing enterprise credit investigation systems often have difficulty effectively identifying key risk factors, which further affects the accuracy of credit assessment.
[0022] In response to the above problems, this embodiment provides a big data-based enterprise credit investigation analysis method, which realizes efficient, accurate, and dynamic enterprise credit assessment through data collection, analysis, model construction, evaluation, and visualization display.
[0023] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0024] Please refer to Figure 1 The flowchart of the big data-based enterprise credit investigation analysis method in a specific embodiment is shown. The method includes the following steps: S1. Construct a big database according to the data sources of enterprise credit investigation analysis, use the kettle tool to collect the required enterprise credit investigation data from the big database, and perform cleaning and sorting; It should be noted that by determining the data source to build a large database, relevant data on enterprise credit investigation can be comprehensively collected, providing a data basis for data analysis; using the kettle tool to collect and clean data can ensure data quality and reduce analysis deviations caused by data problems; classifying and organizing the data, querying, calling, and managing the data can improve the data usage efficiency and facilitate data analysis and model training; S2. Analyze the collected enterprise credit investigation data, identify risk factors, evaluate the risk level, and issue risk warnings for enterprises with credit risks; It should be noted that analyzing the collected enterprise credit investigation data, identifying risk factors, and evaluating the risk level provide a basis for risk warnings; issuing risk warnings for enterprises with credit risks enables financial institutions and investor users to take timely measures to avoid risks; S3. Determine the business requirements for enterprise credit investigation analysis, select algorithms from the algorithm library according to the business requirements to build a credit assessment model, and use the collected enterprise credit investigation data and data analysis results to train the credit assessment model; It should be noted that selecting algorithms from the algorithm library according to the business requirements to build a credit assessment model makes the model more in line with the actual application scenario; using the collected enterprise credit investigation data and data analysis results to train the credit assessment model improves the accuracy and reliability of the model; S4. Respond to the user's credit assessment request for the target enterprise, input the enterprise credit investigation data of the target enterprise into the trained credit assessment model to obtain the credit assessment result; It should be noted that by verifying the user's identity, the legality of the credit assessment request and the security of the data are ensured, preventing unauthorized access and data leakage; quickly organizing and preprocessing the data of the target enterprise and inputting it into the model for evaluation improves the system's response speed and user experience; S5. Determine the data display method according to the business requirements and the data type of the credit assessment result, and visually output the credit assessment result according to the determined data display method; It should be noted that presenting the complex credit assessment result in an intuitive form reduces the difficulty for users to understand and apply the assessment result, improving the user experience; determining the display method according to the business requirements and data type can meet the needs of different users, enhancing the flexibility and adaptability of the system; through visual display, users can quickly obtain key information to assist them in making decisions.
[0025] This embodiment can automatically collect, clean, organize, and analyze a large amount of enterprise credit investigation data, identify risk factors, evaluate the risk level, and issue risk warnings for enterprises with credit risks. At the same time, it can flexibly select algorithms according to business needs to build a credit assessment model, improving the accuracy and reliability of credit assessment. Through a visual display method, it presents complex credit assessment results to users in an intuitive and easy-to-understand form, facilitating users to quickly understand the enterprise credit status and assisting users in decision-making.
[0026] Further, as a refinement and extension of the specific implementation manner of the above embodiment, in order to fully illustrate the specific implementation process in this embodiment, another enterprise credit investigation analysis method based on big data is provided. This method includes the following steps: S1. Construct a big database according to the data sources for enterprise credit investigation analysis, use the kettle tool to collect the required enterprise credit investigation data from the big database, and perform cleaning and organization. The specific steps of step S1 are as follows: S11. Determine the data sources for enterprise credit investigation analysis and apply for access rights to the corresponding data in the data sources. Exemplarily, the data sources for enterprise credit investigation analysis include authoritative credit investigation agencies, enterprise credit publicity platforms, banks, and third-party credit investigation agencies. Authoritative credit investigation agencies are responsible for collecting, organizing, and managing enterprise credit investigation information, and credit investigation can be carried out through authoritative websites or offline agencies. Enterprise credit publicity platforms provide enterprise registration information and information related to business operations. Banks provide credit investigation services for enterprises with which they have business dealings. Third-party credit investigation agencies can provide credit investigation report services externally in a paid manner, such as industrial and commercial registration information, business operations, and judicial risks. S12. Construct a big database using the data for which the access right application in the data source has passed. S13. Construct a relational target database, use the kettle tool to create a timed data update task, and regularly export the required enterprise credit investigation data from the big database and import it into the relational target database. Exemplarily, the required enterprise credit investigation data includes the basic information, financial status, credit records, and legal litigation records of the enterprise. S14. Clean the enterprise credit investigation data to remove duplicate, invalid, and incorrect data. S15. Classify and organize the enterprise credit investigation data according to the target fields and fill them into the corresponding database tables, and establish associations between different database tables. S2. Analyze the collected enterprise credit investigation data, identify risk factors, evaluate the risk level, and issue risk warnings for enterprises with credit risks. The specific steps of step S2 are as follows: S21. Summarize the basic characteristics of the enterprise credit investigation data that has been cleaned and sorted, and complete the descriptive statistical analysis; Exemplarily, calculate and statistically analyze the average registered capital, average operating income, and average net profit of the enterprise as basic characteristics, and calculate the standard deviation or coefficient of variation index of the basic characteristics to evaluate the degree of dispersion and volatility of the basic characteristics; S22. Select the independent variable field and the dependent variable field from the target fields, fit the corresponding data of the selected fields through a pre-constructed regression model to obtain inferential characteristics, and analyze the relationship between the two fields through the inferential characteristic values; It should be noted that the inferential characteristics include regression coefficient, standard error, t-value, and p-value; Exemplarily, if the selected independent variable field is the registered capital and the selected dependent variable field is the operating income, analyze the relationship between the registered capital and the operating income through a linear regression model; Taking the sample data of 100 enterprises as an example, each enterprise has two variables, registered capital and operating income. Use a linear regression model to analyze the relationship between these two variables; through regression analysis, we get the following results: Regression coefficient (β1): It means that for every 1 unit increase in the registered capital, the operating income increases by an average of β1 units; Standard error: It represents the estimation error of the regression coefficient; t-value: It is used to test whether the regression coefficient is significantly different from 0; p-value: It represents the probability of observing the current data or more extreme data under the condition that the null hypothesis (that is, the regression coefficient is equal to 0) is true; if the p-value is less than a certain significance level, such as 0.05, then reject the null hypothesis and consider that the registered capital has a significant impact on the operating income; If the regression coefficient (β1) is 0.5, the standard error is 0.1, the t-value is 5, and the p-value is less than 0.05, it means that for every 1 unit increase in the registered capital, the operating income increases by an average of 0.5 units, and this impact is statistically significant; S23. Compare the enterprise credit investigation data with the preset index thresholds, and cluster according to financial information, operating information, and credit history to obtain different risk levels; S24. Select the decision fields in the enterprise credit investigation data, and screen out the field combinations from the target fields according to the support and confidence, and extract the association rules; S3. Determine the business requirements of the enterprise credit investigation analysis, select an algorithm from the algorithm library according to the business requirements to build a credit assessment model, and use the collected enterprise credit investigation data and data analysis results to train the credit assessment model; The specific steps of step S3 are as follows: S31. Construct a credit assessment algorithm library, which includes a logistic regression algorithm, a decision tree algorithm, a random forest algorithm, and a neural network algorithm; It should be noted that the logistic regression algorithm is based on a linear regression model. By mapping the results of linear regression to a probability value, it is used to handle binary classification problems. Specifically, in enterprise credit investigation, it is used to predict whether an enterprise will default. By constructing a linear equation, multiple input features are weighted and summed, and then the result is transformed into a probability value between 0 and 1 through a logistic function to judge the possibility of an enterprise being in a certain credit state; The decision tree algorithm divides data in a tree-like structure. Each internal node is a test on an attribute, the branches are the test outputs, and the leaf nodes are the categories or values. Specifically, in enterprise credit investigation, according to the various financial indicators and credit record information of an enterprise, decision rules are constructed to show the influence paths of different factors on the credit assessment results. Exemplarily, by dividing the debt ratio and net profit rate indicators of an enterprise multiple times, several decision rules are formed to judge the credit rating of the enterprise; The random forest algorithm is based on the decision tree algorithm. By constructing multiple decision trees and comprehensively voting or averaging the results of these decision trees, the accuracy of the model is improved. Specifically, when dealing with high-dimensional enterprise credit investigation data, it automatically performs feature selection to reduce the risk of overfitting. At the same time, by integrating the results of multiple decision trees, the adaptability of the model is enhanced; The neural network algorithm, such as the multi-layer perceptron MLP and convolutional neural network CNN in the deep learning model, automatically learns the complex patterns and relationships in the data through a large number of neuron nodes. Specifically, in enterprise credit investigation analysis, it processes a large amount of high-dimensional and non-linear data to deeply explore the impact of multi-source data of enterprises on credit assessment; S32. Determine the business requirements for enterprise credit investigation analysis; When the business requirement is to explain the credit assessment results and the data dimension is lower than the dimension threshold, select the logistic regression algorithm or the decision tree algorithm from the credit assessment algorithm library; When the business requirement has an accuracy requirement for the model, the data dimension is higher than the dimension threshold, and there is non-linearity between the data, select the random forest algorithm or the neural network algorithm from the credit assessment algorithm library; S33. Use the selected algorithm to construct a credit assessment model; Exemplarily, if the logistic regression algorithm is selected to construct a credit assessment model, the input features need to be determined first: the basic information of the enterprise, such as the registered capital and the number of years of establishment, the financial status, such as the asset-liability ratio and the net profit growth rate, and the credit record, such as the number of overdue times and the default history; then preprocess these input feature data, such as standardization processing, to make different features have the same scale and avoid affecting model training due to feature scale differences; then use the training data set to train the logistic regression model, and through the iterative optimization algorithm, adjust the parameters of the model, such as the weight coefficient, to minimize the prediction error of the model on the training data. If the decision tree algorithm is selected, first perform feature selection on the enterprise credit investigation data to determine the most critical features for credit assessment; then, according to the selected features, recursively partition the data according to certain splitting criteria, such as information gain and Gini coefficient, to construct a decision tree model; and during the construction process, avoid overfitting caused by an overly deep decision tree by setting control parameters such as the maximum depth and minimum number of samples of the tree. If the random forest algorithm is selected, first generate multiple sub-datasets based on the training data set by sampling with replacement; then construct decision tree models on each sub-dataset separately, and the feature selection and partitioning processes of each decision tree model are independent of each other; finally, by synthesizing the prediction results of multiple decision trees, specifically, using the voting method for classification problems and the averaging method for regression problems, obtain the final credit assessment result. If the neural network algorithm is selected, first design the network structure and determine the number of neuron layers and the number of neurons in each layer; then preprocess the enterprise credit investigation data, such as normalization processing; then train the neural network with the training data set, and continuously adjust the connection weights between neurons using the backpropagation algorithm during the training process to minimize the loss function of the model on the training data; during the training process, use regularization methods to avoid overfitting. S34. Use the enterprise credit investigation data that has been cleaned, sorted, and analyzed as training data to train the constructed credit assessment model. It should be noted that during the training process, continuously adjust the parameters of the model to improve the prediction accuracy of the credit assessment model for the enterprise credit status; exemplarily, for the logistic regression model, adjust the weight coefficient through the optimization method of the gradient descent algorithm to make the prediction of the credit assessment model for the credit status of enterprises in the training data, such as whether they default, as close as possible to the actual situation. Optimize the training process of the enterprise credit assessment model by using the results obtained from the analysis of enterprise credit investigation data, such as the basic characteristics of enterprises obtained from descriptive statistical analysis, the variable relationships obtained from inferential statistical analysis, and the risk association rules obtained from association rule mining; Exemplarily, if it is found in association rule mining that the default risk of an enterprise increases significantly when low net profit and high debt ratio appear simultaneously, this association rule can be incorporated into the model training process as prior knowledge to improve the model's recognition ability for such high-risk enterprises; Adopt a cross-validation method to verify the performance of the credit assessment model during the training process to avoid overfitting; Exemplarily, divide the training data into multiple subsets, each time using a part of the subsets as the training set and the other part as the validation set, and through multiple iterations, verify the performance of the credit assessment model on different subsets; S4. Respond to the user's credit assessment request for the target enterprise, and input the enterprise credit investigation data of the target enterprise into the trained credit assessment model to obtain the credit assessment result; The specific steps of step S4 are as follows: S41. Receive the user's credit assessment request for the target enterprise, parse the user identity and the relevant identifiers of the target enterprise, and verify the credit assessment permission according to the user identity; Exemplarily, for financial institution users, it is necessary to verify their institutional qualifications and user permissions; for ordinary investor users, it is necessary to confirm their registration information and the scope of usage permissions; S42. Query and obtain the corresponding enterprise credit investigation data from the relational target database according to the relevant identifiers of the target enterprise input by the user; S43. Organize and preprocess the obtained enterprise credit investigation data of the target enterprise according to the format and feature order required by the model, and then input it into the trained credit assessment model; Exemplarily, if the credit assessment model is a logistic regression model, then according to the input enterprise feature data, through linear combination and logistic function calculation, output the probability value of the target enterprise defaulting; for example, if the output result is 0.2, it means that the enterprise has a 20% possibility of defaulting; If the credit assessment model is a decision tree model, then according to the path of the data on the decision tree, finally determine the credit rating of the target enterprise; for example, through the decision rules of the decision tree, it is judged that the credit rating of the enterprise is B level; If the credit assessment model is a random forest model, then through the voting or average results of multiple decision trees, give the credit assessment result of the target enterprise; for example, after multiple decision trees vote, it is concluded that the enterprise belongs to a medium-risk enterprise; If the credit assessment model is a neural network model, the output is a numerical value or vector representing the scores or probability distributions of the target enterprise in different credit dimensions. For example, an output of a three-dimensional vector [0.1, 0.7, 0.2] represents the probabilities of the enterprise belonging to low risk, medium risk, and high risk respectively. S5. Determine the data display method according to the business requirements and the data type of the credit assessment result, and visually output the credit assessment result according to the determined data display method. The specific steps of step S5 are as follows: S51. Determine the content of the credit assessment result to be displayed according to the business requirements. Exemplarily, when conducting loan approval or investment decisions, financial institutions are more concerned about the enterprise's credit risk level, default probability, and debt repayment ability-related indicators, such as the asset-liability ratio, current ratio, and loan records, in order to accurately assess the risks of lending or investing. Enterprises themselves are concerned about the changing trends of indicators related to their operating conditions, such as changes in operating income and net profit, business cooperation situations, and public opinion information. Regulatory authorities need to pay attention to the compliance information of enterprises, such as judicial information and administrative penalty records. Specifically, based on these different business requirements, determine the content of the credit assessment result that needs to be highlighted for display. S52. Judge the data type of the credit assessment result, and determine the data display method according to the business requirements and the data type. The data type includes numerical data, categorical data, text data, and chart data. Exemplarily, numerical data includes enterprise credit scores, default probabilities, and numerical values of various financial indicators, such as total assets and net profit. Categorical data, such as credit risk levels, may specifically include high risk, medium risk, low risk, and the industry classification of the enterprise. Text data includes public opinion information, case descriptions in judicial information, etc. Chart data, such as credit change trend charts, business cooperation relationship charts, etc. It should be noted that for numerical data: for enterprise credit scores and default probabilities, they are displayed in a form combining numbers and progress bars, and the scoring levels are marked at the same time. When the business requirement is to compare the credit scores of different enterprises or view the changes in credit scores over time, bar charts or line charts can be used for display. For example, use a bar chart to display the credit scores of different enterprises to visually compare the credit differences between enterprises; use a line chart to display the change trend of an enterprise's credit score in the past few years to help users understand the dynamic changes in the enterprise's credit status. For financial indicator values, bar charts and line charts are used to display the changing trends that need to be presented. For example, when presenting the changes in a company's operating income and net profit in recent years, bar charts or line charts are adopted; when comparing the indicator values of different companies, bar charts are used. For example, in industry comparative analysis, bar charts are used to show the comparison of financial credit scores and risk level indicators between a company and other companies in the same industry. For categorical data: For example, credit ratings, pie charts and bar charts are used to show the distribution of companies with different credit ratings. Specifically, a pie chart is used to show the proportions of companies with high, medium, and low credit ratings in the entire sample, enabling users to quickly understand the overall distribution of credit ratings. For text data: It is presented in the form of text boxes and lists. For example, the legal litigation records, administrative penalty records, etc. of a company are presented in list form, and the specific content of each record is shown in detail. For chart data, the corresponding charts are directly embedded in the platform interface. For example, a relationship chart that can show the correlation between various financial indicators of a company is presented, enabling users to understand the company's financial status and the sources of credit risks. When the business requirements need to comprehensively display various types of data, a multi-dimensional data visualization method is adopted. For example, a dashboard is used to comprehensively display a company's credit score, risk level, and key financial indicator information, enabling users to comprehensively understand the company's credit status on one interface. S53. On the user interface of the enterprise credit investigation visualization platform, the credit assessment results are visually presented according to the determined data display method. Specifically, visualization tools such as Tableau or PowerBI are used to convert the data into intuitive and easy-to-understand graphics and images. For example, a bar chart is created through Tableau to show the comparison of credit scores between a target company and other companies in the same industry, with different colors set to distinguish different companies, and elements such as titles and axis labels are added. The layout of the visualization interface is designed, and the positions of various charts and information are reasonably arranged to ensure that users can quickly find the information they need. For example, the basic information of the company is placed at the top of the page, and the key indicators of the credit assessment results, such as credit scores and risk levels, are prominently shown in the center of the page, and detailed data analysis charts and text information are classified and shown below. Interactive functions are added to the visualization results, such as showing detailed information when the mouse hovers and data drilling when a chart is clicked. For example, when the user's mouse hovers over the credit score bar chart, the specific credit score value of the company and the calculation basis of the score are shown; when the user clicks on the line chart of a certain financial indicator, detailed data for different time periods can be drilled down. Perform real-time updates on the results of the visual output. When the enterprise's credit investigation data changes or the credit assessment model is updated, promptly refresh the visual interface to ensure that users obtain the latest credit assessment information.
[0027] In an embodiment of the present invention, based on step S14, the following will give a possible embodiment to non-restrictively elaborate on its specific implementation.
[0028] The specific steps of step S14 are as follows: S141. Identify duplicate records in the enterprise credit investigation data and delete the duplicate records from the relational target database; Exemplarily, identify whether multiple identical loan records of the same enterprise within the same time period are duplicate records by comparing the keyword fields of the loan amount, loan term, and repayment method, and only retain one record for the identified duplicate records; S142. Correct invalid data in the enterprise credit investigation data; Exemplarily, by comparing the records of banks or financial institutions and the historical data of credit investigation agencies, confirm that the credit card records in the enterprise credit investigation data have expired, such as the credit card has been cancelled or has expired, mark these credit card records as invalid data and delete them, and at the same time update the subsequent credit rating results; S143. Correct incorrect data in the enterprise credit investigation data; Exemplarily, by comparing the records of banks or financial institutions and the information provided by the enterprise, after confirming that the debt amount in the debt record is incorrect, correct these debt records and at the same time update the subsequent credit rating results.
[0029] In an embodiment of the present invention, based on step S15, the following will give a possible embodiment to non-restrictively elaborate on its specific implementation.
[0030] The specific steps of step S15 are as follows: S151. Determine the category of the enterprise credit investigation data; Exemplarily, the category includes basic information, credit information, non-credit information, public information, and risk assessment; S152. Determine the target fields of the enterprise credit investigation data in each category; Exemplarily, the target fields of basic information: enterprise name, enterprise code, registered address, business address, establishment date, enterprise type, and industry classification; The target fields of credit information: loan records, guarantee records, credit card records; The target fields of non-credit information: financial information, business information, payment records, court judgment records, administrative penalty records; Target fields of public information: enterprise qualifications, tax information, and environmental protection information; Target fields of risk assessment: credit score, risk level, and early warning prompt; It should be noted that there is no data in the target fields of risk assessment, which can be used for filling in data after subsequent enterprise credit investigation analysis and assessment; S153. Unify the attributes of enterprise credit investigation data in the relational target database according to the field name, data type, and length of the target fields, and process missing or abnormal data; S154. Fill the processed enterprise credit investigation data into the database tables of corresponding categories; S155. Establish associations between different database tables through the primary and foreign key constraints of the database tables; Exemplarily, loan records, guarantee records, enterprise names, and enterprise codes are usually the associated fields between different database tables; S156. Determine the fields that need to be format-converted in each database table and perform the conversion according to the preset rules; Exemplarily, convert the date format and currency unit according to the determined target format to unify the format.
[0031] In an embodiment of the present invention, based on step S23, a possible embodiment will be given below to non-restrictively elaborate on its specific implementation scheme.
[0032] The specific steps of step S23 are as follows: S231. Compare the financial information and business information in the enterprise credit investigation data with the preset index thresholds, and issue a risk warning when the index thresholds are exceeded; Exemplarily, set warning indicators according to the business characteristics and risk management requirements of the enterprise; Specific warning indicators include financial indicators, such as asset-liability ratio, current ratio, quick ratio, business indicators such as sales growth rate, market share, inventory turnover rate, and credit indicators, such as credit score, overdue rate; Set a reasonable threshold or range for each warning indicator; When the index value in the enterprise credit investigation data exceeds the set threshold or range, a warning is triggered; Through real-time monitoring and calculation of the enterprise credit investigation data, abnormal data or potential risks are timely discovered, and a warning message is generated and pushed to the enterprise; the warning message includes warning type, warning level, warning reason, and recommended measures; S232. Cluster the enterprise credit investigation data according to financial information and credit history to obtain different risk levels, and fill the risk levels into the target fields of the corresponding database tables; Exemplarily, according to the financial information and credit history of enterprises and through clustering, enterprises are classified into three categories: high-risk, medium-risk, and low-risk; It should be noted that through cluster analysis, enterprises with similar credit status are grouped into one category; the classification results of risk levels can provide financial institutions with differentiated risk management strategies.
[0033] In an embodiment of the present invention, based on step S24, a possible embodiment will be given below to non-restrictively elaborate on its specific implementation scheme.
[0034] The specific steps of step S24 are as follows: S241. Using the risk level as the decision field, relevant fields related to the decision field are selected from the enterprise credit investigation data according to the data type and business logic to construct a candidate item set; Exemplarily, taking the credit investigation data of 50 enterprises as an example, the enterprise credit investigation data includes the net profit, debt ratio, and default risk of the enterprise, and the high and low of the default risk are used as the decision field; According to the business logic, the financial condition of the enterprise is closely related to the default risk, so the net profit and debt ratio are selected as the fields related to the decision field; It should be noted that the net profit and debt ratio are classified. The net profit lower than 30% of the industry average level is defined as low net profit, and the debt ratio higher than 50% of the industry average level is defined as high debt ratio; The candidate item set constructed thereby is as follows: Single-item candidate item set: {low net profit}, {high debt ratio}, {high default risk}, {low default risk} Multi-item candidate item set: {low net profit, high debt ratio}, {low net profit, high default risk}, {high debt ratio, high default risk}, {low net profit, high debt ratio, high default risk}; S242. Count the number of records in the enterprise credit investigation data that contain each candidate item set as the number of transactions of the candidate item set; Exemplarily, through traversing and counting the credit investigation data of 50 enterprises, the number of transactions of each candidate item set is obtained as follows: Candidate item set {low net profit}: After screening, it is found that 15 enterprises have a net profit lower than 30% of the industry average level, so the number of transactions of this candidate item set is 15; Candidate item set {high debt ratio}: 12 enterprises have a debt ratio higher than 50% of the industry average level, and the number of transactions is 12; Candidate item set {high default risk}: 8 enterprises have defaulted, and the number of transactions is 8; Candidate itemset {low net profit, high debt ratio}: There are 6 enterprises that simultaneously meet the conditions of having a net profit lower than 30% of the industry average and a debt ratio higher than 50% of the industry average, and the number of transactions is 6; Candidate itemset {low net profit, high default risk}: There are 4 enterprises that belong to both low net profit and high default risk, and the number of transactions is 4; Candidate itemset {high debt ratio, high default risk}: There are 3 enterprises that simultaneously have a high debt ratio and default situations, and the number of transactions is 3; Candidate itemset {low net profit, high debt ratio, high default risk}: There are 2 enterprises that simultaneously meet the conditions of low net profit, high debt ratio and high default risk, and the number of transactions is 2; S243. Calculate the support of each candidate itemset based on the number of transactions of each candidate itemset; It should be noted that the calculation formula for support is:
[0035] Among them, represents the transaction of candidate set X, N represents the total number of transactions, represents the support of candidate itemset X; Exemplarily, taking the total number of the above enterprises as 50, that is, the total number of transactions is 50 as an example; According to the number of transactions counted above, calculate the support of each candidate itemset as follows:
[0036]
[0037]
[0038]
[0039]
[0040]
[0041]
[0042] S244. Determine the minimum support threshold according to the business requirements of enterprise credit investigation analysis, and take the candidate itemset with a support greater than the minimum support threshold as the frequent itemset; Exemplarily, according to the business requirements of enterprise credit investigation analysis, determine the minimum support threshold to be 0.1; Compare the support of each candidate itemset with the minimum support threshold: The support of candidate itemset {low net profit} is 0.3, which is greater than 0.1, so {low net profit} is a frequent itemset; The support of the candidate item set {high debt-to-asset ratio} is 0.24, which is greater than 0.1. Therefore, {high debt-to-asset ratio} is a frequent item set; The support of the candidate item set {low net profit, high debt-to-asset ratio} is 0.12, which is greater than 0.1. Therefore, {low net profit, high debt-to-asset ratio} is a frequent item set; However, the supports of the candidate item sets {high default risk}, {low net profit, high default risk}, {high debt-to-asset ratio, high default risk}, and {low net profit, high debt-to-asset ratio, high default risk} are all less than 0.1 and they are not frequent item sets; S245. Determine the association rule set based on the frequent item sets and calculate the confidence of each association rule; Exemplarily, association rules are generated from the above frequent item sets, mainly generating rules related to low net profit, high debt-to-asset ratio, and high default risk. The following are the association rules: {low net profit, high debt-to-asset ratio} → {high default risk} {low net profit} → {high debt-to-asset ratio} {high debt-to-asset ratio} → {low net profit} It should be noted that the formula for calculating confidence is:
[0043] Thus, for the association rule {low net profit, high debt-to-asset ratio} → {high default risk}: Since it is known that: ,
[0044] The confidence of this association rule is: ; For the association rule {low net profit} → {high debt-to-asset ratio}: Since it is known that: ,
[0045] The confidence of this association rule is: ; For the association rule {high debt-to-asset ratio} → {low net profit}: Since it is known that: ,
[0046] The confidence of this association rule is: ; S246. Determine the minimum confidence threshold according to the business requirements of enterprise credit investigation analysis, and regard the management rules with confidence greater than the minimum confidence threshold as valid association rules; Exemplarily, if the minimum confidence threshold is determined to be 0.3 according to the business requirements of enterprise credit investigation analysis; Compare the confidence of each association rule with the minimum confidence threshold: The confidence of the association rule {low net profit, high debt ratio} → {default} is about 0.33, which is greater than 0.3, so it is a valid association rule; The confidence of the association rule {low net profit} → {high debt ratio} is 0.4, which is greater than the minimum confidence threshold, so it is a valid association rule; The confidence of the association rule {high debt ratio} → {low net profit} is 0.5, which is greater than 0.3, so it is a valid association rule.
[0047] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0048] Such as Figure 2 As shown, the following is an embodiment of an enterprise credit investigation analysis system based on big data provided by the embodiments of the present disclosure. This system and the enterprise credit investigation analysis method based on big data in the above embodiments belong to the same inventive concept. For the details not described in detail in the embodiment of the enterprise credit investigation analysis system based on big data, reference can be made to the embodiments of the enterprise credit investigation analysis method based on big data.
[0049] The system includes: A data acquisition module, which is used to construct a big database according to the data sources of enterprise credit investigation analysis, collect the required enterprise credit investigation data from the big database using the kettle tool, and perform cleaning and sorting; A data analysis module, which is used to analyze the collected enterprise credit investigation data, identify risk factors, evaluate the risk level, and give risk warnings to enterprises with credit risks; A credit assessment model construction module, which is used to determine the business requirements of enterprise credit investigation analysis, select an algorithm from the algorithm library according to the business requirements to construct a credit assessment model, and use the collected enterprise credit investigation data and the data analysis results to train the credit assessment model; A credit assessment module, which is used to respond to the user's credit assessment request for the target enterprise, input the enterprise credit investigation data of the target enterprise into the trained credit assessment model to obtain a credit assessment result; An enterprise credit investigation analysis display module, which is used to determine the data display method according to the business requirements and the data type of the credit assessment result, and visually output the credit assessment result according to the determined data display method.
[0050] In this embodiment, through the interactive cooperation of the data acquisition module, the data analysis module, the credit assessment model construction module, the credit assessment module, and the enterprise credit investigation analysis and display module, efficient enterprise credit assessment is achieved.
Claims
1. A corporate credit analysis method based on big data, characterized in that: The steps include: S1. Build a big database based on the data source of corporate credit analysis, use the kettle tool to collect the required corporate credit data from the big database, and clean and organize it; S2. Analyze the collected corporate credit data, identify risk factors, assess risk levels, and issue risk warnings for companies with credit risks; S3. Determine the business needs of corporate credit analysis, select algorithms from the algorithm library to build a credit assessment model based on the business needs, and use the collected corporate credit data and data analysis results to train the credit assessment model; S4. responding to the user's request for credit assessment of the target enterprise, inputting the target enterprise's credit data into the trained credit assessment model to obtain a credit assessment result; S5. Determine the data display method based on business needs and the data type of the credit assessment results, and visualize the credit assessment results according to the determined data display method.
2. The enterprise credit analysis method based on big data according to claim 1 is characterized in that: The specific steps of step S1 are as follows: S11. Determine the data source for corporate credit analysis and apply for access rights to the corresponding data in the data source; S12. Use the data obtained from the data source to obtain the approved permission to build a large database; S13. Build a relational target database, use the kettle tool to create a scheduled data update task, export the required corporate credit data from the big database regularly, and import it into the relational target database; S14. Clean the corporate credit data and remove duplicate, invalid and erroneous data; S15. Classify and sort the corporate credit data according to the target fields and fill them into the database tables of the corresponding categories, and establish associations between different database tables.
3. The enterprise credit analysis method based on big data according to claim 2 is characterized in that: The specific steps of step S15 are as follows: S151. Determine the category of corporate credit data; S152. Determine the target field of the enterprise credit data in each category; S153. Unify the attributes of the corporate credit data in the relational target database according to the field name, data type and length of the target field, and process the missing or abnormal data; S154. Fill the processed corporate credit data into the database table of the corresponding category; S155. Establish associations between different database tables through primary and foreign key constraints of database tables; S156. Determine the fields in each database table that need to be formatted, and perform the conversion according to preset rules.
4. The enterprise credit analysis method based on big data according to claim 2 is characterized in that: The specific steps of step S2 are as follows: S21. Summarize the basic characteristics of the cleaned and sorted corporate credit data and complete descriptive statistical analysis; S22. Select an independent variable field and a dependent variable field from the target field, fit the corresponding data of the selected fields through a pre-built regression model to obtain inferred features, and analyze the relationship between the two fields through the inferred feature values; S23. Compare the enterprise credit data with the preset index threshold, and cluster them according to financial information, operating information and credit history to obtain different risk levels; S24. Select decision fields from the corporate credit data, and filter out field combinations from the target fields based on support and confidence, and extract association rules.
5. The enterprise credit analysis method based on big data according to claim 4 is characterized in that: The specific steps of step S23 are as follows: S231. Compare the financial information and operating information in the enterprise credit data with the preset index threshold, and issue a risk warning when the index threshold is exceeded; S232. Cluster the corporate credit data according to financial information and credit history to obtain different risk levels, and fill the risk levels into the target fields of the corresponding database tables.
6. The enterprise credit analysis method based on big data according to claim 4 is characterized in that: The specific steps of step S24 are as follows: S241. Taking the risk level as the decision field, filter out the fields related to the decision field from the enterprise credit data according to the data type and business logic, and construct a candidate item set; S242. Count the number of records in the enterprise credit investigation data that contain each candidate item set as the number of transactions in the candidate item set; S243. Calculate the support of each candidate item set according to the number of transactions of each candidate item set; S244. Determine the minimum support threshold according to the business needs of the enterprise credit analysis, and take the candidate item sets with support greater than the minimum support threshold as frequent item sets; S245. Determine a set of association rules based on the frequent item sets, and calculate the confidence of each association rule; S246. Determine the minimum confidence threshold based on the business needs of the enterprise credit analysis, and take the management rules with confidence greater than the minimum confidence threshold as valid association rules.
7. The enterprise credit analysis method based on big data according to claim 4 is characterized in that: The specific steps of step S3 are as follows: S31. Build a credit assessment algorithm library, which includes a logistic regression algorithm, a decision tree algorithm, a random forest algorithm, and a neural network algorithm; S32. Determine the business needs of corporate credit analysis; When the business requirement is to explain the credit assessment results and the data dimension is lower than the dimension threshold, a logistic regression algorithm or a decision tree algorithm is selected from the credit assessment algorithm library; When the business needs require accuracy of the model, the data dimension is higher than the dimension threshold, and there is nonlinearity between the data, select the random forest algorithm or neural network algorithm from the credit assessment algorithm library; S33. construct a credit assessment model using the selected algorithm; S34. Use the cleaned, organized and analyzed corporate credit data as training data to train the constructed credit assessment model.
8. The enterprise credit analysis method based on big data according to claim 7 is characterized in that: The specific steps of step S4 are as follows: S41. Receive a user's credit assessment request for a target enterprise, resolve the user's identity and the target enterprise's related identification, and verify the credit assessment authority based on the user's identity; For example, for financial institution users, their institution qualifications and user permissions need to be verified; For ordinary investor users, their registration information and scope of use rights need to be confirmed; S42. Query and obtain the corresponding enterprise credit data from the relational target database according to the relevant identification of the target enterprise input by the user; S43. The acquired corporate credit data of the target enterprise is sorted and preprocessed according to the format and feature order required by the model, and then input into the trained credit assessment model.
9. The enterprise credit analysis method based on big data according to claim 8 is characterized in that: The specific steps of step S5 are as follows: S51. Determine the content of the credit assessment results displayed based on business needs; S52. Determine the data type of the credit assessment results and determine the data display method based on business needs and data type; S53. On the user interface of the enterprise credit investigation visualization platform, the credit assessment results are visualized according to the determined data display method.
10. A corporate credit analysis system based on big data, characterized in that: include: The data collection module is used to build a big database based on the data source of corporate credit analysis, use the kettle tool to collect the required corporate credit data from the big database, and clean and organize it; The data analysis module is used to analyze the collected corporate credit data, identify risk factors, assess risk levels, and issue risk warnings for companies with credit risks; The credit assessment model building module is used to determine the business needs of corporate credit analysis, select algorithms from the algorithm library to build a credit assessment model based on business needs, and use the collected corporate credit data and data analysis results to train the credit assessment model; The credit assessment module is used to respond to the user's credit assessment request for the target enterprise, input the target enterprise's corporate credit data into the trained credit assessment model to obtain the credit assessment result; The enterprise credit analysis and display module is used to determine the data display method according to business needs and the data type of the credit assessment results, and to visualize the credit assessment results according to the determined data display method.