Text analysis and machine learning embedded bond default prediction method
By embedding text analysis and machine learning methods, we obtain a set of qualitative and quantitative indicators of bonds, conduct embedded text analysis and probability distribution fixed-time analysis, build a target bond default indicator pool, use machine learning models to extract key features, and train a logistic regression model. This solves the problem of incomplete bond default prediction in existing technologies and achieves more accurate default prediction.
Patent Information
- Application Number
- CN202510793182.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-16
AI Technical Summary
Existing technologies rely on structured financial data in bond default risk assessment and fail to effectively capture external risk factors such as macroeconomic cycles, industry policy adjustments, and market interest rate changes. This leads to systematic deviations in prediction results in complex economic environments. The lack of unstructured data sources also limits the integrated analysis of multidimensional risk factors.
By embedding text analysis and machine learning methods, we obtain a set of qualitative and quantitative indicators of bonds, conduct embedded text analysis and probability distribution fixed-time analysis, build a target bond default indicator pool, use machine learning models to extract key features, and train a logistic regression model to predict bond defaults.
It improves the accuracy and explainability of bond default predictions, reduces the risk of bond defaults, and provides a more scientific basis for default predictions.
Smart Images

Figure CN120655415A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of financial risk management, and in particular to a bond default prediction method embedded with text analysis and machine learning. Background Art
[0002] In the field of bond default risk assessment, existing technical frameworks primarily rely on structured financial data to construct predictive models, but their input feature systems are limited. Current mainstream models rely on traditional financial indicators such as debt-to-asset ratios and cash flow indicators as core inputs. While these features can reflect a company's static debt repayment capacity, they lack the ability to capture dynamic external risk factors such as macroeconomic cyclical fluctuations, industry policy adjustments, and market interest rate fluctuations. Furthermore, the lack of unstructured data sources (such as public opinion and supply chain relationship networks) results in incomplete coverage of risk dimensions, limiting the model's ability to integrate and analyze multidimensional risk factors. Regarding data quality, corporate financial reports are subject to disclosure lags and room for subjective adjustments, and differences in data collection standards across institutions further exacerbate raw data heterogeneity. Furthermore, the inherent high-frequency noise characteristics of market transaction data (such as short-term price fluctuations) create a scale mismatch with the low-frequency characteristics of fundamental data, making it difficult to effectively filter out interfering signals during feature engineering. This limits the model's ability to fully capture the drivers of bond defaults and makes predictions susceptible to systematic bias in complex economic environments.
[0003] Chinese patent publication number CN111583012A discloses a method for assessing the default risk of credit bond issuers that integrates text information. It merely uses news, public opinion, macroeconomic, and financial data to construct an assessment model, without analyzing the qualitative and quantitative characteristics of the bonds themselves. This results in incomplete default prediction indicators, large deviations in prediction results, and consequently insufficient prediction accuracy. Summary of the Invention
[0004] To this end, the present invention provides a bond default prediction method embedded with text analysis and machine learning, so as to overcome the problem in the prior art that the evaluation model is constructed by merely using news public opinion, macroeconomic and financial data, without analyzing the qualitative and quantitative characteristics of the bonds themselves, resulting in incomplete default prediction indicators, large deviations in prediction results, and thus insufficient prediction accuracy.
[0005] To achieve the above objectives, the present invention provides a bond default prediction method embedded with text analysis and machine learning, comprising the following steps: S1. Obtain bond qualitative indicator set and bond quantitative indicator set based on market bond sample data; S2. Perform embedded text analysis on the bond qualitative indicator set to obtain the corresponding bond qualitative characteristic indicator set, perform probability distribution fixed-time analysis on the bond quantitative indicator set to obtain the bond quantitative characteristic indicator set, and construct a target bond default indicator pool based on the bond qualitative characteristic indicator set and the bond quantitative characteristic indicator set; S3. Extract features from the target bond default indicator pool based on the machine learning model to obtain a set of key target bond default indicators, and construct a bond default prediction dataset based on the set of key target bond default indicators; S4. Train the Logistic regression model based on the bond default prediction dataset, and output the Logistic regression model that meets the preset analysis and test accuracy as a bond default prediction model, and perform bond default prediction for the target bond-issuing enterprise based on the bond default prediction model.
[0006] In this solution, based on market bond sample data, the qualitative and quantitative indicator sets of bonds are obtained, and embedded text analysis is performed on the qualitative indicator set to extract the qualitative feature indicator set; probability distribution fixed-time analysis is performed on the quantitative indicator set to obtain the quantitative feature indicator set. The two are used together to construct a target bond default indicator pool. The machine learning model is used to extract key features from the indicator pool to form a key target bond default indicator set, and a bond default prediction dataset is constructed. The Logistic regression model is trained based on the dataset. After testing to meet the accuracy requirements, it is output as a bond default prediction model, which is used to predict bond defaults for target bond-issuing enterprises.
[0007] Compared with the existing technology, the beneficial effect of this application is that by obtaining a set of qualitative and quantitative bond indicators based on market bond sample data, a comprehensive foundation is provided for subsequent analysis, and embedded text analysis is performed on qualitative indicators and probability distribution fixed-point analysis is performed on quantitative indicators. The indicator characteristics can be deeply mined, and the constructed target bond default indicator pool is more representative. The key target bond default indicator set is extracted using a machine learning model, effectively focusing on the core features, and the constructed bond default prediction data set is of higher quality. On this basis, a logistic regression model is trained to output a bond default prediction model that meets the preset analysis test accuracy, providing a basis for predicting the probability of bond default, while enhancing the interpretability of the model results, which can significantly improve the accuracy of bond default prediction for target bond-issuing enterprises and reduce the risk of bond default.
[0008] Furthermore, the S1 comprises the following steps: S11. Filtering first bond issuance information of multiple bond issuers that have defaulted on bonds from bond market data, and classifying and labeling the first bond issuance information to obtain a bond default sample; S12. Determine the selection ratio of non-defaulting bond issuers and bond samples based on the corporate size and industry attributes of each bond issuer in the bond default sample, and select normal bond samples from the second bond issuance information of the non-defaulting bond issuers based on the bond sample selection ratio; S13. Construct market bond sample data based on the bond default sample and the bond normal sample, and determine the bond qualitative indicator set and the bond quantitative indicator set based on the market bond sample data.
[0009] In this plan, the first bond issuance information refers to the original information of the bond issuer that has defaulted at the time of bond issuance (such as issuance size, face interest rate, term, prospectus, etc.); the bond default sample refers to the dataset that has been classified and labeled by default status (such as material default, technical default); the bond sample selection ratio refers to the sampling weight of non-defaulted bonds determined based on enterprise size (such as total assets) and industry attributes (such as manufacturing / finance); the second bond issuance information refers to the bond issuance data of non-defaulting issuers (such as the issuance records of credit bonds and interest-bearing bonds that have been normally redeemed); and the normal bond sample refers to the dataset representing no default risk, extracted from the second bond issuance information according to the bond sample selection ratio.
[0010] By screening the first bond issuance information of issuers that have defaulted on bonds from bond market data and classifying and labeling them to obtain bond default samples, we can accurately locate the characteristics of the defaulting entities. Based on the corporate scale and industry attributes of the issuers in the bond default samples, we can determine the selection ratio of non-defaulting bond issuers and bond samples. Then, we can select normal bond samples from the second bond issuance information of non-defaulting bond issuers, ensuring the scientificity and rationality of the selection of normal samples. Finally, we construct market bond sample data based on the bond default samples and the normal bond samples, and determine the bond qualitative data and bond quantitative data accordingly. The constructed market bond sample data takes into account both default and normal situations, making the determined bond qualitative indicator set and bond quantitative indicator set more comprehensive.
[0011] Furthermore, the step S2 includes the following steps: S21. Performing embedded text analysis on each bond qualitative indicator in the bond qualitative indicator set to obtain embedded text analysis results, determining a quantification rule for each bond qualitative indicator based on the embedded text analysis results, and quantifying each bond qualitative indicator based on the quantification rule for each bond qualitative indicator to obtain a first indicator score corresponding to each bond qualitative indicator; S22. Add the first indicator score and the first mapping relationship between the first indicator score and the bond qualitative indicator to the bond qualitative characteristic indicator set; S23. Perform a probability distribution fixed-point analysis on each bond quantitative indicator in the bond quantitative indicator set to obtain a second indicator score corresponding to each bond quantitative indicator, and add the second indicator score and a second mapping relationship between the second indicator score and the bond quantitative indicator to the bond quantitative characteristic indicator set; S24. Construct a target bond default indicator pool based on the bond qualitative characteristic indicator set and the bond quantitative characteristic indicator set.
[0012] In this scheme, the first indicator score refers to the score obtained after performing embedded text analysis on each bond qualitative indicator in the bond qualitative indicator set, which is used to quantify the characteristics of the qualitative indicator. The first mapping relationship refers to the correlation between the bond qualitative indicator and the corresponding first indicator score. The second indicator score refers to the score obtained after performing probability distribution fixed-point analysis on each bond quantitative indicator in the bond quantitative indicator set, which is used to quantify the characteristics of the quantitative indicator. The second mapping relationship refers to the correlation between the bond quantitative indicator and the corresponding second indicator score. The quantification rule of each bond qualitative indicator refers to a set of rule systems formulated based on the results of embedded text analysis to convert bond qualitative indicators into quantifiable numerical values.
[0013] By performing embedded text analysis on bond qualitative indicators, the first indicator score is obtained and the first mapping relationship is determined. By performing probability distribution fixed-time analysis on bond quantitative indicators, the second indicator score is obtained and the second mapping relationship is determined. Then, a set of qualitative and quantitative characteristic indicators of bonds is constructed and a target bond default indicator pool is formed, making the indicators more quantitative and improving the accuracy of bond default analysis.
[0014] Furthermore, the step S21 includes the following steps: S211. Performing embedded text analysis on each bond qualitative indicator in the bond qualitative indicator set using natural language processing technology to obtain embedded text analysis results, and obtaining word frequency feature parameters and sentiment tendency parameters corresponding to each bond qualitative indicator based on the embedded text analysis results; S212: performing normalized weight calculation on the word frequency feature parameter to obtain a normalized weight calculation result, and mapping the normalized weight calculation result to a preset scoring range to obtain a word frequency feature index; S213, performing dictionary matching on the sentiment tendency parameter to generate an intensity parameter, and mapping the intensity parameter to a preset scoring range to obtain an intensity parameter index; S214. Fusing the word frequency feature index and the intensity parameter index to obtain a first indicator score corresponding to each bond qualitative indicator.
[0015] In this solution, embedded text analysis of bond qualitative indicators is performed through natural language processing technology to obtain word frequency feature parameters and sentiment tendency parameters. The word frequency feature parameters are then normalized and weighted and mapped to scores. The sentiment tendency parameter dictionary is matched to generate intensity parameters and mapped to scores. Finally, the two are combined to obtain the first indicator score, which can more comprehensively and accurately quantify bond qualitative indicators and improve the accuracy of bond qualitative feature assessment.
[0016] Furthermore, the S23 includes the following steps: S231. Standardize each bond quantitative indicator in the bond quantitative indicator set to obtain a standardized data sequence, and fit a normal distribution model based on the standardized data sequence; S232. Extract distribution characteristic parameters according to the normal distribution model, and determine the tiered data intervals and the bond default risk levels corresponding to the tiered data intervals according to the distribution characteristic parameters; S233. Score the bond default risk level to obtain a second indicator score corresponding to each bond quantitative indicator.
[0017] In this scheme, by standardizing each bond quantitative indicator in the bond quantitative indicator set and fitting a normal distribution model, extracting the distribution characteristic parameters to determine the tiered data interval and the corresponding bond default risk level, and then scoring the bond default risk level to obtain the second indicator score, it is possible to standardize the quantitative indicator data, clarify its distribution pattern, accurately divide the risk interval and quantify the degree of default risk, thereby improving the standardization of the assessment of bond quantitative indicators.
[0018] Furthermore, the step S24 includes the following steps: S241. Structurally integrate the bond qualitative characteristic indicator set and the bond quantitative characteristic indicator set to obtain a target bond indicator set including a first indicator score, a first mapping relationship, a second indicator score, and a second mapping relationship; S242. Preset indicator logic mapping rules between bond qualitative indicators and bond quantitative indicators based on bond default business domain knowledge; S243. According to the indicator logic mapping rules, the target bond indicator set is divided into a bottom original indicator layer, a middle composite indicator layer and a top default identification layer according to the bond risk identification priority. The bottom original indicator layer stores bond qualitative indicators and bond quantitative indicators, the middle composite indicator layer drives the indicator combination calculation, and the top default identification layer generates the target bond default indicator pool.
[0019] In this plan, the preset bond default data standard refers to a set of pre-set data features or rule sets used to measure bond default situations, such as the financial indicator thresholds and credit rating lower limits of defaulted bonds. The bond default business domain knowledge refers to the professional knowledge, experience, rules and patterns related to bond defaults in the bond market, including the definition, causes, influencing factors, prediction methods and processing procedures of bond defaults. The indicator logic mapping rules refer to the logical relationship rules between preset bond qualitative indicators and bond quantitative indicators based on the bond default business domain knowledge. For example, there is a certain logical relationship between certain qualitative indicators (such as the issuer's credit rating) and quantitative indicators (such as the bond's yield to maturity). The bond risk identification priority refers to the order in which different indicators or indicator layers are evaluated when assessing bond default risk.
[0020] Through structured integration, a target bond indicator set containing multiple scores and mapping relationships is formed, so that the indicator information is ordered, the indicator logic mapping rules are preset, and the indicator association is clarified. After layered processing, it can accurately generate a target bond default indicator pool based on bond risk identification priority analysis, thereby improving the efficiency and accuracy of bond risk identification and analysis.
[0021] Furthermore, the step S3 includes the following steps: S31. Training a convolutional neural network model in the machine learning model based on a preset bond default feature extraction data set, and outputting the convolutional neural network model that meets a preset accuracy rate as a bond default feature recognition model; S32. Construct multidimensional feature data based on each target bond default indicator in the target bond default indicator pool, as well as the first indicator score and the second indicator score. Input the multidimensional feature data into the bond default feature recognition model for forward propagation calculation to obtain a weighted response value for each feature dimension. S33. Based on the local receptive field characteristics of the convolutional layer in the bond default feature recognition model, perform feature map convolution and maximum pooling operations on the weighted response values to obtain a reduced-dimensional intermediate feature vector. S34. Based on the statistical distribution of the intermediate eigenvectors and the gradient back propagation result of the model loss function, screen out key target bond default indicators with a contribution level not less than a preset bond default prediction level, and add the key target bond default indicators to the key target bond default indicator set; S35. Perform time series alignment, missing value filling, and data format unification processing on each key target bond default indicator in the key target bond default indicator set to obtain a bond default prediction data set.
[0022] In this solution, the preset bond default feature extraction dataset refers to a preset dataset for training a convolutional neural network model, which is stored in the form of historical bond default feature extraction data - key target bond default indicators. The convolutional neural network model refers to a machine learning model used to extract features from each target bond default indicator in the target bond default indicator pool and predict the corresponding key target bond default indicators. The local receptive field characteristic refers to the bond default feature identification model focusing on the local relationship of the bond default indicators, thereby extracting local features that have a significant impact on the bond default risk. The variance distribution refers to the method used to understand the volatility of each intermediate feature vector. The correlation matrix refers to the method used to understand the degree of mutual influence between each intermediate feature vector. The preset bond default prediction contribution refers to a pre-set threshold, such as 0.1, used to measure the importance of the bond default indicator to the bond default prediction result when performing a bond default prediction task.
[0023] By training a convolutional neural network model based on a preset bond default feature extraction dataset, a bond default feature recognition model that meets the preset accuracy is output, and a target bond default indicator pool is used to construct multidimensional feature data and input it into the bond default feature recognition model for calculation to obtain a weighted response value. Convolution and pooling operations are performed based on the convolutional layer characteristics of the bond default feature recognition model to obtain an intermediate feature vector after dimensionality reduction. Based on the statistical distribution and gradient back propagation results, key target bond default indicators are screened out and a key target bond default indicator set is constructed. Finally, the key target bond default indicator set is processed to obtain a bond default prediction dataset, which can make the bond default prediction dataset more targeted and accurate.
[0024] Further, the S4 includes the following steps: S41. Construct a logistic regression model based on the bond default prediction dataset; S42. Divide 60% of the bond default prediction dataset into a bond default prediction training set, a 20% bond default prediction validation set, and a 20% bond default prediction test set; S43, inputting the bond default prediction training set into the Logistic regression model to train the Logistic regression model, and inputting the bond default prediction validation set into the trained Logistic regression model to iteratively optimize the trained Logistic regression model; S44, inputting the bond default prediction test set into the iteratively optimized Logistic regression model to perform analysis and testing on the iteratively optimized Logistic regression model, and outputting the Logistic regression model that meets the preset analysis and testing accuracy as the bond default prediction model; S45. Obtain the bond issuance data of the target bond-issuing enterprise, pre-process the bond issuance data to obtain actual bond issuance data, and input the actual bond issuance data into the bond default prediction model for analysis to obtain the target bond default prediction probability corresponding to the target bond-issuing enterprise.
[0025] In this solution, the logistic regression model refers to a generalized linear regression model that maps linear combinations to probability values (0-1) using the sigmoid function. It is suitable for binary classification problems (such as default / non-default). The preset analysis test accuracy refers to the minimum performance threshold that the model must achieve on the test set, for example, 85%.
[0026] By constructing a Logistic regression model based on the bond default prediction dataset and reasonably dividing it into training set, validation set and test set, the model is trained with the training set, iteratively optimized with the help of the validation set, and then the optimized model is analyzed and tested with the test set. Finally, a bond default prediction model that meets the preset analysis and test accuracy is output, ensuring the accuracy and reliability of the model. At the same time, the bond issuance data of the target bond issuing enterprise is obtained and pre-processed before being input into the model for analysis to obtain the target bond default prediction probability, which can provide a scientific and effective basis for evaluating the default risk of the target bond issuing enterprise.
[0027] Furthermore, the S41 includes the following steps: S411. Obtain each key target bond default indicator based on the bond default prediction data set, and set the bond qualitative indicator and the bond quantitative indicator in each key target bond default indicator as model independent variables; S412. Determine the corresponding bond issuing enterprise based on each key target bond default indicator, set the bond default prediction probability corresponding to the bond issuing enterprise as the model dependent variable, and construct a logistic regression model based on the model independent variables and the model dependent variables.
[0028] In this scheme, by obtaining various key target bond default indicators based on the bond default prediction data set, and setting the bond qualitative indicators and bond quantitative indicators as the model independent variables, the issuing companies are determined based on the key target bond default indicators, and the bond default prediction probability corresponding to the issuing companies is set as the model dependent variable, and then a logistic regression model is constructed. This can clearly establish the quantitative relationship between the key target bond default indicators and the bond default prediction probability, which helps to more accurately assess the bond default risk.
[0029] Furthermore, the mathematical expression of the Logistic regression model is: ; in, represents the bond default probability of the issuing enterprise i in the bond default year t, represents the logarithmic probability of bond default by bond-issuing enterprise i in bond default year t, , represents the probability of bond default for bond issuing enterprise i in bond default year t, represents the natural logarithm function, represents the base value of the logarithmic probability of bond default when all independent variables are 0, represents the bond qualitative index, m represents the total number of bond qualitative indexes, represents the regression coefficient of the j-th bond qualitative indicator, represents the qualitative index of the jth bond of the bond issuing enterprise i in the bond default year t, represents the bond quantitative index, n represents the total number of bond quantitative indexes, represents the regression coefficient of the k-th bond quantitative indicator, represents the kth bond quantitative indicator of bond issuing enterprise i in the bond default year t-1, represents the random disturbance term.
[0030] In this scheme, the mathematical expression of the Logistic regression model can clearly quantify the relationship between the independent variable and the dependent variable, intuitively present the influence of each factor on the probability of bond default, provide a basis for predicting the probability of bond default, and enhance the interpretability of the model results. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 This is a structural diagram of a bond default prediction method that embeds text analysis and machine learning in an embodiment of the present invention. DETAILED DESCRIPTION
[0032] The following is further described in detail through specific implementation methods: See also Figure 1 As shown in FIG, it is a flow chart of a bond default prediction method embedded with text analysis and machine learning according to an embodiment of the present invention, which includes the following steps: S1. Obtain bond qualitative indicator set and bond quantitative indicator set based on market bond sample data; S2. Perform embedded text analysis on the bond qualitative indicator set to obtain the corresponding bond qualitative characteristic indicator set, perform probability distribution fixed-time analysis on the bond quantitative indicator set to obtain the bond quantitative characteristic indicator set, and construct a target bond default indicator pool based on the bond qualitative characteristic indicator set and the bond quantitative characteristic indicator set; S3. Extract features from the target bond default indicator pool based on the machine learning model to obtain a set of key target bond default indicators, and construct a bond default prediction dataset based on the set of key target bond default indicators; S4. Train the Logistic regression model based on the bond default prediction dataset, and output the Logistic regression model that meets the preset analysis and test accuracy as a bond default prediction model, and perform bond default prediction for the target bond-issuing enterprise based on the bond default prediction model.
[0033] Specifically, S1 includes the following steps: S11. Filtering first bond issuance information of multiple bond issuers that have defaulted on bonds from bond market data, and classifying and labeling the first bond issuance information to obtain a bond default sample; S12. Determine the selection ratio of non-defaulting bond issuers and bond samples based on the corporate size and industry attributes of each bond issuer in the bond default sample, and select normal bond samples from the second bond issuance information of the non-defaulting bond issuers based on the bond sample selection ratio; S13. Construct market bond sample data based on the bond default sample and the bond normal sample, and determine the bond qualitative indicator set and the bond quantitative indicator set based on the market bond sample data.
[0034] This embodiment does not limit the implementation method of classifying and labeling the first bond issuance information. For example, it can be set to use NLP to parse the default clauses (such as cross default and early maturity) in the prospectus, and combine the historical default records of a third-party rating agency to label the bond issuance information with a default type label.
[0035] Specifically, S2 includes the following steps: S21. Performing embedded text analysis on each bond qualitative indicator in the bond qualitative indicator set to obtain embedded text analysis results, determining a quantification rule for each bond qualitative indicator based on the embedded text analysis results, and quantifying each bond qualitative indicator based on the quantification rule for each bond qualitative indicator to obtain a first indicator score corresponding to each bond qualitative indicator; S22. Add the first indicator score and the first mapping relationship between the first indicator score and the bond qualitative indicator to the bond qualitative characteristic indicator set; S23. Perform a probability distribution fixed-point analysis on each bond quantitative indicator in the bond quantitative indicator set to obtain a second indicator score corresponding to each bond quantitative indicator, and add the second indicator score and a second mapping relationship between the second indicator score and the bond quantitative indicator to the bond quantitative characteristic indicator set; S24. Construct a target bond default indicator pool based on the bond qualitative characteristic indicator set and the bond quantitative characteristic indicator set.
[0036] Specifically, S21 includes the following steps: S211. Performing embedded text analysis on each bond qualitative indicator in the bond qualitative indicator set using natural language processing technology to obtain embedded text analysis results, and obtaining word frequency feature parameters and sentiment tendency parameters corresponding to each bond qualitative indicator based on the embedded text analysis results; S212: performing normalized weight calculation on the word frequency feature parameter to obtain a normalized weight calculation result, and mapping the normalized weight calculation result to a preset scoring range to obtain a word frequency feature index; S213, performing dictionary matching on the sentiment tendency parameter to generate an intensity parameter, and mapping the intensity parameter to a preset scoring range to obtain an intensity parameter index; S214. Fusing the word frequency feature index and the intensity parameter index to obtain a first indicator score corresponding to each bond qualitative indicator.
[0037] Specifically, natural language processing technology refers to the use of computers to process and analyze text data (such as word segmentation, part-of-speech tagging, and semantic understanding). Embedded text analysis results refer to the analysis results of converting text into vector representations using word embedding models (such as Word2Vec) and extracting keywords and sentiment polarity. The word frequency feature parameter refers to the frequency statistics of preset risk keywords (such as default, litigation, and restructuring) in the text. The sentiment tendency parameter refers to the intensity of the text sentiment polarity obtained by matching the sentiment lexicon. For example, in a bond clause, "default" appears three times and "guarantor's qualifications are good" appears once, with a sentiment score of 0.6 (predominantly negative). The normalized weight calculation result refers to the conversion of raw word frequencies into standardized weights in the range of 0-1 using TF-IDF or L2 normalization. The preset scoring range refers to a pre-set score mapping interval (such as mapping word frequency weights to 0-10 points and sentiment intensity to 0-10 points). The word frequency feature index refers to the structured score value after normalizing the word frequency weights. The intensity parameter refers to the sum of the intensity values of the sentiment words matched in the sentiment lexicon. The intensity parameter index refers to the normalized score after the intensity parameter is mapped to the preset scoring range.
[0038] This embodiment does not limit the implementation method of normalizing the weight calculation of the word frequency feature parameters. For example, it can be set to use the TF-IDF algorithm to calculate the keyword weight, and scale the weight to the interval [0,1] through Min-Max normalization; this embodiment does not limit the implementation method of dictionary matching of the sentiment tendency parameter. For example, it can be set to use the financial field sentiment dictionary (such as the Loughran-McDonald dictionary) to match the sentiment words in the text, and calculate the total score according to the preset strength value of the dictionary (such as "bankruptcy" is 3, "sound" is 2); this embodiment does not limit the implementation method of data fusion of word frequency feature indicators and strength parameter indicators. For example, it can be set to combine the word frequency feature indicators and the strength parameter indicators through weighted summation to obtain a comprehensive score.
[0039] Specifically, S23 includes the following steps: S231. Standardize each bond quantitative indicator in the bond quantitative indicator set to obtain a standardized data sequence, and fit a normal distribution model based on the standardized data sequence; S232. Extract distribution characteristic parameters according to the normal distribution model, and determine the tiered data intervals and the bond default risk levels corresponding to the tiered data intervals according to the distribution characteristic parameters; S233. Score the bond default risk level to obtain a second indicator score corresponding to each bond quantitative indicator.
[0040] This embodiment does not limit the implementation method of standardizing each bond quantitative indicator in the bond quantitative indicator set, such as setting it to standardize each bond quantitative indicator in the bond quantitative indicator set using the Z-Score standard method; this embodiment does not limit the implementation method of the bond default risk level score, such as setting it to be based on the normal distribution model. The bond default risk level is scored based on the principle of
[0041] Specifically, S24 includes the following steps: S241. Structurally integrate the bond qualitative characteristic indicator set and the bond quantitative characteristic indicator set to obtain a target bond indicator set including a first indicator score, a first mapping relationship, a second indicator score, and a second mapping relationship; S242. Preset indicator logic mapping rules between bond qualitative indicators and bond quantitative indicators based on bond default business domain knowledge; S243. According to the indicator logic mapping rules, the target bond indicator set is divided into a bottom original indicator layer, a middle composite indicator layer and a top default identification layer according to the bond risk identification priority. The bottom original indicator layer stores bond qualitative indicators and bond quantitative indicators, the middle composite indicator layer drives the indicator combination calculation, and the top default identification layer generates the target bond default indicator pool.
[0042] In this example, a database of bond default cases is used to refine risk transmission pathways (e.g., financial deterioration leading to refinancing difficulties) and extract key drivers, forming qualitative rules (e.g., a credit rating downgrade triggers an early warning) and quantitative rules (e.g., a lower interest coverage ratio below a threshold triggers a crisis). These rules are stored as a logic library using formal expressions (e.g., an "IF-THEN" structure). Weights are assigned to these rules through historical data backtesting, ensuring dynamic updates to reflect market changes. Based on risk identification priorities, the target bond indicator set is divided into three tiers: a bottom-level raw indicator tier stores raw qualitative and quantitative bond indicators, retaining the original scores; a middle-level composite indicator tier drives indicator calculations through weighted synthesis or conditional combinations (e.g., financial risk index = 0.4 × leverage ratio + 0.6 × cash flow coverage ratio) to generate composite indicators (e.g., solvency score); and a top-level default indicator tier generates final risk indicators (e.g., high / medium / low risk) based on composite indicators and threshold rules. Tier weights or thresholds are dynamically adjusted through real-time data feedback to form a target bond default indicator pool.
[0043] Specifically, S3 includes the following steps: S31. Training a convolutional neural network model in the machine learning model based on a preset bond default feature extraction data set, and outputting the convolutional neural network model that meets a preset accuracy rate as a bond default feature recognition model; S32. Construct multidimensional feature data based on each target bond default indicator in the target bond default indicator pool, as well as the first indicator score and the second indicator score. Input the multidimensional feature data into the bond default feature recognition model for forward propagation calculation to obtain a weighted response value for each feature dimension. S33. Based on the local receptive field characteristics of the convolutional layer in the bond default feature recognition model, perform feature map convolution and maximum pooling operations on the weighted response values to obtain a reduced-dimensional intermediate feature vector. S34. Based on the statistical distribution of the intermediate eigenvectors and the gradient back propagation result of the model loss function, screen out key target bond default indicators with a contribution level not less than a preset bond default prediction level, and add the key target bond default indicators to the key target bond default indicator set; S35. Perform time series alignment, missing value filling, and data format unification processing on each key target bond default indicator in the key target bond default indicator set to obtain a bond default prediction data set.
[0044] This embodiment does not limit the method for training the convolutional neural network model. Those skilled in the art can freely set it, as long as it meets the requirement of identifying each target bond default indicator in the target bond default indicator pool. For example, it can be set to divide 75% of the preset bond default feature extraction data set into a feature extraction data training set and 25% into a feature extraction data test set. The feature extraction data training set is input into the convolutional neural network model for training, and the feature extraction data test set is input into the trained convolutional neural network model. The parameters in the convolutional neural network model are optimized iteratively until the accuracy of the output result of the feature extraction data test set of the convolutional neural network model reaches a preset accuracy rate. The convolutional neural network model is then output as a bond default feature recognition model. The preset accuracy rate refers to a preset value of the accuracy rate that reflects the training status of the convolutional neural network model. This embodiment does not limit the value of the preset accuracy rate. Those skilled in the art can freely set it, as long as it meets the requirement of reflecting the training status of the convolutional neural network model. For example, the preset accuracy rate can be set to 95%.
[0045] In this example, based on each target bond default indicator in the target bond default indicator pool, as well as the first and second indicator scores, the scores for each indicator are aligned in time series to form a multidimensional array that includes the time dimension. For example, the "credit rating score" and "normalized debt-to-asset ratio score" are used as two feature dimensions, and together with other indicators such as the "cash flow coverage ratio score," an N×M matrix is constructed (N is the number of time points, M is the number of indicators). The data format is uniform (for example, all scores are normalized to the range of 0-1), and missing values are filled using interpolation or industry averages, ultimately generating structured multidimensional feature data.
[0046] After inputting multidimensional feature data into the convolutional layer of the bond default feature recognition model, the model leverages local receptive field properties to perform a weighted summation of local regions of the input data (e.g., a 3×3 window) to generate a feature map. The weight matrix of each convolution kernel is dot-producted with the local region of the input data to obtain weighted responses (i.e., the value of each element in the feature map). Subsequently, the feature map is reduced in dimensionality through a max pooling operation (e.g., taking the maximum value of a 2×2 window), retaining the most significant feature responses while reducing computational complexity. For example, the maximum of four adjacent weighted responses in the feature map is taken to generate a more compact feature vector, ultimately outputting the reduced-dimensional intermediate feature vector. Based on the statistical distribution of the intermediate feature vector (e.g., variance contribution) and the gradient backpropagation results of the model loss function, the contribution of each indicator to default prediction (e.g., gradient magnitude) is calculated. A preset bond default prediction contribution is set (e.g., the top 20% of indicators). Indicators with a contribution not below the preset bond default prediction contribution are selected as key target bond default indicators and added to the set of key target bond default indicators.
[0047] This embodiment does not limit the specific implementation methods for time series alignment, missing value filling, and data format unification. These methods may include using a dynamic time warping (DTW) algorithm to perform nonlinear timeline alignment on multiple asynchronous indicators (such as monthly financial data and daily market transaction data) using key time nodes (such as financial report release dates and interest payment dates) as anchor points. For example, the quarterly EBITDA data and monthly public opinion index of a corporate bond can be mapped to a unified time grid. Multiple interpolation (MICE) is used for quantitative bond indicators, combining historical means, industry means, and related indicators (such as using the debt-to-asset ratio to interpolate the cash flow coverage ratio). For qualitative bond indicators, a BERT-based text generation model is used to predict missing clause descriptions based on contextual semantics (such as completing the blank in "If __ clause is triggered, then..."). A financial domain ontology library is constructed to convert unstructured data (such as PDF financial reports) into RDF triple format. One-hot encoding (such as industry classification) and min-max normalization (such as for numerical indicators) are applied to structured data to ensure that all features fit within the [0, 1] input range.
[0048] Specifically, S4 includes the following steps: S41. Construct a logistic regression model based on the bond default prediction dataset; S42. Divide 60% of the bond default prediction dataset into a bond default prediction training set, a 20% bond default prediction validation set, and a 20% bond default prediction test set; S43, inputting the bond default prediction training set into the Logistic regression model to train the Logistic regression model, and inputting the bond default prediction validation set into the trained Logistic regression model to iteratively optimize the trained Logistic regression model; S44, inputting the bond default prediction test set into the iteratively optimized Logistic regression model to perform analysis and testing on the iteratively optimized Logistic regression model, and outputting the Logistic regression model that meets the preset analysis and testing accuracy as the bond default prediction model; S45. Obtain the bond issuance data of the target bond-issuing enterprise, pre-process the bond issuance data to obtain actual bond issuance data, and input the actual bond issuance data into the bond default prediction model for analysis to obtain the target bond default prediction probability corresponding to the target bond-issuing enterprise.
[0049] This embodiment does not limit the implementation method of training the Logistic regression model. For example, it can be set to input the bond default prediction training set into the Logistic regression model, calculate the bond default prediction probability, optimize the loss function through maximum likelihood estimation, and use the gradient descent method to iteratively update the weight W and bias b until the loss function converges.
[0050] This embodiment does not limit the implementation method of iterative optimization of the trained Logistic Regression model. For example, it can be set to input the bond default prediction validation set into the trained Logistic Regression model and calculate the validation accuracy. If overfitting occurs (for example, the training accuracy is 95% but the validation accuracy is 70%), L2 regularization is used to suppress complexity, and hyperparameters are adjusted (for example, the regularization coefficient λ = 0.1). The training-validation cycle is repeated until the validation accuracy stabilizes.
[0051] This embodiment does not limit the implementation method of analyzing and testing the iteratively optimized logistic regression model. For example, the bond default prediction test set can be input into the iteratively optimized logistic regression model, and indicators such as test accuracy and AUC-ROC are calculated. If the test accuracy is greater than or equal to 85%, the iteratively optimized logistic regression model meets the requirements. Otherwise, the features or structure need to be readjusted, and the final model parameters (for example, the weight vector W = [0.8, −1.2, 0.5] corresponds to the debt-to-asset ratio, cash flow, industry risk, etc.) are output.
[0052] Specifically, S41 includes the following steps: S411. Obtain each key target bond default indicator based on the bond default prediction data set, and set the bond qualitative indicator and the bond quantitative indicator in each key target bond default indicator as model independent variables; S412. Determine the corresponding bond issuing enterprise based on each key target bond default indicator, set the bond default prediction probability corresponding to the bond issuing enterprise as the model dependent variable, and construct a logistic regression model based on the model independent variables and the model dependent variables.
[0053] The mathematical expression of the Logistic regression model is: ; in, represents the bond default probability of the issuing enterprise i in the bond default year t, represents the logarithmic probability of bond default by bond-issuing enterprise i in bond default year t, , represents the probability of bond default for bond issuing enterprise i in bond default year t, represents the natural logarithm function, represents the base value of the logarithmic probability of bond default when all independent variables are 0, represents the bond qualitative index, m represents the total number of bond qualitative indexes, represents the regression coefficient of the j-th bond qualitative indicator, represents the qualitative index of the jth bond of the bond issuing enterprise i in the bond default year t, represents the bond quantitative index, n represents the total number of bond quantitative indexes, represents the regression coefficient of the k-th bond quantitative indicator, represents the kth bond quantitative indicator of bond issuing enterprise i in the bond default year t-1, represents the random disturbance term.
[0054] The above are only embodiments of the present invention. Common knowledge such as the known specific structures and characteristics in the scheme are not described in detail here. Ordinary technicians in the relevant field are aware of all the common technical knowledge in the technical field of the invention before the application date or priority date, can obtain all the existing technologies in the field, and have the ability to apply conventional experimental means before that date. Ordinary technicians in the relevant field can improve and implement this scheme in combination with their own abilities under the inspiration given by this application. Some typical known structures or known methods should not become obstacles for ordinary technicians in the relevant field to implement this application. It should be pointed out that for those skilled in the art, several variations and improvements can be made without departing from the structure of the present invention. These should also be regarded as the scope of protection of the present invention. These will not affect the effect of the implementation of the present invention and the practicality of the patent. The scope of protection required by this application shall be based on the content of its claims, and the specific implementation methods and other records in the specification can be used to interpret the content of the claims.
Claims
1. A bond default prediction method that embeds text analysis and machine learning, characterized by: The following steps are involved: S1. Obtain bond qualitative indicator set and bond quantitative indicator set based on market bond sample data; S2. Perform embedded text analysis on the bond qualitative indicator set to obtain the corresponding bond qualitative characteristic indicator set, perform probability distribution fixed-time analysis on the bond quantitative indicator set to obtain the bond quantitative characteristic indicator set, and construct a target bond default indicator pool based on the bond qualitative characteristic indicator set and the bond quantitative characteristic indicator set; S3. Extract features from the target bond default indicator pool based on the machine learning model to obtain a set of key target bond default indicators, and construct a bond default prediction dataset based on the set of key target bond default indicators; S4. Train the Logistic regression model based on the bond default prediction dataset, and output the Logistic regression model that meets the preset analysis and test accuracy as a bond default prediction model, and perform bond default prediction for the target bond-issuing enterprise based on the bond default prediction model.
2. The bond default prediction method embedded with text analysis and machine learning according to claim 1, characterized in that: Said S1 comprises the following steps: S11. Filtering first bond issuance information of multiple bond issuers that have defaulted on bonds from bond market data, and classifying and labeling the first bond issuance information to obtain a bond default sample; S12. Determine the selection ratio of non-defaulting bond issuers and bond samples based on the corporate size and industry attributes of each bond issuer in the bond default sample, and select normal bond samples from the second bond issuance information of the non-defaulting bond issuers based on the bond sample selection ratio; S13. Construct market bond sample data based on the bond default sample and the bond normal sample, and determine the bond qualitative indicator set and the bond quantitative indicator set based on the market bond sample data.
3. The bond default prediction method embedded with text analysis and machine learning according to claim 1, characterized in that: The S2 comprises the following steps: S21. Performing embedded text analysis on each bond qualitative indicator in the bond qualitative indicator set to obtain embedded text analysis results, determining a quantification rule for each bond qualitative indicator based on the embedded text analysis results, and quantifying each bond qualitative indicator based on the quantification rule for each bond qualitative indicator to obtain a first indicator score corresponding to each bond qualitative indicator; S22. Add the first indicator score and the first mapping relationship between the first indicator score and the bond qualitative indicator to the bond qualitative characteristic indicator set; S23. Perform a probability distribution fixed-point analysis on each bond quantitative indicator in the bond quantitative indicator set to obtain a second indicator score corresponding to each bond quantitative indicator, and add the second indicator score and a second mapping relationship between the second indicator score and the bond quantitative indicator to the bond quantitative characteristic indicator set; S24. Construct a target bond default indicator pool based on the bond qualitative characteristic indicator set and the bond quantitative characteristic indicator set.
4. The bond default prediction method embedded with text analysis and machine learning according to claim 3, characterized in that: The S21 includes the following steps: S211. Performing embedded text analysis on each bond qualitative indicator in the bond qualitative indicator set using natural language processing technology to obtain embedded text analysis results, and obtaining word frequency feature parameters and sentiment tendency parameters corresponding to each bond qualitative indicator based on the embedded text analysis results; S212: performing normalized weight calculation on the word frequency feature parameter to obtain a normalized weight calculation result, and mapping the normalized weight calculation result to a preset scoring range to obtain a word frequency feature index; S213, performing dictionary matching on the sentiment tendency parameter to generate an intensity parameter, and mapping the intensity parameter to a preset scoring range to obtain an intensity parameter index; S214. Fusing the word frequency feature index and the intensity parameter index to obtain a first indicator score corresponding to each bond qualitative indicator.
5. The bond default prediction method embedded with text analysis and machine learning according to claim 3, characterized in that: The S23 includes the following steps: S231. Standardize each bond quantitative indicator in the bond quantitative indicator set to obtain a standardized data sequence, and fit a normal distribution model based on the standardized data sequence; S232. Extract distribution characteristic parameters according to the normal distribution model, and determine the tiered data intervals and the bond default risk levels corresponding to the tiered data intervals according to the distribution characteristic parameters; S233. Score the bond default risk level to obtain a second indicator score corresponding to each bond quantitative indicator.
6. The bond default prediction method embedded with text analysis and machine learning according to claim 3, characterized in that: The S24 includes the following steps: S241. Structurally integrate the bond qualitative characteristic indicator set and the bond quantitative characteristic indicator set to obtain a target bond indicator set including a first indicator score, a first mapping relationship, a second indicator score, and a second mapping relationship; S242. Preset indicator logic mapping rules between bond qualitative indicators and bond quantitative indicators based on bond default business domain knowledge; S243. According to the indicator logic mapping rules, the target bond indicator set is divided into a bottom original indicator layer, a middle composite indicator layer and a top default identification layer according to the bond risk identification priority. The bottom original indicator layer stores bond qualitative indicators and bond quantitative indicators, the middle composite indicator layer drives the indicator combination calculation, and the top default identification layer generates the target bond default indicator pool.
7. The bond default prediction method embedded with text analysis and machine learning according to claim 1, characterized in that: The S3 includes the following steps: S31. Training a convolutional neural network model in the machine learning model based on a preset bond default feature extraction data set, and outputting the convolutional neural network model that meets a preset accuracy rate as a bond default feature recognition model; S32. Construct multidimensional feature data based on each target bond default indicator in the target bond default indicator pool, as well as the first indicator score and the second indicator score. Input the multidimensional feature data into the bond default feature recognition model for forward propagation calculation to obtain a weighted response value for each feature dimension. S33. Based on the local receptive field characteristics of the convolutional layer in the bond default feature recognition model, perform feature map convolution and maximum pooling operations on the weighted response values to obtain a reduced-dimensional intermediate feature vector. S34. Based on the statistical distribution of the intermediate eigenvectors and the gradient back propagation result of the model loss function, screen out key target bond default indicators with a contribution level not less than a preset bond default prediction level, and add the key target bond default indicators to the key target bond default indicator set; S35. Perform time series alignment, missing value filling, and data format unification processing on each key target bond default indicator in the key target bond default indicator set to obtain a bond default prediction data set.
8. The bond default prediction method embedded with text analysis and machine learning according to claim 1, characterized in that: The S4 comprises the following steps: S41. Construct a logistic regression model based on the bond default prediction dataset; S42. Divide 60% of the bond default prediction dataset into a bond default prediction training set, a 20% bond default prediction validation set, and a 20% bond default prediction test set; S43, inputting the bond default prediction training set into the Logistic regression model to train the Logistic regression model, and inputting the bond default prediction validation set into the trained Logistic regression model to iteratively optimize the trained Logistic regression model; S44, inputting the bond default prediction test set into the iteratively optimized Logistic regression model to perform analysis and testing on the iteratively optimized Logistic regression model, and outputting the Logistic regression model that meets the preset analysis and testing accuracy as the bond default prediction model; S45. Obtain the bond issuance data of the target bond-issuing enterprise, pre-process the bond issuance data to obtain actual bond issuance data, and input the actual bond issuance data into the bond default prediction model for analysis to obtain the target bond default prediction probability corresponding to the target bond-issuing enterprise.
9. The bond default prediction method embedded with text analysis and machine learning according to claim 8, characterized in that: The S41 includes the following steps: S411. Obtain each key target bond default indicator based on the bond default prediction data set, and set the bond qualitative indicator and the bond quantitative indicator in each key target bond default indicator as model independent variables; S412. Determine the corresponding bond issuing enterprise based on each key target bond default indicator, set the bond default prediction probability corresponding to the bond issuing enterprise as the model dependent variable, and construct a logistic regression model based on the model independent variables and the model dependent variables.
10. The bond default prediction method embedded with text analysis and machine learning according to claim 9, characterized in that: The mathematical expression of the Logistic regression model is: ; in, represents the bond default probability of the issuing enterprise i in the bond default year t, represents the logarithmic probability of bond default by bond-issuing enterprise i in bond default year t, , represents the probability of bond default for bond issuing enterprise i in bond default year t, represents the natural logarithm function, represents the base value of the logarithmic probability of bond default when all independent variables are 0, represents the bond qualitative index, m represents the total number of bond qualitative indexes, represents the regression coefficient of the j-th bond qualitative indicator, represents the qualitative index of the jth bond of the bond issuing enterprise i in the bond default year t, represents the bond quantitative index, n represents the total number of bond quantitative indexes, represents the regression coefficient of the k-th bond quantitative indicator, represents the kth bond quantitative indicator of bond issuing enterprise i in the bond default year t-1, represents the random disturbance term.
Citation Information
Patent Citations
Text information fused default risk assessment method for credit debt issuer
CN111583012A