Financial health condition automatic assessment method based on clustering analysis

Through the automatic assessment method of financial health status based on cluster analysis, combined with the autoencoder and K-means clustering, the clustering results are adjusted using the enterprise relationship diagram to solve the problem that traditional static analysis methods are difficult to adapt to the rapidly changing market environment, and the accurate and dynamic assessment of the financial health status of enterprises is achieved.

CN120013692APending Publication Date: 2025-05-16HEBEI UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510100979.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Traditional static analysis methods are difficult to adapt to the rapidly changing market environment and cannot comprehensively and accurately reflect the financial health of the company.

Method used

The automatic financial health status assessment method based on clustering analysis is adopted. By obtaining and cleaning the real-time financial data, historical financial data and external economic data of the enterprise, the financial ratio is calculated and historical trend characteristics are extracted. The data is compressed into low-dimensional feature vectors using the autoencoder, and K-means clustering and relationship diagram adjustment is performed based on external economic data and enterprise relationship diagrams.

Benefits of technology

It realizes accurate and dynamic assessment of the financial health status of enterprises, can respond to changes in the market environment in a timely manner, and dynamically update clustering and grouping based on real-time data, overcoming the problem that traditional static analysis methods cannot adapt to the changing market environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013692A_ABST
    Figure CN120013692A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of enterprise financial health analysis, and more specifically relates to a clustering analysis-based financial health condition automatic assessment method. The method comprises the following steps: acquiring and cleaning real-time financial data, historical financial data and external economic data of an enterprise, calculating a financial ratio and extracting historical trend features, compressing high-dimensional complex financial data into low-dimensional feature vectors by using an auto-encoder, and extracting potential financial health information; according to the method, low-dimensional feature vectors, external economic data and an enterprise relation graph are combined, K-means clustering is adopted to preliminarily classify enterprises, and a clustering result is further optimized according to the relation between the enterprises, so that final classification is more accurate and reasonable, the change of a market environment can be responded in time, and the classification efficiency is improved. In addition, clustering groups can be dynamically updated according to real-time data, an accurate and dynamic evaluation system is provided for financial health conditions of enterprises, and the problem that a traditional static analysis method cannot adapt to changing market environments is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the technical field of enterprise financial health analysis, and more specifically, to an automatic financial health assessment method based on cluster analysis. Background Art

[0002] With the increasing globalization of the economy and market competition, the financial health of enterprises has become one of the key indicators for assessing their sustainable development and profitability. The financial health of an enterprise not only reflects its existing operational capabilities, but also provides important decision-making basis for management, investors, creditors and other stakeholders. Traditionally, the assessment of the financial health of an enterprise relies on financial statement analysis, mainly measured by financial ratios (such as debt-to-asset ratio, current ratio, gross profit margin, etc.). However, relying solely on these static indicators often fails to fully and accurately reflect the financial health of an enterprise, especially when dealing with a rapidly changing market environment.

[0003] In the real world, the financial status of an enterprise is affected by many factors, including the external economic environment, industry changes, and market dynamics. The timeliness, dynamism, and complexity of these factors make it difficult for traditional static analysis methods to adapt. Summary of the invention

[0004] The present invention provides an automatic financial health assessment method based on cluster analysis, which is intended to solve the problem that the current traditional static analysis method is difficult to cope with the rapidly changing market environment.

[0005] The automatic assessment method of financial health status based on cluster analysis includes the following steps:

[0006] Step 1: Obtain external economic data, real-time financial data of the enterprise, and historical financial data, clean the acquired data, and standardize the cleaned data;

[0007] Step 2: Calculate financial ratios based on the standardized real-time financial data and historical financial data of the enterprise, and extract historical trend features based on the standardized historical financial data;

[0008] Step 3: The standardized real-time financial data, historical financial data, extracted historical trend features, and financial ratio calculation results are used as inputs of the constructed autoencoder, and the input data is compressed into a low-dimensional feature vector based on the autoencoder;

[0009] Step 4: Combining the low-dimensional feature vector, external economic data and enterprise relationship diagram, first use K-means clustering to preliminarily classify the enterprises, and then further adjust the clustering results through the enterprise relationship diagram to obtain the adjusted clustering results;

[0010] Step 5: Based on the adjusted clustering results, each company is assigned to a category, each category representing a different financial health status; the cluster grouping is then updated based on new real-time data.

[0011] The present invention obtains and cleans the real-time financial data, historical financial data and external economic data of the enterprise, and adopts standardized processing to ensure the consistency of the data; secondly, by calculating the financial ratio and extracting the historical trend characteristics, the changing trend of the enterprise's financial status is fully reflected, avoiding the one-sidedness of a single financial indicator; on this basis, the autoencoder is used to compress the high-dimensional and complex financial data into a low-dimensional feature vector, extract the potential financial health information, reduce the information redundancy, and improve the analysis efficiency; combining the low-dimensional feature vector, external economic data and enterprise relationship diagram, K-means clustering is used to preliminarily classify the enterprises, and the clustering results are further optimized according to the relationship between the enterprises, so that the final classification is more accurate and reasonable. Based on this, the present invention can not only respond to the changes in the market environment in a timely manner, but also dynamically update the clustering grouping according to the real-time data, and provide an accurate and dynamic evaluation system for the financial health status of the enterprise, thereby overcoming the problem that the traditional static analysis method cannot adapt to the changing market environment.

[0012] Preferably, the financial ratio calculation includes current ratio, current assets, current load, quick ratio, debt-to-asset ratio, total liabilities, total assets, gross profit margin, net profit margin and accounts receivable turnover.

[0013] Preferably, the historical trend feature extraction comprises the following steps:

[0014] Cyclical component decomposition: STL decomposition is used to decompose the standardized historical financial data into trend components, seasonal components and cyclical fluctuation components;

[0015] Apply the Hodrick-Prescott filter: HP filter the trend component based on the Hodrick-Prescott filter to obtain the smoothed long-term trend;

[0016] Growth rate characteristic calculation: Calculate annual growth rate and compound annual growth rate based on long-term trends:

[0017]

[0018] Where: Growth Rate t represents the annual growth rate in year t; T t represents the long-term trend in year t; T t―1 represents the long-term trend in year t-1;

[0019]

[0020] Where: T T represents the final trend value; T0 represents the initial value; T represents the number of years;

[0021] Volatility characteristics: Volatility characteristics are extracted based on periodic volatility components, where volatility characteristics include standard deviation and volatility;

[0022] Periodic features: Periodic features are extracted based on seasonal components, where periodic features include period length, seasonal fluctuation amplitude, and seasonal peak and valley values.

[0023] Preferably, the step of compressing the input data into a low-dimensional feature vector based on the autoencoder is as follows:

[0024] Based on the bidirectional LSTM layer, the historical financial data and historical trend features are processed to obtain the hidden states of the historical financial data and historical trend features;

[0025] Based on the fully connected layer, the financial ratio calculation results and the real-time financial data are processed to obtain the hidden state of the financial ratio calculation results and the real-time financial data;

[0026] Splicing historical financial data, historical trend features, financial ratio calculation results, and hidden states of real-time financial data to obtain spliced ​​features;

[0027] The self-attention mechanism is used to perform weighted summation on the concatenated features to generate a refined feature representation; the refined feature representation is mapped to a low-dimensional latent space through a fully connected layer to obtain a low-dimensional feature vector as the output of the autoencoder.

[0028] Preferably, the steps of using K-means clustering to preliminarily classify enterprises are as follows:

[0029] Initialize cluster centers: randomly select K enterprises as initial cluster centers, and the dimension of the cluster centers is d latent +d enconomy , which combines low-dimensional feature vectors with external economic data;

[0030] Calculate the distance: For each enterprise i, the low-dimensional feature vector h i and external economic data i Make up a d latent +d enconomy dimensional feature vector, and then calculate the Euclidean distance between each enterprise and all cluster centers:

[0031]

[0032] Where: d(h i ,c k ) represents the Euclidean distance between the enterprise and the cluster center; Represents the spatial coordinates of the cluster center in the low-dimensional feature vector; Represents the spatial coordinates of the cluster center in the external economic data;

[0033] Assign to the nearest cluster center: Each enterprise i is assigned to the cluster center k closest to it;

[0034] Update cluster centers: After clustering is completed, update the center c of each cluster k is the average characteristic of all enterprises belonging to this cluster:

[0035]

[0036] Where: S k represents all enterprises in cluster k; |S k | represents the size of cluster k;

[0037] By taking the average vector of all enterprises in each cluster as the new cluster center, the steps of calculating the distance to updating the cluster center are repeated until the predetermined number of iterations is reached to obtain a preliminary clustering result.

[0038] Preferably, the specific steps of further adjusting the clustering results through the enterprise relationship diagram are as follows:

[0039] Relationship diagram construction: construct a weighted relationship matrix through the relationships between enterprises;

[0040] Propagation adjustment: Update the labels of each enterprise obtained after preliminary clustering based on the weighted propagation model;

[0041] Convergence judgment: If the value of the label obtained in this round of iteration minus the label obtained in the previous round of iteration is less than the preset threshold, the iteration is stopped; or the iteration is stopped after the preset number of iterations is reached. After the iteration is completed, the category label of the enterprise is used for the final clustering result of each enterprise.

[0042] Preferably, the weighted propagation model is as follows:

[0043]

[0044] Where: represents the category label of enterprise i in the t+1th round of dissemination; represents the preliminary clustering label of enterprise i; w ij represents the strength of the relationship between enterprise i and enterprise j; α represents the relative importance parameter for controlling the initial clustering labels and the propagation of the relationship network; j represents all enterprises related to enterprise i.

[0045] Preferably, the specific steps of constructing the relationship graph are as follows:

[0046] Definition of nodes and edges: Each enterprise is a node in the graph, and the number of nodes is N; the edges between enterprises represent the relationship between enterprises;

[0047] The types of relationships between enterprises include positive relationships, negative relationships, and neutral relationships;

[0048] The positive relationship refers to enterprises with cooperative and joint investment relationships; the negative relationship refers to enterprises with competitive relationships; the neutral relationship refers to those that are neither positive nor negative, and the relationship strength between two enterprises in a neutral relationship is 0;

[0049] Calculate the relationship strength of each enterprise based on the relationship type between them, and fill the relationship strength into the weighted relationship matrix to obtain a constructed weighted relationship matrix;

[0050] The relationship strength of the positive relationship is calculated based on the following formula:

[0051] w ij =α·f ij ;

[0052]

[0053] Where: α represents the scaling factor, which is used to adjust the scale of the weight; f ij represents the cooperation intensity function value; T ij represents the cooperation duration between enterprise i and enterprise j; T max V represents the maximum cooperation duration between all enterprise pairs; ij represents the scale of cooperation between enterprise i and enterprise j; V max represents the maximum cooperation scale between all enterprise pairs; R ij represents the degree of shared resources between enterprises i and j; R max represents the maximum value of resource sharing between all enterprise pairs; α1, α2, and α3 represent adjustable parameters, which are used to adjust the weights of the impact of duration, scale, and resource sharing on the intensity of cooperation;

[0054] The relationship strength of the negative relationship is calculated based on the following formula:

[0055] w ij =―β·g ij ;

[0056]

[0057] Where: S ij represents the overlap of market shares held by enterprise i and enterprise j in the same market; S max represents the maximum value of market share overlap between all pairs of companies; P ijrepresents the similarity between the products or services of enterprise i and enterprise j; P max It represents the maximum value of the similarity of products or services between all enterprise pairs; Q ij represents the intensity of price competition between enterprises i and j; Q max It represents the maximum value of price competition among all pairs of enterprises; β1, β2 and β3 are all adjustable adoption numbers, which are used to adjust the weights of the impact of market share overlap, product or service similarity and price competition on competition intensity.

[0058] The beneficial effects of the present invention include:

[0059] The present invention obtains and cleans the real-time financial data, historical financial data and external economic data of the enterprise, and adopts standardized processing to ensure the consistency of the data; secondly, by calculating the financial ratio and extracting the historical trend characteristics, the changing trend of the enterprise's financial status is fully reflected, avoiding the one-sidedness of a single financial indicator; on this basis, the autoencoder is used to compress the high-dimensional and complex financial data into a low-dimensional feature vector, extract the potential financial health information, reduce the information redundancy, and improve the analysis efficiency; combining the low-dimensional feature vector, external economic data and enterprise relationship diagram, K-means clustering is used to preliminarily classify the enterprises, and the clustering results are further optimized according to the relationship between the enterprises, so that the final classification is more accurate and reasonable. Based on this, the present invention can not only respond to the changes in the market environment in a timely manner, but also dynamically update the clustering grouping according to the real-time data, and provide an accurate and dynamic evaluation system for the financial health status of the enterprise, thereby overcoming the problem that the traditional static analysis method cannot adapt to the changing market environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0061] Figure 1 An overall step block diagram provided for an embodiment of the present invention.

[0062] Figure 2 A schematic diagram showing the structure of an autoencoder provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0063] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0064] See also Figure 1 As shown, the automatic assessment method of financial health status based on cluster analysis includes the following steps:

[0065] Step 1: Obtain external economic data, real-time financial data of the enterprise, and historical financial data, clean the acquired data, and standardize the cleaned data;

[0066] The external economic data include macroeconomic data, industry index, interest rate, exchange rate and other information;

[0067] The real-time financial data of the enterprise includes cash flow statement, income statement, balance sheet, etc.;

[0068] The historical financial data includes the company's financial statement data for the past few years, such as the balance sheet, income statement, cash flow statement, etc. for the past three or five years.

[0069] In this embodiment, data cleaning includes missing value processing and outlier detection; wherein the missing value processing is filled by mean filling; a statistical method is used to identify and process outliers, and outliers are replaced with values ​​within a reasonable range, or outliers are deleted; the technical means for the above data cleaning are conventional technical means in the field, and therefore will not be described in detail in this embodiment.

[0070] The standardization process adopts Z-score standardization to eliminate the dimensional differences between different financial data dimensions and external economic data.

[0071] Step 2: Calculate financial ratios based on the standardized real-time financial data and historical financial data of the enterprise, and extract historical trend features based on the standardized historical financial data;

[0072] The financial ratio calculations include current ratio, current assets, current load, quick ratio, debt-to-asset ratio, total liabilities, total assets, gross profit margin, net profit margin, and accounts receivable turnover;

[0073] Current Ratio: Where CA represents current assets; CL represents current liabilities;

[0074] Quick Ratio Among them, Inventory means inventory;

[0075] Asset Load Ratio Where TL represents total liabilities; TA represents total assets;

[0076] Total load TL = CL + NCL, where NCL represents non-current liabilities;

[0077] Total assets TA = CA + NCA, where NCA represents non-current assets;

[0078] Gross profit margin Sales Revenue means sales revenue; COGS means cost of goods sold;

[0079] Net Profit Margin Where Net Profit means net profit;

[0080] Accounts receivable turnover Where Net Sales represents net sales revenue; Average AR represents the average accounts receivable balance; AR beginning Indicates the beginning balance of accounts receivable; AR end It represents the balance of accounts receivable at the end of the period;

[0081] Current assets CA = Cash + AR + Inventory + Prepaid Expenses + ... Wherein Cash represents cash; AR represents accounts receivable; Prepaid Expenses represents prepaid expenses; ... represents other unlisted current assets; the definition of current assets is common knowledge in this field, so how to calculate current assets is a conventional technical means in this field;

[0082] The definition of current liabilities CL is also common knowledge in the field, so the specific calculation formula of current liabilities belongs to the conventional technical means in the field.

[0083] The historical trend feature extraction comprises the following steps:

[0084] Cyclical component decomposition: STL decomposition is used to decompose the standardized historical financial data into trend components, seasonal components and cyclical fluctuation components;

[0085] The exemplary specific steps are as follows:

[0086] Preliminary smoothing: Let the original time series data Y t , Loess regression is used to smooth the data and estimate the trend component. For each time point t, a weighting function is used to determine its relationship with the domain data point. The weighting function is:

[0087]

[0088] Where: w t,i represents the weight of the t-th time point to the i-th time point; h t Indicates the width of the smoothing window at time point t, and controls the size of the domain;

[0089] The width of the smoothing window is dynamically adjusted based on the local volatility:

[0090]

[0091] h t =α·σ t ;

[0092] Where: t represents the local volatility at time point t; k represents the local window size; μ t represents the local mean; Y i represents the data point at the i-th time point in the time series; α represents the adjustment factor;

[0093] For each point t, a local regression fit is performed based on the surrounding weighted data points to obtain the local trend:

[0094]

[0095] Where: T t represents the trend component at time t; P k (Y i ) indicates Y i polynomial transformation; k represents the order of the polynomial; Y i Represents the data point at time i in the time series;

[0096] In this embodiment, preliminary smoothing (such as Loess regression and STL decomposition) makes the data more stable and reliable by removing short-term fluctuations, noise and seasonal components. It plays an important role in the subsequent use of the HP filter. Since the HP filter relies on the stability of the data to accurately extract long-term trends, preliminary smoothing reduces the HP filter's misjudgment of short-term fluctuations, enabling it to more accurately decompose long-term trends and cyclical fluctuations, thereby enhancing the accuracy and predictability of trend analysis.

[0097] Seasonal component extraction: From the original time series data Y t The trend component is removed from the equation to obtain the residual of seasonal fluctuations, which includes seasonal fluctuations and cyclical fluctuations.

[0098] Based on the residuals obtained, modeling is performed with monthly and quarterly cycles respectively. An independent Loess regression model is used to extract seasonal components in each cycle, and then the seasonal components are obtained by weighted merging:

[0099]

[0100] Where: λ1 and λ2 represent weight factors; represents the seasonal component extracted with a monthly period; represents the seasonal component extracted with a quarterly cycle; S t represents the seasonal component after weighted merger;

[0101] Extracting periodic components: From the original time series data Y t After removing the trend and seasonal components, we get the cyclical fluctuation component C t ;

[0102] Apply the Hodrick-Prescott filter: HP filter the trend component based on the Hodrick-Prescott filter to obtain the smoothed long-term trend; the details are as follows:

[0103] Decompose the trend component T from STL t As the initial data of the HP filter, the HP filter is applied to the trend component and the objective function is constructed for smoothing:

[0104]

[0105] Where: Y t represents the original time series data; T t represents the trend component; λ represents the smoothing parameter; T represents the length of the time series; (T t+1 ―T t )―(T t ―T t―1 ) represents the acceleration of trend change, and the second-order difference term is used to penalize drastic changes in trend;

[0106] The exemplary solution process of the HP filter is as follows:

[0107] First, convert the objective function into matrix form as follows:

[0108]

[0109] Where: Y represents the T×1 original data column vector; T represents the T×1 trend component column vector; W Y represents the weight matrix; λ represents the smoothing parameter (selected based on experience or cross-validation); L represents the difference matrix;

[0110] Solve the optimization problem: The fitting part of the objective function can be expressed as (Y-T) T W Y (Y―T), is a standard least squares problem; T T L T The LT penalty term forces the second-order difference of the trend component to be smoothed, thereby constraining the rate of change of the trend; the objective function is converted to the following quadratic form:

[0111]

[0112] A=W Y +L T L;

[0113] b=W Y Y;

[0114] Where: b represents the relationship between the weighted original data and the trend; A represents the coefficient matrix; ·T represents the transpose of the corresponding item;

[0115] Solve the above optimization problem to obtain the smoothed trend component t;

[0116] Growth rate characteristic calculation: Calculate annual growth rate and compound annual growth rate based on long-term trends:

[0117]

[0118] Where: Growth Rate t represents the annual growth rate in year t; T t represents the long-term trend in year t; T t―1 represents the long-term trend in year t-1;

[0119]

[0120] Where: T T represents the final trend value; T0 represents the initial value; T represents the number of years;

[0121] Volatility characteristics: Volatility characteristics are extracted based on periodic volatility components, where volatility characteristics include standard deviation and volatility;

[0122] Calculate the standard deviation:

[0123]

[0124] Where: C Represents the periodic fluctuation component C t The standard deviation of C represents the mean of the periodic fluctuation component; T represents the total number of time points of the periodic fluctuation component;

[0125] Calculate annualized volatility:

[0126]

[0127] Where: Annualized Volatility represents annualized volatility; N represents the number of data points per year;

[0128] Volatility Kurtosis:

[0129]

[0130] Where: Kurtosis represents the kurtosis of the cyclical fluctuation component;

[0131] Calculate the volatility index:

[0132]

[0133] Where: V t represents the volatility index at time point t; C i represents the value of the periodic fluctuation component; k represents the weighted window size;

[0134] Periodic features: Periodic features are extracted based on seasonal components, where periodic features include period length, seasonal fluctuation amplitude, and seasonal peak and valley values.

[0135] By performing periodic analysis on the seasonal component, the period length is extracted based on the autocorrelation function:

[0136]

[0137] Where: k represents the seasonal component S t Autocorrelation coefficient at lag k; S t represents the seasonal component; μ S represents the mean of the seasonal component;

[0138] Find the lag value k corresponding to the maximum autocorrelation coefficient as the period length of the periodic fluctuation;

[0139] Amplitude of seasonal fluctuation: based on the difference between the maximum and minimum values ​​of seasonal fluctuation;

[0140] Seasonal peaks and valleys are the maximum and minimum values ​​of seasonal fluctuations.

[0141] In this embodiment, the purpose of combining STL decomposition with HP filter is to further remove the cyclical fluctuation components through HP filter; because in many financial data, seasonal factors have a significant impact on company performance; for example, retail sales usually rise sharply during holiday seasons, and travel companies may be affected by demand in different seasons. The STL method can effectively decompose the seasonal components in the time series; unlike traditional seasonal adjustment methods, STL can handle complex nonlinear seasonal patterns in data.

[0142] However, after obtaining the trend component through STL decomposition, there may still be cyclical fluctuation components in the data. These fluctuations are usually a reflection of macroeconomic cycles, industry cycles or other external cyclical factors; HP filters can help extract cyclical fluctuations from the data and remove the impact of these cyclical fluctuations on the trend.

[0143] STL extracts seasonal components and local trends of data through non-parametric methods. In practical applications, many changes in financial data are strongly affected by seasonal factors, such as monthly income, quarterly profits, etc. The original intention of the HP filter is to extract long-term trends from time series. Therefore, if the HP filter is directly applied, seasonal components and cyclical components may be mixed into the long-term trend, thereby affecting the accuracy of trend extraction. Therefore, in this embodiment, by first using STL to decompose the seasonal components and then performing HP filtering, it can be ensured that the HP filter only processes the remaining cyclical components (i.e., the part without seasonal fluctuations), thereby improving the accuracy of the HP filter in extracting long-term trends.

[0144] Step 3: The standardized real-time financial data, historical financial data, extracted historical trend features, and financial ratio calculation results are used as inputs of the constructed autoencoder, and the input data is compressed into a low-dimensional feature vector based on the autoencoder;

[0145] See also Figure 2 As shown, the steps of compressing the input data into a low-dimensional feature vector based on the autoencoder are as follows:

[0146] Based on the bidirectional LSTM layer, the historical financial data and historical trend features are processed to obtain the hidden states of the historical financial data and historical trend features;

[0147] Suppose that the historical data X=X1,X2,…,X T , where each X t represents the feature vector at time step t;

[0148] The bidirectional LSTM layer generates hidden states in two directions, from front to back and from back to front; assuming represents the hidden state of the forward LSTM at time t; represents the hidden state of the reverse LSTM;

[0149]

[0150] The final bidirectional LSTM hidden state represents the concatenation of the forward and backward hidden states:

[0151]

[0152] Process the financial ratio calculation results and real-time financial data based on the fully connected layer to obtain the hidden state of the financial ratio calculation results and the real-time financial data;

[0153] Assume that the result of the financial ratio calculation is R = [r1, r2, ..., r M ], where M represents the number of ratios; real-time financial data is R = [d1, d2, …, d P ], where P represents the dimension of the data;

[0154] Each input r i and d j By weight W r and W d Perform a linear transformation and then obtain the hidden state through an activation function (such as ReLU);

[0155] h r =ReLU(W r R+b r );

[0156] h d =ReLU(W d R+b d );

[0157] Where: b r and b d All represent bias terms;

[0158] The historical financial data, historical trend features, financial ratio calculation results and the hidden state of real-time financial data are spliced ​​to obtain the spliced ​​feature h concat =[h t ,h r ,h d ];

[0159] The self-attention mechanism is used to perform weighted summation on the concatenated features to generate a refined feature representation. The refined feature representation is mapped to a low-dimensional latent space through a fully connected layer to obtain a low-dimensional feature vector as the output of the autoencoder.

[0160] Calculate the attention weight of each feature, assuming h concat is a feature vector of length L, where each h i Represents the concatenated features, and the self-attention mechanism calculates the attention weight a of each feature i :

[0161]

[0162] Where: score(h i ,Q,K) represents the feature h iThe similarity between the query vector and the key vector is calculated using the dot product:

[0163]

[0164] Where: Represents the feature h i The transpose of

[0165] The weighted feature representation is obtained by weighted summation:

[0166]

[0167] Finally, the weighted feature h att Through the fully connected layer, it is mapped to a low-dimensional latent space z to obtain a low-dimensional feature vector.

[0168] Step 4: Combining the low-dimensional feature vector, external economic data and enterprise relationship diagram, first use K-means clustering to preliminarily classify the enterprises, and then further adjust the clustering results through the enterprise relationship diagram to obtain the adjusted clustering results;

[0169] The specific steps of constructing the relationship graph are as follows:

[0170] Definition of nodes and edges: Each enterprise is a node in the graph, and the number of nodes is N; the edges between enterprises represent the relationship between enterprises;

[0171] The types of relationships between enterprises include positive relationships, negative relationships, and neutral relationships;

[0172] The positive relationship refers to enterprises with cooperative and joint investment relationships; the negative relationship refers to enterprises with competitive relationships; the neutral relationship refers to those that are neither positive nor negative, and the relationship strength between two enterprises in a neutral relationship is 0;

[0173] Calculate the relationship strength of each enterprise based on the relationship type between them, and fill the relationship strength into the weighted relationship matrix to obtain a constructed weighted relationship matrix;

[0174] The relationship strength of the positive relationship is calculated based on the following formula:

[0175] w ij =α·f ij ;

[0176]

[0177] Where: α represents the scaling factor, which is used to adjust the scale of the weight; f ij represents the cooperation intensity function value; T ij represents the cooperation duration between enterprise i and enterprise j; Tmax V represents the maximum cooperation duration between all enterprise pairs; ij represents the scale of cooperation between enterprise i and enterprise j; V max represents the maximum cooperation scale between all enterprise pairs; R ij represents the degree of shared resources between enterprises i and j; R max represents the maximum value of resource sharing between all enterprise pairs; α1, α2, and α3 represent adjustable parameters, which are used to adjust the weights of the impact of duration, scale, and resource sharing on the intensity of cooperation;

[0178] The relationship strength of the negative relationship is calculated based on the following formula:

[0179] w ij =―β·g ij ;

[0180]

[0181] Where: S ij represents the overlap of market shares held by enterprise i and enterprise j in the same market; S max represents the maximum value of market share overlap between all pairs of companies; P ij represents the similarity between the products or services of enterprise i and enterprise j; P max It represents the maximum value of the similarity of products or services between all enterprise pairs; Q ij represents the intensity of price competition between enterprises i and j; Q max It represents the maximum value of price competition among all pairs of enterprises; β1, β2 and β3 are all adjustable adoption numbers, which are used to adjust the weights of the impact of market share overlap, product or service similarity and price competition on competition intensity.

[0182] In actual enterprise data, there may be some data sparsity, such as incomplete characteristic information of some enterprises, or lack of specific economic data. The use of enterprise relationship graphs can make up for the deficiencies of these data; for example, through adjacency propagation, some information may be filled in by similarities with neighboring enterprises, thereby solving the problem of data sparsity; and because in K-means clustering, isolated points may be misclassified into a certain cluster, resulting in inaccurate clustering results, by introducing the mutual connections in the enterprise relationship graph, the impact of isolated points can be effectively reduced. The relationship graph can capture the correlation between enterprises, help identify real isolated points, and avoid them affecting the clustering results.

[0183] The steps of using K-means clustering to preliminarily classify enterprises are as follows:

[0184] Initialize cluster centers: randomly select K enterprises as initial cluster centers, and the dimension of the cluster centers is dlatent +d enconomy , which combines low-dimensional feature vectors with external economic data;

[0185] Calculate the distance: For each enterprise i, the low-dimensional feature vector h i and external economic data i Make up a d latent +d enconomy dimensional feature vector, and then calculate the Euclidean distance between each enterprise and all cluster centers:

[0186]

[0187] Where: d(h i ,c k ) represents the Euclidean distance between the enterprise and the cluster center; Represents the spatial coordinates of the cluster center in the low-dimensional feature vector; Represents the spatial coordinates of the cluster center in the external economic data;

[0188] Assign to the nearest cluster center: Each enterprise i is assigned to the cluster center k closest to it;

[0189] Update cluster centers: After clustering is completed, update the center c of each cluster k is the average characteristic of all enterprises belonging to this cluster:

[0190]

[0191] Where: S k represents all enterprises in cluster k; |S k | represents the size of cluster k;

[0192] By taking the average vector of all enterprises in each cluster as the new cluster center, repeating the steps of calculating the distance to update the cluster center until the predetermined number of iterations is reached, a preliminary clustering result is obtained; in the clustering result, we can assign multiple labels, such as financially sound enterprises, growth enterprises, troubled enterprises, etc.; through clustering, we can preliminarily obtain the label category to which each enterprise belongs.

[0193] The specific steps to further adjust the clustering results through the enterprise relationship diagram are as follows:

[0194] Relationship diagram construction: construct a weighted relationship matrix through the relationships between enterprises;

[0195] Propagation adjustment: Update the labels of each enterprise obtained through preliminary clustering based on the weighted propagation model, that is, map the labels obtained through preliminary clustering to digital representations;

[0196] Convergence judgment: If the value of the label obtained in this round of iteration minus the label obtained in the previous round of iteration is less than the preset threshold, the iteration is stopped; or the iteration is stopped after the preset number of iterations is reached. After the iteration is completed, the category label of the enterprise is used for the final clustering result of each enterprise.

[0197] The weighted propagation model is as follows:

[0198]

[0199] Where: represents the category label of enterprise i in the t+1th round of dissemination; represents the preliminary clustering label of enterprise i; w ij represents the strength of the relationship between enterprise i and enterprise j; α represents the relative importance parameter for controlling the initial clustering labels and the propagation of the relationship network; j represents all enterprises related to enterprise i;

[0200] When the final calculation result is obtained, the calculated value is converted to the closest positive integer label. For example, the calculated value is 1.6; if the mapping label value of a growth enterprise is 2, the adjusted enterprise category label is a growth enterprise.

[0201] In this embodiment, the adjustment based on the relationship graph realizes local update through the weighted propagation model. There is no need to fully re-cluster the entire data set, but only to fine-tune the clustering results, thereby improving the efficiency of clustering. The computational amount of each iterative propagation is relatively small, avoiding the high computational cost of global optimization, and is particularly suitable for processing real-time and large-scale data sets.

[0202] Step 5: Based on the adjusted clustering results, each company is assigned to a category, each category representing a different financial health status; the cluster grouping is then updated based on new real-time data.

[0203] In this embodiment, external economic data (such as macroeconomic indicators, industry trends, etc.) are incorporated into the clustering process, further enriching the dimension of clustering. External economic factors have an important impact on the business performance of enterprises, and can help distinguish enterprises that are more affected by the economic environment from other enterprises, thereby improving the recognition and accuracy of clustering; by utilizing information such as cooperation and competition relationships between enterprises, the relationship diagram provides an external constraint based on the interaction between enterprises. This makes clustering not only rely on the internal characteristics of enterprises, but also integrates actual business relationships, ensuring that the clustering results are more in line with the distribution of enterprises in the real business environment.

[0204] Through preliminary clustering with K-means, enterprises can be quickly divided into different categories. The K-means algorithm has good local optimality and is suitable for processing large-scale data sets. However, K-means may be affected by the selection of initial clustering centers, resulting in unstable clustering results. By combining external economic data and low-dimensional feature vectors, preliminary clustering can more accurately reflect the actual differences between enterprises.

[0205] Based on the preliminary clustering results, using the enterprise relationship graph to adjust the clustering results helps to improve the stability and consistency of the clustering. The strength of the relationship between enterprises (cooperation or competition) can provide additional information about the enterprise category; even if there are some deviations in the preliminary clustering, the propagation adjustment based on the relationship graph can correct the error through information transmission, thereby effectively reducing the noise and misclassification in the clustering.

[0206] The above are only preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the protection scope of the present application.

Claims

1. An automatic financial health assessment method based on cluster analysis, characterized in that: The following steps are involved: Step 1: Obtain external economic data, real-time financial data of the enterprise, and historical financial data, clean the acquired data, and standardize the cleaned data; Step 2: Calculate financial ratios based on the standardized real-time financial data and historical financial data of the enterprise, and extract historical trend features based on the standardized historical financial data; Step 3: The standardized real-time financial data, historical financial data, extracted historical trend features, and financial ratio calculation results are used as inputs of the constructed autoencoder, and the input data is compressed into a low-dimensional feature vector based on the autoencoder; Step 4: Combining the low-dimensional feature vector, external economic data and enterprise relationship diagram, first use K-means clustering to preliminarily classify the enterprises, and then further adjust the clustering results through the enterprise relationship diagram to obtain the adjusted clustering results; Step 5: Based on the adjusted clustering results, each enterprise is assigned to a category, each category representing a different financial health status; Then update the clustering grouping based on new real-time data.

2. The automatic financial health assessment method based on cluster analysis according to claim 1 is characterized in that: The financial ratio calculations include current ratio, current assets, current load, quick ratio, debt-to-asset ratio, total liabilities, total assets, gross profit margin, net profit margin, and accounts receivable turnover.

3. The automatic financial health assessment method based on cluster analysis according to claim 1 is characterized in that: The historical trend feature extraction comprises the following steps: Cyclical component decomposition: STL decomposition is used to decompose the standardized historical financial data into trend components, seasonal components and cyclical fluctuation components; Apply the Hodrick-Prescott filter: HP filter the trend component based on the Hodrick-Prescott filter to obtain the smoothed long-term trend; Growth rate characteristic calculation: Calculate annual growth rate and compound annual growth rate based on long-term trends: Where: Growth Rate t represents the annual growth rate in year t; T t represents the long-term trend in year t; T t―1 represents the long-term trend in year t-1; Where: T T represents the final trend value; T0 represents the initial value; T represents the number of years; Volatility characteristics: Volatility characteristics are extracted based on periodic volatility components, where volatility characteristics include standard deviation and volatility; Periodic features: Periodic features are extracted based on seasonal components, where periodic features include period length, seasonal fluctuation amplitude, and seasonal peak and valley values.

4. The automatic financial health status assessment method based on cluster analysis according to claim 1 is characterized in that: The steps of compressing the input data into a low-dimensional feature vector based on the autoencoder are as follows: Based on the bidirectional LSTM layer, the historical financial data and historical trend features are processed to obtain the hidden states of the historical financial data and historical trend features; Process the financial ratio calculation results and real-time financial data based on the fully connected layer to obtain the hidden state of the financial ratio calculation results and the real-time financial data; Splicing historical financial data, historical trend features, financial ratio calculation results, and hidden states of real-time financial data to obtain spliced ​​features; The self-attention mechanism is used to perform weighted summation on the concatenated features to generate a refined feature representation; the refined feature representation is mapped to a low-dimensional latent space through a fully connected layer to obtain a low-dimensional feature vector as the output of the autoencoder.

5. The automatic financial health assessment method based on cluster analysis according to claim 1 is characterized in that: The steps of using K-means clustering to preliminarily classify enterprises are as follows: Initialize cluster centers: randomly select K enterprises as initial cluster centers, and the dimension of the cluster centers is d latent +d enconmy , which combines low-dimensional feature vectors with external economic data; Calculate the distance: For each enterprise i, the low-dimensional feature vector h i and external economic data i Make up a d latent +d enconomy dimensional feature vector, and then calculate the Euclidean distance between each enterprise and all cluster centers: Where: d(h i ,c k ) represents the Euclidean distance between the enterprise and the cluster center; Represents the spatial coordinates of the cluster center in the low-dimensional feature vector; Represents the spatial coordinates of the cluster center in the external economic data; Assign to the nearest cluster center: Each enterprise i is assigned to the cluster center k closest to it; Update cluster centers: After clustering is completed, update the center c of each cluster k is the average characteristic of all enterprises belonging to this cluster: Where: S k represents all enterprises in cluster k; |S k | represents the size of cluster k; By taking the average vector of all enterprises in each cluster as the new cluster center, the steps of calculating the distance to updating the cluster center are repeated until the predetermined number of iterations is reached to obtain a preliminary clustering result.

6. The automatic financial health assessment method based on cluster analysis according to claim 1 is characterized in that: The specific steps to further adjust the clustering results through the enterprise relationship diagram are as follows: Relationship diagram construction: construct a weighted relationship matrix through the relationships between enterprises; Propagation adjustment: Update the labels of each enterprise obtained after preliminary clustering based on the weighted propagation model; Convergence judgment: If the value of the label obtained in this round of iteration minus the label obtained in the previous round of iteration is less than the preset threshold, the iteration is stopped; or the iteration is stopped after the preset number of iterations is reached. After the iteration is completed, the category label of the enterprise is used for the final clustering result of each enterprise.

7. The automatic financial health assessment method based on cluster analysis according to claim 6 is characterized in that: The weighted propagation model is as follows: Where: represents the category label of enterprise i in the t+1th round of dissemination; represents the preliminary clustering label of enterprise i; w ij represents the strength of the relationship between firm i and firm j; α represents the relative importance parameter controlling the propagation of preliminary clustering labels and relational networks; j represents all enterprises related to enterprise i.

8. The method for automatically assessing financial health status based on cluster analysis according to claim 6, characterized in that: The specific steps of constructing the relationship graph are as follows: Definition of nodes and edges: Each enterprise is a node in the graph, and the number of nodes is N; the edges between enterprises represent the relationship between enterprises; The types of relationships between enterprises include positive relationships, negative relationships, and neutral relationships; The positive relationship refers to enterprises with cooperative and joint investment relationships; the negative relationship refers to enterprises with competitive relationships; The neutral relationship is a relationship that is neither positive nor negative, and the relationship strength between two enterprises in a neutral relationship is 0; Calculate the relationship strength of each enterprise based on the relationship type between them, and fill the relationship strength into the weighted relationship matrix to obtain a constructed weighted relationship matrix; The relationship strength of the positive relationship is calculated based on the following formula: w ij =α·f ij ; Where: α represents the scaling factor, which is used to adjust the scale of the weight; f ij represents the cooperation intensity function value; T ij represents the cooperation duration between enterprise i and enterprise j; T max represents the maximum cooperation duration between all enterprise pairs; V ij represents the scale of cooperation between enterprise i and enterprise j; V max represents the maximum cooperation scale between all enterprise pairs; R ij represents the degree of shared resources between enterprises i and j; R max represents the maximum value of resource sharing between all enterprise pairs; α1, α2, and α3 represent adjustable parameters, which are used to adjust the weights of the impact of duration, scale, and resource sharing on the intensity of cooperation; The relationship strength of the negative relationship is calculated based on the following formula: w ij =―β·g ij ; Where: S ij represents the overlap of market shares held by enterprise i and enterprise j in the same market; S max represents the maximum value of market share overlap between all pairs of companies; P ij represents the similarity between the products or services of enterprise i and enterprise j; P max It represents the maximum value of the similarity of products or services between all enterprise pairs; Q ij represents the intensity of price competition between enterprises i and j; Q max It represents the maximum value of price competition among all pairs of enterprises; β1, β2 and β3 are all adjustable adoption numbers, which are used to adjust the weights of the impact of market share overlap, product or service similarity and price competition on competition intensity.

Citation Information

Cited By

  • Geographic information-based land reclamation planning surveying and mapping collaborative operation method

    CN120278681A