Listed company performance early warning method based on support vector machine
Through the performance warning method of listed companies based on support vector machines, the problems of high data quality requirements and complex feature selection in the existing technology are solved, and a more efficient and stable performance warning effect is achieved.
Patent Information
- Application Number
- CN202510241390.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art requires high data quality in the performance warning of listed companies, is susceptible to noise, inconsistency and missing values, and has complex feature selection, which makes it easy to overfit problems.
The performance warning method of listed companies based on support vector machines is adopted, and the support vector machine model is optimized by obtaining historical transaction data and basic factor data, a factor database is constructed, outliers, missing values are eliminated, missing values are filled, and the support vector machine model is standardized. Cross-verification combined with grid search is used to optimize the support vector machine model.
It improves computing efficiency, reduces the risk of overfitting the model, enhances the robustness and reliability of the model, and can more accurately warn of the performance of listed companies.
Smart Images

Figure BDA0005294255540000031 
Figure BDA0005294255540000071 
Figure FDA0005294255530000021
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of prospect evaluation of listed companies, and in particular to a performance early warning method for listed companies based on support vector machines. Background Art
[0002] The performance of listed companies is the focus of the current trading market. The data disclosed by listed companies and market fundamentals are the basis for their performance warnings. Performance warnings need to extract reliable information from a large amount of data to judge the company's future performance and risks. At present, there are machine learning based on multi-factor models for early warning of listed companies, but machine learning models have high requirements for data quality. Noise, inconsistency and missing values in the data will seriously affect the performance of the model. At the same time, the selection and construction of factors related to listed companies are key steps. Machine learning models need to select the most useful features from a large number of factors, which requires complex feature selection and engineering processes. If the features are not selected properly, the performance of the model may be degraded and overfitting may occur. Summary of the invention
[0003] The purpose of the present invention is to provide a performance early warning method for listed companies based on support vector machines. The present invention can simply and efficiently process listed companies and factors related to listed companies, and can use support vector machines as a data processing model to improve calculation efficiency and reduce the risk of overfitting of the model.
[0004] The technical solution of the present invention is a method for early warning of listed company performance based on support vector machine, comprising the following steps:
[0005] Step 1: Obtain the historical transaction data and basic factor data of listed companies, and build a factor database based on the collected historical transaction data and basic factor data;
[0006] Step 2: Eliminate outliers and extreme values in the factor database, fill in missing values, and then standardize the factor database. Then screen the standardized factor database to obtain the principal component factors related to the performance of listed companies, and merge the principal component factors into a feature matrix.
[0007] Step 3: Build a support vector machine model, use cross-validation combined with grid search to optimize the parameters of the support vector machine model, and train the support vector machine model with the feature matrix in step 2;
[0008] Step 4: Use the trained support vector machine model to provide early warning on the performance of listed companies.
[0009] In the above-mentioned performance early warning method for listed companies based on support vector machines, in step 1, the historical transaction data includes company financial data, industry indicator data, and macroeconomic data; the basic factor data consists of fundamental factors and technical factors, where the fundamental factors cover economic, policy, and industry factors; the technical factors include technological breakthroughs, intellectual property rights, and industry technology proportion factors; the factor database includes technical factors, valuation factors, profitability factors, growth ability factors, operation ability factors, and solvency factors.
[0010] In the aforementioned performance early warning method for listed companies based on support vector machines, in step 2, the method for filling missing values is to fill the missing values of the previous trading day with the data of the next trading day by traversing forward.
[0011] In the aforementioned performance early warning method for listed companies based on support vector machines, in step 2, the screening of the factors is carried out by using the IC analysis method and the IR analysis method;
[0012] The calculation formula of the IC analysis method is as follows:
[0013] IC = corr(factor t , r t+1 );
[0014] In the formula: factor t represents the specific factor value in the current t period, r t+1 represents the next period's return rate; corr represents the Pearson correlation coefficient;
[0015] The calculation formula of the IR analysis method is as follows:
[0016]
[0017] In the formula: mean represents the multi-period mean of the IC value, and std represents the multi-period standard deviation of the IC value.
[0018] In the aforementioned performance early warning method for listed companies based on support vector machines, in step 3, cross-validation combined with grid search is used to optimize the parameters of the support vector machine model, and the optimization process is as follows:
[0019] Step 3.1: Set the SVM hyperparameters to be tuned and their possible value ranges;
[0020] Step 3.2: Combine the different values of all parameters into a grid to form all possible parameter combinations;
[0021] Step 3.3: Use K-fold cross-validation to evaluate the model for each parameter combination to ensure that the evaluation results have high robustness;
[0022] Step 3.4: Train the SVM model within each fold of cross-validation and calculate the performance metrics of the validation set;
[0023] Step 3.5: Compare the average cross-validation performance of all parameter combinations and select the parameter configuration with the best performance;
[0024] Step 3.6: Retrain the SVM model on all the training data using the optimal parameters.
[0025] Compared with the prior art, the present invention has the following beneficial effects:
[0026] (1) The present invention uses cross-validation to divide the dataset into multiple subsets, and trains and validates with different subsets in turn. This can effectively prevent the model from overfitting on the training set and ensure that the model also has good performance on unseen data. The cross-validation of the present invention uses the average value of multiple training and validation results, which can reduce the accidental error caused by data splitting, thereby improving the robustness and reliability of the model. And cross-validation evaluates the model performance through multiple trainings and validations, and the results are relatively stable and not easily affected by the randomness of data splitting. The grid search method of the present invention can traverse the predefined parameter combinations, and combined with cross-validation, it can more accurately evaluate the performance of each group of parameters. Through this method, the parameter combination that makes the model performance optimal can be found. In addition, cross-validation and grid search can run fully automatically without human intervention. This makes it more efficient and convenient in large-scale parameter tuning.
[0027] (2) The present invention selects the support vector machine model, which can handle high-dimensional data and cope with the complexity of factors. At the same time, cross-validation combined with grid search can help optimize the weights and combinations of these factors, thereby improving the overall performance of the strategy, and can automatically update the corresponding model parameters with different data to achieve automatic update. Specific Embodiments
[0028] The present invention will be further described below in conjunction with embodiments, but it shall not be used as a basis for limiting the present invention.
[0029] Embodiment: A method for predicting the performance of listed companies based on support vector machines, comprising the following steps:
[0030] Step 1: Obtain the historical trading data and basic factor data of listed companies, and construct a factor database based on the collected historical trading data and basic factor data;
[0031] In this step, the historical trading data and basic factor data of listed companies in the stock market are obtained by using network data sources or professional data providers.
[0032] The historical transaction data includes company financial data, industry indicator data, and macroeconomic data. Among them, the company financial data can be obtained from company announcements, the industry indicator data is collected through big data, and the macroeconomic data is officially announced.
[0033] The basic factor data includes fundamental factors and technical factors; the fundamental factors include economic factors, policy factors, and industry factors; the technical factors include technology breakthrough factors, intellectual property factors, and industry technology proportion factors.
[0034] Based on the above historical transaction data and basic factor data, the constructed factor database includes technical factors, valuation factors, profitability factors, growth ability factors, operation ability factors, and solvency factors.
[0035] In this step, the fundamental factors include major economic policies of the country, such as industrial, tax, and monetary policies, which will directly affect the performance of listed companies. The support or restriction of the development of specific industries by the state affects the performance trend of related companies. For example, the restriction of the development of an industry will lead to a decline in the performance of related companies. For example, the price limit on public utility products such as transportation, gas, and water and electricity will reduce the company's profit level. The change in market interest rates caused by the adjustment of monetary policy will also affect the performance trend. In terms of tax policy, the performance of companies enjoying tax reduction incentives usually rises, while the increase in personal income tax may lead to a decline in consumption, affecting the company's profit. These policies cause performance fluctuations by affecting the company's profit and market interest rates.
[0036] The industry factors can be analyzed from the perspective of products to determine whether the company's products are productive resources or consumer resources. Productive resources meet production needs, and consumer resources are directly consumed by people. Productive resources are greatly affected by the economic environment. When the economy is booming, the demand increases rapidly, and when the economy is sluggish, the demand drops rapidly. Consumer resources are less affected, and the market demand and price fluctuations of necessities and luxury goods are also different. In addition, by analyzing the demand pattern, the proportion of domestic and foreign sales of the company's products can be understood. Domestic sales are affected by the domestic economy and events, and foreign sales are affected by the international economy and trade environment. At the same time, the degree to which the company's products meet the needs of different demand objects should be analyzed. Different demand objects have different requirements for product performance and quality, and the company needs to customize products according to the demand to avoid a decline in sales, a reduction in profit, and a decline in performance.
[0037] Through the analysis of production patterns, it is possible to determine whether a company is labor-intensive, capital-intensive, or knowledge and technology-intensive, and assign weights to each factor. Labor-intensive enterprises rely on labor, capital-intensive enterprises rely on capital, and knowledge and technology-intensive enterprises focus on knowledge and technology. In less developed regions, there are more labor-intensive enterprises; in developed regions, capital-intensive enterprises have more advantages. With the development of new technologies, technology-intensive enterprises are gradually replacing capital-intensive enterprises. The labor productivity and competitiveness of different types of enterprises affect product sales and profitability, resulting in differences in performance returns.
[0038] Step 2: Eliminate outliers and extreme values in the factor database, fill in the missing values, then standardize the factor database, and then screen the standardized factor database to obtain the principal component factors related to the performance of listed companies, and combine the principal component factors into a feature matrix;
[0039] In this step, eliminate the sample data in the factor database where the outliers, extreme values, and missing values are greater than 30%, and then use the method of filling the missing values of the previous trading day with the data of the next trading day in a forward traversal manner. Finally, perform Z-score standardization on the factor data to ensure the accuracy and consistency of the data, and standardize these factors into a feature matrix.
[0040] In this step, the principal component factors of the factor database are screened through backtesting, that is, using historical data to evaluate the correlation and predictive ability between each factor and performance, and screening out the factors with better performance as the objects for subsequent analysis. The backtesting method is carried out using the IC analysis method and the IR analysis method;
[0041] The calculation formula of the IC analysis method is as follows:
[0042] IC = corr(factor t , r t+1 );
[0043] In the formula: factor t represents the specific factor value in the current t period, r t+1 represents the next period's return rate; corr represents the Pearson correlation coefficient;
[0044] The calculation formula of the IR analysis method is as follows:
[0045]
[0046] In the formula: mean represents the multi-period mean of the IC value, and std represents the multi-period standard deviation of the IC value.
[0047] In the above way, redundant information between factors can be reduced and the main influencing factors can be extracted, thereby reducing the model complexity and improving the prediction accuracy.
[0048] Step 3: Construct a support vector machine model, optimize the parameters of the support vector machine model using cross-validation combined with grid search, and train the support vector machine model with the feature matrix in Step 2.
[0049] In this step, the steps of optimizing the parameters of the support vector machine model using cross-validation combined with grid search are as follows:
[0050] Step 3.1: Select the parameter range of the support vector machine model, determine the parameters of the support vector machine model that need to be tuned, such as C (regularization parameter) and gamma (RBF kernel parameter), and their candidate value ranges.
[0051] Step 3.2: Combine different values of all parameters into a grid to form all possible parameter combinations.
[0052] Step 3.3: Use K-fold cross-validation (K = 5 or 10) to evaluate the model for each parameter combination to ensure that the evaluation results are highly robust.
[0053] Step 3.4: Train the SVM model within each fold of cross-validation and calculate the performance metrics of the validation set.
[0054] Step 3.5: Compare the average cross-validation performance of all parameter combinations and select the parameter configuration with the best performance.
[0055] Step 3.6: Retrain the SVM model on all training data using the optimal parameters.
[0056] In this step, grid search systematically explores the parameter space to ensure finding the optimal parameter combination, thereby improving the performance of the model; compared with manual parameter tuning, grid search covers a wider parameter range and avoids falling into local optimal solutions.
[0057] Cross-validation divides the dataset into multiple subsets, and different subsets are cyclically used for training and validation, effectively preventing model overfitting and improving the generalization ability of the model on unseen data. Through multiple trainings and validations, the evaluation bias caused by different data partitions is reduced, making the selection of model parameters more reliable.
[0058] Grid search automatically adjusts parameters, reducing the need for manual intervention and making the parameter optimization process more efficient and accurate. At the same time, parallel computing significantly reduces the computing time.
[0059] The support vector machine model is good at handling high-dimensional data. By combining grid search to optimize parameters, it can better handle complex data in fields such as quantitative investment. Cross-validation can handle different data distributions, ensuring the robustness of the model on various datasets. Through systematic parameter optimization, it provides a data-driven decision-making basis and avoids biases caused by human subjective judgment. The results of cross-validation and grid search are clear, making the parameter selection process transparent and interpretable, which helps with model deployment and maintenance.
[0060] Step 4: Use the trained support vector machine model to give early warnings about the performance of listed companies.
[0061] In this step, in practical applications, by setting up a real-time data interface, historical trading data and basic factor data of current listed companies are collected and processed to obtain a feature matrix, and then the feature matrix is input into the trained support vector machine model for performance output. Thus, early warnings can be given about the subsequent development of the company based on the output performance, which can reduce investment risks in the trading market.
[0062] In the present invention, the present invention mainly overcomes the defects in parameter selection and the need for a large amount of computing resources, as follows:
[0063] Advantages of parameter selection:
[0064] (1) Systematic search for optimal parameters: Grid search systematically searches all possible parameter combinations through a predefined parameter grid, ensuring that potential best parameter combinations are not missed.
[0065] (2) Determinacy: Compared with random search, grid search provides a deterministic optimization process and can accurately find the best parameters within a specific range.
[0066] (3) Stability of cross-validation: Cross-validation reduces the parameter selection bias caused by different data partitions through multiple trainings and validations, making the selected parameters more reliable and stable.
[0067] (4) No need for manual debugging: Grid search automatically adjusts parameters, reducing the need for manual intervention and making the parameter selection process more efficient and accurate.
[0068] Advantages of reducing computing resource requirements:
[0069] (1) Parallel computing to improve efficiency: Many modern machine learning libraries (such as Scikit-Learn) support parallel computing and can evaluate multiple parameter combinations simultaneously, greatly reducing the computing time.
[0070] (2) Gradually narrow the search scope: First, a rough grid search can be carried out to determine a better parameter range, and then a more refined search can be carried out. This phased search strategy can effectively reduce the computational amount.
[0071] (3) Hierarchical grid search: After initially screening out the parameter combinations with better performance, further refine the search scope of these parameters to improve the search efficiency.
[0072] (4) K-value optimization: Reasonably select the number of folds for cross-validation (such as 5 or 10 folds), which can reduce unnecessary computational overhead while ensuring the stability of model evaluation.
[0073] (5) Performance early termination: During the grid search process, if the performance of some parameter combinations is significantly poor, the evaluation of these combinations can be terminated early, and resources can be concentrated on evaluating potential high-quality parameter combinations.
[0074] In summary, the present invention can use the support vector machine model with optimized parameters to predict the enterprise bankruptcy risk and business crisis. The best model parameter combination is found through grid search, and at the same time, cross-validation is combined to select the characteristic factors for evaluating the performance of listed companies that have a significant impact on classification, reducing the complexity of model processing and improving the accuracy of enterprise crisis prediction.
Claims
1. A performance early warning method for listed companies based on support vector machines, comprising the following steps: Step 1: Obtain the historical transaction data and basic factor data of listed companies, and build a factor database based on the collected historical transaction data and basic factor data; Step 2: Eliminate outliers and extreme values in the factor database, fill in missing values, and then standardize the factor database. Then screen the standardized factor database to obtain the principal component factors related to the performance of listed companies, and merge the principal component factors into a feature matrix. Step 3: Build a support vector machine model, use cross-validation combined with grid search to optimize the parameters of the support vector machine model, and train the support vector machine model with the feature matrix in step 2; Step 4: Use the trained support vector machine model to provide early warning on the performance of listed companies.
2. The performance early warning method for listed companies based on support vector machines according to claim 1 is characterized by: In step 1, the historical transaction data includes company financial data, industry indicator data and macroeconomic data; the basic factor data is composed of basic factors and technical factors, wherein the basic factors cover economic, policy and industry factors; the technical factors include technological breakthroughs, intellectual property rights and industry technology share factors; the factor database includes technical factors, valuation factors, profitability factors, growth capacity factors, operating capacity factors and repayment capacity factors.
3. The performance early warning method for listed companies based on support vector machine according to claim 1 is characterized by: In step 2, the method of filling missing values is to fill the missing values of the data of the previous trading day with the data of the next trading day in a forward traversal manner.
4. The performance early warning method for listed companies based on support vector machine according to claim 1 is characterized by: In step 2, the screening of the factors is carried out by IC analysis and IR analysis; The calculation formula of the IC analysis method is as follows: IC=corr(factor t ,r t+1 ); Where: factor t represents the specific factor value of the current period t, r t+1 represents the next period rate of return; corr represents the Pearson correlation coefficient; The calculation formula of the IR analysis method is as follows: Where: mean represents the multi-period mean of the IC value, and std represents the multi-period standard deviation of the IC value.
5. The performance early warning method for listed companies based on support vector machine according to claim 1 is characterized by: In step 3, cross validation combined with grid search is used to optimize the parameters of the support vector machine model. The optimization process is as follows: Step 3.1, set the tuned SVM hyperparameters and their possible value ranges; Step 3.2, combine all different values of parameters into a grid to form all possible parameter combinations; Step 3.3: Use K-fold cross validation to evaluate the model for each parameter combination to ensure that the evaluation results have high robustness; Step 3.4: Train the SVM model in each cross-validation fold and calculate the performance indicators of the validation set; Step 3.5: Compare the average cross-validation performance of all parameter combinations and select the best performing parameter configuration; Step 3.6: Retrain the SVM model on all training data using the optimal parameters.