Big data mining method and system for stock market volatility analysis and prediction
By preprocessing and enhancing stock market data, calculating feature similarity and constructing Laplace matrix, optimizing parameters to determine factor and weight targets, the problem of insufficient impact of time lag in the existing technology is solved, and a more comprehensive data mining effect is achieved.
Patent Information
- Application Number
- CN202510343225.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-04
AI Technical Summary
In the stock market volatility analysis and prediction, the method based on non-negative matrix decomposition has insufficient impact on time lag and has failed to effectively pay attention to the relationship between local features, resulting in insufficient comprehensive data mining.
By preprocessing and enhancing the stock-related indicators, calculating feature similarity, constructing Laplace matrix and regularization terms, optimizing parameters using the least squares method, combining the distribution histogram and key events to determine factors and weight goals, comprehensive data mining is achieved.
It improves the comprehensiveness of stock market data mining, can better capture the data's internal manifold structure, analyze the lag effect at time points, and observe the dynamic changes of factors and weight parameters.
Smart Images

Figure CN120258985A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a big data mining method and system for analyzing and predicting the volatility of the stock market, belonging to the technical field of big data mining. Background Art
[0002] The big data mining method for analyzing and predicting the volatility of the stock market refers to the process of analyzing the index data of the stock market within a continuous period and then mining the factors that cause the aggregation of the stock market.
[0003] Currently, the prior art has proposed a patent for an anomaly detection method and process in the stock market based on non-negative matrix factorization, which mainly identifies the local features of data to a certain extent and quantitatively describes the potential and additive non-linear combination relationship between the local and the whole. However, the anomaly detection in the stock market based on non-negative matrix factorization lacks consideration of the lagging impact generated by time points. Secondly, the anomaly detection in the stock market based on non-negative matrix factorization mainly focuses on local features and weakens the attention to the relationships between various local features. Therefore, relying solely on the method of non-negative matrix factorization to achieve the comprehensiveness of data mining in the stock market is insufficient. Summary of the Invention
[0004] The present invention provides a big data mining method and system for analyzing and predicting the volatility of the stock market, and its main purpose is to improve the comprehensiveness of data mining in the stock market.
[0005] To achieve the above object, a big data mining method for analyzing and predicting the volatility of the stock market provided by the present invention includes:
[0006] Collect stock-related indicators of the stock market, perform data preprocessing on the stock-related indicators to obtain preprocessed data, and perform data enhancement on the preprocessed data to obtain enhanced data features, where the stock-related indicators include closing price, trading volume, and volatility;
[0007] Calculate the feature similarity between each enhanced data feature in the enhanced data features, generate a similarity matrix of the feature similarity, perform matrix sparsification on the similarity matrix to obtain an adjacency matrix, and convert the adjacency matrix into a Laplacian matrix;
[0008] Construct an error term between the enhanced data features and preset mining parameters, and construct a regularization term between the Laplacian matrix and the preset mining parameters, and generate an objective term of the preset mining parameters by using the error term and the regularization term;
[0009] After initializing the preset mining parameters, divide the parameters to be optimized and fixed parameters from the preset mining parameters. According to the parameters to be optimized and the fixed parameters, use the target item to perform parameter alternating optimization on the preset mining parameters to obtain alternating optimization parameters;
[0010] Obtain the factor parameters and weight parameters in the alternating optimization parameters, draw a distribution histogram of the factor parameters in the stock market, use the distribution histogram to determine the factor target of the factor parameters, obtain the key events within the time period corresponding to the weight parameters, and use the key events to determine the weight target of the weight parameters;
[0011] Take the factor parameters, the weight parameters, the factor target, and the weight target as the big data mining results of the stock market.
[0012] Optionally, the data preprocessing of the stock-related indicators to obtain preprocessed data includes:
[0013] Obtain the target closing price, target trading volume, and target volatility that belong to the same stock type in the stock-related indicators;
[0014] Construct a closing price time series, a trading volume time series, and a volatility time series of the target closing price, the target trading volume, and the target volatility respectively within a continuous time period;
[0015] Use the closing price time series, the trading volume time series, and the volatility time series to generate matrix row values;
[0016] Take the numbers of the target closing price, the target trading volume, and the target volatility as the matrix dimensions;
[0017] Take the number of stock types as the number of matrix rows;
[0018] Take the time period length of the closing price time series as the number of matrix columns;
[0019] Determine a multi-dimensional matrix through the matrix row values, the matrix dimensions, the number of matrix rows, and the number of matrix columns;
[0020] Perform sequence segmentation on the matrix row values in the same row of the multi-dimensional matrix at a preset time interval to obtain segmentation sequences;
[0021] Use the segmentation sequences to extract the preprocessed data in different rows and the same time period in the multi-dimensional matrix.
[0022] Optionally, the data enhancement of the preprocessed data to obtain enhanced data features includes:
[0023] Obtain the linear mapping layer, activation function, time-delay weight, and cell membrane potential in the preset non-linear time-delay network;
[0024] Use the linear mapping layer to perform linear mapping on each row of data in the preprocessed data to obtain a linear mapping result;
[0025] Use the activation function to perform non-linear mapping on the linear mapping result to obtain a non-linear mapping result;
[0026] According to the time-delay weight, perform time-delay capture on the non-linear mapping result to obtain a time-delay capture result;
[0027] According to the cell membrane potential, perform data enhancement on the time-delay capture result to obtain enhanced data features.
[0028] Optionally, calculating the feature similarity between each enhanced data feature in the enhanced data features includes:
[0029] Use the preset structural similarity method to calculate the dimensional similarity of each enhanced data feature in the matrix dimension;
[0030] Based on the dimensional similarity, calculate the feature similarity.
[0031] Optionally, constructing the error term between the enhanced data features and the preset mining parameters includes:
[0032] Construct the original product between the original factor and the original weight in the preset mining parameters;
[0033] According to the original product, construct the error term between the enhanced data features and the preset mining parameters.
[0034] Optionally, constructing the regularization term between the Laplacian matrix and the preset mining parameters includes:
[0035] Obtain the original factor in the preset mining parameters;
[0036] Use the original factor to perform congruent transformation on the Laplacian matrix to obtain a congruent transformation matrix;
[0037] Take the product of the trace of the matrix and the preset regularization strength as the regularization term.
[0038] Optionally, according to the parameter to be optimized and the fixed parameter, using the objective term to perform parameter alternating optimization on the preset mining parameters to obtain alternating optimization parameters includes:
[0039] When the parameter to be optimized is the original weight and the fixed parameter is the original factor, obtain the error term in the objective term;
[0040] Optimize the original weights in the error term using the preset least squares method to obtain weight parameters;
[0041] When the parameter to be optimized is the original factor and the fixed parameter is the weight parameter, optimize the original factor in the target term using the least squares method to obtain factor parameters;
[0042] Use the weight parameters and the factor parameters as alternating optimization parameters.
[0043] Optionally, the factor target for determining the factor parameters using the distribution histogram includes:
[0044] Determine the highest ordinate from the distribution histogram;
[0045] Use the factor parameter corresponding to the highest ordinate as the factor target.
[0046] Optionally, the weight target for determining the weight parameters using the key event includes:
[0047] Query the occurrence time period of the key event;
[0048] Query the target parameters of the weight parameters during the occurrence time period and the non-target parameters outside the occurrence time period;
[0049] Analyze the parameter change of the target parameters relative to the non-target parameters;
[0050] Determine the weight target of the weight parameters from the key event through the parameter change.
[0051] To solve the above problems, the present invention also provides a big data mining system for stock market volatility analysis and prediction, the system includes:
[0052] A data enhancement module, configured to collect stock-related indicators of the stock market, perform data preprocessing on the stock-related indicators to obtain preprocessed data, and perform data enhancement on the preprocessed data to obtain enhanced data features, wherein the stock-related indicators include closing price, trading volume, and volatility;
[0053] A matrix conversion module, configured to calculate the feature similarity between each enhanced data feature in the enhanced data features, generate a similarity matrix of the feature similarity, perform matrix sparsification on the similarity matrix to obtain an adjacency matrix, and convert the adjacency matrix into a Laplacian matrix;
[0054] A target generation module, configured to construct an error term between the enhanced data features and preset mining parameters, and construct a regularization term between the Laplacian matrix and the preset mining parameters, and generate a target term of the preset mining parameters by using the error term and the regularization term;
[0055] A parameter optimization module, configured to, after initializing the preset mining parameters, divide the preset mining parameters into parameters to be optimized and fixed parameters, and perform alternating parameter optimization on the preset mining parameters by using the target term according to the parameters to be optimized and the fixed parameters, so as to obtain alternately optimized parameters;
[0056] A target determination module, configured to obtain a factor parameter and a weight parameter in the alternately optimized parameters, draw a distribution histogram of the factor parameter in the stock market, determine a factor target of the factor parameter by using the distribution histogram, obtain key events within a time period corresponding to the weight parameter, and determine a weight target of the weight parameter by using the key events;
[0057] A result determination module, configured to use the factor parameter, the weight parameter, the factor target, and the weight target as big data mining results of the stock market.
[0058] Compared with the problems described in the background art, in the embodiments of the present invention, the stock-related indicators are preprocessed to matrixize the stock data within a continuous time period, which is convenient for subsequent calculations between matrices. Further, in the embodiments of the present invention, the preprocessed data is enhanced to analyze the lag effect generated by time points on the preprocessed data. In the embodiments of the present invention, the feature similarity between each enhanced data feature in the enhanced data features is calculated to cluster the data of different stock types. In the embodiments of the present invention, an error term between the enhanced data features and the preset mining parameters is constructed to ensure the minimization of the error between z(t) and ew(t) when optimizing the parameters using the least squares method in the subsequent process. Further, in the embodiments of the present invention, a regularization term between the Laplacian matrix and the preset mining parameters is constructed to help the algorithm better capture the internal manifold structure of the data through the method of graph regularization, and ensure that the relative distances between similar data points are maintained in the low-dimensional representation of the data, thereby enhancing the clustering relationship between the clustered data. Further, in the embodiments of the present invention, the preset mining parameters are alternately optimized using the objective term according to the parameters to be optimized and the fixed parameters to solve the specific values of the original factors and the original weights. In the embodiments of the present invention, the factor target of the factor parameters is determined using the distribution histogram to observe in which stock industry the factor parameters are most active. In the embodiments of the present invention, the weight target of the weight parameters is determined using the key events to observe the dynamic weight change of the weight parameters with respect to the factor parameters. Therefore, the big data mining method and system for stock market volatility analysis and prediction provided by the embodiments of the present invention can improve the comprehensiveness of data mining in the stock market. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 FIG. is a schematic flowchart of a big data mining method for stock market volatility analysis and prediction provided by an embodiment of the present invention;
[0060] Figure 2 FIG. is a schematic block diagram of a module of a big data mining system for implementing stock market volatility analysis and prediction provided by an embodiment of the present invention.
[0061] The implementation, functional features and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0062] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0063] An embodiment of the present application provides a big data mining method for analyzing and predicting the volatility of the stock market. The execution subject of the big data mining method for analyzing and predicting the volatility of the stock market includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiment of the present application. In other words, the big data mining method for analyzing and predicting the volatility of the stock market can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc.
[0064] Embodiment 1:
[0065] Refer to Figure 1 As shown, it is a schematic flowchart of the big data mining method for analyzing and predicting the volatility of the stock market provided by an embodiment of the present invention. In this embodiment, the big data mining method for analyzing and predicting the volatility of the stock market includes:
[0066] S1. Collect stock-related indicators of the stock market, perform data preprocessing on the stock-related indicators to obtain preprocessed data, and perform data enhancement on the preprocessed data to obtain enhanced data features, where the stock-related indicators include closing price, trading volume, and volatility.
[0067] In the embodiment of the present invention, by performing data preprocessing on the stock-related indicators, the stock data within a continuous period is matrixed, which is convenient for subsequent calculations between matrices.
[0068] In an embodiment of the present invention, the performing data preprocessing on the stock-related indicators to obtain preprocessed data includes: obtaining the target closing price, target trading volume, and target volatility belonging to the same stock type among the stock-related indicators; respectively constructing a closing price time series, a trading volume time series, and a volatility time series of the target closing price, the target trading volume, and the target volatility within a continuous period; generating matrix row values using the closing price time series, the trading volume time series, and the volatility time series; taking the number of the target closing price, the target trading volume, and the target volatility as the matrix dimension; taking the number of stock types as the number of matrix rows; taking the time period length of the closing price time series as the number of matrix columns; determining a multi-dimensional matrix through the matrix row values, the matrix dimension, the number of matrix rows, and the number of matrix columns; performing sequence segmentation on the matrix row values in the same row of the multi-dimensional matrix at a preset time interval to obtain a segmentation sequence; and extracting preprocessed data in different rows and the same period in the multi-dimensional matrix using the segmentation sequence.
[0069] Among them, the preset time interval refers to a time window. For example, if the time series of the closing price has a period length of 10 minutes and the preset time interval is one minute, the stock types refer to different types of stocks, such as agricultural stocks, technology stocks, etc.
[0070] Furthermore, in the embodiment of the present invention, data enhancement is performed on the preprocessed data to analyze the lag effect generated by the time point on the preprocessed data.
[0071] In one embodiment of the present invention, performing data enhancement on the preprocessed data to obtain enhanced data features includes: obtaining a linear mapping layer, an activation function, a time lag weight, and a cell membrane potential in a preset nonlinear time-delay network; using the linear mapping layer to perform a linear mapping on each row of data in the preprocessed data to obtain a linear mapping result; using the activation function to perform a nonlinear mapping on the linear mapping result to obtain a nonlinear mapping result; according to the time lag weight, using the following formula to perform time lag capture on the nonlinear mapping result to obtain a time lag capture result:
[0072]
[0073] Among them, represents the time lag capture result of the m-th matrix dimension of the i-th stock type at the t-th moment, represents the preprocessed data of the m-th matrix dimension of the i-th stock type at the (t - k)-th moment, δ k represents the time lag weight, W k represents the weight of the linear mapping layer, b k represents the bias of the linear mapping layer, represents the linear mapping result, σ represents the activation function, represents the nonlinear mapping result, T represents the preset time interval;
[0074] According to the cell membrane potential, using the following formula to perform data enhancement on the time lag capture result to obtain enhanced data features:
[0075]
[0076] Among them, represents the enhanced data feature, α represents the attenuation coefficient, T represents the preset time interval, represents the time lag capture result of the m-th matrix dimension of the i-th stock type at the (t - k)-th moment, represents the cell membrane potential at the previous moment of, represents the cell membrane potential at the (t - k)-th moment.
[0077] Among them, the time-delay weight refers to the weight parameter obtained after training the non-linear time-delay network.
[0078] S2. Calculate the feature similarity between each enhanced data feature in the enhanced data features, generate a similarity matrix of the feature similarity, sparsify the similarity matrix to obtain an adjacency matrix, and convert the adjacency matrix into a Laplacian matrix.
[0079] In the embodiment of the present invention, by calculating the feature similarity between each enhanced data feature in the enhanced data features, data of different stock types are clustered.
[0080] In an embodiment of the present invention, the calculating the feature similarity between each enhanced data feature in the enhanced data features includes: calculating the dimensional similarity of each enhanced data feature in the matrix dimension by using a preset structural similarity method; the structural similarity method includes:
[0081]
[0082] Among them, represents the dimensional similarity between the i-th stock type and the j-th stock type, C1 and C2 represent decimals to avoid a denominator of 0, and μ i represents the mean of the time series of the m-th matrix dimension of the i-th stock type, and μ j represents the mean of the time series of the m-th matrix dimension of the j-th stock type, and ρ ij represents the covariance between the m-th matrix dimension of the i-th stock type and the m-th matrix dimension of the j-th stock type, and ρ i represents the variance of the m-th matrix dimension of the i-th stock type, and ρ j represents the variance of the m-th matrix dimension of the j-th stock type;
[0083] Based on the dimensional similarity, calculate the feature similarity by using the following formula:
[0084]
[0085] Among them, S ij represents the feature similarity, M represents the number of matrix dimensions, represents the dimensional similarity between the i-th stock type and the j-th stock type.
[0086] Among them, the similarity matrix refers to a matrix of feature similarities with the row index i and the column index j, the adjacency matrix refers to the representation result of the adjacency matrix of the similarity matrix, and the Laplacian matrix refers to the difference between the degree matrix corresponding to the adjacency matrix and the adjacency matrix.
[0087] S3. Construct an error term between the enhanced data features and the preset mining parameters, and construct a regularization term between the Laplacian matrix and the preset mining parameters, and generate a target term of the preset mining parameters by using the error term and the regularization term.
[0088] In an embodiment of the present invention, by constructing an error term between the enhanced data features and the preset mining parameters, the error between z(t) and ew(t) is minimized when optimizing the parameters by using the least squares method subsequently.
[0089] In an embodiment of the present invention, constructing the error term between the enhanced data features and the preset mining parameters includes: constructing an original product between the original factors and the original weights in the preset mining parameters; and constructing the error term between the enhanced data features and the preset mining parameters according to the original product by using the following formula:
[0090]
[0091] where, Δ1 represents the error term, z(t) represents the enhanced data features at time t, the matrix dimension of z is M, the number of rows of the matrix is the number of stock types, the number of columns of the matrix is the time period length of the closing price time series, ew(t) represents the original product, e represents the original factor, w(t) represents the original weight, and F represents the Frobenius norm.
[0092] Furthermore, in an embodiment of the present invention, by constructing a regularization term between the Laplacian matrix and the preset mining parameters, the method of graph regularization is used to help the algorithm better capture the intrinsic manifold structure of the data, and ensure that the relative distances between similar data points are maintained in the low-dimensional representation of the data, thereby enhancing the clustering relationship between the clustered data.
[0093] In an embodiment of the present invention, constructing the regularization term between the Laplacian matrix and the preset mining parameters includes: obtaining the original factors in the preset mining parameters; performing a congruence transformation on the Laplacian matrix by using the original factors to obtain a congruence transformation matrix; and taking the product of the trace of the matrix and the preset regularization strength as the regularization term.
[0094] Optionally, the congruence transformation matrix refers to the process of e T Ae, where A represents the Laplacian matrix and e represents the original factor.
[0095] where, the target term refers to the mean value of the difference between the error term and the regularization term within a continuous time period.
[0096] S4. After initializing the preset mining parameters, divide the parameters to be optimized and the fixed parameters from the preset mining parameters. According to the parameters to be optimized and the fixed parameters, use the target item to perform parameter alternating optimization on the preset mining parameters to obtain the alternately optimized parameters.
[0097] Optionally, the process of initializing the preset mining parameters refers to assigning random values to the parameters to be optimized and the fixed parameters. The parameters to be optimized and the fixed parameters include the original factors and the original weights.
[0098] Furthermore, in the embodiment of the present invention, by performing parameter alternating optimization on the preset mining parameters according to the parameters to be optimized and the fixed parameters and using the target item, the specific values of the original factors and the original weights are solved.
[0099] In an embodiment of the present invention, the performing parameter alternating optimization on the preset mining parameters according to the parameters to be optimized and the fixed parameters and using the target item to obtain the alternately optimized parameters includes: when the parameter to be optimized is the original weight and the fixed parameter is the original factor, obtaining the error term in the target item; using the preset least squares method to optimize the original weight in the error term to obtain the weight parameter; when the parameter to be optimized is the original factor and the fixed parameter is the weight parameter, using the least squares method to optimize the original factor in the target item to obtain the factor parameter; using the weight parameter and the factor parameter as the alternately optimized parameters.
[0100] Optionally, the algorithm for using the preset least squares method to optimize the original weight in the error term is as follows: where w(t)′ represents the weight parameter, and ∈ represents a decimal to ensure that the denominator is not zero; furthermore, the algorithm for using the least squares method to optimize the original factor in the target item is as follows: where L(t) represents the Laplacian matrix, and e′ represents the factor parameter.
[0101] S5. Obtain the factor parameter and the weight parameter in the alternately optimized parameters, draw a distribution histogram of the factor parameter in the stock market, use the distribution histogram to determine the factor target of the factor parameter, obtain the key events during the period corresponding to the weight parameter, and use the key events to determine the weight target of the weight parameter.
[0102] In the embodiment of the present invention, the distribution histogram refers to a distribution diagram with the factor parameter as the abscissa and the weight parameter as the ordinate. Each column of the factor parameter represents a potential factor, and each row represents the loading of a stock on this factor, that is, the contribution degree of this factor to the stock characteristics. A high loading indicates that the stock is significantly affected by this factor, and a low loading indicates that the stock has a weak association with this factor.
[0103] Furthermore, in the embodiment of the present invention, the factor target of the factor parameter is determined by using the distribution histogram to observe in which stock industry the factor parameter is the most active.
[0104] Wherein, the factor target refers to a specific industry in the factor parameter, and the factor parameter refers to categories of different industries, such as technology, finance, energy and other industries.
[0105] In an embodiment of the present invention, determining the factor target of the factor parameter by using the distribution histogram includes: determining the highest ordinate from the distribution histogram; using the factor parameter corresponding to the highest ordinate as the factor target.
[0106] Furthermore, in the embodiment of the present invention, the weight target of the weight parameter is determined by using the key event to observe the dynamic weight change of the weight parameter with respect to the factor parameter.
[0107] Wherein, the key event refers to an event that promotes the change of the factor parameter, such as economic recession, policy change, etc.
[0108] In an embodiment of the present invention, determining the weight target of the weight parameter by using the key event includes: querying the occurrence time period of the key event; querying the target parameter of the weight parameter during the occurrence time period and the non-target parameter outside the occurrence time period; analyzing the parameter change of the target parameter relative to the non-target parameter; determining the weight target of the weight parameter from the key event through the parameter change.
[0109] Optionally, the process of determining the weight target of the weight parameter from the key event through the parameter change is, for example: if the non-target parameter is distributed between 1 and 10, and the target parameter is distributed above 20, it means that the target parameter has changed drastically relative to the non-target parameter, which is consistent with the time period when the key event occurs, then the key event is used as the weight target.
[0110] S6. Using the factor parameter, the weight parameter, the factor target and the weight target as the big data mining results of the stock market.
[0111] Compared with the problems described in the background art, in the embodiments of the present invention, the stock-related indicators are preprocessed to matrixize the stock data within a continuous period, facilitating subsequent calculations between matrices. Further, in the embodiments of the present invention, the preprocessed data is enhanced to analyze the lag effect generated by time points on the preprocessed data. In the embodiments of the present invention, the feature similarity between each enhanced data feature in the enhanced data features is calculated to cluster the data of different stock types. In the embodiments of the present invention, an error term between the enhanced data features and the preset mining parameters is constructed to minimize the error between z(t) and ew(t) when optimizing the parameters using the least squares method subsequently. Further, in the embodiments of the present invention, a regularization term between the Laplacian matrix and the preset mining parameters is constructed to help the algorithm better capture the intrinsic manifold structure of the data through graph regularization and ensure that the relative distances between similar data points are maintained in the low-dimensional representation of the data, thereby enhancing the clustering relationship between the clustered data. Further, in the embodiments of the present invention, the preset mining parameters are alternately optimized using the objective term according to the parameters to be optimized and the fixed parameters to solve the specific values of the original factors and the original weights. In the embodiments of the present invention, the factor target of the factor parameters is determined using the distribution histogram to observe in which stock industry the factor parameters are most active. In the embodiments of the present invention, the weight target of the weight parameters is determined using the key events to observe the dynamic weight changes of the weight parameters with respect to the factor parameters. Therefore, the big data mining method and system for stock market volatility analysis and prediction provided by the embodiments of the present invention can improve the comprehensiveness of data mining in the stock market.
[0112] Embodiment 2:
[0113] As Figure 2 shown, it is a functional module diagram of a big data mining system for stock market volatility analysis and prediction according to the present invention.
[0114] The big data mining system 200 for stock market volatility analysis and prediction according to the present invention can be installed in an electronic device. According to the functions achieved, the big data mining system for stock market volatility analysis and prediction can include a data enhancement module 201, a matrix conversion module 202, a target generation module 203, a parameter optimization module 204, a target determination module 205, and a result determination module 206. The modules in the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.
[0115] In the embodiments of the present invention, the functions of each module / unit are as follows:
[0116] The data enhancement module 201 is used to collect stock-related indicators in the stock market, perform data preprocessing on the stock-related indicators to obtain preprocessed data, and perform data enhancement on the preprocessed data to obtain enhanced data features. Among them, the stock-related indicators include closing price, trading volume, and volatility;
[0117] The matrix conversion module 202 is used to calculate the feature similarity between each enhanced data feature in the enhanced data features, generate a similarity matrix of the feature similarity, perform matrix sparsification on the similarity matrix to obtain an adjacency matrix, and convert the adjacency matrix into a Laplacian matrix;
[0118] The target generation module 203 is used to construct an error term between the enhanced data features and preset mining parameters, and construct a regularization term between the Laplacian matrix and the preset mining parameters, and generate a target term of the preset mining parameters by using the error term and the regularization term;
[0119] The parameter optimization module 204 is used to divide the preset mining parameters into parameters to be optimized and fixed parameters after initializing the preset mining parameters, and perform parameter alternating optimization on the preset mining parameters by using the target term according to the parameters to be optimized and the fixed parameters to obtain alternately optimized parameters;
[0120] The target determination module 205 is used to obtain the factor parameters and weight parameters in the alternately optimized parameters, draw a distribution histogram of the factor parameters in the stock market, determine the factor target of the factor parameters by using the distribution histogram, obtain key events during the period corresponding to the weight parameters, and determine the weight target of the weight parameters by using the key events;
[0121] The result determination module 206 is used to use the factor parameters, the weight parameters, the factor target, and the weight target as the big data mining results of the stock market.
[0122] Specifically, each module in the big data mining system 200 for stock market volatility analysis and prediction in the embodiments of the present invention adopts the same technical means as the big data mining method for stock market volatility analysis and prediction described above Figure 1 and can produce the same technical effects, which will not be elaborated here.
[0123] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.
[0124] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A big data mining method for analyzing and predicting the volatility of the stock market, characterized in that, The method includes: Collecting stock-related indicators in the stock market, performing data preprocessing on the stock-related indicators to obtain preprocessed data, and performing data augmentation on the preprocessed data to obtain augmented data features, where the stock-related indicators include closing price, trading volume, and volatility; Calculating the feature similarity between each augmented data feature in the augmented data features, generating a similarity matrix of the feature similarity, performing matrix sparsification on the similarity matrix to obtain an adjacency matrix, and converting the adjacency matrix into a Laplacian matrix; Constructing an error term between the augmented data features and preset mining parameters, and constructing a regularization term between the Laplacian matrix and the preset mining parameters, and generating an objective term of the preset mining parameters by using the error term and the regularization term; After initializing the preset mining parameters, dividing the parameters to be optimized and fixed parameters from the preset mining parameters, and alternately optimizing the preset mining parameters by using the objective term according to the parameters to be optimized and the fixed parameters to obtain alternately optimized parameters; Obtaining the factor parameters and weight parameters in the alternately optimized parameters, plotting a distribution histogram of the factor parameters in the stock market, determining a factor target of the factor parameters by using the distribution histogram, obtaining key events during the time period corresponding to the weight parameters, and determining a weight target of the weight parameters by using the key events; Taking the factor parameters, the weight parameters, the factor target, and the weight target as the big data mining results of the stock market.
2. The big data mining method for stock market volatility analysis and prediction according to claim 1, characterized in that, The performing data preprocessing on the stock-related indicators to obtain preprocessed data includes: Obtaining the target closing price, target trading volume, and target volatility that belong to the same stock type in the stock-related indicators; Respectively constructing a closing price time series, a trading volume time series, and a volatility time series of the target closing price, the target trading volume, and the target volatility in a continuous time period; Generating matrix row values by using the closing price time series, the trading volume time series, and the volatility time series; Taking the numbers of the target closing price, the target trading volume, and the target volatility as matrix dimensions; Taking the number of stock types as the number of matrix rows; Taking the time period length of the closing price time series as the number of matrix columns; Determining a multi-dimensional matrix through the matrix row values, the matrix dimensions, the number of matrix rows, and the number of matrix columns; Performing sequence segmentation on the matrix row values in the same row of the multi-dimensional matrix at a preset time interval to obtain segmented sequences; Extracting preprocessed data in different rows and the same time period in the multi-dimensional matrix by using the segmented sequences.
3. The big data mining method for analyzing and predicting the volatility of the stock market according to claim 1, characterized in that The performing data augmentation on the preprocessed data to obtain augmented data features includes: Obtaining a linear mapping layer, an activation function, a time-delay weight, and a membrane potential in a preset non-linear time-delay network; Performing linear mapping on each row of data in the preprocessed data by using the linear mapping layer to obtain a linear mapping result; Performing non-linear mapping on the linear mapping result by using the activation function to obtain a non-linear mapping result; Perform time-delay capture on the non-linear mapping result according to the time-delay weight to obtain a time-delay capture result; Perform data enhancement on the time-delay capture result according to the cell membrane potential to obtain enhanced data features.
4. The big data mining method for analyzing and predicting the volatility of the stock market according to claim 1, wherein Calculating the feature similarity between each enhanced data feature in the enhanced data features includes: Calculating the dimensional similarity of each enhanced data feature in the matrix dimension by using a preset structural similarity method; Calculating the feature similarity based on the dimensional similarity.
5. The big data mining method for stock market volatility analysis and prediction according to claim 1, characterized in that Constructing the error term between the enhanced data feature and the preset mining parameter includes: Constructing the original product between the original factor and the original weight in the preset mining parameter; Constructing the error term between the enhanced data feature and the preset mining parameter according to the original product.
6. The big data mining method for stock market volatility analysis and prediction according to claim 1, wherein Constructing the regularization term between the Laplacian matrix and the preset mining parameter includes: Obtaining the original factor in the preset mining parameter; Performing a congruent transformation on the Laplacian matrix by using the original factor to obtain a congruent transformation matrix; Taking the product of the trace of the matrix and the preset regularization strength as the regularization term.
7. The big data mining method for stock market volatility analysis and prediction according to claim 1, characterized in that According to the parameter to be optimized and the fixed parameter, using the objective term to perform parameter alternating optimization on the preset mining parameter to obtain an alternating optimization parameter, including: When the parameter to be optimized is the original weight and the fixed parameter is the original factor, obtaining the error term in the objective term; Optimizing the original weight in the error term by using a preset least squares method to obtain a weight parameter; When the parameter to be optimized is the original factor and the fixed parameter is the weight parameter, using the least squares method to optimize the original factor in the objective term to obtain a factor parameter; Taking the weight parameter and the factor parameter as the alternating optimization parameter.
8. The big data mining method for analyzing and predicting the volatility of the stock market according to claim 1, wherein, Using the distribution histogram to determine the factor target of the factor parameter includes: Determining the highest ordinate from the distribution histogram; Taking the factor parameter corresponding to the highest ordinate as the factor target.
9. The big data mining method for analyzing and predicting the volatility of the stock market according to claim 1, wherein Using the key event to determine the weight target of the weight parameter includes: Querying the occurrence time period of the key event; Querying the target parameter of the weight parameter in the occurrence time period and the non-target parameter outside the occurrence time period; Analyzing the parameter change of the target parameter relative to the non-target parameter; Determining the weight target of the weight parameter from the key event through the parameter change.
10. A big data mining system for stock market volatility analysis and prediction, characterized in that, The system includes: A data enhancement module, configured to collect stock-related indicators in the stock market, perform data preprocessing on the stock-related indicators to obtain preprocessed data, and perform data enhancement on the preprocessed data to obtain enhanced data features, where the stock-related indicators include closing price, trading volume, and volatility; A matrix conversion module, configured to calculate the feature similarity between each enhanced data feature in the enhanced data features, generate a similarity matrix of the feature similarity, perform matrix sparsification on the similarity matrix to obtain an adjacency matrix, and convert the adjacency matrix into a Laplacian matrix; A target generation module, configured to construct an error term between the enhanced data features and preset mining parameters, and construct a regularization term between the Laplacian matrix and the preset mining parameters, and generate a target term of the preset mining parameters by using the error term and the regularization term; A parameter optimization module, configured to, after initializing the preset mining parameters, divide the preset mining parameters into parameters to be optimized and fixed parameters, and perform parameter alternating optimization on the preset mining parameters by using the target term according to the parameters to be optimized and the fixed parameters, so as to obtain alternately optimized parameters; A target determination module, configured to obtain factor parameters and weight parameters in the alternately optimized parameters, draw a distribution histogram of the factor parameters in the stock market, determine a factor target of the factor parameters by using the distribution histogram, obtain key events within a time period corresponding to the weight parameters, and determine a weight target of the weight parameters by using the key events; A result determination module, configured to use the factor parameters, the weight parameters, the factor target and the weight target as big data mining results of the stock market.