Abnormal behavior detection and risk prevention and control method and system in stock transaction
By optimizing the decision tree feature selection and information entropy calculation in the random forest algorithm, a subset of features is constructed, the risk prevention and control capabilities of the stock trading data classification model are improved, and the detection accuracy problem caused by the non-balance of stock trading data is solved, and more efficient abnormal behavior detection is achieved.
Patent Information
- Application Number
- CN202510536136.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-08-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, due to the unbalance of stock trading data, the accuracy of the abnormal data detection by the stock trading data classification model decreases, which reduces the model's risk prevention and control capabilities.
The random forest algorithm is used to construct a stock trading data classification model, and the feature selection of optimized decision tree features is used to construct a subset of features, calculate the decision value, and improve the accuracy of detection of stock trading anomaly data.
The risk prevention and control ability of the stock trading data classification model for abnormal behaviors of multiple types of stock trading is improved, the impact of non-equilibrium data in the training set on model training is reduced, and the accuracy of abnormal behavior detection is enhanced.
Smart Images

Figure CN120450863A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of machine learning technology, and specifically to a method and system for detecting abnormal behavior and preventing and controlling risks in stock trading. Background Art
[0002] Abnormal behaviors in stock trading, such as stock price manipulation, insurance and credit fraud, occur frequently in both mature and emerging stock markets. They disrupt market order, reduce market efficiency, harm the interests of ordinary investors, and in serious cases, cause financial risks, which is not conducive to the long-term stable development of the stock market.
[0003] Currently, AI-based stock trading data classification models outperform traditional multivariate statistical regression models in detecting abnormal stock trading behavior. The random forest algorithm, an AI classification model based on decision trees, offers improved interpretability and provides a foundation for further exploring the micro-manifestations of abnormal stock trading behavior and market structure. However, due to the small proportion of abnormal stock trading data in the securities market, there is a significant imbalance in the distribution of normal and abnormal stock trading data. This results in a decrease in the accuracy of the stock trading data classification model in detecting abnormal data, weakening the model's risk prevention and control capabilities. Summary of the Invention
[0004] To address the technical issue of decreased accuracy in detecting abnormal data, this application provides a method and system for detecting abnormal behavior and preventing and controlling risks in stock trading. The technical solutions employed are as follows:
[0005] In one aspect, the present application proposes a method for detecting abnormal behavior and preventing and controlling risks in stock trading, the method comprising the following steps:
[0006] Collecting a number of stock transaction data, forming a feature set from a number of features constituting the stock transaction data, and numbering the features in the feature set to obtain a feature sequence;
[0007] A training set is constructed by extracting features from all collected stock trading data. Within the training set, features are extracted from the feature set to construct a feature subset, and the feature sequences are sorted to obtain a feature subsequence. The eigenvalues of each data feature are classified into normal data and abnormal data. The information entropy of the data feature is calculated by combining the frequency of occurrence of the eigenvalue in normal data and abnormal data with all eigenvalues corresponding to the data feature. The amount of abnormal information in the feature subset is calculated based on the difference between the information entropy of all data features in the feature subset and the mean information entropy of all data features in the feature set.
[0008] Extract feature subsequences from T feature subsets to form a feature sequence; use the difference between the number of intersections of the data features of the T feature subsets and the number of data features of the feature subsets as the repetition weight; calculate the objective function of the feature sequence based on the repetition weight and the amount of abnormal information of the feature subsets in the feature sequence; the smaller the repetition weight, the greater the amount of abnormal information and the larger the objective function; use an optimization algorithm to process multiple feature sequences to obtain the optimal feature sequence and optimal fitness;
[0009] A decision tree is constructed based on the optimal feature sequence of the training set to obtain a classification model. The optimal fitness corresponding to the decision tree and the frequency of occurrence of all data features in the decision tree are forward fused to obtain the weight of the decision tree. The stock trading data to be detected is input into the classification model to obtain the decision value of each decision tree. The decision value of each decision tree is weighted by the weight to obtain the decision value of the classification model. Anomaly detection is completed based on the decision value of the classification model.
[0010] In the above scheme, this application adopts the random forest algorithm to construct a stock trading data classification model, analyzes various stock trading abnormal behaviors in the securities market, and improves the risk prevention and control capabilities of the model; according to the characteristic distribution of stock trading data, the selection of decision tree features in the random forest algorithm is optimized to reduce the impact of the imbalance of stock trading data on the model detection capability, and improve the accuracy of detecting abnormal stock trading data; by analyzing the distribution of stock trading data features, using the optimization algorithm to construct a feature subset, and further calculating the decision value of the random forest algorithm, on the one hand, the risk prevention and control capabilities of the stock trading data classification model for multiple types of stock trading abnormal behaviors are improved, and on the other hand, the impact of unbalanced data in the training set on model training is reduced, and the accuracy of detecting abnormal behaviors in stock trading is improved to strengthen risk prevention and control.
[0011] In one embodiment, the number of the training sets is N, the number of stock trading data in the training sets is the same as the number of all collected stock trading data, and the stock trading data in the training sets are repeatable.
[0012] In one embodiment, the method for calculating the information entropy of a data feature by using the frequency of occurrence of the feature value in normal data and abnormal data as a weight and combining all the feature values corresponding to the data feature is:
[0013] pu i Indicates the frequency of occurrence of the i-th eigenvalue of the data feature in abnormal data; pv i Indicates the frequency of occurrence of the i-th eigenvalue of the data feature in normal data, L k represents the number of eigenvalues in the kth data feature, a1 and a2 represent the weights of the eigenvalues appearing in abnormal data and normal data respectively, and H kRepresents the information entropy of the kth data feature in the feature subset.
[0014] In one embodiment, the method for calculating the abnormal information amount of the feature subset based on the difference between the information entropy of all data features in the feature subset and the mean information entropy of all data features in the training set is:
[0015] The mean of the information entropy of all data features in the feature set is recorded as the first information mean;
[0016] The amount of abnormal information is positively correlated with the information entropy of the data feature and negatively correlated with the first information mean.
[0017] In one embodiment, the target component is obtained by calculating the ratio of the number of data features in the intersection of the data features of the T feature subsets corresponding to the feature sequence to the number of data features in the feature subset according to the repetition weight.
[0018] In one embodiment, the method for calculating the objective function of the feature sequence based on the repetition weight and the abnormal information content of the feature subset in the feature sequence is:
[0019] α represents the repetition weight of the feature sequence, A t represents the amount of abnormal information of the t-th feature subset, exp() represents the exponential function with a natural constant as the base, and J represents the objective function of the feature sequence.
[0020] In one embodiment, the method for obtaining the weight of the decision tree by forward fusing the optimal fitness corresponding to the decision tree and the frequencies of occurrence of all data features in the decision tree is:
[0021] J n Indicates the optimal fitness corresponding to the nth decision tree, F m represents the frequency of the mth data feature in N decision trees, M represents the number of data features in the nth decision tree, norm() represents the normalization function, and w n Represents the weight of the nth decision tree.
[0022] In one embodiment, the anomaly detection method is:
[0023] Input the stock trading data to be detected into the classification model; obtain the decision value in each decision tree at this time; weight the decision value by the weight of the decision tree to obtain the decision value of the current classification model; when the decision value is not less than zero, classify the stock trading data as normal data; when the decision value is negative, classify the stock trading data as abnormal data.
[0024] In one embodiment, the decision value of the classification model is:
[0025] w n Represents the weight of the nth decision tree, s n Represents the decision value of the nth decision tree, N represents the number of decision trees, and S represents the decision value of the random forest algorithm output classification model.
[0026] On the other hand, an embodiment of the present application also provides a system for detecting abnormal behavior and preventing and controlling risks in stock transactions, comprising a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of any one of the above-mentioned methods for detecting abnormal behavior and preventing and controlling risks in stock transactions are implemented.
[0027] The beneficial effects of this application are:
[0028] This application uses a random forest algorithm to construct a stock trading data classification model, analyzes various stock trading abnormal behaviors in the securities market, and improves the risk prevention and control capabilities of the model; based on the characteristic distribution of stock trading data, the selection of decision tree features in the random forest algorithm is optimized to reduce the impact of stock trading data imbalance on the model detection capability and improve the accuracy of detecting abnormal stock trading data; by analyzing the distribution of stock trading data features, using an optimization algorithm to construct a feature subset, and further calculating the decision value of the random forest algorithm, on the one hand, the risk prevention and control capabilities of the stock trading data classification model for multiple types of stock trading abnormal behaviors are improved, and on the other hand, the impact of unbalanced data in the training set on model training is reduced, and the accuracy of detecting abnormal behaviors in stock trading is improved to strengthen risk prevention and control. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0030] Figure 1 A flow chart of a method for detecting abnormal behavior and preventing and controlling risks in stock trading provided in one embodiment of the present application. DETAILED DESCRIPTION
[0031] In order to further illustrate the technical means and effects adopted by this application to achieve the predetermined invention purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail the specific implementation method, structure, features and effects of a method and system for detecting abnormal behavior and preventing and controlling risks in stock transactions proposed in this application. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics of one or more embodiments may be combined in any suitable form.
[0032] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.
[0033] An embodiment of a method and system for detecting abnormal behavior and preventing and controlling risks in stock trading:
[0034] The following describes in detail a specific scheme of a method for detecting abnormal behavior and preventing and controlling risks in stock transactions provided by this application with reference to the accompanying drawings.
[0035] See also Figure 1 , which shows a flow chart of a method for detecting abnormal behavior and preventing and controlling risks in stock trading provided by one embodiment of the present application, the method comprising the following steps:
[0036] Step S001: Collect stock transaction data to form a feature set and feature sequence.
[0037] High-frequency trading generates a large amount of fine-grained, high-speed data in the stock market. This high-frequency data quantitatively depicts the stock market's trading mechanisms, reveals changes in the market's microstructure, and details trader behavior. It contains a wealth of information that is difficult to capture in traditional low-frequency, coarse-grained data.
[0038] This application first constructs a training set. The data of the training set is high-frequency and ultra-high-frequency trading data in the stock market, and the data comes from the official website of the China Securities Regulatory Commission (CSRC). The number of stock trading data selected in this embodiment is 30,000. Then, data with multiple different features are selected. In order to better capture the high-frequency and micro information in the market, this embodiment selects three categories of data features of 22 stock trading data. Each stock trading data is a sequence composed of the data features of 22 stock trading data, specifically:
[0039] Based on the characteristics of market transactions: rate of return, excess rate of return, standard deviation of rate of return per minute, trading volume per minute, rise or fall, excess rise or fall, and number of transactions per minute.
[0040] Characteristics based on the micro-information of the limit order book: the rate of return of the middle quote per minute, the rate of return of the weighted price per minute, the total number of orders in the limit order book, the relative spread of the best bid-ask price, the height imbalance of the limit order book, the length imbalance of the limit order book, the buyer slope of the limit order book, the seller slope of the limit order book, the difference between the buyer and seller slopes of the limit order book, the dispersion of the limit order book, the total buyer entrustment volume of the order book, and the total seller entrustment volume of the order book. Among them, the height imbalance of the limit order book and the length imbalance of the limit order book are used to measure the relative value of the imbalance in the quotes of buyers and sellers and the relative value of the imbalance in the number of transactions willing to be made by buyers and sellers, respectively.
[0041] Based on trading behavior: the number of large orders, the amount of aggressive buy orders, and the total number of order cancellations per minute.
[0042] The method for calculating and obtaining the characteristic values of the above-mentioned data features is a well-known technology. All data features are grouped into a feature set, and the data features in the feature set are numbered using an arbitrary method to form a sequence recorded as a feature sequence.
[0043] At this point, different features and their serial numbers are obtained.
[0044] Step S002: Select a training set, construct a feature subset therein, calculate the information entropy of the data features based on the frequency of the feature values in the feature subset in different data and the feature values of the data features; and calculate the abnormal information amount of the feature subset based on the information entropy difference.
[0045] All stock transaction data are divided into N training sets. In this embodiment, the Bootstrap sampling method is used to extract the training sets, and the value of N is 200. In this embodiment, each training set contains 30,000 stock transaction data, and the stock transaction data can be repeated.
[0046] For each training set, all data features are divided into different feature subsets to comprehensively characterize different types of stock trading anomalies and improve the risk prevention and control capabilities of the stock trading data classification model. Considering that some features have a low correlation with data anomalies, constructing feature subsets based on them would actually reduce the classification accuracy of the classification model. Therefore, feature subsets for the training set are obtained based on the distribution of data features within the training set.
[0047] K data features are randomly selected from the feature set to construct a feature subset. In this embodiment, K is set to 5. The serial numbers of the corresponding features are arranged in the order of selection to obtain a feature subsequence of the feature subset. The length of the feature subsequence is K. In this embodiment, each data feature corresponds to 30,000 eigenvalues. The information entropy of the eigenvalues corresponding to a data feature reflects the amount of information that the data feature contains in identifying abnormal data. The greater the information entropy, the greater the effectiveness in identifying abnormal stock trading behavior.
[0048] In the training set, the eigenvalues of each data feature are labeled and artificially divided into outliers and non-outliers. All outliers constitute abnormal data, and all non-outliers constitute normal data. To reduce the impact of data imbalance in the training set on the training classification model, when calculating information entropy, a higher weight is assigned to the eigenvalues of abnormal stock trading data, so that the amount of information provided by abnormal and normal stock trading data in the training set for model training tends to be balanced. For each eigenvalue of a data feature, the frequency of its occurrence in abnormal data and non-abnormal data is calculated, and then the information entropy of each data feature is calculated based on its corresponding weight.
[0049] In this embodiment, the specific expression of the information entropy of the data feature is:
[0050] pu i Indicates the frequency of occurrence of the i-th eigenvalue of the data feature in abnormal data; pv i Indicates the frequency of occurrence of the i-th eigenvalue of the data feature in normal data, L k represents the number of eigenvalues in the kth data feature, a1 and a2 represent the weights of the eigenvalues in abnormal data and normal data, a1 is 0.3, a2 is 0.7, H k Represents the information entropy of the kth data feature in the feature subset.
[0051] The information entropy of each data feature in the feature subset is calculated in the above manner, the mean of the information entropy of all data features in the feature set is recorded as the first information mean, and the abnormal information amount of the feature subset is obtained based on the difference between all information entropies in the feature subset and the first information mean.
[0052] The amount of abnormal information is positively correlated with the information entropy of each data feature and negatively correlated with the first information mean.
[0053] It should be noted that positive correlation means that when one variable increases, the other variable also increases, and the two variables change in the same direction. When one variable changes from large to small or from small to large, the other variable also changes from large to small or from small to large; the specific relationship is determined by actual application and this application does not impose any special restrictions.
[0054] It should be noted that negative correlation means that when one variable increases, the other variable decreases accordingly, and the two variables change in opposite directions. When one variable changes from large to small or from small to large, the other variable also changes from small to large or from large to small. The specific details are determined by actual applications and are not particularly limited in this application.
[0055] Preferably, the expression of the abnormal information amount of the feature subset is:
[0056] H k represents the information entropy of the kth data feature in the feature subset, represents the first information mean, K represents the number of data features in the feature subset, and A represents the amount of abnormal information in the feature subset. The abnormal information reflects the relative amount of information provided by all data features in the feature subset for detecting abnormal stock trading data.
[0057] At this point, the amount of abnormal information for each feature subset has been obtained.
[0058] Step S003: extract feature subsets to form a feature sequence, calculate the objective function of the feature sequence based on the repetition and abnormal information amount of the feature sequence, and obtain the optimal feature sequence and optimal fitness.
[0059] After obtaining the abnormal information of each feature subset through the above steps, randomly select T feature subsets and arrange the feature subsequences of the T feature subsets in sequence to form a feature sequence. The length of the feature sequence is 3K.
[0060] Because repeated data features exist in different feature subsets, a larger number of repeated data features reduces their utilization, exacerbating the impact of data imbalance and hindering subsequent classification model training. Therefore, to minimize the impact of repeated data features on the identification of abnormal stock trading data, feature sequences with a large number of repeated data features are assigned smaller weights. This results in a smaller calculated objective function, reflecting a poorer classification effect on stock trading data.
[0061] Take the intersection of the data features of the T feature subsets corresponding to the feature sequence, and take the ratio of the number of data features in the intersection to the number of data features in the feature subset as the repetition weight; the smaller the repetition weight, the fewer intersections, the less repetition, and the larger the objective function should be. Therefore, the objective function of the feature sequence is calculated based on the repetition weight and the abnormal information amount of the T feature subsets corresponding to the feature sequence. The expression is:
[0062] α represents the repetition weight of the feature sequence, A t represents the amount of abnormal information from the tth feature subset, exp() represents an exponential function with a natural constant as its base, and J represents the objective function of the feature sequence. The larger the objective function, the greater the overall information provided by the T feature subsets constructed from the feature sequence for detecting abnormal stock trading data, and the better the effect of reducing the impact of data imbalance.
[0063] According to the above steps, 30 feature sequences are obtained as the initial population, and the objective function of each is calculated. Using the objective function as the fitness of the feature sequence, an optimization algorithm is used to obtain the optimal feature sequence. The fitness corresponding to the optimal feature sequence is recorded as the optimal fitness, and its corresponding T feature subsets. The optimization algorithm includes but is not limited to the particle swarm optimization algorithm, the fireworks optimization algorithm, the whale optimization algorithm, etc. Specifically, one implementation method of this application uses the particle swarm optimization algorithm, and its maximum number of iterations is set to 50.
[0064] At this point, the optimal feature sequence and its corresponding optimal fitness are obtained.
[0065] Step S004: construct a decision tree based on the data, weight the decision tree according to the optimal fitness and the frequency of occurrence of the data features in the decision tree to obtain the decision value of the classification model; and complete anomaly detection.
[0066] A random forest algorithm was used to build a classification model for stock trading data. By analyzing the data distribution within the training set, the weights of different decision tree values were adjusted to further reduce the impact of data imbalance.
[0067] An optimal feature sequence is obtained for each training set, and each optimal feature sequence is used to form a decision tree. Each decision tree determines T combinations of abnormal stock trading behaviors. In this embodiment, 200 decision trees are constructed in sequence to determine multiple combinations of abnormal stock trading behaviors.
[0068] The training set labels are divided into normal stock trading data and abnormal stock trading data, represented by "-1" and "1" respectively, where "-1" indicates abnormal stock trading data and "1" indicates normal stock trading data. The maximum depth of the decision tree is set to 10. At this point, the stock trading data classification model training is complete.
[0069] Taking into account that the optimal fitness reflects the overall abnormal information amount of data features for the abnormal data in the current training set, the decision value weight is calculated based on the optimal fitness of the corresponding feature sequence to improve the detection accuracy of abnormal stock trading data and reduce the impact of data imbalance.
[0070] Since optimal fitness reflects the overall amount of anomaly information a data feature provides about the abnormal data in the current training set, the greater the fitness, the less impact the decision tree is affected by data imbalance and the higher the classification accuracy. Furthermore, by counting the frequency of each feature value appearing in N decision trees, the greater the frequency of a data feature appearing in a single decision tree, the greater the amount of anomaly information it provides for anomaly detection. This allows it to identify a wider range of stock trading anomalies, assigning a higher weight to the decision tree in which it resides, and improving its ability to detect trading anomalies.
[0071] The weight of the decision value of each decision tree is obtained based on the optimal fitness corresponding to each decision tree and the frequency of each data feature in the decision tree in all decision trees.
[0072] The weight of the decision value of the decision tree is positively correlated with the optimal fitness and the frequency of the data feature appearing in all decision trees.
[0073] Preferably, in this embodiment, the expression of the weight of the decision value of the decision tree is:
[0074] J n Indicates the optimal fitness corresponding to the nth decision tree, F m represents the frequency of the mth data feature in N decision trees, M represents the number of data features in the nth decision tree, norm() represents the normalization function, and w n Represents the weight of the nth decision tree.
[0075] Furthermore, the decision value of each decision tree is weighted based on the weight of each decision tree to obtain the decision value of the classification model.
[0076] Specifically, the expression of the decision value of the classification model is:
[0077] w n Represents the weight of the nth decision tree, s n It represents the decision value of the nth decision tree, N represents the number of decision trees, and S represents the decision value of the classification model output by the random forest algorithm, which represents the classification result of the classification model on stock trading data.
[0078] For each stock trading data to be detected for anomalies, it is input into the classification model, and the decision value in each decision tree at this time is obtained. The decision value is weighted by the weight of the above decision tree to obtain the decision value of the current classification model, which is the decision value of this stock trading data; when the decision value S is not less than zero, the stock trading data is classified as normal data; when the decision value S is a negative number, the stock trading data is classified as abnormal data, and the stock trading data is marked by Dolphin software to remind users that there are abnormal behaviors in stock trading and to pay attention to risk prevention and control.
[0079] At this point, the detection and prevention of abnormal behavior in stock trading have been completed.
[0080] Based on the same inventive concept as the above method, an embodiment of the present invention also provides a system for detecting abnormal behavior and preventing and controlling risks in stock trading, comprising a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of any one of the above-mentioned methods for detecting abnormal behavior and preventing and controlling risks in stock trading are implemented.
[0081] It should be noted that the above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application, and should all be included in the scope of protection of the present application.
[0082] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.
Claims
1. A method for detecting abnormal behavior and preventing and controlling risks in stock trading, characterized in that: The method comprises the following steps: Collecting a number of stock transaction data, forming a feature set from a number of features constituting the stock transaction data, and numbering the features in the feature set to obtain a feature sequence; A training set is constructed by extracting features from all collected stock trading data. Within the training set, features are extracted from the feature set to construct a feature subset, and the feature sequences are sorted to obtain a feature subsequence. The eigenvalues of each data feature are classified into normal data and abnormal data. The information entropy of the data feature is calculated by combining the frequency of occurrence of the eigenvalue in normal data and abnormal data with all eigenvalues corresponding to the data feature. The amount of abnormal information in the feature subset is calculated based on the difference between the information entropy of all data features in the feature subset and the mean information entropy of all data features in the feature set. Extract feature subsequences from T feature subsets to form a feature sequence; use the difference between the number of intersections of the data features of the T feature subsets and the number of data features of the feature subsets as the repetition weight; calculate the objective function of the feature sequence based on the repetition weight and the amount of abnormal information of the feature subsets in the feature sequence; the smaller the repetition weight, the greater the amount of abnormal information and the larger the objective function; use an optimization algorithm to process multiple feature sequences to obtain the optimal feature sequence and optimal fitness; A decision tree is constructed based on the optimal feature sequence of the training set to obtain a classification model. The optimal fitness corresponding to the decision tree and the frequency of occurrence of all data features in the decision tree are forward fused to obtain the weight of the decision tree. The stock trading data to be detected is input into the classification model to obtain the decision value of each decision tree. The decision value of each decision tree is weighted by the weight to obtain the decision value of the classification model. Anomaly detection is completed based on the decision value of the classification model.
2. The method for detecting abnormal behavior and preventing and controlling risks in stock trading according to claim 1, wherein: The number of the training set is N. The number of stock trading data in the training set is the same as the number of all collected stock trading data, and the stock trading data in the training set is repeatable.
3. The method for detecting abnormal behavior and preventing and controlling risks in stock trading according to claim 1, wherein: The method for calculating the information entropy of data features by combining the frequencies of occurrence of eigenvalues in normal data and abnormal data as weights and all eigenvalues corresponding to the data features is: pu i Indicates the frequency of occurrence of the i-th eigenvalue of the data feature in abnormal data; pv i Indicates the frequency of occurrence of the i-th eigenvalue of the data feature in normal data, L k represents the number of eigenvalues in the kth data feature, a1 and a2 represent the weights of the eigenvalues appearing in abnormal data and normal data respectively, and H k Represents the information entropy of the kth data feature in the feature subset.
4. The method for detecting abnormal behavior and preventing and controlling risks in stock trading according to claim 1, wherein: The method for calculating the abnormal information amount of the feature subset based on the difference between the information entropy of all data features in the feature subset and the information entropy mean of all data features in the training set is: The mean of the information entropy of all data features in the feature set is recorded as the first information mean; The amount of abnormal information is positively correlated with the information entropy of the data feature and negatively correlated with the first information mean.
5. The method for detecting abnormal behavior and preventing and controlling risks in stock trading according to claim 1, wherein: The target component is obtained by calculating the ratio of the number of data features in the intersection of the data features of the T feature subsets corresponding to the feature sequence to the number of data features in the feature subset according to the repetition weight.
6. The method for detecting abnormal behavior and preventing and controlling risks in stock trading according to claim 1, wherein: The method for calculating the objective function of the feature sequence based on the repetition weight and the abnormal information amount of the feature subset in the feature sequence is: α represents the repetition weight of the feature sequence, A t represents the amount of abnormal information of the t-th feature subset, exp() represents the exponential function with a natural constant as the base, and J represents the objective function of the feature sequence.
7. The method for detecting abnormal behavior and preventing and controlling risks in stock trading according to claim 1, wherein: The method for obtaining the weight of the decision tree by forward fusing the optimal fitness corresponding to the decision tree and the frequency of occurrence of all data features in the decision tree is: J n Indicates the optimal fitness corresponding to the nth decision tree, F m represents the frequency of the mth data feature in N decision trees, M represents the number of data features in the nth decision tree, norm() represents the normalization function, and w n Represents the weight of the nth decision tree; each training set corresponds to a decision tree, and the number of decision trees is the same as the number of training sets.
8. The method for detecting abnormal behavior and preventing and controlling risks in stock trading according to claim 1, wherein: The methods for anomaly detection are: Input the stock trading data to be detected into the classification model; obtain the decision value in each decision tree at this time; weight the decision value by the weight of the decision tree to obtain the decision value of the current classification model; when the decision value is not less than zero, classify the stock trading data as normal data; when the decision value is negative, classify the stock trading data as abnormal data.
9. The method for detecting abnormal behavior and preventing and controlling risks in stock trading according to claim 1, wherein: The method for the decision value of the classification model is: w n Represents the weight of the nth decision tree, s n Represents the decision value of the nth decision tree, N represents the number of decision trees, and S represents the decision value of the random forest algorithm output classification model.
10. A system for detecting abnormal behavior and preventing and controlling risks in stock trading, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method for detecting abnormal behavior and preventing and controlling risks in stock trading as described in any one of claims 1 to 9 are implemented.