Enterprise Data Management Platform
By using isolated forest algorithm and target clustering algorithm in the enterprise data management platform, abnormal judgments are made on enterprise data, which solves the problem that abnormal data in enterprise data is difficult to judge, and improves the accuracy and reliability of abnormal data detection.
Patent Information
- Application Number
- CN202510272054.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-10
AI Technical Summary
Abnormal data in enterprise data is not easy to judge, and the judgment of the existing technology is low in accuracy and is easy to misjudgment.
It provides an enterprise data management platform, including a first exception judgment module, an outlier value determination module and a second exception judgment module, and uses an isolated forest algorithm and a target clustering algorithm to judge and verify the abnormal data.
By quickly identifying potential abnormal data and accurately capturing abnormal abnormalities related to logic or data patterns, the accuracy and reliability of abnormal data detection are improved.
Smart Images

Figure CN119783008B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data processing, and particularly to an enterprise data management platform. Background Art
[0002] With the rapid development of information technology, the amount of data in various industries has grown explosively. At the enterprise operation level, data covers multiple key aspects such as production, sales, finance, and customer relationship, and has become an important basis for enterprise decision-making and development.
[0003] However, among the numerous data of an enterprise, the existence of abnormal data cannot be ignored. It may be due to the failure of data collection devices, such as sensors malfunctioning and recording incorrect values; or it may be due to human factors, such as negligence during data entry. Abnormal data will seriously interfere with the results of data analysis and lead to incorrect decisions. Existing technologies can perform preliminary data processing, but there are still many deficiencies in the monitoring of abnormal data, and the accuracy of judging abnormal data is relatively low, and misjudgment is likely to occur.
[0004] Therefore, there is an urgent need for a stable and reliable enterprise data management platform. Summary of the Invention
[0005] Embodiments of the present disclosure provide an enterprise data management platform to solve the problem that it is difficult to judge abnormal data in enterprise data.
[0006] Embodiments of the present disclosure provide an enterprise data management platform, including: a first abnormal data judgment module, configured to judge target data based on an isolation forest algorithm model to determine first abnormal data;
[0007] An abnormal value determination module, configured to, in response to the first abnormal data satisfying a first relevant condition, determine a first target abnormal value based on relevant data of the first abnormal data; and further configured to, in response to the first abnormal data not satisfying the first relevant condition, determine a second target abnormal value based on a target clustering algorithm;
[0008] A second abnormal data judgment module, configured to judge whether the first abnormal data is real data based on the first target abnormal value, or judge whether the first abnormal data is real data based on the second target abnormal value, and send the judgment result to a target device.
[0009] In an exemplary embodiment of the present disclosure, the enterprise data management platform further includes: a relevant data determination module;
[0010] The relevant data determination module is configured to determine multiple relevant data of the first abnormal data, and weights corresponding to the multiple relevant data;
[0011] The abnormal value determination module is specifically further configured to determine a first predicted data based on the multiple relevant data and the weights corresponding to the multiple relevant data;
[0012] Determine a first target outlier based on the difference value between the first prediction data and the first outlier data.
[0013] In an exemplary embodiment of the present disclosure, the relevant data determination module is further specifically configured to:
[0014] Introduce the data in the first relevant data set into the first linear regression model one by one as independent variables, and determine a second linear regression model based on a first criterion; introduce the independent variables in the second linear regression model one by one, and determine a third linear regression model based on a second criterion; until a first condition is met to obtain a target linear regression model;
[0015] The first linear regression model is a preset linear regression model; the dependent variable of the first linear regression model is the first outlier data; the independent variables in the target linear regression model are the relevant data of the first outlier data;
[0016] Normalize the parameters of the independent variables in the target linear regression model to obtain the weights corresponding to the relevant data; wherein, the first relevant data set is a data set composed of data pre-determined to be related to the first outlier data.
[0017] In an exemplary embodiment of the present disclosure, the relevant data determination module is further specifically configured to:
[0018] Determine a first significance probability value of the introduced variable based on a target test rule;
[0019] In response to the first significance probability value being less than a first probability value, add the introduced variable to the first linear regression model to obtain a second linear regression model;
[0020] The relevant data determination module is further specifically configured to:
[0021] Determine a second significance probability value of the removed variable based on a target test rule;
[0022] In response to the second significance probability value being greater than a second probability value, remove the removed variable from the second linear regression model to obtain a third linear regression model.
[0023] In an exemplary embodiment of the present disclosure, the enterprise data management platform further includes: a standard adjustment module;
[0024] The standard adjustment module is configured to, in response to the number of data in the first relevant data set being greater than a first relevant number, reduce the first probability value based on a first adjustment step;
[0025] In response to the number of independent variables in the second linear regression model being greater than a second relevant number, increase the second probability value based on a second adjustment step.
[0026] In an exemplary embodiment of the present disclosure, the enterprise data management platform further includes: a relevant condition determination module;
[0027] The relevant condition determination module is configured to determine that the first abnormal data meets the first relevant condition in response to the number of relevant data of the first abnormal data being greater than the first relevant number or the importance level of the first abnormal data being greater than the preset level.
[0028] In an exemplary embodiment of the present disclosure, the outlier determination module is specifically further configured to:
[0029] Perform clustering processing on the data set composed of the first abnormal data and the historical first abnormal data based on the target clustering algorithm to obtain a target clustering result;
[0030] In response to the first abnormal data being in the target cluster and the distance between the first abnormal data and the clustering center of the target cluster being less than the first distance, determine a second target outlier based on the first method;
[0031] In response to the first abnormal data not being in the target cluster and / or the distance between the first abnormal data and the clustering center of the target cluster being greater than or equal to the first distance, determine a second target outlier based on the second method;
[0032] The calculation criteria of the first method and the second method are different.
[0033] In an exemplary embodiment of the present disclosure, the enterprise data management platform further includes: a distance adjustment module;
[0034] The distance adjustment module is configured to reduce the first distance based on the first distance adjustment step in response to the density of the target cluster being greater than the first density;
[0035] In response to the density of the target cluster being less than the second density, increase the first distance based on the second distance adjustment step.
[0036] In an exemplary embodiment of the present disclosure, the enterprise data management platform further includes: a first data processing module;
[0037] The first data processing module is configured to send abnormal information to the first device in response to the first abnormal data not being real data.
[0038] In an exemplary embodiment of the present disclosure, the enterprise data management platform further includes: a second data processing module;
[0039] The second data processing module is configured to control the second device to store the first abnormal data and the relevant data of the first abnormal data in response to the first abnormal data not being real data.
[0040] The beneficial effects of the enterprise data management platform provided by the embodiments of the present disclosure are as follows:
[0041] Through the first anomaly judgment module, the present disclosure uses the Isolation Forest algorithm model to preliminarily screen the target data, and can quickly and effectively identify potential first anomaly data. The present disclosure takes into account the diversity of anomaly data. When the first anomaly data meets specific first related conditions, the first target anomaly value is determined based on the relevant information of these data, which helps to accurately capture anomalies related to logical or data patterns. When the first anomaly data does not meet these conditions, the target clustering algorithm is used to determine the second target anomaly value, improving the accuracy and reliability of anomaly data detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0043] Figure 1 It is a schematic structural diagram of the first enterprise data management platform provided by the embodiments of the present disclosure;
[0044] Figure 2 It is a schematic structural diagram of the second enterprise data management platform provided by the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] In order to enable those skilled in the art of this technology to better understand this solution, the following will clearly describe the technical solutions in the embodiments of this solution in conjunction with the drawings in the embodiments of this solution. Obviously, the described embodiments are part of the embodiments of this solution, rather than all of the embodiments. Based on the embodiments in this solution, all other embodiments obtained by those of ordinary skill in the art without creative efforts should fall within the scope of protection of this solution.
[0046] The terms "including" and any other variations in the specification and claims of this solution, as well as the above drawings, mean "including but not limited to", intending to cover non-exclusive inclusion and not limited to the examples listed in the text. In addition, terms such as "first" and "second" are used to distinguish different objects, rather than to describe a specific order.
[0047] The following will describe the implementation of the present disclosure in detail with reference to specific drawings:
[0048] Figure 1 It is a schematic structural diagram of the first enterprise data management platform provided by the embodiments of the present disclosure. Referring to Figure 1 , the enterprise data management platform includes:
[0049] The first anomaly judgment module 10 is used to judge the target data based on the isolation forest algorithm model to determine the first anomaly data.
[0050] The outlier determination module 11 is used to, in response to the first anomaly data satisfying the first relevant condition, determine the first target outlier based on the relevant data of the first anomaly data; and is also used to, in response to the first anomaly data not satisfying the first relevant condition, determine the second target outlier based on the target clustering algorithm.
[0051] The second anomaly judgment module 12 is used to judge whether the first anomaly data is real data based on the first target outlier, or judge whether the first anomaly data is real data based on the second target outlier, and send the judgment result to the target device.
[0052] In this embodiment, the isolation forest algorithm model can process and judge the data in the target data and its historical data. The training process of the isolation forest algorithm is relatively simple and efficient. It mainly constructs a tree structure through random sampling and recursive partitioning, and does not require a large number of parameter iterative updates like some deep learning models. Therefore, in the face of a large-scale data set, the isolation forest algorithm can complete training in a relatively short time and quickly give the anomaly detection result. Therefore, in this article, the isolation forest algorithm is selected to process a large amount of enterprise data and make a preliminary judgment of anomalies. The parameters in the training process of the isolation forest algorithm can be set according to the inherent parameters of the isolation forest model or the parameters when solving similar problems.
[0053] The target data refers to one of the data received by the enterprise central data platform, or can also be the data in a pre-determined data list. Since the amount of data collected by the enterprise is relatively large, judging the outlier of each data will consume more resources and time. Therefore, only some of the data can be judged and processed.
[0054] The data obtained by processing and judging the target data and its historical data through the above isolation forest algorithm model is the first anomaly data.
[0055] Considering that some data itself has relatively strong volatility or has large fluctuations due to certain objective reasons, rather than due to manual filling errors or the subjective intention of the submitter to hide, the first anomaly data is only initially judged as data that may have anomalies, not necessarily abnormal data. The purpose of the outlier determination module 11 is to determine the outlier of the first anomaly data. The larger the outlier, the greater the possibility that the first anomaly data is abnormal false data.
[0056] For different data, different methods should be adopted to determine outliers. If the first abnormal data meets the first relevant condition, the first target outlier can be determined based on the relevant data of the first abnormal data, and the first target outlier is the outlier corresponding to the first abnormal data. Conversely, if the first abnormal data does not meet the first relevant condition, the second target outlier can be determined based on the target clustering algorithm, and the second target outlier is also the outlier corresponding to the first abnormal data. The only difference between the first target outlier and the second target outlier lies in the different determination methods.
[0057] The first relevant condition can be judged in the following way: In an embodiment of the present disclosure, the enterprise data management platform further includes: a relevant condition determination module 15.
[0058] The relevant condition determination module 15 is configured to determine that the first abnormal data meets the first relevant condition in response to the number of relevant data of the first abnormal data being greater than the first relevant number, or the importance level of the first abnormal data being greater than the preset level.
[0059] In this embodiment, the first relevant condition is set based on the characteristics of two different means. Considering that a large number of relevant data means more information dimensions to describe and understand the first abnormal data. These rich data can reflect the characteristics, relationships, trends, etc. of the first abnormal data from different angles, and thus the analysis and prediction based on these data are more comprehensive and accurate. For example, in enterprise sales data, when the amount of data such as customer information, product information, and market environment information related to a certain abnormal sales data is very large, the reason for the occurrence of this abnormal sales data can be analyzed more accurately, and it can be judged whether it is really abnormal and the degree of abnormality.
[0060] Secondly, if the importance level of the first abnormal data itself is high, it indicates that it plays a key role in the enterprise's business decisions, analysis, etc., and its abnormality or non-abnormality may have a significant impact. For example, the core financial index data of an enterprise, the key performance data of major projects, etc. For these important data, more refined and targeted analysis methods are required. Determining outliers based on their relevant data can more fully explore the information behind the data, ensure accurate judgment of their abnormal conditions, and avoid major decision-making mistakes caused by incorrect judgments.
[0061] Therefore, when the number of related data of the first abnormal data is greater than the first related number, or the importance level of the first abnormal data is greater than the preset level, the first abnormal data meets the first related condition. The first related number and the preset level can be determined in advance according to experience or the actual scenario. The importance level of the first abnormal data can be preset by relevant personnel during the operation of the enterprise. For example, data related to finance, key performance, or sales can be appropriately preset with a higher importance level, while data that is outdated or invalid can be preset with a lower importance level.
[0062] In this embodiment, the purpose of the second abnormal judgment module 12 is to determine whether the first abnormal data is real data. Real data should be understood as error-free data that is not caused by human error. The data opposite to it is false data.
[0063] Considering that the abnormal value of the first abnormal data may be the first target abnormal value or the second target abnormal value, and the sources of the first target abnormal value and the second target abnormal value are different, that is, the obtaining methods are different, different thresholds should also be set to determine whether the first abnormal data is real data.
[0064] For example, in response to the first target abnormal value being greater than or equal to the first abnormal threshold, the first abnormal data is not real data.
[0065] In response to the first target abnormal value being less than the first abnormal threshold, the first abnormal data is real data.
[0066] In response to the second target abnormal value being greater than or equal to the second abnormal threshold, the first abnormal data is not real data.
[0067] In response to the second target abnormal value being less than the second abnormal threshold, the first abnormal data is real data.
[0068] The first abnormal threshold and the second abnormal threshold can be adjusted according to experience or during the experimental process.
[0069] Through the judgment of the second abnormal judgment module 12, it can be obtained whether the first abnormal data is real data, and the judgment result can also be sent to relevant devices, which can be a computer or mobile phone dedicated to enterprise managers.
[0070] As can be seen from the above, through the first anomaly judgment module 10 of the present disclosure, the target data is preliminarily screened using the isolation forest algorithm model, and potential first anomaly data can be quickly and effectively identified. The present disclosure takes into account the diversity of anomaly data. When the first anomaly data meets specific first relevant conditions, the first target anomaly value is determined based on the relevant information of these data, which helps to accurately capture anomalies related to logical or data patterns. When the first anomaly data does not meet these conditions, the target clustering algorithm is used to determine the second target anomaly value, improving the accuracy and reliability of anomaly data detection.
[0071] In the foregoing description, it can be obtained that if the number of relevant data of the first anomaly data is greater than the first relevant number or the importance level of the first anomaly data is greater than the preset level, the judgment of the first anomaly data is determined according to the data related to it. The relevant data of the first anomaly data and the first target anomaly value can be determined in the following manner:
[0072] Figure 2 It is a schematic structural diagram of the second enterprise data management platform provided by the embodiments of the present disclosure. Refer to Figure 2 In an embodiment of the present disclosure, the enterprise data management platform further includes: a relevant data determination module 13.
[0073] The relevant data determination module 13 is specifically configured to determine multiple relevant data of the first anomaly data and the weights corresponding to the multiple relevant data.
[0074] The anomaly value determination module 11 is further specifically configured to determine first prediction data based on the multiple relevant data and the weights corresponding to the multiple relevant data.
[0075] Determine the first target anomaly value based on the difference value between the first prediction data and the first anomaly data.
[0076] In this embodiment, it is possible to first determine which relevant data of the first anomaly data and their corresponding weights. The weight is the degree of influence of the relevant data on the first anomaly data. Then, a predicted value is determined according to the relevant data and its corresponding weight, and then the difference value between the predicted value and the first anomaly data is calculated, and the first target anomaly value is determined according to this difference value.
[0077] The relevant data of the first anomaly value and its corresponding weight can be determined in the following manner:
[0078] In an embodiment of the present disclosure, the relevant data determination module 13 is further specifically configured to:
[0079] One by one, introduce the data in the first relevant data set as independent variables into the first linear regression model, and determine the second linear regression model based on the first criterion; one by one, introduce the independent variables in the second linear regression model, and determine the third linear regression model based on the second criterion; until the first condition is met, the target linear regression model is obtained.
[0080] The first linear regression model is a preset linear regression model; the dependent variable of the first linear regression model is the first abnormal data; the independent variables in the target linear regression model are the relevant data of the first abnormal data.
[0081] Normalize the parameters of the independent variables in the target linear regression model to obtain the weights corresponding to the relevant data; among them, the first relevant data set is a data set composed of data determined in advance and related to the first abnormal data.
[0082] In this embodiment, it should be noted that the execution time of the relevant data determination module 13 is before all steps, that is, the relevant data and corresponding weights of each data should be determined in advance by the relevant data determination module 13, otherwise the judgment of the first relevant condition in the outlier determination module cannot be executed.
[0083] The first relevant data set is a data set composed of data determined in advance and related to the first abnormal data. It should be understood as a data set composed of data that may be related to the first abnormal data determined artificially through common sense or experience. It is a data set that may contain irrelevant data. It should be noted that when determining the first relevant data set, if there is data whose relevance cannot be determined, it should also be included in the first relevant data set to avoid inaccurate subsequent judgments.
[0084] The first linear regression model is a preset linear regression model. The dependent variable of the first linear regression model is the first abnormal data and has no independent variables. The first linear regression model can be , where is the first abnormal data, is the intercept, is the error term.
[0085] Introduction process: Select independent variables one by one from the first relevant data set , introduce it into the model to obtain , use the least squares method to estimate the model parameters , determine whether the introduced variable should remain in the first linear regression model based on the first criterion until the last data in the first relevant data set is introduced, and the obtained linear regression model is the second linear regression model.
[0086] Ejection process: Based on the second linear regression model with multiple independent variables already introduced, introduce each independent variable one by one , determine whether the introduced variable should be removed from the second linear regression model based on the second criterion until the last independent variable introduced into the second linear regression model is obtained, and obtain the third linear regression model;
[0087] Determine whether the first condition is satisfied. If the first condition is not satisfied, repeat the above introduction process and extraction process until the first condition is satisfied to obtain the target linear regression model.
[0088] In this process, the first criterion refers to determining whether to add the introduced variable to the variables of the first linear regression model through the target test rule, and the second criterion refers to determining whether to add the extracted variable to the variables of the first linear regression model through the target test rule. Specifically:
[0089] In an embodiment of the present disclosure, the relevant data determination module 13 is further specifically configured to determine the first significance probability value of the introduced variable based on the target test rule.
[0090] In response to the first significance probability value being less than the first probability value, add the introduced variable to the first linear regression model to obtain the second linear regression model.
[0091] The relevant data determination module 13 is further specifically configured to determine the second significance probability value of the extracted variable based on the target test rule.
[0092] In response to the second significance probability value being greater than the second probability value, remove the extracted variable from the second linear regression model to obtain the third linear regression model.
[0093] In this embodiment, the target test rule can be a t-test or an F-test. In this embodiment, considering that the t-test is mainly used to test whether a single independent variable has a significant effect on the dependent variable, and the F-test is mainly used to test the significance of the entire regression model, that is, to determine whether all independent variables as a whole have a significant effect on the dependent variable. Therefore, the t-test method is selected as the target test rule to obtain the significance probability value, which is the first significance probability value of the introduced variable or the second significance probability value of the extracted variable.
[0094] The first significance probability value and the second significance probability value are essentially the same, both being the significance probability values of their corresponding data.
[0095] In the introduction process, starting from a model without any independent variables, consider introducing independent variables into the model one by one. For each independent variable to be selected, calculate its corresponding value. If the first significance probability value is less than the preset first probability value for introduction, introduce the introduced variable into the model. Then continue to perform the same test on the remaining independent variables until no independent variable can meet the introduction condition.
[0096] During the derivation process, the independent variables in the obtained second linear regression model are derived one by one. Calculate the significance probability value corresponding to each independent variable. If the second significance probability value is greater than the preset second probability value for elimination, the variable is eliminated from the model. Then continue to test the remaining independent variables until no independent variable meets the elimination condition.
[0097] The smaller the significance probability value, the more significant the influence of the variable on the independent variable, and it can be introduced into the model. On the contrary, the larger the significance probability value, the less significant the influence of the variable on the independent variable, and it can be derived from the model.
[0098] After normalizing the parameters of each independent variable in the finally obtained target regression model, they can be used as their corresponding weights. The positive or negative sign of the parameter estimation value of the independent variable can intuitively reflect the relationship direction between the independent variable and the dependent variable. If the parameter estimation value is positive, it indicates a positive correlation between the independent variable and the dependent variable, that is, when the independent variable increases, the dependent variable also tends to increase; if the parameter estimation value is negative, it indicates a negative correlation. The magnitude of the absolute value of the parameter reflects to a certain extent the relative magnitude of the influence of the independent variable on the dependent variable. The larger the absolute value, the more significant the influence of the independent variable on the dependent variable. Therefore, it can be normalized and used as its corresponding weight.
[0099] Through the above derivation and description, the target regression model of each data can be obtained. When calculating the first target outlier, the target regression model corresponding to the first abnormal data can be obtained according to the pre-stored target regression model corresponding to each data, and then the relevant data of the current first abnormal data is input into the target regression model, and the value of the dependent variable obtained is the first predicted data. The first target outlier can be determined according to the difference value between the first predicted data and the first abnormal data. The difference value between the first predicted data and the first abnormal data can be the absolute value of the data difference after normalizing the first predicted data and the first abnormal data. The purpose of normalization is to eliminate the problem of inconsistent evaluation criteria caused by different units or data base quantities of different data. For example, the first target outlier can be obtained through Table 1:
[0100]
[0101] As can be seen from the above, the present disclosure can further process the first abnormal data by introducing the relevant data determination module 13 and the outlier determination module 11, and further considers other data related to the abnormal data and their weights, so as to more accurately determine the outlier. In this embodiment, the relevant data determination module 13 uses a linear regression model (including the introduction and derivation processes) to determine the data related to the first abnormal data and their weights. In this embodiment, the significance probability value is used to judge which data have a significant impact on the abnormal data, so as to allocate reasonable weights to them, improving the accuracy and reliability of abnormal data detection.
[0102] In one embodiment of the present disclosure, the enterprise data management platform further includes: a standard adjustment module 14.
[0103] The standard adjustment module 14 is configured to, in response to the number of data in the first related data set being greater than the first related number, reduce the first probability value based on the first adjustment step.
[0104] In response to the number of independent variables in the second linear regression model being greater than the second related number, increase the second probability value based on the second adjustment step.
[0105] In this embodiment, considering that when the number of data in the first related data set is greater than the first related number, it means that there are more independent variables to choose from. To avoid introducing too many insignificant variables and causing model overfitting, the first probability value can be appropriately reduced, and the adjusted value is the first adjustment step. The first adjustment step can be determined by the first formula.
[0106] The first formula can be: , where represents the first adjustment step, is the first adjustment coefficient, and its value range can be adjusted according to the actual situation, and is used to control the overall size of the adjustment step, represents the number of data in the first related data set, is the first related number, represents the variance of the first related data set, represents the number of existing independent variables in the current regression model. The first related number can be preset according to experience, is a preset maximum number of independent variables, which is used to normalize the impact on the model complexity. It should be noted that the value of should be related to the number of data in the first related data set. The larger the number of data in the first related data set, the larger the value of
[0107] The logic of the first formula is: the degree of excess of the number of data in the data set is , when the number of data in the first related data set exceeds the first related number When there are more, it means there are more independent variables to choose from. In stepwise regression, if not strictly controlled, it is easy to introduce too many insignificant variables, leading to overfitting of the model. An overfitted model performs well on the training data but has poor generalization ability on new data. Therefore, the greater the degree of excess, the larger the adjustment step size is required to reduce the first probability value, making the criteria for introducing variables more stringent, so as to screen out the variables that truly have a significant impact on the dependent variable.
[0108] Variance of the dataset It reflects the degree of dispersion of the data. The larger the variance, the greater the fluctuation of the data, and the more complex and unstable the relationship between variables. In this case, more caution is needed when introducing variables because a seemingly significant variable may only show a correlation with the dependent variable due to random fluctuations in the data, rather than a true causal relationship. Therefore, the larger the variance, the corresponding larger the adjustment step size should be to reduce the risk of introducing falsely significant variables.
[0109] Complexity of the model , measured by the number of independent variables already existing in the current model to measure the complexity of the model. The more complex the model (i.e., is larger), the more likely it is to have the problem of overfitting. As the number of independent variables increases, the model will overfit the noise and outliers in the training data and ignore the true patterns of the data. Therefore, when the model complexity is relatively high, a larger adjustment step size is needed to reduce the first probability value and strictly control the introduction of new variables to avoid the model becoming too complex.
[0110] Similarly, the second adjustment step size is used to increase the second probability value when the number of independent variables in the second linear regression model is greater than the second relevant quantity. The second adjustment step size can be determined according to the second formula.
[0111] The second formula can be: , where represents the second adjustment step size, is the second adjustment coefficient, and its value range can be adjusted according to the actual situation, used to control the overall size of the adjustment step size, represents the number of independent variables in the second linear regression model, is the second relevant quantity, represents the covariance matrix of the independent variables in the second linear regression model, and its determinant is , the coefficient of determination after each introduction in the current regression model is (which can be calculated based on, for example, the Akaike information criterion or the Bayesian information criterion), is a very small positive number used to avoid the case where the determinant of the covariance matrix is zero, for example, it can be 0.001. The second relevant quantity can be preset based on experience.
[0112] The logic of the second formula is: the degree to which the number of data in the dataset exceeds , when the number of independent variables in the second linear regression model exceeds the second correlation quantity by a large amount, it indicates that some unimportant independent variables are included in the current model. In the backward elimination process of stepwise regression, if the criterion is too strict, some variables that contribute to the model may be wrongly eliminated, resulting in the model losing important information and the fitting effect deteriorating. Therefore, the greater the degree of excess, the larger the adjustment step size is needed to increase the second probability value, making the criterion for eliminating variables more lenient and retaining more potentially useful variables.
[0113] The determinant of the covariance matrix of the dataset , the determinant of the covariance matrix reflects the degree of correlation between independent variables. The smaller the determinant value, the stronger the correlation between independent variables and the more likely the problem of multicollinearity occurs. Multicollinearity will lead to unstable estimation of model parameters and make the significance test results of some variables unreliable. In this case, more caution is needed when eliminating variables because a variable may seem insignificant due to the influence of multicollinearity, but in fact it has an impact on the dependent variable. Therefore, the smaller the determinant value, the larger the adjustment step size should be to increase the second probability value and reduce the possibility of wrongly eliminating variables.
[0114] The goodness of fit of the model , which is measured by the adjusted (i.e., each time introduced) . The higher the goodness of fit, the stronger the ability of the model to explain the data. When the model already has a good fitting effect, eliminating variables may destroy this good fit and lead to a decline in model performance. Therefore, the higher the goodness of fit, the larger the adjustment step size should be to increase the second probability value and eliminate variables more cautiously to ensure the stability and accuracy of the model.
[0115] It can be concluded from the above that in this embodiment, when the number of data in the first correlation dataset exceeds the preset first correlation number, the standard adjustment module 14 will appropriately reduce the first probability value according to the first adjustment step. Considering the scale, variance of the dataset and the complexity of the current model, it ensures that only variables that truly have a significant impact on the dependent variable will be introduced into the model, effectively avoiding the overfitting problem caused by introducing too many insignificant variables and improving the generalization ability of the model. In this embodiment, when the number of independent variables in the second linear regression model exceeds the preset second correlation number, the standard adjustment module 14 will increase the second probability value according to the second adjustment step. Considering the degree of correlation of independent variables in the dataset and the goodness of fit of the model, it ensures that important variables will not be wrongly excluded, helps to retain useful information in the model, prevents the performance degradation of the model caused by excluding too many variables, and improves the accuracy and reliability of abnormal data detection.
[0116] In an embodiment of the present disclosure, the outlier determination module 11 is further specifically configured to:
[0117] Perform clustering processing on the data set composed of the first abnormal data and the historical first abnormal data based on the target clustering algorithm to obtain a target clustering result;
[0118] In response to the first abnormal data being in the target cluster and the distance between the first abnormal data and the clustering center of the target cluster being less than the first distance, determine the second target outlier based on the first method;
[0119] In response to the first abnormal data not being in the target cluster and / or the distance between the first abnormal data and the clustering center of the target cluster being greater than or equal to the first distance, determine the second target outlier based on the second method;
[0120] The calculation criteria of the first method and the second method are different.
[0121] In this embodiment, the target clustering algorithm may be a density-based clustering algorithm, and the density-based clustering algorithm performs clustering operations on the data set composed of the first abnormal data and the historical first abnormal data. The purpose of clustering is to divide data points into different clusters so that data points within the same cluster have high similarity, while data points in different clusters are quite different.
[0122] Considering that the Isolation Forest mainly judges anomalies based on the isolation degree of data points, when the data distribution is relatively complex, especially when there are regions with uniform density, data points at the edges or relatively sparse positions of these regions will be misjudged as anomaly points. Because the Isolation Forest is not very good at capturing the local density characteristics of data, for those data points that seem isolated globally but are normal within the locally uniform density regions, it may give incorrect judgments. The Density-Based Spatial Clustering of Applications with Noise (DBSCAN) is a density-based spatial clustering algorithm that identifies clusters and noise points by defining the density of data points. In regions with uniform density, DBSCAN can accurately divide the data points within these regions into a cluster and can well distinguish whether the points at the density change boundary belong to the cluster or are anomalies. It can discover clusters of any shape and is very sensitive to local density changes in the data, and can effectively identify the regions with uniform density in the data and the anomalies within them. Therefore, in this embodiment, it is considered to use the DBSCAN algorithm, and through clustering processing, the target clustering result is obtained, that is, the data set is divided into several clusters. Among them, the target cluster is the cluster with the smallest standard deviation of the data in the cluster. The smaller the standard deviation, the more uniform the density. Therefore, the clustering result of DBSCAN and the anomaly detection result of the Isolation Forest are comprehensively evaluated. If the point is within the density-uniform cluster identified by DBSCAN and the distance from the cluster center is within a certain range, then it can be considered that this point may not be a real anomaly point but a misjudgment caused by the limitations of the Isolation Forest algorithm, and it is removed from the anomaly point set. On the contrary, if the noise points identified by DBSCAN also have a high anomaly score in the Isolation Forest algorithm, then these points can be further confirmed as anomaly points to improve the accuracy of anomaly detection.
[0123] Therefore, when the first anomaly data is in the target cluster and the distance between the first anomaly data and the cluster center of the target cluster is less than the first distance, the second target anomaly value is determined based on the first method. When the first anomaly data is not in the target cluster and / or the distance between the first anomaly data and the cluster center of the target cluster is greater than or equal to the first distance, the second target anomaly value is determined based on the second method. The first distance can be determined according to experience.
[0124] Both the first method and the second method can determine the second anomaly value based on the distance between the first anomaly data and the cluster center of the target cluster, but the difference is that the calculation criteria of the first method and the second method are different, which can be specifically reflected in the different slopes of the formulas or the different corresponding relationships of the mapping tables.
[0125] For example, the first method is the third formula, specifically it can be , the second method may be the fourth formula, specifically it may be , where represents the second outlier, respectively represent the first slope and the second slope, which can be determined according to experiments, represents the initial outlier, which can be determined according to experiments.
[0126] In an embodiment of the present disclosure, the enterprise data management platform further includes: a distance adjustment module 16;
[0127] The distance adjustment module 16 is configured to, in response to the density of the target cluster being greater than the first density, reduce the first distance based on the first distance adjustment step;
[0128] In response to the density of the target cluster being less than the second density, increase the first distance based on the second distance adjustment step.
[0129] In this embodiment, considering that for clusters with higher density, the distance between data points is relatively small, the distance threshold should be set smaller to set the first distance more strictly; while for clusters with lower density, the distance threshold should be increased accordingly to set the first distance more loosely.
[0130] Therefore, when the density of the target cluster is greater than the first density, the first distance is reduced based on the first distance adjustment step; when the density of the target cluster is less than the second density, the first distance is increased based on the second distance adjustment step. The first distance adjustment step and the second distance adjustment step can be determined according to experiments.
[0131] It can be concluded from the above that the present disclosure uses the DBSCAN algorithm through the outlier determination module 11 to cluster the first abnormal data and the historical first abnormal data, and combines the outlier detection results of the isolation forest algorithm, which can make up for the misjudgment that may occur in the isolation forest under complex data distributions. Especially for data points in a region with uniform density, DBSCAN can accurately identify their clustering attributes, thus avoiding misjudging these normal points as outliers. In this embodiment, the distance adjustment module 16 can dynamically adjust the first distance according to the density of the target cluster. For clusters with higher density, the module will reduce the first distance to screen outliers more strictly; while for clusters with lower density, the module will increase the first distance to set the outlier standard more loosely, making the outlier detection more in line with the actual situation of the data and improving the accuracy and reliability of the outlier data detection.
[0132] In an embodiment of the present disclosure, the enterprise data management platform further includes: a first data processing module 17;
[0133] The first data processing module 17 is configured to, in response to the first abnormal data not being real data, send outlier information to the first device.
[0134] The enterprise data management platform further includes: a second data processing module 18;
[0135] The second data processing module 18 is configured to control the second device to store the first abnormal data and the related data of the first abnormal data in response to the first abnormal data not being real data.
[0136] In this embodiment, after the first data processing module 17 obtains that the first abnormal data is not real data, it can send the related abnormal information to the first device, which can be the mobile phone of the relevant person or the related device of the data uploader, to remind them to make corrections or conduct a second review.
[0137] The second data processing module 18 can store the first abnormal data and the related data of the first abnormal data as evidence. In this application, the related data of the first abnormal data is defaulted to be real data.
[0138] It can be concluded from the above that when the first data processing module 17 in the present disclosure determines that the first abnormal data is not real data, it can send abnormal information to the first device, which is crucial for timely discovering and correcting data errors, because it can ensure that relevant personnel (such as data uploaders, data reviewers, or data administrators) obtain error information in a timely manner, so as to take necessary correction measures. The second data processing module 18 can store the first abnormal data and its related data as evidence, which is convenient for enterprise management.
[0139] The above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them; although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present disclosure.
Claims
1. An enterprise data management platform, characterized in that: include: A first abnormality judgment module, used to judge the target data based on the isolation forest algorithm model to determine the first abnormal data; an abnormal value determination module, configured to determine a first target abnormal value based on the correlation data of the first abnormal data in response to the first abnormal data satisfying a first correlation condition; further configured to determine, in response to the first abnormal data not satisfying the first correlation condition, a second target abnormal value based on a target clustering algorithm; A second abnormality judgment module, configured to judge whether the first abnormal data is real data based on the first target abnormal value, or judge whether the first abnormal data is real data based on the second target abnormal value, and send the judgment result to a target device; The outlier determination module is further specifically used for: Performing clustering processing on a data set consisting of the first abnormal data and the historical first abnormal data based on the target clustering algorithm to obtain a target clustering result; In response to the first abnormal data being in a target cluster and the distance between the first abnormal data and the cluster center of the target cluster being less than a first distance, determining the second target abnormal value based on a first manner; In response to the first abnormal data not being in the target cluster and / or the distance between the first abnormal data and the cluster center of the target cluster being greater than or equal to the first distance, determining the second target abnormal value based on a second manner; The calculation standards of the first method and the second method are different.
2. The enterprise data management platform according to claim 1, characterized in that: Also includes: Related data determination module; The related data determination module is used to determine a plurality of related data of the first abnormal data, and weights corresponding to the plurality of related data; The outlier determination module is further configured to determine first prediction data based on the plurality of related data and weights corresponding to the plurality of related data; A first target abnormal value is determined based on a difference value between the first predicted data and the first abnormal data.
3. The enterprise data management platform according to claim 2, characterized in that: The relevant data determination module is specifically used for: Introducing the data in the first related data set as independent variables into the first linear regression model one by one, and determining the second linear regression model based on the first standard; deriving the independent variables in the second linear regression model one by one, and determining the third linear regression model based on the second standard; until the first condition is met, the target linear regression model is obtained; The first linear regression model is a preset linear regression model; the dependent variable of the first linear regression model is the first abnormal data; The independent variable in the target linear regression model is the related data of the first abnormal data; The parameters of the independent variables in the target linear regression model are normalized to obtain weights corresponding to the relevant data; wherein the first relevant data set is a data set consisting of predetermined data related to the first abnormal data.
4. The enterprise data management platform according to claim 3, characterized in that: The relevant data determination module is also specifically used for: Determine the first significance probability value of the introduced variable based on the target test rule; In response to the first significance probability value being less than a first probability value, adding the introduced variable to the first linear regression model to obtain the second linear regression model; The relevant data determination module is also specifically used for: Determining a second significance probability value of the induced variable based on the target test rule; In response to the second significance probability value being greater than the second probability value, the derived variable is removed from the second linear regression model to obtain the third linear regression model.
5. The enterprise data management platform according to claim 4, characterized in that: Also includes: Standard adjustment module; The standard adjustment module is configured to reduce the first probability value based on a first adjustment step size in response to the amount of data in the first correlation data set being greater than a first correlation amount; In response to the number of independent variables in the second linear regression model being greater than a second correlation number, the second probability value is increased based on a second adjustment step size.
6. The enterprise data management platform according to claim 1, characterized in that: Also includes: Related condition determination module; The correlation condition determination module is configured to respond that, in response to the number of correlation data of the first abnormal data being greater than a first correlation number, or the importance level of the first abnormal data being greater than a preset level, the first abnormal data satisfies a first correlation condition.
7. The enterprise data management platform according to claim 1, characterized in that: Also includes: Distance adjustment module; The distance adjustment module is configured to reduce the first distance based on a first distance adjustment step in response to the density of the target cluster being greater than a first density; In response to the density of the target cluster being less than a second density, the first distance is increased based on a second distance adjustment step size.
8. The enterprise data management platform according to claim 1, characterized in that: Also includes: A first data processing module; The first data processing module is configured to send abnormal information to the first device in response to the first abnormal data not being real data.
9. The enterprise data management platform according to claim 1, characterized in that: Also includes: A second data processing module; The second data processing module is used for controlling the second device to store the first abnormal data and related data of the first abnormal data in response to the first abnormal data not being real data.
Citation Information
Patent Citations
Abnormal transaction enterprise recognition method and device
CN112711577A
Method and device for determining abnormal threshold value of performance index, equipment and storage medium
CN114358581A