A data risk detection method, device and equipment
By constructing an abnormal risk model, using the K-Means algorithm to cluster and calculate the deviation degree of fund data, the problem of insufficient accuracy of fund data risk estimates is solved, and more efficient and accurate fund abnormal warning is achieved.
Patent Information
- Application Number
- CN202010234432.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-03-30
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2040-03-30
AI Technical Summary
The existing technology has insufficient accuracy in the risk estimate of fund data, which has greatly reduced the value of early warning at the macro level and requires a lot of manpower to investigate the cost.
By constructing a retrieval layer, a median calculation layer, a deviation degree calculation layer and a risk assessment layer, the K-Means algorithm is used to cluster and group the fund data, calculate the median value and deviation degree, determine the safety interval, and form an abnormal risk model to detect the risks of the fund data.
It improves the accuracy of fund data risk estimates, reduces the cost of labor investigation, and improves the rate and accuracy of fund abnormal warnings.
Smart Images

Figure CN111553383B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data processing, and particularly to a data risk detection method, apparatus and device. Background Art
[0002] All along, the cost supervision of various funds has been the focus of key supervision and work by national regulatory departments, the Ministry of Human Resources and Social Security, etc. How to avoid fraud, abuse, illegal waste of various funds in specific institutions and their subordinate departments at all levels is of great concern to government departments. Traditionally, whether it is the provincial regulatory department or the city regulatory department, they only monitor and analyze the doubts about the fund costs at the macro level. When facing suspected abnormal cost problems, they only conduct targeted analysis and tracking for a single object. For example, alarm for the fund expenditure and balance in XX city, and give early warning for the increase in the fund treatment expenditure of XX institution...
[0003] However, in the actual supervision process, the influencing factors affecting abnormal cost fluctuations are various, and the causes of each abnormal problem are also different. Therefore, warning a single object at the macro level often requires a large amount of human investigation costs in the later stage, and the value of the warning itself is correspondingly greatly reduced.
[0004] Therefore, how to improve the accuracy of risk prediction of fund data has become a technical problem to be solved urgently at present. Summary of the Invention
[0005] In view of this, the present application provides a data risk detection method, apparatus and device. The main purpose is to solve the technical problem of how to improve the accuracy of risk prediction of fund data.
[0006] According to the first aspect of the present application, a data risk detection method is provided. The steps of the method include:
[0007] Construct an acquisition layer, and use the acquisition layer to retrieve normal fund data within a predetermined time period from the storage databases of each fund institution, and use the normal fund data as sample data;
[0008] Construct a median calculation layer, and the median calculation layer uses the K-Means algorithm to cluster and group the sample data, and calculates the median M of each group;
[0009] Construct a deviation calculation layer, and calculate the deviation P of each sample data in each group from the median M;
[0010] Construct a risk assessment layer, obtain the maximum value P of the deviation P of the sample data, max , and 0 to P maxThe value range is determined as a safe interval, and the safe interval is stored in the risk assessment layer;
[0011] The median number calculation layer, the deviation degree calculation layer and the risk assessment layer are combined to form an abnormal risk model;
[0012] The obtained fund data to be detected is input into the abnormal risk model;
[0013] The abnormal risk model processes the fund data to be detected. After being processed by the median number calculation layer and the deviation degree calculation layer, the deviation degree to be detected is obtained. The risk assessment layer determines whether the deviation degree to be detected is within the safe interval. If not, an alarm message is formed and displayed.
[0014] According to the second aspect of the present application, a data risk detection device is provided. The device includes:
[0015] A retrieval module, configured to construct a retrieval layer, and use the retrieval layer to retrieve normal fund data within a predetermined time period from the storage databases of each fund institution, and use the normal fund data as sample data;
[0016] A calculation module, configured to construct a median number calculation layer, and the median number calculation layer uses the K-Means algorithm to cluster and group the sample data, and calculate the median number M of each group;
[0017] The calculation module is further configured to construct a deviation degree calculation layer, and calculate the deviation degree P of each sample data in each group from the median number M;
[0018] A risk assessment module, configured to obtain the maximum value P of the deviation degree P of the sample data max , and determine the value range of 0 to P max as a safe interval, and store the safe interval in the risk assessment layer;
[0019] A combination module, configured to combine the median number calculation layer, the deviation degree calculation layer and the risk assessment layer to form an abnormal risk model;
[0020] An input module, configured to input the obtained fund data to be detected into the abnormal risk model;
[0021] A processing module, configured to use the abnormal risk model to process the fund data to be detected, and obtain the deviation degree to be detected after passing through the median number calculation layer and the deviation degree calculation layer. The risk assessment layer determines whether the deviation degree to be detected is within the safe interval. If not, an alarm message is formed and displayed.
[0022] According to a third aspect of the present application, there is provided a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the data risk detection method described in the first aspect are implemented.
[0023] According to a fourth aspect of the present application, there is provided a computer storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the data risk detection method described in the first aspect are implemented.
[0024] By means of the above technical solutions, a data risk detection method, device and equipment provided by the present application can use the K-Means algorithm to cluster and group the sample data of each fund institution retrieved, calculate the median of each group, and calculate the deviation degree of each sample data according to the median. Then, the dangerous interval value is determined according to the deviation degree, and the dangerous interval value is saved to obtain an abnormal risk model. In this way, when a user wants to monitor the risk of a certain fund data, the fund data is directly input into the abnormal risk model to determine whether the fund data belongs to dangerous data. If so, the user is reminded to process it in time. If not, it proves that the fund data is normal and no further processing is required. Through the above solution, it is possible to accurately judge whether there is a risk in the fund data, and then when it is determined that there is a risk in the fund data, the user can be reminded to give an early warning of the fund data in time, improving the speed and accuracy of the fund abnormal early warning.
[0025] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically given below. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0027] Figure 1 It is a flowchart of an embodiment of the data risk detection method of the present application;
[0028] Figure 2 It is a structural block diagram of an embodiment of the data risk detection device of the present application;
[0029] Figure 3 It is a schematic structural diagram of the computer device of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0031] The embodiments of the present application provide a data risk detection method, which can embed the K-Means algorithm into the abnormal risk model. The result of clustering and grouping using the K-Means algorithm is more accurate, thereby effectively improving the accuracy of the abnormal risk model in identifying and judging fund data. In this way, users can timely perform early warning processing on the dangerous fund data determined by the abnormal risk model, improving the efficiency of fund abnormal early warning.
[0032] As Figure 1 shown, the embodiments of the present application provide a data risk detection method, including the following steps:
[0033] Step 101, construct a retrieval layer, and use the retrieval layer to retrieve the normal fund data within a predetermined time period from the storage databases of each fund institution, and use the normal fund data as sample data.
[0034] In this step, each fund institution stores information such as the relevant fees of various funds and the corresponding fund purchase time in the storage database set within the institution. For example, for medical funds, information such as the medical treatment time, medical expenses, and types of medical treatment when users use medical funds is stored in the storage database of the medical institution.
[0035] The constructed retrieval layer is provided with the network addresses of the storage databases of each fund institution, and directly retrieves the normal fund data within a predetermined time period (for example, within one month, within one year) from each storage database according to the network address, and uses the retrieved normal fund data as sample data. Among them, all the retrieved data are normal fund data to ensure that the safe interval can be accurately determined after processing these normal fund data below.
[0036] Step 102, construct a median number calculation layer, and the median number calculation layer uses the K-Means algorithm to cluster and group the sample data, and calculates the median number M of each group.
[0037] In this step, the retrieval layer sends the retrieved sample data to the median number calculation layer, uses the K-Means algorithm to divide the sample data into several groups, and then draws a circle with the distance between the two farthest data in each group as the diameter. The center of the circle is the median number M. If R groups are divided, there are R corresponding median numbers.
[0038] The calculated median is the normal value representing the set of data.
[0039] Step 103: Construct a deviation degree calculation layer to calculate the deviation degree P of each sample data in each group from the median M.
[0040] In this step, the deviation degree is the deviation distance between each fund data and the median, which is used to judge the deviation of each sample data from the normal value. The larger the deviation degree, the greater the deviation of the sample data from the normal value, and the more dangerous it is.
[0041] Step 104: Construct a risk assessment layer to obtain the maximum value P of the deviation degree P of the sample data. max and determine the value range from 0 to P max as the safe interval and save the safe interval in the risk assessment layer.
[0042] In this step, since all the sample data are normal fund data, the maximum value P of the obtained deviation degree max is determined as the maximum value of the safe interval, so that the data corresponding to the deviation degree exceeding this safe interval can be determined as dangerous data.
[0043] Then the obtained safe interval is saved in the risk assessment layer for subsequent assessment of the risk level of the deviation degree.
[0044] In addition, if there is an error in this safe interval, the user can adjust it according to their actual experience.
[0045] Step 105: Combine the median calculation layer, the deviation degree calculation layer and the risk assessment layer to form an abnormal risk model.
[0046] In this step, since there is no need to retrieve data again when detecting fund data, the retrieved layer obtained above needs to be sealed, and the remaining layers are combined. That is, the output port of the median calculation layer is connected to the input port of the deviation degree calculation layer, and the output port of the deviation degree calculation layer is connected to the input port of the risk assessment layer, so as to obtain an abnormal risk model.
[0047] Step 106: Input the obtained fund data to be detected into the abnormal risk model.
[0048] In this step, the fund data to be detected can be one or more. If there are multiple, they are input into the abnormal risk model in sequence.
[0049] Step 107: The abnormal risk model processes the fund data to be detected. After being processed by the median calculation layer and the deviation degree calculation layer, the deviation degree to be detected is obtained. The risk assessment layer judges whether the deviation degree to be detected is within the safe interval. If not, an alarm message is formed and displayed.
[0050] In this step, if the detection result of the obtained abnormal risk model for fund data is not accurate enough, the retrieval layer is re-embedded into the abnormal risk model, and the output port of the retrieval layer is connected to the output port of the median number calculation layer. Then, the required fund data is retrieved and the abnormal risk model is adjusted according to the above steps, so that the detection accuracy of the abnormal risk model is effectively improved.
[0051] Through the above solution, the K-Means algorithm can be used to cluster and group the sample data of each retrieved fund institution, calculate the median number of each group, calculate the deviation degree of each sample data according to the median number, then determine the dangerous interval value according to the deviation degree, save the dangerous interval value, and thus obtain the abnormal risk model. In this way, when a user wants to monitor the risk of a certain fund data, the fund data is directly input into the abnormal risk model to determine whether the fund data belongs to dangerous data. If it does, the user is timely reminded to process it. If not, it proves that the fund data is normal and no further processing is required. In this way, it can accurately judge whether there is a risk in the fund data, and then when it is determined that there is a risk in the fund data, the user can be timely reminded to give an early warning for the fund data, improving the speed and accuracy of fund abnormal early warning.
[0052] In a specific embodiment, step 102 specifically includes:
[0053] Step 1021, arbitrarily select K sample data from multiple sample data as the initial clustering centers.
[0054] In this step, the value of K can be set according to actual needs.
[0055] Step 1022, calculate the distances between the remaining sample data and each initial clustering center.
[0056] Step 1023, allocate the remaining sample data to the group corresponding to the initial clustering center with the closest distance.
[0057] In this step, after the allocation, K groups are obtained. Then, the groups with only initial clustering centers are extracted, the distances between the sample data of the initial clustering center and other initial clustering centers are calculated, and the sample data of the initial clustering center is allocated to the group of the other initial clustering center with the closest distance. This avoids the influence of the grouped null values on the calculation result of the safe interval.
[0058] Step 1024, calculate the center points of each group of sample data, use the sample data closest to the center point as the clustering center, calculate the distances between the remaining sample data and each clustering center, and allocate the remaining sample data to the group corresponding to the clustering center with the closest distance.
[0059] In this step, a circle is drawn with the distance between the two most distant data in each group as the diameter, and the center of this circle is the center point. The sample data closest to the center point is used as the clustering center.
[0060] Step 1025: Determine whether the clustering center has changed. If it has changed, recalculate the new center point of the sample data after clustering for each group, use the sample data closest to the new center point as the new clustering center, calculate the distances between the remaining sample data and each new clustering center, and assign the remaining sample data to the group corresponding to the new clustering center with the closest distance until the new clustering center no longer changes. If it remains unchanged, calculate the median number M of the sample data after clustering for each group.
[0061] In this step, through the above solution, the sample data can be clustered repeatedly, so that the obtained clustering center is more accurate, and the similarity of each group of sample data obtained is higher.
[0062] When calculating the median number M, a circle is drawn with the distance between the two most distant data in each group as the diameter, and the center of this circle is the median number M.
[0063] The above step process is the implementation process of the K-Means algorithm. Embedding the K-Means algorithm into the abnormal risk model, the result of clustering and grouping using the K-Means algorithm is more accurate, thereby effectively improving the accuracy of the abnormal risk model in identifying and judging fund data.
[0064] In a specific embodiment, step 1022 specifically includes:
[0065] Step 10221: Establish a coordinate system with the time data in the sample data as the horizontal axis x and the numerical data as the vertical axis y.
[0066] Step 10222: Mark each sample data in the coordinate system.
[0067] Step 10223: Calculate the Euclidean distance between each sample data and each initial clustering center:
[0068] where i is the sequence value of the sample data and n is the number of sample data.
[0069] Through the above solution, all sample data is presented in a two-dimensional plane, so that the process of clustering and grouping is more convenient and fast, effectively improving the clustering process of the K-Means algorithm and ensuring the accuracy of clustering and grouping.
[0070] In a specific embodiment, step 103 specifically includes:
[0071] Step 1031: Calculate the difference Δ between each sample data F and the median M in each group.
[0072] Step 1032: Calculate the deviation value T of each sample data F, where T = Δ / M.
[0073] Step 1033: Use the normalization algorithm to convert the obtained deviation value T into a deviation degree P within a predetermined value range.
[0074] Among them, for the difference F - M, take the absolute value of this difference to obtain Δ.
[0075] Through the above solution, normalizing the deviation degree to a predetermined value range can avoid the problem that it is difficult to compare and calculate the deviation degree due to large differences in the magnitude of the deviation degree P, improve the evaluation efficiency of fund data, and ensure that users can timely give early warnings for dangerous fund data.
[0076] In a specific embodiment, step 104 specifically includes:
[0077] Step 1041: Construct a risk assessment layer and obtain the maximum value P of the deviation degree P of the sample data. max 。
[0078] Step 1042: Divide the low-risk interval [0, P max / 2], the medium-risk interval (P max / 2, P max , and the high-risk interval (P max , ∞), where the low-risk interval and the medium-risk interval are both safe intervals, and the high-risk interval is a dangerous interval.
[0079] Step 1043: Save the low-risk interval, the medium-risk interval, and the high-risk interval to the risk assessment layer.
[0080] Through the above solution, when the fund data to be detected is input into the abnormal risk model, after the evaluation of the risk assessment layer, the corresponding risk level of the fund data to be detected can be determined. For the fund data within the low-risk interval, no processing is required; for the fund data within the medium-risk interval, obtain the user contact information of this fund data and give an early warning prompt to the user through the contact information; for the fund data within the high-risk interval, the user corresponding to this fund data has the behavior of maliciously using medical funds, obtain the personal information of the user, and freeze the user, prohibiting the user from using medical funds again until the system unfreezes the user.
[0081] In a specific embodiment, there are N retrieval paths for sample data. The retrieval layer retrieves N groups of sample data according to the N paths. The N groups of sample data respectively correspond to N abnormal risk models. Mark the corresponding paths for the input ports of each abnormal risk model, and integrate the input ports of the N abnormal risk models.
[0082] Then step 106 is specifically as follows: Extract the retrieval path of the fund data to be detected, select the matching input port according to the retrieval path, and input the fund data to be detected from the matching input port into the matching abnormal risk model.
[0083] In the above solution, different paths correspond to different sample data. After the sample data of one path goes through the above steps 101 - 106, an abnormal warning model capable of identifying the risk level of the fund data of this path is obtained. Then, if there are N paths, N abnormal warning models are obtained.
[0084] Set the corresponding path identifier for each abnormal warning model to distinguish different abnormal warning models. Sort the N abnormal warning models with path identifiers according to the first letter of the path identifier, and integrate the N abnormal warning models in this order together as the total warning model.
[0085] In this way, add a path recognition function at the entrance of the total warning model. Directly according to the path of the fund data to be detected obtained, retrieve the corresponding abnormal warning model from the total warning model, and input the fund data to be detected into this abnormal warning model for processing.
[0086] Through the above solution, different paths correspond to different abnormal warning models, which can improve the accuracy and efficiency of risk detection of the abnormal warning model.
[0087] In a specific embodiment, step 107 specifically includes:
[0088] Step 1071, input the fund data to be detected into the median number calculation layer, compare the fund data to be detected with the value ranges of each group, and obtain the median number M of the group where the fund data to be detected is located. 待检测 , send M 待检测 to the deviation degree calculation layer.
[0089] Step 1072, the deviation degree calculation layer calculates the deviation degree P of the fund data to be detected from M 待检测 , and send P 待检测 to the risk assessment layer. 待检测 to the risk assessment layer.
[0090] Step 1073, the risk assessment layer sends P 待检测Compare with the danger range. When it is determined that the fund data to be detected is not within the safe range and belongs to dangerous data, an alarm message is formed and displayed.
[0091] Through the above solution, when a user wants to monitor the risk of a certain fund data, directly input the fund data into the abnormal risk model to determine whether the fund data belongs to dangerous data. If it does, remind the user to handle it in time. If not, it proves that the fund data is normal and no further processing is required.
[0092] In another embodiment of the data risk detection method of the present application, which is applicable to medical funds, the following steps are included:
[0093] I. Build an abnormal warning model
[0094] 1. The first layer, the retrieval layer, retrieves the medical insurance expenses within a predetermined time period (for example, 1 month, 6 months or 1 year, specifically selected according to the actual situation of the user) in the storage systems of each institution, and the total number of medical visits of each institution.
[0095] Set the corresponding data selection paths (this path is different from the first step of the first part, and each path corresponds to a group of samples): For example: Path A: Disease - Institution - Department - Inspection and Examination, Path B: Disease - Physician - Medication, Path C: Disease - Institution - Department - Physician - Diagnosis and Treatment.
[0096] Select samples according to different paths (Note: The user needs to detect which one or several paths, then select several samples.). The samples include medical expenses and the number of medical visits, and the user artificially pre - marks the corresponding risk levels for each sample data according to the actual situation of each sample data.
[0097] Among them, each path corresponds to a group of sample data.
[0098] 2. The second layer, the median calculation layer, adds the K - Means algorithm for clustering and grouping, specifically as follows:
[0099] 1) Arbitrarily select k objects from the sample data objects as the initial clustering centers.
[0100] 2) Calculate the Euclidean distance between the remaining sample data and each initial clustering center.
[0101] Among them, take the acquisition time of each object in the sample data as the horizontal axis x and the specific value as the vertical axis y to establish a coordinate system.
[0102] Assign the remaining sample data to the group of the nearest clustering center to obtain K clustering groups.
[0103] 3) Recalculate the cluster center of each group, specifically: calculate the center point of the coordinate range where each group is located in the coordinates, and use the object closest to the center point as the new cluster center of the group.
[0104] 4) Loop through steps 2) to 3) until each cluster group no longer changes.
[0105] Then, arrange each cluster group obtained after clustering in ascending order and calculate the median number of each cluster group.
[0106] For example, the corresponding groups are Group A, Group B, Group C, and Group D, and the obtained median numbers are: MedianA, MedianB, MedianC, MedianD.
[0107] 3. Third layer, deviation degree calculation layer, calculate the difference between each sample data F in each group and the corresponding median number: △ A (x) = F A - MedianA, △ B (x) = F B - MedianB, △ C (x) = F C - MedianC, △ D (x) = F D - MedianD.
[0108] Calculate the deviation degree of the secondary average cost of each group
[0109] T A (x) = △ A (x) / MedianA
[0110] T B (x) = △ B (x) / MedianB
[0111] T C (x) = △ C (x) / MedianC
[0112] T D (x) = △ D (x) / MedianD.
[0113] And use the normalization algorithm to convert the deviation degree T into a value P between 0 and 100.
[0114] 4. Fourth layer, risk assessment layer, the user sets the interval of the risk value according to their own needs. For example, 0 - 50 is low risk, 50 - 80 is medium risk, and 80 - 100 is high risk.
[0115] Compare the value P calculated in the fourth layer with the set range of dangerous values, and output the corresponding degree of danger.
[0116] Compare with the pre-marked degree of danger. If they are different, readjust the range of dangerous values of the abnormal warning model, and then complete the correction of the abnormal warning model.
[0117] Different paths correspond to different sample data. After the sample data of one path goes through the above steps 1-4, an abnormal warning model capable of identifying the risk degree of the medical data of this path is obtained. Then, if there are N paths, N abnormal warning models are obtained.
[0118] Set a corresponding path identifier for each abnormal warning model to distinguish different abnormal warning models;
[0119] Sort the N abnormal warning models with path identifiers according to the first letter of the path identifier, and integrate the N abnormal warning models in this order to form a total warning model.
[0120] In this way, the corresponding abnormal warning can be selected from the total warning model according to the path of the medical data for risk detection.
[0121] II. Detection using the abnormal warning model
[0122] 1. When performing detection, seal the first layer in the integrated abnormal warning model above.
[0123] 2. Add a path recognition function at the entrance of the total warning model. Directly according to the path of the medical data X to be detected (i.e., medical expenses) obtained, retrieve the corresponding abnormal warning model from the total warning model according to the path identifier, and input the medical data X to be detected into this abnormal warning model.
[0124] 3. The second layer of this abnormal warning model compares the medical data X to be detected with the ranges of each group of data divided in this abnormal warning model, determines the group corresponding to the medical data X to be detected, and obtains the median number M of this group.
[0125] 4. The second layer sends the medical data X to be detected and the corresponding median number M to the third layer. The third layer calculates the deviation degree T between the data to be detected and the median number. T = (X - M) / M, and uses the normalization algorithm to convert the deviation degree T into a value between 0 and 100.
[0126] 5. After the fourth layer receives the converted value, compare it with the corresponding range of dangerous values to determine the degree of danger of the data to be detected.
[0127] If the risk level is high, the institutional storage system corresponding to the data to be detected is retrieved, and relevant information (such as the source of the data to be detected, personal information of relevant persons, etc.) is retrieved from the storage system and displayed for the user to process in a timely manner.
[0128] Further, as Figure 1 a specific implementation of the method, an embodiment of the present application provides a data risk detection device, as Figure 2 shown, the device includes: a retrieval module 21, a calculation module 22, a risk assessment module 23, a combination module 24, an input module 25, and a processing module 26 that are connected in sequence.
[0129] The retrieval module 21 is used to construct a retrieval layer, and use the retrieval layer to retrieve normal fund data within a predetermined time period from the storage databases of each fund institution, and use the normal fund data as sample data;
[0130] The calculation module 22 is used to construct a median number calculation layer, and the median number calculation layer uses the K-Means algorithm to cluster and group the sample data, and calculate the median number M of each group;
[0131] The calculation module 22 is also used to construct a deviation degree calculation layer, and calculate the deviation degree P of each sample data in each group from the median number M;
[0132] The risk assessment module 23 is used to obtain the maximum value P of the deviation degree P of the sample data, max and determine the value range of 0 to P max as the safe interval, and save the safe interval in the risk assessment layer;
[0133] The combination module 24 is used to combine the median number calculation layer, the deviation degree calculation layer, and the risk assessment layer to form an abnormal risk model;
[0134] The input module 25 is used to input the obtained fund data to be detected into the abnormal risk model;
[0135] The processing module 26 is used to process the fund data to be detected by using the abnormal risk model, and obtain the deviation degree to be detected after passing through the median number calculation layer and the deviation degree calculation layer. The risk assessment layer judges whether the deviation degree to be detected is within the safe interval. If not, an alarm message is formed and displayed.
[0136] In a specific embodiment, the calculation module 22 includes:
[0137] A selection unit for arbitrarily selecting K sample data from multiple sample data as the initial clustering centers;
[0138] A distance calculation unit for calculating the distances between the remaining sample data and each initial clustering center;
[0139] A clustering and grouping unit, configured to allocate the remaining sample data to the group corresponding to the nearest initial clustering center;
[0140] The clustering and grouping unit is further configured to calculate the center point of each group of sample data, use the sample data closest to the center point as the clustering center, calculate the distances between the remaining sample data and each clustering center, and allocate the remaining sample data to the group corresponding to the nearest clustering center;
[0141] A judgment unit, configured to judge whether the clustering center has changed. If it has changed, recalculate the new center point of each group of clustered sample data, use the sample data closest to the new center point as the new clustering center, calculate the distances between the remaining sample data and each new clustering center, and allocate the remaining sample data to the group corresponding to the nearest new clustering center until the obtained new clustering center no longer changes. If it remains unchanged, calculate the median number M of each group of clustered sample data.
[0142] In a specific embodiment, the distance calculation unit specifically includes:
[0143] A coordinate system establishment unit, configured to establish a coordinate system with the time data in the sample data as the horizontal axis x and the numerical data as the vertical axis y;
[0144] A marking unit, configured to mark each sample data in the coordinate system;
[0145] A Euclidean distance calculation unit, configured to calculate the Euclidean distances between each sample data and each initial clustering center:
[0146] where i is the sequence value of the sample data and n is the number of sample data.
[0147] In a specific embodiment, the calculation module 22 further includes:
[0148] A difference calculation unit, configured to calculate the difference △ between each sample data F in each group and the median number M;
[0149] A deviation value calculation unit, configured to calculate the deviation value T of each sample data F as T = △ / M;
[0150] A normalization unit, configured to convert the obtained deviation value T into a deviation degree P within a predetermined value range by using a normalization algorithm.
[0151] In a specific embodiment, the risk assessment module 23 specifically includes:
[0152] An acquisition unit, configured to construct a risk assessment layer and acquire the maximum value P of the deviation degree P of the sample data max ;
[0153] Partitioning unit, used to partition the low-risk interval [0, P max / 2], medium-risk interval (P max / 2, P max , and high-risk interval (P max , ∞), where the low-risk interval and the medium-risk interval are both safe intervals, and the high-risk interval is a dangerous interval;
[0154] Storage unit, used to store the low-risk interval, medium-risk interval and high-risk interval to the risk assessment layer.
[0155] In a specific embodiment, there are N paths for retrieving sample data. The retrieval layer retrieves N groups of sample data according to N paths. The N groups of sample data correspond to N abnormal risk models. Mark the corresponding paths for the input ports of each abnormal risk model, and integrate the input ports of the N abnormal risk models;
[0156] Then the input module 25 is specifically used for: extracting the retrieval path of the fund data to be detected, selecting a matching input port according to the retrieval path, and inputting the fund data to be detected from the matching input port into the matching abnormal risk model.
[0157] In a specific embodiment, the processing module 26 specifically includes:
[0158] Comparison unit, used to input the fund data to be detected into the median calculation layer, compare the fund data to be detected with the value ranges of each group, and obtain the median M of the group where the fund data to be detected is located 待检测 , and send M 待检测 to the deviation calculation layer;
[0159] Deviation calculation unit, used to calculate the deviation P of the fund data to be detected from M 待检测 using the deviation calculation layer, and send P 待检测 to the risk assessment layer; 待检测 Comparison unit, is also used to compare P
[0160] with the dangerous interval using the risk assessment layer. When it is determined that the fund data to be detected is not within the safe interval and belongs to dangerous data, an alarm message is formed and displayed. 待检测 Based on the above
[0161] method and Figure 1 device embodiments shown, to achieve the above purpose, an embodiment of the present application also provides a computer device, as Figure 2 shown, including a memory 32 and a processor 31, where both the memory 32 and the processor 31 are arranged on a bus 33. The memory 32 stores a computer program, and when the processor 31 executes the computer program, it implements Figure 3 shown, including a memory 32 and a processor 31, where both the memory 32 and the processor 31 are arranged on a bus 33. The memory 32 stores a computer program, and when the processor 31 executes the computer program, it implementsFigure 1 The data risk detection method shown
[0162] Based on such understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile memory (such as a CD-ROM, USB flash drive, mobile hard disk, etc.), and includes several instructions to enable a computer device (such as a personal computer, server, or network device, etc.) to execute the methods described in various implementation scenarios of this application.
[0163] Optionally, the device can also be connected to a user interface, a network interface, a camera, a radio frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, etc. The user interface can include a display screen, an input unit such as a keyboard, etc., and the optional user interface can also include a USB interface, a card reader interface, etc. The network interface can optionally include a standard wired interface, a wireless interface (such as a Bluetooth interface, a WI-FI interface), etc.
[0164] Those skilled in the art can understand that the structure of a computer device provided in this embodiment does not limit the entity device, and it can include more or fewer components, or combine certain components, or have different component arrangements.
[0165] Based on the above as Figure 1 shown in the method and Figure 2 shown in the device embodiment, correspondingly, this application embodiment also provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the data risk detection method as Figure 1 shown above.
[0166] The storage medium can also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of a computer device and supports the operation of information processing programs and other software and / or programs. The network communication module is used to implement communication between components inside the storage medium and communication between other hardware and software in the computer device.
[0167] Through the description of the above implementation manners, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary general hardware platform, or can also be implemented by hardware.
[0168] By applying the technical solution of the present application, it is possible to use the K-Means algorithm to cluster and group the sample data of each retrieved fund institution, calculate the median of each group, calculate the deviation degree of each sample data based on the median, then determine the dangerous interval value according to the deviation degree, save the dangerous interval value to obtain an abnormal risk model. In this way, when a user wants to monitor the risk of a certain fund data, directly input the fund data into the abnormal risk model to determine whether the fund data belongs to dangerous data. If it does, the user is timely reminded to process it. If not, it proves that the fund data is normal and no further processing is required. In this way, it is possible to accurately judge whether there is a risk in the fund data, and then when it is determined that there is a risk in the fund data, the user can be timely reminded to perform early warning processing on the fund data, improving the speed and accuracy of fund abnormal early warning.
[0169] Those skilled in the art can understand that the attached drawings are only schematic diagrams of a preferred implementation scenario, and the modules or processes in the attached drawings are not necessarily essential for implementing the present application. Those skilled in the art can understand that the modules in the device in the implementation scenario can be distributed in the device of the implementation scenario according to the description of the implementation scenario, or can be correspondingly changed and located in one or more devices different from the present implementation scenario. The modules in the above implementation scenario can be combined into one module, or further split into multiple sub-modules.
[0170] The above serial numbers of the present application are only for description and do not represent the advantages or disadvantages of the implementation scenario. The above disclosure is only several specific implementation scenarios of the present application. However, the present application is not limited thereto, and any changes that can be thought of by those skilled in the art should fall within the protection scope of the present application.
Claims
1. A method for detecting data risks, characterized in that, The steps of the method include: Construct an extraction layer, and use the extraction layer to extract normal fund data within a predetermined time period from the storage databases of each fund institution, and use the normal fund data as sample data; Construct a median calculation layer, and the median calculation layer uses the K-Means algorithm to cluster and group the sample data, and calculates the median M of each group; Construct a deviation calculation layer, and calculate the deviation P of each sample data in each group from the median M; Construct a risk assessment layer to obtain the maximum deviation degree P of the sample data, denoted as P max , and determine the value range from 0 to P max as the safe interval, and save the safe interval in the risk assessment layer; Combine the median calculation layer, the deviation calculation layer and the risk assessment layer to form an abnormal risk model; Input the obtained fund data to be detected into the abnormal risk model; The abnormal risk model processes the fund data to be detected, and after being processed by the median calculation layer and the deviation calculation layer, obtains the deviation to be detected. The risk assessment layer determines whether the deviation to be detected is within the safe range. If not, an alarm message is formed and displayed; The abnormal risk model processes the fund data to be detected, and after being processed by the median calculation layer and the deviation calculation layer, obtains the deviation to be detected. The risk assessment layer determines whether the deviation to be detected is within the safe range. If not, an alarm message is formed and displayed, specifically including: Input the fund data to be detected into the median number calculation layer, compare the fund data to be detected with the value ranges of each group, and obtain the median number M of the group where the fund data to be detected is located 待检测 , and send M 待检测 to the deviation degree calculation layer; The deviation degree calculation layer calculates the deviation degree P of the fund data to be detected and M 待检测 and sends P 待检测 to the risk assessment layer; 待检测 The risk assessment layer compares P 待检测 with the danger range. When it is determined that the fund data to be detected is not within the safe range and belongs to dangerous data, an alarm message is formed and displayed. Among them, the value range of the danger range is P max ~∞.
2. The method according to claim 1, characterized in that, The construction of the median calculation layer, and the median calculation layer uses the K-Means algorithm to cluster and group the sample data, and calculates the median M of each group, specifically including: Arbitrarily select K sample data from the multiple sample data as the initial clustering centers; Calculate the distances between the remaining sample data and each initial clustering center; Assign the remaining sample data to the group corresponding to the initial clustering center with the closest distance; Calculate the center points of each group of sample data, use the sample data closest to the center point as the clustering center, calculate the distances between the remaining sample data and each clustering center, and assign the remaining sample data to the group corresponding to the clustering center with the closest distance; Determine whether the clustering center has changed. If it has changed, recalculate the new center points of each group of clustered sample data, use the sample data closest to the new center point as the new clustering center, calculate the distances between the remaining sample data and each new clustering center, and assign the remaining sample data to the group corresponding to the new clustering center with the closest distance until the obtained new clustering center no longer changes. If it does not change, calculate the median M of each group of clustered sample data.
3. The method according to claim 2, characterized in that, The calculation of the distances between the sample data and each initial clustering center specifically includes: Establish a coordinate system with the time data in the sample data as the horizontal axis x and the numerical data as the vertical axis y; Mark each sample data in the coordinate system; Calculate the Euclidean distances between each sample data and each initial clustering center: Wherein, i is the sequential value of the sample data, and n is the number of the sample data.
4. The method according to claim 1, characterized in that, The construction of the deviation calculation layer, and the calculation of the deviation P of each sample data in each group from the median M, specifically includes: Calculate the difference △ between each sample data F in each group and the median M; Calculate the deviation value T of each sample data F = △ / M; The obtained deviation value T is converted into a deviation degree P within a predetermined value range by using a normalization algorithm.
5. The method according to claim 1, characterized in that, Construct a risk assessment layer to obtain the maximum deviation P of the sample data, P max , and determine the value range from 0 to P max as the safe interval, and save the safe interval in the risk assessment layer, specifically including: Construct a risk assessment layer to obtain the maximum value P of the deviation degree P of the sample data max ; Divide the low-risk interval [0, P max / 2], the medium-risk interval (P max / 2, P max , and the high-risk interval (P max , ∞), where the low-risk interval and the medium-risk interval are both safe intervals, and the high-risk interval is a dangerous interval; The low-risk interval, the medium-risk interval, and the high-risk interval are stored in the risk assessment layer.
6. The method according to claim 1, characterized in that, There are N retrieval paths for the sample data. The retrieval layer retrieves N groups of sample data according to the N paths. N abnormal risk models are correspondingly obtained from the N groups of sample data. Each input port of the abnormal risk models is marked with the corresponding path, and the input ports of the N abnormal risk models are integrated. Then, inputting the obtained fund data to be detected into the abnormal risk model specifically includes: Extracting the retrieval path of the fund data to be detected, selecting a matching input port according to the retrieval path, and inputting the fund data to be detected into the matching abnormal risk model from the matching input port.
7. A data risk detection device, characterized in that The device includes: A retrieval module, configured to construct a retrieval layer, and use the retrieval layer to retrieve normal fund data within a predetermined time period from the storage databases of each fund institution, and use the normal fund data as sample data. A calculation module, configured to construct a median number calculation layer, and the median number calculation layer uses the K-Means algorithm to perform clustering grouping on the sample data and calculate the median number M of each group. The calculation module is further configured to construct a deviation degree calculation layer to calculate the deviation degree P of each sample data in each group from the median number M. A risk assessment module for obtaining the maximum deviation P of the sample data, denoted as P max , and determining the value range from 0 to P max as the safety interval, and storing the safety interval in the risk assessment layer; A combination module, configured to combine the median number calculation layer, the deviation degree calculation layer, and the risk assessment layer to form an abnormal risk model. An input module, configured to input the obtained fund data to be detected into the abnormal risk model. A processing module, configured to process the fund data to be detected by using the abnormal risk model, obtain the deviation degree to be detected after passing through the median number calculation layer and the deviation degree calculation layer, and the risk assessment layer determines whether the deviation degree to be detected is within the safe interval. If not, an alarm message is formed and displayed. The processing module specifically includes: A comparison unit is configured to input the fund data to be detected into the median number calculation layer, compare the fund data to be detected with the value ranges of each group, and obtain the median number M of the group where the fund data to be detected is located 待检测 , and send M 待检测 to the deviation degree calculation layer; A deviation degree calculation unit is used to calculate the deviation degree P of the fund data to be detected and M by using a deviation degree calculation layer 待检测 and send P 待检测 to the risk assessment layer; 待检测 The comparison unit is further configured to use the risk assessment layer to compare P 待检测 with the danger range. When it is determined that the fund data to be detected is not within the safe range and belongs to dangerous data, an alarm message is formed and displayed, where the value range of the danger range is P max to ∞.
8. A computer device, including a memory and a processor, the memory stores a computer program, characterized in that When the processor executes the computer program, the steps of the data risk detection method according to any one of claims 1 to 6 are implemented.
9. A computer storage medium, on which a computer program is stored, characterized in that When the computer program is executed by the processor, the steps of the data risk detection method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Industrial experiment data abnormal point detecting method and device
CN108829878A