Target data grade classification method and device
By calculating the differential volatility value and determining the segmentation point, and dividing the data levels in combination with the data quantity ratio, the problem of inaccurate big data classification results in the existing technology is solved, and higher classification accuracy is achieved.
Patent Information
- Application Number
- CN202510137747.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, the classification methods, bucketing method and hierarchical method of big data have problems of inaccurate classification results, especially the difference between critical data in adjacent level areas is not obvious.
A method of dividing data levels by calculating the differential volatility value and determining the segmentation point based on its size, and combining the data quantity ratio. Specific steps include: sorting index data, calculating differential volatility sequence, determining segmentation points, dividing data segments, and calculating data quantity ratios to confirm data level.
It realizes more accurately identifying critical data in adjacent rank areas, avoiding the problem of insignificant differences between critical data, thereby improving the accuracy of classification results.
Smart Images

Figure CN120067753A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a method and device for classifying the levels of target data. Background Art
[0002] With the advent of the big data era, big data contains rich information. Classifying big data into different levels can better mine the information contained in the data and improve the utilization efficiency of the data. Therefore, it is very necessary to classify big data into different levels. In the related art, there are mainly two methods for classifying big data into different levels: one is the bucketing method, and the other is the layering method.
[0003] The bucketing method divides big data into several buckets, and the data within each bucket has similar characteristics, while the data between different buckets has differences. The layering method divides big data into several layers, and the data within each layer has similar characteristics, while the data between different layers has differences. When classifying big data into different levels, both the bucketing method and the layering method are based on the characteristics of the data for division, and can achieve the classification of big data.
[0004] However, when using the bucketing method and the layering method in the related art to classify big data, in the obtained classification results, there is a problem that the difference between the critical data in adjacent level regions is not obvious, resulting in inaccurate classification results. Summary of the Invention
[0005] Embodiments of the present invention provide a method and device for classifying the levels of target data, so as to at least solve the problem of inaccurate classification results in the related art.
[0006] According to an embodiment of the present invention, a method for classifying the levels of target data is provided, including: arranging the index data to be classified in ascending order of data values; calculating the difference volatility values of two adjacent ones of the sorted index data to obtain a difference volatility sequence; arranging the difference volatility sequence in descending order of difference volatility values; using two index data with difference volatility values greater than a first preset threshold as a first segmentation point; arranging the index data in descending order of data values, and dividing the index data into multiple data segments based on the first segmentation point; in the direction of decreasing data values, calculating a first ratio of the data volume of the first data segment to the total data volume of the index data; in the case where the first ratio is less than or equal to a first proportional value, confirming that the first data segment is a first level, and calculating a second ratio of the sum of the data volume of the second data segment and the data volume of the first data segment to the total data volume of the index data; in the case where the second ratio is less than or equal to a second proportional value, confirming that the second data segment is a second level.
[0007] According to another embodiment of the present invention, a hierarchical classification system for target data is provided, including: a first sorting module for sorting the index data to be classified in ascending order of data values; a first calculation module for calculating the differential volatility values of two adjacent ones of the sorted index data to obtain a differential volatility sequence; a second sorting module for sorting the differential volatility sequence in descending order of differential volatility values; a segmentation point determination module for using two of the index data with differential volatility values greater than a first preset threshold as a first segmentation point; a third sorting module for sorting the index data in descending order of data values and dividing the index data into multiple data segments based on the first segmentation point; a second calculation module for calculating a first ratio of the data volume of the first data segment to the total data volume of the index data in the direction of decreasing data values; a third calculation module for, when the first ratio is less than or equal to a first proportional value, confirming that the first data segment is a first level and calculating a second ratio of the sum of the data volume of the second data segment and the data volume of the first data segment to the total data volume of the index data; and a processing module for, when the second ratio is less than or equal to a second proportional value, confirming that the second data segment is a second level.
[0008] According to still another embodiment of the present invention, a computer-readable storage medium is further provided. A computer program is stored in the computer-readable storage medium, wherein the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0009] According to still another embodiment of the present invention, an electronic device is further provided, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0010] According to still another embodiment of the present invention, a computer program product is further provided, including computer instructions, and the computer instructions, when executed by a processor, implement the steps in any one of the above method embodiments.
[0011] Through one of the embodiments of the present invention, since the method of calculating the differential volatility value, determining the segmentation point according to its size, and then combining the data volume ratio to divide the data level is adopted, compared with the bucketing method and the layering method in the related art, it can more accurately identify the critical data in the adjacent level regions, and based on the first preset threshold, two adjacent index data (critical data) to be classified with larger differences can be determined, so as to avoid the problem that the difference between the critical data is not obvious. Therefore, the problem of inaccurate classification results in the related art can be solved, and the effect of improving the accuracy of the classification results can be achieved. Brief Description of the Drawings
[0012] The drawings described herein are used to provide a further understanding of the present invention and form a part of the present invention. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0013] Figure 1 is a flowchart of a method for classifying the levels of target data according to an embodiment of the present invention;
[0014] Figure 2 is a flowchart of a method for re-classifying all data when the first segment of data does not meet the preset requirements according to an embodiment of the present invention;
[0015] Figure 3 is a flowchart of a method for re-classifying the second segment of data and the data after the second segment when the first segment of data meets the preset requirements but the second segment of data does not meet the preset requirements according to an embodiment of the present invention;
[0016] Figure 4 is a schematic structural diagram of a level classification system for target data according to an embodiment of the present invention;
[0017] Figure 5 is a hardware structure block diagram of a computer terminal for a method of classifying the levels of target data according to an embodiment of the present invention. Detailed Description of the Embodiments
[0018] The present invention will be described in detail below with reference to the drawings and in conjunction with the embodiments. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.
[0019] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence.
[0020] In an embodiment of the present invention, a method for classifying the levels of target data is provided, Figure 1 is a flowchart of a method for classifying the levels of target data according to an embodiment of the present invention, as Figure 1 shown, and the process includes:
[0021] Step S101, arranging the index data to be classified in ascending order of data values;
[0022] In an exemplary embodiment, all the metric data to be classified is collected, and these data are sorted from smallest to largest using a sorting algorithm (such as bubble sort, quick sort, etc.). By sorting, the distribution of the data becomes more orderly, facilitating the subsequent calculation of differential volatility and providing a basis for the classification of the data.
[0023] Step S102, calculate the differential volatility values of two adjacent metric data after sorting to obtain a differential volatility sequence;
[0024] In one embodiment, the differential volatility value is calculated by taking the difference between the data value of the latter metric data and the data value of the adjacent former metric data as the numerator, and taking the data value of the latter metric data as the denominator; and when the denominator is zero, the differential volatility value is determined to be zero.
[0025] In an exemplary embodiment, for example, in a pair of sorted data, for two adjacent data: 469884 and 991795, the corresponding differential volatility is: For two adjacent data: 265221 and 369786, the corresponding differential volatility is: For two adjacent data: 0 and 1, since it is detected that the denominator is 0 when calculating the differential volatility, the corresponding differential volatility can be directly confirmed to be 0.
[0026] Therefore, the differential volatility sequence reflects the relative change amplitude between the data. A larger differential volatility value indicates a significant interval between the data, which helps to identify important break points in the data.
[0027] Step S103, sort the differential volatility sequence in descending order of the differential volatility value;
[0028] In an exemplary embodiment, the differential volatility sequence calculated in step S102 is sorted in descending order. For example, sorting 0.282772, 0.526229, 0, etc. in descending order to obtain: {0.526229, 0.282772, 0}. Therefore, by sorting, it is easier to identify significant fluctuation points in the data, and these points may be potential segmentation points for data segmentation.
[0029] Step S104, use two metric data with differential volatility values greater than the first preset threshold as the first segmentation point;
[0030] In an exemplary embodiment, a preset threshold K is selected. For example, K = max{the nth largest difference volatility value, M}, where n is a natural number and M is a rational number, both of which can be set based on actual situations. For example, when n = 6, M% = 1%, and the 6th largest difference volatility value is determined to be 0.333323, then K = max{0.33332, 31%} = 0.33332. For the difference volatilities: 0.526229, 0.282772, 0, it can be determined that the two adjacent data corresponding to 0.526229: 469884 and 991795 can be used as the first splitting point. Of course, the above example is just a simple one. In an actual classification scenario, there may be multiple such first splitting points, that is, there are multiple first splitting points. Therefore, determining the splitting points can divide the data into multiple segments. The data within each segment has similar characteristics, and the data between different segments has significant differences, thus providing a clear boundary for data grading.
[0031] Step S105: Arrange the index data in descending order of data values, and divide the index data into multiple data segments based on the first splitting point;
[0032] In an exemplary embodiment, in this step, the original data is rearranged in descending order of data values. According to the splitting points determined in step S104, the data is divided into multiple segments. For example, if 6 splitting points are determined, then the data will be divided into 7 data segments. Therefore, through reverse sorting and splitting, the data distribution in different segments can be seen more intuitively, which is convenient for subsequent calculation of the proportion of the data volume, and further refines the data grading.
[0033] Step S106: In the direction of descending data values, calculate the first ratio of the data volume of the first segment of data to the total data volume of the index data;
[0034] In an exemplary embodiment, in this step, calculate the data volume of the first segment of data, calculate the total data volume (i.e., the number of all data), and calculate the ratio of the data volume of the first segment of data to the total data volume. For example, if there are 10 data in the first segment of data and the total data volume is 100, then the first ratio is: Among them, the first ratio reflects the proportion of the first segment of data in the total data, helps to judge the importance of this segment of data, provides a basis for determining the data level, and makes the selected first segment of data have a reasonable data volume for further analysis.
[0035] Step S107: When the first ratio is less than or equal to the first proportion value, confirm that the first segment of data is of the first level, and calculate the data volume of the second segment of data and the second ratio of the sum of the data volumes of the first segment of data to the total data volume of the index data;
[0036] In an exemplary embodiment, in this step, a first ratio value is set, for example, 5%. If the first ratio is less than or equal to 5%, it indicates that the data volume of the first segment of data is reasonable, and then the first segment of data is confirmed as the first level. Calculate the sum of the data volume of the second segment of data and the data volume of the first segment of data. Calculate the ratio of this sum to the total data volume, that is, the second ratio. For example, if the data volume of the first segment is 10, the data volume of the second segment is 5, and the total data volume is 100, then the second ratio is: Therefore, by calculating the first ratio and the second ratio, the level of the data can be gradually determined, at least ensuring that the data volumes of the first level and the second level are statistically significant, and avoiding affecting the accuracy of classification due to too little data volume.
[0037] Step S108, in the case where the second ratio is less than or equal to the second ratio value, confirm that the second segment of data is the second level.
[0038] In an exemplary embodiment, in this step, a second ratio value is set, for example, 20%. Since the second ratio (0.15) is less than or equal to 20%, it can be confirmed that the first segment of data is the first level and the second segment of data is the second level. Therefore, through the judgment of the second ratio, the classification of the data can be further refined, ensuring the rationality and accuracy of the data volume and characteristics of each level, and finally realizing the scientific classification of the data.
[0039] Through the above steps S101 to S108, since the method of calculating the differential volatility value, determining the segmentation point according to its size, and then combining the data volume ratio to divide the data level is adopted, compared with the bucketing method and the layering method in the related art, it can more accurately identify the critical data in the adjacent level regions, and based on the first preset threshold, it can determine two adjacent index data (critical data) to be classified with larger differences, so as to avoid the problem that the difference between the critical data is not obvious. Therefore, the problem of inaccurate classification results in the related art can be solved, and the effect of improving the accuracy of the classification results can be achieved.
[0040] In an embodiment, after confirming that the second segment of data is the second level in the case where the second ratio is less than or equal to the second ratio value, the method further includes:
[0041] Confirm that the third segment of data is the third level;
[0042] Or, merge the third segment of data and the index data after the third segment of data, and confirm it as the third level.
[0043] In an exemplary embodiment, for example, during a classification process, if in a case where only three groups of grades of data need to be classified, and based on steps S101 to S108, the data of the first grade and the data of the second grade have been determined, then the divided third segment of data can be directly used as the third grade, or the third segment of data and the index data after the third segment of data can be combined as the data of the third grade. Therefore, flexible means can be adopted to meet different scenario requirements.
[0044] In one embodiment, after confirming that the third segment of data is the third grade, the method further includes:
[0045] Combine the fourth segment of data and the index data after the fourth segment of data, and confirm it as the fourth grade.
[0046] In an exemplary embodiment, for example, during a classification process, if in a case where only four groups of grades of data need to be classified, and based on steps S101 to S108, the data of the first grade and the data of the second grade have been determined, and the data of the third grade has been determined, then the divided fourth segment of data can be directly used as the fourth grade, or the fourth segment of data and the index data after the fourth segment of data can be combined as the data of the fourth grade. Therefore, flexible means can be adopted to meet different scenario requirements.
[0047] In one embodiment, Figure 2 is a flowchart of a method for re-classifying all data when the first segment of data does not meet the preset requirements according to an embodiment of the present invention. As Figure 2 shown, after calculating the first ratio of the data volume of the first segment of data to the total data volume of the index data in the direction from large to small of the data values, it further includes:
[0048] Step S201, when the first ratio is greater than the first proportional value, arrange the index data to be classified in ascending order of data values;
[0049] In an exemplary embodiment, if there is a case where the first ratio is greater than the first proportional value, it indicates that the data volume of the first segment of data obtained from the initial classification does not meet the statistical requirements (the data volume is too large), and re-sorting and classification are required. Therefore, re-arrange the index data to be classified in ascending order of data values.
[0050] Step S202, calculate the differential volatility values of two adjacent index data after sorting to obtain a differential volatility sequence;
[0051] In an exemplary embodiment, this step may be similar to step S102 and the same technical means may be adopted. Alternatively, when the result obtained in step S102 is known, the result obtained in step S102 may be directly called as the result of this step.
[0052] Step S203: Arrange the difference volatility sequences in descending order of the difference volatility values.
[0053] In an exemplary embodiment, this step may be similar to step S103.
[0054] Step S204: Use the two index data with difference volatility values greater than the second preset threshold as the second segmentation points, where the second preset threshold is greater than the first preset threshold.
[0055] In an exemplary embodiment, if there is a first ratio greater than the first proportional value, it indicates that the data volume of the first segment of data obtained by the initial classification does not meet the statistical requirements (the data volume is too large). Therefore, in this step, the control value (the second preset threshold) needs to be increased. For example, n = 4 and M% = 0 can be directly set. Then, the fourth largest difference volatility value will be determined as the second preset threshold to select the second segmentation point based on the second preset threshold.
[0056] Step S205: Arrange the index data in descending order of the data values and divide the index data into multiple data segments based on the second segmentation points.
[0057] In an exemplary embodiment, this step is similar to step S105.
[0058] Step S206: In the direction from the largest to the smallest data value, calculate the third ratio of the data volume of the new first segment of data to the total data volume of the index data until the third ratio is less than or equal to the first proportional value, and then confirm the new first segment of data as the first level, and calculate the fourth ratio of the sum of the data volume of the new second segment of data and the data volume of the new first segment of data to the total data volume of the index data.
[0059] In an exemplary embodiment, based on steps S201 to S205, when the second segmentation points are reselected based on the new second preset threshold, the original data is re-segmented so that the data volume of the first segment of data is reduced to meet the classification and statistical requirements. If the third ratio is still greater than the first proportional value, steps S201 to S205 are continued to be repeated (further increasing the second preset threshold) until the third ratio is less than or equal to the first proportional value, and then the new first segment of data is confirmed as the first level. Therefore, the first segment of data that meets the requirements can be confirmed as the data of the first level.
[0060] Step S207, when the fourth ratio is less than or equal to the second ratio value, confirm that the new second-segment data is of the second level.
[0061] In an exemplary embodiment, this step is similar to step S108 to confirm the new second-segment data as data of the second level. In this way, the data of the reclassified first level and second level can meet the requirements.
[0062] Through the above steps S201 to S207, the second segmentation point of the original data can be reconfirmed by using the new second preset threshold, so that the first-segment data and the second-segment data re-segmented based on the second segmentation point meet the preset requirements, that is, the data of the first level and the second level that meet the requirements are confirmed.
[0063] In one embodiment, Figure 3 is a flowchart of a method for reclassifying the second-segment data and the data after the second-segment data when the first-segment data meets the preset requirements but the second-segment data does not meet the preset requirements according to an embodiment of the present invention. As Figure 3 shown, the method further includes:
[0064] Step S301, when the second ratio is greater than the second ratio value, use the second-segment data and the index data after the second-segment data as new data to be classified, and arrange the new data to be classified in ascending order of data values;
[0065] In an exemplary embodiment, if the second ratio is greater than the second ratio value (for example, 20%), it means that the data volume of the current second-segment data is too large and needs to be reclassified. Extract the second-segment data and all the subsequent data to form a new dataset of data to be classified. Use a sorting algorithm (such as bubble sort, quick sort, etc.) to sort these new data to be classified in ascending order. Therefore, by re-extracting and sorting, the second-segment data with too large a data volume is reclassified to ensure the rationality and accuracy of the data volume and characteristics of each level. This step is similar to step S101, and the re-sorting is for subsequent steps to calculate the differential volatility and determine the segmentation point based on the new dataset.
[0066] Step S302, calculate the differential volatility values of adjacent data values after sorting to obtain a differential volatility sequence;
[0067] In an exemplary embodiment, for the sorted new classification target data, the differential volatility of each pair of adjacent data is calculated. The specific calculation method is as follows: The difference between the latter data value and the former data value is used as the numerator, and the latter data value is used as the denominator for calculation. If the denominator is zero, the differential volatility value is determined to be zero. Therefore, the differential volatility sequence reflects the relative change amplitude between the data. A larger differential volatility value indicates a significant interval between the data, which helps to identify important break points in the data. This step is the same as step S102 and provides the differential volatility sequence for step S303, which is a key step in identifying data segments.
[0068] Step S303: Arrange the differential volatility sequence in descending order of the differential volatility value.
[0069] In an exemplary embodiment, in this step, the differential volatility sequence calculated in step S302 can be sorted to be in descending order. Therefore, through sorting, it is easier to identify significant fluctuation points in the data, and these points may be potential segmentation points for data segmentation. This step is similar to step S103 and provides the sorted differential volatility sequence for step S304 to help determine new segmentation points.
[0070] Step S304: Use two index data with differential volatility values greater than the third preset threshold as the third segmentation point.
[0071] In an exemplary embodiment, a preset threshold K is selected 3 , for example, K 3 = max (the differential volatility value of the nth largest, M, where, for example, n = 2 can be directly set, and M% = 0. Then, the second largest differential volatility value is determined as the third preset threshold to select the third segmentation point based on the third preset threshold. It should be noted that since this step re-classifies the data in the second segment and the data after the second segment, there is no fixed size relationship between the third preset threshold and the first preset threshold and the second preset threshold, and it can be set based on the actual situation. For example, the third preset threshold can be less than the first preset threshold or the second preset threshold, or the third preset threshold can be equal to the first preset threshold or the second preset threshold, or the third preset threshold can be greater than the first preset threshold or the second preset threshold.
[0072] Therefore, through step S304, the third segmentation point of the data in the second segment and the data after the second segment can be re-determined based on the new third preset threshold, so that the data within each segment has similar characteristics, and the data between different segments has significant differences, thereby providing a clear boundary for data classification. This step is similar to step S104 but uses different preset thresholds to handle the situation of excessive data volume and ensure the accuracy of classification.
[0073] Step S305: Arrange the new data of the to-be-classified metrics in descending order of data values, and divide the metric data into multiple data segments based on the third splitting point.
[0074] In an exemplary implementation, this step is similar to step S105.
[0075] Step S306: In the direction of descending data values, take the new first data segment as the new second data segment.
[0076] In an exemplary implementation, in this step, the first data segment in the new segmentation obtained in step S305 can be regarded as the new second data segment. Therefore, by adjusting the numbering of the data segments, it is convenient to calculate the ratio of the data volume in subsequent steps, ensuring the coherence and accuracy of the grading. This step is similar to step S106, but for the new data segments, and is used to recalculate the ratio of the data volume.
[0077] Step S307: Calculate the second ratio of the sum of the data volume of the new second data segment and the data volume of the first data segment to the total data volume of the metric data.
[0078] In an exemplary implementation, in this step, the data volume of the new second data segment can be calculated, the data volume of the first data segment can be calculated, and the total data volume (i.e., the number of all data) can be calculated. Calculate the ratio of the sum of the data volume of the new second data segment and the data volume of the first data segment to the total data volume, that is, the second ratio. Therefore, the second ratio reflects the total proportion of the new second data segment and the first data segment, helps to judge the importance of these data segments, provides a basis for determining the data level, and ensures that the data volume of these data segments has statistical significance. This step is similar to step S107, but for the new data segments, and is used to reconfirm the data level.
[0079] Step S308: When the second ratio is less than or equal to the second proportional value, confirm that the new second data segment is the second level.
[0080] In an exemplary implementation, in this step, a second proportional value is set, for example, 20%. If the second ratio is less than or equal to 20%, then confirm that the new second data segment is the second level. Therefore, through the judgment of the second ratio, the grading of the data can be further refined, ensuring the rationality and accuracy of the data volume and characteristics of each level, and finally realizing the scientific classification of the data. This step is similar to step S108 and is the final confirmation step for data grading, ensuring the accuracy and rationality of the grading result.
[0081] Through the above steps S301 to S308, the embodiments of the present invention provide a flexible method for classifying target data levels. By reclassifying and adjusting the preset thresholds, this method can handle the situation of excessive data volume and ensure the rationality and accuracy of the data volume and features of each level. Specifically, through multiple adjustments and reclassifications, it is ensured that the final classification result meets the statistical requirements, improving the accuracy and reliability of the classification result.
[0082] It should be noted that, as an example, the above solution can also be applied to the scenario of classifying certain business data of a bank. Of course, the above solution is not limited to this scenario.
[0083] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software adding the necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0084] In the embodiments of the present invention, a system for classifying target data levels is also provided. This system is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0085] Figure 4 is a schematic structural diagram of a system for classifying target data levels according to an embodiment of the present invention, as Figure 4 shown, this system includes:
[0086] A first sorting module 41, configured to sort the index data to be classified in ascending order of data values;
[0087] A first calculation module 42, configured to calculate the differential volatility values of two adjacent index data after sorting to obtain a differential volatility sequence;
[0088] A second sorting module 43, configured to sort the differential volatility sequence in descending order of differential volatility values;
[0089] A split point determination module 44, configured to use two index data with differential volatility values greater than a first preset threshold as the first split point;
[0090] A third sorting module 45, for arranging the indicator data in descending order of data value, and dividing the indicator data into a plurality of data segments based on the first segmentation point;
[0091] A second calculation module 46 is used to calculate a first ratio of the amount of data of the first segment of data to the total amount of data of the indicator data in a direction from large to small data values;
[0092] A third calculation module 47 is used to determine that the first segment of data is of the first level when the first ratio is less than or equal to the first ratio value, and calculate a second ratio of the sum of the data volume of the second segment of data and the data volume of the first segment of data to the total data volume of the indicator data;
[0093] The processing module 48 is used to confirm that the second segment of data is of the second level when the second ratio is less than or equal to the second ratio value.
[0094] By adopting the above technical solution, since the embodiment of the present invention uses the first sorting module 41 to sort the indicator data from small to large according to the data value, the first calculation module 42 calculates the differential volatility sequence of adjacent indicator data, the second sorting module 43 sorts from large to small according to the differential volatility value, the segmentation point determination module 44 uses two indicator data with differential volatility values greater than the first preset threshold as the first segmentation point, the third sorting module 45 divides the indicator data into multiple data segments based on the first segmentation point, the second calculation module 46 calculates the data volume ratio of the first segment data, and the third calculation module 47 confirms that the first segment data is the first level when certain conditions are met, and calculates The ratio of the sum of the data volume of the first segment of data and the data volume of the second segment of data is calculated, and the processing module 48 confirms that the second segment of data is of the second level when certain conditions are met. This multi-step division method that comprehensively considers the differential volatility and the data volume ratio can more accurately identify the critical data of adjacent level areas, and can determine two adjacent indicator data (critical data) to be classified with large differences based on the first preset threshold, thereby avoiding the problem of unclear distinction between critical data in related technologies. Therefore, the problem of inaccurate classification results in related technologies can be solved, thereby achieving the effect of improving the accuracy of the classified results.
[0095] In one embodiment, the system is further used to: confirm that the third segment of data is a third level;
[0096] Alternatively, the third segment of data and the indicator data after the third segment of data are combined and confirmed as the third level.
[0097] In one embodiment, the system is also used to: merge the fourth end data and the indicator data after the fourth segment data, and confirm them as the fourth level.
[0098] In one embodiment, the system is further configured to: when the first ratio is greater than the first proportional value, arrange the index data to be classified in ascending order of the data values;
[0099] Calculate the differential volatility values of two adjacent index data after sorting to obtain a differential volatility sequence;
[0100] Arrange the differential volatility sequence in descending order of the differential volatility values;
[0101] Use two index data with differential volatility values greater than the second preset threshold as the second segmentation point; wherein, the second preset threshold is greater than the first preset threshold;
[0102] Arrange the index data in descending order of the data values, and divide the index data into multiple data segments based on the second segmentation point;
[0103] In the direction from large to small of the data values, calculate the third ratio of the data volume of the new first data segment to the total data volume of the index data until the third ratio is less than or equal to the first proportional value, then confirm the new first data segment as the first level, and calculate the fourth ratio of the sum of the data volume of the new second data segment and the data volume of the new first data segment to the total data volume of the index data;
[0104] When the fourth ratio is less than or equal to the second proportional value, confirm the new second data segment as the second level.
[0105] In one embodiment, the system is further configured to: when the second ratio is greater than the second proportional value, use the second data segment and the index data after the second data segment as the new index data to be classified, and arrange the new index data to be classified in ascending order of the data values;
[0106] Calculate the differential volatility values of adjacent data values after sorting to obtain a differential volatility sequence;
[0107] Arrange the differential volatility sequence in descending order of the differential volatility values;
[0108] Use two index data with differential volatility values greater than the third preset threshold as the third segmentation point;
[0109] Arrange the new index data to be classified in descending order of the data values, and divide the index data into multiple data segments based on the third segmentation point;
[0110] In the direction from large to small of the data values, use the new first data segment as the new second data segment;
[0111] Calculate the second ratio of the sum of the data volume of the new second data segment and the data volume of the first data segment to the total data volume of the index data;
[0112] When the second ratio is less than or equal to the second proportional value, confirm that the new second-segment data is of the second level.
[0113] In one implementation, the differential volatility value is calculated by using the difference between the data value of the subsequent metric data and the data value of the adjacent previous metric data as the numerator, and the data value of the subsequent metric data as the denominator; and when the denominator is zero, determine that the differential volatility value is zero.
[0114] The method embodiments provided in the embodiments of the present invention can be executed on a mobile terminal, a computer terminal or a similar computing device. Taking running on a computer terminal as an example, Figure 5 is a hardware structure block diagram of a computer terminal for the hierarchical classification method of the target data in the embodiments of the present invention. As Figure 5 shown, the computer terminal may include one or more ( Figure 5 only one is shown in the figure) processors 102 (the processors 102 may include, but are not limited to, processing devices such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above computer terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 5 the structure shown is only schematic and does not limit the structure of the above computer terminal. For example, the computer terminal may further include more or fewer components than those shown in Figure 5 the figure, or have a different configuration from that shown in Figure 5 the figure.
[0115] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to the hierarchical classification method of the target data in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, a flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories may be connected to the computer terminal through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0116] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include wired / wireless networks provided by communication providers of computer terminals. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one example, the transmission device 106 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0117] It should be noted that the above-mentioned modules can be implemented by software or hardware. For the latter, it can be achieved in the following ways, but not limited to: the above-mentioned modules are all located in the same processor; or, the above-mentioned modules are respectively located in different processors in any combination form.
[0118] An embodiment of the present invention also provides a computer-readable storage medium, in which a computer program is stored. Among them, the computer program is set to execute the steps in any one of the above method embodiments when running.
[0119] In an exemplary embodiment, the above-mentioned computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (abbreviated as ROM), random access memories (abbreviated as RAM), mobile hard disks, magnetic disks or optical disks, etc., various media that can store computer programs.
[0120] An embodiment of the present invention also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is set to run the computer program to execute the steps in any one of the above method embodiments.
[0121] In an exemplary embodiment, the above-mentioned electronic device may further include a transmission device and an input / output device. Among them, the transmission device is connected to the above-mentioned processor, and the input / output device is connected to the above-mentioned processor.
[0122] An embodiment of the present invention also provides a computer program product, including a computer program, which implements the steps in any one of the above method embodiments when executed by a processor.
[0123] Specific examples in this embodiment can refer to the examples described in the above embodiments and exemplary embodiments, and will not be repeated here.
[0124] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed over a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a sequence different from that here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the present invention is not limited to any specific combination of hardware and software.
[0125] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for hierarchical classification of target data, characterized in that: include: Arrange the indicator data to be classified in order from small to large data values; Calculating the differential volatility values of two adjacent index data after sorting to obtain a differential volatility sequence; Arrange the differential volatility sequence in descending order of differential volatility values; The two indicator data whose difference volatility values are greater than the first preset threshold are used as the first segmentation points; Arrange the indicator data in descending order of data value, and divide the indicator data into a plurality of data segments based on the first segmentation point; In the direction of the data values from large to small, calculating a first ratio of the data volume of the first segment of data to the total data volume of the indicator data; When the first ratio is less than or equal to the first ratio value, confirm that the first segment of data is at the first level, and calculate a second ratio of the sum of the data volume of the second segment of data and the data volume of the first segment of data to the total data volume of the indicator data; When the second ratio is less than or equal to the second ratio value, the second segment data is confirmed to be of the second level.
2. The method according to claim 1, characterized in that When the second ratio is less than or equal to the second ratio value, after confirming that the second segment of data is of the second level, the method further includes: Confirm that the third segment of data is the third level; Alternatively, the third segment of data and the indicator data after the third segment of data are combined and confirmed as the third level.
3. The method according to claim 2, characterized in that After confirming that the third segment of data is the third level, it also includes: The fourth end data and the indicator data after the fourth segment data are combined and confirmed as the fourth level.
4. The method according to claim 1, characterized in that: After calculating a first ratio of the amount of the first segment of data to the total amount of the indicator data in a direction from large to small of the data value, the method further includes: When the first ratio is greater than the first proportion value, the indicator data to be classified are arranged in order of data value from small to large; Calculating the differential volatility values of two adjacent index data after sorting to obtain a differential volatility sequence; Arrange the differential volatility sequence in descending order of differential volatility values; The two indicator data whose differential volatility values are greater than a second preset threshold are used as second segmentation points; wherein the second preset threshold is greater than the first preset threshold; Arrange the indicator data in descending order of data value, and divide the indicator data into a plurality of data segments based on the second segmentation point; In the direction of the data values from large to small, a third ratio of the data volume of the new first segment of data to the total data volume of the indicator data is calculated, until the third ratio is less than or equal to the first ratio value, the new first segment of data is confirmed to be the first level, and a fourth ratio of the sum of the data volume of the new second segment of data and the data volume of the new first segment of data to the total data volume of the indicator data is calculated; When the fourth ratio is less than or equal to the second ratio, the new second segment data is confirmed to be of the second level.
5. The method according to claim 1, characterized in that Also includes: When the second ratio is greater than the second proportion value, the second segment of data and the indicator data after the second segment of data are used as new indicator data to be classified, and the new indicator data to be classified are arranged in ascending order of data value; Calculating the differential volatility values of the sorted adjacent data values to obtain a differential volatility sequence; Arrange the differential volatility sequence in descending order of differential volatility values; The two indicator data whose differential volatility values are greater than the third preset threshold are used as the third split point; Arrange the new indicator data to be classified in descending order of data value, and divide the indicator data into a plurality of data segments based on the third segmentation point; In the direction from large to small of the data value, the new first segment of data is used as the new second segment of data; Calculate a second ratio of the sum of the data volume of the new second segment of data and the data volume of the first segment of data to the total data volume of the indicator data; When the second ratio is less than or equal to the second ratio value, the new second segment data is confirmed to be of the second level.
6. The method according to claim 1, characterized in that The differential volatility value is calculated based on taking the difference between the data value of the latter indicator data and the data value of the adjacent previous indicator data as the numerator and the data value of the latter indicator data as the denominator; and when the denominator is zero, the differential volatility value is determined to be zero.
7. A hierarchical classification system for target data, characterized in that: include: The first sorting module is used to sort the indicator data to be classified in order from small to large data values; A first calculation module, used for calculating the differential volatility values of two adjacent index data after sorting, so as to obtain a differential volatility sequence; A second sorting module is used to sort the differential volatility sequence in descending order of differential volatility values; A split point determination module, used for taking the two indicator data whose differential volatility values are greater than a first preset threshold as the first split point; A third sorting module, used to arrange the indicator data in descending order of data value, and divide the indicator data into a plurality of data segments based on the first segmentation point; A second calculation module is used to calculate a first ratio of the amount of data of the first segment of data to the total amount of data of the indicator data in a direction from large to small of the data value; A third calculation module is used to determine that the first segment of data is of the first level when the first ratio is less than or equal to the first ratio value, and calculate a second ratio of the sum of the data volume of the second segment of data and the data volume of the first segment of data to the total data volume of the indicator data; The processing module is used to confirm that the second segment of data is of the second level when the second ratio is less than or equal to the second ratio value.
8. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program executes the steps of the method described in any one of claims 1 to 6 when executed by a processor.
9. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the steps of the method according to any one of claims 1 to 6.
10. A computer program product, characterized in that The method comprises computer instructions, which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 6.