A data processing method and device, electronic equipment and storage medium

By acquiring imbalance monitoring data, establishing an initial imbalanced dataset, determining the imbalance ratio, performing data preprocessing and rotation operations, and expanding the minority class dataset, the problems of overfitting and information loss in traditional methods are solved, and the balance optimization of the dataset and the improvement of minority class sample diversity are achieved.

CN120610958BActive Publication Date: 2025-11-07CHINA MOBILE QUANTONG SYST INTEGRATION CO LTD +3
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511108080.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-11-07
Estimated Expiration
2045-08-08

AI Technical Summary

Technical Problem

In supervised data analysis, due to the small amount of minority class data, traditional oversampling techniques are prone to model overfitting, while undersampling techniques may lead to the loss of useful information. There is an urgent need for a method that can effectively improve the diversity of minority class samples and reduce the risk of overfitting.

Method used

By acquiring imbalance monitoring data, an initial imbalance dataset is established, the imbalance ratio is determined, data preprocessing operations are performed, an intermediate imbalance dataset and a data rotation interval are generated, the minority class dataset is expanded, its representativeness is enhanced, and the risk of overfitting is reduced.

Benefits of technology

It effectively improves the representativeness of minority class samples in classification tasks, reduces the risk of overfitting, and achieves balanced optimization of the dataset.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120610958B_ABST
    Figure CN120610958B_ABST
Patent Text Reader

Abstract

The application discloses a data processing method and device, electronic equipment and storage medium, and relates to the technical field of data processing. The data processing method comprises the following steps: acquiring unbalanced monitoring data, and establishing an initial unbalanced data set of the unbalanced monitoring data; wherein the difference between the majority class data and the minority class data in the unbalanced monitoring data exceeds a preset threshold; determining the unbalance ratio of the initial unbalanced data set, performing a data preprocessing operation on the initial unbalanced data set to obtain an intermediate unbalanced data set and a data rotation interval; acquiring the minority class data set and the majority class data set in the intermediate unbalanced data set, and determining the intermediate minority class data set corresponding to the minority class data set based on the data rotation interval and the unbalance ratio; determining a target minority class data set according to the majority class data set and the intermediate minority class data set, and combining the majority class data set and the target minority class data set to generate a target unbalanced data set, so that the diversity of the minority class sample is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and in particular to a data processing method and device, electronic equipment and storage medium. BACKGROUND

[0002] In supervision data analysis, the class distribution of supervision objects or problem clues is often uneven. One class of data can be much larger than another class, which is called class imbalance of supervision data. Among them, the class with more data is called the majority class, which usually reflects common and frequent behaviors; while the class with less data is called the minority class, which may represent rare but serious problem clues. Generally speaking, the minority class data can represent abnormal behavior patterns, and the majority class data can represent normal data. Many major problems often hide in a small number of samples, and traditional models are easily dominated by a large number of "normal behavior" samples, ignoring a small number of abnormal behaviors.

[0003] The technical problems to be solved by the present application include: in the oversampling technology, since the data elements are generated comprehensively, overfitting problems are easily generated in the model construction process; and in the undersampling technology, in order to balance the data set and eliminate part of the data elements, useful information may be lost. The traditional and simplest oversampling method-random oversampling technology realizes oversampling by randomly selecting and duplicating original minority class data elements, but since the same information is used multiple times, it is easy to cause model overfitting. Therefore, there is an urgent need for a method that can effectively improve the diversity of minority class samples and reduce the risk of overfitting. SUMMARY

[0004] The present application provides a data processing method and device, electronic equipment and storage medium to solve the problem of small amount of minority class data in the prior art and easy overfitting in the data analysis process.

[0005] According to an aspect of the present application, a data processing method is provided, wherein the method comprises:

[0006] Obtaining unbalanced supervision data, and establishing an initial unbalanced data set of the unbalanced supervision data; wherein the difference between the majority class data and the minority class data in the unbalanced supervision data exceeds a preset threshold;

[0007] Determining the imbalance ratio of the initial unbalanced data set, performing a data preprocessing operation on the initial unbalanced data set to obtain an intermediate unbalanced data set, and determining a data rotation interval;

[0008] Obtaining the minority class data set and the majority class data set in the intermediate unbalanced data set, and determining the intermediate minority class data set corresponding to the minority class data set based on the data rotation interval and the imbalance ratio;

[0009] determine a target minority class data set according to the majority class data set and the intermediate minority class data set, and combine the majority class data set and the target minority class data set to generate a target unbalanced data set.

[0010] According to another aspect of the present application, there is provided a data processing apparatus, wherein the apparatus comprises:

[0011] a data acquisition module configured to acquire unbalanced supervision data, and establish an initial unbalanced data set of the unbalanced supervision data; wherein a difference between majority class data and minority class data in the unbalanced supervision data exceeds a preset threshold;

[0012] an interval determination module configured to determine an unbalance ratio of the initial unbalanced data set, and perform a data preprocessing operation on the initial unbalanced data set to obtain an intermediate unbalanced data set and a data rotation interval;

[0013] a data processing module configured to acquire a majority class data set and a minority class data set in the intermediate unbalanced data set, and determine an intermediate minority class data set corresponding to the minority class data set based on the data rotation interval and the unbalance ratio;

[0014] a data set adjustment module configured to determine a target minority class data set according to the majority class data set and the intermediate minority class data set, and combine the majority class data set and the target minority class data set to generate a target unbalanced data set.

[0015] According to another aspect of the present application, there is provided an electronic device, comprising:

[0016] at least one processor; and

[0017] a memory communicatively connected to the at least one processor; wherein

[0018] the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform a data processing method according to any one of the embodiments of the present application.

[0019] According to another aspect of the present application, there is provided a computer readable storage medium storing computer instructions for enabling a processor to perform a data processing method according to any one of the embodiments of the present application when executed by the processor.

[0020] The technical scheme of the embodiment of the application comprises the following steps: obtaining unbalanced monitoring data, establishing an initial unbalanced data set of the unbalanced monitoring data, determining an unbalanced ratio of the initial unbalanced data set, performing a data preprocessing operation on the initial unbalanced data set to obtain an intermediate unbalanced data set and a data rotation interval, and realizing simulation of virtual samples; obtaining a minority class data set and a majority class data set in the intermediate unbalanced data set, determining an intermediate minority class data set corresponding to the minority class data set based on the data rotation interval and the unbalanced ratio, and realizing expansion of the minority class data; determining a target minority class data set according to the majority class data set and the intermediate minority class data set, and combining the majority class data set and the target minority class data set to generate a target unbalanced data set, thereby enhancing the representativeness of the minority class sample in a classification task and effectively improving the diversity of the minority class sample and reducing the risk of overfitting.

[0021] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the application, nor is it used to limit the scope of the application. Other features of the application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can also be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0023] Figure 1 is a flow chart of a data processing method according to an embodiment of the application;

[0024] Figure 2 is a flow chart of a data processing method according to an embodiment of the application;

[0025] Figure 3 is a structural schematic diagram of a data processing device according to an embodiment of the application;

[0026] Figure 4 is a structural schematic diagram of an electronic device for implementing the data processing method of the embodiment of the application. DETAILED DESCRIPTION

[0027] In the following, the technical solutions in the embodiments of the present application will be described clearly and completely with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all the other embodiments obtained by a person of ordinary skill in the art without creative work should belong to the protection scope of the present application.

[0028] It should be noted that the terms "first", "second" and the like in the description and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in other than the order illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a list of steps or units need not be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to such processes, methods, products or devices.

[0029] In the technical solutions of the present application, the acquisition, storage, use, processing and the like of data comply with the relevant provisions of national laws and regulations.

[0030] It should be noted that in the embodiments of the present application, some software, components, models and the like may be mentioned, which should be considered as exemplary, and the purpose is only to illustrate the feasibility of the implementation of the technical solutions of the present application, but does not mean that the applicant has or will necessarily use the scheme.

[0031] Embodiment one

[0032] Figure 1 is a flowchart of a data processing method according to the first embodiment of the present application. The present embodiment can be applicable to the case of effectively overfitting the minority class data in the unbalanced data set. The method can be executed by a data processing device, which can be realized in the form of hardware and / or software, and can be configured in an electronic device. As shown in the figure, the method comprises: Figure 1

[0033] S110, acquiring unbalanced monitoring data, and establishing an initial unbalanced data set of the unbalanced monitoring data.

[0034] ​The unbalanced supervision data refers to data in which the number of data in each category is significantly different in the supervision field. In actual operation, the unbalanced supervision data can be discipline inspection and supervision data. Generally, the unbalanced supervision data can be binary class data, that is, divided into majority class data and minority class data, and the difference between the majority class data and the minority class data in the unbalanced supervision data exceeds a preset threshold, that is, the majority class data is much more than the minority class data. Generally, the minority class data can represent abnormal behavior patterns, and the majority class data can represent normal data. The preset threshold can be understood as a threshold for defining the majority class data and the minority class data, which can be set according to the needs of business personnel. The initial unbalanced data set refers to a data set composed of unbalanced supervision data, used to store all unbalanced supervision data.

[0035] In the embodiment, the supervision data can be obtained through a public website or business needs. If the difference between the majority class data and the minority class data exceeds the preset threshold, it can be confirmed that the supervision data is unbalanced supervision data, and the initial unbalanced data set of the unbalanced supervision data is established. In actual operation, data sets of the majority class data and the minority class data in the unbalanced supervision data can be established respectively, and the two data sets are combined as the initial unbalanced data set; or the unbalanced supervision data can be directly stored in the preset data set as the initial unbalanced data set. The arrangement order of the unbalanced supervision data in the initial unbalanced data set can be preset or random, which is not limited.

[0036] S120, determine the imbalance ratio of the initial unbalanced data set, perform data preprocessing operation on the initial unbalanced data set to obtain an intermediate unbalanced data set and a data rotation interval.

[0037] The imbalance ratio refers to a core index for measuring the degree of class imbalance of a data set, used to quantify the proportional relationship between the number of majority class samples and the number of minority class samples in the data set. In actual application, the determination of the imbalance ratio is the ratio of the number of majority class samples to the number of minority class samples. The data preprocessing operation refers to a process of normalizing, rotating and other operations on data to meet the requirements of subsequent tasks. The intermediate unbalanced data set refers to a new data set obtained after the initial unbalanced data set performs data preprocessing operation. The data rotation interval refers to the angle interval of data rotation, which is convenient for subsequent two-dimensional rotation operation on the data in the data set.

[0038] In an embodiment, the number of majority class samples and the number of minority class samples in the initial unbalanced data set can be determined respectively, and the ratio of the number of majority class samples to the number of minority class samples is determined as the imbalance ratio of the initial unbalanced data set. Normalization, rotation and other data preprocessing operations are performed on the initial unbalanced data set to determine the intermediate unbalanced data set and the data rotation interval. In actual operation, the normalization operation can be performed on the data in the initial unbalanced data set first to effectively eliminate the difference in numerical range of different features. Further, the normalized data can be subjected to two-dimensional rotation operation to facilitate the generation of more representative data to alleviate the imbalance problem.

[0039] In an embodiment, a two-dimensional rotation matrix can be extracted, the angle interval of the two-dimensional rotation matrix is determined, and the data in the initial unbalanced data set is divided into continuous attribute pairs one by one. The data rotation operation is performed on the continuous attribute pairs according to the two-dimensional rotation matrix, and a new attribute pair is obtained as a rotated attribute pair. In actual application, the rotated attribute pair can be determined according to each angle of the angle interval, and the variance of the rotated attribute pair and the continuous attribute pair is determined. The angle corresponding to the variance that meets the preset threshold interval is taken as the data rotation interval. When the number of data in the initial unbalanced data set is odd, the last data and the previous data form a continuous attribute pair, and after rotating the last continuous attribute pair, the data is discarded because it is reused in the rotation of the previous attribute pair. Then, any angle in the data rotation interval is selected as a target angle, and the data set formed by the rotated attribute pair is rotated again according to the target angle, and the rotated data is placed in the order of the original data set to obtain the intermediate unbalanced data set.

[0040] S130, obtaining the minority class data set and the majority class data set in the intermediate unbalanced data set, and determining the intermediate minority class data set corresponding to the minority class data set based on the data rotation interval and the imbalance ratio.

[0041] Among them, the minority class data set refers to the data set composed of minority class data in the intermediate unbalanced data set, and the majority class data set refers to the data set composed of majority class data in the intermediate unbalanced data set. The intermediate minority class data set refers to the minority class data set determined according to the data rotation interval, the imbalance ratio and the minority class data set adjustment.

[0042] In an embodiment, minority class data in the intermediate unbalanced data set can be extracted to form a minority class data set, and majority class data in the intermediate unbalanced data set can be extracted to form a majority class data set. Based on each angle in the data rotation interval, the data in the minority class data set is rotated in pairs according to the two-dimensional rotation matrix for the continuous attribute pairs, to obtain a plurality of rotated data pairs. The minority rotated attribute pairs belonging to the same rotation number are arranged in the order of the minority continuous attribute pairs to obtain a temporary minority class data set, and the minority class data set is expanded according to the temporary minority class data set to obtain an intermediate minority class data set. In order to ensure that the minority class data is at least twice the majority class data, further rotating the minority class data will definitely make the minority class data more, that is, the data is more balanced. The rotation number of each minority continuous attribute pair can be set to twice the imbalance ratio. Then, the minority rotated attributes are inserted into the intermediate unbalanced data set in the order of the corresponding continuous attribute pairs to obtain the intermediate minority class data set. Alternatively, the minority rotated attribute pairs belonging to the same rotation number can be arranged in the order of the minority continuous attribute pairs to obtain a temporary minority class data set, and the minority class data set is expanded according to the temporary minority class data set to obtain an intermediate minority class data set.

[0043] In S140, a target minority class data set is determined according to the majority class data set and the intermediate minority class data set, and the majority class data set and the target minority class data set are combined to generate a target unbalanced data set.

[0044] In an embodiment, the target minority class data set refers to a minority class data set that has completed oversampling. In actual operation, the data quantity of the target minority class data set can be the same as the data quantity in the majority class data set. The target unbalanced data set refers to a data set that has completed data unbalance optimization.

[0045] In an embodiment, the quantity of the majority class data set can be determined as the clustering quantity, and the data in the intermediate minority class data set is clustered according to the clustering quantity to obtain a final target centroid as the oversampled minority class data. The oversampled minority class data is stored as the target minority class data set, and the majority class data set and the target minority class data set are combined to obtain the target unbalanced data set. In actual operation, the clustering operation can adopt a K-means clustering algorithm, a partition clustering, etc., and each cluster represents a specific type of abnormal behavior pattern.

[0046] The embodiment of the application realizes simulation of virtual samples by acquiring unbalanced monitoring data, establishing an initial unbalanced data set of the unbalanced monitoring data, determining an imbalance ratio of the initial unbalanced data set, and performing a data preprocessing operation on the initial unbalanced data set to obtain an intermediate unbalanced data set and a data rotation interval; the embodiment realizes expansion of the minority class data by acquiring a minority class data set and a majority class data set in the intermediate unbalanced data set and determining an intermediate minority class data set corresponding to the minority class data set based on the data rotation interval and the imbalance ratio; and the embodiment enhances the representativeness of the minority class samples in the classification task by determining a target minority class data set according to the majority class data set and the intermediate minority class data set, and generating a target unbalanced data set by combining the majority class data set and the target minority class data set, thereby effectively improving the diversity of the minority class samples and reducing the overfitting risk.

[0047] Embodiment Two

[0048] Figure 2 is a flowchart of a data processing method according to the embodiment two of the application. The embodiment is further optimized and expanded based on the above-mentioned embodiments and can be combined with each optional technical solution in the above-mentioned embodiments. As shown in the figure, the method comprises: Figure 2

[0049] S201, acquiring unbalanced monitoring data and determining majority class data and minority class data in the unbalanced monitoring data.

[0050] In the embodiment, the unbalanced monitoring data can be divided to determine the number of data of different categories, and the data of a category with a large amount of data is taken as the majority class data, and the data of a category with a small amount of data is taken as the minority class data.

[0051] S202, storing the majority class data as an initial majority class data set, storing the minority class data as an initial minority class data set, and combining the initial majority class data set and the initial minority class data set as an initial unbalanced data set.

[0052] S203, respectively determining the number of majority class data and the number of minority class data in the initial unbalanced data set, and taking the ratio of the number of majority class data to the number of minority class data as an imbalance ratio of the initial unbalanced data set.

[0053] In the embodiment, the number of majority class data and the number of minority class data in the initial unbalanced data set can be determined, the ratio of the number of majority class data to the number of minority class data can be determined, and the ratio is taken as the imbalance ratio of the initial unbalanced data set. That is, the imbalance ratio (IR) is wherein T maj and T min respectively represent the number of majority class data elements and the number of minority class data elements. ​

[0054] S204, normalize the data in the initial unbalanced data set to obtain a normalized data set.

[0055] In an embodiment, the mean of the data in the initial unbalanced data set and the standard deviation of each data can be determined, and the data in the initial unbalanced data set is normalized according to the mean and the standard deviation to obtain a normalized data set. In actual operation, the expression is as follows: ; wherein, and are the i-th original data and normalized data, is the mean, is the standard deviation of the i-th data.

[0056] S205, performing a two-dimensional rotation operation on the normalized data set to determine an intermediate unbalanced data set and a data rotation interval.

[0057] In an embodiment, a preset rotation matrix can be extracted, and a two-dimensional rotation operation is performed on the normalized data set according to the preset rotation matrix to obtain an intermediate unbalanced data set and a data rotation interval.

[0058] wherein performing a two-dimensional rotation operation on the normalized data set to determine an intermediate unbalanced data set and a data rotation interval comprises:

[0059] composing adjacent data in the normalized data set into a continuous attribute pair, and performing a two-dimensional rotation operation on the continuous attribute pair in the angle range of the preset rotation matrix to obtain a rotated attribute pair;

[0060] determining the variance of the rotated attribute pair and the continuous attribute pair, and taking the angle corresponding to the variance that meets the preset threshold interval as the data rotation interval;

[0061] extracting any angle in the data rotation interval as a first angle, taking the rotated attribute pair corresponding to the first angle as a first rotated attribute pair, and performing a two-dimensional rotation operation on the first rotated attribute pair according to the first angle and the preset rotation matrix to obtain a second rotated attribute pair;

[0062] arranging the second rotated attribute pair according to the order of the continuous attribute pair to obtain an intermediate unbalanced data set.

[0063] The continuous attribute pair refers to a data pair composed of two adjacent data in the normalized data set. When the number of data in the initial unbalanced data set is odd, the last data and the previous data form a continuous attribute pair. The angle range refers to the angle range of two-dimensional rotation. For example, the angle range can be 0° to 360°. In actual operation, the continuous attribute pair can form a two-dimensional vector, which can be rotated and transformed by using a preset rotation matrix, so as to obtain a new two-dimensional vector. The new two-dimensional vector is the rotated attribute pair. The preset threshold interval refers to a critical value for determining whether the data rotation interval is met, which can be set according to business requirements. Since the continuous attribute pair includes two parameters, the preset threshold interval can be set for each parameter. The two preset threshold intervals can be the same or different. Generally, the preset threshold interval can be greater than 0.

[0064] In an embodiment, the adjacent data in the normalized data set can be formed into a continuous attribute pair, the angle range of the preset rotation matrix is selected, and the continuous attribute pair is two-dimensionally rotated according to the preset rotation matrix for each angle of the angle range to obtain a rotated attribute pair. In an embodiment, when the normalized data set can generate q continuous attribute pairs , and , and ; for the kth continuous attribute pair , the angle range from 0° to 360° is determined. The preset rotation matrix can adopt a clockwise two-dimensional rotation matrix: At this time, the rotated attribute pair is . The continuous attribute pair is traversed in the angle range. The variance of the rotated attribute pair and the continuous attribute pair is calculated, that is, the variance of each data in the continuous attribute pair after rotation and the original data is determined, , The angle corresponding to the variance meeting the preset threshold interval is taken as the data rotation interval. For example, when the preset threshold interval is greater than 0, the data rotation interval is , wherein ; any angle in the data rotation interval is extracted as a first angle, the rotated attribute pair corresponding to the first angle is taken as a first rotated attribute pair, and the first rotated attribute pair is two-dimensionally rotated according to the preset rotation matrix based on the first angle , that is, , wherein The second rotated attribute pair is arranged in the order of the continuous attribute pair to obtain an intermediate unbalanced data set.

[0065] S206, extracting the minority class data set and the majority class data set in the intermediate unbalanced data set.

[0066] S207, form a minority continuous attribute pair by adjacent data in the minority class data set, and perform a two-dimensional rotation operation on the minority continuous attribute pair according to a preset rotation matrix based on a data rotation interval to obtain a minority rotated attribute pair.

[0067] The rotation number of each minority continuous attribute pair is twice the unbalance ratio.

[0068] In an embodiment, the minority continuous attribute pair can be determined by the same way of forming a continuous attribute pair by adjacent data in the normalized data set. The angle range of the preset rotation matrix is set as the data rotation interval. A two-dimensional rotation operation is performed on each minority continuous attribute according to the data rotation interval to obtain a minority rotated attribute pair. The minority rotated attribute pair is taken as the minority continuous attribute pair until the rotation number reaches twice the unbalance ratio.

[0069] S208, arrange the minority rotated attribute pairs belonging to the same rotation number according to the order of the minority continuous attribute pairs to obtain a temporary minority class data set, and expand the minority class data set according to the temporary minority class data set to obtain an intermediate minority class data set.

[0070] In an embodiment, the minority rotated attribute pairs belonging to the same rotation number can be determined, and the order of the minority continuous attribute pairs can be determined. The minority rotated attribute pairs belonging to the same rotation number are arranged according to the order of the minority continuous attribute pairs as a temporary minority class data set. The temporary minority class data set is inserted according to the original data order in the minority class data set to realize the expansion of the minority class data set to obtain an intermediate minority class data set.

[0071] S209, determine the number of majority class data in the majority class data set as the clustering number.

[0072] S210, cluster the intermediate minority class data set according to the clustering number to obtain a target minority class data set.

[0073] In an embodiment, the data in the intermediate minority class data set with the clustering number can be determined as an initial centroid, and a clustering operation is performed based on the initial centroid until the target minority class data set is obtained.

[0074] The clustering operation of the intermediate minority class data set according to the clustering number to obtain the target minority class data set includes:

[0075] Randomly select the clustering number of minority class data in the intermediate minority class data set as an initial centroid, and determine the Euclidean distance between each minority class data in the target minority class data set and each initial centroid in turn. The initial centroid corresponding to the minimum Euclidean distance is taken as the temporary cluster of the minority class data.

[0076] determining the feature mean of all minority data in each temporary cluster, taking the feature mean as a temporary centroid, until the change of the temporary centroid is less than a preset value, taking the temporary centroid with the change less than the preset value as a target centroid;

[0077] taking the target centroid as a target minority data and storing the target minority data as a target minority data set.

[0078] In an embodiment, first, minority data in the intermediate minority data set is randomly selected as an initial centroid , , k = 1, 2, 3, …, Q, the number of initial centroids is the number of clusters, and the Euclidean distance between each minority data in the target minority data set and each initial centroid is determined in turn, that is , wherein x i is each minority data in the target minority data set, Ci and Ci+1 are the Euclidean distance squares of the centroids, and sample points xi satisfying the distance square to Ci less than or equal to the distance square to Ci+1 are divided into a set S i , that is, the temporary cluster of the minority data. The feature mean of all minority data in each temporary cluster is determined, that is, the feature mean ; x l is each minority data S i in the temporary cluster, and the feature mean is taken as a temporary centroid, until the change of the temporary centroid is less than a preset value, that is, the temporary centroid tends to be stable, and the temporary centroid with the change less than the preset value is taken as a target centroid.

[0079] S211, combining the majority data set and the target minority data set to generate a target unbalanced data set.

[0080] In the embodiment of the application, by acquiring unbalanced monitoring data, majority class data and minority class data in the unbalanced monitoring data are determined to establish an initial unbalanced data set, the unbalanced ratio of the initial unbalanced data set is determined, the data in the initial unbalanced data set is normalized to obtain a normalized data set, and the normalized data set is subjected to a two-dimensional rotation operation to determine an intermediate unbalanced data set and a data rotation interval, so as to facilitate subsequent data enhancement of the minority class data; by grouping adjacent data in the minority class data set into a minority continuous attribute pair, the minority continuous attribute pair is subjected to a two-dimensional rotation operation according to a preset rotation matrix based on the data rotation interval to obtain a minority rotated attribute pair, the minority rotated attribute pairs belonging to the same rotation number are arranged according to the order of the minority continuous attribute pairs to obtain a temporary minority class data set, and the minority class data set is expanded according to the temporary minority class data set to obtain an intermediate minority class data set, so as to ensure the data amount of the minority class data; by determining the number of the majority class data in the majority class data set, the number is taken as a clustering number, and the intermediate minority class data set is subjected to a clustering operation according to the clustering number to obtain a target minority class data set, so as to ensure the number and quality of the minority class data, unify the number of the minority class data and the majority class data, and facilitate subsequent analysis of the unbalanced monitoring data.

[0081] Embodiment three

[0082] In an embodiment, the embodiment is a further description of a data processing method based on the above-mentioned embodiments. The method comprises data normalization, two-dimensional rotation of normalized data, rotation and enhancement of minority class data, and clustering and expansion of the minority class data.

[0083] Step 1, data normalization.

[0084] Specifically, the initial unbalanced data set can be normalized by using the standard score (z-score), because it provides the best classification performance for the full feature set compared with other normalization methods, can effectively process outliers and generate attributes. The expression is as follows: ; wherein, and are the i-th original data and the normalized data, is the mean, is the standard deviation of the i-th data.

[0085] Step 2, two-dimensional rotation of normalized data. In this step, the proposal performs multiple rotation operations on the known minority class data to simulate more "virtual samples" with similar features but slight differences. These samples not only retain the key features of the original behavior pattern, but also enhance the data diversity, thereby effectively alleviating the recognition difficulty caused by the scarcity of samples. After two rotation enhancements, the number of minority class samples even exceeds that of the majority class, breaking the original data imbalance state.

[0086] In this step, a clockwise two-dimensional rotation matrix is ​​used: Rotation can preserve geometric properties such as length and inner product. Since the dimension of a two-dimensional rotation matrix is ​​2×2, rotation can be performed using two attributes of the dataset at a time to satisfy the column-row rules of matrix multiplication. Taking monitoring data as an example, suppose we want to analyze employee travel expense reimbursement behavior and select the following two key attributes as features: x1 is the number of days on a single business trip, and x2 is the amount of travel expense reimbursement. These two attributes can form a two-dimensional vector. We can obtain a new two-dimensional variable by rotating it using a two-dimensional rotation matrix. This achieves rotation enhancement of the original data points in two-dimensional space, which helps generate more representative "suspected outlier samples," thereby alleviating the class imbalance problem. The specific algorithm details are shown in Table 1.

[0087] Table 1. RotationOrgImb Algorithm Flowchart

[0088]

[0089] The first step of the RotationOrgImb algorithm (rotation optimization algorithm) is from... Generate q consecutive attribute pairs When n is even, q = n / 2, and attribute pairs can be generated directly. When n is odd, q = (n+1) / 2, and attribute pairs are generated from the last attribute and the previous attribute. When n is odd, after rotating the last attribute pair, reused attributes are discarded because they are rotated within the previous attribute pair. Each attribute pair is rotated through steps 2 to 8. In step 4, the k-th attribute pair is rotated... Get In step 5, the variances between the original attribute pairs and the rotated attribute pairs are calculated as v1 and v2. In step 6, the variances are derived for each attribute pair. ,in This is called a paired threshold (preset threshold range). These values ​​are typically randomly set, relatively small, and can be chosen differently for each attribute pair. For the k-th attribute pair, the variances v1 and v2 satisfy the inequality... The value is stored in the data rotation interval In. Stored in In The range is called the valid range. In step 8, from... Randomly select one The k-th attribute pair is rotated and placed in the order it appears in the original dataset, ultimately generating a complete two-dimensional rotated dataset (an unbalanced dataset in the middle). .

[0090] Step 3: Rotation and Augmentation of Minority Class Data

[0091] In the rotated training dataset (the imbalanced dataset) In this algorithm, a threshold-based two-dimensional rotation transformation is used to rotate the minority class data twice according to the imbalance ratio, ensuring that the minority class data is at least twice the size of the majority class data. Further rotation of the minority class data would certainly increase the number of minority class data, i.e., make the data more balanced, but the computational complexity of the algorithm would also increase. Therefore, we keep the number of rotations at [(2×IR)] (i.e., rounded up to 2×IR). Here, the imbalance ratio (IR) is... T maj and T min These represent the number of elements in the majority and minority classes, respectively.

[0092] The specific algorithm details are shown in Table 2:

[0093] Table 2 RotationAugmentation Algorithm Flowchart

[0094]

[0095] The RotationAugmentation algorithm takes minority class data from a rotated training dataset (an imbalanced dataset) as input. Where g is the number of minority class data points and n is the number of attributes. The output is the rotated and augmented minority class dataset. Where p = [(2×IR)] + g. Steps 2 to 6 of the algorithm... Similar to the generation method of Algorithm 1. In steps 7 and 8, The value from The data is generated by randomly selecting [(2×IR)] times, and is the k-th attribute pair. Therefore, the k-th attribute pair uses The rotations were performed [(2×IR)] times. In step 9, the final k-th rotation attribute pair was generated. Therefore, through steps 2 to 9, all minority class attribute pairs are rotated. In step 10, all rotated attribute pairs... Arrange them in their original order to generate [(2×IR)] rotated minority class datasets. In step 11, using right The expanded and generated data expanded minority class data set (intermediate minority class data set) .

[0096] Step 4, clustering the expanded minority class data

[0097] The last step uses the K-means clustering algorithm to extract multiple cluster centers from the enhanced minority class samples, and each cluster represents a specific type of abnormal behavior pattern (such as high-frequency small false reporting, cross-departmental collusion fraud, etc.). The number of clusters is consistent with the number of majority class samples, ensuring that the data amount of each category is balanced when training the model, which helps to improve the sensitivity of the classifier to abnormal behavior. CARBO algorithm is an oversampling method, in the first two steps, the rotated and expanded minority class instances are greater than the majority class instances, so it is necessary to generate several clusters from the expanded minority class data set In this proposal, the k-means clustering algorithm is used, and the number of clusters is set to the number of majority class data sets of the rotated training data set , where w is the number of majority class data. The centroid of each cluster is considered as a new minority class data, where the centroid is the average of several expanded minority class data. This processing method will be data balanced. The overall process of CARBO algorithm is shown in Table 3:

[0098] Table 3 CARBO algorithm flow table

[0099]

[0100] First, determine the intermediate unbalanced data set and the intermediate minority class data set, and determine the number of clusters. Randomly select a minority class data in the intermediate minority class data set as the initial centroid C k , , k = 1, 2, 3, …, Q, the number of initial centroids is the number of clusters, and the Euclidean distance between each minority class data in the target minority class data set and each initial centroid is determined, that is , where x i is each minority class data in the target minority class data set, Ci and Ci+1 are the Euclidean distance of the centroid, and the sample point xi that satisfies the distance square to Ci is less than or equal to the distance square to Ci+1 is divided into set S i , that is, the temporary cluster of the minority class data. Determine the feature mean of all minority class data in each temporary cluster, that is, the feature mean ; x l is each minority class data in the temporary cluster S i is the temporary cluster, and the feature mean is taken as the temporary centroid, until the change of the temporary centroid is less than the preset value, that is, the temporary centroid tends to be stable, and the temporary centroid with a change less than the preset value is taken as the target centroid.

[0101] In one embodiment, data is first randomly selected as initial centroids, i.e., center points. The distance of each data point from the center point is calculated, and each object is assigned to the nearest cluster center (i.e., step 7). Step 8 is to recalculate the cluster centers based on the data contained in the cluster. Example: Suppose there are 6 data points, A(1,2), B(1,4), C(1,0), D(4,2), E(4,4), F(4,0), and the goal is to divide these data points into two clusters. Step 1: Initialize centroids. Two data points are randomly selected as initial centroids: centroid 1(1,2) and centroid 2(4,4). Step 2: Assign data points to clusters. The Euclidean distance from each data point to the two centroids is calculated, and each data point is assigned to the nearest centroid. It is calculated that centroid 1 contains data (1,2), (1,4), (1,0), and (4,0). Centroid 2 contains data (4,2) and (4,4). Step 3: Recalculate the centroids. The new centroid coordinates are the average coordinates of all points within the cluster; therefore, centroid 1 is (1.75, 1.5), and centroid 2 is (4, 3). Step 4: Reassign data points to the new centroids, and so on, until the centroids no longer change.

[0102] In one embodiment, the time complexity of the CARBO algorithm is calculated below. Algorithm 3 describes the overall process of the CARBO algorithm. The input to this algorithm is an imbalanced dataset. Here, m and n represent the number of instances and the number of attributes, respectively. In step 1, Algorithm 1 is called. In Algorithm 1, the first line simply selects attribute pairs. When n is even, the computational complexity is n / 2; when n is odd, the computational complexity is (n+1) / 2. Therefore, the time cost of the first line is O(n). In the fourth line, the matrix multiplication is 2×2×m, meaning the time complexity of the fourth line is O(m). Each vector subtraction in the fifth line is... Therefore, the time complexity of line 5 is O(m). The calculation in line 6 is simple, with a time complexity of O(1). Line 7 is similar to line 4, so we take O(n). Loop 2 iterates at most n times, so the time complexity of loop 2 is O(n×(n+m+m+1+m)), which is O(n×m). Therefore, the time complexity of algorithm 1 is the sum of the time complexities of each step, O(n+n×m), which is O(n×m). Algorithm 2 is called in line 2 of algorithm 3. In steps 2 to 6 of algorithm 2, Similar to the generation method of Algorithm 1, line 8 is similar to line 7 of Algorithm 1, being rotated [(2×IR)] times (a constant). Lines 9 and 11 both have a time complexity of O(1). Therefore, the time complexity of Algorithm 2 is O(m+1+1+(m×n)). Lines 3 to 8 of Algorithm 3 are k-means clustering, with a time complexity of O(m... 2), where m is the number of instances. The time complexity of lines 9 and 10 is both O(l). Thus the total time complexity of Algorithm 3 is O((m x n) + (m x n) + m 2 +1+1), i.e., O(m 2 x n), achieving the boosting of the algorithm.

[0103] The embodiment aims at the binary class imbalance phenomenon of supervision data, and proposes a binary class imbalance learning method based on clustering and rotation: normalizing the imbalance data set; performing multiple two-dimensional rotation operations on the known minority class samples to simulate more "virtual samples" with similar characteristics but slight differences; for the rotated training data set, using a two-dimensional rotation transformation based on a threshold to rotate the minority class data twice according to the imbalance ratio, to ensure that the minority class data is at least twice the majority class data; using a clustering algorithm to extract multiple cluster centers from the enhanced minority class samples, each cluster represents a specific type of abnormal behavior pattern, and the representativeness of the minority class samples in the classification task is enhanced. At the same time, the rotation angle is dynamically adjusted to adapt to different data sets, and has good adaptability and generalization ability.

[0104] Embodiment Four

[0105] Figure 3 is a structural schematic diagram of a data processing device provided according to Embodiment Four of the present application. As shown in the figure, the device comprises a data acquisition module 31, an interval determination module 32, a data processing module 33 and a data set adjustment module 34. Figure 3

[0106] The data acquisition module 31 is configured to acquire imbalance supervision data, and establish an initial imbalance data set of the imbalance supervision data; wherein the difference between the majority class data and the minority class data in the imbalance supervision data exceeds a preset threshold.

[0107] The interval determination module 32 is configured to determine an imbalance ratio of the initial imbalance data set, perform a data preprocessing operation on the initial imbalance data set to obtain an intermediate imbalance data set, and determine a data rotation interval.

[0108] The data processing module 33 is configured to acquire a minority class data set and a majority class data set in the intermediate imbalance data set, and determine an intermediate minority class data set corresponding to the minority class data set based on the data rotation interval and the imbalance ratio.

[0109] The data set adjustment module 34 is configured to determine a target minority class data set according to the majority class data set and the intermediate minority class data set, and combine the majority class data set and the target minority class data set to generate a target imbalance data set.

[0110] ​The technical scheme of the embodiment of the application obtains unbalanced monitoring data through the data acquisition module, establishes an initial unbalanced data set of the unbalanced monitoring data, determines an unbalance ratio of the initial unbalanced data set through the interval determination module, performs a data preprocessing operation on the initial unbalanced data set to obtain an intermediate unbalanced data set and a data rotation interval, and realizes simulation of virtual samples; the data processing module is used to obtain a minority class data set and a majority class data set in the intermediate unbalanced data set, determine an intermediate minority class data set corresponding to the minority class data set based on the data rotation interval and the unbalance ratio, and realize expansion of the minority class data; the data set adjustment module is used to determine a target minority class data set according to the majority class data set and the intermediate minority class data set, combine the majority class data set and the target minority class data set to generate a target unbalanced data set, and enhance the representativeness of the minority class sample in the classification task, effectively improve the diversity of the minority class sample, and reduce the overfitting risk.

[0111] In an embodiment, the data acquisition module 31 comprises:

[0112] The data determination unit is configured to determine the majority class data and the minority class data in the unbalanced monitoring data.

[0113] The data set determination unit is configured to store the majority class data as an initial majority class data set, store the minority class data as an initial minority class data set, and combine the initial majority class data set and the initial minority class data set as an initial unbalanced data set.

[0114] In an embodiment, the interval determination module 32 comprises:

[0115] The ratio determination unit is configured to determine the number of the majority class data and the number of the minority class data in the initial unbalanced data set respectively, and take the ratio of the number of the majority class data to the number of the minority class data as the unbalance ratio of the initial unbalanced data set.

[0116] The normalization unit is configured to normalize the data in the initial unbalanced data set to obtain a normalized data set.

[0117] The interval determination unit is configured to perform a two-dimensional rotation operation on the normalized data set to determine an intermediate unbalanced data set and a data rotation interval.

[0118] In an embodiment, the interval determination unit is specifically configured to:

[0119] The interval determination unit is specifically configured to:

[0120] The interval determination unit is specifically configured to:

[0121] extract any angle in the data rotation interval as a first angle, take the rotation attribute pair corresponding to the first angle as a first rotation attribute pair, perform a two-dimensional rotation operation on the first rotation attribute pair according to a preset rotation matrix based on the first angle to obtain a second rotation attribute pair;

[0122] arrange the second rotation attribute pair according to the order of the continuous attribute pairs to obtain an intermediate unbalanced data set.

[0123] In an embodiment, the data processing module 33 comprises:

[0124] a data set extraction unit configured to extract a minority class data set and a majority class data set in the intermediate unbalanced data set;

[0125] an attribute pair determination unit configured to form adjacent data in the minority class data set into a minority continuous attribute pair, and perform a two-dimensional rotation operation on the minority continuous attribute pair according to a preset rotation matrix based on the data rotation interval to obtain a minority rotation attribute pair; wherein the rotation number of each minority continuous attribute pair is twice the unbalanced ratio;

[0126] a data expansion unit configured to arrange the minority rotation attribute pairs belonging to the same rotation number according to the order of the minority continuous attribute pairs to obtain a temporary minority class data set, and expand the minority class data set according to the temporary minority class data set to obtain an intermediate minority class data set.

[0127] In an embodiment, the data set adjustment module 34 comprises:

[0128] a cluster number determination unit configured to determine the number of majority class data in the majority class data set as a cluster number;

[0129] a data clustering unit configured to perform a clustering operation on the intermediate minority class data set according to the cluster number to obtain a target minority class data set.

[0130] In an embodiment, the data clustering unit is specifically configured to:

[0131] randomly select a cluster number of minority class data in the intermediate minority class data set as an initial centroid, and sequentially determine the Euclidean distance between each minority class data in the target minority class data set and each initial centroid, and take the initial centroid corresponding to the minimum value of the Euclidean distance as a temporary cluster of the minority class data;

[0132] determine the feature mean of all minority class data in each temporary cluster, and take the feature mean as a temporary centroid, until the change amount of the temporary centroid is less than a preset value, and take the temporary centroid with the change amount less than the preset value as a target centroid;

[0133] store the target centroid as the target minority class data set.

[0134] The data processing apparatus provided by the embodiments of the present application can execute the data processing method provided by any of the embodiments of the present application, and has the corresponding function modules and beneficial effects of executing the method.

[0135] Embodiment Five

[0136] Figure 4 is a structural schematic diagram of an electronic device implementing the data processing method of the embodiments of the present application. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices (such as headsets, glasses, watches, etc.), and other similar computing devices. The components shown here, their connections, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present application described and / or claimed herein.

[0137] As shown in Figure 4 The electronic device 10 includes at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11, wherein the memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0138] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunications networks.

[0139] The processor 11 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, and the like. The processor 11 performs various methods and processes described above, such as a data processing method.

[0140] In some embodiments, the data processing method can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded onto the RAM 13 and executed by the processor 11, one or more steps of the data processing method described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the data processing method by any other suitable means, such as by means of firmware.

[0141] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0142] Computer programs used to implement the methods of the application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program, when executed by the processor, implements the functions / acts specified in the flowcharts and / or block diagrams. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, and partially on a machine or entirely on a remote machine or server.

[0143] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0144] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0145] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0146] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.

[0147] It should be understood that the various forms of flow shown above can be reordered, additional steps added, or steps deleted. For example, the steps described in the present application can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions of the present application can be achieved, which are not limited herein.

[0148] The above detailed description does not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A data processing method, characterized by, The method comprises the following steps: obtaining unbalanced monitoring data, and establishing an initial unbalanced data set of the unbalanced monitoring data; wherein the difference between the majority class data and the minority class data in the unbalanced monitoring data exceeds a preset threshold value; determining an unbalance ratio of the initial unbalanced data set, performing a data preprocessing operation on the initial unbalanced data set to obtain an intermediate unbalanced data set and a data rotation interval; obtaining a minority class data set and a majority class data set in the intermediate unbalanced data set, and determining an intermediate minority class data set corresponding to the minority class data set based on the data rotation interval and the unbalance ratio; determining a target minority class data set according to the majority class data set and the intermediate minority class data set, and combining the majority class data set and the target minority class data set to generate a target unbalanced data set; wherein the step of obtaining the minority class data set and the majority class data set in the intermediate unbalanced data set, and determining the intermediate minority class data set corresponding to the minority class data set based on the data rotation interval and the unbalance ratio comprises the following steps: extracting the minority class data set and the majority class data set in the intermediate unbalanced data set; grouping adjacent data in the minority class data set into a minority continuous attribute pair, and performing a two-dimensional rotation operation on the minority continuous attribute pair according to a preset rotation matrix to obtain a minority rotated attribute pair; wherein the rotation number of each minority continuous attribute pair is twice the unbalance ratio; arranging the minority rotated attribute pairs belonging to the same rotation number according to the order of the minority continuous attribute pairs to obtain a temporary minority class data set, and expanding the minority class data set according to the temporary minority class data set to obtain an intermediate minority class data set.

2. The method of claim 1, wherein, The step of establishing the initial unbalanced data set of the unbalanced monitoring data comprises the following steps: determining the majority class data and the minority class data in the unbalanced monitoring data; storing the majority class data as an initial majority class data set, storing the minority class data as an initial minority class data set, and combining the initial majority class data set and the initial minority class data set as an initial unbalanced data set.

3. The method of claim 1, wherein, The step of determining the unbalance ratio of the initial unbalanced data set, performing a data preprocessing operation on the initial unbalanced data set to obtain an intermediate unbalanced data set and a data rotation interval comprises the following steps: respectively determining the number of majority class data and the number of minority class data in the initial unbalanced data set, and taking the ratio of the number of majority class data to the number of minority class data as the unbalance ratio of the initial unbalanced data set; normalizing the data in the initial unbalanced data set to obtain a normalized data set; performing a two-dimensional rotation operation on the normalized data set to determine an intermediate unbalanced data set and a data rotation interval.

4. The method of claim 3, wherein, The step of performing a two-dimensional rotation operation on the normalized data set to determine an intermediate unbalanced data set and a data rotation interval comprises the following steps: grouping adjacent data in the normalized data set into a continuous attribute pair, and performing a two-dimensional rotation operation on the continuous attribute pair in the angle range of a preset rotation matrix to obtain a rotated attribute pair; Determine the variance of the rotation attribute pair and the continuous attribute pair, and take the angle corresponding to the variance satisfying the preset threshold interval as the data rotation interval; Extract any angle in the data rotation interval as a first angle, take the rotation attribute pair corresponding to the first angle as a first rotation attribute pair, and perform a two-dimensional rotation operation on the first rotation attribute pair according to a preset rotation matrix based on the first angle to obtain a second rotation attribute pair; Arrange the second rotation attribute pair according to the order of the continuous attribute pair to obtain an intermediate unbalanced data set.

5. The method of claim 1, wherein, The target minority class data set is determined according to the majority class data set and the intermediate minority class data set, comprising: Determine the number of majority class data in the majority class data set, and take the number as the clustering number; The intermediate minority class data set is clustered according to the clustering number to obtain a target minority class data set.

6. The method of claim 5, wherein, The intermediate minority class data set is clustered according to the clustering number to obtain a target minority class data set, comprising: Randomly select a clustering number of minority class data in the intermediate minority class data set as an initial centroid, and sequentially determine the Euclidean distance between each minority class data in the target minority class data set and each initial centroid, and take the initial centroid corresponding to the minimum value of the Euclidean distance as the temporary cluster of the minority class data; Determine the feature mean of all minority class data in each temporary cluster, and take the feature mean as a temporary centroid, until the change amount of the temporary centroid is less than a preset value, and take the temporary centroid with a change amount less than the preset value as a target centroid; The target centroid is stored as a target minority class data set.

7. A data processing apparatus, characterized by, Comprising: A data acquisition module is configured to acquire unbalanced monitoring data and establish an initial unbalanced data set of the unbalanced monitoring data; wherein the difference between the majority class data and the minority class data in the unbalanced monitoring data exceeds a preset threshold value; An interval determination module is configured to determine the imbalance ratio of the initial unbalanced data set, perform data preprocessing on the initial unbalanced data set to obtain an intermediate unbalanced data set and a data rotation interval; A data processing module is configured to acquire a minority class data set and a majority class data set in the intermediate unbalanced data set, and determine an intermediate minority class data set corresponding to the minority class data set based on the data rotation interval and the imbalance ratio; A data set adjustment module is configured to determine a target minority class data set according to the majority class data set and the intermediate minority class data set, and combine the majority class data set and the target minority class data set to generate a target unbalanced data set; The data processing module comprises: A data set extraction unit is configured to extract a minority class data set and a majority class data set in the intermediate unbalanced data set; An attribute pair determination unit is configured to form a minority continuous attribute pair by combining adjacent data in the minority class data set, and perform a two-dimensional rotation operation on the minority continuous attribute pair according to a preset rotation matrix based on the data rotation interval to obtain a minority rotation attribute pair; wherein the rotation number of each minority continuous attribute pair is twice the imbalance ratio. The data expansion unit is configured to arrange the few rotation attribute pairs belonging to the same rotation number in the order of the few continuous attribute pairs to obtain a temporary few-class data set, and expand the few-class data set according to the temporary few-class data set to obtain an intermediate few-class data set.

8. An electronic device, comprising: The electronic device comprises: at least one processor; and a memory connected to the at least one processor in communication; wherein The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the data processing method of any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for enabling the processor to implement the data processing method of any one of claims 1-6 when executed.

Citation Information

Patent Citations

  • A multi-classification oriented unbalanced data preprocessing method and device and an apparatus

    CN109033148A

  • Unbalanced data set preprocessing method based on neighborhood information

    CN112835955A