Data Processing Method, Device, Electronic Device and Storage Medium
The data processing method addresses prolonged training times in noisy data sets by dimensionality reduction and noise classification, enhancing training efficiency through noise data removal.
Patent Information
- Application Number
- CN202011632027.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-31
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2040-12-31
AI Technical Summary
In the classification data set, the presence of a large amount of noise data causes the training time to be extended, and the data that determines the classification plane exists only in the data around the classification boundary, which is difficult to efficiently process in the prior art.
By processing the data set by dimensionality reduction, a search index is established, a noise ratio is calculated, and a data point is judged based on the noise ratio, a noise data point is marked and cleared, and the training data set is simplified.
Without affecting the classification results, simplify the training data set, shorten the training time, and improve the data processing speed.
Smart Images

Figure CN114764872B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a data processing method, apparatus, electronic device, and storage medium. Background Art
[0002] When using classification algorithms, the most time-consuming part is often the training time, which is related to the size of the dataset and the design of the algorithm. In a classification dataset, the tolerable noise calculation often needs to be adjusted repeatedly to find the optimal classification plane for different labeled classes. Conducting classification training in a dataset with a large amount of noisy data results in a significant extension of the training time. However, the data determining the classification plane only exists in the data around the classification boundary. Summary of the Invention
[0003] In view of the above problems, this application proposes a data processing method, apparatus, electronic device, and storage medium to shorten the time for training data.
[0004] The first aspect of this application provides a data processing method, which includes: performing dimensionality reduction on a dataset and obtaining the dimension of the dataset after dimensionality reduction; establishing a search index based on the dataset after dimensionality reduction; setting the denominator d for calculating the noise ratio based on the dimension of the dataset; selecting an unlabeled data point P from the dataset and searching for the neighbors of the data point P according to the established index to obtain a neighbor dataset; counting the number of data categories in the neighbor dataset that are different from the category of the data point P, and setting the counted number as the numerator c for calculating the noise ratio; further calculating the noise ratio A = c / d based on the denominator d of the noise ratio and the numerator c of the noise ratio; determining whether the calculated noise ratio is less than or equal to a preset noise ratio; and when the calculated noise ratio is greater than the preset noise ratio, marking the data point P as noise data.
[0005] According to some embodiments of this application, the method further includes: when the calculated noise ratio is less than or equal to the preset noise ratio, marking the data point P as retained data.
[0006] According to some embodiments of this application, the method further includes: determining whether all data points in the dataset have been marked; and when all data points in the dataset have been marked, clearing all data points marked as noise data in the dataset.
[0007] According to some embodiments of the present application, searching for neighbors of the data point P according to the established index, the obtained neighbor data set includes: centering on the data point P, finding the first data point with the greatest similarity to the data point P from each dimension in the data set, obtaining a plurality of first data points; using the plurality of first data points as neighbors of the data point P; determining whether the number of neighbors of the data point P reaches the denominator d of the noise ratio; when the number of neighbors of the data point P reaches the denominator d of the noise ratio, setting the plurality of first data points as the neighbor data set.
[0008] According to some embodiments of the present application, searching for neighbors of the data point P according to the established index, the obtained neighbor data set further includes: when the number of neighbors of the data point P does not reach the denominator d of the noise ratio, continuing to find a plurality of second data points with the second greatest similarity to the data point P from other dimensions of the data set until the number of neighbors of the data point P reaches the denominator d of the noise ratio.
[0009] According to some embodiments of the present application, setting the denominator d of the noise ratio to twice the dimension dim of the data set, or setting the denominator d of the noise ratio to
[0010] According to some embodiments of the present application, centering on the data point P, finding the first data point with the greatest similarity to the data point P from each dimension in the data set includes: from the first dimension of the data set, finding the data point corresponding to the data with the greatest similarity to the first data in the first dimension of the data point P as the first data point, where data with the greatest similarity to the first data is found in the positive direction of the first data and data with the greatest similarity to the first data is found in the negative direction of the first data; continuing to find the data point corresponding to the data with the greatest similarity to the second data in the second dimension of the data point P from the second dimension of the data set as the first data point, where data with the greatest similarity to the second data is found in the positive direction of the second data and data with the greatest similarity to the second data is found in the negative direction of the second data; repeating the operation of finding the data point corresponding to the data with the greatest similarity to the first data in the first dimension of the data point P from the first dimension of the data set as the first data point and continuing to find the data point corresponding to the data with the greatest similarity to the second data in the second dimension of the data point P from the second dimension of the data set as the first data point until all dimensions in the data set are searched, and finding the first data point with the greatest similarity to the data point P, obtaining a plurality of first data points.
[0011] The second aspect of the present application provides a data processing device, which includes: a processing module for performing dimensionality reduction processing on a data set and obtaining the dimension of the data set after dimensionality reduction processing; a building module for building a search index based on the data set after dimensionality reduction processing; a setting module for setting the denominator d for calculating the noise ratio based on the dimension of the data set; the processing module is further configured to select unlabeled data points P from the data set and search for neighbors of the data point P according to the established index to obtain a neighbor data set; the setting module is further configured to count the number of data categories in the neighbor data set that are different from the category of the data point P and set the counted number as the numerator c for calculating the noise ratio; the processing module is further configured to calculate the noise ratio A = c / d based on the denominator d of the noise ratio and the numerator c of the noise ratio; a judgment module for judging whether the calculated noise ratio is less than or equal to a preset noise ratio; and a marking module for marking the data point P as noise data when the calculated noise ratio is greater than the preset noise ratio.
[0012] The third aspect of the present application provides an electronic device, which includes: a processor; and a memory in which a plurality of program modules are stored, and the plurality of program modules are loaded and executed by the processor to perform the data processing method as described above.
[0013] The fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the data processing method as described above.
[0014] The data processing method, device, electronic device and medium provided by the present application. The data processing method provided by the present application calculates the noise ratio by searching the number of neighbors of the data point, and judges whether the data point is noise according to the comparison result of the calculated noise ratio and the preset noise ratio. During classification training, the number of the training data set can be simplified without affecting the classification result, so as to accelerate the overall training time. Description of the Drawings
[0015] Figure 1 is a schematic flowchart of the data processing method provided by an embodiment of the present application.
[0016] Figure 2 is a schematic diagram of the data set provided by an embodiment of the present application.
[0017] Figure 3 is a schematic diagram of the data points to be deleted in the data set provided by an embodiment of the present application.
[0018] Figure 4 is a functional module diagram of the data processing device provided by an embodiment of the present application.
[0019] Figure 5It is a schematic diagram of the electronic device architecture provided by an embodiment of the present application. Detailed implementation manners
[0020] In order to more clearly understand the purposes, features, and advantages of the present application, the present application will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments may be combined with each other. In the following description, many specific details are set forth in order to fully understand the present application. The described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.
[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used in the specification of the present application herein are only for the purpose of describing specific embodiments, and are not intended to limit the present application.
[0022] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of the data processing method provided by an embodiment of the present application. According to different requirements, the order of the steps in the flowchart may be changed, and some steps may be omitted. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown. The data processing method of the embodiments of the present application is applied to an electronic device. For an electronic device that needs to perform data processing, the data processing function provided by the method of the present application can be directly integrated on the electronic device, or a client for implementing the data processing method of the present application can be installed. Again, the data processing method provided by the present application can also run on devices such as a server in the form of a Software Development Kit (SDK), and provide an interface for the data processing function in the form of the SDK. An electronic device or other devices can implement the data processing function through the provided interface. The data processing method includes the following steps.
[0023] Step S1: Perform dimensionality reduction processing on the data set, and obtain the dimension of the data set after dimensionality reduction processing.
[0024] In the present embodiment, in order to solve the problem that the large data dimension in the data set leads to a large amount of calculation and a long training time, the data set is first subjected to dimensionality reduction processing. Specifically, the dimensionality reduction processing includes: selecting preset-dimensional data from the data set based on a feature selection method, where the preset-dimensional data is important data characterizing user information.
[0025] In this embodiment, in order to shorten the training time and avoid excessive computational complexity caused by the high dimensionality of the data, several dimensions with important attributes are selected from the data set through a feature selection method, thereby simplifying the model, reducing overfitting, and improving generality. For example, when the data recorded in the data set is patient information, the patient information includes multi-dimensional information such as height, weight, address, phone number, heart rate, body temperature, etc. In order to analyze the physical condition of the patient, height, weight, heart rate, and body temperature with important attributes can be selected from the patient information.
[0026] In this embodiment, the feature selection method includes filter methods, wrapper methods, and embedded methods. The filter method is used to remove features with small value changes. Among them, the filter method further includes variance selection method, correlation coefficient method, chi-squared test, and mutual information method. The wrapper method is used to determine whether to add features through an objective function. Among them, the wrapper method includes a recursive feature elimination algorithm. The embedded method is automatically selected by the learner.
[0027] Step S2: Establish a search index based on the data set after dimensionality reduction.
[0028] In this embodiment, in order to speed up the search when searching for neighbors, a search index can be established for the data in the data set.
[0029] In this embodiment, the search index for the data set after dimensionality reduction can be established by using the K-D tree algorithm and the ball tree algorithm. The K-D tree algorithm and the ball tree algorithm are both existing technologies and will not be elaborated here.
[0030] Step S3: Set the denominator d for calculating the noise ratio based on the dimension of the data set.
[0031] In this embodiment, the denominator of the noise ratio is the number of neighbors to be selected. Set the denominator d of the noise ratio to twice the dimension dim of the data set, or set the denominator d of the noise ratio to
[0032] Step S4: Select an unlabeled data point P from the data set, and search for the neighbors of the data point P according to the established index to obtain a neighbor data set.
[0033] In this embodiment, the method for searching for neighbors of the data point P according to the established index to obtain a neighbor data set includes:
[0034] (1) Centering on the data point P, find the first data point with the greatest similarity to the data point P in each dimension of the data set, and obtain multiple first data point categories.
[0035] In one embodiment, it is assumed that the data set includes a first dimension, a second dimension, a third dimension... an Nth dimension. Similarly, each data point in the data set also includes a first dimension, a second dimension, a third dimension... an Nth dimension. First, in the first dimension of the data set, find multiple first data points with the greatest similarity to the data point P. Specifically, in the first dimension of the data set, find the data point corresponding to the data with the greatest similarity to the first data of the first dimension of the data point P as the first data point, where the data with the greatest similarity to the first data is searched in the positive direction of the first data, and the data with the greatest similarity to the first data is searched in the negative direction of the first data. Then continue to find in the second dimension of the data set the first data point corresponding to the data with the greatest similarity to the second data of the second dimension of the data point P, where the data with the greatest similarity to the second data is searched in the positive direction of the second data, and the data with the greatest similarity to the second data is searched in the negative direction of the second data; until all dimensions in the data set are searched, find the first data point with the greatest similarity to the data point P to obtain multiple first data points.
[0036] For example, the data set includes data points P, P1, P2... PM. It is assumed that the dimension of the data set is 4, and the data point P = {P 00 ,P 01 ,P 02 ,P 03}, the data point P1 = {P 10 ,P 11 ,P 12 ,P 13}, the data point P2 = {P 20 ,P 21 ,P 22 ,P 23}, and the data point PM = {P m0 ,P m1 ,P m2 ,P m3}. Then the first data of the first dimension of the data point P is P 00 , and searching in the positive direction of the first data P 00 for the data with the greatest similarity to the first data, assume the found data is P 20 ; searching in the negative direction of the first data P00 Search for the data with the highest similarity to the first data in the negative direction, and let the found data be P. 10 Then, use the data points P1 and P2 as the first data points.
[0037] Continue to search in the second dimension of the data set for the first data point corresponding to the data with the highest similarity to the second data in the second dimension of the data point P. The second data in the second dimension of the data point P is P. 01 Using the second data P 01 Search for the data with the highest similarity to the second data in the positive direction, and let the found data be P. m1 Using the second data P 01 Search for the data with the highest similarity to the second data in the negative direction, and let the found data be P. 11 Then, use the data points P1 and PM as the first data points.
[0038] Continue to search in the third dimension of the data set for the first data point corresponding to the data with the highest similarity to the third data in the third dimension of the data point P. The third data in the third dimension of the data point P is P. 02 Using the third data P 012 Search for the data with the highest similarity to the third data in the positive direction, and let the found data be P. m2 Using the third data P 02 Search for the data with the highest similarity to the third data in the negative direction, and let the found data be P. 12 Then, use the data points P1 and PM as the first data points. Repeat the process of searching in the first dimension of the data set for the data point corresponding to the data with the highest similarity to the first data in the first dimension of the data point P as the first data point, and continue to search in the second dimension of the data set for the data point corresponding to the data with the highest similarity to the second data in the second dimension of the data point P as the first data point until all dimensions in the data set are searched, and find the first data point with the highest similarity to the data point P to obtain multiple first data points.
[0039] In this embodiment, the Euclidean distance between the data point P and the first data point can be calculated to confirm whether the first data point is the point with the highest similarity to the data point P. When the Euclidean distance between the data point P and the first data point is smaller, it is confirmed that the first data point has a greater similarity to the data point P.
[0040] It should be noted that, in addition to the Euclidean distance, the parameters for confirming the similarity between data points in the dataset may also be parameters such as Hamming distance and cosine similarity. This application does not make any restrictions in this regard.
[0041] (2) Take the multiple first data points as the neighbors of the data point P;
[0042] (3) Determine whether the number of neighbors of the data point P reaches the denominator d of the noise ratio;
[0043] (4) When the number of neighbors of the data point P reaches the denominator d of the noise ratio, set the multiple first data points as the neighbor dataset;
[0044] (5) When the number of neighbors of the data point P does not reach the denominator d of the noise ratio, continue to find multiple second data points with the second highest similarity to the data point P from other dimensions of the dataset until the number of neighbors of the data point P reaches the denominator d of the noise ratio. It should be noted that when the number of data in the neighbor dataset composed of multiple first data points and multiple second data points still does not reach the denominator d of the noise ratio, continue to find multiple third data points with the third highest similarity to the data point P from other dimensions of the dataset. And so on until the number of neighbors of the data point P reaches the denominator d of the noise ratio.
[0045] In one embodiment, when the data points satisfying the number d of the denominator of the noise ratio cannot be found from the data in the current dimension of the data point P, the neighbors of the data point P can be continuously found from the data in other dimensions.
[0046] (6) Set the multiple first data points and the multiple second data points as the neighbor dataset.
[0047] For example, as Figure 2 shown, the dimension of the dataset category is 2. Among them, the dataset category includes data point a, data point b, data point c, data point d, data point e, data point f, data point g, and data point h, which are represented by circles in Figure 2 ; the dataset category also includes data point i, data point j, data point k, data point l, data point m, data point n, data point o, data point p, data point q, and data point r, which are represented by triangles in Figure 2 . Set the calculation of the Euclidean distance between data points as the basis for judging the similarity between data points. Set the allowable noise ratio a to 0.75 and the number of neighbors d to 2.
[0048] First, select the unlabeled data point a and search for neighbors along the x-axis and y-axis relative to data point a respectively. Start from the positive direction of the x-axis relative to data point a. When data point d is found, since there is no closest neighbor in this direction at this time, data point d is the closest neighbor visited so far. The Euclidean distance between data point a and data point d can be calculated as 11.2, as shown in Table 1.
[0049] Table 1
[0050]
[0051] Continue to search for neighbors in the dataset. When data point b is found, calculate the Euclidean distance between data point a and data point b as 7.07. Since the current closest neighbor in this direction is data point d, but after comparing the Euclidean distance between data point a and data point b and the Euclidean distance between data point a and data point d, it is confirmed that data point b is closer to data point a. Therefore, replace data point d in Table 1 with data point b to obtain Table 2.
[0052] Table 2
[0053]
[0054] Continue to search for neighbors in the dataset. Since no neighbor is found in the negative direction of the x-axis for data point a, data point c is found in the positive direction of the y-axis and data point e is found in the negative direction of the y-axis, obtaining Table 3.
[0055] Table 3
[0056]
[0057] Since no neighbor is found in the negative direction of the x-axis for data point a and the number of neighbors of data point a has not reached 4. Therefore, select data point f, which is the closest to data point a in the negative direction of the x-axis of data point a, as the neighbor of data point a, obtaining Table 4. The neighbor dataset of data point a includes data points b, f, c, and e as shown in Table 4.
[0058] Table 4
[0059]
[0060] Then, select the unlabeled data point d and search for neighbors along the x-axis and y-axis relative to data point d respectively. The obtained neighbor dataset includes data points e, k, c, and n, as shown in Table 5.
[0061] Table 5
[0062]
[0063] Then select the unlabeled data point r, and search for neighbors on the x-axis and y-axis relative to the data point r respectively. The first data points with the same category as the data point r obtained include the data point i and the data point l, as shown in Table 6. Since no neighbors are found in the negative direction of the X-axis and the positive direction of the Y-axis of the data point r, the number of neighbors obtained is less than 4. It is necessary to continue searching for data points with different categories from the data point r in the dataset.
[0064] Table 6
[0065]
[0066] Search for the data points that are the same as the data point r in category but the second closest to the data point r in similarity on the negative direction of the X-axis and the positive direction of the Y-axis relative to the data point r, which are the data point i and the data point m respectively, and obtain Table 7.
[0067] Table 7
[0068]
[0069] Then select the unlabeled data point s, and search for neighbors on the x-axis and y-axis relative to the data point s respectively. The obtained neighbor dataset includes the data point g, the data point f, the data point e, and the data point h, as shown in Table 8.
[0070] Table 8
[0071]
[0072] Step S5: Count the number of data categories in the neighbor dataset that are different from the category of the data point P, and set the counted number as the numerator c for calculating the noise ratio.
[0073] For example, as shown in Table 4, the number of data categories in the neighbor dataset corresponding to the data point a that are different from the category of the data point a is 0; as shown in Table 5, the number of data categories in the neighbor dataset corresponding to the data point d that are different from the category of the data point d is 2, such as the data point k and the data point n; as shown in Table 7, the number of data categories in the neighbor dataset corresponding to the data point r that are different from the category of the data point r is 0; as shown in Table 8, the number of data categories in the neighbor dataset corresponding to the data point s that are different from the category of the data point s is 4, such as the data point g, the data point f, the data point e, and the data point h.
[0074] Step S6: Calculate the noise ratio A = c / d based on the denominator d of the noise ratio and the numerator c of the noise ratio.
[0075] For example, the noise ratio corresponding to the data point a is calculated to be 0; the noise ratio corresponding to the data point d is calculated to be 0.5; the noise ratio corresponding to the data point r is calculated to be 0; the noise ratio corresponding to the data point s is calculated to be 1.
[0076] Step S7: Determine whether the calculated noise ratio is less than or equal to the preset noise ratio or equal to zero. When the calculated noise ratio is less than or equal to the preset noise ratio or not equal to zero, the process proceeds to step S8; when the calculated noise ratio is greater than the preset noise ratio or equal to zero, the process proceeds to step S9.
[0077] For example, the preset noise ratio is set to 0.75.
[0078] Step S8: Mark the data point P as reserved data, and then the process proceeds to step S10.
[0079] For example, mark the data point d as reserved data.
[0080] Step S9: Mark the data point P as noise data, and then the process proceeds to step S10.
[0081] For example, mark the data points a, r, and s as noise data.
[0082] Step S10: Determine whether all data points in the dataset have been marked. When there are still unmarked data points in the dataset, the process returns to step S4; when all data points in the dataset have been marked, the process proceeds to step S11.
[0083] Step S11: Remove all data points marked as noise data from the dataset.
[0084] After traversing all data points in the dataset using the data processing method of the present application, the data points marked as noise data can be obtained, including the data points a, r, s, l, m, and q, as Figure 3 the gray-marked data points shown in.
[0085] Figure 1 The data processing method of the present application is introduced in detail. Through this method, the data processing speed can be improved. The following combines Figure 4 and Figure 5 to introduce the functional modules and hardware device architectures for implementing the data processing device. It should be understood that the embodiments are for illustrative purposes only and are not limited by this structure in the scope of the patent application.
[0086] Figure 4 It is a functional module diagram of a data processing device provided by an embodiment of the present application.
[0087] In some embodiments, the data processing device 20 may include a plurality of functional modules composed of program code segments. The program code of each program segment in the data processing device 20 may be stored in the memory of the electronic device and executed by at least one processor in the electronic device to achieve the function of quickly processing data.
[0088] Reference Figure 4 , in this embodiment, according to the functions it executes, the data processing device 20 may be divided into a plurality of functional modules, and each of the functional modules is used to execute Figure 1 the respective steps in the corresponding embodiment to achieve the function of data processing. In this embodiment, the functional modules of the data processing device 20 include: a processing module 201, a building module 202, a setting module 203, a judging module 204, and a marking module 205.
[0089] The processing module 201 is used to perform dimensionality reduction processing on the data set and obtain the dimension of the data set after dimensionality reduction processing; the building module 202 builds a search index based on the data set after dimensionality reduction processing by the processing module 201; the setting module 203 sets the denominator d for calculating the noise ratio based on the dimension of the data set obtained by the processing module 201; the processing module 201 is further used to select unlabeled data points P from the data set and search for the neighbors of the data point P according to the index built by the building module 202 to obtain a neighbor data set; the setting module 203 is further used to count the number of data categories in the neighbor data set obtained by the processing module 201 that are different from the category of the data point P and set the counted number as the numerator c for calculating the noise ratio; the processing module 201 is further used to calculate the noise ratio A = c / d based on the denominator d of the noise ratio and the numerator c of the noise ratio set by the setting module 203; the judging module 204 is used to judge whether the noise ratio calculated by the processing module 201 is less than or equal to the default noise ratio; and the marking module 205 is used to mark the data point P as noise data when the noise ratio calculated by the processing module 201 is greater than the default noise ratio.
[0090] Figure 5 It is a schematic diagram of the functional modules of an electronic device provided in an embodiment of the present application. The electronic device 1 includes a memory 11, a processor 12, and a computer program 13 stored in the memory 11 and executable on the processor 12, such as a program for data processing.
[0091] In this embodiment, the electronic device 1 may be, but is not limited to, a smart phone, a tablet computer, a computer device, a server, etc.
[0092] When the processor 12 executes the computer program 13, the steps of the data processing method in the method embodiment are implemented. Alternatively, when the processor 12 executes the computer program 13, the functions of each module / unit in the system embodiment are implemented.
[0093] Exemplarily, the computer program 13 may be divided into one or more modules / units, and the one or more modules / units are stored in the memory 11 and executed by the processor 12 to complete this application. The one or more modules / units may be a series of computer-readable instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 13 in the electronic device 1. For example, the computer program 13 may be divided into Figure 4 modules 201-205 in
[0094] Those skilled in the art can understand that the schematic Figure 5 is only an example of the electronic device 1 and does not constitute a limitation on the electronic device 1. The electronic device 1 may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the electronic device 1 may further include input / output devices, etc.
[0095] The so-called processor 12 may be a central processing unit (CPU), and may also include other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor 12 is the control center of the electronic device 1 and connects various parts of the entire electronic device 1 through various interfaces and lines.
[0096] The memory 11 can be used to store the computer program 13 and / or modules / units. By running or executing the computer program and / or modules / units stored in the memory 11, and invoking the data stored in the memory 11, the processor 12 realizes various functions of the electronic device 1. The memory 11 can include external storage media and can also include memory. In addition, the memory 11 can include volatile / non-volatile memory, such as a hard disk, memory, plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, FlashCard, at least one magnetic disk storage device, flash device, or other storage devices.
[0097] If the modules / units integrated in the electronic device 1 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by instructing related hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be realized. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, Read-Only Memory (ROM), random access memory, etc.
[0098] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit them. Although the present application has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present application can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present application.
Claims
1. A data processing method, characterized in that, The method includes: Reducing the dimension of the data set and obtaining the dimension of the data set after dimension reduction; Establishing a search index based on the data set after dimension reduction; Setting the denominator d for calculating the noise ratio based on the dimension of the data set; Selecting unlabeled data points P from the data set and searching for the neighbors of the data points P according to the established index to obtain a neighbor data set; Counting the number of data categories in the neighbor data set that are different from the category of the data point P, and setting the counted number as the numerator c for calculating the noise ratio; Calculating the noise ratio A = c / d based on the denominator d of the noise ratio and the numerator c of the noise ratio; Judging whether the calculated noise ratio is less than or equal to a preset noise ratio; and When the calculated noise ratio is greater than the preset noise ratio, marking the data point P as noise data; Judging whether all the data points in the data set have been marked; and When all the data points in the data set have been marked, clearing all the data points marked as noise data in the data set.
2. The data processing method according to claim 1, wherein, The method further includes: When the calculated noise ratio is less than or equal to the preset noise ratio, marking the data point P as retained data.
3. The data processing method according to claim 1, wherein The searching for the neighbors of the data point P according to the established index to obtain a neighbor data set includes: Centering on the data point P, finding the first data points with the greatest similarity to the data point P in each dimension of the data set to obtain a plurality of first data points; Taking the plurality of first data points as the neighbors of the data point P; Judging whether the number of neighbors of the data point P reaches the denominator d of the noise ratio; When the number of neighbors of the data point P reaches the denominator d of the noise ratio, setting the plurality of first data points as the neighbor data set.
4. The data processing method according to claim 3, wherein The searching for the neighbors of the data point P according to the established index to obtain a neighbor data set further includes: When the number of neighbors of the data point P does not reach the denominator d of the noise ratio, continuing to find a plurality of second data points with the second greatest similarity to the data point P from other dimensions of the data set until the number of neighbors of the data point P reaches the denominator d of the noise ratio.
5. The data processing method according to claim 1, wherein Set the denominator d of the noise ratio to twice the dimension dim of the dataset, or set the denominator d of the noise ratio to .
6. The data processing method according to claim 3, wherein Wherein, The centering on the data point P and finding the first data points with the greatest similarity to the data point P in each dimension of the data set includes: From the first dimension of the data set, finding the data point corresponding to the data with the greatest similarity to the first data in the first dimension of the data point P as the first data point, wherein, searching for the data with the greatest similarity to the first data in the positive direction of the first data and searching for the data with the greatest similarity to the first data in the negative direction of the first data; Continuing from the second dimension of the data set, finding the data point corresponding to the data with the greatest similarity to the second data in the second dimension of the data point P as the first data point, wherein, searching for the data with the greatest similarity to the second data in the positive direction of the second data and searching for the data with the greatest similarity to the second data in the negative direction of the second data; Repeatedly execute the steps of finding, from the first dimension of the dataset, the data point corresponding to the data with the largest similarity to the first dimension of the data point P as the first data point, and continuing to find, from the second dimension of the dataset, the data point corresponding to the data with the largest similarity to the second dimension of the data point P as the first data point, until all dimensions in the dataset are processed, to find the first data point with the largest similarity to the data point P, and obtain multiple first data points.
7. A data processing device, characterized in that, The device includes: A processing module, configured to perform dimensionality reduction processing on the dataset and obtain the dimensions of the dataset after dimensionality reduction processing; A building module, configured to build a search index based on the dataset after dimensionality reduction processing; A setting module, configured to set the denominator d for calculating the noise ratio based on the dimensions of the dataset; The processing module is further configured to select an unlabeled data point P from the dataset and search for the neighbors of the data point P according to the established index to obtain a neighbor dataset; The setting module is further configured to count the number of data categories in the neighbor dataset that are different from the category of the data point P and set the counted number as the numerator c for calculating the noise ratio; The processing module is further configured to calculate the noise ratio A = c / d based on the denominator d of the noise ratio and the numerator c of the noise ratio; A judgment module, configured to judge whether the calculated noise ratio is less than or equal to a preset noise ratio; and A marking module, configured to mark the data point P as noise data when the calculated noise ratio is greater than the preset noise ratio; The processing module is further configured to judge whether all data points in the dataset have been marked; and when all data points in the dataset have been marked, clear all data points marked as noise data in the dataset.
8. An electronic device, characterized in that, The electronic device includes: A processor; and A memory, in which multiple program modules are stored, and the multiple program modules are loaded and executed by the processor to perform the data processing method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by the processor, implements the data processing method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Data processing method and apparatus, computer readable storage medium, and electronic apparatus
CN109408583A
Text classification noise monitoring method and device, equipment and computer readable medium
CN110717033A