Data oversampling method and system, electronic equipment and storage medium
By modifying the Euro-based distance calculation formula, considering the Spearman correlation coefficient between each eigenvalue of the sample point and the result, the unbalanced data set is oversampled, which solves the problem of the same sample weight in the existing technology, improves the training effect of the machine learning model and reduces the generation of noise data.
Patent Information
- Application Number
- CN202510083484.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-06-10
AI Technical Summary
When the prior art oversampling of unbalanced training sample sets, each sample is assigned the same weight, which affects the training effect of the machine learning model.
Based on the Spearman correlation coefficient between each eigenvalue of the sample point in the uneven dataset and the result, the Euro-like distance calculation formula is modified to oversample the preset minority sample points in the uneven dataset.
By assigning different weights to different features, oversampling of unbalanced data sets improves the training effect of machine learning models and reduces the generation of noise data.
Smart Images

Figure CN120124772A_ABST
Abstract
Description
Background Art
[0002] When training a machine learning model, if the positive and negative samples in the training sample set are unbalanced, it will affect the training effect of the machine learning model. Therefore, oversampling is often performed on the unbalanced training sample set. Currently, oversampling is often based on the Euclidean distance, but the disadvantage is that the weights assigned to each sample in the training sample set are the same, which will affect the training effect of the machine learning model. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a data oversampling method, system, electronic device and storage medium in view of the deficiencies of the prior art, specifically as follows:
[0004] 1) In the first aspect, the present invention provides a data oversampling method, and the specific technical solution is as follows:
[0005] Based on the Spearman correlation coefficient between each eigenvalue of the sample points in the unbalanced data set and the result, modify the Euclidean distance calculation formula to determine the modified Euclidean distance calculation formula;
[0006] Use the distance calculated by the modified Euclidean distance calculation formula to perform oversampling on the preset minority class sample points in the unbalanced data set.
[0007] The beneficial effects of a data oversampling method provided by the present invention are as follows:
[0008] Consider the Spearman correlation coefficient between each feature and the result, so as to assign different weights to different features. Using the present invention to perform oversampling on the unbalanced data set can improve the training effect of the machine learning model.
[0009] On the basis of the above solution, a data oversampling method of the present invention can also be improved as follows.
[0010] Further, using the distance calculated by the modified Euclidean distance calculation formula to perform oversampling on the preset minority class sample points in the unbalanced data set includes:
[0011] Use the modified Euclidean distance calculation formula to calculate the distance between any preset minority class sample point and each remaining preset minority class sample point, and select M preset minority class sample points from the top K preset minority class sample points with the smallest distance as the M target preset minority class sample points corresponding to the preset minority class sample point until the M target preset minority class sample points corresponding to each preset minority class sample point are obtained, where K and M are both positive integers;
[0012] Generate a new minority-class sample point by using any preset minority-class sample point and any target preset minority-class sample point corresponding to the preset minority-class sample point. Calculate the distances between the new minority-class sample point and each sample point in the imbalanced dataset by using the modified Euclidean distance calculation formula, and determine whether to save the new minority-class sample point according to the distances between the new minority-class sample point and each sample point in the imbalanced dataset. If yes, save it; if not, regenerate the new minority-class sample point and continue to determine whether the regenerated new minority-class sample point is saved;
[0013] Traverse each preset minority-class sample point and traverse each target preset minority-class sample point corresponding to each preset minority-class sample point, and save multiple new minority-class sample points.
[0014] The beneficial effect of adopting the above further solution is that currently, after oversampling based on the Euclidean distance, it is easy to generate some noise data (noise sample points), which affects the effect of oversampling and also affects the training effect of the machine learning model. In the present invention, only new minority-class sample points are generated through preset minority-class sample points, so it is not easy to generate noise data and can ensure the training effect of the machine learning model.
[0015] Further, select M preset minority-class sample points from the top K preset minority-class sample points with the smallest distances, including:
[0016] Randomly select M preset minority-class sample points from the top K preset minority-class sample points with the smallest distances.
[0017] Further, it further includes:
[0018] Add all the saved new minority-class sample points to the imbalanced dataset.
[0019] 2) In the second aspect, the present invention also provides a data oversampling system, and the specific technical solution is as follows:
[0020] It includes a formula modification module and an oversampling processing module;
[0021] The formula modification module is used for: modifying the Euclidean distance calculation formula based on the Spearman correlation coefficient between each feature value and the result of the sample points in the imbalanced dataset, and determining the modified Euclidean distance calculation formula;
[0022] The oversampling processing module is used for: performing oversampling processing on the preset minority-class sample points in the imbalanced dataset by using the distances calculated by the modified Euclidean distance calculation formula.
[0023] On the basis of the above solution, a data oversampling system of the present invention can also be improved as follows.
[0024] Further, the oversampling processing module is specifically configured to:
[0025] Using the modified Euclidean distance calculation formula, calculate the distance between any preset minority class sample point and each of the remaining preset minority class sample points, and select M preset minority class sample points from the top K preset minority class sample points with the smallest distance as the M target preset minority class sample points corresponding to the preset minority class sample point, until the M target preset minority class sample points corresponding to each preset minority class sample point are obtained, where K and M are both positive integers;
[0026] Using any preset minority class sample point and any target preset minority class sample point corresponding to the preset minority class sample point, generate a new minority class sample point. Using the modified Euclidean distance calculation formula, calculate the distance between the new minority class sample point and each sample point in the imbalanced dataset, and determine whether to save the new minority class sample point according to the distance between the new minority class sample point and each sample point in the imbalanced dataset. If so, save it; if not, regenerate the new minority class sample point and continue to determine whether the regenerated new minority class sample point is saved;
[0027] Traverse each preset minority class sample point and traverse each target preset minority class sample point corresponding to each preset minority class sample point, and save multiple new minority class sample points.
[0028] Further, the oversampling processing module is also specifically configured to: randomly select M preset minority class sample points from the top K preset minority class sample points with the smallest distance.
[0029] Further, it further includes a data addition module, and the data addition module is configured to: add all the saved new minority class sample points to the imbalanced dataset.
[0030] 3) In a third aspect, the present invention also provides an electronic device, which includes a processor. The processor is coupled to a memory, and at least one computer program is stored in the memory. The at least one computer program is loaded and executed by the processor so that the electronic device implements any one of the above data oversampling methods.
[0031] 4) In a fourth aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements any one of the above data oversampling methods.
[0032] It should be noted that for the beneficial effects obtained by the technical solutions of the second to fourth aspects of the present invention and their corresponding possible implementation manners, reference may be made to the technical effects of the first aspect and its corresponding possible implementation manners described above, which will not be elaborated here. Description of the Drawings
[0033] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments of the present invention:
[0034] Figure 1 It is a schematic flowchart of a data oversampling method according to an embodiment of the present invention;
[0035] Figure 2 It is a schematic structural diagram of a data oversampling system according to an embodiment of the present invention;
[0036] Figure 3 It is a schematic structural diagram of an electronic device according to an embodiment of the present invention. Detailed implementation manners
[0037] The following describes the principles and features of the present invention. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention.
[0038] The following uses specific embodiments to elaborate in detail on the technical solutions of the present invention and how the technical solutions of the present invention solve the above technical problems. These several specific embodiments can be combined with each other. For the same or similar concepts or processes, they may not be repeated in some embodiments. The following will describe the embodiments of the present invention in conjunction with the accompanying drawings.
[0039] As Figure 1 shown, a data oversampling method according to an embodiment of the present invention includes the following steps:
[0040] S1. Modify the Euclidean distance calculation formula based on the Spearman correlation coefficient between each feature value and the result of the sample points in the imbalanced dataset, and determine the modified Euclidean distance calculation formula;
[0041] Among them, after feature quantization of the imbalanced sample set, an imbalanced dataset is obtained. Specifically:
[0042] 1) After feature quantization of the imbalanced sample set for training the disease recurrence probability prediction model, an imbalanced dataset is obtained. Specifically:
[0043] The features of the samples in the imbalanced sample set for training a disease recurrence probability prediction model include: age, gender, disease stage (such as the early stage, middle stage or late stage of the disease), time of initial diagnosis, treatment methods (such as surgery, radiotherapy and chemotherapy, etc.), treatment response (such as complete remission, partial remission, stable or progressive, etc.), genetic characteristics (such as gene mutations, etc.), lifestyle (such as smoking and drinking, eating habits), living environmental factors (such as environmental factors like air pollution, occupational exposure, etc.) and mental state (such as anxiety and depression, etc.). Different codes or numerical values can be assigned to different features according to a preset first assignment rule to achieve the quantification of each feature. Thus, the feature values of each feature in each sample can be obtained, and the feature values of each feature in each sample form corresponding sample points for subsequent calculations. All the sample points form the imbalanced data set corresponding to the imbalanced sample set for training the disease recurrence probability prediction model.
[0044] Among them, the preset first assignment rule can be set according to the actual situation. For example, the code or numerical value of "male" is 0, the code or numerical value of "female" is 1, the code or numerical value of the early stage of the disease is 1, the code or numerical value of the middle stage of the disease is 2, and the code or numerical value of the late stage of the disease is 3.
[0045] It should be noted that when the imbalanced data set can be the imbalanced data set corresponding to the imbalanced sample set for training the disease recurrence probability prediction model, the result (which can also be called the outcome) is whether the disease recurs.
[0046] 2) After quantifying the features of the imbalanced sample set for training a vehicle operation anomaly prediction model, an imbalanced data set is obtained. Specifically:
[0047] The features of the samples in the imbalanced sample set for training the vehicle operation anomaly prediction model include: vehicle speed, engine speed, engine temperature, vehicle vibration frequency, vehicle vibration amplitude, vehicle vibration position, tire pressure, tire wear condition, wear degree of brake pads, liquid level and quality of brake fluid, steering angle of the steering wheel, steering force of the steering wheel, road conditions, weather conditions, driver behavior data (such as hard acceleration, hard braking, frequent lane changes, etc.), maintenance records and mileage. Different codes or numerical values can be assigned to different features according to a preset second assignment rule to achieve the quantification of each feature. Thus, the feature values of each feature in each sample can be obtained, and the feature values of each feature in each sample form corresponding sample points for subsequent calculations. All the sample points form the imbalanced data set corresponding to the imbalanced sample set for training the vehicle operation anomaly prediction model.
[0048] Among them, the preset second assignment rule can be set according to the actual situation. For example, one percent of the vehicle speed is used as the code or value corresponding to the vehicle speed, and one percent of the engine speed is used as the code or value corresponding to the engine speed.
[0049] It should be noted that when the imbalanced dataset is the imbalanced dataset corresponding to the imbalanced sample set for training the vehicle operation anomaly prediction model, the result (which can also be called the outcome) is whether the vehicle has an anomaly.
[0050] 3) After feature quantization of the imbalanced sample set for training the nuclear magnetic resonance imaging device anomaly prediction model, an imbalanced dataset is obtained. Specifically:
[0051] The features of the samples in the imbalanced sample set for training the nuclear magnetic resonance imaging device anomaly prediction model include: image resolution, the ratio of signal to noise in the image (signal-to-noise ratio), magnetic field strength, scan sequence, scan time, operating temperature, cooling system status, power supply stability, gradient system status, radio frequency system status, maintenance date, maintenance content, usage frequency, and environmental factors (such as temperature, humidity, electromagnetic interference, etc.). Different codes or values can be assigned to different features according to the preset third assignment rule to achieve quantization of each feature. Thus, the feature values of each feature in each sample can be obtained, and the feature values of each feature in each sample form corresponding sample points, which is convenient for subsequent calculations. All sample points form the imbalanced dataset corresponding to the imbalanced sample set for training the nuclear magnetic resonance imaging device anomaly prediction model.
[0052] Among them, the preset third assignment rule can be set according to the actual situation. For example, the value of the image resolution is directly used as the code or value corresponding to the image resolution, and the value of the operating temperature is directly used as the code or value corresponding to the operating temperature.
[0053] It should be noted that when the imbalanced dataset is the imbalanced dataset corresponding to the imbalanced sample set for training the nuclear magnetic resonance imaging device anomaly prediction model, the result (which can also be called the outcome) is whether the nuclear magnetic resonance imaging device has an anomaly.
[0054] Other imbalanced datasets can also be obtained according to the actual situation.
[0055] Among them, the following method can be used to determine whether a sample set is an imbalanced sample set. Specifically:
[0056] 1) The first determination method:
[0057] The proportion between the numbers of samples of various categories in the sample set is directly calculated to determine the balance of the sample set. When the proportion of the number of samples of a certain category in the sample set does not exceed the first preset proportion threshold, it can be determined that the sample set is an unbalanced sample set, and the samples of this category are called minority-class samples, and the sample points corresponding to the minority-class samples are called minority-class sample points. Among them, the first preset proportion threshold can be set according to the actual situation. For example, the preset proportion threshold can be 30% or 40%, etc., which can be set according to the actual situation.
[0058] 2) The second judgment method:
[0059] A bar chart or a pie chart of the numbers of samples of various categories is drawn, which can intuitively display the distribution of the numbers of samples of various categories and can intuitively judge whether the sample set is an unbalanced sample set, and then determine the minority-class samples.
[0060] Among them, for any two sample points x and y in the unbalanced dataset, the modified Euclidean distance calculation formula is defined as:
[0061]
[0062] Among them, D(x, y) represents: the distance between sample point x and sample point y, x = (s 1 , s 2 ,..., s n ), y = (t 1 , t 2 ,..., t n ), x and y are two sample points in the unbalanced dataset, s 1 represents: the first eigenvalue in sample point x, s 2 represents: the second eigenvalue in sample point x, s n represents: the nth eigenvalue in sample point x, t 1 represents: the first eigenvalue in sample point y, t 2 represents: the second eigenvalue in sample point y, t n represents: represents the nth eigenvalue in sample point y, n represents: the number of eigenvalues in a sample point, which is equal to the number of features in the sample, c q represents: the Spearman correlation coefficient between the qth feature and the result (calculated from the qth eigenvalue and the result), c 1 represents: the Spearman correlation coefficient between the first feature and the result, c 2 represents: the Spearman correlation coefficient between the second feature and the result, c n represents: the Spearman correlation coefficient between the nth feature and the result.
[0063] S2. Use the distance calculated by the modified Euclidean distance calculation formula to perform oversampling on the preset minority class sample points in the imbalanced dataset.
[0064] Optionally, in S2, using the distance calculated by the modified Euclidean distance calculation formula to perform oversampling on the preset minority class sample points in the imbalanced dataset includes:
[0065] Use the modified Euclidean distance calculation formula to calculate the distance between any preset minority class sample point and each of the remaining preset minority class sample points, and select M preset minority class sample points from the top K preset minority class sample points with the smallest distance as the M target preset minority class sample points corresponding to this preset minority class sample point until the M target preset minority class sample points corresponding to each preset minority class sample point are obtained, where both K and M are positive integers;
[0066] Use any preset minority class sample point and any target preset minority class sample point corresponding to this preset minority class sample point to generate a new minority class sample point. Use the modified Euclidean distance calculation formula to calculate the distance between this new minority class sample point and each sample point in the imbalanced dataset, and determine whether to save this new minority class sample point according to the distance between this new minority class sample point and each sample point in the imbalanced dataset. If so, save it; if not, regenerate the new minority class sample point and continue to judge whether the regenerated new minority class sample point is saved;
[0067] Traverse each preset minority class sample point and traverse each target preset minority class sample point corresponding to each preset minority class sample point, and save multiple new minority class sample points.
[0068] Optionally, selecting M preset minority class sample points from the top K preset minority class sample points with the smallest distance includes:
[0069] Randomly select M preset minority class sample points from the top K preset minority class sample points with the smallest distance.
[0070] Optionally, it further includes: adding all the saved new minority class sample points to the imbalanced dataset.
[0071] The oversampling process for the preset minority class sample points in the imbalanced dataset is elaborated in detail as follows:
[0072] S21. Use the modified Euclidean distance calculation formula to calculate the distance between any minority class sample point and each of the remaining sample points. Determine whether the proportion of minority class sample points among the top K1 sample points with the smallest distances is not less than the second preset proportion threshold. If so, determine this minority class sample point as a preset minority class sample point, which can also be called a type-A minority class sample point. If not, do not determine this minority class sample point as a preset minority class sample point. Traverse each minority class sample point to determine multiple preset minority class sample points, and record the number of preset minority class sample points as T.
[0073] Among them, both K1 and T are positive integers. The second preset proportion threshold can be 60%, or it can be set according to the actual situation.
[0074] Among them, the top K1 sample points with the smallest distances refer to: Sort the distances between any minority class sample point calculated and each of the remaining sample points in ascending order to obtain a first sequence, and use the sample points corresponding to the first K1 distances in the first sequence as the top K1 sample points with the smallest distances.
[0075] S22. Use the modified Euclidean distance calculation formula to calculate the distance between any preset minority class sample point and each of the remaining preset minority class sample points, and select M preset minority class sample points from the top K preset minority class sample points with the smallest distances as the M target preset minority class sample points corresponding to this preset minority class sample point, until the M target preset minority class sample points corresponding to each preset minority class sample point are obtained. Among them, both K and M are positive integers, and K ≥ M > 0;
[0076] Among them, the top K preset minority class sample points with the smallest distances refer to: Sort the distances between any preset minority class sample point calculated and each of the remaining preset minority class sample points in ascending order to obtain a second sequence, and use the preset minority class sample points corresponding to the first K distances in the second sequence as the top K preset minority class sample points with the smallest distances.
[0077] S23. Generate a new minority class sample point using any preset minority class sample point and any target preset minority class sample point corresponding to this preset minority class sample point. Use the modified Euclidean distance calculation formula to calculate the distance between this new minority class sample point and each sample point in the imbalanced dataset, and determine whether to save this new minority class sample point based on the distance between this new minority class sample point and each sample point in the imbalanced dataset. If so, save it. If not, regenerate a new minority class sample point and continue to determine whether the regenerated new minority class sample point is saved;
[0078] For the convenience of description, denote the i-th preset minority class sample point as x i, for i = 1, 2, 3... T, the sample point x i The j-th target preset minority class sample point among the randomly selected M target preset minority class sample points corresponding to is denoted as x ij , j = 1, 2, 3... M.
[0079] For the preset minority class sample point x i and the target preset minority class sample point x ij Randomly generate a random number r ij , and generate a new minority class sample point y using the following formula ij :
[0080] y ij = x i + r ij (x ij - x i )
[0081] where the random number r ij has a value range of: [0, 1].
[0082] Using the modified Euclidean distance calculation formula, calculate the distance between the new minority class sample point y ij and each sample point in the imbalanced dataset, and judge whether the number of minority class sample points among the top K2 sample points with the smallest distance is not less than the third preset proportion threshold to obtain the first judgment result. When the first judgment result is yes, save the new minority class sample point y ij , when the first judgment result is no, regenerate a random number r within the range of [0, 1] ij , regenerate a new minority class sample point y ij , until the first judgment result corresponding to the regenerated minority class sample point y ij is yes, and save the regenerated minority class sample point y ij .
[0083] where the third preset proportion threshold can be 40%, or can be set according to the actual situation.
[0084] where the top K2 sample points with the smallest distance refer to: sorting the distances between the calculated new minority class sample point y ij and each sample point in the imbalanced dataset in ascending order to obtain the third sequence, and taking the sample points corresponding to the first K2 distances in the third sequence as the top K2 sample points with the smallest distance.
[0085] S24. Traverse each preset minority class sample point and traverse each target preset minority class sample point corresponding to each preset minority class sample point, and save multiple new minority class sample points. The number of saved new minority class sample points is: T × M.
[0086] Among them, K, K1, and K2 can be equal, and the second preset proportion threshold and the third preset proportion threshold can be equal.
[0087] Then, add the saved T×M new minority class sample points to the imbalanced data set.
[0088] Optionally, in the above technical solution, it further includes:
[0089] After adding the saved T×M new minority class sample points to the imbalanced data set, determine whether the current imbalanced data set is an imbalanced data set to obtain a second judgment result. When the second judgment result is no, it means that the current imbalanced data set has become a balanced data set, and this data set can be used for model training. When the second judgment result is yes, continue to perform oversampling processing using S20 to S24 until the second judgment result is no.
[0090] In the present invention, in the modified Euclidean distance calculation formula, the Spearman correlation coefficient between each feature and the result is considered, and a weight of 1 + |c is assigned to the qth feature value q |, obviously 1 ≤ 1 + |c q | ≤ 2. A feature with a relatively large absolute value of the Spearman correlation coefficient means it is more important, so a larger weight is assigned. A feature with a relatively small absolute value of the Spearman correlation coefficient may also be helpful for prediction, so the weight is: 1 + |c q |, rather than |c i |, to make them also contribute to some extent, which helps to improve the training effect of the machine learning model. Moreover, a preset minority class sample point (type A minority class sample point) is defined, and only new minority class sample points are generated through the preset minority class sample points, so it is not easy to generate "noisy data". Further, for the generated new minority class sample points, it is required whether the number of minority class sample points among the top K2 sample points with the smallest distance is not less than the third preset proportion threshold, which is even less likely to generate "noisy data" and effectively reduces the generation of "noisy data".
[0091] In the above embodiments, although the steps are numbered S1, S2, etc., these are only specific embodiments given by the present invention. Those skilled in the art can adjust the execution order of S1, S2, etc. according to the actual situation, and this is also within the protection scope of the present invention. It can be understood that in some embodiments, it may include some or all of the above embodiments.
[0092] As Figure 2 shown, a data oversampling system 200 according to an embodiment of the present invention includes a formula modification module 201 and an oversampling processing module 202;
[0093] The formula modification module 201 is used to: modify the Euclidean distance calculation formula based on the Spearman correlation coefficient between each feature value of the sample points in the imbalanced dataset and the result, and determine the modified Euclidean distance calculation formula;
[0094] The oversampling processing module 202 is used to: perform oversampling processing on the preset minority class sample points in the imbalanced dataset using the distances calculated by the modified Euclidean distance calculation formula.
[0095] Optionally, in the above technical solution, the oversampling processing module 202 is specifically used to:
[0096] Use the modified Euclidean distance calculation formula to calculate the distances between any preset minority class sample point and each of the remaining preset minority class sample points, and select M preset minority class sample points from the top K preset minority class sample points with the smallest distances as the M target preset minority class sample points corresponding to this preset minority class sample point, until the M target preset minority class sample points corresponding to each preset minority class sample point are obtained, where both K and M are positive integers;
[0097] Use any preset minority class sample point and any target preset minority class sample point corresponding to this preset minority class sample point to generate a new minority class sample point, use the modified Euclidean distance calculation formula to calculate the distances between this new minority class sample point and each sample point in the imbalanced dataset, and determine whether to save this new minority class sample point according to the distances between this new minority class sample point and each sample point in the imbalanced dataset. If so, save it; if not, regenerate the new minority class sample point and continue to determine whether the regenerated new minority class sample point is saved;
[0098] Traverse each preset minority class sample point and traverse each target preset minority class sample point corresponding to each preset minority class sample point, and save multiple new minority class sample points.
[0099] Optionally, in the above technical solution, the oversampling processing module 202 is also specifically used to: randomly select M preset minority class sample points from the top K preset minority class sample points with the smallest distances.
[0100] Optionally, in the above technical solution, it further includes a data addition module, and the data addition module is used to: add all the saved new minority class sample points to the imbalanced dataset.
[0101] It should be noted that the beneficial effects of the data oversampling system 200 provided in the above embodiments are the same as those of the above data oversampling method, and will not be elaborated here. In addition, when the system provided in the above embodiments realizes its functions, only the division of the above functional modules is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the system can be divided into different functional modules according to the actual situation to complete all or part of the functions described above. In addition, the system and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be seen in the method embodiments, which will not be elaborated here.
[0102] Among them, the data oversampling system of the present invention can be a computer program (including program code) running on a computer device. For example, the data oversampling system of the present invention is an application software and can be used to execute the corresponding steps in the data oversampling method of the present invention.
[0103] In some embodiments, the data oversampling system of the present invention can be implemented in a combination of software and hardware. As an example, the data oversampling system of the present invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the data oversampling method of the present invention. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs, Application Specific Integrated Circuits), DSPs, programmable logic devices (PLDs, Programmable Logic Devices), complex programmable logic devices (CPLDs, Complex Programmable Logic Devices), field-programmable gate arrays (FPGAs, Field-Programmable Gate Arrays) or other electronic components.
[0104] Among them, the modules described in the embodiments of the present invention can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the module itself in some cases.
[0105] An electronic device according to an embodiment of the present invention includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements any one of the above data oversampling methods. That is to say, an electronic device according to an embodiment of the present invention may include, but is not limited to: a processor and a memory; the memory is used to store a computer program; the processor is used to execute the data oversampling method shown in any one of the embodiments of the present invention by calling the computer program.
[0106] In an alternative embodiment, an electronic device is provided, such asFigure 3 As shown Figure 3 The electronic device 4000 shown in Figure 3 includes: a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as being connected through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 can be used for data interaction between this electronic device and other electronic devices, such as data sending and / or data receiving, etc. It should be noted that in practical applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation to the embodiments of the present invention.
[0107] The processor 4001 can be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in combination with the disclosure of the present invention. The processor 4001 can also be a combination that realizes computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0108] The bus 4002 may include a path for transmitting information between the above components. The bus 4002 can be a PCI (Peripheral Component Interconnect, peripheral component interconnect standard) bus or an EISA (Extended Industry Standard Architecture, extended industry standard architecture) bus, etc. The bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 3 only a thick line is used to represent the bus 4002 in Figure 3 , but it does not mean that there is only one bus or one type of bus.
[0109] The memory 4003 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0110] The memory 4003 is used to store the application program code (computer program) for executing the solution of the present invention and is controlled by the processor 4001 for execution. The processor 4001 is used to execute the application program code stored in the memory 4003 to implement the content shown in the foregoing method embodiments.
[0111] Among them, the electronic device can also be a terminal device, and the terminal device can be any device that can install applications, including at least one of a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a smart TV, and a smart vehicle device.
[0112] It should be noted that Figure 3 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.
[0113] A computer-readable storage medium according to an embodiment of the present invention has a computer program stored thereon, and when the computer program is executed by a processor, it implements any one of the above data oversampling methods.
[0114] Optionally, the computer-readable storage medium can be a Read-Only Memory (ROM), a Random Access Memory (RAM), a Compact Disc Read-Only Memory (CD-ROM), magnetic tape, a floppy disk, and an optical data storage device, etc.
[0115] In an exemplary embodiment, a computer program product or a computer program is further provided. The computer program product or the computer program includes computer instructions that are stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device performs any one of the above data oversampling.
[0116] Computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0117] It should be understood that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of methods and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0118] The computer-readable storage medium provided by the embodiments of the present invention may be, but is not limited to, a system, device, or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or component.
[0119] The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device is caused to execute the method shown in the above embodiments.
[0120] The above description is only the preferred embodiments of the present invention and the description of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present invention is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present invention.
[0121] It should be noted that the terms "first", "second", etc. in the specification and claims of this application are used to distinguish similar objects, and are used to limit a specific order or sequence. Under appropriate circumstances, the order of use of similar objects can be interchanged so that the embodiments of this application described here can be implemented in an order other than the order shown or described.
[0122] Those skilled in the art know that the present invention can be implemented as a system, method, or computer program product. Therefore, the present invention can be specifically implemented in the following forms: it can be completely hardware, can also be completely software (including firmware, resident software, microcode, etc.), and can also be in the form of a combination of hardware and software, which is generally referred to as "circuit", "module", or "system" in this article. In addition, in some embodiments, the present invention can also be implemented in the form of a computer program product in one or more computer-readable media, and the computer-readable media contains computer-readable program code.
[0123] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A data oversampling method, characterized in that: include: Based on the Spearman correlation coefficient between each characteristic value of the sample point in the unbalanced data set and the result, the Euclidean distance calculation formula is modified to determine the modified Euclidean distance calculation formula; The distance calculated by the modified Euclidean distance calculation formula is used to perform oversampling processing on the preset minority class sample points in the imbalanced data set.
2. A data oversampling method according to claim 1, characterized in that: Using the distance calculated by the modified Euclidean distance calculation formula, oversampling is performed on the preset minority class sample points in the imbalanced data set, including: The modified Euclidean distance calculation formula is used to calculate the distance between any preset minority class sample point and each remaining preset minority class sample point, and M preset minority class sample points are selected from the first K preset minority class sample points with the smallest distance as the M target preset minority class sample points corresponding to the preset minority class sample points, until M target preset minority class sample points corresponding to each preset minority class sample point are obtained, where K and M are both positive integers; Generate a new minority class sample point by using any preset minority class sample point and any target preset minority class sample point corresponding to the preset minority class sample point, calculate the distance between the new minority class sample point and each sample point in the imbalanced data set by using the modified Euclidean distance calculation formula, and determine whether to save the new minority class sample point according to the distance between the new minority class sample point and each sample point in the imbalanced data set, if yes, save it, if not, regenerate a new minority class sample point, and continue to determine whether the regenerated new minority class sample point is saved; Each preset minority class sample point is traversed and each target preset minority class sample point corresponding to each preset minority class sample point is traversed to save a plurality of new minority class sample points.
3. A data oversampling method according to claim 2, characterized in that: From the first K preset minority class sample points with the smallest distance, select M preset minority class sample points, including: Randomly select M preset minority class sample points from the first K preset minority class sample points with the smallest distance.
4. A data oversampling method according to claim 2 or 3, characterized in that: Also includes: All new minority class sample points saved are added to the imbalanced data set.
5. A data oversampling system, characterized in that: Including formula modification module and oversampling processing module; The formula modification module is used to: modify the Euclidean distance calculation formula based on the Spearman correlation coefficient between each characteristic value of the sample point in the unbalanced data set and the result, and determine the modified Euclidean distance calculation formula; The oversampling processing module is used to: perform oversampling processing on preset minority class sample points in the imbalanced data set using the distance calculated by the modified Euclidean distance calculation formula.
6. A data oversampling system according to claim 5, characterized in that: The oversampling processing module is specifically used for: The modified Euclidean distance calculation formula is used to calculate the distance between any preset minority class sample point and each remaining preset minority class sample point, and M preset minority class sample points are selected from the first K preset minority class sample points with the smallest distance as the M target preset minority class sample points corresponding to the preset minority class sample points, until M target preset minority class sample points corresponding to each preset minority class sample point are obtained, where K and M are both positive integers; Generate a new minority class sample point by using any preset minority class sample point and any target preset minority class sample point corresponding to the preset minority class sample point, calculate the distance between the new minority class sample point and each sample point in the imbalanced data set by using the modified Euclidean distance calculation formula, and determine whether to save the new minority class sample point according to the distance between the new minority class sample point and each sample point in the imbalanced data set, if yes, save it, if not, regenerate a new minority class sample point, and continue to determine whether the regenerated new minority class sample point is saved; Each preset minority class sample point is traversed and each target preset minority class sample point corresponding to each preset minority class sample point is traversed to save a plurality of new minority class sample points.
7. The data oversampling system according to claim 6, characterized in that: The oversampling processing module is also specifically used for: Randomly select M preset minority class sample points from the first K preset minority class sample points with the smallest distance.
8. A data oversampling system according to claim 6 or 7, characterized in that: It also includes a data adding module, which is used to add all the new minority class sample points saved to the unbalanced data set.
9. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements a data oversampling method according to any one of claims 1 to 4 when executing the computer program.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the data oversampling method according to any one of claims 1 to 4 is implemented.