Unbalanced data regression method and system, storage medium and terminal

By using a multi-grained spatial generation model and dynamic oversampling strategy based on grid division in unbalanced data regression, the problem of taking into account data distribution and noise samples in large sample scenarios is solved, and real-time adaptation to dynamic data flow is achieved, which improves the model's learning ability and prediction performance.

CN120180385AActive Publication Date: 2025-06-20YIBIN VOCATIONAL & TECH COLLEGE
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510637692.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-06-20
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

When processing large sample sizes, existing unbalanced data regression methods are difficult to take into account the overall distribution of the data, are easily disturbed by noise samples, and destroy the continuity and intrinsic correlation of the samples. At the same time, existing methods cannot effectively handle dynamic data flow, resulting in the model being unable to adapt to data changes in real time.

Method used

A multi-grained spatial generation model based on mesh division is proposed. By dividing similar samples in the same particle size space, and oversampling them with spatial importance and dynamic neighborhood radius, the generation of noise samples is reduced. At the same time, the granular space and sampling strategies are dynamically updated to adapt to changes in data flow.

Benefits of technology

In larger sample size scenarios, the generation of noise samples is effectively reduced, the data continuity and intrinsic correlation are maintained, and the processing efficiency is improved. At the same time, it can adapt to changes in dynamic data flow in real time, improving the learning ability and prediction performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180385A_ABST
    Figure CN120180385A_ABST
Patent Text Reader

Abstract

The invention discloses an unbalanced data regression method and system, a storage medium and a terminal, and belongs to the technical field of data regression, and the method comprises the steps: constructing a grid division-based multi-granularity space generation model, and dividing similar samples in the same granularity space; oversampling is independently carried out in each granularity space, and an oversampling range is defined by combining the space importance degree and the dynamic neighborhood radius. According to the method, under the scene of large sample size, the data stream and imbalance problems are comprehensively considered, the processing efficiency is improved, the overall distribution of data is considered, the interference of noise samples is reduced, the continuity and internal correlation of the samples are reserved, and a new thought is provided for the field of imbalance data stream regression.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data regression, and particularly to an imbalanced data regression method, system, storage medium and terminal. Background Art

[0002] In traditional standard machine learning algorithms, the number of samples in each class considered is approximately the same, and each class is equally important. However, there are a large number of imbalance problems in the data of various fields in the real world. Among them, the industrial data collected by mechanical equipment in the industrial field is more likely to be imbalanced and complex and diverse. For example, the fault information of mechanical equipment, the abnormal data of disaster monitoring systems, the unqualified data of product quality inspection, etc. In these actual application scenarios, the prediction problem of rare events must be solved, and the data imbalance problem will seriously affect the accuracy of rare event prediction. Traditional machine learning algorithms are usually designed to reduce the overall error. In the case of imbalanced data, they tend to predict the majority class, thus ignoring the rare events of the minority class. For example, when the standard logistic regression method processes an imbalanced dataset, it usually underestimates the probability of rare events, resulting in a low recall rate and insufficient event prediction. The data characteristics of imbalanced regression also increase the difficulty of model learning, making it difficult for the model to capture the patterns and rules of rare event occurrence, thereby affecting the prediction performance, resulting in a higher monitoring difficulty for abnormal situations in various fields, and false predictions may lead to huge cost waste. Therefore, the regression of imbalanced data is particularly important.

[0003] Most of the existing methods for solving the imbalanced regression problem divide the samples in the regression task into different classes. There are two drawbacks to this type of method: 1. Dividing the samples with continuous target values into different classes destroys the continuity of the data and loses some important information on the class boundaries. It is easy to introduce noise samples during oversampling, thereby reducing the prediction accuracy and stability of the model: 2. Data oversampling requires multiple calculations of the distance between samples. When the number of samples is large, it will lead to a reduction in processing efficiency. Therefore, how to reduce the sample size, reduce the sample neighbor search space while taking into account the overall distribution of the samples, retain the continuity of the data, and avoid the interference of noise samples will be an effective way to solve imbalanced regression.

[0004] In addition, the data processed by existing imbalanced problems is usually static, that is, once the model is trained and determined, its parameters and structure will not be dynamically adjusted, making it impossible to learn the characteristics of new data in real time. However, most of the data in reality is dynamic data that appears in the form of streams, and its data distribution will continue to change over time. Unbalanced stream data is widely used in fields such as network monitoring and fault detection. Online learning is a key technology for processing stream data, but online real-time learning of data streams will face some difficulties: 1. Data distribution continues to change over time; 2. The inherent structure is unstable and concept drift occurs; 3. The data imbalance state is unstable; 4. The model cannot process new data in real time. For the problem of imbalanced data stream classification, there are few solutions that comprehensively consider data streams and imbalance problems, and they are not suitable for direct improvement for regression problems.

[0005] To this end, researchers have proposed a variety of algorithms to deal with problems such as uneven category distribution, concept drift, and scarcity of minority class samples in dynamic data streams. These methods cover a variety of technologies from data sampling to classifier optimization. By dynamically adjusting the model or data, they effectively improve the learning ability of unbalanced data, such as the C-SMOTE algorithm, CSARF algorithm, VFCSMOTE algorithm, ROSE algorithm, etc. However, while these unbalanced data regression methods improve processing efficiency, they still have the following defects: 1. The idea of ​​most methods to solve imbalanced regression is to divide samples into rare samples and common samples based on the correlation scale assigned to each sample by the correlation function, so as to be suitable for imbalanced classification algorithms. This approach destroys the continuity and intrinsic correlation of the sample target value, and when the threshold is not selected properly, the resulting sample misclassification will affect the prediction performance of the model; 2. The existing clustering-based methods to reduce the sample size to improve processing efficiency are highly dependent on clustering parameters and clustering results. When the clustering parameters are not selected properly or the data distribution is irregular, the clustering effect will be poor, affecting the subsequent sampling process; 3. The data distribution in actual application scenarios such as industry and manufacturing is very complex. The traditional method of oversampling boundary samples is difficult to identify complex decision boundaries and is prone to ignoring some important samples. Summary of the invention

[0006] The object of the present invention is to study the regression method for imbalanced data in the scenario of a large sample size. In the face of a large sample size scenario, while improving the processing efficiency, the existing imbalanced data regression methods are difficult to take into account the overall distribution of data, are easily interfered by noise samples, and damage the continuity and internal correlation of samples. An improved granular sphere model (grid self-organizing granular sphere model) is proposed to process the imbalanced regression data with a large sample size. At the same time, around the problem that the existing imbalanced data regression methods are limited to processing static data and cannot process dynamic data that appears in the form of a stream, an imbalanced data stream regression method that comprehensively considers the data stream and the imbalanced problem in the regression task scenario is realized.

[0007] The object of the present invention is achieved by the following technical solutions: In the first aspect, an imbalanced data regression method is provided, including: Construct a multi-granularity space generation model based on grid division, and divide similar samples into the same granularity space; Perform oversampling independently in each granularity space, where the oversampling range is delimited by combining the spatial importance and the dynamic neighborhood radius.

[0008] In some embodiments, the construction of the multi-granularity space generation model based on grid division includes: Determine the division area of the granularity space according to the spatial area formed by the boundary ranges of each dimension of the samples; Subdivide the division area into several small rectangles; According to the number of samples sparsely distributed in each small rectangle, divide the small rectangles into different types; Calculate the distribution degree and the minimum distribution degree of the granularity space. If the distribution degree of the granularity space is less than the minimum distribution degree, continue to divide the granularity space until the distribution degree of all granularity spaces is not less than the minimum distribution degree.

[0009] In some embodiments, the types of the small rectangles include blank rectangles with a sample number of 0, high-density rectangles with the largest sample number, and sparse rectangles between the high-density rectangles and the blank rectangles.

[0010] In some embodiments, the step of continuing to divide the granularity space if the distribution degree of the granularity space is less than the minimum distribution degree includes: For the granularity space that needs to be further divided, select the center of the granularity space and the centers of the samples in the high-density rectangle as the centers of two new granularity spaces respectively. Calculate the distances from each sample to the two centers, and assign the samples to the granularity space belonging to the nearest center to obtain two new granularity spaces.

[0011] In some embodiments, the calculation of the spatial importance includes: Calculate the sample density of the continuous target values of the samples; Use a weighting function to assign weights to the sample density to obtain the density weight of each sample; Calculate the spatial importance according to the density weight of each sample.

[0012] In some embodiments, combining the spatial importance and the dynamic neighborhood radius to delimit the oversampling range includes: Determine the oversampling quantity by the following formula: , where represents the oversampling quantity, represents the neighborhood radius, represents the spatial importance.

[0013] In some embodiments, the specific process of oversampling includes: When , take the samples with higher importance in the granularity space as the candidate seed samples for oversampling; Select a seed sample, and select a neighbor sample within the range of its neighborhood radius for linear interpolation to synthesize a new sample; Loop to select seed samples for new sample synthesis until the number of synthesized new samples reaches the oversampling quantity.

[0014] In a second aspect, there is provided an imbalanced data regression system, including: A multi-granularity space division module, configured to construct a multi-granularity space generation model based on grid division, and divide similar samples into the same granularity space; An oversampling module, configured to perform oversampling independently in each granularity space, wherein the oversampling range is delimited by combining the spatial importance and the dynamic neighborhood radius.

[0015] In a third aspect, there is provided a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements an imbalanced data regression method according to any one of the first aspect.

[0016] In a fourth aspect, there is provided a terminal, including a memory and a processor, where the memory stores computer instructions that can run on the processor, and when the processor runs the computer instructions, it executes an imbalanced data regression method according to any one of the first aspect.

[0017] It should be further noted that the corresponding technical features of the above embodiments can be combined or replaced with each other without conflict to form a new technical solution.

[0018] Compared with the prior art, the beneficial effects of the present invention are: 1. In a scenario with a relatively large sample size, the present invention constructs a multi-granularity space generation model based on grid partitioning, divides similar samples into the same granularity space, takes into account the overall distribution of data while improving processing efficiency, and the multi-granularity space generation model can iteratively granulate large-scale samples into multiple granularity spaces based on the overall distribution of data, uses the overall space to replace multiple similar samples, and makes the sample distribution of each final division uniform. The subsequent sampling process is separately processed within each space, which can effectively avoid sampling in the overlapping areas between different sample distributions and avoid the generation of noise samples.

[0019] 2. The present invention combines the spatial importance and the dynamic neighborhood radius to delimit the oversampling range, and subsequent samples are oversampled within the delimited oversampling range, which can reduce the possibility of generating noise samples due to oversampling within the granule sphere and avoid destroying the continuity and inherent correlation of the samples.

[0020] 3. The multi-granularity space generation model of the present invention can dynamically update the granularity space and the sampling strategy as the data flows, adapting to the dynamic change characteristics of the data stream itself. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 It is a flowchart of an imbalanced data regression method shown in an embodiment of the present invention; Figure 2 It is a schematic diagram of sample distribution shown in an embodiment of the present invention; Figure 3 It is a schematic diagram of the grid area of the granularity space shown in an embodiment of the present invention; Figure 4 It is the blank part where there are no samples in the rectangular area shown in an embodiment of the present invention; Figure 5 It is the result of dividing the grid area of the granularity space by using small rectangles shown in an embodiment of the present invention; Figure 6 It is a schematic diagram of three types of rectangles shown in an embodiment of the present invention; Figure 7 It is the initial granularity space shown in an embodiment of the present invention; Figure 8 It is based on in an embodiment of the present invention Figure 7 A schematic diagram of two new granularity spaces divided from the initial granularity space; Figure 9 It is based on in an embodiment of the present invention Figure 8 A schematic diagram of four new granularity spaces divided from the two new granularity spaces; Figure 10 It is based on in an embodiment of the present invention Figure 9 A schematic diagram of eight new granularity spaces divided from the four new granularity spaces; Figure 11 As shown in the embodiments of the present invention, based on Figure 10 the eight new granularity spaces divided in Figure 12 As shown in the embodiments of the present invention, based on Figure 11 the sixteen new granularity spaces divided in Figure 13 As shown in the embodiments of the present invention, it is a schematic diagram of oversampling within the sample neighborhood range. Detailed implementation manners

[0022] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. The components of the embodiments of the present application described and illustrated herein can be arranged and designed in various different configurations. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0023] It should be noted that the defects existing in the above prior art solutions are all the results obtained by the inventors through practice and careful research. Therefore, the discovery process of the above problems and the solutions proposed by the embodiments of the present application in the following text for the above problems should be the contributions made by the inventors to the present application during the invention creation process, rather than being understood as the technical content known to those skilled in the art.

[0024] In response to the technical problems pointed out in the background art, the embodiments provided by the present invention are as follows: Referring to Figure 1 , in an exemplary embodiment, an unbalanced data regression method is provided, including: Constructing a multi-granularity space generation model based on grid division to divide similar samples into the same granularity space; Performing oversampling independently in each granularity space, where the oversampling range is delimited by combining the spatial importance and the dynamic neighborhood radius.

[0025] Collecting sensor data of a certain mechanical device, and given a sample set of a granularity space GS , where represents the m-dimensional feature vector of the sample, represents the class label of the sample, and the center O of GS is calculated by the following formula: .

[0026] Assuming that the sample distribution of the granularity space GS is as Figure 2 shown, the definition of the space division strategy is as follows: Definition 1 (Partition Region): The partition region of the granularity space GS is the spatial region formed by the boundary ranges of each dimension of the samples. The boundary of the j-th dimension of the sample space is defined as follows: , where represents the value of the sample center O in the j-th dimension, represents the maximum distance from the samples in the partitioned space to the space center O, and its calculation formula is as follows: .

[0027] According to Definition 1, the coverage range of the partitioned space region in each dimension is , Figure 3 Figure gives a schematic diagram of the grid region of the granularity space GS. In the figure, represents the -th dimension of the sample space.

[0028] Definition 2 (Side Length of Partitioned Space): The partitioned space region is subdivided into several small rectangles, and the side lengths of each small rectangle are defined as follows: .

[0029] According to Definition 1 and Definition 2, the number of small grids that the space region can be subdivided into is . Figure 4 The gray area in can roughly represent the blank part where there are no samples in the rectangular area. The space region partitioning strategy can accurately represent the proportion of samples in the grid region using small grids. Figure 5 Figure shows the result of partitioning the grid region of the granularity space GS using small rectangles.

[0030] Definition 3: After the granularity space region is partitioned by grids, according to the number of samples sparsely distributed in each small rectangle, the small rectangles are divided into: 1) Rectangles with a sample number of 0 are called "blank rectangles"; 2) The grid with the largest number of samples is called "high-density rectangle"; 3) The remaining grids are called "sparse rectangles". As Figure 6 shown.

[0031] Definition 4 (Granularity Distribution Degree and Minimum Distribution Degree): After the partition region of the granularity space is partitioned by grids, the number of blank rectangle grids obtained is, and the definitions of the granularity distribution degree F and the minimum distribution degree are as follows: , .

[0032] Based on Definition 4, the partitioning strategy of the granularity space of the present invention is as follows: 1) Given a regression data set , where Denote the sample, Denote the class label corresponding to the sample. Initially, use this dataset as the sample set of the initial granularity space GS, and adopt a space partitioning strategy to divide its area into a rectangular grid; 2) Calculate the distribution degree F of the granularity space GS. If , then the division of GS terminates and serves as the final granularity space structure; if , then GS needs to continue the division; 3) For the granularity space GS that needs to continue the division, select the center O of GS and the center of the samples within the "high-density rectangle" The two centers are used to perform subsequent division of the space. Take the center point O and as the center points of two new spaces and respectively. Calculate the distance from each sample to the two center points, and assign the sample to the granule ball belonging to the nearest center point to obtain two new granularity division spaces and ; 4) Repeat steps 2) - 4) for and respectively until all granularity spaces no longer continue to be divided.

[0033] According to the rectangular grid partitioning strategy of the granularity space, Figures 7 - 12 shows the schematic diagram of the granulation process of the new granularity space partitioning strategy on a four-class dataset, Figure 7 represents the initial granularity space, Figure 8 is the schematic diagram of two new granularity spaces obtained by partitioning the initial granularity space in Figure 7 , Figure 9 is the schematic diagram of four new granularity spaces obtained by partitioning the two new granularity spaces in Figure 8 , Figure 10 is the schematic diagram of eight new granularity spaces obtained by partitioning the four new granularity spaces in Figure 9 , Figure 11 is the schematic diagram of sixteen new granularity spaces obtained by partitioning the eight new granularity spaces in Figure 10 , Figure 12 is the schematic diagram of thirty-two new granularity spaces obtained by partitioning the sixteen new granularity spaces in Figure 11 .

[0034] In the imbalanced regression problem, users pay more attention to rare samples. Therefore, it is very important to measure the rarity of samples. Calculate the density function of the continuous target values of the samples, and use a weighting function to assign weights to the sample densities to obtain the density weights of each sample, which are used to reflect the importance of the samples.

[0035] Based on the regression dataset , the related definitions of the importance of the granularity space are as follows.

[0036] Definition 5 (Sample Density): The density of a sample is as follows: , where KDE(*) represents the kernel function, which is used to estimate the probability density function of continuous target values, and h represents the bandwidth. Further, the density of the sample is min-max normalized, and is used to represent the value range of, then the normalized sample density is as follows: .

[0037] Definition 6 (Sample Importance): The importance of a sample is as follows: , where is a hyperparameter that determines the degree of weight increase for important samples, is a very small positive constant to avoid the case where the weight is 0 or negative.

[0038] Definition 7 (Spatial Importance): The spatial importance is a weighting function, and the importance of the granularity space GS is as follows: .

[0039] According to the importance of the granularity space, the present invention can quickly find samples with a higher degree of importance. After representing the samples as multiple granularity spaces based on the rectangular grid partitioning model, it can be considered that the samples within the granularity space are evenly distributed and highly similar, so the samples also have approximately the same degree of importance. According to the spatial importance, the present invention can quickly find samples with a higher degree of importance and perform corresponding oversampling processing on these samples in units of partitions.

[0040] Definition 8 (Sample Neighborhood Radius): The neighborhood radius R of a sample is as follows: , where is the side length of the rectangular region of the granularity space in Definition 2.

[0041] Further, the calculation formula for the oversampling quantity of the granularity space GS is given: , where represents the oversampling quantity, represents the neighborhood radius, represents the spatial importance.

[0042] Combined with the importance of the granularity space and the neighborhood radius of the samples, the oversampling strategy of the granularity space GS is as follows: 1) When it indicates that the importance of most samples in GS is higher than the average level. These samples are used as candidate seed samples for oversampling, and the importance of the samples is used as the selection weight, which can reduce the probability of samples with lower importance in GS becoming seed samples; 2) Select a seed sample, such as Figure 13 the sample at the center of the dashed circle in and select a neighbor sample within the range of its neighborhood radius R to linearly interpolate and synthesize a new sample 3) Repeat step 2) until the number of synthesized new samples reaches .

[0043] The present invention aims at processing unbalanced regression data sets with a large number of samples, divides similar samples into the same rectangle to reduce the scale of the samples, improves the processing efficiency by adopting the divide-and-conquer idea, and at the same time avoids the possibility of generating noise samples due to sample oversampling between rectangles. And it improves the existing granularity space division model, quantifies the distribution of samples based on the rectangular grid division of the sample set, so that the sample distribution of each final division is uniform; on the other hand, it uses the probability density function of the continuous target value of the sample to calculate the importance of the sample, and the subsequent oversampling range is based on the dynamically calculated neighborhood radius and sample importance of the sample, avoiding destroying the continuity and internal correlation of the sample.

[0044] In another exemplary embodiment, based on the same inventive concept as the method embodiment, an unbalanced data regression system is provided, including: A multi-granularity space division module, configured to construct a multi-granularity space generation model based on grid division and divide similar samples into the same granularity space; An oversampling module, configured to perform oversampling independently in each granularity space, wherein the oversampling range is delimited by combining the space importance and the dynamic neighborhood radius.

[0045] In another exemplary embodiment, based on the same inventive concept as the method embodiment, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements an unbalanced data regression provided by an embodiment of the present invention. Based on such an understanding, the technical solution of this embodiment, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.

[0046] In another exemplary embodiment, based on the same inventive concept as the method embodiment, a terminal is provided, including a memory and a processor. A computer instruction that can run on the processor is stored on the memory, and when the processor runs the computer instruction, it executes an unbalanced data regression method provided by an embodiment of the present invention.

[0047] The processor can be a single-core or multi-core central processing unit or a specific integrated circuit, or an integrated circuit configured to implement one or more of the present invention.

[0048] The embodiments of the subject matter and the functional operations described in this specification can be implemented in the following: tangibly embodied computer software or firmware, computer hardware including the structures disclosed in this specification and their structural equivalents, or a combination of one or more of them. The embodiments of the subject matter described in this specification can be implemented as one or more computer programs, that is, one or more modules in computer program instructions encoded on a tangible non-transitory program carrier to be executed by a data processing device or to control the operation of a data processing device. Alternatively or additionally, the program instructions can be encoded on an artificially generated propagated signal, such as a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode and transmit information to a suitable receiver device for execution by a data processing device.

[0049] The processing and logical flows described in this specification can be executed by one or more programmable computers executing one or more computer programs to perform corresponding functions by operating on input data and generating outputs. The processing and logical flows can also be executed by dedicated logic circuits, such as FPGAs (Field Programmable Gate Arrays) or ASICs (Application Specific Integrated Circuits), and the device can also be implemented as dedicated logic circuits.

[0050] Processors suitable for executing computer programs include, for example, general and / or special-purpose microprocessors, or any other type of central processing unit. Generally, the central processing unit will receive instructions and data from read-only memory and / or random access memory. The basic components of a computer include a central processing unit for implementing or executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, etc., or the computer will be operatively coupled to such a mass storage device to receive data therefrom or transfer data thereto, or both. However, a computer is not necessarily required to have such devices. In addition, a computer may be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name just a few examples.

[0051] It should be understood that each block in a flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0052] The above specific embodiments are detailed descriptions of the present invention. It cannot be determined that the specific embodiments of the present invention are limited only to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions and substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A method for unbalanced data regression, characterized in that: include: Construct a multi-granularity space generation model based on grid division to divide similar samples into the same granularity space; Oversampling is performed independently in each granularity space, where the oversampling range is defined by combining spatial importance and dynamic neighborhood radius.

2. The unbalanced data regression method according to claim 1, characterized in that: The construction of a multi-granularity space generation model based on grid division includes: Determine the division area of ​​the granular space according to the spatial area formed by the boundary range of each dimension of the sample; Subdivide the partitioned area into several small rectangles; According to the number of sparsely distributed samples in each small rectangle, the small rectangles are divided into different types; The distribution degree and the minimum distribution degree of the particle size space are calculated. If the distribution degree of the particle size space is less than the minimum distribution degree, the particle size space is further divided until the distribution degree of all particle size spaces is not less than the minimum distribution degree.

3. The unbalanced data regression method according to claim 2, characterized in that: The types of the small rectangles include a blank rectangle with a sample number of 0, a high-density rectangle with a maximum sample number, and a sparse rectangle between the high-density rectangle and the blank rectangle.

4. The unbalanced data regression method according to claim 3, characterized in that: If the distribution degree of the granular space is less than the minimum distribution degree, the granular space is further divided, including: For the granular space that needs to be further divided, select its center and the center of the sample in the high-density rectangle as the center points of two new granular spaces, calculate the distance from each sample to the two center points, and assign the sample to the granular space to which the nearest center point belongs to obtain two new granular spaces.

5. The unbalanced data regression method according to claim 1, characterized in that: The calculation of the spatial importance includes: Calculate the sample density of the sample continuous target value; Use the weighting function to assign weights to the sample density and obtain the density weight of each sample; The spatial importance is calculated based on the density weight of each sample.

6. The unbalanced data regression method according to claim 1, characterized in that: The oversampling range is defined by combining the spatial importance and the dynamic neighborhood radius, including: The amount of oversampling is determined by the following formula: ,in, represents the amount of oversampling, represents the neighborhood radius, Indicates the importance of space.

7. The unbalanced data regression method according to claim 6, characterized in that: The specific process of oversampling includes: when When , the samples with higher importance in the granularity space are used as candidate seed samples for oversampling; Select a seed sample, select a neighbor sample within its neighborhood radius, and perform linear interpolation to synthesize a new sample; The seed samples are selected cyclically to synthesize new samples until the number of synthesized new samples reaches the oversampling number.

8. An imbalanced data regression system, characterized in that: include: The multi-granularity space partitioning module is used to build a multi-granularity space generation model based on grid partitioning, dividing similar samples into the same granularity space; The oversampling module is used to perform oversampling independently in each granularity space, wherein the oversampling range is defined in combination with the spatial importance and the dynamic neighborhood radius.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, an unbalanced data regression method as described in any one of claims 1 to 7 is implemented.

10. A terminal comprising a memory and a processor, wherein the memory stores computer instructions that can be executed on the processor, characterized in that: When the processor runs the computer instructions, it executes the imbalanced data regression method described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • High-dimensional unbalanced data classification method based on feature learning and ensemble learning

    CN113344075A

  • Fuel cell fault data sampling method and system

    CN114595742A

  • Apparatus and method for clustering data in streaming clustering without reducing precision

    US20160357840A1

  • Method for classifying high-dimensional imbalanced data based on svm

    WO2019041629A1