Data enhancement method, device, equipment and storage medium
During the point cloud data enhancement process, the data distribution is adjusted according to the statistical classification proportion of objects in the sample database, and the problem of sample imbalance is solved and the detection accuracy of three-dimensional object detection is improved.
Patent Information
- Application Number
- CN202111581724.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-22
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2041-12-22
AI Technical Summary
In the prior art, in the three-dimensional object detection method based on point cloud data, in the data enhancement process, the uneven distribution of sample data leads to training underfitting or overfitting, affecting the detection accuracy.
By statistically stating the proportion of objects in the sample database under multiple statistical categories, the distribution of objects under different categories during the data enhancement process is adjusted to make them more balanced.
Improve the data enhancement effect, ensure that the distribution of training samples is more balanced, and improve the detection accuracy of three-dimensional object detection.
Smart Images

Figure CN114529778B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a data enhancement method, apparatus, device and storage medium. Background Art
[0002] Point cloud data refers to a collection of vectors in a three-dimensional coordinate system. Three-dimensional object detection based on point cloud data has great potential in vision applications such as autonomous driving and robotic navigation, and has become an active research topic in the field of three-dimensional computer vision. When training 3D object detection methods based on point cloud data, data augmentation is often performed on the training sample data to improve the accuracy of the 3D object detection method.
[0003] In related technologies, random sampling is used to enhance the training sample data. The original number of samples is still unbalanced after data enhancement, resulting in underfitting or overfitting of the training, and the effect of data enhancement is very poor. Summary of the Invention
[0004] The present invention provides a data enhancement method, device, equipment and storage medium, which improve the data enhancement effect.
[0005] In a first aspect, the present invention provides a data enhancement method, comprising:
[0006] Obtain a dataset to be enhanced and a sample database; the dataset to be enhanced includes point cloud data and label information of M types of objects, and the sample database includes point cloud data and label information of N types of objects, where M is a positive integer, N is an integer greater than 1, and N is greater than or equal to M;
[0007] Determining, based on the sample database, the proportion of each type of object in the N types of objects under multiple statistical categories; wherein the multiple statistical categories are determined based on parameters in the label information;
[0008] Determining target objects to be added to the data set to be enhanced and a target number of the target objects;
[0009] Target data is obtained from the sample database according to the proportion of the target objects in the multiple statistical categories; the target data includes point cloud data and label information of the target objects of the target number.
[0010] Optionally, the multiple statistical categories include:
[0011] Multiple statistical categories determined based on the distance between the object and the acquisition device in the forward direction of the object; and / or multiple statistical categories determined based on the height, occlusion degree and truncation degree of the object.
[0012] Optionally, obtaining target data from the sample database according to the proportions of the target objects in the multiple statistical categories includes:
[0013] Determining a sampling ratio of the target object in each statistical category according to the proportion of the target object in the multiple statistical categories;
[0014] Obtaining the number of samples of the target object in each statistical category according to the target number and the sampling ratios corresponding to the multiple statistical categories;
[0015] The target data is acquired according to the number of samples corresponding to the multiple statistical categories.
[0016] Optionally, determining the sampling ratio of the target object in each statistical category according to the proportion of the target object in the multiple statistical categories includes:
[0017] According to the proportion of the target object in the multiple statistical categories and the number of the multiple statistical categories, the sampling ratio of the target object in each of the statistical categories is determined.
[0018] Optionally, determining the sampling ratio of the target object in each statistical category according to the proportion of the target object in the multiple statistical categories and the number of the multiple statistical categories includes:
[0019] Using the formula Determining a sampling ratio of the target object under the statistical classification;
[0020] Among them, α represents the sampling ratio of the target object under the statistical classification, e represents a natural constant, k represents the proportion of the target object under the statistical classification, and n represents the number of the multiple statistical classifications.
[0021] Optionally, obtaining the number of samples of the target object in each statistical category according to the number of targets and the sampling ratios corresponding to the multiple statistical categories includes:
[0022] Using the formula Obtaining the number of samples of the target object under each of the statistical categories;
[0023] Wherein, number represents the number of samples of the target object under the statistical classification, round() represents the rounding function, N represents the number of targets, k represents the proportion of the target object under the statistical classification, n represents the number of the multiple statistical classifications, and α represents the sampling ratio of the target object under the statistical classification.
[0024] Optionally, the multiple statistical categories include I×J statistical categories, where I represents the number of first statistical categories determined according to the first rule, and J represents the number of second statistical categories determined according to the second rule, and each statistical category in the I×J statistical categories represents the second statistical category under the first statistical category.
[0025] Optionally, the sample database includes multiple frames of training sample sets, the data set to be enhanced is any one frame of training sample sets in the multiple frames of training sample sets, and the training sample set includes point cloud data and label information of at least one type of object.
[0026] In a second aspect, the present invention provides a data enhancement device, comprising:
[0027] An acquisition module is configured to acquire a dataset to be enhanced and a sample database; the dataset to be enhanced includes point cloud data and label information of M types of objects, and the sample database includes point cloud data and label information of N types of objects, where M is a positive integer and N is an integer greater than 1, and N is greater than or equal to M;
[0028] A statistical classification module, configured to determine, based on the sample database, the proportion of each of the N types of objects in a plurality of statistical classifications; wherein the plurality of statistical classifications are determined based on parameters in the label information;
[0029] a determination module, configured to determine a target object to be added to the dataset to be enhanced and a target number of the target objects;
[0030] A data enhancement module is used to obtain target data from the sample database according to the proportion of the target objects in the multiple statistical categories; the target data includes point cloud data and label information of the target objects of the target number.
[0031] In a third aspect, the present invention provides a data enhancement device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the data enhancement method provided by the present invention when executing the computer program.
[0032] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which implements the data enhancement method provided by the present invention when the computer program is executed by a processor.
[0033] The present invention provides a data enhancement method, apparatus, device and storage medium. By counting the proportions of objects in a sample database under multiple statistical categories, data enhancement is performed on a to-be-enhanced dataset according to different proportions, so that the distribution of objects added under different statistical categories is more balanced, thereby improving the data enhancement effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0035] Figure 1 A flow chart of a data enhancement method provided by an embodiment of the present invention;
[0036] Figure 2 An object distribution diagram of a sample database provided by an embodiment of the present invention;
[0037] Figure 3A A schematic diagram of the proportion of cars in multiple distance categories provided by an embodiment of the present invention;
[0038] Figure 3B A schematic diagram of the proportion of pedestrians in multiple distance categories provided by an embodiment of the present invention;
[0039] Figure 4A A schematic diagram of the proportion of cars in multiple difficulty classifications provided by an embodiment of the present invention;
[0040] Figure 4B A schematic diagram of the proportion of pedestrians in multiple difficulty levels provided by an embodiment of the present invention;
[0041] Figure 5 Another flow chart of the data enhancement method provided by an embodiment of the present invention;
[0042] Figure 6 A schematic structural diagram of a data enhancement device provided by an embodiment of the present invention;
[0043] Figure 7 A schematic diagram of the structure of a data enhancement device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0044] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0045] It will be understood that the terms "first", "second", "third", "fourth", etc. (if any) in the embodiments of the present application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0046] The data augmentation method, apparatus, device, and storage medium provided by the present invention are applicable to scenarios involving data augmentation of point cloud data. Data augmentation is a data expansion technique that maximizes the value of limited data. For example, when training 3D object detection based on point cloud data, the amount of manually annotated point cloud data is limited, while a larger amount of point cloud data is typically required to achieve good training results. Therefore, data augmentation can be performed based on the annotated point cloud data.
[0047] In related technologies, data enhancement is performed through random sampling. The proportion of random sampling is basically consistent with the proportion of existing data. For example. Suppose that the existing data includes sample point cloud data of 1,000 cars, of which there are 200 cars at close range, 700 cars at medium range, and 100 cars at long range, with the largest number of cars at medium range. If 10 cars are randomly selected from the sample point cloud data of 1,000 cars for data enhancement, then among the 10 cars, the number of cars at medium range will still be the largest, for example, 6, and the number of cars at close range and long range will still be smaller, for example, 3 cars at close range and 1 car at long range. The proportion of random sampling is basically the same as the proportion of existing data. In this way, by performing data enhancement through random sampling, the number of samples in the range with a small number of original samples will still be small after random sampling, and the model will easily underfit during training. Conversely, the number of samples in the range with a large number of original samples will be larger after random sampling, and the model will easily overfit during training, resulting in low detection accuracy and poor data enhancement effect.
[0048] In response to the above technical problems, the present invention provides a data enhancement method, which performs data enhancement based on the proportion of different types of objects in the existing data under different statistical classifications, so that the data volume of objects added under different statistical classifications is more balanced. For example, taking the sample point cloud data of 1,000 cars as an example, the proportion of medium-distance cars is the largest, and the proportion of close-range and long-distance cars is relatively small. Then, when performing data enhancement, the data enhancement ratio of medium-distance cars can be reduced, and the data enhancement ratio of close-range and long-distance cars can be increased. Assume that the data of 10 cars are selected for data enhancement. Through the data enhancement method provided by the present invention, among the 10 cars, there can be 3 cars each for close-range and long-distance cars, and 4 cars for medium distance. It can be seen that the data enhancement method provided by the present invention makes the data under different statistical classifications in the enhanced data more balanced, thereby improving the data enhancement effect.
[0049] For easier understanding, the point cloud data and label information are first explained.
[0050] A point cloud is a collection of numerous points that represent the surface characteristics of an object. Point cloud data can be collected using detection equipment such as LiDAR. Alternatively, LiDAR can be installed on vehicles, drones, and other equipment.
[0051] Optionally, the point cloud data may include point cloud coordinates x (front-back distance), y (left-right offset), z (up-down offset), and reflection intensity.
[0052] The point cloud data may have label information. The point cloud data and label information may be stored in a KITTI dataset format, for example, as a bin file.
[0053] For example, Table 1 shows the label information in the KITTI dataset format. As shown in Table 1, the label information can include the following parameters: object category (type), truncated state (truncated), occluded state (occluded), observation angle (alpha), two-dimensional size information (bbox), three-dimensional size information (dimensions), three-dimensional position information (location), three-dimensional angle (rotation_y), and detection confidence (score). The KITTI dataset format has 16 bits, and the number of bits corresponding to each parameter is shown in Table 1. For example, the parameter type occupies 1 bit, and the parameter bbox occupies 4 bits. The parameter type indicates the category to which the object belongs, for example, car (Car), van (Van), truck (Truck), pedestrian (Pedetrain), seat (Person_sitting), bicycle (Cyclist), tram (Tram), other (Misc), or unlabeled (Dontcare). For example, point cloud data collected when the target object is too far away from the lidar can be marked as Dontcare.
[0054] Table 1 Label information
[0055]
[0056] The technical solution of the present invention is described in detail below with reference to specific embodiments.
[0057] Figure 1 This is a flow chart of a data enhancement method provided by an embodiment of the present invention. The data enhancement method provided by this embodiment can be executed by a data enhancement device or a data enhancement equipment. Figure 1 As shown, the data enhancement method provided in this embodiment may include:
[0058] S101: Obtain a dataset to be augmented and a sample database. The dataset to be augmented includes point cloud data and label information of M types of objects, and the sample database includes point cloud data and label information of N types of objects, where M is a positive integer and N is an integer greater than 1, and N is greater than or equal to M.
[0059] This embodiment does not limit the values of M and N, and does not limit the number of each type of objects in the dataset to be enhanced and the sample database.
[0060] The following examples illustrate the sample database and the dataset to be enhanced.
[0061] Assume that N = 5 and M = 2. The N class of objects includes cars, vans, trucks, pedestrians, and bicycles. The M class of objects includes cars and pedestrians. The sample database may include point cloud data and label information for 1000 cars, 200 vans, 100 trucks, 600 pedestrians, and 200 bicycles. The dataset to be augmented may include point cloud data and label information for 10 cars and 3 pedestrians.
[0062] Optionally, the sample database may include a dataset to be enhanced.
[0063] Optionally, the sample database includes a multi-frame training sample set, the data set to be enhanced is any one frame training sample set in the multi-frame training sample set, and the training sample set includes point cloud data and label information of at least one type of object.
[0064] Optionally, an implementation method for obtaining the training sample set, the dataset to be enhanced, and the sample database may include:
[0065] Collect a frame of point cloud data through the lidar and obtain the label information of the frame of point cloud data.
[0066] Identify at least one target box and count the number of points within each target box.
[0067] For each target box, if the number of points in the target box is less than the preset threshold, the target box is deleted.
[0068] For each valid target frame, the point cloud data and label information within the valid target frame are stored as the point cloud data and label information of the object. The target frames remaining after deleting the target frames with a number of points less than a preset threshold from at least one target frame are all valid target frames.
[0069] The point cloud data and label information within all valid target frames are determined as a frame of training sample set.
[0070] Repeat the above process to obtain a multi-frame training sample set, and form a sample database with the point cloud data and label information in all valid target boxes in the multi-frame training sample set.
[0071] For example. Assume that the first frame training sample set includes 3 valid target frames, corresponding to 2 cars and 1 pedestrian. The first frame training sample set includes point cloud data and label information of 2 cars and 1 pedestrian. The second frame training sample set includes 10 valid target frames, corresponding to 5 cars, 3 pedestrians and 2 trucks. The second frame training sample set includes point cloud data and label information of 5 cars, 3 pedestrians and 2 trucks. And so on. Assume that a total of 3721 frames of training sample sets are obtained, then the sample database includes point cloud data and label information of all objects in the 3721 frames of training sample sets. For example, Figure 2 An object distribution diagram of a sample database provided by an embodiment of the present invention. Figure 2 As shown, the sample database includes point cloud data and label information for 2,207 pedestrians, 14,357 cars, 734 bicycles, 1,297 vans, 488 trucks, 224 trams, 56 seats, and 337 other objects. In this case, N = 8. The dataset to be augmented can be any frame in the 3,721-frame training sample set, for example, the second frame of the training sample set. In this case, M = 3.
[0072] S102: Determine the proportion of each type of object in N types of objects under multiple statistical categories based on the sample database.
[0073] The multiple statistical categories are determined based on the parameters in the tag information. This embodiment does not limit the number of the multiple statistical categories and the specific statistical category rules.
[0074] Optionally, in one implementation, the multiple statistical categories may include multiple statistical categories determined based on the distance between the object and the acquisition device in the forward direction of the object.
[0075] In this implementation, the number of objects of this type in the sample database is classified based on the distance between the objects.
[0076] For example, assume there are three statistical categories: distance < 20 meters, 20 meters ≤ distance ≤ 40 meters, and distance > 40 meters. The distance between the object and the acquisition device in the direction of the object's movement can be determined based on the 3D position information in the tag information. Optionally, the object's movement direction can be the X-axis direction in the camera coordinate system.
[0077] For the object "car" in the sample database, the percentage results are shown in Figure 3A .like Figure 3AAs shown in Figure 3, the sample database includes point cloud data and label information of 14,357 cars, of which 4,580 cars have a distance (represented by X) < 20 meters, accounting for 31.9%; 6,176 cars have a distance of 20 meters ≤ X ≤ 40 meters, accounting for 43.02%; and 3,601 cars have a distance of X > 40 meters, accounting for 25.08%.
[0078] For the object "pedestrian" in the sample database, the percentage results are shown in Figure 3B .like Figure 3B As shown in Figure 3, the sample database includes point cloud data and label information of 2207 pedestrians, of which 1359 pedestrians have a distance (represented by X) < 20 meters, accounting for 61.58%; 739 pedestrians have a distance of 20 meters ≤ X ≤ 40 meters, accounting for 33.48%; and 109 pedestrians have a distance of X > 40 meters, accounting for 4.94%.
[0079] It can be seen that the distribution of objects in different distance ranges is very uneven.
[0080] Optionally, in another implementation, the multiple statistical categories may include multiple statistical categories determined based on the height, occlusion degree, and truncation degree of the object.
[0081] In this implementation, the number of objects of this type in the sample database is categorized based on their height, degree of occlusion, and degree of truncation, representing the difficulty of object acquisition. Typically, when LiDAR and other detection equipment collect point cloud data, the higher the object's height, the less occlusion, and the fewer truncations, the more point cloud data corresponding to the object, the more complete the surface point cloud, and the easier it is to collect point cloud data. Conversely, the lower the object's height, the greater the occlusion, and the more truncations, the more difficult it is to collect point cloud data.
[0082] Let's take an example. Assume there are three statistical categories, represented by easy, medium, and difficult. The classification rules are shown in Table 2. The object's height, occlusion level, and truncation level can be determined based on the two-dimensional size information, truncation status, and occlusion status parameters in the label information. Optionally, the object's height can be expressed in pixels. For example, the object's height can be the difference between the upper and lower pixel coordinates of the target box in the two-dimensional size information.
[0083] Table 2 Difficulty classification
[0084] Difficulty Height of the object Degree of occlusion Degree of truncation Simple >40 <=0 <=0.15 medium >25 <=1 <=0.3 difficulty >25 <=2 <=0.5
[0085] For the object "car" in the sample database, the percentage results are shown in Figure 4A .like Figure 4AAs shown in Figure 1, the sample database includes point cloud data and label information of 14,357 cars, of which 29.3% are simply classified cars, 44.86% are medium classified cars, and 25.84% are difficult classified cars.
[0086] For the object "pedestrian" in the sample database, the percentage results are shown in Figure 4B .like Figure 4B As shown in Figure 1, the sample database includes point cloud data and label information of 2,207 pedestrians, of which 65.06% are pedestrians with simple classification, 29.09% are pedestrians with medium classification, and 5.85% are pedestrians with difficult classification.
[0087] It can be seen that the distribution of objects corresponding to different difficulty levels is very uneven.
[0088] Optionally, in yet another implementation, the multiple statistical categories may include multiple statistical categories determined based on the distance between the object and the acquisition device in the forward direction of the object and based on the height, occlusion degree, and truncation degree of the object.
[0089] In this implementation, the distance of the object and the difficulty of object collection are comprehensively considered, and multiple statistical classifications are more detailed, which can further improve the data enhancement effect.
[0090] Optionally, the multiple statistical categories may include I×J statistical categories, where I represents the number of first statistical categories determined according to the first rule, J represents the number of second statistical categories determined according to the second rule, and each statistical category in the I×J statistical categories represents the second statistical category under the first statistical category.
[0091] This embodiment does not limit the specific values of I and J.
[0092] For example, assume that the first rule classifies objects based on the distance between the object and the acquisition device in the direction of the object's travel, and the second rule classifies objects based on their height, degree of occlusion, and degree of truncation. I = 4, J = 3. The four first statistical categories are: distance < 10 meters, 10 meters ≤ distance < 20 meters, 20 meters ≤ distance ≤ 40 meters, and distance > 40 meters. See Table 2 for the three second statistical categories.
[0093] Then, multiple statistical categories can include 4×3=12, specifically including: simple classification within distance <10 meters, medium classification within distance <10 meters, difficult classification within distance <10 meters; simple classification within 10 meters≤distance <20 meters, medium classification within 10 meters≤distance <20 meters, difficult classification within 10 meters≤distance <20 meters; simple classification within 20 meters≤distance≤40 meters, medium classification within 20 meters≤distance≤40 meters, difficult classification within 20 meters≤distance≤40 meters; simple classification within distance >40 meters, medium classification within distance >40 meters, difficult classification within distance >40 meters.
[0094] S103: Determine the target objects to be added to the dataset to be enhanced and the target number of target objects.
[0095] This embodiment does not limit the type and number of target objects.
[0096] Optionally, in one implementation, the target object may be of the same type as the object included in the dataset to be enhanced. For example, taking the second frame training sample set as the dataset to be enhanced, the target object may include the objects in the second frame training sample set, that is, the target objects include cars, pedestrians, and trucks.
[0097] In this implementation, the target object is determined according to the data set to be enhanced, which effectively increases the data volume of various objects in the data set to be enhanced and improves the data enhancement effect.
[0098] Optionally, in another implementation, the target object may be a certain type of object in the dataset to be enhanced. In this implementation, data enhancement may be performed on the specific type of object in the dataset to be enhanced, thereby improving the data enhancement effect.
[0099] Optionally, in another implementation, the types of target objects may be greater than the types of objects in the dataset to be enhanced. In this implementation, the complexity of the dataset to be enhanced may be increased, thereby expanding the application scenarios.
[0100] Optionally, if the target objects include multiple types, the number of targets for different target objects can be the same or different. For example, taking the second frame training sample set as the dataset to be enhanced, the number of car targets can be 15, the number of pedestrian targets can be 10, and the number of truck targets can be 5.
[0101] S104: Obtain target data from the sample database based on the proportion of target objects in multiple statistical categories. The target data includes point cloud data and label information of the target objects.
[0102] Specifically, the proportion of target objects in multiple statistical categories is usually unbalanced. The target data is obtained from the sample database based on the proportion of target objects in multiple statistical categories. The impact of the uneven distribution of objects in the sample database on the data enhancement effect is taken into consideration, so that the distribution of objects added during the data enhancement process is as balanced as possible, thereby improving the effect of data enhancement.
[0103] Optionally, if there are multiple target objects, the statistical classification used for each target object can be the same or different.
[0104] For example, let's assume that the second frame training sample set is used as the dataset to be enhanced. The target objects include three categories: cars, pedestrians, and trucks. The number of car targets can be 15, the number of pedestrian targets can be 10, and the number of truck targets can be 5.
[0105] In one example, cars, pedestrians, and trucks all use multiple statistical categories determined based on the distance between the object and the collection device in the direction of the object's travel. For cars, the number of cars to be added to each distance category is determined based on the proportion of cars in the sample database at each distance category. Similarly, for pedestrians, the number of pedestrians to be added to each distance category is determined based on the proportion of pedestrians in the sample database at each distance category. For trucks, the number of trucks to be added to each distance category is determined based on the proportion of trucks in the sample database at each distance category.
[0106] In another example, cars and pedestrians use multiple statistical categories determined based on the distance between the object and the acquisition device in the direction of the object's advance, while trucks use multiple statistical categories determined based on the object's height, degree of occlusion, and degree of truncation. For cars, the number of cars that need to be added to each distance category is determined based on the proportion of cars in the sample database at each distance category. For pedestrians, the number of pedestrians that need to be added to each distance category is determined based on the proportion of pedestrians in the sample database at each distance category. For trucks, the number of trucks that need to be added to each difficulty category is determined based on the proportion of trucks in the sample database at each difficulty category.
[0107] It can be seen that this embodiment provides a data enhancement method, which counts the proportions of objects in the sample database under multiple statistical categories, and performs data enhancement on the enhanced data set according to different proportions, so that the data volume of objects added under different statistical categories is more balanced, thereby improving the data enhancement effect.
[0108] Based on the above embodiment, another embodiment of the present invention provides an implementation method of the data enhancement method, specifically providing an implementation method of obtaining target data from a sample database according to the proportion of the target object in multiple statistical categories in S104.
[0109] Figure 5 Another flow chart of the data enhancement method provided by the embodiment of the present invention. Figure 5 As shown, in S104, obtaining target data from a sample database according to the proportion of target objects in multiple statistical categories may include:
[0110] S501: Determine a sampling ratio of the target object in each statistical category according to the proportion of the target object in multiple statistical categories.
[0111] S502: Obtain the number of samples of the target object in each statistical category according to the number of targets and the sampling ratios corresponding to the multiple statistical categories.
[0112] S503: Obtain target data according to the number of samples corresponding to the multiple statistical categories.
[0113] In this embodiment, the sampling ratio of the target object in each statistical category is determined based on the target object's proportion in multiple statistical categories. The sampling ratio is related to the proportion; different proportions result in different sampling ratios. Therefore, considering the impact of the uneven distribution of objects in the sample database on the data enhancement effect, data enhancement is performed based on the different sampling ratios of the target object in multiple statistical categories. This ensures that the distribution of the objects added during the data enhancement process is as balanced as possible, thereby improving the data enhancement effect.
[0114] Optionally, in S501, determining the sampling ratio of the target object in each statistical category according to the proportion of the target object in multiple statistical categories may include:
[0115] According to the proportion of the target object in multiple statistical categories and the number of statistical categories, the sampling ratio of the target object in each statistical category is determined.
[0116] The sampling ratios for different statistical categories are determined by combining the number of statistical categories and the proportions of each category, which improves the rationality of determining the sampling ratios. For example, the greater the number of statistical categories, the more detailed the statistical classifications, and the more refined the sampling ratios for different statistical categories can be.
[0117] Optionally, determining the sampling ratio of the target object in each statistical category based on the proportion of the target object in multiple statistical categories and the number of statistical categories may include:
[0118] Using the formula Determine the sampling ratio of the target object under the statistical classification.
[0119] Among them, α represents the sampling ratio of the target object under the statistical classification, e represents the natural constant, k represents the proportion of the target object under the statistical classification, and n represents the number of multiple statistical classifications.
[0120] In this implementation, the sampling ratios under different statistical categories are determined by the average value (1 / n) of the proportion of each statistical category and the proportion k under different statistical categories, thereby improving the rationality of determining the sampling ratios.
[0121] Optionally, obtaining the number of samples of the target object in each statistical category based on the number of targets and the sampling ratios corresponding to the multiple statistical categories may include:
[0122] Using the formula Get the number of samples of the target object in each statistical category.
[0123] Among them, number represents the number of samples of the target object under the statistical classification, round() represents the rounding function, N represents the number of targets, k represents the proportion of the target object under the statistical classification, n represents the number of multiple statistical classifications, and α represents the sampling ratio of the target object under the statistical classification.
[0124] Optionally, the number of samples of the target object under multiple statistical categories can be determined by using the above formula for each statistical category.
[0125] Give an example.
[0126] For the first statistical classification, the number of samples is determined as:
[0127]
[0128] For the second statistical classification, the number of samples is determined as:
[0129]
[0130] And so on.
[0131] For the nth statistical classification, the number of samples is determined as:
[0132]
[0133] Optionally, to determine the number of samples of the target object under multiple statistical categories, the above formula can be used to determine the number of samples for each of the n-1 statistical categories, and then the number of samples for the last statistical category can be determined based on the target number N and the number of samples for the n-1 statistical categories.
[0134] Give an example.
[0135] For the first statistical classification, the number of samples is determined as:
[0136]
[0137] For the second statistical classification, the number of samples is determined as:
[0138]
[0139] And so on.
[0140] For the n-1th statistical classification, the number of samples is determined as:
[0141]
[0142] For the nth statistical classification, the number of samples is determined as:
[0143] number n =N-number1-number2-…-number n-1
[0144] Figure 6 This is a structural diagram of a data enhancement device provided by an embodiment of the present invention. The data enhancement device provided by this embodiment is used to execute the data enhancement method provided by the present invention. Figure 6 As shown, the data enhancement device provided in this embodiment may include:
[0145] The acquisition module 61 is configured to acquire a dataset to be enhanced and a sample database. The dataset to be enhanced includes point cloud data and label information of M types of objects, and the sample database includes point cloud data and label information of N types of objects, where M is a positive integer and N is an integer greater than 1, and N is greater than or equal to M.
[0146] The statistical classification module 62 is configured to determine the proportion of each type of object in the N types of objects under multiple statistical classifications based on the sample database. The multiple statistical classifications are determined based on parameters in the tag information.
[0147] The determination module 63 is configured to determine the target objects to be added to the dataset to be enhanced and the target number of the target objects.
[0148] The data enhancement module 64 is used to obtain target data from the sample database based on the proportion of target objects in multiple statistical categories. The target data includes point cloud data and label information of the target objects.
[0149] Optionally, the multiple statistical categories include:
[0150] Multiple statistical categories determined based on the distance between the object and the acquisition device in the forward direction of the object; and / or multiple statistical categories determined based on the height, occlusion degree and truncation degree of the object.
[0151] Optionally, the data enhancement module 64 is configured to:
[0152] Determining a sampling ratio of the target object in each statistical category according to the proportion of the target object in the multiple statistical categories;
[0153] Obtaining the number of samples of the target object in each statistical category according to the target number and the sampling ratios corresponding to the multiple statistical categories;
[0154] The target data is acquired according to the number of samples corresponding to the multiple statistical categories.
[0155] Optionally, the data enhancement module 64 is configured to:
[0156] According to the proportion of the target object in the multiple statistical categories and the number of the multiple statistical categories, the sampling ratio of the target object in each of the statistical categories is determined.
[0157] Optionally, the data enhancement module 64 is configured to:
[0158] Using the formula Determining a sampling ratio of the target object under the statistical classification;
[0159] Among them, α represents the sampling ratio of the target object under the statistical classification, e represents a natural constant, k represents the proportion of the target object under the statistical classification, and n represents the number of the multiple statistical classifications.
[0160] Optionally, the data enhancement module 64 is configured to:
[0161] Using the formula Obtaining the number of samples of the target object under each of the statistical categories;
[0162] Wherein, number represents the number of samples of the target object under the statistical classification, round() represents the rounding function, N represents the number of targets, k represents the proportion of the target object under the statistical classification, n represents the number of the multiple statistical classifications, and α represents the sampling ratio of the target object under the statistical classification.
[0163] Optionally, the multiple statistical categories include I×J statistical categories, where I represents the number of first statistical categories determined according to the first rule, and J represents the number of second statistical categories determined according to the second rule, and each statistical category in the I×J statistical categories represents the second statistical category under the first statistical category.
[0164] Optionally, the sample database includes multiple frames of training sample sets, the data set to be enhanced is any one frame of training sample sets in the multiple frames of training sample sets, and the training sample set includes point cloud data and label information of at least one type of object.
[0165] Figure 7 A structural diagram of a data enhancement device provided by an embodiment of the present invention. Figure 7 As shown, the data enhancement device provided in this embodiment may include a processor 702, a memory 704, and a communication interface 703 connected to a system bus 701. The processor 702 is used to provide computing and control capabilities. The memory 704 includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The communication interface 703 of the detection device is used to communicate with other devices. When the computer program is executed by the processor 702, the data enhancement method provided by the present invention is implemented.
[0166] Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the scheme of the present application, and does not constitute a limitation on the data enhancement device provided by the present application. The specific detection device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0167] It should be clear that the process of executing the computer program by the processor in the embodiment of the present application is consistent with the execution process of each step in the above method. For details, please refer to the description above.
[0168] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the data enhancement method provided in the above method embodiment of the present application can be implemented.
[0169] It should be clear that the process of executing the computer program by the processor in the embodiment of the present application is consistent with the execution process of each step in the above method. For details, please refer to the description above.
[0170] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct RAM bus dynamic RAM (DRDRAM), and RAM bus dynamic RAM (RDRAM), etc.
[0171] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0172] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. A data enhancement method, characterized in that: include: Obtain the dataset and sample database to be enhanced; The dataset to be enhanced includes point cloud data and label information of M types of objects, and the sample database includes point cloud data and label information of N types of objects, where M is a positive integer, N is an integer greater than 1, and N is greater than or equal to M; Determining, based on the sample database, the proportion of each type of object in the N types of objects under multiple statistical categories; wherein the multiple statistical categories are determined based on parameters in the label information; Determining target objects to be added to the data set to be enhanced and a target number of the target objects; Acquire target data from the sample database according to the proportion of the target objects in the multiple statistical categories; the target data includes point cloud data and label information of the target objects of the target number; The acquiring target data from the sample database according to the proportion of the target object in the multiple statistical categories includes: Using the formula Determining a sampling ratio of the target object under the statistical classification; Wherein, α represents the sampling ratio of the target object under the statistical classification, e represents a natural constant, k represents the proportion of the target object under the statistical classification, and n represents the number of the multiple statistical classifications; Obtaining the number of samples of the target object in each statistical category according to the target number and the sampling ratios corresponding to the multiple statistical categories; The target data is obtained according to the number of samples corresponding to the multiple statistical categories.
2. The method according to claim 1, characterized in that The plurality of statistical categories include: Multiple statistical categories determined based on the distance between the object and the acquisition device in the forward direction of the object; and / or multiple statistical categories determined based on the height, occlusion degree and truncation degree of the object.
3. The method according to claim 1, characterized in that The acquiring, based on the target number and the sampling ratios corresponding to the plurality of statistical categories, the sampling number of the target object in each statistical category includes: Using the formula Obtaining the number of samples of the target object under each of the statistical categories; Among them, number represents the number of samples of the target object under the statistical classification, round() represents the rounding function, and N represents the number of targets.
4. The method according to claim 1 or 2, characterized in that The multiple statistical classifications include I×J statistical classifications, where I represents the number of first statistical classifications determined according to the first rule, and J represents the number of second statistical classifications determined according to the second rule. Each of the I×J statistical classifications represents the second statistical classification under the first statistical classification.
5. The method according to claim 1 or 2, characterized in that The sample database includes multiple frames of training sample sets, the data set to be enhanced is any one frame of training sample sets in the multiple frames of training sample sets, and the training sample set includes point cloud data and label information of at least one type of object.
6. A data enhancement device, characterized in that: include: Acquisition module, used to obtain the dataset to be enhanced and the sample database; The dataset to be enhanced includes point cloud data and label information of M types of objects, and the sample database includes point cloud data and label information of N types of objects, where M is a positive integer, N is an integer greater than 1, and N is greater than or equal to M; A statistical classification module, configured to determine, based on the sample database, the proportion of each of the N types of objects in a plurality of statistical classifications; wherein the plurality of statistical classifications are determined based on parameters in the label information; a determination module, configured to determine a target object to be added to the dataset to be enhanced and a target number of the target objects; A data enhancement module is configured to obtain target data from the sample database according to the proportion of the target objects in the multiple statistical categories; the target data includes point cloud data and label information of the target objects of the target number; The data enhancement module is used to: Using the formula Determining a sampling ratio of the target object under the statistical classification; Wherein, α represents the sampling ratio of the target object under the statistical classification, e represents a natural constant, k represents the proportion of the target object under the statistical classification, and n represents the number of the multiple statistical classifications; Obtaining the number of samples of the target object in each statistical category according to the target number and the sampling ratios corresponding to the multiple statistical categories; The target data is obtained according to the number of samples corresponding to the multiple statistical categories.
7. A data enhancement device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 5 when executing the computer program.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Three-dimensional point cloud data instance segmentation method and system in automatic driving scene
CN111968133A
Data enhancement method and system for unbalanced text classification data
CN113076424A