Sampling method, device, electronic device and storage medium
In machine learning model training, the number of requirements is determined based on the difficulty factors of the sample and the number of regional clusters, and difficult cases are hierarchical sampling, which solves the problem of uneven sample sampling and improves the effect of model training.
Patent Information
- Application Number
- CN202210786592.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-04
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-07-04
AI Technical Summary
During the machine learning model training process, samples are prone to imbalance problems, resulting in insufficient model performance.
By obtaining sample sets of different types of samples, determining the difficulty factors of each sample, generating sample area clusters, determining the number of requirements of each type of samples based on the difficulty factors and the number of area clusters, and sampling them to achieve difficult-to-example hierarchical sampling.
The balance of sample sampling is achieved, the richness and representativeness of sampled samples are ensured, and the training effect of the target detection model is improved.
Smart Images

Figure CN115147593B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a sampling method, device, electronic device and storage medium. Background Art
[0002] With the development of computer vision technology, more and more fields have begun to use image processing and recognition technology based on target detection. In addition, with the rapid development of machine learning, it has become a trend to identify targets in pictures by training machine learning models. When training a model, it is necessary to make the model learn various samples, but in the process of sampling samples, it is easy to have an imbalance in the number of samples of different types, resulting in insufficient model performance.
[0003] Currently, technicians use hard example mining samplers or IoU (Intersection over Union) balanced samplers to make the sampling of samples more balanced. However, as the sampling increases, sample bias is prone to occur, resulting in a small number of samples of a certain type, which is not conducive to model training. Summary of the invention
[0004] The present application provides a sampling method, device, electronic device and storage medium to balance the number of samples of different types.
[0005] According to one aspect of the present application, a sampling method is provided, the method comprising:
[0006] Obtaining a sample set including samples of different types, and determining a difficulty factor of each sample in the sample set;
[0007] Generate different types of sample area clusters according to each difficulty factor;
[0008] Determine the required number of samples of each type based on the difficulty factor, the number of samples of different types, and the number of regional clusters of samples of different types;
[0009] Different types of samples are sampled according to the quantity required.
[0010] According to another aspect of the present application, a sampling device is provided, comprising:
[0011] A difficulty factor determination module, used to obtain a sample set including samples of different types and determine the difficulty factor of each sample in the sample set;
[0012] A region cluster determination module is used to generate different types of sample region clusters according to various difficulty factors;
[0013] A required quantity determination module is used to determine the required quantity of each type of samples according to the difficulty factor, the sample quantity of different types of samples and the number of different types of sample area clusters;
[0014] The sampling module is used to sample different types of samples according to the required quantity.
[0015] According to another aspect of the present application, an electronic device is provided, the electronic device comprising:
[0016] at least one processor; and
[0017] a memory communicatively connected to the at least one processor; wherein,
[0018] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the sampling method described in any embodiment of the present application.
[0019] According to another aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the sampling method described in any embodiment of the present application when executed.
[0020] The technical solution of the embodiment of the present application divides each sample in the sample set into regional clusters, determines the required number of samples to be collected, and collects samples from different regional clusters. This can achieve stratified sampling of difficult examples, so that the sampling results are more representative, ensures the richness of the sampled samples, and facilitates the training of the target detection model.
[0021] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0023] Figure 1 is a flow chart of a sampling method provided according to Embodiment 1 of the present application;
[0024] Figure 2 is a flow chart of a sampling method provided according to Embodiment 2 of the present application;
[0025] Figure 3 is a schematic diagram of a sampling sequence provided according to Embodiment 3 of the present application;
[0026] Figure 4 It is a structural schematic diagram of a sampling device provided according to the fourth embodiment of the present application;
[0027] Figure 5 It is a schematic diagram of the structure of an electronic device for implementing the sampling method of an embodiment of the present application. DETAILED DESCRIPTION
[0028] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this application.
[0029] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0030] Embodiment 1
[0031] Figure 1 A flowchart of a sampling method is provided for the first embodiment of the present application. This embodiment is applicable to the case of balancing the number of samples. The method can be performed by a sampling device, which can be implemented in the form of hardware and / or software, and can be configured in an electronic device. Figure 1 As shown, the method includes:
[0032] S110: Acquire a sample set including samples of different types, and determine the difficulty factor of each sample in the sample set.
[0033] Among them, the sample can be any sample needed to train a machine learning model (such as a target detection model, etc., which is not limited in this embodiment; wherein the target detection model includes detection tasks and classification tasks). It is understandable that the richness of sample types can help the machine learning model to be better trained and thus improve performance. Therefore, the number of samples of different types should be balanced to make the training of the model more balanced. In an optional embodiment, the sample can be an image recognition frame sample. The embodiments of the present application use the target detection recognition frame in the image processing technology as an example to illustrate the content of the technical solution of the present application, and therefore it cannot be understood as a limitation on the technical solution of the present application. The difficulty factor can be used to characterize the importance of a sample to model training, and the difficulty factor can be obtained according to any difficulty factor determination algorithm in the prior art.
[0034] Specifically, taking the target detection recognition box as an example, the IoU (Intersection over Union) of the candidate recognition box and the marked recognition box can be used to determine the negative samples used for target detection training to form a negative sample set (including at least two types of negative samples). All negative samples in the negative sample set can be passed through a preset classification and regression algorithm, and the classification loss of each negative sample can be calculated using a cross entropy loss function as the difficulty factor L.
[0035] S120 . Generate different types of sample region clusters according to each difficulty factor.
[0036] Among them, the sample region cluster can be a combination of samples with related and similar attributes used to train the model. Continuing with the previous example, in target detection, a negative sample can be a recognition box that does not include the target to be identified. Different recognition boxes appear in different positions on the image, and the image information included in different recognition boxes has different importance to model training (or other performance indicators). In target detection, recognition boxes with similar importance are combined into a region cluster. Samples that meet the same parameter value range can be divided into the same region cluster according to the parameter value of the difficulty factor. It should be noted that different types of samples can be assigned to different region clusters, that is, each sample in the same region cluster can belong to the same type.
[0037] S130: Determine the required number of samples of each type according to the difficulty factor, the number of samples of different types, and the number of regional clusters of samples of different types.
[0038] The required number of samples of each type may be the required number of samples needed to train the machine learning model. Samples of each type are selected in each regional cluster according to the difficulty factor corresponding to each sample and the number of samples of different types. A model may be pre-set to determine the required number, and the model inputs the difficulty factor of each sample, the number of samples of each type, and the number of regional clusters. The required number of samples of each type is output.
[0039] Continuing with the previous example, in target recognition, we can first calculate the difficulty coefficient of each negative sample type based on the difficulty factor of each negative sample recognition box and the number of negative samples of the same type, and then calculate the number of samples of each type that need to be sampled based on the difficulty coefficient of each type of negative sample and the number of each type of regional clusters.
[0040] S140. Sample different types of samples according to the required quantity.
[0041] According to the required number of samples of each type determined in the previous steps, a corresponding number of samples are selected from all the determined samples to be used for training the machine learning model.
[0042] In an optional implementation, sampling each type of sample according to the required quantity may include: determining a sampling order of different types of samples according to a difficulty factor; and sampling different types of samples according to the sampling order and the required quantity.
[0043] The sampling sequence may include the sampling sequence for each regional cluster, and may also include the sampling sequence for different samples in the same regional cluster. According to the sampling sequence, samples of a corresponding required number are collected in each type of regional cluster. The method of collecting a certain required number of samples according to the sampling sequence is conducive to balancing the number of samples of each type.
[0044] In an optional embodiment, determining the sampling order of each type of samples based on the difficulty factor may include: determining the difficulty factor mean of different types of sample area clusters based on the difficulty factor; determining the cluster sampling order of different types of sample area clusters based on the difficulty factor mean; determining the sample sampling order of samples in different types of sample area clusters based on the difficulty factor; determining the sampling order of different types of samples based on the cluster sampling order and the sample sampling order.
[0045] Among them, a regional cluster includes at least one sample, and the average value of the difficulty factor in the regional cluster is calculated according to the difficulty factor corresponding to each sample. The corresponding sampling order is arranged for different regional clusters according to the size of the average value, that is, the cluster sampling order. For example, the sampling order of the regional clusters is arranged from large to small according to the average value of the difficulty factor of the samples in each regional cluster, that is, the samples in the regional clusters with larger average difficulty factor are sampled first. Then, the samples are sampled from large to small according to the difficulty factor of each sample in the same regional cluster. That is, each regional cluster is sampled from large to small according to the mean value of the difficulty factor, and each sample in the same regional cluster is sampled from large to small according to the difficulty factor. Different types of samples are in different regional clusters, so they will not affect the sampling order and sampling quantity between each other.
[0046] In the above implementation, by determining the cluster sampling order of the regional cluster and the sample sampling order of each sample in the same regional cluster, the sampling method in the sample set is converted into stratified sampling in the form of regional clusters. Corresponding samples can be sampled in different types and different regional clusters, which improves the balance of samples of each type and provides richer and more balanced samples for the training of machine learning models.
[0047] In an optional embodiment, generating different types of sample area clusters according to each difficulty factor may include: selecting a current cluster center sample from the unclustered samples in the sample set according to the difficulty factor; selecting neighbor samples of the current cluster center sample from the unclustered samples of the same type as the current cluster center sample in the sample set; constructing a sample area cluster of the same type as the current cluster center sample using the selected neighbor samples and the current cluster center sample; and removing clustered samples from the sample set to update the sample set.
[0048] Among them, the clustered samples can be samples that have been assigned to the regional cluster in the sample set, and the remaining samples in the sample set at this time can all be unclustered samples. The cluster center sample can be the first sample of the regional cluster selected in the sample set when constructing the regional cluster. Based on this sample, a regional cluster is constructed with samples close to the cluster center sample according to preset rules. In particular, since the samples used for model training are essentially different, they can be samples with similar attributes or samples with similar actual distances. In the embodiment of the present application, target recognition is taken as an example. In the set of negative samples of the recognition box used for model training, a negative sample with the largest difficulty factor is selected as the cluster center sample of the regional cluster. The IoU of the remaining negative samples (i.e., unclustered samples) in the set of negative samples is calculated with the cluster center sample, and the negative samples with a preset IoU value greater than t and the cluster center sample are constructed as a regional cluster. The size of the t value can be determined by relevant technical personnel based on manual experience or a large number of experiments. For example, t=0.5 can be set, and the embodiment of the present application does not limit this. Of course, you can also directly set a distance threshold based on the distance between each remaining negative sample and the identification box of the cluster center negative sample, and all negative samples that meet the distance threshold and the cluster center sample form a regional cluster. All samples that constitute the regional cluster are removed from the sample set, and the sample with the largest difficulty factor is selected as the cluster center from the remaining samples. Repeat the above regional cluster construction operation until there are no unclustered samples in the sample set, that is, all samples belong to the corresponding regional cluster.
[0049] In the above implementation, regional clusters are constructed and all samples in the sample set are divided into respective regional clusters, which provides a basis for the subsequent collection of samples in different regional clusters. The different regional clusters divided represent different levels or types of all samples. The subsequent collection of samples in different regional clusters helps to provide samples of different levels or types for model training, expands the breadth and universality of sampling, and improves the richness of samples.
[0050] The technical solution of the embodiment of the present application divides each sample in the sample set into regional clusters, determines the required number of samples to be collected, and collects samples from different regional clusters. This can achieve stratified sampling of difficult examples, so that the sampling results are more representative, ensures the richness of the sampled samples, and facilitates the training of machine learning models.
[0051] Embodiment 2
[0052] Figure 2 This is a flow chart of a sampling method provided in Example 2 of the present application. Based on the above-mentioned embodiments, this embodiment further refines the operation of determining the required quantity. For the contents not detailed in this embodiment of the present application, please refer to other embodiments of the present application, and this embodiment of the present application will not be repeated here. Figure 2 As shown, the method includes:
[0053] S210: Obtain a sample set including samples of different types, and determine the difficulty factor of each sample in the sample set.
[0054] S220 . Generate different types of sample region clusters according to the difficulty factors.
[0055] S230. Determine the difficulty coefficient corresponding to each type according to the difficulty factor and the number of samples of different types.
[0056] Among them, the difficulty coefficient can be used to characterize the learning difficulty of different sample types when used for model training, and the difficulty coefficient can be used to calculate the number of samples required for different types. The difficulty coefficient only depends on the number of samples of a certain type and the difficulty factor of each sample of that type, and can be calculated using the following formula:
[0057]
[0058] Among them, H is the difficulty coefficient of a certain type of sample, n is the total number of samples of this type, a is the index of the sample, and L is the difficulty factor of each sample.
[0059] S240. Determine the required quantity according to the difficulty coefficient and the number of different types of sample area clusters.
[0060] Since different sample types have different difficulty coefficients, in order to better determine how many samples are needed for different types of samples, the required number is calculated by the difficulty coefficient and the number of different types of sample area clusters. For example, the required number can be calculated by a pre-trained required number determination model, which inputs the difficulty coefficient of each type of sample and the number of different types of sample area clusters, and outputs the required number corresponding to each type. For another example, it can be determined by the following formula:
[0061]
[0062] Among them, E is the sum of the number of samples of all types, E 1 is the required number of samples for the first type, C 1 is the number of sample region clusters of the first sample, H 1 is the difficulty coefficient of the first sample, C 2 is the number of sample area clusters of the second sample, H 2 is the difficulty coefficient of the second sample, and so on. n is the number of sample region clusters of the nth sample, H n is the difficulty coefficient of the nth sample.
[0063] In an optional implementation, determining the required number of samples of each type based on the difficulty coefficient and the number of sample area clusters of different types may include: determining the required number based on the difficulty coefficient and the number of sample area clusters of different types, and a preset minimum sampling ratio.
[0064] Among them, the minimum sampling ratio can be the minimum ratio of the required number of samples of a certain type to the total number of sample requirements. It is understandable that in order to ensure that the number of samples of various types used for model training remains balanced, each type of sample should be guaranteed to occupy a certain proportion of the total number of samples, so the minimum sampling ratio is set to ensure that the required number of samples of each type is not too small. For example, the minimum sampling ratio can be set to 10%. Of course, the minimum sampling ratio can be determined by relevant personnel based on manual experience or a large number of experiments, and the embodiments of the present application are not limited to this.
[0065] Continuing with the previous example, the required number of samples of a certain type can be calculated using the following formula:
[0066]
[0067] Among them, δ is the minimum sampling ratio. The larger one of δ and δ is selected to ensure the practicality of the minimum sampling ratio. In the above implementation, the minimum sampling ratio is set to ensure the diversity of sample types and the balance of the number of samples of each type, which is helpful for model training.
[0068] In an optional embodiment, different types of samples may include first-category samples and second-category samples, and determining the required quantity based on the difficulty coefficient and the number of regional clusters of different types of samples, and a preset minimum sampling ratio may include: determining the required quantity of first-category samples based on the difficulty coefficient of the first-category samples, the number of regional clusters of the first-category samples, and the preset minimum sampling ratio; and taking the difference between the total sample demand and the required quantity of the first-category samples as the required quantity of the second-category samples.
[0069] In this embodiment, there are only two different types of samples, namely, first type samples and second type samples. The required number of first type samples is first determined by the method of the above embodiment, and then the required number of second type samples is obtained by subtracting the required number of first type samples from the total number of sample requirements.
[0070] In a specific implementation, taking the sample of target detection as an example, in target detection, the target object on the image is identified. In order to enhance the accuracy of the model recognition sample, in addition to setting positive samples (recognition boxes with target objects) and negative samples (recognition boxes without target objects), the types of negative samples must also be diverse. In this implementation, the case of two negative samples is taken as an example, including foreground negative samples and background negative samples. The foreground negative sample can be a recognition box in an image area with a target object, but the recognition box does not include the target object; the background negative sample can be a recognition box in an image area without a target object. In the above steps, regional clusters of foreground negative samples and regional clusters of background negative samples are generated, and the difficulty coefficients of foreground negative samples and background negative samples are calculated. Then, according to the preset minimum sampling ratio, the required number of foreground negative samples and background negative samples is calculated, for example:
[0071]
[0072] Among them, E b is the number of background negative samples required, E is the total number of negative samples required, C b is the number of regional clusters of background negative samples, H b is the difficulty coefficient of background negative samples, C f is the number of regional clusters of foreground negative samples, H f is the difficulty coefficient of foreground negative samples, and δ is the minimum sampling ratio.
[0073] After determining the required number of background negative samples, calculate the required number of foreground negative samples:
[0074] E f =EE b ;
[0075] Among them, E f is the required number of foreground negative samples.
[0076] S250. Sample different types of samples according to the required quantity.
[0077] In the technical solution of the embodiment of the present application, the required number of samples of different types is determined by the difficulty coefficient and the number of regional clusters, which provides a practical and effective way to determine the sampling number of each type of samples and ensures the diversity, richness of sample types and the balance of the number of samples of each type.
[0078] Embodiment 3
[0079] The embodiment of the present application is a feasible preferred embodiment provided on the basis of the aforementioned embodiments. The embodiment of the present application takes the sample of target detection as an example, in which the target object on the image is identified. In order to enhance the accuracy of the model recognition sample, in addition to setting the positive sample (recognition frame with the target object) and the negative sample (recognition frame without the target object), the type of negative sample should also be diverse. In this embodiment, the case of two negative samples is taken as an example, including foreground negative samples and background negative samples. The foreground negative sample can be a recognition frame in an image area with a target object, but the recognition frame does not include the target object; the background negative sample can be a recognition frame in an image area without a target object. It can be understood that the foreground area can be an image area including the target object, and the background area can be an image area that does not include the target object. In actual situations, relevant technicians can splice the foreground area and the background area in the same image for sample recognition and collection.
[0080] The recognition frame is divided into a candidate recognition frame and annotated recognition frame. The candidate recognition frame may be a frame that is not labeled with the presence or absence of the target object, while the annotated recognition frame is a frame that has been manually annotated. First, the IoU of the candidate recognition frame and the annotated recognition frame is used to determine the negative sample set {N}. All candidate recognition frames in the set are classified and regressed, and the classification loss of each candidate recognition frame (i.e., each sample) is calculated using the cross entropy loss function as the difficulty factor L. The type is determined based on the area where the center point of the candidate recognition frame is located. If the center falls in the foreground, it is a foreground negative sample; if it falls in the background, it is a background negative sample, and a foreground set {N} is generated. f} and background set {N b}.
[0081] Then, respectively, {N f} and {N b} are sorted from large to small according to the difficulty factor; the negative sample with the largest difficulty factor is selected as the cluster center sample of a regional cluster, and the IoU with other samples in the set is calculated; the samples with IoU greater than the preset threshold t are added to the same regional cluster, and the already clustered samples are removed from {N f} and {N b}; continue to generate new regional clusters from the remaining samples in their respective sets in the above manner; until {N f} and {N b All negative samples in} are assigned to the corresponding regional clusters. At this time, the number of regional clusters recording foreground negative samples is C f , record the number of regional clusters of background negative samples as C b .
[0082] Calculate the difficulty coefficients of foreground negative samples and background negative samples respectively, for example:
[0083]
[0084] Among them, H f is the difficulty coefficient of the foreground negative sample, n is the number of foreground negative samples, i is the index of the foreground negative sample, L i is the difficulty factor of a foreground negative sample.
[0085]
[0086] Among them, H b is the difficulty coefficient of the background negative sample, m is the number of background negative samples, j is the index of the background negative sample, L j is the difficulty factor of a background negative sample.
[0087] According to the difficulty coefficient of foreground negative samples, the difficulty coefficient of background negative samples, the number of foreground sample area clusters and the number of background sample area clusters, the required number of samples is determined:
[0088]
[0089] Among them, E b is the number of background negative samples required, E is the total number of negative samples required, C b is the number of regional clusters of background negative samples, H b is the difficulty coefficient of background negative samples, C f is the number of regional clusters of foreground negative samples, H f is the difficulty coefficient of foreground negative samples, and δ is the minimum sampling ratio.
[0090] After determining the required number of background negative samples, calculate the required number of foreground negative samples:
[0091] E f =EE b ;
[0092] Among them, E f is the required number of foreground negative samples.
[0093] Finally, in the regional clusters of foreground negative samples and background negative samples, the negative samples in each regional cluster are sorted from large to small according to the mean difficulty factor, and the negative samples in the cluster are also sorted from large to small according to the difficulty factor; the samples with the largest difficulty factor in the regional clusters are taken out in order until all samples are taken out, and the sorted negative sample lists in the foreground and background are obtained; the top E in the foreground negative sample list are taken out f samples, take the first E samples from the background sample list b The two groups of samples are combined together to form the final negative sample sampling result.
[0094] like Figure 3 As shown in the figure, Fg is the foreground area and Bg is the background area. In Fg, Fa1, Fa2 and Fa3 are in the same area cluster (set as Fa area cluster), and Fb1 and Fb2 are in another area cluster (set as Fb area cluster). Among them, the difficulty factor of Fa1 is 0.56, the difficulty factor of Fa2 is 0.42, the difficulty factor of Fa3 is 0.35, the difficulty factor of Fb1 is 0.39, and the difficulty factor of Fb2 is 0.28.
[0095] According to the average values of the difficulty factors in the two region clusters, the average value of the difficulty factor of the Fa region cluster is greater than that of the Fb region cluster. Therefore, in the order of collecting region clusters, the Fa region cluster should be collected first with Fb (corresponding to the "cluster sampling order" in the above embodiment). In addition, in each region cluster, the collection order should be from large to small according to the difficulty factor of each negative sample (corresponding to the "sample sampling order" in the above embodiment), so the final sampling order should be Fa1, Fb1, Fa2, Fb2, Fa3, and the sampling method of background negative samples is the same. According to the above sampling method, the foreground negative samples are selected from E f , select the background negative samples E b They are finally combined into the total negative samples needed for model training.
[0096] The sampling results of this implementation cover different sample types, making the sample features used for model training richer and more balanced. Different regional clusters are sampled in a balanced manner, which increases the sampling range and helps model training.
[0097] Embodiment 4
[0098] Figure 4 This is a schematic diagram of the structure of a sampling device provided in Example 4 of the present application. Figure 4 As shown, the sampling device 400 includes:
[0099] A difficulty factor determination module 410 is used to obtain a sample set including samples of different types and determine the difficulty factor of each sample in the sample set;
[0100] A region cluster determination module 420, for generating different types of sample region clusters according to various difficulty factors;
[0101] A required quantity determination module 430 is used to determine the required quantity of samples of each type according to the difficulty factor, the sample quantity of different types of samples and the number of different types of sample region clusters;
[0102] The sampling module 440 is used to sample different types of samples according to the required quantity.
[0103] The technical solution of the embodiment of the present application divides each sample in the sample set into regional clusters, determines the required number of samples to be collected, and collects samples from different regional clusters. This can achieve stratified sampling of difficult examples, so that the sampling results are more representative, ensures the richness of the sampled samples, and facilitates the training of the target detection model.
[0104] In an optional implementation, the demand quantity determination module 430 may include:
[0105] A difficulty coefficient determination unit, used to determine the difficulty coefficient corresponding to each type according to the difficulty factor and the number of samples of different types;
[0106] The required quantity determination unit is used to determine the required quantity according to the difficulty coefficient and the quantity of different types of sample area clusters.
[0107] In an optional implementation, the required quantity determination unit is specifically used to determine the required quantity according to the difficulty coefficient and the number of different types of sample area clusters, and a preset minimum sampling ratio.
[0108] In an optional implementation manner, the different types of samples include first type of samples and second type of samples, and the required quantity determination unit may include:
[0109] A first category quantity determination subunit, used to determine the required quantity of first category samples according to the difficulty coefficient of the first category samples and the quantity of first category sample regional clusters, as well as a preset minimum sampling ratio;
[0110] The second category quantity determination subunit is used to take the difference between the total sample demand and the demand quantity of the first category samples as the demand quantity of the second category samples.
[0111] In an optional implementation, the sampling module 440 may include:
[0112] A sampling order determination unit, used to determine the sampling order of different types of samples according to the difficulty factor;
[0113] The sequential sampling unit is used to sample different types of samples according to the sampling sequence and required quantity.
[0114] In an optional implementation, the sampling order determination unit may include:
[0115] A difficulty factor mean value determination subunit is used to determine the difficulty factor means of different types of sample area clusters according to the difficulty factor;
[0116] A cluster sampling order determination subunit is used to determine the cluster sampling order of different types of sample area clusters according to the difficulty factor mean;
[0117] A sample sampling order determination subunit is used to determine the sample sampling order of samples in different types of sample area clusters according to the difficulty factor;
[0118] The sampling order determination subunit is used to determine the sampling order of different types of samples according to the cluster sampling order and the sample sampling order.
[0119] In an optional implementation, the region cluster determination module 420 may include:
[0120] A cluster center determination unit, used to select a current cluster center sample from the unclustered samples in the sample set according to a difficulty factor;
[0121] A cluster screening unit, used to select neighboring samples of the current cluster center sample from the unclustered samples of the same type as the current cluster center sample in the sample set;
[0122] A cluster construction unit, used to construct a sample region cluster of the same type as the current cluster center sample by combining the selected neighbor samples and the current cluster center sample;
[0123] The sample set updating unit is used to remove clustered samples from the sample set to update the sample set.
[0124] In an optional implementation, the sample may be an image recognition frame sample.
[0125] The sampling device provided in the embodiments of the present application can execute the sampling method provided in any embodiment of the present application, and has the corresponding functional modules and beneficial effects for executing each sampling method.
[0126] Embodiment 5
[0127] Figure 5 A schematic diagram of an electronic device 10 that can be used to implement an embodiment of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.
[0128] like Figure 5As shown, the electronic device 10 includes at least one processor 11, and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0129] A number of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0130] The processor 11 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs the various methods and processes described above, such as the sampling method.
[0131] In some embodiments, the sampling method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the sampling method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to perform the sampling method in any other suitable manner (e.g., by means of firmware).
[0132] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0133] The computer programs for implementing the methods of the present application may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer programs are executed by the processor, the functions / operations specified in the flow charts and / or block diagrams are implemented. The computer programs may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0134] In the context of the present application, a computer readable storage medium may be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, device or equipment. A computer readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer readable storage medium may be a machine readable signal medium. A more specific example of a machine readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0135] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0136] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0137] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.
[0138] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this application can be executed in parallel, sequentially or in different orders, as long as the expected results of the technical solution of this application can be achieved, and this document is not limited here.
[0139] The above specific implementations do not constitute a limitation on the protection scope of this application. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of this application should be included in the protection scope of this application.
Claims
1. A sampling method, characterized in that: The method comprises: Acquire a sample set including samples of different types, and determine a difficulty factor of each sample in the sample set; the difficulty factor represents the importance of the sample for model training; generating different types of sample region clusters according to the difficulty factors; Determine the required number of samples of each type according to the difficulty factor, the number of samples of different types and the number of regional clusters of samples of different types; the required number of samples is the required number of samples required for training the machine learning model; According to the required quantity, different types of samples are sampled; Generating different types of sample region clusters according to the difficulty factors includes: Selecting a current cluster center sample from the unclustered samples in the sample set according to the difficulty factor; selecting a neighboring sample of the current cluster center sample from the unclustered samples of the same type as the current cluster center sample in the sample set; constructing a sample region cluster of the same type as the current cluster center sample by combining the selected neighboring samples and the current cluster center sample; removing clustered samples from the sample set to update the sample set; Determining the required number of samples of each type according to the difficulty factor, the number of samples of different types and the number of regional clusters of samples of different types includes: Determine the difficulty coefficient corresponding to each type according to the difficulty factor and the number of samples of different types; the difficulty coefficient is used to characterize the learning difficulty of different sample types when used for model training; determine the required number according to the difficulty coefficient and the number of regional clusters of different types of samples; The sampling of each type of samples according to the required quantity includes: According to the difficulty factor, the sampling order of the different types of samples is determined; according to the sampling order and the required quantity, the different types of samples are sampled; the sampling order includes: the cluster sampling order of different types of sample area clusters and the sample sampling order of samples in different types of sample area clusters.
2. The method according to claim 1, characterized in that Determining the required number of samples of each type according to the difficulty coefficient and the number of sample area clusters of different types includes: The required quantity is determined based on the difficulty coefficient and the number of different types of sample area clusters, as well as a preset minimum sampling ratio.
3. The method according to claim 2, characterized in that The different types of samples include first type samples and second type samples, and the required number is determined according to the difficulty coefficient and the number of different type sample area clusters, and a preset minimum sampling ratio, including: Determining the required number of the first category of samples according to the difficulty coefficient of the first category of samples, the number of the first category of sample regional clusters, and a preset minimum sampling ratio; The difference between the total sample demand and the demand quantity of the first type of samples is used as the demand quantity of the second type of samples.
4. The method according to claim 1, characterized in that: The step of determining the sampling order of each type of samples according to the difficulty factor includes: Determining the difficulty factor means of the different types of sample region clusters according to the difficulty factor; Determining a cluster sampling order of the different types of sample region clusters according to the difficulty factor mean; Determining a sample sampling order of samples in different types of sample region clusters according to the difficulty factor; The sampling order of the different types of samples is determined according to the cluster sampling order and the sample sampling order.
5. A sampling device, characterized in that: include: A difficulty factor determination module, used to obtain a sample set including samples of different types, and determine the difficulty factor of each sample in the sample set; The difficulty factor represents the importance of the sample to the model training; A region cluster determination module, used for generating different types of sample region clusters according to the difficulty factors; A required quantity determination module, used to determine the required quantity of samples of each type according to the difficulty factor, the number of samples of different types and the number of regional clusters of different types of samples; the required quantity of samples is the required quantity of samples needed to train the machine learning model; A sampling module, used for sampling different types of samples according to the required quantity; The regional cluster determination module includes: A cluster center determination unit, configured to select a current cluster center sample from the unclustered samples in the sample set according to the difficulty factor; A cluster screening unit, configured to select neighboring samples of the current cluster center sample from the non-clustered samples of the same type as the current cluster center sample in the sample set; A cluster construction unit, used to construct a sample region cluster of the same type as the current cluster center sample by using the selected neighbor samples and the current cluster center sample; A sample set updating unit, used to remove clustered samples from the sample set to update the sample set; The demand quantity determination module includes: A difficulty coefficient determination unit, used to determine the difficulty coefficient corresponding to each type according to the difficulty factor and the number of samples of different types; the difficulty coefficient is used to characterize the learning difficulty of different sample types when used for model training; A required quantity determination unit, used for determining the required quantity according to the difficulty coefficient and the quantity of the different types of sample area clusters; The sampling module comprises: A sampling order determination unit, used to determine the sampling order of the different types of samples according to the difficulty factor; The sequential sampling unit is used to sample different types of samples according to the sampling sequence and the required quantity; the sampling sequence includes: the cluster sampling sequence of different types of sample area clusters and the sample sampling sequence of samples in different types of sample area clusters.
6. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the sampling method according to any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the sampling method according to any one of claims 1 to 4 when executed.
Citation Information
Patent Citations
Clustering-based adaptive weighted oversampling method
CN113378927A