System and method for training machine learning-based models

By evaluating and filtering the contribution value of data samples, omitting redundant or low-contribution data samples, forming a high-quality target data set, solving the problem of insufficient model performance reliability caused by data imbalance in the prior art, and achieving higher quality training data sets and more reliable model performance.

CN119948504APending Publication Date: 2025-05-06GENESIS CLOUD SERVICES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380069022.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-10
Filing Date
2023-10-10
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When the prior art deals with the data imbalance problem in the training data set, it is impossible to effectively evaluate the expected contribution of each data sample to the subsequent training process, resulting in insufficient reliability of model performance.

Method used

By training a first-level model based on machine learning, the contribution values ​​of each data sample to the training second-level model are calculated, and redundant or low-contributing data samples are omitted based on these contribution values, thus forming a high-quality target data set for training the second-level model.

Benefits of technology

Improves the quality and consistency of the training dataset and enhances the performance reliability of the trained machine learning model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948504A_ABST
    Figure CN119948504A_ABST
Patent Text Reader

Abstract

A system and method of training a machine learning (ML)-based model by at least one processor may include receiving an initial data set, the initial data set including a plurality of annotated data samples; training at least one ML-based first-level model to execute a first-level task based on the initial data set; calculating, based on training the at least one ML-based first-level model, at least one characteristic representing, for each data sample, a contribution value to training an ML-based second-level model to perform a second-level task; omitting a subset of data samples from the initial data set based on the at least one characteristic to obtain a target data set; and training an ML-based second-level model based on the target data set to execute the second-level task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates generally to machine learning and artificial intelligence. More specifically, the present invention relates to techniques for preparing training data sets to be used in machine learning (ML) processes. Background Art

[0002] As is known, the development of mathematical models that can learn from data and make predictions about the data is a general purpose of machine learning. Specifically, supervised and semi-supervised machine learning includes model training using a so-called "training data set" (or "supervised data set") and testing using an "inference data set". The term "training data set" generally refers to a set of annotated data samples, where the annotations provide the association of the data samples with multiple categories. In other words, the annotated data samples represent input and output vector (or scalar) pairs for the machine learning model. The model iteratively analyzes the data samples of the training data set to produce results, which are then compared to the target results (the corresponding annotations of the data samples in the training data set). Based on this comparison, the supervised learning algorithm determines the best combination of variables that will provide the highest prediction reliability. Finally, a well-trained model must show sufficiently reliable results when analyzing unknown data.

[0003] Therefore, the quality of the training dataset is reasonably considered a key aspect of machine learning.When exploring and analyzing training datasets for specific machine learning tasks, various problems may emerge and be considered to be solved.

[0004] For example, the uneven distribution of classes within a dataset, known as data imbalance, is considered one of the most common problems in classification machine learning tasks. When the data is highly imbalanced, the trained model will likely suffer from low prediction accuracy for the minority class. There are many methods known from the prior art that are intended to handle imbalanced classes, such as: random oversampling and undersampling, Cluster Centroid based Majority Under-sampling Technique (CCMUT), SMOTE (Synthetic Minority Oversampling Technique), applying higher penalties for misclassification in the minority class, and the like.

[0005] However, in addition to the problem of the number of training data sets, there is also a quality problem. As is known, different data samples in the same data set may have different degrees of influence on the training process. It is important to find an appropriate combination of data samples for the training data set in order to develop the ability of the model to correctly summarize and distinguish data. Therefore, when solving the data imbalance problem by reducing a certain amount of data samples from the majority category, it is key not to lose valuable data samples and to keep redundant data samples. However, existing undersampling techniques cannot consider the expected contribution of each data sample to the next training process in a comprehensive manner or not at all. Therefore, the ML-based model trained on the data set implemented by such methods lacks reliability, particularly in view of the contribution to the training process that the corresponding initial data set can potentially provide. Summary of the invention

[0006] Therefore, there is a need for a system and method for training ML-based models that will incorporate an improved process for balancing training datasets, thereby improving the quality of the training datasets and, therefore, increasing the reliability of the performance of the trained ML-based models. More specifically, there is a need to create a data balancing method that provides for the omission of data samples based on an assessment of their expected contribution to the subsequent training process.

[0007] In order to overcome the shortcomings of the prior art, the following invention is provided.

[0008] In a general aspect, the present invention may relate to a method for training a machine learning (ML)-based model by at least one processor, the method comprising: receiving an initial dataset comprising a plurality of annotated data samples; based on the initial dataset, training at least one ML-based first-level model to perform a first-level task; based on training the at least one ML-based first-level model, calculating at least one feature, the at least one feature representing, for each data sample, a contribution value to training an ML-based second-level model to perform a second-level task; based on the at least one feature, omitting a subset of data samples from the initial dataset to obtain a target dataset; and training the ML-based second-level model based on the target dataset to perform the second-level task.

[0009] In another general aspect, the present invention may relate to a system for training an ML-based model, the system comprising: a non-volatile memory device in which an instruction code module is stored; and at least one processor, the at least one processor being associated with the memory device and configured to execute the instruction code module, wherein when executing the instruction code module, the at least one processor is configured to: receive an initial data set, the initial data set comprising a plurality of data samples; based on the initial data set, train at least one ML-based first-level model to perform a first-level task; based on training the at least one ML-based first-level model, calculate at least one feature, the at least one feature representing, for each data sample, a contribution value to training an ML-based second-level model to perform a second-level task; based on the feature, omit a subset of data samples from the initial data set to obtain a target data set for training the ML-based second-level model to perform the second-level task.

[0010] In some embodiments, training the at least one ML-based first-level model, calculating the at least one feature, and omitting a subset of the data samples are performed iteratively.

[0011] In some embodiments, the annotated data sample includes annotations providing associations of the data sample with a plurality of categories; and the method further includes selecting at least one of the plurality of categories based on an amount of data samples associated with each category in the initial data set.

[0012] In some embodiments, the at least one processor is further configured to select at least one of the plurality of categories based on an amount of data samples associated with each category in the initial data set.

[0013] In some embodiments, training at least one ML-based first-level model includes: dividing the initial data set into a training part of data samples and an inference part of data samples; training the at least one ML-based first-level model based on the training part to perform the first-level task; based on the training of the at least one ML-based first-level model, inferring at least one trained ML-based first-level model on the inference part to perform the first-level task.

[0014] In some embodiments, the at least one processor is configured to train the at least one ML-based first-level model by: dividing the initial data set into a training part of data samples and an inference part of data samples; training the at least one ML-based first-level model based on the training part to perform the first-level task; based on the training of the at least one ML-based first-level model, inferring at least one trained ML-based first-level model on the inference part to perform the first-level task.

[0015] In some embodiments, the partitioning of the initial data set into a training portion of the data samples and an inference portion of the data samples is performed by having a random ratio of the data samples of each of the plurality of categories in the training portion or in the inference portion.

[0016] In some embodiments, the first level task includes classifying the data samples of the initial data set according to the multiple categories; the at least one characteristic includes a confidence value, which represents the relevance of the one or more data samples of the inferred part to their corresponding associated categories in the inferred result; and the method also includes selecting a subset of the data samples from the inferred part based on a set of omission conditions, the omission conditions including (i) the association of the selected data samples in the inferred result with at least one selected category, and (ii) the selection of the data samples based on the calculated confidence value.

[0017] In some embodiments, the at least one processor is further configured to select a subset of the data samples from the inference portion based on a set of omission conditions, wherein the omission conditions include (i) an association of the selected data samples with at least one selected category in the inferred result, and (ii) a selection of the data samples based on a calculated confidence value.

[0018] In some embodiments, the selection of the data sample based on the calculated confidence value includes selection of the data sample having (a) a confidence value with the highest score and (b) a confidence value exceeding a predefined threshold.

[0019] In some embodiments, the at least one processor is further configured to select the data sample having (a) a confidence value with the highest score and (b) a confidence value exceeding a predefined threshold based on the calculated confidence value.

[0020] In some alternative embodiments, the at least one ML-based first-level model includes a plurality of ML-based first-level models.

[0021] In some alternative embodiments, training at least one ML-based first-level model includes training multiple ML-based first-level models; and the selection of the data sample based on the calculated confidence value includes selecting the data sample among the multiple first-level models based on a function of the calculated confidence value of a particular data sample.

[0022] In some further alternative embodiments, the first-level task includes clustering at least one selected category, and forming a cluster set of the data samples associated with at least one selected category in the clustering result; the at least one characteristic includes the distance between the one or more data samples in the inferred result and the centroid of the cluster in the cluster set to which the one or more data samples belong; and the method also includes selecting a subset of the data samples from the inferred part based on a set of omission conditions, wherein the omission conditions include selecting the data samples based on the calculated distance.

[0023] In some alternative embodiments, the at least one processor is further configured to select a subset of the data samples from the inferred portion based on a set of omission conditions, the omission conditions comprising selecting the data samples based on the calculated distance.

[0024] In some embodiments, the selection of the data sample based on the calculated distance includes selecting the data sample having (a) the shortest distance to the centroid of the cluster to which the data sample belongs, and (b) a distance below a predefined threshold.

[0025] In some alternative embodiments, the selection of the data sample based on the calculated distance includes: determining, for at least one cluster, a range of distances corresponding to the densest distribution of the data samples in the cluster; and selecting at least one data sample having a distance from the determined range.

[0026] In some embodiments, the method further includes: calculating an omission percentage for at least one selected category, the omission percentage representing the percentage of data samples to be omitted from the initial data set; and selecting a subset of the data samples from the inferred portion based on a set of omission conditions, including selection of the data samples up to the calculated omission percentage.

[0027] In some embodiments, the second level task includes classification of data samples of the incoming data set according to the plurality of categories. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The subject matter which is regarded as the invention is particularly pointed out and distinctly claimed in the concluding portion of the specification. However, the invention as to its organization and method of operation, together with objects, features, and advantages thereof, may be best understood by reference to the following detailed description when read in connection with the accompanying drawings, in which:

[0029] Figure 1 is a block diagram depicting a computing device that may be included in a system for training an ML-based model according to some embodiments;

[0030] Figure 2is a block diagram depicting a system for training an ML-based model according to some embodiments;

[0031] FIG. 3A to FIG. 3C is a series of diagrams depicting examples of selecting and omitting subsets of data samples from an inference portion of a training data set according to some alternative embodiments; and

[0032] Figure 4 is a flowchart depicting a method of training an ML-based model according to some embodiments.

[0033] It should be understood that for simplicity and clarity of illustration, the elements shown in the figures are not necessarily drawn to scale. For example, for clarity, the size of some elements may be enlarged relative to other elements. In addition, where deemed appropriate, reference numerals may be repeated in the drawings to indicate corresponding or similar elements. DETAILED DESCRIPTION

[0034] Those skilled in the art will recognize that the present invention may be embodied in other specific forms without departing from the spirit or essential features of the present invention. Therefore, the foregoing embodiments are to be considered in all respects as illustrative rather than limiting the present invention as described herein. Therefore, the scope of the present invention is indicated by the appended claims rather than by the foregoing description, and therefore all changes within the meaning and scope of the claim equivalents are intended to be included therein.

[0035] In the detailed description below, many specific details are provided to provide a thorough understanding of the present invention. However, it will be appreciated by those skilled in the art that the present invention can be put into practice without these specific details. In other cases, well-known methods, procedures and parts are not described in detail to avoid blurring the present invention. Some features or elements described about an embodiment may be combined with features or elements described about other embodiments. For the sake of clarity, the discussion of the same or similar features or elements may not be repeated.

[0036] Although embodiments of the invention are not limited in this regard, discussions utilizing terms such as, for example, "process," "compute," "predict," "determine," "establish," "analyze," "examine," "select," "choose," "eliminate," "train," and the like may refer to operations and / or processes of a computer, computing platform, computing system, or other electronic computing device that manipulates and / or transforms data represented as physical (e.g., electronic) quantities within a computer's registers and / or memory into other data similarly represented as physical quantities within a computer's registers and / or memory or other information non-transitory storage media that may store instructions for performing the operations and / or processes.

[0037] Although the embodiments of the present invention are not limited in this regard, the term "plurality" as used herein may include, for example, "multiple" or "two or more". Throughout the specification, the term "plurality" may be used to describe two or more components, devices, elements, units, parameters, etc. The term "set" when used herein may include one or more items.

[0038] Unless explicitly stated, the method embodiments described herein are not limited to a particular order or sequence. In addition, some of the described method embodiments or elements thereof may occur or be performed simultaneously, at the same time point, in parallel, iteratively or repeatedly.

[0039] In some embodiments of the invention, the ML-based model may be an artificial neural network (ANN).

[0040] Neural network (NN) or artificial neural network (ANN), for example, a neural network that implements machine learning (ML) or artificial intelligence (AI) functions may refer to an information processing paradigm that may include nodes referred to as neurons organized into layers with links between neurons. Links may transmit signals between neurons and may be associated with weights. NN may be configured or trained for specific tasks, such as pattern recognition or classification. Training NN for specific tasks may involve adjusting these weights based on examples. Each neuron in an intermediate or final layer may receive an input signal, for example, a weighted sum of output signals from other neurons, and may process the input signal using a linear or nonlinear function (e.g., an activation function). The results of the input layer and intermediate layers may be passed to other neurons, and the results of the output layer may be provided as the output of the NN. Typically, neurons and links within a NN are represented by mathematical constructs, such as activation functions and matrices of data elements and weights. A processor (e.g., a CPU or a graphics processing unit (GPU)) or a dedicated hardware device may perform the relevant calculations.

[0041] It is obvious to those skilled in the art that various ML-based models can be implemented without departing from the essence of the present invention. It should also be understood that in some embodiments, the ML-based model can be a single ML-based model or a group (collection) of ML-based models that implement the same functionality as a single ML-based model as a whole. Therefore, considering the scope of the present invention, the above variations should be considered equivalent.

[0042] In some aspects, the following description of the claimed invention is provided with respect to the task of solving the problem of unbalanced data in a training data set. When annotated data sets have significantly unequal data sample distributions between specified categories, the unbalanced problem occurs. Such specific purposes and embodiments of the present invention related thereto are provided so as to make the description fully illustrative, and they are not intended to limit the scope of the claimed invention. It will be appreciated by those of ordinary skill in the art that the specific implementations of the claimed invention according to such tasks are provided as non-exclusive examples, and other actual specific implementations may be covered by the claimed invention, such as any specific implementation of the method of omitting data samples utilizing the claimed protection, regardless of whether the purpose of such omissions is to solve the problem of unbalanced data or a different purpose.

[0043] As is known, artificial intelligence and machine learning techniques are commonly implemented because they prove to be indispensable tools for solving tasks where the connections and dependencies between input data and target output data are complex and uncertain. Therefore, the proposed invention incorporates the use of ML techniques for evaluating the expected contribution of each data sample of an annotated data set into the subsequent training process. The claimed invention thus proposes a method for defining subsets of data samples in a training data set that are redundant to some extent and can therefore be reasonably omitted. Therefore, such incorporation of ML techniques into the process of preparing a training data set provides the realization of improved technical effects by improving the quality of the training data set and thereby increasing the reliability of the performance of the trained ML-based model.

[0044] Reference now Figure 1 , which is a block diagram depicting a computing device that may be included within an embodiment of a system for training an ML-based model according to some embodiments.

[0045] The computing device 1 may include a processor or controller 2, which may be, for example, a central processing unit (CPU) processor, a chip or any suitable computing device, an operating system 3, a memory device 4, an instruction code 5, a storage system 6, an input device 7, and an output device 8. The processor 2 (or one or more controllers or processors, possibly on multiple units or devices) may be configured to perform the methods described herein, and / or to perform or act as various modules, units, etc. More than one computing device 1 may be included in a system according to an embodiment of the present invention, and one or more computing devices 1 may act as a component of a system according to an embodiment of the present invention.

[0046] The operating system 3 may be or may include any code segment (e.g., a code segment similar to the instruction code 5 described herein) designed and / or configured to perform tasks involving coordination, scheduling, arbitration, supervision, control, or otherwise managing the operation of the computing device 1 (e.g., scheduling the execution of software programs or tasks or enabling software programs or other modules or units to communicate). The operating system 3 may be a commercial operating system. It should be noted that the operating system 3 may be an optional component, for example, in some embodiments, the system may include a computing device that does not require or include an operating system 3.

[0047] The memory device 4 may be or may include, for example, a random access memory (RAM), a read-only memory (ROM), a dynamic RAM (DRAM), a synchronous DRAM (SD-RAM), a double data rate (DDR) memory chip, a flash memory, a volatile memory, a non-volatile memory, a cache memory, a buffer, a short-term memory unit, a long-term memory unit, or other suitable memory unit or storage unit. The memory device 4 may be or may include a plurality of possibly different memory units. The memory device 4 may be a computer or processor non-transitory readable medium, or a computer non-transitory storage medium, such as RAM. In one embodiment, a non-transitory storage medium such as a memory device 4, a hard drive, another storage device, etc. may store instructions or codes that, when executed by a processor, may cause the processor to perform the method as described herein.

[0048] The instruction code 5 may be any executable code, such as an application, programming, process, task, or script. The instruction code 5 may be executed by the processor or controller 2 under the control of the operating system 3. For example, the instruction code 5 may be a stand-alone application or API module that may be configured to train an ML-based model, as further described herein. Although for clarity, in Figure 1 A single item of instruction code 5 is shown in the figure, but the system according to some embodiments of the present invention may include multiple executable code segments or modules similar to instruction code 5, which can be loaded into the memory device 4 and cause the processor 2 to execute the method described herein.

[0049] The storage system 6 may be or may include, for example, a flash memory, internal or embedded memory, a microcontroller or chip, a hard drive, a CD-Recordable (CD-R) drive, a Blu-ray Disc (BD), a Universal Serial Bus (USB) device, or other suitable removable and / or fixed storage unit as known in the art. Various types of input and output data may be stored in the storage system 6 and may be loaded from the storage system 6 into the memory device 4 where they may be processed by the processor or controller 2. In some embodiments, the storage system 6 may be omitted. Figure 1 For example, memory device 4 may be a non-volatile memory having the storage capacity of storage system 6. Thus, although shown as a separate component, storage system 6 may be embedded or included in memory device 4.

[0050] Input device 7 may be or may include any suitable input device, component or system, such as a detachable keyboard or keypad, a mouse, etc. Output device 8 may include one or more (possibly detachable) displays or monitors, speakers and / or any other suitable output device. Any applicable input / output (I / O) device may be connected to computing device 1, as shown in blocks 7 and 8. For example, a wired or wireless network interface card (NIC), a universal serial bus (USB) device, or an external hard drive may be included in input device 7 and / or output device 8. It should be understood that any suitable number of input devices 7 and output devices 8 may be operably connected to computing device 1, as shown in blocks 7 and 8.

[0051] Systems according to some embodiments of the present invention may include components such as, but not limited to, multiple central processing units (CPUs) or any other suitable multi-purpose or special-purpose processors or controllers (e.g., similar to element 2), multiple input units, multiple output units, multiple memory units, and multiple storage units.

[0052] Reference now Figure 2 , which depicts a system 10 for training an ML-based model according to some embodiments.

[0053] According to some embodiments of the present invention, system 10 may be implemented as a software module, a hardware module, or any combination thereof. For example, system 10 may be or may include a computing device such as Figure 1 In addition, the system 10 may be adapted to execute one or more instruction code modules (e.g., Figure 1 As described in further detail herein, the system 10 may be adapted to execute one or more instruction code modules (e.g., Figure 1 Element 5) in order to evaluate the expected contribution of each data sample of the annotated dataset to the subsequent training process, calculate at least one characteristic of the data sample, omit a subset of the data samples based on the characteristic, and train an ML-based model, etc.

[0054] like Figure 2 As shown, arrows may represent the flow of one or more data elements to and from system 10 and / or between modules or components of system 10. For clarity, Figure 2 Some arrows are omitted.

[0055] In some embodiments, the system 10 may be logically divided into two levels: a first level, which involves ML data preprocessing (e.g., solving the imbalanced data problem), and a second level, which involves training an ML-based target model (ML-based second level model 20) based on the preprocessed data of the first level.

[0056] In some embodiments, the first level of system 10 may include the following instruction code modules (e.g., Figure 1 The instruction code 5) of the computing device 1 shown: a data set analysis and division module 30, a first level training module 40 and an omission module 50.

[0057] The data set analysis and partitioning module 30 may be configured to request and receive an initial data set 61A extracted, for example, from a database 60A provided by a third-party data provider 60 .

[0058] Initial data set 61A may include multiple annotated data samples. Annotated data samples may include annotations that provide associations of data samples with multiple categories, as is typically done in supervised or semi-supervised machine learning practices. Initial data set 61A may have various structures and contents depending on the specific ML task to be solved (e.g., speech recognition, email spam detection, plant species classification, etc.). The objects of the claimed invention are not constrained by the specific structure or content of the annotated data sets, and since such annotated data sets are considered common aspects of machine learning, it will be clear to one of ordinary skill in the art what structure or content they may have.

[0059] In some embodiments, the data set analysis and partitioning module 30 may be further configured to analyze the received initial data set 61A to determine whether it has an unbalanced data problem, that is, whether it has a significantly unequal distribution of data samples between categories. The data set analysis and partitioning module 30 may be further configured to select at least one category (e.g., the majority category) from the plurality of categories based on the amount of data samples associated with each category in the initial data set 61A determined in the results of the analysis.

[0060] The data set analysis and partitioning module 30 may be further configured to calculate an omission percentage for at least one selected category (eg, the majority category). The omission percentage may represent the percentage of data samples that must be omitted from the initial data set 61A in order to balance the distribution of data samples between categories.

[0061] The data set analysis and partitioning module 30 may be further configured to generate an omission setting 30A including information about the selected at least one category and the calculated omission percentage.

[0062] The data set analysis and partitioning module 30 may be further configured to partition the initial data set 61A into a training portion 31A of data samples and an inference portion 31B of data samples. In some embodiments, the partitioning may be performed by having a random ratio of data samples of each category in a plurality of categories in the training portion 31A or in the inference portion 31B. In some alternative embodiments, the ratio may be predetermined. In yet other alternative embodiments, the data set analysis and partitioning module 30 may be further configured to calculate the ratio, for example, as a function of such ratio in the initial data set 61A (if considered as a whole). The ratio of the sizes of the training portion 31A and the inference portion 31B may also be different, for example, as commonly used, the training portion 31A is 70% and the inference portion 31B is 30%.

[0063] In some embodiments, the first level training module 40 may be configured to train at least one ML-based first level model (eg, first level model 70 ) based on the initial data set 61A to perform a first level task.

[0064] Therefore, in some embodiments, the first level training module 40 may be configured to receive the training part 31A and the inference part 31B. The first level training module 40 may be further configured to train at least one ML-based first level model (e.g., first level model 70) based on the training part (e.g., training part 31A) to perform the first level task. The first level training module 40 may be further configured to infer at least one trained ML-based first level model (e.g., first level model 70) on the inference part (e.g., inference part 31B) based on the training of at least one ML-based first level model (e.g., first level model 70) to perform the first level task.

[0065] In some embodiments, the first-level training module 40 may be further configured to calculate at least one characteristic based on training at least one ML-based first-level model (e.g., the first-level model 70), the at least one characteristic representing, for each data sample, a contribution value to the subsequent training of the ML-based second-level model (e.g., the second-level model 20) to perform the second-level task. Therefore, in the result of the inference, the first-level model 70 may be configured to output a characterized inference portion 70A of the data sample, the characterized inference portion including the data sample of the inference portion 31B accompanied by the value of the calculated characteristic.

[0066] In some embodiments, the omission module 50 may be configured to receive the characterized inference portion 70A and the omission setting 30A. The omission module 50 may be further configured to select a subset of data samples from the characterized inference portion 70A based on a set of omission conditions, including selection of data samples up to a calculated omission percentage and association of the selected data samples with at least one selected category (e.g., a majority category), as specified by the omission setting 30A.

[0067] The omission module 50 may be further configured to omit a subset of data samples from the characterized inference portion 70A of the initial data set 61A based on at least one characteristic specified in the characterized inference portion 70A. The omission module 50 may be further configured to obtain a target data set 51A, which is a portion of the initial data set 61A that remains after the omission. The obtained target data set 51A may also be used in a second level process of the system 10, as further described herein, or transferred to and used by a third party system.

[0068] In some embodiments, the system 10 may be configured to perform training of at least one ML-based first-level model (e.g., first-level model 70) by the first-level training module 40, calculation of at least one feature by the first-level training module 40, and iterative omission of a subset of data samples from the characterized inference portion 70A by the omission module 50. Depending on the specific embodiment, all three processes (the training, calculation, and omission) may be included in the iterative process together, or individually, or in any combination thereof. It should be understood that in such embodiments, any of the processes of training, calculation, and omission may provide intermediate results at intermediate iterations of the iterative process and provide final (target) results at the final iteration.

[0069] Specifically, the data set analysis and partitioning module 30 may be configured to divide the calculated omission percentage into intermediate omission percentages so as to distribute the omission process in proportion to the number of iterations and obtain the required number of data samples omitted as a whole. The data set analysis and partitioning module 30 may be further configured to generate an omission setting 30A, which includes information about the calculated omission percentage of the data samples to be omitted as a whole and the omission percentage of the data samples to be omitted at each iteration. The omission module 50 may be further configured to select a subset of data samples from the characterized inference portion 70A based on a set of omission conditions, including the selection of data samples up to the calculated intermediate omission percentage, as specified by the omission setting 30A. The omission module 50 may be further configured to omit a subset of data samples from the characterized inference portion 70A of the initial data set 61A according to each intermediate iteration and be configured to obtain an intermediate reduced inference portion 50A. In the next iteration, the data set analysis and partitioning module 30 may be further configured to receive the intermediate reduced inference part 50A of the previous iteration and replace the inference part 31B of the previous iteration with the intermediate reduced inference part 50A, thereby achieving an intermediate reduced initial data set (not specified in the figure).

[0070] The data set analysis and partitioning module 30 may be further configured to perform a random partitioning of the intermediate reduced initial data set into a training portion 31A of data samples and an inference portion 31B of data samples at each intermediate iteration. Thus, at each iteration, the first level model 70 is trained and inferred on corresponding portions of data samples having a combination of data samples that vary from one iteration to another. Thus, additional technical effects may be provided by enabling more reliable selection and omission of data samples for the initial data set 61A as a whole.

[0071] It should be understood that the iterative process is not limited to only combining the training, calculation and omission. The iterative process may therefore additionally include other processes related or unrelated to the purpose of the claimed invention.

[0072] In some embodiments, the second level of system 10 may include the following instruction code modules (e.g., Figure 1 The instruction code 5 of the computing device 1 shown: a data set partitioning module 80 and a second level training module 90.

[0073] In some embodiments, the dataset partitioning module 80 may be configured to receive the target dataset 51 A. The dataset partitioning module 80 may be further configured to partition the target dataset 51 A into a training portion 80A of data samples and an inference portion 80B of data samples.

[0074] In some embodiments, the second level training module 90 may be configured to train an ML-based second level model (eg, the second level model 20 ) based on the target dataset 51A to perform a second level task.

[0075] More specifically, the second level training module 90 may be configured to receive a training portion 80A and an inference portion 80B. The second level training module 90 may be further configured to train an ML-based second level model (e.g., second level model 20) based on the training portion (e.g., training portion 80A) of the target data set 51A to perform a second level task. The second level training module 90 may be further configured to infer a trained ML-based second level model (e.g., second level model 20) on an inference portion (e.g., inference portion 80B) of the target data set 51A based on the training of the ML-based second level model (e.g., second level model 20) to perform a second level task. Through the inference, the performance of the second level model (e.g., second level model 20) may be evaluated.

[0076] It should be clear to those skilled in the art that the present invention is neither limited to combining with a specific type of ML-based model (e.g., convolutional neural network (CNN), multi-layer perceptron, naive Bayes algorithm, support vector machine (SVM) algorithm, etc.), nor limited to combining with a specific known ML method (e.g., reinforcement learning, back propagation, stochastic gradient descent (SGD), etc.) nor limited to combining with a specific ML task (classification, regression, clustering, etc.).

[0077] The essence of the present invention can be generalized by the idea of ​​utilizing machine learning techniques to optimize annotated datasets for their further utilization in a subsequent machine learning process. More specifically, the essence of the present invention is based on the idea that training a specific ML-based model (e.g., the first-level model 70) based on a certain dataset (e.g., the initial dataset 61A) can provide specific information that characterizes the data samples of the dataset with respect to their expected contribution in the training of another ML-based model (e.g., the second-level model 20) or their usefulness for the training of the other ML-based model.

[0078] Thus, it should be understood that the combination of these models (e.g., the first level model 70 and the second level model 20) itself is essential for the purposes of the claimed invention rather than explicitly defining any particular implementation of each of the ML-based models or any of the ML-based models utilized. Thus, the particular implementation of the combination of ML-based models may vary from examples where the first level model 70 and the second level model 20 are the same extrapolation to examples where they are completely different or even use different sets of first level models 70.

[0079] However, in order to make this description sufficiently illustrative and non-abstract, by providing separate references FIG. 3A to FIG. 3C Three alternative non-exclusive embodiments are described in further detail to describe the invention.

[0080] FIG. 3A to FIG. 3C A series of diagrams are included that depict examples of selecting and omitting subsets of data samples from an inferred portion (e.g., inferred portion 31B) of an initial data set 61A according to respective alternative embodiments. The diagrams depict the distribution of data samples of the inferred portion 31B in a feature space defined by features A and B.

[0081] According to the annotations provided in the initial data set 61A, the data samples are divided into two categories: the first category 610A (or 610B or 610C, respectively) is marked by a triangle, and the second category 611A (or 611B or 611C, respectively) is marked by a circle. As can be seen, category 610A (or 610B or 610C, respectively) includes a significantly larger number of data samples than category 611A (or 611B or 611C, respectively), so it can be selected by the data analysis and partitioning module 30 as the majority category that needs to omit a certain percentage of data samples in order to solve the data imbalance problem.

[0082] refer to Figure 3A The first specific embodiment is further described.

[0083] According to a first specific embodiment, the first level model 70 can be a classification model, and the first level task can include classification of data samples of the initial data set 61A according to multiple categories (e.g., categories 610A and 611A). In the results of the first level model 70 training, the first level training module 40 can be configured to calculate a decision boundary 700A, which is illustrated in a series of diagrams. The decision boundary 700A can define the division of the feature space into categories, such as it can be completed by the trained first level model 70.

[0084] In this embodiment, the second level model 20 may also be a classification model. The second level task may include classification of data samples of an incoming data set according to multiple categories (eg, categories 610A and 611A).

[0085] In a first specific embodiment, the first level training module 40 may be configured to calculate the at least one characteristic, wherein the characteristic includes a confidence value that indicates the relevance of one or more data samples of the inference portion 31B to their corresponding associated categories (e.g., categories 610A and 611A) in the inference results of the first level model 70. With respect to the provided illustration, it should be understood that the farther a particular data sample is from the decision boundary 700A, the higher the confidence value it receives in the inferred results.

[0086] In a first specific embodiment, the omission module 50 may be further configured to select a subset of data samples (e.g., subset 612A) from the inference portion 31B based on a set of omission conditions, the set of omission conditions including: an association of the selected data sample with at least one selected category (e.g., category 610A, as specified by the omission setting 30A) in the inferred result; and a selection of the data sample based on the calculated confidence value. More specifically, the omission module 50 may be further configured to select data samples having (a) the highest scoring confidence value and (b) a confidence value exceeding a predefined threshold (e.g., threshold 701A illustrated by a line located at a specific distance from the decision boundary 700A). In some embodiments, the data set analysis and partitioning module 30 may be configured to predefine (manually or automatically) the threshold 701A and include information about the threshold in the omission setting 30A.

[0087] Thus, as can be seen in this series of diagrams, the omission module 50 selects a subset 612A that includes three data samples that are selected because: they belong to the selected majority class 610A; are located farther relative to the decision boundary 700A than the other data samples; and are above the threshold 701A. The last diagram in the series represents the contents of the reduced inference portion formed as a result of the omission of the subset 612A by the omission module 50. The reduced inference portion is also included in the target data set 51A.

[0088] As can be seen, this particular embodiment may assume that data samples having the highest confidence values ​​calculated during training of a particular classification model (e.g., first level model 70) have an insignificant impact on the ability of another classification model (e.g., second level model 20) to classify incoming data samples by category (e.g., categories 610A and 611A).

[0089] refer to Figure 3B The second specific embodiment is further described.

[0090] According to the second specific embodiment, the first level model 70 can be a clustering model. The first level task may include clustering of at least one selected category (e.g., majority category 610B), and forming a cluster 700B set of data samples associated with at least one selected category (e.g., category 610B) in the clustering result. The first level training module 40 can be configured to calculate the position of the centroid 701B of the cluster 700B in the feature space in the inferred result of the first level model 70.

[0091] In this embodiment, the second level model 20 may be a classification model. The second level task may include classification of data samples of an incoming data set according to a plurality of categories (eg, categories 610B and 611B).

[0092] In a second specific embodiment, the first level training module 40 may be configured to calculate the at least one characteristic, wherein the characteristic includes the distance between one or more data samples of the inferred part 31B in the inference result of the first level model 70 and the centroid of the cluster in the cluster set to which the one or more data samples belong (for example, the centroid 701B of the corresponding cluster 700B).

[0093] In a second specific embodiment, the omission module 50 may be further configured to select a subset of data samples from the inference portion 31B based on a set of omission conditions, including the selection of data samples based on the calculated distance. More specifically, the omission module 50 may be further configured to select data samples with (a) the shortest distance to the centroid of the cluster to which the data sample belongs, (b) a distance below a predefined threshold (not shown in the figure). In some embodiments, the data set analysis and partitioning module 30 may be configured to predefine (manually or automatically) the threshold and include information about the threshold into the omission setting 30A. In the illustrated example, the selected data sample is marked by a contour line with a dashed type.

[0094] Thus, as can be seen in the series of diagrams, the omission module 50 selects a subset comprising four data samples (one data sample from each cluster 700B) that are selected because: they are located closer to the respective centroids 701B of the clusters 700B than the other data samples; and the respective distances to the centroids 701B are below a predefined threshold. The last diagram in the series represents the content of the reduced inference portion achieved in the result of the omission of the subset by the omission module 50. The reduced inference portion is also included in the target data set 51A.

[0095] As can be seen, this particular embodiment may assume that the data sample having the shortest distance to the centroid 701B of the cluster 700B to which the data sample belongs is the most easily defined member of the corresponding cluster and therefore the overall majority class (e.g., class 610A). Therefore, such data samples have an insignificant impact on the ability of the trained classification model (e.g., the second level model 20) to classify the incoming data samples by class (e.g., classes 610A and 611A). As in Figure 3B As can be seen in the last diagram in the series of , the reduced inference portion has a more uniform distribution of data samples in the feature space than the inference portion 31B.

[0096] In yet another alternative embodiment, the omission module 50 may be configured to determine, for at least one cluster (e.g., for at least one cluster 700B), a range of distances to the centroid 701B of the corresponding cluster corresponding to the densest distribution of data samples in the corresponding cluster 700B. The omission module 50 may be further configured to select at least one data sample having a distance from the determined range. In other words, in such embodiments, the selection of data samples for further omission is performed from the region of the feature space having the densest distribution of data samples.

[0097] This particular embodiment may assume that the density of the data sample distribution in the feature space does not contribute significantly to the ability of a training classification model (e.g., the second level model 20) to classify incoming data samples by category (e.g., categories 610B and 611B). Therefore, with respect to the majority category (e.g., category 610B), reducing the density but maintaining the diversity of the data sample distribution in the feature space may be considered an effective way to balance the data.

[0098] Reference includes a series of diagrams Figure 3C Describing a third specific embodiment, the figure depicts a third example of selecting and omitting a subset of data samples from the inferred portion 31B of the initial data set 61A.

[0099] According to a third specific embodiment, the system 10 includes multiple first level models, for example three different first level models 70. The first level training module 40 can be further configured to train multiple ML-based first level models 70. A series of diagrams for each specific first level model 70 are provided in separate rows.

[0100] According to a third specific embodiment, the first level model 70 can be a classification model, and the first level task can include classification of data samples of the initial data set 61A according to multiple categories (e.g., categories 610C and 611C). The first level training module 40 can be further configured to calculate decision boundaries 701C, 702C, and 703C, respectively, in the results of training the corresponding first level model 70. The decision boundaries 701C, 702C, and 703C define the division of the feature space into categories, as it can be completed by the corresponding first level model 70 in the results of training.

[0101] In this embodiment, the second level model 20 may also be a classification model. The second level tasks may include classification of data samples of an incoming data set according to multiple categories (eg, categories 610C and 611C).

[0102] In a third specific embodiment, the first level training module 40 may be configured to calculate the at least one characteristic, wherein the characteristic includes a confidence value representing the relevance of one or more data samples of the inference part 31B to their corresponding associated categories (e.g., categories 610C and 611C) in the inference results of each first level model 70.

[0103] It should be appreciated with respect to the illustrations provided that the further a particular data sample is from the corresponding decision boundary 701C, 702C, or 703C, the higher the confidence value it receives in the inferred result.

[0104] In a third specific embodiment, with respect to each first level model 70, the omission module 50 may be further configured to select a subset of data samples (e.g., subsets 612C, 613C, and 614C, respectively) from the inference portion 31B based on a set of omission conditions. The set of omission conditions may include an association of the selected data sample with at least one selected category (e.g., the majority category 610C as specified by the omission setting 30A) in the inferred result; and a selection of data samples based on the calculated confidence value. More specifically, with respect to each first level model 70, the omission module 50 may be further configured to select data samples having (a) the highest scoring confidence value and (b) a confidence value exceeding a predefined threshold (e.g., thresholds 704C, 705C, and 706C illustrated by the corresponding line at a specific distance from the corresponding decision boundary 701C, 702C, or 703C). In some embodiments, the data set analysis and partitioning module 30 may be configured to predefine (manually or automatically) the thresholds 704C, 705C, and 706C and include information about the thresholds into the omission settings 30A.

[0105] The omission module 50 may be further configured to complete the selection of a subset of data samples from the inference portion 31B by selecting data samples based on a function of the calculated confidence values ​​of specific data samples in the plurality of first level models 70. For example, the omission module 50 may apply a logical AND to subsets 612C, 613C, and 614C and obtain a resulting subset (not shown in the figure), which includes only those data samples that appear in each of the subsets 612C, 613C, and 614C. It should be understood that the claimed invention is not limited to using only a logical AND as the function. Any other suitable function may be used to determine the resulting subset of data samples for further omission.

[0106] Thus, the omission module 50 selects a resulting subset that includes two data samples that are selected because: they belong to the selected majority class 610C; are located farther relative to the decision boundaries 701C, 702C, and 703C than the other data samples; are located above the respective thresholds 704C, 705C, and 706C; and are included in each of the selected subsets 612C, 613C, and 614C. The last diagram in the series represents the contents of the reduced initial portion achieved in the result of the omission of the resulting subset by the omission module 50. The reduced inferred portion is also included in the target data set 51A.

[0107] As can be seen, this particular embodiment may assume that data samples having the highest confidence values ​​calculated during training of multiple classification models (e.g., first level model 70) have an insignificant impact on the ability of another classification model (e.g., second level model 20) to classify incoming data samples by category (e.g., categories 610C and 611C).

[0108] Reference now Figure 4 , presents a flowchart depicting a method for training an ML-based model by at least one processor according to some embodiments.

[0109] As shown in step S1005, at least one processor (eg, Figure 1 The processor 2 of the embodiment of the present invention may receive an initial data set (e.g., the initial data set 61A), the initial data set comprising a plurality of annotated data samples, wherein the annotated data samples include annotations that provide associations between the data samples and a plurality of categories (e.g., categories 610A and 611A). Step S1005 may be implemented by the data set analysis and segmentation module 30 (as shown in FIG. 1 ). Figure 2 and FIG. 3A to FIG. 3C described).

[0110] As shown in step S1010, at least one processor (eg, Figure 1The processor 2 of the embodiment may perform the selection of at least one category from a plurality of categories (e.g., categories 610A and 611A) based on the amount of data samples associated with each category in the initial data set (e.g., initial data set 61A). Step S1010 may be implemented by the data set analysis and segmentation module 30 (e.g., reference 1). Figure 2 and FIG. 3A to FIG. 3C described).

[0111] As shown in step S1015, at least one processor (eg, Figure 1 The processor 2 of the embodiment may be executed to calculate the omission percentage for the selected at least one category (e.g., category 610A), which indicates the percentage of data samples to be omitted from the initial data set (e.g., initial data set 61A). Step S1015 may be implemented by the data set analysis and segmentation module 30 (e.g., referring to Figure 2 and FIG. 3A to FIG. 3C described).

[0112] As shown in step S1020, at least one processor (eg, Figure 1 The processor 2 of the processor 2 may perform the division of the initial data set (e.g., the initial data set 61A) into a training portion of data samples (e.g., the training portion 31A) and an inference portion of data samples (e.g., the inference portion 31B), wherein the division is performed by having a random ratio of data samples of each of a plurality of categories (e.g., categories 610A and 611A) in the training portion (e.g., the training portion 31A) or in the inference portion (e.g., the inference portion 31B). Step S1020 may be implemented by the data set analysis and division module 30 (as shown in reference Figure 2 and FIG. 3A to FIG. 3C described).

[0113] As shown in step S1025, at least one processor (eg, Figure 1 The processor 2 of the embodiment may perform training of at least one ML-based first-level model (e.g., first-level model 70) based on the training part (e.g., training part 31A) to perform the first-level task. Step S1025 may be implemented by the first-level training module 40 (e.g., referring to Figure 2 and FIG. 3A to FIG. 3C described).

[0114] As shown in step S1030, at least one processor (eg, Figure 1The processor 2) may perform inference of at least one trained ML-based first-level model (e.g., first-level model 70) on the inference part (e.g., inference part 31B) based on the training of at least one ML-based first-level model (e.g., first-level model 70) to perform the first-level task. Step S1030 may be implemented by the first-level training module 40 (as shown in FIG. 1 ). Figure 2 and FIG. 3A to FIG. 3C described).

[0115] As shown in step S1035, at least one processor (eg, Figure 1 The processor 2 of the embodiment of the present invention may perform calculation of at least one characteristic, which represents, for each data sample, a contribution value to training a second-level model based on ML (e.g., the second-level model 20) to perform the second-level task. Step S1035 may be implemented by the first-level training module 40 (e.g., referring to Figure 2 and FIG. 3A to FIG. 3C described).

[0116] As shown in step S1040, at least one processor (eg, Figure 1 The processor 2 of the embodiment of the present invention may perform the omission of a subset of data samples (e.g., subset 612A) from the initial data set based on at least one characteristic to obtain a target data set (e.g., target data set 51A), wherein the subset of data samples (e.g., subset 612A) is selected from the inference part (e.g., characterized inference part 70A) based on a set of omission conditions, the set of omission conditions including: the association of the selected data samples with at least one selected category (e.g., category 610A) in the inference result; and the selection of data samples up to the calculated omission percentage. Step S1040 may be implemented by the omission module 50 (as shown in reference to Figure 2 and FIG. 3A to FIG. 3C described).

[0117] As shown in step S1045, at least one processor (eg, Figure 1 The processor 2 of the embodiment may perform training of the second level model (e.g., the second level model 20) based on the ML based on the target data set (e.g., the target data set 51A) to perform the second level task. Step S1045 may be implemented by the second level training module 90 (e.g., referring to Figure 2 and FIG. 3A to FIG. 3C described).

[0118] As can be seen from the provided description, the claimed invention represents a system and method for training an ML-based model that incorporates an improved process for balancing a training dataset, thereby improving the quality of the training dataset and, therefore, increasing the reliability of the trained ML-based model performance. More specifically, the claimed invention includes a data balancing method that provides for the omission of data samples based on an assessment of their expected contribution to the subsequent training process.

[0119] Unless explicitly stated, the method embodiments described herein are not limited to a particular order or sequence. In addition, all formulas described herein are intended to be examples only, and other or different formulas may be used. In addition, some of the described method embodiments or their elements may occur or be executed at the same time point.

[0120] Although certain features of the present invention have been illustrated and described herein, many modifications, substitutions, changes and equivalents may occur to those skilled in the art. It should therefore be understood that the appended claims are intended to cover all such modifications and changes that fall within the true spirit of the invention.

[0121] Various embodiments have been presented. Each of these embodiments may of course include features from other embodiments presented, and embodiments not specifically described may include various features described herein.

Claims

1. A method for training a machine learning (ML) based model by at least one processor, the method comprising: receiving an initial dataset comprising a plurality of annotated data samples; Based on the initial data set, training at least one first-level ML-based model to perform a first-level task; Based on training the at least one ML-based first-level model, calculating at least one feature, the at least one feature representing, for each data sample, a contribution value to training the ML-based second-level model to perform the second-level task; omitting a subset of data samples from the initial data set based on the at least one characteristic to obtain a target data set; as well as The ML-based second-level model is trained based on the target dataset to perform the second-level task. 2 . The method of claim 1 , wherein training the at least one ML-based first-level model, calculating the at least one feature, and omitting a subset of the data samples are performed iteratively.

3. The method of claim 1 , wherein the annotated data samples include annotations providing associations of the data samples with a plurality of categories, and wherein the method further comprises selecting at least one of the plurality of categories based on an amount of data samples associated with each category in the initial data set.

4. The method of claim 3, wherein training the at least one ML-based first level model comprises: Dividing the initial data set into a training portion of data samples and an inference portion of data samples; training the at least one ML-based first-level model based on the training portion to perform the first-level task; Based on the training of the at least one ML-based first-level model, at least one trained ML-based first-level model is inferred on the inference portion to perform the first-level task.

5. The method of claim 4, wherein the partitioning of the initial data set into a training portion of the data samples and an inference portion of the data samples is performed by having a random ratio of the data samples of each of the multiple categories in the training portion or in the inference portion.

6. The method according to claim 4, wherein The first level task includes classifying the data samples of the initial data set according to the plurality of categories; The at least one characteristic comprises a confidence value, the confidence value indicating, in a result of the inference, a relevance of the one or more data samples of the inference portion to their respective associated categories; and The method further comprises A subset of the data samples is selected from the inferred portion based on a set of omission conditions, the omission conditions comprising (i) an association of the selected data samples with at least one selected category in the inferred result, and (ii) a selection of the data samples based on a calculated confidence value.

7. The method of claim 6, wherein the selection of the data samples based on the calculated confidence values ​​comprises selection of the data samples having (a) the highest scoring confidence values ​​and (b) confidence values ​​exceeding a predefined threshold.

8. The method according to claim 6, wherein Training at least one ML-based first level model includes training a plurality of ML-based first level models; and The selecting of the data sample based on the calculated confidence value comprises selecting the data sample based on a function of the calculated confidence value for a particular data sample in a plurality of first level models.

9. The method according to claim 4, wherein The first level task includes clustering the selected at least one category, and forming a cluster set of the data samples associated with the selected at least one category in the clustering result; The at least one characteristic comprises a distance between the one or more data samples and a centroid of a cluster in the set of clusters to which the one or more data samples belong in the inferred result; and The method further comprises A subset of the data samples is selected from the inferred portion based on a set of omission conditions, the omission conditions including a selection of the data samples based on the calculated distances.

10. The method of claim 9, wherein the selection of the data samples based on the calculated distance comprises selection of the data samples having (a) the shortest distance to the centroid of the cluster to which the one or more data samples belong, and (b) a distance below a predefined threshold.

11. The method of claim 9, wherein the selecting of the data samples based on the calculated distance comprises, for at least one cluster, determining a range of distances corresponding to a densest distribution of the data samples in the cluster; and A selection of at least one data sample having a distance from the determined range.

12. The method of claim 4, wherein the second level task comprises classification of data samples of an incoming data set according to the plurality of categories.

13. The method according to claim 4, wherein the method further comprises For the selected at least one category, calculating an omission percentage, the omission percentage indicating a percentage of data samples to be omitted from the initial data set; and A subset of the data samples is selected from the inferred portion based on a set of omission conditions, including selection of the data samples up to a calculated omission percentage.

14. A system for training an ML-based model, the system comprising: a non-transitory memory device having instruction code modules stored therein; and at least one processor, the at least one processor being associated with the memory device and configured to execute the instruction code module, wherein when executing the instruction code module, the at least one processor is configured to: receiving an initial data set including a plurality of data samples; Based on the initial data set, training at least one first-level ML-based model to perform a first-level task; Based on training the at least one ML-based first-level model, calculating at least one feature, the at least one feature representing, for each data sample, a contribution value to training the ML-based second-level model to perform the second-level task; A subset of data samples is omitted from the initial dataset based on the characteristics to obtain a target dataset for training the second-level ML-based model to perform the second-level task.

15. The system of claim 14, wherein the annotated data samples include annotations providing associations of the data samples with a plurality of categories; and wherein the at least one processor is further configured to select at least one of the plurality of categories based on an amount of data samples associated with each category in the initial data set.

16. The system of claim 15, wherein the at least one processor is configured to train the at least one ML-based first level model by: Dividing the initial data set into a training portion of data samples and an inference portion of data samples; training the at least one ML-based first-level model based on the training portion to perform the first-level task; Based on the training of the at least one ML-based first-level model, at least one trained ML-based first-level model is inferred on the inference portion to perform the first-level task.

17. The system of claim 16, wherein The first level task includes classifying the data samples of the initial data set according to the plurality of categories; The at least one characteristic comprises a confidence value, the confidence value indicating, in a result of the inference, a relevance of the one or more data samples of the inference portion to their respective associated categories; and wherein the at least one processor is configured to A subset of the data samples is selected from the inferred portion based on a set of omission conditions, the omission conditions comprising (i) an association of the selected data samples with at least one selected category in the inferred result, and (ii) a selection of the data samples based on a calculated confidence value.

18. The system of claim 17, wherein the at least one processor is configured to select the data sample having (a) a confidence value with the highest score and (b) a confidence value exceeding a predefined threshold based on the calculated confidence values.

19. A system according to claim 17, wherein the at least one ML-based first-level model includes multiple ML-based first-level models; and the selection of the data sample based on the calculated confidence value includes the selection of the data sample based on a function of the calculated confidence value of a specific data sample in the multiple first-level models.

20. The system of claim 16, wherein The first level task includes clustering the selected at least one category, and forming a cluster set of the data samples associated with the selected at least one category in the clustering result; The at least one characteristic comprises a distance between the one or more data samples and a centroid of a cluster in the set of clusters to which the one or more data samples belong in the inferred result; and wherein the at least one processor is further configured to: A subset of the data samples is selected from the inferred portion based on a set of omission conditions, the omission conditions including a selection of the data samples based on the calculated distances.