Selecting a subset of data for training a neural network
By applying a clustering algorithm and nearest neighbor algorithm to select a subset of data for training neural networks, the method addresses the computational challenges of training DNNs, achieving faster and more efficient training while improving generalization and reducing overfitting.
Patent Information
- Application Number
- PCT/EP2023/084244
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-05
- Publication Date
- 2025-06-12
AI Technical Summary
Training deep neural networks (DNNs) is computationally intensive and challenging, especially with large datasets, due to the need for extensive computation and memory resources, which can lead to slow training processes and overfitting.
A method is proposed to select a representative subset of data for training neural networks by applying a clustering algorithm to a latent representation of the dataset, selecting specific cluster points, and using a nearest neighbor algorithm to add these points to the subset, thereby reducing the computational burden and improving generalization.
This approach allows for faster and more efficient training of neural networks with fewer computational resources, reduces the risk of overfitting, and improves the model's ability to generalize to new data by selecting a subset that is more representative and less prone to noise or bias.
Smart Images

Figure EP2023084244_12062025_PF_FP_ABST
Abstract
Description
[0001] SELECTING A SUBSET OF DATA FOR TRAINING A NEURAL NETWORK
[0002] TECHNICAL FIELD
[0003] The present disclosure relates, in general, to selecting a subset of data for training a neural network. Aspects of the disclosure relate to selecting a representative subset of data for a given cluster.
[0004] BACKGROUND
[0005] Deep neural networks (DNN) are a popular class of machine learning algorithms that have achieved remarkable success in various fields, including computer vision (CV), natural language processing (NLP), and speech recognition. However, training DNNs is a computationally intensive and challenging task, especially when dealing with large datasets.
[0006] The DNN training problem involves optimising the weights of the network to minimise the difference between its predictions and the actual labels of the training data. This process requires computing the gradients of the loss function with respect to the network weights, which can be time-consuming and computationally expensive.
[0007] Supervised learning is a type of machine learning where the algorithm learns to map input data to output labels based on a labelled training dataset. In this case, the DNN training problem is to minimise the difference between the network's predictions and the known labels in the training data. On the other hand, unsupervised learning involves learning patterns and relationships in the input data without explicit output labels. In this case, the DNN training problem is to find a lower-dimensional representation of the input data that preserves the important features and structure of the original data.
[0008] High-dimensional input data requires more computation and memory resources, which can slow down the training process and make it computationally expensive. In the recent years, specifically in NLP domain, the models have become huge and hard to train on existing hardware. Specifically, those models use large datasets where training on the full dataset can be time-consuming and resource-intensive. SUMMARY
[0009] An objective of the present disclosure is to provide a fast, efficient method for selecting a subset of data for training a neural network.
[0010] The foregoing and other objectives are achieved by the features of the independent claims.
[0011] Further implementation forms are apparent from the dependent claims, the description and the Figures.
[0012] A first aspect of the present disclosure provides a method of selecting a subset of data for training a neural network, the method comprising applying a clustering algorithm to a latent representation of a dataset for training the neural network, whereby to acquire a data cluster, wherein the data cluster comprises multiple cluster points including a central point, selecting a first cluster point of the multiple cluster points, wherein the first cluster point is separated from the central point by a first distance, wherein the first cluster point is located in a first direction with respect to the central point, generating a logical point, the logical point separated from the central point by the first distance, wherein the logical point is located in a second direction with respect to the central point, wherein the second direction is opposite the first direction, applying a nearest neighbour algorithm to the logical point, whereby to select a second cluster point of the multiple cluster points, and adding the selected first cluster point and the selected second cluster point to the subset of data.
[0013] Accordingly, the training of the neural network can be completed in a shorter amount of time and with fewer computational resources. By selecting a representative subset of data, the model can be trained to better generalise new data, instead of overfitting and not performing well on new / unseen data, as would be the case if the model was trained on the entire dataset. In addition, datasets often contain noise or irrelevant data points. By selecting a subset of data, those data points can be removed, therefore improving training performance. Furthermore, some datasets may contain bias or imbalances in the data distribution. By selecting a subset of data, it may be possible to reduce the bias or imbalance in the data distribution, leading to a more accurate and fairer model.
[0014] A number of cluster points to be added to the subset of data may be equal to a number of the multiple cluster points in the data cluster multiplied by a pruning factor. The first distance may comprise the greatest distance between a central point and a cluster point of the multiple cluster points.
[0015] The latent representation of the dataset for training the neural network may comprise a number of dimensions associated therewith, wherein a number of cluster points added to the subset of data is smaller than the number of dimensions.
[0016] The method may further comprise calculating, based on the subset of data, an average cluster point, and applying the nearest neighbour algorithm to the calculated average cluster point, whereby to select the first cluster point.
[0017] The method may further comprise subtracting the selected first cluster point and the selected second cluster point from each cluster point of the multiple cluster points.
[0018] The method may further comprise adding the selected first cluster point and the selected second cluster point to a temporary subset of data.
[0019] The method may further comprise calculating a nullspace based on the temporary subset of data, wherein the subtracting the selected first cluster point and the selected second cluster point from each cluster point of the multiple cluster points comprises projecting the multiple cluster points onto the calculated nullspace.
[0020] A second aspect of the present disclosure provides a computer readable storage medium comprising computer program code, accessible by an apparatus comprising a processor, to provide instructions and / or data to the apparatus, the computer program code configured to, with the processor, cause the apparatus to apply a clustering algorithm to a latent representation of a dataset for training the neural network, whereby to acquire a data cluster, wherein the data cluster comprises multiple cluster points including a central point, select a first cluster point of the multiple cluster points, wherein the first cluster point is separated from the central point by a first distance, wherein the first cluster point is located in a first direction with respect to the central point, generate a logical point, the logical point separated from the central point by the first distance, wherein the logical point is located in a second direction with respect to the central point, wherein the second direction is opposite the first direction, apply a nearest neighbour algorithm to the logical point, whereby to select a second cluster point of the multiple cluster points, and add the selected first cluster point and the selected second cluster point to the subset of data.
[0021] The first distance may comprise the greatest distance between the central point and a cluster point of the multiple cluster points.
[0022] The computer readable storage medium may further comprise the computer program code configured to, with the processor, cause the apparatus to calculate, based on the subset of data, an average cluster point, and apply the nearest neighbour algorithm to the calculated average cluster point, whereby to select the first cluster point.
[0023] The computer readable storage medium may further comprise the computer program code configured to, with the processor, cause the apparatus to subtract the selected first cluster point and the selected second cluster point from each cluster point of the multiple cluster points.
[0024] A third aspect of the present disclosure provides a computing system, comprising a processor, a memory coupled to the processor, configured to store program code executable by the processor, the program code comprising one or more instructions to cause the computing system to apply a clustering algorithm to a latent representation of a dataset for training the neural network, whereby to acquire a data cluster, wherein the data cluster comprises multiple cluster points including a central point, select a first cluster point of the multiple cluster points, wherein the first cluster point is separated from the central point by a first distance, wherein the first cluster point is located in a first direction with respect to the central point, generate a logical point, the logical point separated from the central point by the first distance, wherein the logical point is located in a second direction with respect to the central point, wherein the second direction is opposite the first direction, apply a nearest neighbour algorithm to the logical point, whereby to select a second cluster point of the multiple cluster points, and add the selected first cluster point and the selected second cluster point to the subset of data.
[0025] These and other aspects of the invention will be apparent from the embodiment s) described below. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order that the present invention may be more readily understood, embodiments of the invention will now be described, by way of example, with reference to the accompanying drawings, in which:
[0027] Figure l is a flow chart of a method of selecting a subset of data for training a neural network according to an example;
[0028] Figures 2a and 2b are schematic representations of selected cluster points according to an example; and
[0029] Figure 3 is schematic depiction of a computing system according to an example.
[0030] DETAILED DESCRIPTION
[0031] Example embodiments are described below in sufficient detail to enable those of ordinary skill in the art to embody and implement the systems and processes herein described. It is important to understand that embodiments can be provided in many alternate forms and should not be construed as limited to the examples set forth herein.
[0032] Accordingly, while embodiments can be modified in various ways and take on various alternative forms, specific embodiments thereof are shown in the drawings and described in detail below as examples. There is no intent to limit to the particular forms disclosed. On the contrary, all modifications, equivalents, and alternatives falling within the scope of the appended claims should be included. Elements of the example embodiments are consistently denoted by the same reference numerals throughout the drawings and detailed description where appropriate.
[0033] The terminology used herein to describe embodiments is not intended to limit the scope. The articles “a,” “an,” and “the” are singular in that they have a single referent, however the use of the singular form in the present document should not preclude the presence of more than one referent. In other words, elements referred to in the singular can number one or more, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises,” “comprising,” “includes,” and / or “including,” when used herein, specify the presence of stated features, items, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, items, steps, operations, elements, components, and / or groups thereof. Unless otherwise defined, all terms (including technical and scientific terms) used herein are to be interpreted as is customary in the art. It will be further understood that terms in common usage should also be interpreted as is customary in the relevant art and not in an idealized or overly formal sense unless expressly so defined herein.
[0034] As mentioned before, machine learning can be divided into, among others, supervised learning, and unsupervised learning. Supervised solutions involve inputting objects and a desired output value to train a model. Unfortunately, these solutions rely on trained models where the user needs to find the appropriate network, hyper parameters, time and computational resources. In contrast, unsupervised learning relies on teaching algorithms patterns exclusively from unlabelled data. However, random solutions cannot ensure the selection of comprehensive data representation. Furthermore, random sampling can be problematic as the number of samples will decrease.
[0035] One of the major challenges associated with machine learning is loss of information, i.e., being able to reduce the dimensionality (which can negatively impact model accuracy) while preserving the most important information. Furthermore, it can be difficult to determine which features are important for the machine learning task and which features are redundant, especially when dealing with high-dimensional data. In addition, it is difficult to develop efficient algorithms that can scale to large datasets and reduce the computational burden. Finally, reduced feature sets can be difficult to interpret, thus being unsuitable for providing insights into the underlying patterns in the data.
[0036] According to an example, there is provided a mechanism to select a subset of data for training a neural network, the training of the neural network can be completed in a shorter amount of time and with fewer computational resources. By selecting a representative subset of data, the model can be trained to better generalise new data, instead of overfitting and not performing well on new / unseen data, as would be the case if the model was trained on the entire dataset. In addition, datasets often contain noise or irrelevant data points. By selecting a subset of data, those data points can be removed, therefore improving training performance. Furthermore, some datasets may contain bias or imbalances in the data distribution. By selecting a subset of data, it may be possible to reduce the bias or imbalance in the data distribution, leading to a more accurate and fairer model. Examples in the present disclosure can be provided as methods, systems or machine-readable instructions, such as any combination of software, hardware, firmware or the like. Such machine-readable instructions may be included on a computer readable storage medium (including but not limited to disc storage, CD-ROM, optical storage, etc.) having computer readable program codes therein or thereon.
[0037] The present disclosure is described with reference to flow charts and / or block diagrams of the method, devices and systems according to examples of the present disclosure. Although the flow diagrams described above show a specific order of execution, the order of execution may differ from that which is depicted. Blocks described in relation to one flow chart may be combined with those of another flow chart. In some examples, some blocks of the flow diagrams may not be necessary and / or additional blocks may be added. It shall be understood that each flow and / or block in the flow charts and / or block diagrams, as well as combinations of the flows and / or diagrams in the flow charts and / or block diagrams can be realized by machine readable instructions.
[0038] The machine-readable instructions may, for example, be executed by a machine such as a general-purpose computer, user equipment such as a smart device, e.g., a smart phone, a special purpose computer, an embedded processor or processors of other programmable data processing devices to realize the functions described in the description and diagrams. In particular, a processor or processing apparatus may execute the machine-readable instructions. Thus, modules of apparatus (for example, a module implementing a comparator unit, or a firewall structure and so on) may be implemented by a processor executing machine readable instructions stored in a memory, or a processor operating in accordance with instructions embedded in logic circuitry. The term 'processor' is to be interpreted broadly to include a CPU, processing unit, ASIC, logic unit, or programmable gate set etc. The methods and modules may all be performed by a single processor or divided amongst several processors.
[0039] Such machine-readable instructions may also be stored in a computer readable storage that can guide the computer or other programmable data processing devices to operate in a specific mode. For example, the instructions may be provided on a non-transitory computer readable storage medium encoded with instructions, executable by a processor. Figure l is a flow chart of a method of selecting a subset of data for training a neural network according to an example. The method comprises, in block 101, applying a clustering algorithm to a latent representation of a dataset for training the neural network, whereby to acquire a data cluster. Here, the term “latent representation” refers to a simplified model of the input data, i.e., a simplified model of the dataset to be used for training the neural network. The latent representation of the dataset for training the neural network may comprise a number of dimensions associated therewith, wherein a number of cluster points to be selected and added to the subset of data is smaller than the number of dimensions. The data cluster comprises multiple cluster points, including a central point.
[0040] In block 102, the method comprises selecting a first cluster point of the multiple cluster points. The first cluster point is separated from the central point by a first distance. The first distance may comprise the greatest distance between the central point and a cluster point of the multiple cluster points. That is, the first cluster point may be a cluster point that is located the farthest from the central point. The first cluster point is located in a first location with respect to the central point.
[0041] In block 103, a logical point is generated. The logical point is separated from the central point by the first distance, i.e., it is located equidistantly from the central point as the first cluster point. The logical point is located in a second direction with respect to the central point, wherein the second direction is a direction opposite to the first direction.
[0042] The method comprises, in block 104, applying a nearest neighbour algorithm to the logical point, whereby to select a second cluster point of the multiple cluster points. In block 105, both the selected first cluster point and the selected second cluster point are added to the subset of data for training the neural network. The number of cluster points to be added to the subset of data may be equal to a number of the multiple cluster points in the data cluster multiplied by a pruning factor.
[0043] Generally speaking, the method can be thought of as having two stages. During the first stage, the dataset, the pruning factor and the model for data may be used as an input. The given dataset may be transformed in order to obtain the latent representation. Then, a clustering algorithm may be applied for each class of data. In the case where only a single class of data is present in the dataset, the clustering algorithm may be applied to a single class. As mentioned above, the number of cluster points (i.e., the number of required samples) to be added to the subset of data may be equal to a number of the cluster points in the data cluster multiplied by a pruning factor. The second stage may be executed until the number of cluster points added to the subset of data is equal to the number of required samples, i.e., until the required number of samples has been added to the subset of data.
[0044] Figures 2a and 2b are schematic representations of selected cluster points according to an example. Figure 2a represents a scenario where points that are orthogonal to the central point are collected, whereas Figure 2b represents a scenario where points that are not orthogonal to the central point are selected, in order to improve the diversity of the selected samples.
[0045] In the scenario shown in Figure 2a, a first cluster point 202 may be selected, i.e., the cluster point having the maximum distance from the central point 201. Following the selection of the first cluster point 202, a logical point 203 may be generated. The logical point 203 may be located in the opposite direction to the first cluster point 202, separated from the central point 201 by the same distance as the first cluster point 202. Following this, a second cluster point 204 may be selected by applying a nearest neighbour algorithm to the logical point 203. Then, a nullspace may be calculated. A nullspace is the set of all solutions to a linear system of equations that satisfy the equation Ax = 0, where A is a matrix and x is a column vector. It is also known as the kernel of the linear transformation represented by the matrix A. All of the vectors x are orthogonal to A. Following the calculation of the nullspace, by using a projection of the multiple cluster points onto the nullspace, the selected first cluster point 202 and the selected second cluster point 204 may be subtracted from each cluster point of the multiple cluster points. A final set of points may then be returned.
[0046] In the scenario shown in Figure 2b, a first cluster point 212 may be selected instead. To select the first cluster point 212, an average cluster point 207 may be calculated based on the samples selected so far. The average cluster point 207 may be calculated by dividing by two the sum of a first average cluster point 205 and a second average cluster point 206. The first cluster point 212 may be the nearest neighbour to the average cluster point 207, as shown in Figure 2b. Similarly to the scenario described in relation to Figure 2a, a logical point 213 may be generated. The logical point 213 may be located in the opposite direction to the first cluster point 212, separated from the central point 201 by the same distance as the first cluster point 212. Following this, a second cluster point 214 may be selected by applying a nearest neighbour algorithm to the logical point 213. As has been described above in relation to Figure 2a, a nullspace may then be calculated. Following the calculation of the nullspace, by using a projection of the multiple cluster points onto the nullspace, the selected first cluster point 212 and the selected second cluster point 214 may be subtracted from each cluster point of the multiple cluster points. A final set of points may then be returned.
[0047] To aid the processes described above in relation to Figures 2a and 2b, the method may additionally comprise adding the selected cluster points to a temporary subset of data. In particular, all of the cluster points that have been selected as the nearest points to the currently selected cluster point may be added to the temporary subset of data in order to create the nullspace. Following this, a projection of the central point may be created on that nullspace.
[0048] Figure 3 is a schematic of a computing system according to an example. The computing system 300 comprises a processor 303, and a memory 305 coupled to the processor 303 and configured to store instructions or program code 307, executable by the processor 303. The computing system 300 can be, e.g., a computing system or apparatus, user equipment, a network device (physical or virtual), or part thereof. The computing system 300 comprises the program code 307 arranged to cause the computing system to perform the method of selecting the subset of data for training a neural network as described above in relation to Figures 1, 2a and 2b.
[0049] According to an example, machine-readable instructions can be loaded onto a computer or other programmable data processing devices, so that the computer or other programmable data processing devices perform a series of operations to produce computer-implemented processing, thus the instructions executed on the computer or other programmable devices provide an operation for realizing functions specified by flow(s) in the flow charts and / or block(s) in the block diagrams.
[0050] Further, the teachings herein may be implemented in the form of a computer or software product, such as a non-transitory machine-readable storage medium, the computer software or product being stored in a storage medium and comprising a plurality of instructions, e.g., machine readable instructions, for making a computer device implement the methods recited in the examples of the present disclosure. In some examples, some methods can be performed in a cloud-computing or network-based environment. Cloud-computing environments may provide various services and applications via the Internet. These cloud-based services (e.g., software as a service, platform as a service, infrastructure as a service, etc.) may be accessible through a web browser or other remote interface of the user equipment for example. Various functions described herein may be provided through a remote desktop environment or any other cloud-based computing environment.
[0051] While various embodiments have been described and / or illustrated herein in the context of fully functional computing systems, one or more of these exemplary embodiments may be distributed as a program product in a variety of forms, regardless of the particular type of computer- readable-storage media used to actually carry out the distribution. The embodiments disclosed herein may also be implemented using software modules that perform certain tasks. These software modules may include script, batch, or other executable files that may be stored on a computer-readable storage medium or in a computing system. In some embodiments, these software modules may configure a computing system to perform one or more of the exemplary embodiments disclosed herein. In addition, one or more of the modules described herein may transform data, physical devices, and / or representations of physical devices from one form to another.
[0052] The preceding description has been provided to enable others skilled in the art to best utilize various aspects of the exemplary embodiments disclosed herein. This exemplary description is not intended to be exhaustive or to be limited to any precise form disclosed. Many modifications and variations are possible without departing from the spirit and scope of the instant disclosure. The embodiments disclosed herein should be considered in all respects illustrative and not restrictive. Reference should be made to the appended claims and their equivalents in determining the scope of the instant disclosure.
Claims
CLAIMS1. A method of selecting a subset of data for training a neural network, the method comprising: applying a clustering algorithm to a latent representation of a dataset for training the neural network, whereby to acquire a data cluster, wherein the data cluster comprises multiple cluster points including a central point (201); selecting a first cluster point of the multiple cluster points, wherein the first cluster point is separated from the central point by a first distance, wherein the first cluster point is located in a first direction with respect to the central point (202); generating a logical point, the logical point separated from the central point by the first distance, wherein the logical point is located in a second direction with respect to the central point, wherein the second direction is opposite the first direction (203); applying a nearest neighbour algorithm to the logical point, whereby to select a second cluster point of the multiple cluster points (204); and adding the selected first cluster point and the selected second cluster point to the subset of data (205).
2. The method of claim 1, wherein a number of cluster points to be added to the subset of data is equal to a number of the multiple cluster points in the data cluster multiplied by a pruning factor.
3. The method of claim 1, wherein the first distance comprises the greatest distance between the central point and a cluster point of the multiple cluster points.
4. The method of claim 3, wherein the latent representation of the dataset for training the neural network comprises a number of dimensions associated therewith, wherein a number of cluster points added to the subset of data is smaller than the number of dimensions.
5. The method of claim 1, further comprising: calculating, based on the subset of data, an average cluster point; and applying the nearest neighbour algorithm to the calculated average cluster point, whereby to select the first cluster point.
6. The method of any one of claims 1 to 5, further comprising: subtracting the selected first cluster point and the selected second cluster point from each cluster point of the multiple cluster points.
7. The method of claim 6, further comprising: adding the selected first cluster point and the selected second cluster point to a temporary subset of data.
8. The method of claim 7, further comprising calculating a nullspace based on the temporary subset of data, wherein the subtracting the selected first cluster point and the selected second cluster point from each cluster point of the multiple cluster points comprises projecting the multiple cluster points onto the calculated nullspace.
9. A computer readable storage medium comprising computer program code, accessible by an apparatus comprising a processor, to provide instructions and / or data to the apparatus, the computer program code configured to, with the processor, cause the apparatus to: apply a clustering algorithm to a latent representation of a dataset for training the neural network, whereby to acquire a data cluster, wherein the data cluster comprises multiple cluster points including a central point; select a first cluster point of the multiple cluster points, wherein the first cluster point is separated from the central point by a first distance, wherein the first cluster point is located in a first direction with respect to the central point; generate a logical point, the logical point separated from the central point by the first distance, wherein the logical point is located in a second direction with respect to the central point, wherein the second direction is opposite the first direction; apply a nearest neighbour algorithm to the logical point, whereby to select a second cluster point of the multiple cluster points; andadd the selected first cluster point and the selected second cluster point to the subset of data.
10. The computer readable storage medium of claim 9, wherein the first distance comprises the greatest distance between the central point and a cluster point of the multiple cluster points.
11. The computer readable storage medium of claim 9, further comprising the computer program code configured to, with the processor, cause the apparatus to: calculate, based on the subset of data, an average cluster point; and apply the nearest neighbour algorithm to the calculated average cluster point, whereby to select the first cluster point.
12. The computer readable storage medium of any one of claims 9 to 11, further comprising the computer program code configured to, with the processor, cause the apparatus to: subtract the selected first cluster point and the selected second cluster point from each cluster point of the multiple cluster points.
13. A computing system (300), comprising: a processor (303); a memory (305) coupled to the processor (303), configured to store: program code (307) executable by the processor (303), the program code (307) comprising one or more instructions to cause the computing system (300) to: apply a clustering algorithm to a latent representation of a dataset for training the neural network, whereby to acquire a data cluster, wherein the data cluster comprises multiple cluster points including a central point; select a first cluster point of the multiple cluster points, wherein the first cluster point is separated from the central point by a first distance, wherein the first cluster point is located in a first direction with respect to the central point; generate a logical point, the logical point separated from the central point by the first distance, wherein the logical point is located in a second direction with respect to the central point, wherein the second direction is opposite the first direction;apply a nearest neighbour algorithm to the logical point, whereby to select a second cluster point of the multiple cluster points; and add the selected first cluster point and the selected second cluster point to the subset of data.