Feature selection from remote data sources for machine learning models
The method addresses the challenge of feature selection from remote data sources by using a feature importance and redundancy model to efficiently select and update features, enhancing the accuracy and efficiency of machine learning tasks.
Patent Information
- Application Number
- PCT/IB2024/050428
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-13
- Filing Date
- 2024-01-17
- Publication Date
- 2025-05-22
AI Technical Summary
Existing feature selection techniques face challenges in efficiently selecting features from remote data sources within a limited number of data accesses, particularly due to network bandwidth limitations and computational constraints.
A computer-implemented method that initializes a feature importance model and a feature redundancy model, selects features that maximize output from the feature importance model, and updates model parameters based on true feature importance, all while efficiently accessing and storing feature data in remote sources.
This method enables efficient and accurate feature selection from remote data sources, optimizing machine learning tasks and supporting decision-making in applications such as medicine, smart buildings, and energy management.
Smart Images

Figure IB2024050428_22052025_PF_FP_ABST
Abstract
Description
FEATURE SELECTION FROM REMOTE DATA SOURCES FOR MACHINE LEARNING MODELSCROSS-REFERENCE TO PRIOR APPLICATION
[0001] Priority is claimed to U.S. Provisional Application No. 63 / 548,238, filed on November 13, 2023, the entire contents of which is hereby incorporated by reference herein. FIELD
[0002] The present invention relates to artificial intelligence (Al) and machine learning (ML), and in particular to a method, system, data structure, computer program product and computer-readable medium for selecting features from remote data sources for machine learning models.BACKGROUND
[0003] There are a large number of existing feature selection techniques. According to existing technology, the suitability of features is tested by comparing their data with the target variable to predict and with other available features. There are several classes of techniques to do so, ranging from computation of correlations to training and inspection of preliminary prediction models. Chandrashekar, Girish, and Ferat Sahin. “A survey on feature selection methods.” Computers & Electrical Engineering 40.1 (2014): 16-28, which is hereby incorporated by reference herein, provides a comprehensive survey of existing feature selection techniques.SUMMARY
[0004] In an embodiment, the present disclosure provides a computer-implemented method for efficiently selecting features in a distributed system comprising a plurality of remote data sources and a computing system. The computing system initializes a feature importance model and a feature redundancy model, and selects a feature that maximizes an output from the feature importance model. The computing system accesses feature data associated with the feature, and the feature data is stored in one of the remote data sources. The computing system updates parameters of the feature importance model based on determining a true feature importance associated with the selected feature. The computing system selects one or more further features based on the feature importance model and the feature redundancy model. The method has applications including, but not limited to, use cases in medicine / healthcare, smart buildings and cities, energy distribution and management, public safety and predictive maintenance, for example, to optimize machine learning tasks or to support decision making.BRIEF DESCRIPTION OF THE DRAWINGS
[0005] Embodiments of the present invention will be described in even greater detail below based on the exemplary figures. The present invention is not limited to the exemplary embodiments. All features described and / or illustrated herein can be used alone or combined in different combinations in embodiments of the present invention. The features and advantages of various embodiments of the present invention will become apparent by reading the following detailed description with reference to the attached drawings which illustrate the following:
[0006] FIG. 1 illustrates a simplified block diagram depicting an exemplary computing environment according to an embodiment of the present disclosure;
[0007] FIG. 2 illustrates a visualization of a metadata subspace with a target y and three features; and
[0008] FIG. 3 is a block diagram of an exemplary processing system, which can be configured to perform any and all operations disclosed herein.DETAILED DESCRIPTION
[0009] Embodiments of the present invention provide a method and system for extracting features for machine learning models from remote data sources within a limited number of data accesses in order to automatically train models for prediction-based decision making, for example, in the domains of smart cities, smart buildings, and healthcare. In particular, embodiments of the present invention provide solutions to the technical problem of how to extract features for machine learning models from remote data sources within a limited number of data accesses for training models in an automated fashion by combining feature selection techniques with predictions of feature importance and feature redundancy from metadata.
[0010] Embodiments of the present invention also provide solutions to the technical problem of how to automatically select remote sources of input features for a prediction model. Remote sources are data sources (e.g., sensors) whose data needs to be retrieved from a remote location over a network. In a large distributed system, a large number of such data sources is available, and it is neither necessary nor possible to use all of them as inputs to a prediction model. Embodiments of the present invention provide solutions to this technical problem by selecting a subset of data sources as inputs for the prediction model in an efficient, accurate and automated fashion.
[0011] As mentioned above, there are a large number of existing feature selection techniques. According to existing technology, the suitability of features is tested by comparing their data with the target variable to predict and with other available features. There are several classes of techniques to do so, ranging from computation of correlations to training and inspection of preliminary prediction models. Chandrashekar, Girish, and Ferat Sahin. “A survey on feature selection methods.” Computers & Electrical Engineering 40.1 (2014): 16-28, which ishereby incorporated by reference herein, provide a comprehensive survey of existing feature selection techniques.
[0012] A technical problem of existing feature selection technology is that the feature data is utilized for assessing the given features. This is a technical problem because, as described above, it is prohibitive to download and analyze the data of all given data sources, due to network bandwidth limitations and computational constraints. Embodiments of the present invention provide solutions to overcome this technical problem to enable performing feature selection in such a scenario.
[0013] Cinelli, Matteo, Giovanna Ferraro, and Antonio lovanella. "Evaluating relevance and redundancy to quantify how binary node metadata interplay with the network structure." Scientific reports 9.1 (2019): 11404, which is hereby incorporated by reference herein, present an approach in which the metadata of nodes in a graph is taken into account to evaluate the relevance and redundancy of each graph node, and the authors mention that this process has some similarity to the feature selection process. However, in that work, only the metadata (plus the structure of a graph, which is not given in a technical problem solved by embodiments of the present invention) is taken into account, so the process remains blind about the real feature data during the whole selection process, which limits its effectiveness.
[0014] Embodiments of the present invention enhance computer functionality of machine learning systems to be able to use any available feature metadata instead of the feature data itself. The metadata serves as input to a meta-model that makes (e.g., generates) predictions about the usefulness of a given feature, and to update the meta-model as features are selected and the feature data is inspected.
[0015] According to a first aspect, the present invention provides a computer-implemented method for efficiently selecting features in a distributed system comprising a plurality of remote data sources and a computing system. For instance, the computing system initializes a feature importance model and a feature redundancy model, and selects a feature that maximizes an output from the feature importance model. The computing system accesses feature data associated with the feature, and the feature data is stored in one of the remote data sources. The computing system updates parameters of the feature importance model based on determining a true feature importance associated with the selected feature. The computing system selects one or more further features based on the feature importance model and the feature redundancy model
[0016] According to a second aspect, the method according to the first aspect further comprises that the plurality of remote data sources are a plurality of sensors located at a plurality of different geographical locations, the feature data is sensor data that is obtained by the sensor,selecting the one or more further features is based on metadata associated with the sensor data, and initializing the feature importance model and the feature redundancy model comprises initializing the feature importance model and the feature redundancy model as functions of a feature distance in a metadata space.
[0017] According to a third aspect, the method according to any of the first or the second aspect further comprises that selecting the one or more further features comprises: performing a selection process to select a new feature from a plurality of features based on the feature importance model and the feature redundancy model; obtaining, from a second remote data source of the plurality of remote data sources, second feature data associated with the new feature; updating the parameters of the feature importance model and the feature redundancy model based on the second feature data; and determining whether to perform another selection process based on the second feature data.
[0018] According to a fourth aspect, the method according to any of the first to third aspects further comprises that performing the selection process is based on maximizing a product of a predicted feature importance associated with the feature importance model and a predicted feature redundancy associated with the feature redundancy model.
[0019] According to a fifth aspect, the method according to any of the first to fourth aspects further comprises that the feature importance model comprises an importance (IMP) function, the feature redundancy model comprises a redundancy (RED) function, and performing the selection process comprises: using the IMP function to determine the predicted feature importance; using the RED function to determine the predicted feature redundancy; and selecting the new feature based on maximizing a product of the predicted feature importance and the predicted feature redundancy.
[0020] According to a sixth aspect, the method according to any of the first to fifth aspects further comprises that the parameters comprise a first parameter and a second parameter, and the method further comprises: initializing the first parameter and the second parameter based on a maximum distance between any pair of features, of the plurality of features, in a metadata space, wherein using the IMP function to determine the predicted feature importance is based on the first parameter, and wherein using the RED function to determine the predicted feature redundancy is based on the second parameter.
[0021] According to a seventh aspect, the method according to any of the first to sixth aspects further comprises that using the IMP function comprises inputting a first distance metric associated with a given feature, from the plurality of features, and a prediction target in the metadata space into the IMP function to determine the predicted feature importance, and using the RED function comprises inputting a second distance metric associated with two features,from the plurality of features, into the RED function to determine the predicted feature redundancy.
[0022] According to an eighth aspect, the method according to any of the first through seventh aspects further comprises that selecting the one or more further features further comprises: determining a measured true importance value and a measured true redundancy value based on the second feature data, and updating the parameters of the feature importance model and the parameters of the feature redundancy model are based on the measured true importance value and the measured true redundancy value.
[0023] According to an ninth aspect, the method according to any of the first through eighth aspects further comprises that determining the measured true importance value comprises determining the measure true importance value based on the second feature data and prediction target feature data, and determining the measured true redundancy value is based on the second feature data and the feature data associated with the feature.
[0024] According to a tenth aspect, the method according to any of the first through ninth aspects further comprises that performing the selection process comprises: obtaining first metadata associated with sets of feature data associated with a plurality of features; obtaining second metadata associated with a prediction target; determining one or more distance metrics based on the first metadata, the second metadata, and a meta-distance function; and performing the selection process based on the one or more distance metrics.
[0025] According to an eleventh aspect, the method according to any of the first through tenth aspects further comprises that determining one or more distance metrics comprises: using a pre-trained language model to determine a first vector representation associated with the first metadata and a second vector representation associated with the second metadata; and determining the one or more distance metrics based on the first vector representation, the second vector representation, and the meta-distance function.
[0026] According to a twelfth aspect, the method according to any of the first through eleventh aspects further comprises that performing the selection process further comprises: inputting the one or more distance metrics into the feature importance model to determine a predicted feature importance; inputting the one or more distance metrics into the feature redundancy model to determine a predicted feature redundancy; and determining the new feature based on the predicted feature importance and the predicted feature redundancy.
[0027] According to a thirteenth aspect, the method according to any of the first through twelfth aspects further comprising that determining whether to perform another selection process comprises: comparing an iteration counter indicating a number of iterations that have been executed with a first stopping criterion threshold; comparing a true redundancy value with asecond stopping criterion threshold; and determining to perform another selection process based on the comparisons.
[0028] According to a fourteenth aspect of the present disclosure, a computer system is provided for efficiently selecting features in a distributed system comprising a plurality of remote data sources and the computing system, the computing system comprising one or more hardware processors, which, alone or in combination, are configured to provide for execution of the method according to any of the first to the thirteenth aspects and / or execution of the following steps: initializing a feature importance model and a feature redundancy model; selecting a feature that maximizes an output from the feature importance model; accessing feature data associated with the feature, wherein the feature data is stored in one of the remote data sources; updating parameters of the feature importance model based on determining a true feature importance associated with the selected feature; and selecting one or more further features based on the feature importance model and the feature redundancy model.
[0029] A fifteenth aspect of the present disclosure provides a tangible, non-transitory computer-readable medium having instructions thereon, which, upon being executed by one or more processors, provides for execution of the method according to any of the first to the thirteenth aspects and / or the method comprising the following: initializing a feature importance model and a feature redundancy model; selecting a feature that maximizes an output from the feature importance model; accessing feature data associated with the feature, wherein the feature data is stored in one of the remote data sources; updating parameters of the feature importance model based on determining a true feature importance associated with the selected feature; and selecting one or more further features based on the feature importance model and the feature redundancy model.
[0030] An exemplary embodiment of the present invention is first described with respect to a simplified setup. After that, a number of extensions and further embodiments are described.
[0031] FIG. 1 illustrates a simplified block diagram depicting an exemplary computing environment according to an embodiment of the present disclosure. For instance, FIG. 1 shows a computing environment 100 comprising a plurality of data sources 102 (e.g., remote data sources), a network 104, and a meta-model feature extraction computing system 106 (“computing system). Although certain entities within environment 100 are described below and / or depicted in the FIGs. as being singular entities, it will be appreciated that the entities and functionalities discussed herein can be implemented by and / or include one or more entities. For example, in some instances, the computing system 106 can be and / or include multiple computing devices such as a first computing device and a second computing device.
[0032] The entities within the environment 100 are in communication with other devices and / or systems within the environment 100 via the network 104. The network 104 can be a global area network (GAN) such as the Internet, a wide area network (WAN), a local area network (LAN), or any other type of network or combination of networks. The network 104 can provide a wireline, wireless, or a combination of wireline and wireless communication between the entities within the system 100.
[0033] Each of the data sources 102 is and / or includes one or more computing devices and / or systems that are configured to provide data (e.g., feature data) and / or information associated with the data (e.g., metadata or feature metadata) to the computing system 106. For example, the data sources 102 are and / or include one or more sensors, computing devices, computing platforms, systems, servers, desktops, laptops, tablets, mobile devices (e.g., smartphone device, or other mobile device), or any other type of computing device that generally comprises one or more processing components and / or one or more other components (e.g., memory components and / or communication components).
[0034] The computing system 106 is a computing system that is configured to use a metamodel to determine usefulness of features from the data sources 102, and to update the metamodel as features are selected and the feature data is inspected. The computing system 106 is and / or includes, but is not limited to, a desktop, laptop, tablet, mobile device (e.g., smartphone device, or other mobile device), server, computing system and / or other types of computing entities that generally comprises one or more communication components, one or more processing components, and one or more memory components.
[0035] It will be appreciated that the exemplary system depicted in FIG. 1 is merely an example, and that the principles discussed herein may also be applicable to other situations — for example, including other types of devices, systems, and network configurations.
[0036] In some instances, in operation, given a set X of features, for each feature x G X (i.e., x is an element of X), a metadata mxG M (i.e., mxis an element of M) is available, where M is called the metadata space. For instance, the data sources 102 can include, store, obtain, and / or otherwise be associated with feature data (i.e., features X) and metadata (i.e., M).
[0037] In one or more embodiments, the metadata can include information about the geographical location of the data source 102 (e.g., the sensor) that provides the feature as a measurement, topological information (e.g., room or road segment where the feature relates to), the type of feature (e.g., the type of measurement such as temperature or humidity), and / or other information associated with the feature data. For instance, the data source 102 can be a sensor that obtains (e.g., collect) sensor data (e.g., feature data) such as temperature or humidity of a room. Further, the sensor may obtain and / or be associated with metadata corresponding to thesensor data such as the type of measurements that the sensor can obtain, topological information (e.g., a room of a building where the sensor is located), and / or other types of metadata. In other words, in some variations, the data sources 102 can be multiple different sensors that are located at different geographical locations, can obtain feature data (e.g., sensor data) at their respective geographical location, and can include metadata associated with the feature data.
[0038] The data dxof each feature x G X is initially unknown. The prediction target y also has a metadata description myG M, and its data dyis known (for a training set). For example, in some instances, the computing system 106 might not know the data, dx. of each feature within the data sources 102, but can know the prediction target’s data, dyand / or the metadata description associated with the prediction target’s data. In some instances, the prediction target’s data, dycan be available as a special data source from the data sources 102. As such, the computing system 106 can obtain the prediction target’s data, dy, from the special data source. The computing system 106 then performs one or more embodiments of the present invention described below (e.g., the Algorithm 1) to obtain a subset of features based on the prediction target’s data.
[0039] The prediction target’s data can be data associated with (e.g., part of) a training set that is used to train one or more models or algorithms described below. For instance, the training set can include the data of the target variable (e.g., prediction target’s data) as well as additional or other data such as the data of the features selected through the execution of one or more embodiments of the present invention. Thus, the training set might not only include the target data, but can also include additional data.
[0040] Embodiments of the present invention utilize a suitable meta-distance function dist: M x M -> IR that expresses the similarity of two features (and between a feature and the prediction target) in the metadata space. For instance, the meta-distance function can be a distance function that is applied to pairs of elements in the metadata space. When applying the distance function to a pair of elements of the metadata space, the output of the distance function is a measure of how distant the two elements are. The function aboveis the mathematical notation for the distance function where M is the metadata space, M x M is the set of pairs of elements from the metadata space, IR is the set of real numbers, and dist is the distance function. In an embodiment, the meta-distance function can be defined as the geographical distance between the feature geo-locations. The meta-distance between two features can also be expressed based on the feature category, where the meta-distance is defined to be 0 for two features of the same category and 1 for two features of different categories. Further, when feature descriptions in the form of natural language text are available, a pre-trained language model can be utilized translate the texts into a vector representation and the vector distance between the representations is used as the meta-distance. Finally, several distance metrics can be combined into a single meta-distance by computing a weighted or unweighted sum, obtaining a meta-distance function taking into account multiple criteria.
[0041] For instance, after obtaining metadata associated with data sources 102, the computing system 106 can determine distance metrics based on applying a distance function to the metadata associated with data sources 102, and / or the metadata associated with the prediction target. For example the computing system 106 can use the metadata associated with a data source 102 and the prediction target to determine the similarity of the data source and the prediction target in the metadata space. In some instances, the metadata can indicate geographical locations of the feature. The computing system 106 can use a distance function to determine a distance metric associated with the difference between the geographical locations. Additionally, and / or alternatively, the feature descriptions (metadata) can be in the form of natural language text. The computing system 106 can use a pre-trained language model (e.g., a machine learning (ML) or artificial intelligence (Al) model) to translate the text into vector representation. The computing system 106 can use a distance function and the vector representations (e.g., vector representations associated with the metadata associated with the data sources 102 and the prediction target) to determine a distance metric. Additionally, and / or alternatively, some features can be associated with metadata indicating multiple different types of information (e.g., geographical location and type of measurements obtained by the sensor). The computing system 106 can determine multiple distance metrics associated with the types of information, and combine them into a single meta-distance based on a weighted or unweighted sum (e.g., each distance metric can be weighed the same or weighed different).
[0042] Based on the distance metric, two models are used to estimate the feature importance as well as the feature redundancy.1. The model estimating the feature importance is a parameterized non-increasing function IMPe: IR -> [0,1] mapping the meta-distance between a feature and the prediction target to a number between 0 and 1 (e.g., a predicted feature importance). IMP is a computer-implemented function (i.e., importance (IMP) function) to estimate the importance of a feature. The input to the function is the distance between a given feature and the target in the metadata space. The output of the IMP is the estimated importance of the feature, which is a number between 0 and 1.2. The model estimating the feature redundancy is a parameterized non-decreasing function RED^ IR -> [0,1] mapping the meta-distance between a first feature and a second feature to a number between 0 and 1 (e.g., a predicted feature redundancy). Low function values of REDrepresent high redundancy. RED is a computer-implemented function (i.e., redundancy (RED) function) to estimate the redundancy of a first feature in light of a second feature. The input to the function is the distance between the two features in the metadata space. The output of RED is the estimated redundancy of the first feature, which is a number between 0 and 1.
[0043] In an embodiment, the following concrete function definitions are used:IMPe(z) := 1 — min{0 • z, 1} RED,p z) := min{ • z, 1} where the parameters are initialized as 0 = <p = 0.2
[0044] For example, the computing system 106 can use a first model to determine feature importance. For instance, referring to the above equations, in one or more embodiments of the present invention, the computing system 106 can implement the IMP as a function that first multiplies its input by a parameter Theta (0). If the result of that first operation is larger than one, the result is set to one in a second operation. In the third operation, the distance between 1 and the result of the second operation is computed, and the result of the third operation serves as the output of IMP.
[0045] For example, the computing system 106 can use a first model to determine feature redundancy. For instance, referring to the above equations, in one or more embodiments of the present invention, the computing system 106 can implement the RED as a function that first multiplies its input by a parameter Phi ( ). If the result of that that first operation is larger than one, the result is set to one in a second operation, and the result of the second operation serves as the output of RED.
[0046] Referring to the above equations, in one or more embodiments of the present invention, the computing system 106 can initialize the parameters Theta (0) and Phi ( ) as one fifth of the maximum distance, in the metadata space, between any pair of features among the set of all features.
[0047] The algorithm selects features in iterations i = 1,2, ... where, in each iteration i, one new featureE X is selected. The selection is performed according to:Xj <— arg max
[0048] For example, the computing system 106 can select features from the data sources 102 in iterations based on the feature that maximizes the importance-redundancy product, which is defined as the product of the feature importance with all feature redundancies, i.e., the redundancies computed against all features that have been selected in preceding iterations. In some instances, the computing system 106 can select a single new feature in each iteration.
[0049] After xthas been selected, its data dx. is accessed. Using the data of and y as well as all previously selected features xn, Xj_ , the true features importance and redundancy are measured. In an embodiment, the true feature importance is measured as the absolute correlation between dx. and dy(e.g., the accessed feature data and the prediction target feature data), and the true feature redundancy between xtand Xj is measured as one minus the absolute correlation between dx. and dx.. The absolute correlation of a pair of variables is the absolute value of the correlation of the two variables, where correlation is a well-known statistical function. Further, the previously selected features or the data associated with the previously selected features can be used to determine the true feature importance and / or the true feature redundancy. These correlations can be computed using the data.
[0050] For instance, after selecting a feature (e.g., “xj”) during one iteration of the process, the computing system 106 can access the data associated with the feature. For instance, the computing system 106 can obtain the data associated with the feature from the data source 102. Using the obtained data of Xj, the prediction target, and information associated with the previously selected features, the computing system 106 can determine the true feature importance and the true feature redundancy.
[0051] Using the measured feature importance and redundancy as ground truth, the parameters of the feature importance and feature redundancy model are updated using a machine learning methodology. In an embodiment, the gradient descent method is used for updating the model parameters. For instance, the computing system 106 can update, using a machine learning methodology, the parameters of the feature importance and feature redundancy model using the measured feature importance and redundancy. In some instances, the computing system 106 can use the gradient descent method to update the parameters of the feature importance and feature redundancy model.
[0052] After the parameter update, the algorithm (e.g., the computing system 106) executes the next iteration i + 1. The algorithm continues with more iterations, until a stopping criterion is satisfied. In an embodiment, the stopping criterion is satisfied when some pre-defined number of iterations have been executed. In another embodiment, the stopping criterion is satisfied when during a pre-defined number of iterations, the true redundancy of the selected feature does not surpassing a pre-defined threshold.
[0053] For instance, after executing the parameters, the computing system 106 determines whether to proceed with another iteration of the process. The computing system 106 can use one or more stopping criterions to determine whether to proceed with another iteration. In some embodiments, the stopping criterion is a threshold (e.g., a first stopping criterion threshold)indicating a pre-defined number of iterations that have been executed. After completing each iteration, the computing system 106 updates a counter associated with the number of iterations executed. The computing system 106 compares the counter with a pre-defined threshold, and determines whether to proceed with another iteration of the process based on the comparison. Additionally, and / or alternatively, the computing system 106 compares the true redundancy of the selected feature with a pre-defined threshold (e.g., a second stopping criterion threshold). Based on the comparison, the computing system 106 determines whether to proceed with another iteration of the process.
[0054] An algorithmic description and visualization of the methodology of the exemplary embodiment of the present invention is provided below:Algorithm 1:Input: target data dy. target metadata my. feature metadata mx, x G X.Step 1 : Initialize model parameters 9, pStep 2: i «- 0Step 3 : while (stopping criterion not satisfied)Step 4: i «- i + 1Step 5: x;«- arg maxStep 6: Access data dx.Step 7: Determine true importance and redundancy of xtStep 8: Update parameters 9, p based on true values ENDWHILEStep 9: Return dXi, ... , dXi
[0055] For instance, the computing system 106 can perform Algorithm 1. For example, the computing system 106 can obtain the target data dy. the target metadata my. and / or the feature metadata mx, x G X. At a first step, the computing system 106 initializes model parameters 9, <p (e.g., the computing system 106 initializes the feature importance and redundancy models, and the model parameters are associated with the feature importance model and the feature redundancy model). The feature importance model indicates an importance of the feature and the feature redundancy model indicates the feature’s association (e.g., redundancy) with previously selected features. At a second step, the computing system 106 initializes an iteration counter. At a third step, the computing system 106 checks a stopping criterion, and performs a while loop based on whether the stopping criterion has been satisfied (e.g., based on comparing the iteration counter with a pre-defined threshold). At a fourth step, the computing system 106 increments the iteration counter by 1. At a fifth step, the computing system 106 performs aselection process based on the feature importance model and the feature redundancy model. At a sixth step, the computing system 106 accesses feature data identified by the selection process (e.g., retrieves and / or obtains the feature data from a data source 102). At a seventh step, the computing system 106 determines the true importance and redundancy of the feature, x,. At an eighth step, the computing system 106 updates the model parameters based on the true values (e.g., the true importance and redundancy). The computing system 106 then checks on whether the stopping criterion has been satisfied. If not, the computing system 106 repeats steps 4-8. For instance, if the stopping criterion is not satisfied, the computing system 106 can perform step 5 (e.g., maximized a product of a predicted feature importance associated with the feature importance model and a predicted feature redundancy associated with the feature redundancy model), step 6 (e.g., access new feature data), step 7 (e.g., determine a measured true importance value and a measured true redundancy value based on maximizing the product of the predicted feature importance), and step 8 (e.g., update the feature importance model and the feature redundancy model based on the measured true importance value and the measured true redundancy value).
[0056] Based on satisfying the stopping criterion (e.g., the iteration counter reaches the predefined threshold), at a ninth step, the computing system 106 returns the obtained feature data.
[0057] In some instances, after obtaining the feature data, the computing system 106 can use the feature data for one or more tasks and / or provide the feature data to an external system that can perform the one or more tasks (e.g., predicting physical properties of the building and / or controlling heating / cooling of the building according to the prediction, predicting physical properties such as flood and / or indicating changes / improvements to be infrastructure based on the prediction, predicting properties such as an overheat status and / or performing automated actions such as maintenance or scheduling thereof based on the prediction, predicting physical properties of a hospital and / or controlling heating and cooling or patient booking system based on the predictions and heating, ventilation, and air conditioning (HVAC) control, and / or perform other tasks). For instance, the computing system 106 can input the feature data (as features) into one or more machine learning (ML) - artificial intelligence (Al) models to perform one or more tasks. This will be described in further detail below.
[0058] FIG. 2 shows a visualization of metadata subspace 200 with the target y 202 and three features 204-208. Features with a low distance to the target have a high relevancy, and features close to other (already selected) features have a high redundancy. The algorithm (e.g., computing system 106) first selects x1204, as it is predicted to be highly relevant. In the second iteration, it is likely that the algorithm (e.g., computing system 106) chooses x3208, because x2206 has a high redundancy in face of x1204 already selected.
[0059] A number of extensions to according to embodiments of the present invention are discussed in the following.
[0060] Further parameterizations: It is possible to parameterize also the meta-distance function. For example, the meta-distance function can be the weighted sum of several different distance functions defined on the metadata space, and the weights can be further parameters. In step 8 of the Algorithm 1, those additional parameters can be updated alongside 9, p in an end- to-end manner. This mechanism provides more flexibility to the prediction models.
[0061] For instance, the computing system 106 can use multiple distance functions to determine multiple distance metrics. The computing system 106 can then determine weights corresponding to the distance functions and associated distance metrics. Based on the weights and the distance metrics, the computing system 106 can determine a cumulative distance metric. The computing system 106 can use the cumulative distance metric as described above (e.g., use the cumulative distance metric to determine feature importance and feature redundancy). Additionally, and / or alternatively, the computing system 106 can update the weights corresponding to the distance functions and associated distance metrics. For instance, at step 8, based on the true values (e.g., the true importance and redundancy), the computing system 106 can update the weights (e.g., by using gradient descent).
[0062] Additional feature extraction step: It is possible to perform feature extraction computations on each accessed feature data dxusing feature extraction functions fltAccess to dxleads to having access to k additional features / i(dx), ... fk(dx). In that case, the feature importance and feature redundancy for each of those k derived features can be determined in step 7 of the Algorithm 1, and true redundancy and importance of the best feature among those k can be used to update the prediction model parameters in step 8 of the Algorithm 1.
[0063] For instance, for each accessed data step (e.g., step 6), the computing system 106 can perform feature extraction computations on the accessed feature data based on one or more feature extraction functions. As such, the computing system 106 can determine a number (“£”) of derived features from the accessed feature data. At step 7, the computing system 106 can determine feature importance and feature redundancy for each of the derived features (e.g., true importance and true redundancy of the features). Then, the computing system 106 can update the prediction model parameters 9, p and / or weights corresponding to the distance functions and associated distance metrics based on the true importance and true redundancy of the features.
[0064] Final feature selection: After execution of the Algorithm 1, the set of features can further be reduced using an existing feature selection method, e.g., recursive feature elimination. Such methods are applicable to the set of features already accessed.
[0065] For instance, in a new step 10, the computing system 106 can perform a feature selection method (e.g., recursive feature elimination) to further reduce the set of features.
[0066] Sensor placement: Using the trained model of feature importance and redundancy, the process of selecting sensors can be optimized. The model can be used to compute the metadata that would maximize / minimize the feature importance / redundancy, and then sensors can be installed that exhibit the properties (e.g., type, location) described in the computed optimal metadata.
[0067] For instance, the computing system 106 can use the trained models (e.g., the trained models for feature importance and redundancy) to determine metadata that would maximize / minimize the feature importance / redundancy. For example, after training the models (e.g., by updating the parameters associated with the true values at step 8 one or more iterations), the computing system 106 can determine the models are trained. Then, the computing system 106 can use the trained models to determine metadata that maximizes and / or minimizes the feature importance / redundancy. The computing system 106 can provide this metadata to an external system (e.g., the original data source 102 that provided the sensor information and / or another system that manages and / or is otherwise associated with the original data source 102). The external system can cause display the properties described in the optimal metadata (e.g., the optimal type and / or location of the sensor) and / or perform other tasks based on the optimal metadata.
[0068] Embodiments of the present invention thus provide for general improvements to computers in machine learning systems to provide for efficient and improved feature selection from remote data sources. Moreover, embodiments of the present invention can be practically applied to use cases to effect further improvements in a number of technical fields including, but not limited to, medical (e.g., digital medicine, healthcare, Al-assisted drug or vaccine development, diagnosis and treatments, etc.), material development, public safety and smart cities (e.g., automated traffic or vehicle control, smart districts, smart buildings, smart industrial plants, smart agriculture, energy networks and management, etc.). In particular, embodiments of the present invention can be advantageously applied to improve any machine learning system using remote data sources.
[0069] In an exemplary embodiment, the present invention can be practically applied for smart building carbon management. A use case is that the heating, ventilation and air conditioning (HVAC) system of buildings are to be automatically operated based on predictions of the occupancy of each room, as well as predictions of the outdoor temperature and the window-opening habits of occupants. This is happening, for example, in a large distributed system where the data of several buildings is managed (e.g.. city district, campus). All availabledata sources (e.g., cameras, air quality sensors, indoor and outdoor temperature sensors, window status sensors, light switch status) are providing potential features. Application of the method according to an embodiment of the present invention provides to select relevant features from this large set for each prediction model. A technical effect includes data flowing through the system to connect the prediction models to the feature sources. As output, the model predicts physical properties of the building. As automated actions or technicity, the HVAC system controls the heating and cooling of the building according to the prediction.
[0070] In an exemplary embodiment, the present invention can be practically applied for a disaster resilience index for a city, which is a system that computes performance indicators for disaster preparedness and resilience (e.g., risks of flooding, risks of buildings collapsing, etc.) from available data like weather information, traffic flow data, noise sensors, historical disaster information, building and other infrastructure information. A use case is to improve the coverage of the index, where prediction models are to be used in areas where insufficient data is available. The prediction models can be potentially fed with all existing data sources, but in larger cities this is not feasible. Application of the method according to an embodiment of the present invention provides to select features for the prediction model based on available metadata. A technical effect includes data flowing through the system as a result of the data access determined by the methodology. As output, the model predicts physical properties (e.g., floods). As automated actions or technicity, changes or improvements to infrastructure can be made or scheduled.
[0071] In an exemplary embodiment, the present invention can be practically applied for energy networks. In an energy network, the state of equipment like switches, transformers, etc. is monitored and predictive maintenance is performed to ensure the reliable operation of the network. Prediction models are used to forecast failure of equipment. The prediction models can be potentially fed with all existing data sources, including voltage, current, temperatures as well as conditions like solar irradiations and wind. Application of the method according to an embodiment of the present invention provides to select features for the prediction model about device failure, based on available metadata. A technical effect includes data flowing through the system as a result of the data access determined by the methodology. As output, the model predicts physical properties (e.g. overheat status). As automated actions or technicity, maintenance or scheduling thereof can be performed as a result of the predictions.
[0072] In an exemplary embodiment, the present invention can be practically applied in healthcare. A use case here is that the HVAC and scheduling systems of hospitals are to be automatically operated based on predictions of the usage of each medical unit, as well as predictions of the indoor and outdoor conditions and the needs of patients. The data of severalbuildings of a hospital campus is managed. All available data sources (e.g., cameras, air quality sensors, indoor and outdoor temperature sensors, window status sensors, light switch status, patient booking systems, healthcare records) are providing potential features. Application of the method according to an embodiment of the present invention provides to select relevant features from this large set for each prediction model. A technical effect includes data flowing through the system to connect the prediction models to the feature sources. As output, the model predicts physical properties of the hospital. As automated actions or technicity, the HVAC system controls the heating and cooling accordingly, and the patient booking system can also be adjusted in an automated manner based on the predictions and / or HVAC system control.
[0073] In an embodiment, the present invention provides a method for efficiently selecting features in a distributed system, the method comprising the steps of:1) Initializing models of feature importance and feature redundancy as functions of the feature distance in the metadata space.2) Selecting a feature that maximizes the feature importance and accessing its data.3) Measuring the true feature importance of the selected feature and updating the feature importance model parameters.4) Selecting further features by maximizing the product of predicted feature importance and predicted feature redundancy (against all previously selected features), each time updating the importance and redundancy prediction models based on the measured true importance and redundancy values.
[0074] Embodiments of the present invention provide for the following improvements and technical advantages over existing technology:1) Maintaining a model of feature importance and a model for feature redundancy, using the models for deciding about features to select, and updating the models as feature data is accessed.2) Modeling both feature importance and feature redundancy as trainable functions of the distance in the metadata space.3) Estimating the value of a feature as the product of importance and redundancy against all features previously selected.4) Providing for efficient feature selection in distributed systems where accessing the data of each feature is associated with a cost. In contrast, existing feature selection methods assume that all feature data is available and thus are not applicable in this situation.
[0075] Referring to FIG. 3, a processing system 300 can include one or more processors 302, memory 304, one or more input / output devices 306, one or more sensors 308, one or moreuser interfaces 310, and one or more actuators 312. Processing system 300 can be representative of each computing system disclosed herein.
[0076] Processors 302 can include one or more distinct processors, each having one or more cores. Each of the distinct processors can have the same or different structure. Processors 302 can include one or more central processing units (CPUs), one or more graphics processing units (GPUs), circuitry (e.g., application specific integrated circuits (ASICs)), digital signal processors (DSPs), and the like. Processors 302 can be mounted to a common substrate or to multiple different substrates.
[0077] Processors 302 are configured to perform a certain function, method, or operation (e.g., are configured to provide for performance of a function, method, or operation) at least when one of the one or more of the distinct processors is capable of performing operations embodying the function, method, or operation. Processors 302 can perform operations embodying the function, method, or operation by, for example, executing code (e.g., interpreting scripts) stored on memory 304 and / or trafficking data through one or more ASICs. Processors 302, and thus processing system 300, can be configured to perform, automatically, any and all functions, methods, and operations disclosed herein. Therefore, processing system 300 can be configured to implement any of (e.g., all of) the protocols, devices, mechanisms, systems, and methods described herein.
[0078] For example, when the present disclosure states that a method or device performs task “X” (or that task “X” is performed), such a statement should be understood to disclose that processing system 300 can be configured to perform task “X”. Processing system 300 is configured to perform a function, method, or operation at least when processors 302 are configured to do the same.
[0079] Memory 304 can include volatile memory, non-volatile memory, and any other medium capable of storing data. Each of the volatile memory, non-volatile memory, and any other type of memory can include multiple different memory devices, located at multiple distinct locations and each having a different structure. Memory 304 can include remotely hosted (e.g., cloud) storage.
[0080] Examples of memory 304 include a non-transitory computer-readable media such as RAM, ROM, flash memory, EEPROM, any kind of optical storage disk such as a DVD, a Blu- Ray® disc, magnetic storage, holographic storage, a HDD, a SSD, any medium that can be used to store program code in the form of instructions or data structures, and the like. Any and all of the methods, functions, and operations described herein can be fully embodied in the form of tangible and / or non-transitory machine-readable code (e.g., interpretable scripts) saved in memory 304.
[0081] Input-output devices 306 can include any component for trafficking data such as ports, antennas (i.e., transceivers), printed conductive paths, and the like. Input-output devices 206 can enable wired communication via USB®, DisplayPort®, HDMI®, Ethernet, and the like. Input-output devices 306 can enable electronic, optical, magnetic, and holographic, communication with suitable memory 304. Input-output devices 306 can enable wireless communication via WiFi®, Bluetooth®, cellular (e.g., LTE®, CDMA®, GSM®, WiMax®, NFC®), GPS, and the like. Input-output devices 306 can include wired and / or wireless communication pathways.
[0082] Sensors 308 can capture physical measurements of environment and report the same to processors 302. User interface 310 can include displays, physical buttons, speakers, microphones, keyboards, and the like. Actuators 312 can enable processors 302 to control mechanical forces.
[0083] Processing system 300 can be distributed. For example, some components of processing system 300 can reside in a remote hosted network service (e.g., a cloud computing environment) while other components of processing system 300 can reside in a local computing system. Processing system 300 can have a modular design where certain modules include a plurality of the features / functions shown in FIG. 3. For example, I / O modules can include volatile memory and one or more processors. As another example, individual processor modules can include read-only-memory and / or local caches.
[0084] While subject matter of the present disclosure has been illustrated and described in detail in the drawings and foregoing description, such illustration and description are to be considered illustrative or exemplary and not restrictive. Any statement made herein characterizing the invention is also to be considered illustrative or exemplary and not restrictive as the invention is defined by the claims. It will be understood that changes and modifications may be made, by those of ordinary skill in the art, within the scope of the following claims, which may include any combination of features from different embodiments described above.
[0085] The terms used in the claims should be construed to have the broadest reasonable interpretation consistent with the foregoing description. For example, the use of the article “a” or “the” in introducing an element should not be interpreted as being exclusive of a plurality of elements. Likewise, the recitation of “or” should be interpreted as being inclusive, such that the recitation of “A or B” is not exclusive of “A and B,” unless it is clear from the context or the foregoing description that only one of A and B is intended. Further, the recitation of “at least one of A, B and C” should be interpreted as one or more of a group of elements consisting of A, B and C, and should not be interpreted as requiring at least one of each of the listed elements A, B and C, regardless of whether A, B and C are related as categories or otherwise. Moreover, therecitation of “A, B and / or C” or “at least one of A, B or C” should be interpreted as including any singular entity from the listed elements, e.g., A, any subset from the listed elements, e.g., A and B, or the entire list of elements A, B and C.
Claims
CLAIMSWhat is claimed is:
1. A computer-implemented method for efficiently selecting features in a distributed system comprising a plurality of remote data sources and a computing system, the computer- implemented method comprising: initializing, by the computing system, a feature importance model and a feature redundancy model; selecting, by the computing system, a feature that maximizes an output from the feature importance model; accessing, by the computing system, feature data associated with the feature, wherein the feature data is stored in one of the remote data sources; updating, by the computing system, parameters of the feature importance model based on determining a true feature importance associated with the selected feature; and selecting, by the computing system, one or more further features based on the feature importance model and the feature redundancy model.
2. The computer-implemented method of claim 1, wherein the plurality of remote data sources are a plurality of sensors located at a plurality of different geographical locations, wherein the feature data is sensor data that is obtained by the sensor, wherein selecting the one or more further features is based on metadata associated with the sensor data, and wherein initializing the feature importance model and the feature redundancy model comprises initializing the feature importance model and the feature redundancy model as functions of a feature distance in a metadata space.
3. The computer-implemented method of claim 1 or 2, wherein selecting the one or more further features comprises: performing a selection process to select a new feature from a plurality of features based on the feature importance model and the feature redundancy model; obtaining, from another one of the remote data sources, second feature data associated with the new feature; updating the parameters of the feature importance model and the feature redundancy model based on the second feature data; and determining whether to perform another selection process based on the second feature data.
4. The computer-implemented method of claim 3, wherein performing the selection process is based on maximizing a product of a predicted feature importance associated with the featureimportance model and a predicted feature redundancy associated with the feature redundancy model.
5. The computer-implemented method of claim 4, wherein the feature importance model comprises an importance (IMP) function, wherein the feature redundancy model comprises a redundancy (RED) function, and wherein performing the selection process comprises: using the IMP function to determine the predicted feature importance; using the RED function to determine the predicted feature redundancy; and selecting the new feature based on maximizing a product of the predicted feature importance and the predicted feature redundancy.
6. The computer-implemented method of claim 5, wherein the parameters comprise a first parameter and a second parameter, and wherein the method further comprises: initializing the first parameter and the second parameter based on a maximum distance between any pair of features, of the plurality of features, in a metadata space, wherein using the IMP function to determine the predicted feature importance is based on the first parameter, and wherein using the RED function to determine the predicted feature redundancy is based on the second parameter.
7. The computer-implemented method of claim 5 or 6, wherein using the IMP function comprises inputting a first distance metric associated with a given feature, from the plurality of features, and a prediction target in the metadata space into the IMP function to determine the predicted feature importance, and wherein using the RED function comprises inputting a second distance metric associated with two features, from the plurality of features, into the RED function to determine the predicted feature redundancy.
8. The computer-implemented method of any of claim 3-7, wherein selecting the one or more further features further comprises: determining a measured true importance value and a measured true redundancy value based on the second feature data, and wherein updating the parameters of the feature importance model and the parameters of the feature redundancy model are based on the measured true importance value and the measured true redundancy value.
9. The computer-implemented method of claim 8, wherein determining the measured true importance value comprises determining the measure true importance value based on the second feature data and prediction target feature data, and wherein determining the measured true redundancy value is based on the second feature data and the feature data associated with the feature.
10. The computer-implemented method of any of claims 3-9, wherein performing the selection process comprises: obtaining first metadata associated with sets of feature data associated with a plurality of features; obtaining second metadata associated with a prediction target; determining one or more distance metrics based on the first metadata, the second metadata, and a meta-distance function; and performing the selection process based on the one or more distance metrics.
11. The computer-implemented method of claim 10, wherein determining one or more distance metrics comprises: using a pre-trained language model to determine a first vector representation associated with the first metadata and a second vector representation associated with the second metadata; and determining the one or more distance metrics based on the first vector representation, the second vector representation, and the meta-distance function.
12. The computer-implemented method of claim 10, wherein performing the selection process further comprises: inputting the one or more distance metrics into the feature importance model to determine a predicted feature importance; inputting the one or more distance metrics into the feature redundancy model to determine a predicted feature redundancy; and determining the new feature based on the predicted feature importance and the predicted feature redundancy.
13. The computer-implemented method of any of claims 3-12, wherein determining whether to perform another selection process comprises: comparing an iteration counter indicating a number of iterations that have been executed with a first stopping criterion threshold; comparing a true redundancy value with a second stopping criterion threshold; and determining to perform another selection process based on the comparisons.
14. A computer system for efficiently selecting features in a distributed system comprising a plurality of remote data sources and the computer system, the computer system comprising one or more hardware processors, which, alone or in combination, are configured to provide for execution of the following steps: initializing a feature importance model and a feature redundancy model; selecting a feature that maximizes an output from the feature importance model;accessing feature data associated with the feature, wherein the feature data is stored in one of the remote data sources; updating parameters of the feature importance model based on determining a true feature importance associated with the selected feature; and selecting one or more further features based on the feature importance model and the feature redundancy model.
15. A tangible, non-transitory computer-readable medium having instructions thereon which, upon being executed by one or more processors, alone or in combination, provide for execution of a method for efficiently selecting features in a distributed system comprising a plurality of remote data sources and a computing system, the method comprising the following steps: initializing a feature importance model and a feature redundancy model; selecting a feature that maximizes an output from the feature importance model; accessing feature data associated with the feature, wherein the feature data is stored in one of the remote data sources; updating parameters of the feature importance model based on determining a true feature importance associated with the selected feature; and selecting one or more further features based on the feature importance model and the feature redundancy model.