Generating embedding space from training dataset of machine learning algorithm

By generating embedding spaces and comparing them with weighted induced data points, the problem of difficult performance control when data specificity is not taken into account is solved, and a combination of high performance and low storage computing requirements is achieved.

CN120020830APending Publication Date: 2025-05-20GENERAL ELECTRIC CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411591076.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-17
Filing Date
2024-11-08
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

The lack of performance control of existing machine learning algorithms in the training data sets makes it difficult to effectively control model performance without taking into account data specificity.

Method used

By generating the embedding space, using weighted induced data points and corresponding metadata to represent the portion of the space, each new inferred or manipulated data is compared only with the induced data points to determine reasonable confidence and uncertainty.

Benefits of technology

Achieving significantly reduces storage and compute requirements while maintaining high performance, allowing the use of black box pre-trained models in edge deployments and guaranteeing secure performance in limited computing environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120020830A_ABST
    Figure CN120020830A_ABST
Patent Text Reader

Abstract

The invention relates to generating an embedding space from a training data set of a machine learning algorithm. A system and method for generating an embedding space from a training dataset of a machine learning algorithm. In one embodiment, a method includes receiving input operational data points by one or more layers of a machine learning algorithm implemented on a computer system through an input of the computer system, generating an embedding space using a training data set, the embedding space including a plurality of inductive data points, each inductive data point corresponding to a cluster of training data points, an input operational data point is compared to a plurality of induction data points, whether there is a sufficient set of support historical data points among a plurality of training data points within a fixed distance from the input operational data point is determined, and an alert is generated to enable a user or an electronic controller to take an appropriate action based on the determination, and the alert is output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to machine learning algorithms and, in particular, to methods and systems for generating an embedding space from a training data set of a machine learning algorithm. Background Art

[0002] Machine learning (ML) algorithms are part of the entire field of artificial intelligence and can be used to automatically learn from data (e.g., progressively improve the performance of a specific task) without explicit programming. For example, machine learning algorithms can be used to control physical assets, where a machine learning model can be trained with relevant historical data (e.g., a training data set). However, current ML algorithms can lack any control in the sense that the training data set (even relatively large) does not account for the particularities of the data. Machine learning models include neural networks (where an embedding can be determined for each layer of the network), decision trees / forests (where an embedding can be determined for each depth of the tree), and monolithic models (where an embedding can be determined only for the input space). Brief Description of the Drawings

[0003] As illustrated in the drawings, the foregoing and other features and advantages will become apparent from the following more particular description of various exemplary embodiments, where like reference numerals generally indicate like, functionally similar, and / or structurally similar elements.

[0004] Figure 1 A flowchart of a system 100 for generating an embedding space from a training data set of a machine learning algorithm according to an embodiment of the present disclosure is shown.

[0005] Figure 2 A graph showing the induction of data points from a training data set according to an embodiment of the present disclosure is shown.

[0006] Figure 3 A graph of the scores of trust (i.e., "I know" or IK) data points relative to the number of induced data points according to an embodiment of the present disclosure.

[0007] Figure 4 An illustration of a computer system for implementing a system and method for generating an embedding space from a training data set of a machine learning algorithm according to an embodiment of the present disclosure. Detailed Description

[0008] The features, advantages, and embodiments of the present disclosure are set forth or apparent by considering the following detailed description, the drawings, and the claims. Further, the following detailed description is exemplary and is intended to provide further explanation without limiting the scope of the present disclosure as claimed.

[0009] The various embodiments of the present disclosure are discussed in detail below. Although specific embodiments are discussed, this is for illustrative purposes only. Those skilled in the relevant art will recognize that other components and configurations may be used without departing from the spirit and scope of the present disclosure.

[0010] Here and throughout the specification and claims, range limitations are combined and interchanged. Such ranges are identified and include all sub-ranges subsumed therein unless the context or language indicates otherwise. For example, all ranges disclosed herein include the endpoints, and the endpoints may be combined independently of each other.

[0011] Machine learning models can use cognitive reasoning to control their performance. In cognitive reasoning, an explicit support set is identified for each new inference, and thus the current decision is associated with the concurrent historical decisions. When a sufficient number of concurrent historical decisions exist in the support set, a reasonable confidence is judged to exist, as defined by the historical decisions within a fixed distance of the new inference. Under the control of cognitive reasoning, when reasonable confidence is judged, the model produces a non-zero output, and when it is not reasonable confidence, it returns a zero output.

[0012] Embodiments of the present disclosure further extract, select, or filter a training data set to provide an embedding space. In the embedding space, portions of the space are represented with weighted (representative) induced data points of the training data set and corresponding metadata regarding the local historical point distribution. In this way, each new inference or new operational data is compared only with the induced data points to determine whether reasonable confidence is judged to occur consistently with the training data set and to estimate uncertainty. The cardinality of the induced data points in the embedding space is generally at least one order of magnitude smaller than the cardinality of the full training data set. Cardinality is the number of elements in a given data set. For example, the induced data points may represent a subset of the training data points with some Gaussian representation. The induced data points essentially represent local density regions within the training data set.

[0013] Machine learning models that can be used for security control in this way using the embedding space include neural networks (where an embedding can be determined for each layer of the network), decision trees / forests (where an embedding can be determined for each depth of the tree), and monolithic models (where an embedding can be determined only for the input space).

[0014] In embodiments of the present disclosure, a sparse Gaussian process, a Gaussian mixture model, or a random draw of the training set can be used to generate the embedding space. However, additional methods can also be used. In a sparse Gaussian process, the tunable control is the number of induced data points. In a Gaussian mixture model, the tunable control is the number of mixture components. Each component generates an induced data point as its centroid. In a random draw of the training data set, each induced data point is a selected data point exemplar. Each of these methods shows significant improvements in terms of size, weight, and power (SWAP), which are measured by the number of historical data points required to be stored in a retrievable memory, the number of comparisons between the inferred points and the historical data points, and the total time required to complete all such comparisons, while maintaining cognitive performance relative to the current state-of-the-art baseline. In contrast, machine learning algorithms typically lack explicit performance control, as in traditional black-box neural networks.

[0015] The determination to use an embedding-based method is novel. For example, as in the following case: data points determined by a sparse Gaussian process are specifically selected for the embedding to ensure maximum SWaP gain while maintaining the ability to perform Bayesian uncertainty quantification. Examples of Gaussian processes and Gaussian distributions used in the embedding space are discussed throughout this document. However, the present disclosure is not limited to Gaussian constructs and can include various methods using other functions such as Lorentzian functions or other distribution functions, or a mixture model that determines Gaussian or non-Gaussian distribution functions via alternative approximation methodologies. The present method and system can also use random selection of points. Any representation or model that captures the density distribution of the training data can be used to generate the embedding space. Thus, given a training data set, one or more density centroids within the training data set can be selected and represented by any function, Gaussian, Lorentzian, or other distribution. The density centroids can then be used as representative points of the training data set, and the centroid points corresponding to the induced data points are used instead of the training data points. Instead of comparing the new operational data points (input data points) with each example data point in the training data set, the new operational data points are only compared with the stored induced data points (e.g., centroid points). Those induced data points are representatives of the training data set, and they have stored the corresponding local density information that can be used to determine the sufficiency of the corresponding support set and reasonable confidence with them.

[0016] By learning an embedding space in the form of a set of weighted induced data points (e.g., Gaussian centroid selection points), the inventors are able to represent the cognitive features associated with the training data set (for determining the sufficiency of support). If representative associated historical data points are desired, those historical data points can also be stored while still significantly reducing SWAP, including the representation of the induced points, multiple additional historical data points from the training set corresponding to that induced point, and the local distribution statistics of the training set around the induced point. In other words, the induced data points can provide a sufficient statistical data set to globally characterize the training data set and, if needed, generate representative local examples around the induced points. Thus, new input points are only compared to the induced data points to determine if they are close enough to a sufficient number of historical examples, which includes a support set with reasonable confidence of judgment.

[0017] When considering the application of induced data points in support-based reasoning, the covariance of the induced data points can also be estimated to efficiently estimate the uncertainty of each induced data point.

[0018] The comparison of each new operational data point to the training data set allows for determining whether the model needs to extrapolate in response to the operational data point, whether the training data set corresponds to a region of low sample density, and / or whether the training data set corresponds to a region of output uncertainty. Traditional methods perform an exhaustive comparison of each new operational data point to all training data points, or at least to all points including the convex hull of the training set. In contrast, the present disclosure embeds selected training data points into a density-based generalization via weighted induced data points (e.g., Gaussian-based weighting). In this way, fewer comparisons may be needed since the comparisons are made against the number of induced data points rather than the total number of training data points.

[0019] Evaluations were performed on open-source benchmarks including the COCO (Common Objects in Context) and CIFAR-10 data sets. The CIFAR-10 data set consists of 60,000 32×32 color images in ten classes and is commonly used for training machine learning and computer vision algorithms. The present method consistently shows at least as good accuracy as the original (non-cognitive) implementation, with coverage approaching that of the previous state-of-the-art cognitive systems, while using only 5% of the storage and computational power.

[0020] The present method and system provide approximately a 90% reduction in SWAP required for visual tasks while maintaining high performance. The present method and system enable the immediate use of industry-leading black-box pre-trained models in edge deployments (previously infeasible due to storage and computational requirements).

[0021] For example, the present method provides for low-SWAP trusted deployment of neural networks and other effective models, which are pre-trained and / or packaged for use without providing the full training set, hyperparameters, or other details of their creation. Even when the information is available, it may not be necessary to store it on the edge device, thereby allowing a significant reduction in the computational footprint of trusted execution. As a result, high-performance models can be deployed while ensuring security performance even in very limited computational environments. This can allow for convenient trusted autonomy applications in multi-party settings for commercial and military supply.

[0022] Referring now to the drawings, Figure 1 FIG. 5 shows a flowchart of a system 100 for generating an embedding space from a training data set of a machine learning algorithm, according to an embodiment of the present disclosure. System 100 includes an input 102. Input 102 receives new operational data. The new operational data may be generated by a device 101. Device 101 may be, for example, a sensor device (e.g., an image sensor, an electro-optical sensor, an infrared sensor, a ultraviolet sensor), a radar detection system (e.g., a synthetic aperture radar (SAR)), etc. Thus, the new operational data received by input 102 may be remote sensing data, such as image data, infrared data, electro-optical data, radar data, etc. The sensor device may be disposed on an aircraft (manned or unmanned). In an unmanned aircraft, the sensor may be part of an unmanned aircraft system (UAS). For example, the sensor device may be used for object detection, such as target detection. The machine learning algorithm includes one or more layers 110. Each of the one or more layers 110 provides a transformation from an input space to an output space, where the cardinality of the input space is equal to the cardinality of the output space of the previous layer (if the previous layer exists). Input 102 is connected to a first layer 104 (layer 1) of the one or more layers 110 of the machine learning algorithm implemented on a computer system 103.

[0023] Input 102 receives new operational data from device 101 and provides or passes the new operational data to the first layer 104 (layer 1) of a machine learning algorithm (e.g., a neural network) implemented on computer system 103. The first layer 104 of the machine learning algorithm will produce an output 105. The output 105 is passed to the next layer of one or more layers 110 of the machine learning algorithm, and so on. For example, the output 105 is passed to the second layer of one or more layers 110, and the second layer will in turn produce another output that is passed to the next layer. This process can be repeated until the last layer 106 (layer k) of one or more layers 110 is reached. The last layer 106 will produce an output 108. The output 108 corresponds to the output of the machine learning algorithm (e.g., a neural network) implemented on computer system 103. Thus, input 102 receives new operational data from device 101 and continuously provides or passes the new operational data to one or more layers 110. One or more layers 110 are serially linked and output the result through output 108. The input space of the first layer 104 is the space corresponding to the original data from device 101 (e.g., image data or radar data from sensors). The output space of the last layer 106 is the output space of the machine learning algorithm (e.g., a neural network) and corresponds to output 108. In a neural network, each layer of one or more layers 110 corresponds to parallel neurons (random variables). Each layer of one or more layers 110 performs a functional transformation using a training data set accessed by computer system 103. The training data set is used to train the machine learning algorithm executed or run on computer system 103.

[0024] In system 100, in addition to processing input 102 through the first layer 104 (layer 1), input 102 is also passed to an embedding comparison in an embedding space 112 generated by the first layer 104 (layer 1). The first layer 104 (layer 1) includes the generated embedding space 112. The embedding space 112 is generated using a training data set. The embedding space 112 is used as a test of whether a representative sample from the training data set has sufficient support in a demonstration set from an embedding vector corresponding to input 102 corresponding to the new operational data.

[0025] Each layer of one or more layers 110 can independently perform the same process as the first layer 104. For example, the last layer 106 (layer k) can also include a defined embedding space 114. The embedding space 114 is generated using a training data set. The embedding space 114 corresponds to the output space of the last layer 106 (layer k).

[0026] Thus, at each layer of one or more layers 110, the input is processed as designed for the model structure of a machine learning algorithm. Additionally, in parallel, by determining the embedding vector corresponding to the input (e.g., input 102), the vector of the input (e.g., input 102) corresponding to the new operational data is compared with the representative samples of the training set corresponding to the induced data points to determine whether there is a sufficient (validated) set of historical data points within a fixed distance from the vector, to perform a test to test the input of each layer of one or more layers 110 (e.g., the first layer 104). If this test passes, the operation continues to the next layer. If the test does not pass, the operation is cancelled and a zero value is returned as output 108 (indicating that the attempted model performance exceeds its capabilities).

[0027] In an embodiment, when performing the comparison of the new operational data point with the embedding space 112 and / or the embedding space 114, typically one layer (e.g., the first layer 104) is used. However, if more than one layer is used, the comparison can be performed for all layers. If the test does not pass for any of the layers, a zero value is returned as output 108.

[0028] Although as shown in Figure 1 the generation of the embedding space and the comparison between the new operational data and the embedding space are performed, for example, using the first layer 104 (layer 1) of a machine learning algorithm (e.g., a neural network) or the last layer 106 (layer k) of a machine learning algorithm (e.g., a neural network) to capture the factors relevant to the comparison, each layer of one or more layers 110 can be independently processed to provide the performance of a specific data point.

[0029] Thus, each layer of one or more layers 110 has a primary output (e.g., output 105 of the first layer 104) and a by - product or secondary output 116 of the comparison between the operational data and the embedding space 112. For example, the secondary output 116 of the comparison determines whether there is a sufficient support or justification set 118 for the embedding space 112. The term "justification" as used herein means that there is a sufficient (valid) set of support demonstrations within a fixed distance from the vector of the input 102 corresponding to the new operational data point. The secondary output 116 can also include a predictive distribution 120. The predictive distribution 120 is the result of quantifying a specific selection of sparse Gaussian data points in the embedding space in the first layer 104 to ensure the maximum SWAP gain using, for example, Bayesian uncertainty quantification and the corresponding posterior probability of the operational data.

[0030] The comparison can be implemented at each of one or more layers 110, and a rationale 118 is implemented for each of the one or more layers 110. Similar to the secondary output 116 of the comparison of the embedding space 112, a secondary output 117 can also be generated for the comparison of the embedding space 114. Similarly, a rationale 119 can also be provided within the secondary output 117, similar to the rationale 118 provided in the secondary output 116. In addition, the secondary output 117 can further include a predictive distribution 121. The predictive distribution 121 is the result of quantifying a specific selection of sparse Gaussian data points in the embedding space in the last layer 106 to ensure the maximum SWAP gain by using, for example, Bayesian uncertainty quantification and the corresponding posterior probability of the manipulated data.

[0031] If a rationale 118 for the first layer 104 (layer 1) is found, but no rationale 119 for the last layer 106 (layer k) is found, the output 108 will be set to zero. A rationale (e.g., rationale 118, rationale 119) is required to exist for each of the one or more layers 110 for the output 108 to be set to a non-zero value.

[0032] In an embodiment, for example, if no rationale 118 is found, the computer system 103 generates and sends an alert 150 to the user of the system 100 to indicate that no rationale 118 has been found. In this case, the alert 150 can take the form of, for example, a flashing red light, the letter "NO-GO" displayed on the monitor of the computer system 103, an alarm sound, etc. On the other hand, if a rationale 118 is found, the computer system 103 generates an alert 150 and sends the alert 150 to the user of the system 100 to indicate that a rationale has been found. In this case, the alert 150 can take the form of, for example, a flashing green light, the letter "GO" displayed on the monitor of the computer system 103, another alarm sound, etc. This alert can include historical data points within the distance of the manipulated data points. This alert can also include representative data points within the distance of the manipulated data points, such as those present in the hypothesis support set of the confidence in judging the rationale. The user can then take an action 152 (appropriate or predetermined action) depending on the alert 150.

[0033] For example, for military applications, system 100 can be used to confirm a positive identification of a target. If, upon receiving an alert 150 indicating the finding of a plausibility 118, a user can take an action 152 by activating a trigger to arm the system for remote munitions delivery or other attributable / expendable effects. On the other hand, if, upon receiving an alert 150 indicating the non-finding of a plausibility 118, a user can take an action 152 to abort and not activate the trigger to arm the system for remote munitions delivery. System 100 can also be installed on an unmanned aerial vehicle system (UAS) that utilizes machine learning algorithms to detect a target of interest, based on which an alert 150 is sent to a user (remote operator) who can take an action 152 (appropriate or predefined action).

[0034] In another embodiment, the alert 150 is not sent to the user taking the action 152. Instead, the alert 150 is sent directly to an electronic controller 154 that will take the action 152 without intervention from the user. The electronic controller 154 can be provided, for example, on an unmanned aerial vehicle system (UAS). In this case, the UAS can be said to be autonomous.

[0035] In the embedding space 112, portions of the space are represented by induced data points 130 of a training data set 132 (e.g., class 1 training samples, class 2 training samples). In this way, each new inference or new operational data point 134 (test sample or input 102) is compared to the induced data points 130 to determine whether the training data set 132 provides sufficient support, i.e., plausibility 118, and to estimate uncertainty and provide a predictive distribution 120. The cardinality of the induced data points 130 in the embedding space 112 is generally at least one order of magnitude smaller than the cardinality of the training data set 132. For example, the induced data points 130 represent the training data set 132 with some Gaussian representation. Thus, the induced data points 130 essentially represent local density regions within the training data set 132. For example, as Figure 1 shown, each cluster in the training data set 132 can be represented by one or more induced data points 130. Each cluster of the training data set 132 can be represented by a distribution such as a Gaussian distribution of data points determined by a sparse Gaussian process, where the induced data points 130 can be characterized as, for example, the centroid of the Gaussian distribution.

[0036] Machine learning algorithms that can perform security control in this way using the embedding space 112 include neural networks (where an embedding can be determined for each layer of the network), decision trees / forests (where an embedding can be determined for each depth of the tree), and ensemble models (where an embedding can be determined only for the input space).

[0037] In embodiments of the present disclosure, the embedding space 112 can be generated using a sparse Gaussian process, using a Gaussian mixture model, or using a random draw of the training set. In a sparse Gaussian process, the tunable control is the number of inducing data points. In a Gaussian mixture model, the tunable control is the number of mixture components. Each component generates an inducing data point as its centroid. In a random draw of the training data set, each inducing data point is a randomly selected data point exemplar, and the tunable control is the number of random selections performed. Each of these methods shows significant improvements in terms of size, weight, and power (SWAP), which are measured by the number of historical data points required to be stored in a retrievable memory, the number of comparisons between the inference points and the historical data points, and the total time required to complete all such comparisons, while maintaining cognitive performance relative to the current state-of-the-art baseline. In contrast, traditional black-box neural network machine learning algorithms lack any testing process for performance control.

[0038] The determination to use an embedding-based approach is novel, for example, as in the case where sparse Gaussian data points are specifically selected for the embedding to ensure maximum SWAP gain while maintaining the ability to perform Bayesian uncertainty quantification. As used in the embedding space 112, examples of Gaussian processes and Gaussian distributions are discussed throughout this document. However, the embedding space 112 is not limited to Gaussian constructs and can include various methods using other functions such as Lorentz functions or other distribution functions, or a mixture model that determines Gaussian or non-Gaussian distribution functions via alternative approximation methodologies. The present method and system can also use random selection of points. Any representation or model that captures the density distribution of the training data set 132 can be used to generate the embedding space 112. Thus, given the training data set 132, one or more density centroids corresponding to the inducing data points 130 within the training data set 132 can be selected and represented by any function, Gaussian, Lorentz, or other distribution. Multiple centroids can then be used as representative points for the training data set 132, and the centroid points corresponding to the inducing data points 130 are used instead of the training data set 132. Instead of comparing the new operational data point 134 with each example data point in the training data set 132, the new operational data point 134 is only compared with the stored inducing data points 130 (e.g., centroid points). The inducing data points 130 represent the training data set 132, and they have stored the local density information that can be used with them.

[0039] By learning the embedding space 112 in the form of a set of induced data points 130 (e.g., Gaussian centroid weighted selected data points), the inventors are able to represent the training data set 132 (for determining the sufficiency of support) with relevant cognitive features. If representative associated historical data points are desired, those data points can also be stored while still significantly reducing SWAP, including the representation of the induced points, multiple additional historical data points from the training set corresponding to that induced point, and the local distribution statistics of the training set around the induced point. In other words, the induced data points 130 can provide a sufficient statistical data set to globally represent the training data set 132. Thus, new operational data points 134 are compared with the induced data points 130 to determine whether they are close to multiple historical data points within the training data set, which provide a sufficient support set for reasonable confidence to be judged.

[0040] When considering the application of the induced data points 130 in support-based reasoning, the covariance 140 (e.g., covariance matrix) of the induced data points 130 can also be estimated to efficiently estimate the uncertainty of each induced data point in the induced data points 130.

[0041] The comparison of each new operational data point 134 with the induced data points in the embedding space 112 allows determination of whether the model will need to extrapolate in response to the new operational data point 134, whether the induced data points correspond to regions of low sample density, and / or whether the induced data points correspond to regions of output uncertainty.

[0042] Traditional methods perform an exhaustive comparison of each new operational data point with all of the training data set 132. In contrast, the present disclosure embeds selected training data points into a density-based generalization by way of the induced data points 130 (e.g., Gaussian-based weighting). In this way, fewer comparisons may be needed because the comparisons are made against the number of induced data points 130 rather than the total number of the training data set 132. The total number of the training data set 132 is generally at least an order of magnitude larger than the total number of the induced data points 130.

[0043] Figure 2A graph of induced data points from a training data set according to an embodiment of the present disclosure is shown, where the vertical and horizontal axes indicate that the maximum information from principal component analysis is used to plot historical data points. The horizontal axis corresponds to the "major axis of data variance from principal component analysis". The vertical axis corresponds to the "minor axis of data variance from principal component analysis". The points representing the training data set 200 are distributed throughout the graph and form a "cloud" of points. In the example shown, the number of data points in the training data set 200 is approximately 6700. The training data set 200 is extracted into induced data points 202. In the example shown, the number of induced points is approximately 100. The induced data points 202 represent clusters of data points in the training data set 200. Instead of storing the relatively large (6700 data points) training data set 200, a relatively small set of induced data points 202 is stored while maintaining prediction accuracy.

[0044] Figure 3 A graph of the scores of trust (i.e., "I know" or IK) data points that are determined to have reasonable confidence according to an embodiment of the present disclosure, relative to the number of induced data points. Figure 3 The plotted points in it are obtained from evaluations performed on open-source benchmarks for the COCO (Common Objects in Context) and CIFAR-10 data sets. The CIFAR-10 data set consists of 60,000 32×32 color images in ten classes and is commonly used to train machine learning and computer vision algorithms. The method and system represented by the points and line 300 (labeled "SGP") correspond to a system using sparse Gaussian processing (SGP) for embedding, while the points and line 302 (labeled "Cov-SGP") correspond to a system using a sparse Gaussian process that is penalized by inducing points with high covariance, where the figures always show at least as good accuracy as the original (non-cognitive) random points or line 304. The horizontal dashed line represents the "baseline" performance of a cognitive system that compares all new operational data points with all historical data points, while the vertical dashed line represents the "solution" where a small number of induced points are selected for the implementation of a low-SWAP deployment of the system.

[0045] The method and system provide a reduction of approximately 90% in the SWAP required for visual tasks while maintaining high performance. The method and system enable the immediate use of industry-leading black-box pre-trained models in edge deployments (previously infeasible due to storage and computational requirements).

[0046] For example, the present method provides for the low-SWAP trusted deployment of neural networks and other effective models that are pre-trained and / or packaged for use without providing a full training set, hyperparameters, or other details of their creation. Even when this information is available, it may not be necessary to store it on the edge device, allowing for a significant reduction in the computational footprint of trusted execution. As a result, high-performance models can be deployed while ensuring security performance, even in very limited computational environments. This can allow for convenient trusted autonomy applications in multi-party settings for commercial and military supply.

[0047] Figure 4 is an illustration of a computer system for implementing a system and method for generating an embedding space from a training data set of a machine learning algorithm according to an embodiment of the present disclosure. Referring Figure 4 , the computer system 103 includes a computing device 400 that includes a processing unit (e.g., a CPU or processor) 420 and a system bus 410 that couples various system components including a system memory 430 such as a read-only memory (ROM) 440 and a random access memory (RAM) 450 to the processor 420. The computing device 400 can include a cache of high-speed memory that is directly connected to, adjacent to, or integrated as part of the processor 420. The computing device 400 copies data from the system memory 430 and / or the storage device 460 to the cache for quick access by the processor 420. In this way, the cache provides a performance boost that avoids the latency of the processor 420 waiting for data. These and other modules can control or be configured to control the processor 420 to perform various actions. Other system memory 430 may also be available for use. The system memory 430 can include a variety of different types of memory with different performance characteristics. The present disclosure can operate on a computing device 400 having more than one processor 420 (e.g., a multi-core processor), or on a group or cluster of computing devices networked together in a distributed computing environment to provide greater processing power. The processor 420 can include any general-purpose processor and hardware modules or software modules as well as a dedicated processor, where software instructions are incorporated into the actual processor design. The processor 420 can essentially be a fully self-standing computing system, containing multiple cores or processors, buses, memory controllers, caches, etc. The multi-core processor can be symmetric or asymmetric.

[0048] The system bus 410 can be any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, and a local bus using any of a variety of bus architectures. The basic input / output (BIOS) stored in memory ROM 440, etc., can provide basic routines that assist in transferring information between the components within the computing device 400, such as during startup. The computing device 400 further includes a storage device 460, such as a hard disk drive, a magnetic disk drive, an optical disk drive, a tape drive, etc. Other hardware or software modules are contemplated. The storage device 460 is connected to the system bus 410 through a drive interface. The drive and the associated computer-readable storage medium provide non-volatile storage of computer-readable instructions, data structures, program modules, and other data for the computing device 400. In one aspect, a hardware module that performs a specific function includes software components stored in a tangible computer-readable storage medium, along with the necessary hardware components, such as the processor 420, the system bus 410, the output device 470 (e.g., a display), etc., to perform the function. In another aspect, the system can use a processor and a computer-readable storage medium to store instructions that, when executed by the processor, cause the processor to perform a method or other specific actions. Depending on the type of device, such as whether the computing device 400 is a small handheld computing device, a desktop computer, or a computer server, basic components and appropriate variations are contemplated.

[0049] The storage device 460 can be a hard disk or other computer-readable medium that can store data accessible by a computer, such as a cassette tape, a flash card, a digital versatile disc, a tape cartridge, etc. Tangible computer-readable storage media, computer-readable storage devices, or computer-readable memory devices expressly exclude media such as transient waves, energy, carrier signals, electromagnetic waves, and signals themselves.

[0050] To enable a user to interact with the computing device 400, the input device 480 represents any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice, etc. The output device 470 can be a display or one or more of a variety of output mechanisms known to those skilled in the art. In some instances, a multimode system enables the user to provide multiple types of input to communicate with the computing device 400. There are no limitations on the operation with any particular hardware arrangement, and thus, as hardware or firmware configurations are developed, the basic features here can be easily replaced with improved hardware or firmware arrangements.

[0051] Additional aspects are provided by the subject matter of the following clauses.

[0052] A method for generating an embedding space from a training dataset of a machine learning algorithm implemented on a computer system having one or more processors, comprising: receiving input operation data points by one or more layers of a machine learning algorithm through an input of the computer system having one or more processors; executing the machine learning algorithm on the computer system; generating, by the one or more layers of the machine learning algorithm executed on the computer system, an embedding space using a training dataset, wherein the embedding space includes a plurality of induced data points, each induced data point corresponding to a cluster of training data points among the plurality of training data points, and the number of the plurality of induced data points is less than the number of the plurality of training data points in the training dataset; comparing, by the machine learning algorithm executed on the computer system, the input operation data points with the plurality of induced data points; determining, by the machine learning algorithm executed on the computer system, whether there is reasonable confidence for the operation data points by determining whether there is a sufficient set of supporting historical data points among the plurality of training data points within a fixed parameterizable distance from the input operation data points; generating, by the machine learning algorithm executed on the computer system, an alert based on the reasonable determination so that a user or an electronic controller can take appropriate actions; and outputting the alert.

[0053] The method according to the foregoing clause, wherein generating the embedding space by the machine learning algorithm includes generating the embedding space using one or more layers of the machine learning algorithm, and each of the one or more layers provides a transformation from an input space to an output space, and the cardinality of the input space is equal to the cardinality of the output space of the previous layer.

[0054] The method according to any of the foregoing clauses, wherein the machine learning algorithm includes a neural network, a decision tree, or an ensemble model.

[0055] The method according to the foregoing clause, wherein in the neural network, generating the embedding space is performed in one or more layers of the neural network.

[0056] The method according to any of the foregoing clauses, wherein in the decision tree, generating the embedding space is performed at each depth of the decision tree.

[0057] The method according to any of the foregoing clauses, wherein in the ensemble model, generating the embedding space is performed for the input space.

[0058] A method as described in any of the preceding clauses, wherein generating the embedding space includes generating the plurality of induced data points, each of the plurality of induced data points corresponding to the centroid of a density distribution, capturing the density distribution of the plurality of training data points to provide a plurality of centroids, and the plurality of centroids being representative points of the plurality of training data points of the training data set.

[0059] A method as described in any of the preceding clauses, wherein generating the embedding space includes generating the plurality of induced data points, each cluster of the plurality of training data points being represented by a Gaussian distribution, and each of the plurality of induced data points corresponding to the centroid of the Gaussian distribution.

[0060] A method as described in any of the preceding clauses, further comprising generating, by a sensor device, the input operation data points received through the input of the computer system.

[0061] A method as described in any of the preceding clauses, wherein the input operation data points include remote sensing data.

[0062] A system for generating an embedding space from a training data set of a machine learning algorithm implemented on a computer system having one or more processors, the system comprising: an input for receiving input operation data points through one or more layers of a machine learning algorithm implemented on the computer system. The computer system is configured to: (a) execute the machine learning algorithm, (b) generate an embedding space using the one or more layers of the machine learning algorithm executed on the computer system, wherein the embedding space includes a plurality of induced data points, each induced data point corresponding to a cluster of training data points among the plurality of training data points, and the number of the plurality of induced data points being less than the number of the plurality of training data points in the training data set, (c) compare, by the machine learning algorithm executed on the computer system, the input operation data points with the plurality of induced data points, (d) determine, by the machine learning algorithm executed on the computer system, whether there is reasonable confidence for the operation data points by determining whether there is a sufficient set of supporting historical data points among the plurality of training data points within a fixed distance from the input operation data points, (e) generate, by the machine learning algorithm executed on the computer system, an alert based on the reasonable determination so that a user or an electronic controller can take appropriate actions, and (f) output the alert.

[0063] A system as described in the preceding clause, wherein the computer system is configured to generate the embedding space using the one or more layers of the machine learning algorithm, and each of the one or more layers provides a transformation from an input space to an output space, and the cardinality of the input space is equal to the cardinality of the output space of the previous layer.

[0064] A system as described in any of the foregoing clauses, wherein the machine learning algorithm includes a neural network, a decision tree, or an ensemble model.

[0065] A system as described in the foregoing clause, wherein, in the neural network, generating the embedding space is performed in one or more layers of the neural network.

[0066] A system as described in any of the foregoing clauses, wherein, in the decision tree, generating the embedding space is performed at each depth of the decision tree.

[0067] A system as described in any of the foregoing clauses, wherein, in the ensemble model, generating the embedding space is performed for the input space.

[0068] A system as described in any of the foregoing clauses, wherein the computer system is configured to generate the plurality of induced data points, each of the plurality of induced data points corresponding to the centroid of a density distribution, capturing the density distribution of the plurality of training data points to provide a plurality of centroids, the plurality of centroids being representative points of the plurality of training data points of the training data set.

[0069] A system as described in any of the foregoing clauses, wherein the computer system is configured to generate the plurality of induced data points, each cluster of the plurality of training data points being represented by a Gaussian distribution, wherein each of the plurality of induced data points corresponds to the centroid of the Gaussian distribution.

[0070] A system as described in any of the foregoing clauses, further comprising a sensor device in communication with the computer system, wherein the sensor device generates input operation data received through the input of the computer system.

[0071] A non-transitory computer-readable medium storing instructions that, when executed by a computer system having one or more processors, cause the computer system to: (a) receive input operation data points by one or more layers of a machine learning algorithm implemented on the computer system, (b) execute the machine learning algorithm, (c) generate an embedding space using a training data set, wherein the embedding space includes a plurality of induced data points, each induced data point corresponding to a cluster of training data points among the plurality of training data points, the number of the plurality of induced data points being less than the number of the plurality of training data points in the training data set, (d) compare the input operation data points with the plurality of induced data points by the machine learning algorithm executed on the computer system, (e) determine whether there is reasonable confidence for the operation data points by determining whether there is a sufficient set of supporting historical data points among the plurality of training data points within a fixed distance from the input operation data points, (f) generate an alert based on determining whether there is the reasonableness so that a user or an electronic controller can take appropriate actions, and (g) output the alert.

[0072] Although the foregoing description is directed to preferred embodiments of the present disclosure, other variations and modifications will be apparent to those skilled in the art and may be made without departing from the spirit or scope of the present disclosure. Additionally, features described in connection with one embodiment of the present disclosure may be used in combination with other embodiments, even if not explicitly stated above.

Claims

1. A method for generating an embedding space from a training data set of a machine learning algorithm implemented on a computer system having one or more processors, the method comprising: Receiving, via an input of the computer system, input operational data points by one or more layers of the machine learning algorithm; executing the machine learning algorithm on the computer system; generating, by the one or more layers of the machine learning algorithm executed on the computer system, an embedding space using the training data set, wherein the embedding space includes a plurality of induced data points, each induced data point corresponding to a cluster of training data points in a plurality of training data points, the number of the plurality of induced data points being less than the number of the plurality of training data points in the training data set; comparing, by the machine learning algorithm executing on the computer system, the input operational data point to the plurality of induced data points; determining, by the machine learning algorithm executing on the computer system, whether there is reasonable confidence in the operational data point by determining whether there is a sufficient set of supporting historical data points among the plurality of training data points within a fixed distance from the input operational data point; generating, by the machine learning algorithm executing on the computer system, an alert based on the determination whether reasonable confidence exists to enable a user or an electronic controller to take appropriate action; as well as The alarm is output by the computer system.

2. The method of claim 1, wherein: Generating the embedding space by the machine learning algorithm includes generating the embedding space using one or more layers of the machine learning algorithm, each of the one or more layers providing a transformation from an input space to an output space, wherein the cardinality of the input space is equal to the cardinality of the output space of the previous layer.

3. The method of claim 1, wherein: Generating the embedding space includes generating the multiple induced data points, each of the multiple induced data points corresponds to the centroid of the density distribution, capturing the density distribution of the multiple training data points to provide multiple centroids, and the multiple centroids are representative points of the multiple training data points of the training data set.

4. The method of claim 1, wherein: Generating the embedding space includes generating the plurality of induced data points, each cluster of the plurality of training data points is represented by a Gaussian distribution, and each of the plurality of induced data points corresponds to a centroid of the Gaussian distribution.

5. The method of claim 1, further comprising generating, by a sensor device, the input operation data received via the input of the computer system.

6. The method of claim 1, wherein: The input operational data includes remote sensing data.

7. The method of claim 1, wherein: The machine learning algorithm includes a neural network, a decision tree or an overall model.

8. The method of claim 7, wherein: In the neural network, generating the embedding space is performed in one or more layers of the neural network.

9. The method of claim 7, wherein: In the decision tree, generating the embedding space is performed at each depth of the decision tree.

10. The method of claim 7, wherein: In the overall model, generating the embedding space is performed with respect to the input space.