Optimizing Modular Architectures for Semi-Supervised Incremental Learning

The machine learning framework uses short-shot learning and hybrid replay to train models in real-time with live streaming data, addressing class imbalance and adapting system complexity, overcoming the limitations of traditional methods in domains with costly data collection.

JP7815198B2Active Publication Date: 2026-02-17SRI INTERNATIONAL
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023201995
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-11-28
Filing Date
2023-11-29
Publication Date
2026-02-17
Estimated Expiration
2043-11-29

AI Technical Summary

Technical Problem

Traditional machine learning models require large amounts of labeled data for training, which is often collected offline, leading to overfitting and poor performance in real-world applications where data distribution and characteristics differ from offline data, especially in domains like maritime and medical imaging where data collection is costly and challenging.

Method used

A machine learning framework incorporating short-shot learning, hybrid replay, and architecture optimization techniques to train models in real-time using live streaming data, addressing class imbalance and adapting system complexity based on sensor data.

Benefits of technology

Enables fast, accurate, and robust training of machine learning models in real-world applications with limited labeled data, adapting to changing data distributions and improving performance over time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007815198000001
    Figure 0007815198000001
  • Figure 0007815198000002
    Figure 0007815198000002
  • Figure 0007815198000003
    Figure 0007815198000003
Patent Text Reader

Abstract

To provide a machine learning framework that addresses situations where prior learning cannot be sufficiently performed.SOLUTION: A computing system 100 has a memory 102, and processing circuit 143 that performs communication. The processing circuit executes a machine learning system 104 having at least a data preprocessing module 114, a task-specific network module 116, and an architecture optimization module 120. The machine learning system trains one or more machine learning models 106. The data preprocessing module generates augmented input data based on streaming input data. The task-specific network module has a machine learning model that executes a specific task based at least in part on the augmented input data. The architecture optimization module adapts a network architecture of one or more machine learning models based on changes in the streaming input data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims the benefit of U.S. Patent Application No. 63 / 385,319, filed November 29, 2022, and U.S. Patent Application No. 63 / 447,559, filed February 22, 2023, which are incorporated by reference in their entireties.

[0002] Government Rights This invention was made with government support under Contract No. N65236-20-C-8020 awarded by the U.S. Navy NIWC Atlantic Charleston. The government has certain rights in this invention.

[0003] The present disclosure relates to machine learning systems. [Background technology]

[0004] Because training machine learning models in real time is computationally expensive, machine learning models are typically trained offline on large datasets. Models are typically frozen offline for training to prevent overfitting to the training data. Overfitting occurs when a model learns too many specific details of the training data, resulting in the model not generalizing well to new data. Models are typically deployed in the real world under the assumption that the characteristics and distribution of online data in the real world are identical to those of offline data. However, this is not always the case. Real-world data may differ from offline data in various ways, such as data distribution, noise level, or the presence of outliers. If real-world data differs from offline data, the model may not perform well. Summary of the Invention [Problem to be solved by the invention]

[0005] In real-world applications, it is not always possible to use pre-trained models. In some domains, such as maritime, offline data collection and labeling can be challenging. This is because collecting data in these domains can be costly and time-consuming, and it can be difficult to find experts to label the data. Data modalities other than vision (cameras) are also less explored and new. For example, in the medical field, there is growing interest in using data from medical imaging modalities such as magnetic resonance imaging (MRI) and computed tomography (CT) scans. However, there are not many pre-trained models available for these data modalities. [Means for solving the problem]

[0006] Generally, this disclosure describes techniques that use a machine learning framework with several new features, including but not limited to, short-shot learning, hybrid replay, and architecture optimization. Short-shot learning is a technique that enables machine learning models to learn new tasks or train with just a few examples. Short-shot learning is useful in real-world applications where it is difficult to collect large amounts of labeled data. Short-shot learning contrasts with traditional machine learning, which trains models with a large number of examples. Short-shot learning is becoming increasingly important as the amount of data available for training machine learning models increases.

[0007] The hybrid replay method may be implemented by a hybrid replay module, which may address the problem of class imbalance by generating expanded samples of classes using limited available examples. The hybrid replay module may help improve the performance of machine learning models for imbalanced classes. The hybrid replay module is a technique that addresses the problem of class imbalance. Class imbalance occurs when there are more examples of one class than other classes. Class imbalance can make it difficult to train a machine learning model to accurately classify minority classes. The architecture optimization method may be implemented by an architecture optimization module, which may automatically adapt the complexity of the system based on sensor data evolved for inference. The architecture optimization module may help improve the performance of the machine learning model over time. The architecture optimization module may automatically adapt the complexity of the system based on sensor data evolved for inference.

[0008] The techniques may provide one or more technical advantages that realize at least one practical application. For example, a hybrid replay module may help improve the performance of a machine learning model for an imbalanced class. A task-specific network may enable a machine learning model to be trained with a small number of examples. Some of the advantages of the disclosed techniques may be useful in real-world applications where it is difficult to collect large amounts of labeled data. An architecture optimization module may automatically adapt system complexity based on evolved sensor data for inference. The architecture optimization module may help improve the performance of a machine learning model over time.

[0009] The combination of processing components described above may be used to train one or more machine learning models, such as, but not limited to, neural networks, using live streaming data. Live streaming data is a relatively new concept that may improve the performance of machine learning models in real-world applications. Additional advantages of the disclosed combination of processing components include, but are not limited to, real-time inference, scalability, and robustness. Advantageously, machine learning models may be trained and deployed in real-time, which may be important for many real-world applications. The disclosed techniques are scalable to handle large amounts of data. The disclosed techniques are also robust to changes in data distribution.

[0010] In one example, a system includes a processing circuit in communication with a storage medium. The processing circuit is configured to execute a machine learning system having at least a first module, a second module, and a third module. The machine learning system is configured to train one or more machine learning models. The first module is configured to generate augmented input data based on streaming input data. The second module includes a machine learning model configured to perform a specific task based at least in part on the augmented input data. The third module is configured to adapt a network architecture of the one or more machine learning models based on changes in the streaming input data.

[0011] In one example, a method includes using a first module to generate augmented input data based on streaming input data, using a second module comprising a machine learning model to perform a particular task based at least in part on the augmented input data, and using a third module to adapt a network architecture of one or more machine learning models based on changes in the streaming input data.

[0012] In one example, a non-transitory computer-readable storage medium having encoded instructions configured to cause a processing circuit to: use a first module to generate augmented input data based on streaming input data; use a second module comprising a machine learning model to perform a particular task based at least in part on the augmented input data; and use a third module to adapt a network architecture of one or more machine learning models based on changes in the streaming input data.

[0013] The details of one or more embodiments of the disclosed technology are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the technology will be apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]

[0014] [Figure 1] FIG. 1 is a block diagram illustrating an exemplary system in accordance with the techniques of this disclosure.

[0015] [Figure 2] FIG. 2 is a conceptual diagram illustrating an example of semi-supervised incremental learning according to the techniques of this disclosure.

[0016] [Figure 3] FIG. 3 is a conceptual diagram illustrating an exemplary hybrid replay method in accordance with the techniques of this disclosure.

[0017] [Figure 4] FIG. 4 is a conceptual diagram illustrating an example of an architecture optimization method according to the technique of the present disclosure.

[0018] [Figure 5] FIG. 5 is a conceptual diagram illustrating an example of an incrementally differentiable architecture search (DARTS) optimization method in accordance with the techniques of this disclosure.

[0019] [Figure 6]FIG. 6 is a flow chart illustrating an example mode of operation of a hybrid replay module in accordance with the techniques described in this disclosure.

[0020] [Figure 7] FIG. 7 is a flowchart illustrating an exemplary mode of operation of an architecture optimization module in accordance with the techniques described in this disclosure.

[0021] [Figure 8] FIG. 8 is an exemplary diagram of a distributed data processing system in which aspects of the illustrative techniques may be implemented.

[0022] Like reference characters refer to like elements throughout the drawings and description. DETAILED DESCRIPTION OF THE INVENTION

[0023] Traditional machine learning approaches typically require large amounts of labeled data to train the model, which is often collected offline, before the model is deployed in the real world.

[0024] However, traditional machine learning techniques have several limitations. Collecting large amounts of labeled data in all domains, such as naval, space, and underwater domains, is impractical. There are many emerging data modalities, such as RF signals, radar, and synthetic aperture radar (SAR), that have limited available labeled data. Data from the same class can change over time, and therefore, models trained on offline data may not perform well on new data. Because each application has unique requirements, it is not always feasible to use a single pre-trained model for all applications.

[0025] This disclosure describes a new machine learning framework that addresses the challenges mentioned above. The disclosed framework is designed to train and deploy machine learning models in real time using live streaming data.

[0026] In one aspect, the disclosed framework may use a combination of three processing modules to accomplish this, including but not limited to a hybrid replay module, a task-specific module, and an architecture optimization module.

[0027] The hybrid replay module may address the issue of limited labeled data by generating augmented samples of a small number of classes. The task-specific network module may train a machine learning model with a small number of examples. The architecture optimization module may automatically adapt the system complexity based on evolved sensor data for inference. Advantageously, the combination of the three processing modules described above enables the framework to train machine learning models for real applications without requiring large amounts of offline labeled data. Some examples of how the disclosed framework can be used in real applications are as follows:

[0028] As one example, the disclosed framework can be used to train a model that detects and tracks ships in real time using radar data. As another example, the disclosed framework can be used to train a model that identifies and classifies objects in real time using satellite imagery. As yet another example, the disclosed framework can be used to train a model that classifies fish species in real time using underwater video footage.

[0029] In accordance with the techniques of this disclosure, this disclosure describes a new approach for training machine learning models in real time using live streaming data even when limited or no offline training data is available. The disclosed technique is known as in-situ algorithm training. In-situ algorithm training has several advantages over traditional machine learning techniques.

[0030] In-situ algorithm training can be faster and more efficient because the model is trained on live data as the model is generated. In-situ algorithm training can be more robust to changes in data distribution because the model is constantly updated with new data. In-situ algorithm training may be used to train models for applications where offline training data is not available or where collecting offline training data is impractical.

[0031] This disclosure also describes a hybrid replay module used to address the problem of class imbalance in live streaming data. Class imbalance can occur when there are more examples of one class than other classes. Class imbalance can make it difficult to train a machine learning model to accurately classify the minority class. In one aspect, the hybrid replay module may function by generating augmented samples of the minority class. In one aspect, such augmented samples of the minority class may be generated using techniques such as, but not limited to, data augmentation or synthetic data generation. The augmented samples may then be added to a training dataset, thereby helping to improve the model's performance for the minority class.

[0032] This disclosure also describes an architecture optimization module that may be used to automatically adapt the complexity of a system based on evolved sensor data for inference. Such adaptation may help ensure that the model is always using an optimal amount of resources to achieve a desired level of accuracy. Overall, this disclosure describes a promising new approach for training machine learning models in real time using live streaming data. The disclosed technology has the potential to revolutionize the way machine learning is used in many different applications. Some examples of how in-situ algorithm training can be used in real-world applications are as follows: In-situ algorithm training can be used to train a model that detects fraudulent transactions in real time using live data from financial institutions. In-situ algorithm training can be used to train a model that diagnoses disease in real time using live data from medical devices. In-situ algorithm training can also be used to train a model that predicts when a machine is about to fail using live data from machine sensors.

[0033] In one aspect, a data preprocessing module may be an optional component of a machine learning pipeline responsible for preparing data for training and evaluation. The data preprocessing module may provide an interface for three main functions: format conversion, metadata derivation, and data association. The format conversion function may convert data into a format compatible with the machine learning algorithms that may be used to train the model(s).

[0034] 1 is a block diagram illustrating an exemplary computing system 100. As shown, the computing system 100 includes processing circuitry 143 and memory 102 for executing a machine learning system 104 having one or more modules, including, but not limited to, a data preprocessing module 114, a task-specific network module 116, a hybrid replay module 118, and an architecture optimization module 120. Further, the task-specific module 116 may include one or more machine learning models 106. The ML models 106 may include various types of neural networks, such as, but not limited to, recurrent neural networks (RNNs), convolutional neural networks (CNNs), and deep neural networks (DNNs).

[0035] Computing system 100 may be implemented as any suitable computing system, such as one or more server computers, workstations, laptops, mainframes, appliances, cloud computing systems, high-performance computing (HPC) systems (i.e., supercomputing) and / or other computing systems that may be capable of performing the operations and / or functions described in accordance with one or more aspects of the present disclosure. In some examples, computing system 100 may represent (or be implemented through) a cloud computing system, server farm and / or server cluster that provides services to client devices and other devices or systems. In other examples, computing system 100 may represent or be implemented through one or more virtualized compute instances (e.g., virtual machines, containers, etc.) of a data center, cloud computing system, server farm and / or server cluster.

[0036] The techniques described in this disclosure may be implemented, at least in part, in hardware, software, firmware, or any combination thereof. For example, various aspects of the described techniques may be implemented within processing circuitry 143 of computing system 100, which may include one or more of a microprocessor, controller, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or equivalent discrete or integrated logic circuitry or other types of processing circuitry. Processing circuitry 143 of computing system 100 may perform functions and / or execute instructions associated with computing system 100. Computing system 100 may use processing circuitry 143 to perform operations according to one or more aspects of the present disclosure using software, hardware, firmware, or a mixture of hardware, software, and firmware residing on and / or executed by computing system 100. The terms “processor” or “processing circuitry” may generally refer to any of the logic circuits described above or other equivalent circuitry, alone or in combination with other logic circuitry. The control unit may include hardware that performs one or more of the techniques of this disclosure.

[0037] In another example, computing system 100 may comprise any suitable computing system having one or more computing devices, such as a desktop computer, a laptop computer, a gaming console, a smart television, a handheld device, a tablet, a mobile phone, a smartphone, etc. In some examples, at least a portion of system 100 is distributed across a cloud computing system, a data center, or a network, such as the Internet, another public or private communications network, for transmitting data between computing systems, servers, and computing devices, such as broadband, cellular, Wi-Fi, ZigBee, Bluetooth (or other personal area network—PAN), near field communication (NFC), ultra-wideband, satellite, enterprise, service provider, and / or other types of communications networks.

[0038] The memory 102 may comprise one or more storage devices. One or more components of the computing system 100 (e.g., the processing circuitry 143, the memory 102, the data preprocessing module 114, the task-specific network module 116, the hybrid replay module 118, and the architecture optimization module 120) may be interconnected (physically, communicatively, and / or operationally) to enable communication between the components. In some examples, such connections may be provided by a system bus, a network connection, an inter-process communication data structure, a local area network, a wide area network, or any other method for communicating data. One or more storage devices of the memory 102 may be distributed among multiple devices.

[0039] The memory 102 may store information for processing during operation of the computing system 100. In some examples, the memory 102 comprises temporary memory, meaning that the primary purpose of one or more storage devices in the memory 102 is not long-term storage. The memory 102 may be configured as volatile memory for short-term storage of information and therefore does not retain its stored contents when operation is stopped. Examples of volatile memory include random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), and other forms of volatile memory known in the art. In some examples, the memory 102 may also include one or more computer-readable storage media. The memory 102 may be configured to store larger amounts of information than volatile memory. The memory 102 may be configured as non-volatile memory space for long-term storage of information and retain information after an on / off cycle. Examples of non-volatile memory include magnetic hard disks, optical disks, flash memory, or forms of electrically programmable memory (EPROM) or electrically erasable and programmable (EEPROM) memory. The memory 102 may store program instructions and / or data associated with one or more of the modules described in accordance with one or more aspects of the present disclosure.

[0040] The processing circuitry 143 and memory 102 may provide an operating environment or platform for one or more modules or units (e.g., data preprocessing module 114, task-specific network module 116, hybrid replay module 118, and architecture optimization module 120), which may be implemented as software, although in some examples, the operating environment or platform includes any combination of hardware, firmware, and software. The processing circuitry 143 may execute instructions, and one or more storage devices, such as memory 102, may store instructions and / or data for one or more modules. The combination of the processing circuitry 143 and memory 102 may retrieve, store, and / or execute instructions and / or data for one or more applications, modules, or software. The processing circuitry 143 and / or memory 102 may be operatively coupled to one or more other software and / or hardware components, including, but not limited to, one or more of the components shown in FIG. 1 .

[0041] The processing circuitry 143 may execute the machine learning system 104 using virtualization modules, such as virtual machines or containers, that run on the underlying hardware. One or more of such modules may execute as one or more services of an operating system or computing platform. Aspects of the machine learning system 104 may execute as one or more executable programs at the application layer of a computing platform.

[0042] One or more input devices 144 of computing system 100 may generate, receive, or process input. Such input may include input from a keyboard, pointing device, voice response system, video camera, biometric detection / response system, button, sensor, mobile device, control pad, microphone, presence-sensitive screen, network, or any other type of device for detecting input from a human or machine.

[0043] One or more output devices 146 may generate, transmit, or process output. Examples of output are tactile, audio, visual, and / or video output. Output device(s) 146 may include a display, a sound card, a video graphics adapter card, speakers, a presence-sensitive screen, one or more USB interfaces, a video and / or audio output interface, or any other type of device capable of generating tactile, audio, video, or other output. Output device(s) 146 may include a display device that may function as an output device using technologies including liquid crystal display (LCD), quantum dot display, dot matrix display, light-emitting diode (LED) display, organic light-emitting diode (OLED) display, cathode ray tube (CRT) display, e-ink, or any other type of display capable of generating monochrome, color, or tactile, audio, and / or visual output. In some examples, computing system 100 may include a presence-sensitive display that may function as a user interface device, operating as both one or more input devices 144 and one or more output devices 146.

[0044] One or more communication devices 145 of computing system 100 may communicate with devices external to computing system 100 (or between separate computing devices of computing system 100) by sending and / or receiving data and, in some respects, may act as both an input device and an output device. In some examples, communication device 145 may communicate with other devices over a network. In other examples, communication device 145 may send and / or receive radio signals over a wireless network, such as a cellular wireless network. Examples of communication device 145 may include a network interface card (e.g., an Ethernet card), an optical transceiver, a radio frequency transceiver, a GPS receiver, or any other type of device capable of sending and / or receiving information. Other examples of communication device 145 may include Bluetooth, GPS, 3G, 4G, and Wi-Fi radios, universal serial bus (USB) controllers, and the like found in mobile devices.

[0045] 1 , the data pre-processing module 114 may receive input data from the input dataset 110 and generate output data 112. The input data 110 and the output data 112 may include various types of information. For example, the input data 110 may include live streaming trading data (streaming data 122). Furthermore, the data pre-processing component 114 may output the streaming data 122 for integration with generated data 123 (output from the hybrid replay 118) to generate augmented data 124 for the task-specific network module 116 and the architecture optimization module 120. The task-specific network module 116 may output interference output (e.g., classification, detection, recognition, segmentation, or prediction), which may be part of the output data 112.

[0046] As described above, the ML model 106 may comprise various types of neural networks, such as, but not limited to, RNNs, CNNs, and DNNs, each having a corresponding set of layers. Each set of layers 108 may have a respective set of artificial neurons. The layers 108 may include, for example, an input layer, a feature layer, an output layer, and one or more hidden layers. The layers 108 may include fully connected layers, convolutional layers, pooling layers, and / or other types of layers. In a fully connected layer, the output of each neuron in the previous layer forms the input of each neuron in the fully connected layer. In a convolutional layer, each neuron in the convolutional layer processes input from neurons associated with its receptive field. A pooling layer combines the outputs of neuron clusters in one layer to a single neuron in the next layer.

[0047] Each input of each artificial neuron in each layer of the set of layers may be associated with a corresponding weight in weights 126. Various activation functions are known in the art, such as rectified linear unit (ReLU), TanH, sigmoid, etc.

[0048] The ML system 104 may process the training data 113 to train the ML model 106 according to the techniques described herein. For example, the machine learning system 104 may apply an end-to-end training method that includes processing the training data 113. The machine learning system 104 may process the input data 110, which may include streaming data 122, to generate inference data (output data 112), as described below.

[0049] In one aspect, the machine learning system 104 may generate fast and accurate results while overcoming potential class imbalance issues presented by the input data 110. The machine learning system 104 may be configured for use in real-world applications where the live streaming data 122 may evolve over time. The data preprocessing module 114 may convert the live streaming data 122 into a format compatible with the machine learning algorithm used to train the ML model 106. The data preprocessing module 114 may extract metadata from the input data 110. In one aspect, the task-specific network module 116 may be a machine learning model configured to perform a specific task, such as, but not limited to, image classification, object detection, image generation, or text generation. The task-specific network module 116 may be trained using a combination of few-shot learning techniques, semi-supervised learning, and self-supervised learning. Such training techniques may enable the task-specific network module 116 to learn to perform a task accurately even when available training data 113 is limited.

[0050] In one embodiment, the hybrid replay module 118 may be configured to use a combination of selective memory and generative networks to augment the actual streaming data 122. The hybrid replay module 118 may help address issues of class imbalance and may prevent the machine learning model 106 from overfitting the training data 113.

[0051] In one aspect, the architecture optimization module 120 may be configured to adapt the network architecture of the task-specific network 116 and adapt the replay mode based on the increasing complexity of the streaming data 122. The architecture optimization module 120 may help improve the accuracy of the output of the ML model 106.

[0052] In summary, the machine learning system 104 may be configured to first preprocess the live streaming data 122. In one aspect, the data preprocessing module 114 may convert the input data 110 into a format compatible with the task-specific network module 116 and extract metadata from the input data 110. The task-specific network module 116 may be trained on the preprocessed data using a combination of few-shot learning, semi-supervised learning, and self-supervised learning. The hybrid replay module 118 may be used to augment the actual streaming data 122, and the hybrid replay module 118 may generate augmented data 124. The hybrid replay module 118 may help address class imbalance issues and prevent the ML model 106 from overfitting the training data 113. Finally, the architecture optimization module 120 may adapt the network architecture of the task-specific network module 116 and may adapt the replay mode based on the increasing complexity of the streaming data 122. The architecture optimization module 120 may help improve the accuracy of the output of the ML model 106 .

[0053] The described machine learning system 104 may have several advantages over traditional machine learning techniques. The machine learning system 104 may produce fast and accurate results even when there is limited or no available training data 113. As another advantage, the machine learning system 104 may be robust to class imbalance and may avoid overfitting to the training data 113. As yet another advantage, the machine learning system 104 may be able to adapt to changes in data distribution over time.

[0054] In one aspect, the disclosed technology enables the machine learning system 104 to be adapted for real-world applications such as, but not limited to, fraud detection, medical diagnosis, and predictive maintenance. The machine learning system 104 can be used to train the ML model 106 to detect fraudulent transactions in real time using live input data 110 (e.g., streaming data 122) from financial institutions. The machine learning system 104 can be used to train the ML model 106 to diagnose illnesses in real time using live input data 110 from medical devices. As yet another non-limiting example, the machine learning system 104 can be used to train the ML model 106 to predict when a machine is about to fail using live input data 110 from sensors in the machine.

[0055] Advantageously, the combination of the hybrid replay module 118, the task-specific network module 116, and the architecture optimization module 120 to train the ML model 106 using live streaming data is a novel concept not currently known in the art. Conventional machine learning techniques typically require large amounts of labeled data to train a model. Such labeled data is often collected offline before the model is deployed in the real world.

[0056] However, traditional machine learning techniques have several limitations: for example, in all domains such as the naval, space, and underwater domains, it may be impractical to collect large amounts of labeled data.

[0057] According to the techniques of the present disclosure, the machine learning system 104 may have several new capabilities that address the challenges of training and deploying machine learning models in real-world applications. For example, few-shot learning is a technique that allows the machine learning system 104 to train with a small number of examples. This type of training is useful for real-world applications where collecting large amounts of labeled data is difficult or costly. As another non-limiting example, the hybrid replay module 118 may implement a hybrid replay technique, which is a technique that addresses the problem of class imbalance. Class imbalance occurs when there are more examples of one class than other classes.

[0058] Advantageously, because the present disclosure describes a new machine learning framework that may train and deploy machine learning models in real time using live streaming data 122, technical aspects of the present disclosure enable long-term learning systems or applications in new data domains to avoid or minimize the manual effort of collecting large amounts of training data offline in advance. To achieve this, as described above, the disclosed framework uses a combination of the three techniques described above: few-shot learning, hybrid replay, and architecture optimization. For example, class imbalance may make it difficult for the machine learning system 104 to learn to accurately classify minority classes. The hybrid replay technique works by generating augmented samples of the minority classes. Such augmented samples may help improve the model's performance for the minority classes. The architecture optimization technique automatically adapts the complexity of the machine learning model based on the data. Such optimization is useful for real-world applications where data distributions may change over time.

[0059] In one aspect, the data preprocessing module 114 may output the streaming data 122 for integration with the generated data (output from the hybrid replay module 118) to generate augmented data 124 for the task-specific network module 116 and the architecture optimization module 120. The architecture optimization module 120 may output selected components to both the hybrid replay module 118 and the task-specific network module 116. The task-specific network module 116 may use short-shot learning (also referred to herein as "short-shot technology") for automatic result generation for new / unknown data by generalizing data and / or features from old data to determine new data. By using short-shot learning technology, the task-specific network module 116 can generate fast and accurate inference output in real time, for example, after a user labels a small amount of live streaming data 122, thereby avoiding offline learning techniques previously known in the art. A confidence score for the inference data may be output to the hybrid replay module 118.

[0060] Furthermore, the task-specific network module 116 may be trained using live streaming input data with class imbalance. In one aspect, techniques that may be used to resolve class imbalance in the live streaming input data may include, but are not limited to, self-supervised learning, semi-supervised learning, and calibration of inference results, as described below. For example, self-supervised learning is a technique that allows the machine learning system 104 to learn without requiring labeled data.

[0061] Semi-supervised learning may include a supervised learning module, a self-supervised learning module, an online learning module, an online prediction module, and a prediction accumulation module (as shown in FIG. 2). The supervised learning module may be used to train the ML model 106 on labeled data. The labeled data may be data annotated with labels by an expert user. The supervised learning module may use the labeled data to learn the relationship between the input data 110 and the output labels. The self-supervised learning module is used to train the ML model 106 on unlabeled data. The unlabeled data may be data that has not been annotated with labels by an expert user.

[0062] The hybrid replay module 118 may address potential class imbalance issues, especially for rare but important classes. The hybrid replay module 118 may address the class imbalance issues described above in a number of ways, including, but not limited to, 1) generating augmented samples, 2) augmenting the current training batch data, and 3) maximizing sample diversity. For example, the hybrid replay module 118 may use various techniques to generate augmented samples (e.g., augmented data 124) for classes with limited available examples. In one aspect, a variational autoencoder, a generative adversarial network, or traditional sample-based augmentation techniques may be used to generate the augmented samples.

[0063] The ML model 106 is trained based on data. If the data is incomplete or inaccurate, the corresponding model cannot be accurately trained. One of the challenges in machine learning is that data is often incomplete or inaccurate. The above-mentioned challenges can be due to many factors, such as, but not limited to, rare examples in the training data 113, catastrophic forgetting, storing a large amount of prior data, and high-risk, low-probability events. For example, some examples may be rare in the training data 113. Such rare examples may make it difficult for the ML model 106 to learn to accurately classify these examples.

[0064] In one aspect, the hybrid replay module 118 may implement a method for implementing the above examples, for example, by using a replay memory and a replay generative artificial intelligence (AI) architecture. The replay memory may be a data structure that stores a mixture of representative labeled samples and unlabeled tracks (including associated metadata) from a fixed-size dynamic buffer. The replay memory may be used to store the most useful and representative examples from the training data 113. The replay memory may help the ML model 106 learn to accurately classify even when new examples are rare or difficult to classify. The replay memory may also retroactively propagate future class labels to increase class-labeled data and supplement future batches with previous data. Such propagation may help prevent catastrophic forgetting and improve the accuracy of the ML model 106 for rare and difficult examples. A generative AI architecture is the design of a system that generates new content or data based on existing data. This type of AI system may be trained on a large dataset of examples and use that knowledge to create new outputs similar to the training data. In one aspect, the replay generation AI architecture may be implemented as a replay generation inverse network (GAN).

[0065] For example, the hybrid replay module 118 may combine both the replay memory and the replay GAN with the streaming data 122 to generate an augmented training set that outputs data to a classifier / classifier. The classifier / classifier is a machine learning model trained to distinguish between real data generated by the replay GAN and fake data. The classifier / classifier may be used by the machine learning system 104 to predict the class of a data sample.

[0066] FIG. 2 is a conceptual diagram illustrating an example of semi-supervised incremental learning according to the techniques of this disclosure. Semi-supervised learning is a machine learning technique that uses both labeled and unlabeled data to train a model. Semi-supervised learning is often used when there is limited labeled data available but a large amount of unlabeled data. Semi-supervised incremental learning (SSIL) is a machine learning paradigm that combines the advantages of semi-supervised learning and incremental learning. In SSIL, a model may be trained with a small amount of labeled data and a large amount of unlabeled data. Once the model is trained, it may learn new classes and update its knowledge base without forgetting what it has already learned. The semi-supervised learning system 200 shown in FIG. 2 includes modules such as, but not limited to, a supervised learning module 202, a self-supervised learning module 204, an online learning module 206, an online prediction module 208, and a prediction accumulation module 210.

[0067] The supervised learning module 202 may be used to train the ML model 106 with labeled data.

[0068] A self-supervised learning module 204 may be used to train the ML model 106 with unlabeled data. An online learning module 206 may be used to update the ML model 106 as new data becomes available. An online prediction module 208 may be used to generate predictions for new data. A prediction accumulation module 210 may combine predictions from multiple observations to improve the accuracy of inference. Pre-training data 212 may be input to either or both of the supervised learning module 202 or the self-supervised learning module 204. Pre-training data 212 from a pre-training domain 214 may be used to train the initial ML model 106. The pre-training data 212 may help improve the performance of the ML model 106 in a target domain 216 even when limited labeled data is available. Limited target data 218 may be input to the online training module 206.

[0069] Advantageously, the limited target data 218 may be sparsely annotated with labels by an expert user. In other words, only a small subset of the limited target data 218 needs to be labeled. The online learning module 206 may use this limited, labeled target data 218 to update the ML model 106 and improve accuracy for the target domain 216. The target data 220 may be input to the online prediction module 208. The online prediction module 208 may use the ML model 106 to generate predictions for new data. The prediction accumulation module 210 may combine predictions from multiple observations to improve the accuracy of inference. In one embodiment, the prediction accumulation module 210 may combine predictions by weighting the predictions with corresponding confidence levels.

[0070] 3 is a conceptual diagram illustrating an example hybrid replay method 300 in accordance with the techniques of this disclosure. A hybrid replay module 118 implementing the hybrid replay method may address the problem of potential class imbalance, especially for rare but important classes, in a number of ways.

[0071] The hybrid replay module 118 may use various techniques to generate augmented samples for classes with limited available examples. In various aspects, the hybrid replay module 118 may generate the augmented samples using a variational autoencoder, a GAN, or traditional sample-based augmentation techniques. For example, a variational autoencoder may be used to generate new data samples that are similar to existing data samples in a training set. A GAN may be used by the hybrid replay module 118 to generate new data samples that are indistinguishable from actual data samples.

[0072] The hybrid replay module 118 may use conventional example-based augmentation techniques to generate new data samples by applying random transformations to existing data samples. The hybrid replay module 118 may also augment the current training batch data (e.g., training data 113) with representative examples from previous data, such as from selected component data received from the architecture optimization module 120. Augmenting the current training batch data may help ensure that the ML model 106 is exposed to a wide variety of data, including examples of rare but important classes.

[0073] The hybrid replay module 118 may be used to maximize sample diversity and increase class balance, especially for high-risk, low-probability events, because the hybrid replay module 118 may generate expanded samples of rare but important classes and because the hybrid replay module 118 may augment the current training batch data with representative examples from previous data. Non-limiting examples of how the hybrid replay module 118 can be used to address class imbalance in fraud detection applications are as follows:

[0074] The training data 113 of an ML model 106 configured to perform fraud detection may contain a large number of non-fraudulent transactions and a small number of fraudulent transactions. Such class imbalance may make it difficult to train the ML model 106 to accurately detect fraudulent transactions. To address this class imbalance, the hybrid replay module 118 may be used to generate an augmented sample of fraudulent transactions (e.g., augmented data 124).

[0075] The hybrid replay module 118 may generate augmented samples of fraudulent transactions using various techniques such as, but not limited to, variational autoencoders, GANs, or traditional sample-based augmentation techniques.

[0076] The generated augmented samples of fraudulent transactions may be added to the training data 113. The generated augmented samples help improve the class balance of the training data 113, making it easier for the ML model 106 to learn to accurately detect fraudulent transactions. In addition to generating augmented samples, the hybrid replay module 118 may be used to augment the current training batch data with representative examples from previous data. Such augmentation may be performed by selecting examples from previous data that are similar to the examples in the current training batch.

[0077] In one aspect, the hybrid replay module 118 may help ensure that the ML model 106 is exposed to a wide variety of data, including, but not limited to, examples of fraudulent transactions, so that the ML model 106 can learn to more accurately detect fraudulent transactions.

[0078] Hybrid Replay is a machine learning method that also addresses the challenges of infrequent examples in training data 113, catastrophic forgetting, and the need to store large amounts of prior data. In one aspect, the Hybrid Replay module 118 addresses the challenges of infrequent examples by using a dynamic memory repository to hold useful and representative prior examples and training class-conditional generative networks to supplement memory and increase sample diversity.

[0079] Infrequent examples are examples that occur infrequently in the training data 113. Infrequent examples can make it difficult for the ML model 106 to learn to accurately classify these examples. Catastrophic forgetting is the phenomenon of a machine learning model forgetting what it was trained on when trained on new data. Catastrophic forgetting can occur when the new data is very different from the data the model was originally trained on.

[0080] In one aspect, the hybrid replay module 118 may combine a replay memory 302 and a replay GAN 304. The replay memory 302 may store a mixture of representative labeled samples and unlabeled tracks (including associated metadata) in a fixed-size dynamic buffer. The replay memory 302 may be used to store the most useful and representative examples from the training data 113. The replay memory 302 may help the ML model 106 learn to accurately classify even when new examples are rare or difficult to classify. The replay memory may also retroactively propagate future class labels to increase class-labeled data and supplement future batches with previous data. Such propagation may help prevent catastrophic forgetting and improve the accuracy of the ML model 106 for rare and difficult examples. The replay memory 302 may maintain buffer size by clustering labeled data and storing representative examples. The buffer may help ensure that the replay memory 302 contains the most useful and representative examples as the ML model 106 learns and the training data 113 changes. The replay GAN 304 is a machine learning model that may be trained to generate new data samples similar to the data samples in the replay memory 302. The replay GAN 304 may be used to increase class balance between the minority and majority classes, and the replay GAN 304 may generate high-priority examples more frequently. The replay GAN 304 may utilize an auxiliary classifier 306 to stabilize sample generation and measure the quality of samples relative to those present in the replay memory 302. The hybrid replay module 118 may combine the replay memory 302 and the replay GAN 304 to train the machine learning model 106. The replay memory 302 may be used to store the most useful and representative examples from the training data. The replay GAN 304 may be used to generate new data samples similar to the data samples in the replay memory 302.The ML model 106 may be trained on the labeled data in the replay memory 302 as well as new data samples generated by the replay GAN 304. The ML model 106 may learn to accurately classify data samples even when the data samples are rare or difficult to classify. The hybrid replay module 118 may provide many of the above-mentioned advantages over traditional machine learning techniques.

[0081] In one aspect, the hybrid replay module 118 may combine both the replay memory 302 and the replay GAN 304 with the streaming data 122 to generate an extended training set 308, which outputs data to the classifier / classifier 306. The classifier / classifier 306 may label the data from the extended training set 308 and predict the class as either real or fake. Furthermore, the classifier / classifier 306 may selectively store data (select and store 310) to update the memory. It is not necessary to store all the data. For example, it may store only a more representative data sample from the real data and ignore all fake data. This updated memory (select and store 310) may be input to the replay memory 302 for further refinement. The following describes step-by-step how the hybrid replay module 118 operates with the streaming data 122. First, the hybrid replay module 118 may collect and buffer the streaming data 122. The streaming data 122 may be labeled or unlabeled. The hybrid replay module 118 may then use the replay GAN 304 to generate new data samples similar to the data samples in the buffer. These new data samples may be labeled or unlabeled. The hybrid replay module 118 may then use the discriminator / classifier 306 to label the data samples in the buffer and the generated data samples. The discriminator / classifier 306 may selectively store data to update the replay memory 302 (select and store 310). The hybrid replay module 118 may then update the replay memory 302 with the selected and stored data 310. Rather than storing all data by the hybrid replay module 118, only the most useful and representative data may be stored. The ML model 106 may then be trained with the labeled data in the replay memory 302. Finally, the ML model 106 may be used to make predictions on new data.The above steps are continuously repeated by the hybrid replay module 118, resulting in an ML model 106 that is configured to accurately learn from streaming data 122, even when the data is imbalanced or contains rare or difficult examples.

[0082] FIG. 4 is a conceptual diagram illustrating an example of an architecture optimization method according to the techniques of the present disclosure. Architecture optimization is a process of automatically adjusting the complexity of a machine learning model to improve performance on a given task. In the context of lifelong inference, architecture optimization may be used to adapt the model to changes in sensor data over time. Adapting the model is important because sensor data can change due to multiple factors, such as, but not limited to, changes in the environment, hardware, or software. By adapting the model of the task-specific network module 116 to changes in sensor data, the architecture optimization module 120 may help prevent the task-specific network module 116 from experiencing catastrophic forgetting. As described above, catastrophic forgetting is the phenomenon in which a machine learning model forgets what it has learned when trained on new data.

[0083] In the context of lifelong inference, the architecture optimization module 120 may be used to adapt the task-specific network module 116 to different input requirements and application domains. Such adaptation may be achieved by leveraging differentiable weight-sharing neural architecture search (NAS) to efficiently evaluate architecture choices in parallel with online learning and by utilizing a cell-based approach that restricts the NAS search space based on prior experience and pilot studies to start with a limited set of high-performing modules. Differentiable weight-sharing NAS is a type of NAS that uses gradient descent to search for optimal architectures. Advantageously, weight-sharing NAS allows for the evaluation of architecture choices in parallel with online learning, which is important for lifelong inference because the task-specific network module 116 needs to be able to quickly adapt to changes in data. Cell-based NAS is a type of NAS that restricts the search space to a predefined set of cells. Cell-based NAS allows for the search to begin with a limited set of high-performing modules, thereby expediting the search process. One challenge in the art of architecture optimization is balancing the continuous adaptation of network architectures with increasing computational complexity and the time required to evaluate new choices. Another challenge is that constructing specific architectural components from neurons can be difficult in an online, continuous learning setting. The architecture optimization method 120 may address these challenges by leveraging a differentiable weight-sharing network architecture algorithm (NAS) to efficiently evaluate architecture choices in parallel with online learning, while utilizing a cell-based approach that constrains the NAS search space based on prior experience and pilot studies to start with a limited set of high-performing modules. The architecture optimization module 120 may have multiple advantages for lifelong inference. For example, the architecture optimization module 120 may help prevent catastrophic forgetting by adapting the task-specific network module 116 to changes in data over time.As another non-limiting example, the architecture optimization module 120 may improve the performance of the task-specific network module 116 for rare but important events.

[0084] Incremental Modular Architecture Search (pNAS) is a method for architecture optimization that starts with a limited set of preselected modules and gradually increases the complexity of the network by adding new modules and optimizing graph edges. Figure 4 is a conceptual diagram illustrating an example of how pNAS can be used to optimize the architecture of a neural network for lifelong inference. The search space may be initialized with a limited set of preselected modules, such as Shaped Multilayer Perceptions (MLP), Shaped ResBlock, and Graph Convolutional Networks (GCN). Each module may have a small number of hyperparameters. The NAS space may be modeled as a directed acyclic graph 402 of the current set of candidate modules. Nodes 404 in the acyclic graph 402 may represent modules, and edges 406 in the graph 402 may represent information flow between modules. At each meta-optimization step, edges 406 in the graph 402 (a mixture of modules) may be optimized. The edges 406 in the graph 402 may be optimized using various methods, such as, but not limited to, gradient descent or evolutionary algorithms. If the performance of the graph 402 is unsatisfactory, new nodes 404 may be added to the graph 402. It should be noted that new nodes may increase the complexity of the network. The NAS optimization may be terminated by early stopping against a performance threshold. Once the NAS optimization is terminated, the task-specific network modules 116 may be trained with the selected architecture. Some of the advantages of pNAS in lifelong inference are: pNAS is efficient because it starts with a pre-selected limited set of modules and gradually increases the network complexity; such a technique can quickly find a high-performance architecture; and pNAS is robust to data changes because it can adapt the network architecture to data changes.Such adaptation is important for lifelong inference because data may change over time.

[0085] The method performed by the architecture optimization module 120 may be summarized in the following steps, shown in FIG. 4 . First, the architecture optimization module 120 may establish initial unknown operations 410 for the edges 406 of the nodes 404. In other words, the architecture optimization module 120 may start with a predefined set of nodes (modules) 404 and connect them with edges 406. Each edge 406 may represent an unknown operation that needs to be optimized. Next, the architecture optimization module 120 may perform successive relaxation 412 by placing a mixture of operations 418 on each edge 406. In other words, the architecture optimization module 120 may replace each unknown operation with a mixture of all possible operations 418. Such replacement may enable the architecture optimization module 120 to optimize the architecture of the network in a continuous space. Next, the architecture optimization module 120 may perform bi-level optimization 414 to jointly train the mixture probabilities and weights. Bi-level optimization 414 is a technique that may be used to train complex models with multiple layers of optimization. In this case, the architecture optimization module 120 may use bi-level optimization 414 to jointly train the mixture probabilities and the weights 126 of the task-specific network module 116. The architecture optimization module 120 may then finalize the task-specific network module 116 based on the learned mixtures and probabilities. Once the task-specific network module 116 is trained, the architecture optimization module 120 may perform model finalization 416 by selecting the operator with the highest mixture probability. The method performed by the architecture optimization module 120 is efficient because it uses successive relaxation 412 and bi-level optimization 414 to train the task-specific network module 116. The successive relaxation 412 enables the architecture optimization module 120 to quickly find a high-performance architecture.The method performed by the architecture optimization module 120 is robust to changes in data because the method can optimize the architecture of the task-specific network module 116 in response to changes in data. This advantage is important for lifelong inference because data may change over time. Finally, the method performed by the architecture optimization module 120 is scalable to large datasets and complex tasks because the method may optimize the architecture of the task-specific network module 116 for the particular task at hand.

[0086] FIG. 5 is a conceptual diagram illustrating an example of an incremental DARTS optimization method according to the techniques of the present disclosure.

[0087] DARTS (Differentiable Architecture Search) is a method for NAS that uses gradient descent to optimize the architecture of a neural network. In one aspect, the architecture optimization module 120 may implement the DARTS algorithm, and considerations for such implementation are as follows: The architecture optimization module 120 implementing the DARTS algorithm may first define a supermodel, which is a large neural network that includes all possible architectures of a desired size and complexity. The supermodel may then be trained on the training dataset 113, and a gradient descent algorithm may be used to adjust the supermodel weights 126 so that the supermodel architecture is optimized for performance on the training dataset 113. Once the supermodel is trained, the architecture optimization module 120 may use the supermodel weights 126 to select an optimal architecture for the neural network. The architecture optimization module 120 may make this selection by considering the supermodel weights 126 and identifying the operations that are most important to performance. The architecture optimization module 120 may select a neural network architecture that includes these operations. The DARTS algorithm has many advantages over other NAS algorithms. DARTS is efficient. This is because DARTS uses gradient descent to optimize the neural network architecture. In other words, DARTS may quickly find a high-performance architecture. DARTS is robust to changes in the training dataset. This is because DARTS optimizes the neural network architecture for the performance of the training dataset. In other words, DARTS may find an architecture that works well on a variety of different datasets. DARTS is scalable for large datasets and complex tasks. This is because DARTS may optimize the neural network architecture for the specific task at hand.

[0088] The architecture optimization module 120 may model the operation at each node (i.e., node 404 shown in FIG. 4) as a mixture of candidate operations at that node. In other words, the operation at each node may be a weighted average of the candidate operations. The weights may be parameterized by a vector α(i;j).

[0089] The DARTS algorithm may solve a bi-level optimization problem to optimize the architecture of the neural network. The bi-level optimization problem may be applied iteratively between optimizing the architecture weights w (which parameterize the candidate operations) on the training data and optimizing the mixture weights α (which parameterize the weighting of the candidate operations) on the holdout data. The following is a more detailed description of the binary optimization problem. First, the DARTS algorithm may optimize the architecture weights w on the training data set. Such optimization may be performed by training a supermodel on the training data set. Next, the DARTS algorithm may optimize the mixture weights α on the holdout data. Such optimization may be performed by training the supermodel on the holdout data set and using a different loss function. The loss function used to train the mixture weights may be designed to encourage the supermodel to learn mixtures of candidate operations that are useful for performance on the holdout data. The above steps may be repeated until the architecture weights w and the mixture weights α converge. The resulting architecture weights and mixture weights may define the optimal architecture of the neural network. Although bilevel optimization problems are difficult to solve, the DARTS algorithm may use multiple techniques to solve them more efficiently. For example, the DARTS algorithm may use a gradient descent algorithm specifically designed for bilevel optimization. Furthermore, the DARTS algorithm may use multiple heuristics to reduce the search space of the bilevel optimization problem. DARTS has shown to be effective in finding high-performance neural network architectures for a variety of different tasks. For example, DARTS has been used to find architectures for image classification, object detection, and natural language processing.

[0090] At the end of DARTS training, architecture optimization module 120 may infer a discrete architecture using the argmax of α (i.e., only the o(i;j) at each (i;j) with the highest corresponding α(i;j) is retained). In other words, architecture optimization module 120 may select the operation with the highest weight at each node of the neural network.

[0091] In one aspect, the architecture optimization module 120 may implement a variant of the DARTS algorithm, i.e., I-DARTS. FIG. 5 illustrates a method 500 for implementing the I-DARTS algorithm by the architecture optimization module 120. An intuitive technique for incrementally modifying DARTS might be to simply run the DARTS algorithm with a core set of exemplar data for replay. However, this technique fails to take advantage of the many advances in incremental learning that go far beyond simple replay data. I-DARTS proposes a powerful incremental learning variant of DARTS by leveraging a dynamic memory repository (DMR) 502 to hold useful and representative precedent examples. The DMR 502 is a data structure that can efficiently store and retrieve data. The DMR 502 may allow the DMR 502 to learn over time to identify the most useful and representative examples. In one aspect, the architecture optimization module 120 may use the DMR 502 to store a core set of exemplar data and other significant past examples.

[0092] In incremental class learning (CIL), a task-specific network module 116 may be trained on a series of tasks, each with a different set of classes. The task-specific network module 116 must be able to learn new classes without forgetting previously learned classes. One challenge of CIL is catastrophic forgetting. Catastrophic forgetting occurs when a machine learning model forgets what it learned when trained on new data. Catastrophic forgetting can occur when the new data is significantly different from the data on which the machine learning model was originally trained. One way to address catastrophic forgetting is to use prediction space regularization. Prediction space regularization may encourage the task-specific network module 116 to learn new classes without discarding previously learned class representations. Prediction space regularization may be performed by penalizing the task-specific network module 116 for making changes to its predictions for old classes. Such a penalty may be achieved by using a loss function that compares the task-specific network module's 116 predictions for the old classes on the new data with the task-specific network module's 116 predictions for the old classes on the old data. Model space regularization is another way to address catastrophic forgetting. Model space regularization may impose a penalty on task-specific network modules 116 that make changes to their weights 126. Such a penalty may be achieved by using a loss function that compares the task-specific network module's 116 weights 126 for the new data with the weights 126 for the old data.

[0093] Related to class augmented learning (CIL), knowledge distillation (KD) is a technique that may be used to transfer knowledge from an old model to a new model. The old model may be trained with data from a previous task, while the new model may be trained with data from the current task. KD may be performed by forcing the new model to generate predictions similar to those of the old model. Such predictions may be achieved by using a loss function that compares the predictions of the two models. One way to use KD in CIL is to perform knowledge distillation from the old model to the new model for all of the data, including data from the previous task, by adding a KD loss term to the loss function of the new model.

[0094] With further reference to FIG. 5 , the architecture optimization module 120 may first train a supermodel 504 for all tasks using bi-level DARTS optimization. The DARTS optimization process described above involves searching for the optimal architecture of the supermodel by alternating between training the supermodel to minimize a loss function and updating the architecture to improve the performance of the supermodel 504. Once the supermodel 504 is trained, the architecture optimization module 120 may infer the optimal architecture for the current task 506 from the supermodel 504. In one aspect, the architecture optimization module 120 may perform this step by selecting the architecture with the best performance for the current task. The architecture optimization module 120 may retrain the optimal architecture 508 of the task-specific network module 116 on all training data for the current task, including the core set. The retraining step 508 may help fine-tune the architecture for a specific task. Furthermore, the architecture optimization module 120 may apply a class balance fine-tuning stage 510 to remove bias in the classification head. In one embodiment, the architecture optimization module 120 may perform this step by adjusting the weights of the classification head so that each class has an equal probability of being predicted. Finally, the architecture optimization module 120 may update the core set, which may be stored in the DMR 502, to best represent the training data of the previous task. In one embodiment, the architecture optimization module 120 may select a subset of the training data 113 that best represents the prior task. In one embodiment, the steps shown in FIG. 5 may be repeated until all tasks have been visited. I-DARTS is a powerful incremental learning algorithm that can be used to train a single neural network architecture to perform multiple tasks sequentially. I-DARTS has been shown to achieve state-of-the-art results on various incremental learning benchmarks.

[0095] 6 is a flowchart illustrating an exemplary mode of operation of hybrid replay module 118 in accordance with the techniques described in this disclosure. Although mode of operation 600 is described with respect to computing system 100 of FIG. 1 having processing circuitry 143 executing hybrid replay module 118, mode of operation 600 may also be performed by computing systems associated with other examples of machine learning systems described herein.

[0096] In mode operation 600, the processing circuit 143 executes the hybrid replay module 118. The hybrid replay module 118 may collect and buffer streaming data (602). The streaming data may be labeled or unlabeled. The hybrid replay module 118 may use a replay GAN to generate new data samples similar to the data samples in the buffer (604). These new data samples may be labeled or unlabeled. The hybrid replay module 118 may use a classifier / discriminator to label the data samples in the buffer and the generated data samples (606). The classifier / classifier may selectively store data for updating the replay memory. The hybrid replay module 118 may update the replay memory with the selected and stored data (608). Not all data may be stored by the hybrid replay module 118; only the most useful and representative data may be stored. The machine learning model 106 may then be trained with the labeled data in the replay memory (610). Finally, the machine learning model 106 may be used to make predictions on new data (612).

[0097] 7 is a flowchart illustrating an example mode of operation of architecture optimization module 120 implementing the I-DARTS algorithm according to the techniques described in this disclosure. Operational mode 700 is described with respect to computing system 100 of FIG. 1 having processing circuitry 143 executing architecture optimization module 120, although operational mode 700 may also be performed by computing systems for other example machine learning systems described herein.

[0098] In operational mode 700, processing circuit 143 executes architecture optimization module 120. Architecture optimization module 120 may first train a supermodel for multiple candidate tasks using bi-level DARTS optimization (702). Once the supermodel is trained, architecture optimization module 120 may infer an optimal architecture for the current task from the supermodel (704). Architecture optimization module 120 may retrain the optimal architecture of task-specific network module 116 on all training data for the current task, including the core set (706). Next, architecture optimization module 120 may apply a class balance fine-tuning stage to remove bias in the classification head (708). Finally, architecture optimization module 120 may update the core set, which may be stored in the DMR, to best represent previous task training data (710).

[0099] 8 is an exemplary diagram of a distributed data processing system in which aspects of the exemplary techniques may be implemented. Distributed data processing system 800 may include a network of computers in which aspects of the exemplary embodiments may be implemented. Distributed data processing system 800 may include at least one network 802, which is the medium used to provide communications links between various devices and computers connected together within distributed data processing system 800. Network 802 may include connections such as wired communications links, wireless communications links, or fiber optic cables.

[0100] In the depicted example, server 804 and server 806 are connected to network 802 along with storage device 808. Additionally, clients 810, 812, and 814 are also connected to network 802. These clients 810, 812, and 814 may be, for example, personal computers, network computers, etc. In the depicted example, server 804 provides data, such as live streaming trading data (streaming data 122), to clients 810, 812, and 814. Clients 810, 812, and 814 are clients to server 804 in the depicted example. Distributed data processing system 800 may include additional servers, clients, and other devices not shown.

[0101] In the depicted example, distributed data processing system 800 is the Internet with network 802 representing a worldwide collection of networks and gateways that use the Transmission Control Protocol / Internet Protocol (TCP / IP) suite of protocols to communicate with one another. At the heart of the Internet is a backbone of high-speed data communication lines between major nodes or host computers, which is made up of thousands of commercial, government, educational, and other computer systems that route data and messages. Of course, distributed data processing system 800 may be implemented to include many different types of networks, such as, for example, intranets, local area networks (LANs), wide area networks (WANs), etc. As noted above, FIG. 8 is intended as an illustration, rather than an architectural limitation for different aspects of the present disclosure, and thus the specific elements illustrated in FIG. 8 should not be considered limiting with regard to the environments in which example aspects of the present disclosure may be implemented.

[0102] As shown in FIG. 8 , one or more computing devices, e.g., server 804, may be specifically configured to implement hybrid replay module 118 and architecture optimization module 120 in accordance with one or more aspects described above. In one or more exemplary aspects, hybrid replay module 118 may operate in a manner as described above with respect to FIG. 3 , and architecture optimization module 120 may operate in a manner as described above with respect to FIG. 5 . Configuring a computing device may include providing application-specific hardware, firmware, etc. to facilitate performing the operations and generating output described herein in connection with the exemplary embodiments. Configuring a computing device may additionally or alternatively include providing software applications stored in one or more storage devices and loaded into memory of a computing device, such as server 904, to cause one or more hardware processors of the computing device to execute the software applications, configuring the processors to perform the operations and generate output described herein in connection with the exemplary embodiments. Furthermore, any combination of application-specific hardware, firmware, software applications executing on the hardware, etc. may be used without departing from the spirit and scope of the exemplary aspects.

[0103] The techniques described in this disclosure may be implemented, at least in part, in hardware, software, firmware, or any combination thereof. For example, various aspects of the described techniques may be implemented in one or more processors, including one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or any other equivalent integrated or discrete logic circuitry, and any combination of such components. The terms "processor" or "processing circuitry" may generally refer to the above-described logic circuitry alone or in combination with other logic circuits or other equivalent circuitry. A control unit comprising hardware may also perform one or more of the techniques of this disclosure.

[0104] Such hardware, software, and firmware may be implemented within the same device or within separate devices to support the various operations and functions described in this disclosure. Furthermore, any of the described units, modules, or components may be implemented together or separately as discrete but interoperable logical devices. The depiction of different features as modules or units is intended to emphasize different functional aspects and does not necessarily imply that such modules or units are required to be implemented by separate hardware or software components. Rather, functionality associated with one or more modules or units may be performed by separate hardware or software components, or may be integrated within a common or separate hardware or software component.

[0105] The techniques described in this disclosure may be embodied in or encoded on a computer-readable medium, such as a computer-readable storage medium containing instructions. The instructions embedded in or encoded on one or more computer-readable storage media may, for example, cause a programmable processor or other processor to perform a method when executed. The computer-readable storage medium may include random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electronically erasable programmable read-only memory (EEPROM), flash memory, hard disk, CD-ROM, floppy disk, cassette, magnetic medium, optical medium, or other computer-readable medium. The inventions disclosed herein include the following: [Aspect 1] a processing circuit in communication with a storage medium, the processing circuit configured to execute a machine learning system comprising at least a first module, a second module, and a third module, the machine learning system configured to train one or more machine learning models; the first module is configured to generate augmented input data based on streaming input data; the second module comprises a machine learning model configured to perform a particular task based at least in part on the augmented input data; the third module is configured to adapt a network architecture of the one or more machine learning models based on changes in the streaming input data. [Aspect 2] 2. The system of claim 1, wherein the streaming input data comprises streaming input data having a class imbalance among a plurality of classes represented in the streaming input data. [Aspect 3] 2. The system of claim 1, wherein the augmented input data includes one or more augmented samples of a minority class. [Aspect 4] 2. The system of aspect 1, wherein the machine learning system is configured to train the one or more machine learning models using one or more semi-supervised incremental learning methods. [Aspect 5] 10. The system of aspect 1, further comprising: one or more modules configured to process the streaming input data by performing at least one of a format conversion operation, a metadata derivation operation, or a data association operation. [Aspect 6] the first module further comprises a dynamic memory repository (DMR), a replay generation artificial intelligence (AI) architecture, and a classifier / classifier; the DMR is configured to selectively store one or more representative data samples; the generative AI architecture is configured to generate one or more new data samples similar to the one or more representative data samples stored in the DMR; 2. The system of claim 1, wherein the classifier / classifier is configured to distinguish between real and fake data in the one or more new data samples generated by the generative AI architecture. [Aspect 7] 7. The system of embodiment 6, wherein the identifier / classifier is further configured to select the one or more new data samples to be stored in the DMR. [Aspect 8] 2. The system of claim 1, wherein the third module is further configured to train a supermodel for a plurality of candidate tasks using at least one of a set of training data and input streaming data, and to infer an optimal architecture for a current task based on the trained supermodel. [Aspect 9] 9. The system of embodiment 8, wherein the third module is further configured to optimize weights of one or more architectures with respect to the training data. [Aspect 10] generating augmented input data based on the streaming input data using a first module; using a second module comprising a machine learning model to perform a specific task based at least in part on the augmented input data; using a third module to adapt a network architecture of one or more machine learning models based on changes in the streaming input data; and A method for providing [Aspect 11] 11. The method of aspect 10, wherein the streaming input data comprises streaming input data having a class imbalance among a plurality of classes represented in the streaming input data. [Aspect 12] 11. The method of claim 10, wherein the augmented input data includes one or more augmented samples of a minority class. [Aspect 13] The method of aspect 10, wherein the machine learning system, comprising at least the first module, the second module, and the third module, is configured to train the one or more machine learning models using one or more semi-supervised incremental learning methods. [Aspect 14] 11. The method of aspect 10, further comprising processing the streaming input data by performing at least one of a format conversion operation, a metadata derivation operation, or a data association operation using one or more modules. [Aspect 15] Selectively storing one or more representative data samples in a dynamic memory repository (DMR); generating one or more new data samples similar to the one or more representative data samples stored in the DMR using a generative artificial intelligence (AI) architecture; using a classifier / classifier to distinguish between real and fake data in the one or more new data samples generated by the generative AI architecture; 11. The method of embodiment 10, further comprising: [Aspect 16] 16. The method of embodiment 15, further comprising using the discriminator / classifier to select one or more new data samples to be stored in the DMR. [Aspect 17] using the third module to train a supermodel for a plurality of candidate tasks using at least one of a set of training data and input streaming data; inferring an optimal architecture for the current task based on the trained supermodel; 11. The method of embodiment 10, further comprising: [Aspect 18] 20. The method of embodiment 17, further comprising: optimizing weights of one or more architectures with respect to the training data using a third module. [Aspect 19] 1. A non-transitory computer-readable storage medium having encoded thereon instructions, the instructions comprising: generating augmented input data based on the streaming input data using a first module; using a second module comprising a machine learning model to perform a specific task based at least in part on the augmented input data; using a third module to adapt a network architecture of one or more machine learning models based on changes in the streaming input data; and a non-transitory computer-readable storage medium configured to cause a processing circuit to perform [Aspect 20] 20. The non-transitory computer-readable storage medium of claim 19, wherein the streaming input data includes streaming input data having a class imbalance among a plurality of classes represented in the streaming input data.

Claims

1. a processing circuit in communication with a storage medium, the processing circuit configured to execute a machine learning system comprising at least a first module, a second module, and a third module, the machine learning system configured to train one or more machine learning models; the first module is configured to generate augmented input data based on streaming input data; the second module comprises a machine learning model configured to perform a particular task based at least in part on the augmented input data; the third module is configured to adapt a network architecture of the one or more machine learning models based on changes in the streaming input data.

2. The system of claim 1 , wherein the streaming input data comprises streaming input data having a class imbalance among a plurality of classes represented in the streaming input data.

3. The system of claim 1 , wherein the augmented input data includes augmented samples of one or more minority classes.

4. The system of claim 1 , wherein the machine learning system is configured to train the one or more machine learning models using one or more semi-supervised incremental learning methods.

5. The system of claim 1 , further comprising one or more modules configured to process the streaming input data by performing at least one of a format conversion operation, a metadata derivation operation, or a data association operation.

6. the first module further comprises a dynamic memory repository (DMR), a replay generation artificial intelligence (AI) architecture, and a classifier / classifier; the DMR is configured to selectively store one or more representative data samples; the generative AI architecture is configured to generate one or more new data samples similar to the one or more representative data samples stored in the DMR; 2. The system of claim 1, wherein the classifier / classifier is configured to distinguish between real and fake data in the one or more new data samples generated by the generative AI architecture.

7. The system of claim 6 , wherein the classifier / classifier is further configured to select the one or more new data samples to be stored in the DMR.

8. 2. The system of claim 1 , wherein the third module is further configured to train a supermodel for a plurality of candidate tasks using at least one of a set of training data and input streaming data, and to infer an optimal architecture for a current task based on the trained supermodel.

9. The system of claim 8 , wherein the third module is further configured to optimize weights of one or more architectures for the training data.

10. A method for generating augmented input data based on streaming input data, the method comprising: the processing circuitry using a second module comprising a machine learning model to perform a specific task based at least in part on the augmented input data; the processing circuitry using a third module to adapt a network architecture of one or more machine learning models based on changes in the streaming input data; A method for providing the above.

11. The method of claim 10 , wherein the streaming input data comprises streaming input data having a class imbalance among a plurality of classes represented in the streaming input data.

12. The method of claim 10 , wherein the augmented input data includes augmented samples of one or more minority classes.

13. 11. The method of claim 10, wherein the machine learning system, comprising at least the first module, the second module, and the third module, is configured to train the one or more machine learning models using one or more semi-supervised incremental learning methods.

14. The method of claim 10, further comprising the processing circuit processing the streaming input data by performing at least one of a format conversion operation, a metadata derivation operation, or a data association operation using one or more modules.

15. The method of claim 14, further comprising: selectively storing one or more representative data samples in a dynamic memory repository (DMR); the processing circuitry using a generative artificial intelligence (AI) architecture to generate one or more new data samples similar to the one or more representative data samples stored in the DMR; the processing circuitry using a classifier / classifier to distinguish between real and fake data in the one or more new data samples generated by the generative AI architecture; The method of claim 10 further comprising:

16. The method of claim 15, further comprising the processing circuit using the discriminator / classifier to select one or more new data samples to be stored in the DMR.

17. The method of claim 17, wherein the processing circuit uses the third module to train a supermodel for a plurality of candidate tasks using at least one of a set of training data and input streaming data; the processing circuitry inferring an optimal architecture for a current task based on the trained supermodel; The method of claim 10 further comprising:

18. The method of claim 17, further comprising: the processing circuitry using the third module to optimize weights of one or more architectures for the training data.

19. 1. A non-transitory computer-readable storage medium having encoded thereon instructions, the instructions comprising: generating augmented input data based on the streaming input data using a first module; using a second module comprising a machine learning model to perform a specific task based at least in part on the augmented input data; using a third module to adapt a network architecture of one or more machine learning models based on changes in the streaming input data; and a non-transitory computer-readable storage medium configured to cause a processing circuit to perform

20. 20. The non-transitory computer-readable storage medium of claim 19, wherein the streaming input data comprises streaming input data having a class imbalance among a plurality of classes represented in the streaming input data.

Citation Information

Patent Citations

  • Providing apparatus, providing method, and program

    JP2021043772A

  • Machine learning program, machine learning method and machine learning device

    JP2022105916A

  • Reinforcement Learning in Real-Time Communication

    JP2022540137A

  • Reinforcement learning in real-time communications

    US20210012227A1

  • Automated handling of data drift in edge cloud environments

    WO2022214843A1