Selective acquisition of multi-modal time data

By combining selecting neural networks and prediction models, using reinforcement learning technology, adaptively determine which data modes should be collected at each time point, solving the problems of resource waste and increased risks in multimodal data processing, and achieving the effect of efficient resource use and risk reduction.

CN120077384APending Publication Date: 2025-05-30GDM HOLDING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380074315.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-21
Filing Date
2023-10-21
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, when processing multimodal data, it is difficult to efficiently determine which data modes should be collected at each point in time, resulting in waste of resources and increased risks.

Method used

By using select neural networks and prediction models, combined with reinforcement learning techniques, adaptively determine which data modes should be collected at each time point and optimize the tradeoff between acquisition cost and prediction performance.

Benefits of technology

Efficient use of resources and reduced risks in multimodal data processing while maintaining acceptable predictive performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120077384A_ABST
    Figure CN120077384A_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating predictions characterizing an environment. In one aspect, a method includes obtaining, for each time step in a sequence of a plurality of time steps, a respective observation characterizing a state of an environment, including, for each time step subsequent to a first time step in the sequence of time steps: processing a network input including observations obtained for one or more previous time steps, generating a plurality of acquisition decisions; obtaining an observation for the time step, where the observation includes data corresponding to modalities selected for acquisition at the time step and does not include data corresponding to modalities not selected for acquisition at the time step; and processing a model input to generate a prediction, the model input comprising an observation for each time step in the sequence of time steps.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims the priority of Greek National Application No. 20220100868, filed on October 21, 2022. The disclosure of the prior application is regarded as part of the disclosure of this application and is incorporated herein by reference. Background Art

[0003] This specification relates to using machine learning models to process data.

[0004] Machine learning models receive inputs and generate outputs based on the received inputs, such as predicted outputs. Some machine learning models are parametric models and generate outputs based on the received inputs and the values of the parameters of the model.

[0005] Some machine learning models are deep models that employ multiple - layer models to generate outputs for received inputs. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers, and each of the one or more hidden layers applies a non - linear transformation to the received input to generate an output. Summary of the Invention

[0006] This specification generally describes a system for generating predictions characterizing an environment, which is implemented as a computer program on one or more computers in one or more locations.

[0007] According to one aspect, there is provided a method executed by one or more computers, the method comprising: obtaining, for each time step in a sequence of multiple time steps, a respective observation characterizing a state of the environment, including, for each time step starting from the first time step in the sequence of time steps: using a selection neural network to process a network input including observations obtained for any previous time step to generate a plurality of acquisition decisions, wherein each acquisition decision corresponds to a respective modality from a set of multiple modalities and defines whether data corresponding to the modality is selected for acquisition at that time step; obtaining an observation for that time step, wherein the observation includes only data corresponding to the modalities selected for acquisition at that time step; and using a prediction model to process a model input, the model input including observations for each time step in the sequence of time steps, to generate a prediction characterizing the environment.

[0008] In some implementations, the method further includes: determining an acquisition cost based on a respective modality selected for acquisition at each time step in a sequence of time steps; determining a reward based at least in part on the acquisition cost; and training a selection neural network based on the reward using reinforcement learning techniques.

[0009] In some implementations, each modality in a set of modalities is associated with a respective cost factor, and wherein determining the acquisition cost includes: for each time step in the sequence of time steps, determining a respective acquisition cost for the time step based on the respective cost factor associated with each modality selected for acquisition at the time step; and determining the acquisition cost as a combination of the acquisition costs for the time steps.

[0010] In some implementations, for each time step in the sequence of time steps, determining the acquisition cost for the time step includes: determining the acquisition cost for the time step as a sum of the cost factors associated with each modality selected for acquisition at the time step.

[0011] In some implementations, determining the acquisition cost as a combination of the acquisition costs for the time steps includes: determining the acquisition cost as a sum of the acquisition costs for the time steps.

[0012] In some implementations, for one or more of the modalities, the cost factor for the modality is based at least in part on an amount of resource usage required to capture data corresponding to the modality.

[0013] In some implementations, the resource usage required to capture data corresponding to a modality at least represents the energy usage required to capture data corresponding to the modality.

[0014] In some implementations, the resource usage required to capture data corresponding to a modality at least represents an amount of time required to capture data corresponding to the modality.

[0015] In some implementations, for one or more of the modalities, the cost factor for the modality is based at least in part on a risk associated with capturing data corresponding to the modality.

[0016] In some implementations, the environment includes a patient, and the risk associated with capturing data corresponding to a modality is based at least in part on a medical risk to the patient resulting from capturing data corresponding to the modality.

[0017] In some implementations, the method further includes: determining a prediction error that measures an error in a prediction generated by a prediction model; and determining a reward based on both (i) the acquisition cost and (ii) the prediction error.

[0018] In some implementations, the prediction model is a machine learning model.

[0019] In some implementations, the prediction model includes a neural network.

[0020] In some implementations, the method further includes: training a predictive machine learning model to optimize an objective function that depends on the prediction error of the predictive machine learning model.

[0021] In some implementations, for each time step in a sequence of time steps, selecting the network input of the neural network at that time step further includes: identifying data of acquisition decisions for any modality at any previous time step.

[0022] In some implementations, the method further includes, for each time step in one or more time steps in a sequence of time steps: using the prediction model to process a model input to generate an intermediate prediction characterizing the environment, the model input including observations for that time step and observations for one or more previous time steps in the sequence of time steps; and determining an intermediate prediction error that measures the error in the intermediate prediction generated by the prediction model; and determining a reward at least in part based on the intermediate prediction error.

[0023] In some implementations, the set of modalities includes an imaging modality, and wherein the data corresponding to the imaging modality includes image data.

[0024] In some implementations, the set of modalities includes a medical imaging modality.

[0025] In some implementations, the environment is a medical environment including a patient.

[0026] In some implementations, the prediction characterizing the environment includes a predicted medical diagnosis of the patient.

[0027] In some implementations, the prediction characterizing the environment includes a prediction of a medical treatment to be applied to the patient.

[0028] In some implementations, the method further includes, for each time step after a first time step in a sequence of time steps: determining (i) that data corresponding to the modality selected for acquisition at that time step from the set of modalities will be included in the observations for that time step, and (ii) that data corresponding to the modality not selected for acquisition at that time step from the set of modalities will not be included in the observations for that time step.

[0029] In some implementations, for each of one or more time steps after a first time step in a sequence of time steps: only a proper subset of the modalities in the set of modalities is selected for acquisition at that time step.

[0030] In some implementations, the method further includes, for each time step after a first time step in a sequence of time steps: acquiring data only for the modalities that are selected for acquisition at that time step.

[0031] According to another aspect, there is provided a system including: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of the methods described herein.

[0032] According to another aspect, there is provided one or more non-transitory computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the methods described herein.

[0033] The subject matter described in this specification can be implemented in particular implementations so as to achieve one or more of the following advantages.

[0034] This specification describes a system for processing multi-modal data captured within a sequence of time points to generate predictions characterizing an environment. In many real-world scenarios, capturing data corresponding to modalities can incur significant costs, for example, in terms of resource consumption (such as energy or time consumption), or in terms of risk (such as medical risk, e.g., due to a patient being exposed to radiation when acquiring a medical image of the patient such as an x-ray image or a CT image). Additionally, processing data corresponding to certain modalities can also incur significant costs, for example, in terms of computing resources (such as memory and computing power), for example, for high-dimensional data such as image data, video data, or audio data. The system described in this specification can adaptively determine which data modalities to acquire at each time point, and for some time points, fewer than all available modalities (or even avoid acquiring any modalities) can be acquired.

[0035] Machine learning techniques can be used to train a system to optimize the trade - off between acquisition cost and prediction performance. In particular, the system can be trained to achieve acceptable prediction performance while minimizing the acquisition cost across available modalities, thereby enabling more efficient use of resources (e.g., energy resources or computing resources) and reduction of risks (e.g., medical risks). In some cases, the system can be trained to optimize prediction performance while encouraging (or requiring) an acquisition cost that meets a cost budget.

[0036] Details of one or more implementations of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the specification, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 An example neural network system is shown.

[0038] Figure 2 is a flowchart of an example process for generating predictions characterizing an environment.

[0039] Figure 3 is a flowchart of an example process for training a selection neural network.

[0040] Figure 4 is for Figure 3 a sub - step of one of the steps of a process.

[0041] Figure 5 is for Figure 3 another sub - step of one of the steps of a process.

[0042] Figure 6 is an example illustration of using a selection neural network to generate predictions and determining one or more updates to parameter values of the selection neural network based on the predictions.

[0043] Like reference numerals and names in the various figures indicate like elements. DETAILED DESCRIPTION

[0044] Figure 1 An example neural network system 100 is shown. Neural network system 100 is an example of a system implemented as a computer program on one or more computers in one or more locations, where the systems, components, and techniques described below can be implemented.

[0045] Neural network system 100 includes a selection neural network 110, a prediction model 120, a data acquisition engine 130, and optionally, in some implementations, a training engine 140.

[0046] Generally, the neural network system 100 is a system that collects input data 102 generated within or about an environment at each time step within a sequence of multiple time steps and makes one or more predictions 122 that characterize one or more aspects of the environment. For example, the neural network system 100 may output a prediction 122 after input data 102 has been collected for the last time step in a sequence of multiple time steps.

[0047] At any given time step in the sequence, the input data 102 collected by the system may potentially (but not necessarily) include data from a set of two or more available modalities. In this specification, a data "modality" refers to, for example, a type of data generated using a particular sensor or diagnostic technique (e.g., a medical diagnostic technique).

[0048] The set of modalities can include any suitable modalities. Several examples of possible modalities are described next. In some implementations, the set of modalities includes one or more of these examples.

[0049] In some implementations, the set of modalities includes an imaging modality, and the data corresponding to the imaging modality includes image data (e.g., one-dimensional (1D) image data, two-dimensional (2D) image data, three-dimensional (3D) image data, etc.). The image data can include pixel value data, such as color or monochromatic pixel value data.

[0050] For example, the set of modalities can include one or more medical imaging modalities, such as computed tomography (CT) modality, ultrasound (US) modality, magnetic resonance imaging (MRI) modality, x-ray modality, histology imaging modality, electroencephalogram (EEG) modality, electromyogram (EMG) modality, electrocardiogram (ECG) modality, etc.

[0051] As another example, the set of modalities can include a camera modality, such as in the case of using a camera (e.g., a visible spectrum camera or an infrared spectrum camera) to capture data corresponding to the camera modality.

[0052] In some implementations, the set of modalities can include a genetic data modality, and the data corresponding to the genetic data modality includes genetic data. The genetic data can include, for example, data defining the respective expression levels (in a subject) of each gene in a set of genes. The genetic data can be obtained by performing a suitable diagnostic technique (such as DNA or RNA sequencing) on genetic material obtained from a subject.

[0053] In some implementations, the set of modalities can include proteomics data modalities, and the data corresponding to the proteomics data modalities includes proteomics data. The proteomics data can include, for example, data defining the respective expression levels (in a subject) of each protein in a set of proteins.

[0054] In some implementations, the set of modalities can include a blood testing modality, and the data corresponding to the blood testing modality can include data defining the levels of one or more components (e.g., sodium, potassium, chloride, bicarbonate, blood urea nitrogen, magnesium, creatinine, glucose, calcium, cholesterol, etc.) in the blood of a subject.

[0055] In some implementations, the set of modalities can include an audio modality, and the data corresponding to the audio modality can include audio data, such as audio data characterizing words spoken by a person, audio data characterizing sounds emitted by one or more body parts (e.g., heart, digestive system, lungs, etc.) of a person, etc. The audio data can include data defining an audio waveform, such as a series of values defining the waveform in the time domain and / or frequency domain.

[0056] In some implementations, the set of modalities can include a biopsy modality, and the data corresponding to the biopsy modality can characterize a sample of cells or tissue obtained from a patient by biopsy. For example, the data corresponding to the biopsy modality can include a microscopic image of the sample obtained from the patient.

[0057] In some implementations, the set of modalities can include modalities for measuring one or more of the following: humidity, light, air quality, sound, body temperature, wind speed, pH, etc.

[0058] In a particular implementation, the set of modalities includes at least: a medical imaging modality and a blood testing modality.

[0059] In a particular implementation, the set of modalities includes at least: a medical imaging modality, a blood testing modality, and a biopsy modality.

[0060] In a particular implementation, the set of modalities includes at least: a medical imaging modality, a blood testing modality, a biopsy modality, and a genetic data modality.

[0061] In some of these implementations, these data types not only differ in feature space and dimension, but also in the data capture processes and costs associated with capturing data corresponding to these data types. For example, medical imaging can be at the discretion of a doctor and requires a relatively higher cost (e.g., in terms of resource consumption or risk) to capture, while blood pressure and body temperature can be monitored regularly and can be captured at a relatively lower cost.

[0062] The environment can be any suitable environment, such as a real-world environment, such as a medical environment, an agricultural environment, an aquaculture environment, an industrial environment, or a scientific environment.

[0063] As described above, the medical environment can include patients, and one or more of the modalities can be modalities that generate data representative of the patient.

[0064] The industrial environment can include, for example, a manufacturing facility (e.g., which includes one or more industrial machines for producing a finished product), a chemical processing facility (e.g., which includes one or more industrial machines for chemical processing), a data center facility (e.g., which includes a series of computing units for performing computing tasks, e.g., processors), or an energy production facility (e.g., a nuclear power plant, a hydropower plant, a photovoltaic power plant, etc.). One or more of the modalities can be modalities that generate data representative of the facility (e.g., data generated by one or more sensors located within or around the facility, such as sensors for measuring the state of industrial machines or computing units within the facility). Predictions representative of the industrial environment can include predicted values of one or more properties that can be determined based on sensor values, or predicted values of one or more properties measured by the sensors, e.g., the predicted values can include predicted sensor values.

[0065] The scientific environment can include a series of subjects being studied for scientific purposes, where the subjects can include, for example, plants, animals, cells, tissues, etc.

[0066] The environment can evolve over time, and thus the input data 102 generated within or with respect to the environment at a first time step can have different values than the input data 102 generated within or with respect to the environment at a second time step. A series of input data 102 collected within a sequence of multiple time steps can thus be referred to as "time" series input data because, in some implementations, the input data are arranged according to the time step at which they are captured. For example, the most recent input data is the last input data in the time series of input data, and the least recent input data is the first input data in the time series.

[0067] One or more predictions 122 that characterize one or more aspects of an environment are made by a prediction model 120 based on input data 102. The prediction model 120 can be configured as a machine learning model, which can have any suitable machine learning model architecture. For example, the prediction model 120 can be implemented as a neural network, or a decision tree, or a random forest, or a support vector machine, or a linear regression model, etc. In a particular example, the prediction model 120 can be implemented as a neural network, which can include any suitable type of neural network layer (e.g., fully connected layer, attention layer, convolutional layer, etc.) connected in any suitable number (e.g., 5 layers, or 10 layers, or 100 layers) and in any suitable configuration (e.g., as a directed graph of layers).

[0068] Several examples of possible predictions 122 are described next.

[0069] In some implementations, the environment is a medical environment including a patient, and the prediction defines the predicted medical treatment to be applied to the patient. For example, the prediction can include a respective score for each medical treatment in a set of medical treatments, where the score of a medical treatment defines the likelihood that the medical treatment should be applied to the patient. The set of medical treatments can include medical treatments corresponding to administering a drug to the patient, intervening on the patient (e.g., surgery), etc.

[0070] In some implementations, the environment is a medical environment including a patient, and the prediction defines the predicted medical diagnosis for the patient. For example, the prediction can include a respective score for each medical diagnosis in a set of medical diagnoses, where the score of a medical diagnosis defines the likelihood that the medical diagnosis is applicable to the patient. The set of medical diagnoses can include diagnoses of one or more diseases such as, for example, cancer, diabetes, heart failure, Alzheimer's disease, influenza, measles, streptococcal pharyngitis, sepsis, etc.

[0071] In some implementations, the environment is an agricultural environment (e.g., an environment for cultivating crops) or an aquaculture environment (e.g., an environment for cultivating aquatic organisms), and the prediction defines the predicted yield (e.g., measured in tons of crops or aquatic organisms), or the predicted amount of time until the crops or aquatic organisms in the environment should be harvested (e.g., measured in days).

[0072] In some implementations, the environment is an industrial environment, and the prediction defines the predicted production level of the industrial environment within a predefined time range (e.g., 1 hour, 1 day, or 1 week), e.g., the number of units of a product generated by a manufacturing facility, or the amount of a chemical produced by a chemical processing facility, or the number of computational tasks completed by a data center facility, or the amount of energy generated by an energy production facility.

[0073] In some implementations, the environment is a scientific environment, and the prediction defines the predicted outcome of a scientific study, e.g., the health of a subject of a scientific study, such as the integrity of the cell walls of a series of cells at the end of a study, or the weight of animals in an animal population at the end of a study.

[0074] In these example environments and many other real-world environments, capturing data corresponding to modalities often incurs significant costs, e.g., in terms of resource consumption (e.g., consumption of energy or time), or in terms of risks (e.g., medical risks, such as a patient being exposed to radiation due to the acquisition of a medical image of the patient, e.g., an x-ray image or a CT image). Additionally, processing data corresponding to certain modalities can also incur significant costs, e.g., in terms of computing resources (e.g., memory and computing power), e.g., for high-dimensional data such as image data, video data, or audio data.

[0075] Thus, although the neural network system 100 can potentially receive input data corresponding to each of two or more available modalities at each time step, the system may not actually do so, and instead may only acquire (and thereafter receive) input data corresponding to each modality in a proper subset of the set of two or more available modalities at each of one or more time steps. The proper subset includes at least one modality in the set of two or more available modalities, but less than all modalities in the set.

[0076] In Figure 1 the example of, the neural network system 100 can potentially receive data corresponding to a set of three data modalities available to the system: data 102A corresponding to modality A, data 102B corresponding to modality B, and data 102C corresponding to modality C.

[0077] However, as shown, the actually acquired input data 102 is multimodal data that includes only data corresponding to two of the three available data modalities respectively - data 102A corresponding to modality A and data 102B corresponding to modality B. That is, the data 102C corresponding to modality C is not acquired and thus not received by the system, and the actually acquired input data 102 does not include the data 102C corresponding to modality C.

[0078] In other examples, the input data can potentially (but not necessarily) include data corresponding to a smaller (e.g., two) or larger (e.g., ten, one hundred, or more) set of available modalities. Similarly, in these examples, the actually acquired input data 102 can include data corresponding to each modality in a proper subset of a smaller or larger set of available modalities.

[0079] In particular, for each time step after the first time step in the sequence of time steps, the neural network system 100 uses the selection neural network 110 to make acquisition decisions, where each acquisition decision defines whether data corresponding to each modality in the set of modalities will be selected for acquisition.

[0080] At any given time step after the first time step in the sequence of multiple time steps, the selection neural network 110 processes the network input to generate multiple acquisition decisions for that given time step, where the network input includes (i) previous observations 112 obtained for one or more previous time steps, and optionally includes (ii) data identifying the acquisition decisions for any modality at any previous time step. Each acquisition decision corresponds to a respective modality from the set of multiple modalities and defines whether the data corresponding to that modality should be selected for acquisition at the given time step. As will be further explained below, an "observation" refers to data generated by the data acquisition engine 130 from the acquired input data 102 and provided to the selection neural network 110 and / or the prediction model 120 for further processing.

[0081] The selection neural network 110 can have any suitable neural network architecture that allows the selection neural network 110 to generate acquisition decisions from previous observations. In particular, the selection neural network can include any suitable type of neural network layer (e.g., fully connected layer, attention layer, convolutional layer, etc.) connected in any suitable number (e.g., 5 layers, or 10 layers, or 100 layers) and in any suitable configuration (e.g., as a directed graph of layers).

[0082] As a specific example, the selection neural network 110 and the prediction model 120 can each be configured to have a respective neural network of an architecture described in: Andrew Jaegle et al., Perceiver IO: A General Architecture for Structured Inputs & Outputs, International Conference on Representation Learning, 2022.

[0083] At Figure 1In the example, at a specific time step, the selection neural network 110 generates a total of three acquisition decisions: Decision A corresponding to modality A, Decision B corresponding to modality B, and Decision C corresponding to modality C. Specifically, Decision A defines that the data corresponding to modality A should be selected for acquisition at this time step, Decision B defines that the data corresponding to modality B should be selected for acquisition at this time step, and Decision C defines that the data corresponding to modality C should not be selected for acquisition at this time step. The selection neural network 110 can generate different acquisition decisions at other time steps.

[0084] Each acquisition decision can be deterministically generated, for example, by the output of the selection neural network 110. For example, the output layer of the selection neural network 110 can include corresponding neurons for each modality; and each modality is selected for acquisition only when the activation of the neuron exceeds a predefined threshold. Alternatively, each acquisition decision can be generated randomly. For example, the output of the selection neural network 110 parameterizes a distribution, and the acquisition decision is sampled from this distribution. For example, each acquisition decision can be a binary decision, where 0 indicates that the data corresponding to a specific modality should not be selected for acquisition, and 1 indicates that the data corresponding to a specific modality should be selected for acquisition.

[0085] The neural network system 100 then uses the data acquisition engine 130 to complete the acquisition decisions generated by the selection neural network 110. That is, the neural network system 100 provides the acquisition decisions to the data acquisition engine 130, and the data acquisition engine 130 causes data to be acquired only for the modalities selected for acquisition at a given time step.

[0086] In Figure 1 the example, the data acquisition engine 130 acquires the data corresponding to modality A and the data corresponding to modality B according to the acquisition decisions generated by the selection neural network 110. The data acquisition engine 130 avoids acquiring the data corresponding to modality C.

[0087] In some implementations, the data acquisition engine 130 can complete (e.g., execute or implement) the acquisition decision by transmitting an electronic signal to a sensor or another electronic device communicatively coupled to the system with environmental sensing capabilities to capture data corresponding to one of the selected modalities. In response to receiving the electronic signal, the sensor operates to capture data within the environment or data about the environment.

[0088] In some implementations, the data acquisition engine 130 can complete the acquisition decision by generating and outputting a prompt to be presented to the user via a user interface device. The user interface device can be any suitable fixed or mobile computing device, such as a desktop computer, a workstation in a medical environment, a tablet computer, a smart phone, or a smart watch.

[0089] This prompt can help guide the user to capture data according to the acquisition decisions generated by the selection neural network 110. This prompt can indicate what modality of data the user should capture. For example, the prompt can present in a window text requiring that data corresponding to one of the selected modalities should be captured. The user can interact with the user interface device to view the selected modality and upload the data corresponding to the selected modality after the data is captured.

[0090] The neural network system 100 then generates an observation 112 for a given time step from the input data 102, which is collected by the data acquisition engine 130 according to a plurality of acquisition decisions generated by the selection neural network 110.

[0091] The observation 112 for a given time step: (i) includes data corresponding to the modalities selected for acquisition at the given time step from the set of modalities, and (ii) excludes, i.e., does not include, data corresponding to the modalities in the set of modalities that are not selected for acquisition at the given time step. After being generated, the observation 112 is then provided to the prediction model 120 for further processing.

[0092] In this way, although the potentially available data of the system includes multi-modal data respectively corresponding to the set of modalities, in fact, only data corresponding to a proper subset of the modalities in the set of modalities can be selected by the neural network system 100 for acquisition. For example, only data corresponding to a small number of modalities among a relatively large number of modalities respectively can be selected by the neural network system 100 for acquisition and then used by the prediction model 120 to generate predictions 122.

[0093] By incorporating the selection neural network 110 and collecting data according to the acquisition decisions generated by the selection neural network 110, the neural network system 100 can reduce the amount of computational resources consumed by the prediction process because it is no longer necessary to repeatedly collect data from all the modalities in the set of modalities and then process the data. Instead, at each time step in at least some of the time steps, it is only necessary to collect data from a relatively small number of selected modalities and then process the data.

[0094] The training engine 140, when included, can train the selected neural network 110 and optionally includes a training prediction model 120 to determine the trained parameter values of the selected neural network 110 and optionally determine the trained parameter values of the prediction model 120, the trained parameter values enabling the selected neural network 110 to generate an acquisition decision that can lead to a reduction in the consumption of computing resources by the system while still maintaining prediction performance, e.g., in terms of the accuracy of the prediction 122. Thus, in some implementations, the selected neural network 110 and the prediction model 120 can be jointly trained by the training engine 140.

[0095] In Figure 1 an example of, the training engine 140 includes or has access to a cost calculation engine 145. The cost calculation engine 145 is configured to calculate the acquisition cost associated with the modality selected for acquisition based on the acquisition decision generated by the selected neural network 110.

[0096] The training engine 140 can thus apply reinforcement learning techniques that use the rewards derived from the acquisition costs to jointly train the selected neural network 110 with the prediction model 120 to optimize the trade-off between the acquisition cost and the prediction performance.

[0097] In particular, the training engine 140 can train the selected neural network 110 and the prediction model 120 to achieve acceptable prediction performance while minimizing the acquisition costs across the available modalities, thereby enabling more efficient use of resources (e.g., energy resources or computing resources) and a reduction in risks (e.g., medical risks).

[0098] The cost calculation engine 145 can be configured to calculate the acquisition cost of a modality in a set of modalities based on any suitable criteria. Several examples of possible criteria for setting the acquisition cost of a modality are described next.

[0099] In some implementations, the acquisition cost of a modality can be at least partially based on the amount of resource usage (e.g., energy or time) required to acquire the data corresponding to that modality.

[0100] In some implementations, the acquisition cost of a modality can be at least partially based on the amount of risk required to acquire the data corresponding to that modality. For example, in a medical setting, acquiring data corresponding to a biopsy modality can pose a risk of patient infection, and acquiring data corresponding to an x-ray modality can pose a risk of exposing the patient to unhealthy levels of radiation. When acquiring data corresponding to a modality, the amount of risk can be determined based on statistical data characterizing different outcomes (e.g., patient outcomes).

[0101] In some implementations, the acquisition cost of a modality can be at least partially based on the level of interference caused by acquiring data corresponding to that modality. For example, in an industrial environment, acquiring data corresponding to a modality can include running diagnostic tests that reduce the production of an industrial facility. As another example, in a scientific environment, acquiring data corresponding to a modality can include interfering with conditions in the environment in a way that may compromise the validity or accuracy of the results of an experiment (e.g., by performing tests on one or more subjects in the environment). This will be further described with reference to Figures 3 to 6 training the selection neural network 110.

[0102] Figure 2 is a flowchart of an example process 200 for generating predictions characterizing an environment. For convenience, process 200 will be described as being performed by a system of one or more computers located in one or more locations. For example, a suitably programmed neural network system (e.g., Figure 1 the neural network system 100) can perform process 200.

[0103] The environment can be any suitable environment, such as a real-world environment, such as a medical environment, an agricultural environment, an aquaculture environment, an industrial environment, or a scientific environment.

[0104] The system repeatedly performs steps 202 and 204 to obtain a corresponding observation characterizing the state of the environment for each time step in a sequence of multiple time steps. That is, the system performs one iteration of steps 202 and 204 for each time step in a sequence of multiple time steps.

[0105] In some implementations, the number of time steps is fixed (predefined). For example, the system can generate a sequence-level prediction after a predefined number of time steps have elapsed. In other implementations, the number of time steps is flexible, and different sequences can include different numbers of time steps. For example, the system can repeatedly perform iterations of steps 202 and 204 until a termination signal (e.g., a flag or another indicator) is received at a given time step, which indicates that the given time step is the last time step in the sequence. For example, if a given time step is not the last time step in the sequence, the flag can be set to a first value, and if the given time step is the last time step in the sequence, the flag can be set to a second value. The termination signal can be based on the predictions generated by the system.

[0106] For each time step after the first time step in a sequence of multiple time steps, the system uses a selection neural network to process a network input to generate multiple acquisition decisions for that time step, where the network input includes (i) observations obtained for one or more previous time steps, and optionally includes (ii) data identifying acquisition decisions for any modality at any previous time step (step 202). An "observation" refers to data generated by the data acquisition engine 130 from the acquired input data 102 and provided to the selection neural network 110 and / or the prediction model 120 for further processing. Each acquisition decision corresponds to a respective modality from a set of multiple modalities and defines whether data corresponding to that modality is selected for acquisition at that time step.

[0107] For the first time step, since there are no previous time steps, some implementations of the system can alternatively provide a predetermined network input (i.e., an input with predetermined values) for the selection neural network to process. Some other implementations of the system can alternatively acquire a default (e.g., random or predefined) set of modalities, i.e., not use the selection neural network to generate any acquisition decisions for the first time step.

[0108] The system obtains observations for the time step based on the multiple acquisition decisions generated by the selection neural network (step 204). For example, the system can use the data acquisition engine to acquire data corresponding to each modality selected for acquisition according to the acquisition decisions, and then include the acquired data in the observations.

[0109] Specifically, the observations (i) include data corresponding to modalities from the set of modalities that are selected for acquisition at that time step, and (ii) do not include data corresponding to modalities from the set of modalities that are not selected for acquisition at that time step.

[0110] The modalities selected can and typically will vary from one time step to another. In other words, the system can obtain data corresponding to different modalities at different time steps.

[0111] In some examples, for one or more time steps in a sequence of multiple time steps, the system may obtain an observation that includes data corresponding to all modalities in the set. In another example, for one or more time steps, the system may obtain an observation that includes data corresponding to a proper subset of the set of modalities (and does not include data corresponding to any remaining modalities not in the proper subset). A "proper" subset of a set is a subset that includes one or more, but not all, of the elements in the set. In another example, for one or more time steps, the system may obtain an empty observation that does not include data corresponding to any modality in the set.

[0112] After performing the iterations of steps 202 and 204 at the last time step in the sequence of multiple time steps, the system uses a prediction model to process the model input to generate a prediction characterizing the environment, where the model input includes the observations for each time step in the sequence of time steps (step 206).

[0113] An example algorithm for generating a prediction is shown below.

[0114]

[0115] In Algorithm 1, each input x includes a sequence of observations x i =(x i,1 ,..., x i,T ). At each time step t, the observation x i,t includes data corresponding to M modalities x i,t =(x i,t,1 ,..., x i,t,M ). Each modality can be high-dimensional For example, x i,t,m can be a single frame in a video with dimension d m =H·W·C, where H is the height, W is the width, and C is the number of color channels.

[0116] At each time step t∈(1,..., T), by sampling from the output of a selection neural network (referred to as the agent in Algorithm 1): Multiple acquisition decisions a t =(a t,1 ,..., a t,M ) across modalities are generated. Here, a t,M ∈(0, 1) is a binary indicator of whether modality m was acquired at time step t. Instead of x, is used to emphasize that the input data may contain missing entries, and θ are the parameters of the selection neural network. At each time step t, and for each modality m, only if a m,tCollect data x corresponding to modality m only when = 1 m,t .

[0117] By repeatedly performing process 200, the system can generate different predictions that characterize the same or different aspects of the environment. That is, process 200 can be performed as part of generating predictions from a sequence of observations for which the desired output (i.e., the desired prediction that the system should generate from the sequence of observations) is unknown. One or more actions can be performed based on the predictions. For example, an agent (such as an electromechanical agent) that interacts with the real-world environment to perform a task can select one or more actions to perform in the real-world environment based on the predictions.

[0118] Some of the steps in all of the steps of process 200 can also be performed as part of processing a sequence of observations derived from a set of training data (i.e., a sequence of observations derived from input data for which the prediction that the system should generate is known) in order to train the trainable components of the system to determine the trained values of the parameters of these components.

[0119] Figure 3 is a flowchart of an example process 300 for training a selection neural network. For convenience, process 300 will be described as being performed by a system of one or more computers located at one or more locations. For example, a suitably programmed neural network system (e.g., Figure 1 the neural network system 100) can perform process 300.

[0120] During training, process 300 can be performed for each training input selected from a set of training data after process 200, where the set of training data is derived from multiple time series of input data generated within or with respect to an environment (e.g., one of the physical environments described above or a computer simulation of one of these physical environments). That is, for each training input, the system performs process 200 to generate a prediction characterizing the environment using the selection neural network and according to the current values of the parameters of the selection neural network, and then performs process 300 to determine one or more updates to the parameter values of the selection neural network based on the prediction generated in process 200.

[0121] Process 300 is illustrated in Figure 6 and Figure 6 shows an example of using a selection neural network to generate a prediction and determining one or more updates to the parameter values of the selection neural network based on that prediction.

[0122] As Figure 6As shown, at any given time step, a neural network (“agent”) is selected to generate a total of three acquisition decisions: a first decision corresponding to the text modality, a second decision corresponding to the image modality, and a third decision corresponding to the numerical modality. To generate these acquisition decisions for a given time step, the selected neural network processes a network input that includes observations obtained for one or more previous time steps to generate an output π that parameterizes a distribution from which the acquisition decisions are sampled.

[0123] For example, at the first time step, the first decision defines that text data should be selected for acquisition at that time step, the second decision defines that image data should be selected for acquisition at that time step, and the third decision defines that numerical data should not be selected for acquisition at that time step. As represented by the [masked] token, some input data may contain missing entries.

[0124] After generating predictions based on the modalities selected for acquisition by the selection neural network at each time step, the system determines the acquisition cost based on the corresponding modalities selected for acquisition at each time step in the sequence of time steps (step 302). In some implementations, each modality in the set of modalities is associated with a corresponding cost factor. In these implementations, the system may perform sub-steps 402-404 (as explained in more detail with reference to Figure 4 to determine the acquisition cost.

[0125] Figure 4 is Figure 3 a flowchart of sub-steps 402-404 of step 302 of the process of

[0126] In implementations where each modality in the set of modalities is associated with a corresponding cost factor, the system can determine the corresponding acquisition cost for each time step in the sequence of time steps based on the corresponding cost factors associated with each modality selected for acquisition at that time step (step 402). For example, the acquisition cost for a time step can be calculated as the weighted or unweighted sum of the cost factors associated with each modality selected for acquisition at that time step.

[0127] The system determines the acquisition cost as a combination of the acquisition costs for the time steps in the sequence (step 404). For example, the acquisition cost can be calculated as the weighted or unweighted sum of the corresponding acquisition costs for the time steps in the sequence.

[0128] For each of one or more modalities in the modality, the cost factor for that modality is at least partially based on the amount of resource usage required to capture data corresponding to that modality. For example, the resource usage required to capture data corresponding to the modality at least characterizes the energy usage required to capture data corresponding to that modality. As another example, the resource usage required to capture data corresponding to the modality at least characterizes the amount of time required to capture data corresponding to the modality.

[0129] Additionally or alternatively, for each of one or more modalities in the modality, the cost factor for that modality is at least partially based on the risk associated with capturing data corresponding to that modality. For example, when the environment is a medical environment including a patient, the risk associated with capturing data corresponding to the modality is at least partially based on the medical risk to the patient caused by capturing data corresponding to that modality.

[0130] The system determines a reward (step 304) at least partially based on the acquisition cost of the selected modality. The acquisition cost can be included in the reward in any suitable manner, and the reward is typically a numerical value.

[0131] For example, the system can determine the reward at least partially based on a comparison of the acquisition cost with a threshold called the "cost budget". In some implementations, if the acquisition cost exceeds the cost budget, the system can reduce the reward by a predefined or adaptive amount. The cost budget can indicate, for example, an acceptable level of acquisition cost, such as an acceptable amount of energy usage, or an acceptable amount of medical risk (e.g., based on the amount of radiation exposure tolerable by the patient), or an acceptable amount of computational resources (e.g., memory and computing power) for processing data from the acquired modality.

[0132] In some implementations, the reward depends on both (i) the acquisition cost and (ii) the prediction error, which measures the error in the prediction generated by the prediction model. In these implementations, as Figure 6 shown, the reward can be calculated, for example, as an expected value:

[0133]

[0134] Here, the expectation is taken over the training inputs (x, y), where x is a sequence of observations and y is the true value prediction, a represents the acquisition decision made by the selected neural network; C(a) represents the total acquisition cost of the sequence of observations; C m is the modality-specific cost factor, and is the log-likelihood loss of the prediction generated by the prediction model relative to the true value prediction (but of course other loss functions can be used, i.e., loss functions that compare the prediction generated by the prediction model with the true value prediction).

[0135] Optionally, in some implementations, the system adds the intermediate prediction error to the reward, such as the reward calculated using Equation (1). The intermediate prediction error, when used, encourages the selection neural network to reduce the prediction error. In these implementations, the system may perform sub-steps 502-506 (as explained in more detail with reference to Figure 5 to determine the reward.

[0136] Figure 5 Yes Figure 3 is a flowchart of sub-steps 502-506 of step 304 of the process of

[0137] For each time step in one or more time steps in the sequence of time steps, the system uses a prediction model to process the model input to generate an intermediate prediction characterizing the environment, where the model input includes the observation for that time step and the observations for one or more previous time steps in the sequence of time steps (step 502).

[0138] For each time step in one or more time steps in the sequence of time steps, the system determines an intermediate prediction error that measures the error in the intermediate prediction generated by the prediction model (step 504).

[0139] The system determines the reward at least in part based on the intermediate prediction errors that have been determined for one or more time steps (step 506). For example, the system may add the intermediate prediction error to the reward calculated using Equation (1). The intermediate prediction error can be calculated, for example, as:

[0140]

[0141] where α is a hyperparameter (e.g., a predefined constant value), γ is a discount factor, x is the sequence of observations, and y is the true value prediction, and is the log-likelihood loss of the prediction generated by the prediction model relative to the true value prediction.

[0142] The system uses reinforcement learning techniques to train the selection neural network based on the reward to adjust the values of the parameters of the selection neural network (step 304). In particular, the system trains the selection neural network to generate acquisition decisions that maximize the reward determined at least in part based on the acquisition cost. For example, the reinforcement learning technique can be a policy gradient technique, such as the advantage actor critic (A2C) policy gradient technique, which applies Gumbel parameterization to the (discrete) acquisition decisions.

[0143] In some implementations, the system also trains the prediction model based on a reward (e.g., the reward calculated using Equation (1), which depends on both the acquisition cost and the prediction error) to adjust the values of the parameters of the prediction model simultaneously.

[0144] For example, the system can train the prediction model and the selection neural network together to jointly update the parameter values of both the selection neural network and the prediction model, e.g., to allow the prediction model to be particularly suitable for combinations of modalities frequently selected by the selection neural network. In this example, since the parameter values of the prediction model are updated, the reward received by the selection neural network can thus change.

[0145] Alternatively, in other implementations, the system can train the prediction model independently of the training of the selection neural network, e.g., based on optimizing an objective function that depends on the prediction error of the prediction machine learning model (during which the parameter values of the prediction model remain fixed).

[0146] For example, the system can pre-train the prediction model to process a sequence of masked observations to generate corresponding predictions. The system then trains the selection neural network to update the parameter values of the selection neural network while keeping the pre-trained parameter values of the prediction model fixed.

[0147] This specification uses the term "configured" in connection with systems and computer program components. For a system of one or more computers to be configured to perform particular operations or actions means that the system has installed thereon software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by a data processing device, cause the device to perform the operations or actions.

[0148] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware (including the structures disclosed in this specification and their structural equivalents), or in a combination of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, a data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or additionally, the program instructions can be encoded on an artificially generated propagated signal (e.g., a machine-generated electrical, optical, or electromagnetic signal) that is generated to encode information for transmission to a suitable receiver device for execution by the data processing apparatus.

[0149] The term "data processing device" refers to data processing hardware and includes all kinds of devices, apparatuses and machines for processing data, such as programmable processors, computers or multiple processors or computers. The device may also be or further include dedicated logic circuitry, such as FPGAs (Field Programmable Gate Arrays) or ASICs (Application Specific Integrated Circuits). In addition to the hardware, the device may optionally include code that creates an execution environment for a computer program, for example, code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.

[0150] A computer program—which may also be referred to or described as a program, software, software application, app, module, software module, script or code—can be written in any form of programming language (including compiled or interpreted languages or declarative or procedural languages); and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine or other unit suitable for use in a computing environment. A program may or may not correspond to a file in a file system. A program may be stored in a part of a file that holds other programs or data (such as one or more scripts in a markup language document), in a single file dedicated to the program being discussed, or in multiple coordinated files (such as files that hold one or more modules, subroutines or portions of code). A computer program may be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a data communication network.

[0151] In this specification, the term "engine" is used broadly to refer to a software-based system, subsystem, or process programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines may be installed and run on the same one or more computers.

[0152] The processes and logical flows described in this specification may be performed by one or more programmable computers that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows may also be performed by, for example, FPGA or ASIC dedicated logic circuitry, or by a combination of dedicated logic circuitry and one or more programmed computers.

[0153] A computer suitable for executing a computer program can be based on a general or special purpose microprocessor or both, or any other kind of central processing unit. Generally, the central processing unit will receive instructions and data from a read only memory or a random access memory or both. The basic elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing the instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or will be operatively coupled to receive data from or transfer data to the one or more mass storage devices, or both. However, a computer need not have such devices. In addition, a computer may be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name but a few.

[0154] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices (such as EPROM, EEPROM, and flash memory devices), magnetic disks (such as internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks.

[0155] To provide for interaction with a user, embodiments of the subject matter described in this specification may be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide for interaction with the user; for example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input received from the user may be in any form, including sound, speech, or tactile input. In addition, a computer may interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending a web page in response to a request received from a web browser on the user's device. Further, a computer may interact with the user by sending text messages or other forms of messages to a personal device (e.g., a smart phone running a messaging application) and receiving responsive messages from the user in response.

[0156] The data processing device for implementing a machine learning model may also include, for example, a dedicated hardware accelerator unit for processing the general and computationally intensive parts of machine learning training or production (i.e., inference, workload).

[0157] A machine learning framework (e.g., the TensorFlow framework or the JAX framework) can be used to implement and deploy a machine learning model.

[0158] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes backend components (e.g., as a data server), or includes middleware components (e.g., an application server), or includes frontend components (e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification), or any combination of one or more such backend components, middleware components, or frontend components. The components of the system can be interconnected by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.

[0159] The computing system can include clients and servers. The clients and servers are typically located far apart from each other and typically interact through a communication network. The relationship between the client and the server is created by computer programs that run on the respective computers and have a client - server relationship with each other. In some embodiments, the server transmits data (e.g., an HTML page) to a user device, for example, for the purpose of displaying data to a user interacting with the device acting as a client and receiving user input from it. Data generated at the user device, such as the result of a user interaction, can be received at the server from the device.

[0160] Although this specification contains many specific implementation details, these details should not be construed as limitations on the scope of any invention or the scope that may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. Certain features that are described in the context of separate embodiments in this specification can also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub - combination in multiple embodiments. Additionally, although features may be described as acting in certain combinations and even initially claimed as such, in some cases one or more features from a claimed combination can be deleted from the combination, and the claimed combination can relate to a sub - combination or a variant of a sub - combination.

[0161] Similarly, although operations are depicted in the drawings and recited in the claims in a particular order, this should not be construed as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Additionally, the separation of various system modules and components in the above-described embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the program components and systems described may generally be integrated together in a single software product or packaged into multiple software products.

[0162] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. For example, the acts recited in the claims may be performed in a different order and still achieve the desired result. As one example, the processes depicted in the figures do not necessarily require the particular order shown or sequential order to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous.

Claims

1. A method performed by one or more computers, the method comprising: obtaining, for each time step in a sequence of multiple time steps, a respective observation characterizing a state of an environment, including for each time step after a first time step in the sequence of time steps: using a selection neural network to process a network input including observations obtained for one or more previous time steps to generate a plurality of acquisition decisions, wherein each acquisition decision corresponds to a respective modality from a set of multiple modalities and defines whether data corresponding to the modality is selected for acquisition at that time step; obtaining the observation for that time step, wherein the observation: (i) includes data corresponding to modalities from the set of modalities that are selected for acquisition at that time step, and (ii) does not include data corresponding to modalities from the set of modalities that are not selected for acquisition at that time step; and using a prediction model to process a model input to generate a prediction characterizing the environment, the model input including the observations for each time step in the sequence of time steps.

2. The method according to claim 1, further comprising: determining an acquisition cost based on the respective modalities selected for acquisition at each time step in the sequence of time steps; determining a reward at least in part based on the acquisition cost; and using reinforcement learning techniques to train the selection neural network based on the reward.

3. The method according to claim 2, wherein each modality in the set of modalities is associated with a respective cost factor, and wherein determining the acquisition cost includes: for each time step in the sequence of time steps, determining a respective acquisition cost for that time step based on the respective cost factors associated with each modality selected for acquisition at that time step; and determining the acquisition cost as a combination of the acquisition costs for the time steps.

4. The method according to claim 3, wherein for each time step in the sequence of time steps, determining the acquisition cost for that time step includes: determining the acquisition cost for that time step as the sum of the cost factors associated with each modality selected for acquisition at that time step.

5. The method according to any one of claims 3 to 4, wherein determining the acquisition cost as a combination of the acquisition costs for the time steps includes: determining the acquisition cost as the sum of the acquisition costs for the time steps.

6. The method according to any one of claims 3 to 5, wherein for one or more of the modalities, the cost factor for the modality is at least in part based on the amount of resource usage required to capture data corresponding to the modality.

7. The method according to claim 6, wherein the resource usage required to capture data corresponding to the modality at least characterizes the energy usage required to capture data corresponding to the modality.

8. The method according to any one of claims 6 to 7, wherein The resources required to capture data corresponding to the modality use an amount that at least characterizes the time required to capture data corresponding to the modality.

9. The method according to any one of claims 3 to 8, wherein, for one or more of the modalities, the cost factor for the modality is at least partially based on the risk associated with capturing data corresponding to the modality.

10. The method according to claim 9, wherein, the environment includes a patient, and the risk associated with capturing data corresponding to the modality is at least partially based on the medical risk to the patient caused by capturing data corresponding to the modality.

11. The method according to any one of claims 2 to 10, further comprising: determining a prediction error that measures the error in the prediction generated by the prediction model; and determining the reward based on both (i) the acquisition cost and (ii) the prediction error.

12. The method according to any one of the preceding claims, wherein, the prediction model is a machine learning model.

13. The method according to claim 12, wherein, the prediction model includes a neural network.

14. The method according to any one of claims 12 to 13, further comprising: training the prediction machine learning model to optimize an objective function that depends on the prediction error of the prediction machine learning model.

15. The method according to any one of the preceding claims, wherein, for each time step in the sequence of time steps, the network input of the selection neural network at that time step further includes: identifying the data of the acquisition decision for any modality at any previous time step.

16. The method according to any one of claims 2 to 15, further comprising: for each of one or more time steps in the sequence of time steps: using the prediction model to process the model input to generate an intermediate prediction characterizing the environment, the model input including the observation for the time step and the observations for one or more previous time steps in the sequence of time steps; and determining an intermediate prediction error that measures the error in the intermediate prediction generated by the prediction model; and determining the reward at least partially based on the intermediate prediction error.

17. The method according to any one of the preceding claims, wherein, the set of modalities includes imaging modalities, and wherein the data corresponding to the imaging modalities includes image data.

18. The method according to claim 17, wherein, the set of modalities includes medical imaging modalities.

19. The method according to any one of the preceding claims, wherein, the environment is a medical environment including a patient.

20. The method according to claim 19, wherein, the prediction characterizing the environment includes the predicted medical diagnosis of the patient.

21. The method according to any one of claims 19 to 20, wherein, the prediction characterizing the environment includes a prediction of the medical treatment to be applied to the patient.

22. The method according to any one of the preceding claims, further comprising, for each time step after the first time step in the sequence of time steps: determining: (i) that data corresponding to the modality selected for acquisition at that time step from the set of modalities will be included in the observation for that time step, and (ii) that data corresponding to the modalities from the set of modalities not selected for acquisition at that time step will not be included in the observation for that time step.

23. The method according to claim 22, wherein, for each of one or more time steps after the first time step in the sequence of time steps: only a proper subset of the modalities in the set of modalities is selected for acquisition at that time step.

24. The method according to any one of claims 22 to 23, further comprising, for each time step after the first time step in the sequence of time steps: acquiring data only for the modalities selected for acquisition at that time step.

25. A system, comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of the corresponding method according to any one of claims 1 to 24.

26. One or more non-transitory computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the corresponding method according to any one of claims 1 to 24.