Ai report generation from medical images, and ai report generation from medical images with an expert in the loop

By fine-tuning a pre-trained visual language model with a cost function based on a medical image and text report database, the neural network system effectively generates preferred medical reports, addressing the challenges of current technologies in this domain.

WO2025114445A1PCT designated stage expired Publication Date: 2025-06-05DEEPMIND TECH LTD

Patent Information

Application Number
PCT/EP2024/083931
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-28
Filing Date
2024-11-28
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Current technologies face challenges in generating accurate and efficient medical reports from medical images, often requiring extensive human intervention and resources.

Method used

A neural network system is trained to generate textual reports from medical images by fine-tuning a pre-trained visual language model, using a cost function that includes a prediction cost term based on a training database of medical images and associated text reports.

Benefits of technology

The system is capable of producing medical reports that are preferred by expert evaluators over human-written reports in a majority of cases, and when used in an assistive scenario with clinician revision, results in revised reports that are also preferred.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024083931_05062025_PF_FP_ABST
    Figure EP2024083931_05062025_PF_FP_ABST
Patent Text Reader

Abstract

A neural network system is trained to generate textual reports (that is, medical reports, such as radiology reports) from one or more medical images, by fine-tuning a pre-trained neural network system (a visual language model, "VLM") operative, upon receiving an input comprising at least one image and a textual input, to generate a value indicative of a predicted likelihood of one or more candidate text continuations of the textual input. The fine-tuning of the neural network system is performed to reduce the value of a cost function which includes a first prediction cost term based on a first training database including first training datasets of at least one medical image and an associated text report, the first training datasets corresponding to first individuals. The first prediction cost term further includes a cost value for each individual, inversely dependent on a likelihood value of the associated textual report, conditioned on the at least one medical image, and created by the neural network system.
Need to check novelty before this filing date? Find Prior Art

Description

Al REPORT GENERATION FROM MEDICAL IMAGES, AND Al REPORTGENERATION FROM MEDICAL IMAGES WITH AN EXPERT IN THE LOOPCROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims priority to GR Patent Application No. 20230100982, filed on November 28, 2023. The disclosure of the prior application is considered part of and is incorporated by reference in the disclosure of this application.BACKGROUND

[0002] This specification relates to processing data using machine learning models.

[0003] Machine learning models receive an input and generate an output, e.g., a predicted output, based on the received input. Some machine learning models are parametric models and generate the output based on the received input and on values of the parameters of the model.

[0004] Some machine learning models are deep models that employ multiple layers of models to generate an output for a received input. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers that each apply a non-linear transformation to a received input to generate an output.SUMMARY

[0005] A first aspect of the present disclosure proposes, in general terms, that a neural network system is trained to generate textual reports (that is, medical reports, such as radiology reports) from one or more medical images, by fine-tuning a pre-trained neural network system (a visual language model, “VLM”).

[0006] The VLM may be one operative, upon receiving an input comprising at least one image and a textual input, to generate a value indicative of a predicted likelihood of one or more candidate text continuations of the textual input. The fine-tuning of the neural network system is performed to reduce the value of a cost function, which includes a first prediction cost term based on a first training database including first training datasets of at least one medical image and an associated text report, the first training datasets corresponding to first individuals. The first prediction cost term further includes a cost value for each individual, inversely dependent on a likelihood value of the associated textual report, conditioned on the at least one medical image, and created by the neural network system.

[0007] Particular embodiments of the subject matter described in the specification can be implemented to realise the advantage that the fine-tuned neural network is capable ofgenerating a text report from at least one medical image which in a large proportion (e.g. a majority) of cases is preferred by expert evaluators to a human-written report.

[0008] Furthermore, an assistive scenario, in which a fine-tuned neural network generates the first draft of a report, which is subsequently revised by a clinician, results in revised text report which in a greater proportion of cases is preferred by expert evaluators to human-written reports.

[0009] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Examples of the present disclosure are now explained with reference to the following figures.

[0011] Fig. la shows an encoder network.

[0012] Fig. lb shows a token processing model.

[0013] Fig. 2 shows a neural network system comprising modality networks including the encoder network shown in Fig. la, and a data-item-token processing model including the token processing model shown in Fig. lb.

[0014] Fig. 3 shows a first modality network.

[0015] Fig. 4 shows a second modality network.

[0016] Fig. 5 shows a pair of layers of the data-item-token processing model of the neural network system of Fig. 2.

[0017] Fig. 6 shows a flow chart showing an example process for training a neural network system.

[0018] Fig. 7 shows a schematic of a computer system that may be used for training the neural network system.

[0019] Fig. 8 is a flow chart showing an example process for generating an auxiliary classification loss term.

[0020] Fig. 9 shows a representation of an example neural network system.

[0021] Fig. 10 shows a flow chart showing an example process for generating a medical report.

[0022] Fig. 11, which is composed of Fig. 11(a) and 11(b), shows experimental results characterising the output of the neural network when performing a method according to an example set out herein.

[0023] Fig. 12 shows a flow chart showing an example process for generating and modifying a medical report.

[0024] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION

[0025] Several visual language neural network models, or more simply visual language models (VLMs), are known. Some comprise a neural network system comprising input units which collectively receive a network input comprising image(s) and a text input, e.g. one of more tokens selected from a set of tokens collectively forming a vocabulary. The vocabulary may comprise the letters of a natural language vocabulary, e.g. the English alphabet, other characters such as punctuation and / or control characters, and / or multi-character components of natural language words (“word pieces”)). The neural network system may be configured to generate a value indicative of a respective predicted likelihood for multiple candidate text continuations of the text input (e.g. one or more tokens selected from the vocabulary), and may comprise an output unit configured to generate a network output based on the corresponding predicted likelihoods. For example, the output unit may select the token from the vocabulary with the highest predicted likelihood to be the network output of the VLM; or the output unit may define a probability distribution over the candidate text continuations based on the predicted likelihoods, and select a candidate text continuation from the probability distribution.

[0026] Some VLMs are “autoregressive”. That is, successive network inputs can be formed and successively input to the VLM, one or more or each network input comprising at least one image and (except in some cases for the first network input) a textual input. In the case of the first network input, the textual input is optional and is an “input text”. Each network input except the first can comprise, in addition to the optional input text, the network output(s) generated by the VLM from the preceding network input(s). In this way, an appropriate textual response is generated from the image and the optional input text, as the successive network outputs of the VLM.

[0027] In general terms, the present disclosure proposes that a pre-existing trained neural network system (VLM) is fine-tuned for processing medical images. The fine-tuning isconducted by iteratively refining numerical parameters of the neural network system based on a training database (a “first database”) comprising a plurality of training datasets. Each training dataset comprises at least one medical image of a corresponding “first” individual (e.g. human subjects, although in principle the individuals could be non-human animals), and an associated textual report. A medical image refers to a multi-dimensional array of pixels, with one or more intensity values associated with each pixel, where the intensity values have been captured by imaging a human (or possibly non-human animal) subject; for example, the interior of the subject (various known imaging modalities are listed below). The textual report may be one written by a human expert, based on the at least one medical image, and describing medical characteristics of the first individual which can be determined from the medical image(s). The training is based on a cost function including a prediction cost term based on the training datasets. The prediction cost term comprises a cost value for each first individual which depends inversely on (i.e. is higher for correspondingly lower values of) a likelihood generated by the neural network system, of the associated text report, conditioned on the corresponding medical image(s).

[0028] As described below, the neural network system may operate probabilistically (e.g. each output token is chosen from a respective probability distribution generated by the neural network system, as described above), and in this case the prediction cost value may be a smooth function of corresponding likelihoods of the tokens of the associated text report under the probability distributions. Alternatively, output tokens generated by the neural network system may be generated deterministically, e.g. as the output token for which the corresponding value of a distribution of values over the elements of the vocabulary is highest; particularly in this case, the prediction cost term may in principle have a default value if the neural network system does not generate the associated medical report from the corresponding images; alternatively, it may have a value based on the distribution for each output token, e.g. indicating how close the associated medical report was to being generated by the neural network system.

[0029] Because of the functionality of the pre-existing trained neural network system, e.g. its understanding of written language, and its ability to extract information sensibly from input image(s) and textual inputs, to generate an appropriate response, the amount of computing resources and medical training data which are needed for the fine-tuning are far less than if the neural network system were trained from, for example, an initially random state. The preexisting trained neural network system may have been trained based on a different database (a “second” training database) in which the training examples are mainly not training examples comprising medical images. For example, less than 20% of the training examples in the secondtraining database may comprise medical images, or less than 10%, or even substantially none of the training examples.

[0030] In other words, the present technique can benefit from a pre-training based mainly on training examples which are not specifically medical images, e.g. in order to leam the basic process of extracting information from the (non-medical) images of the training examples in the second training database. Very large databases suitable for fulfilling the role of the second training database are widely available, containing training examples drawn from a wide range of technical and non-technical literature. The pre-trained neural network may then be fine-tuned using a first training database which is much smaller than the second database (e.g. the number of first training datasets in the first database is much less than the number of training examples in the second training database), to enhance its ability to produce valuable textual reports (that is, medical reports) from medical image(s). Medical databases are expensive to obtain and curate (e.g. due to privacy concerns) so the present techniques may make possible the generation of superior textual reports from those medical databases which exist.

[0031] The medical images of the first training datasets of the first database (which may also be termed radiology images) are typically images of the interior of the corresponding first individual. Furthermore, the medical images of the first training datasets of the first database may be frontal view images (e.g., anterior-posterior or posterior-anterior). In some examples, the images may be only a subset of frontal view images, for example, the first training dataset may include only anterior-posterior frontal view images, or alternatively may include only posterior-anterior frontal view images.

[0032] The image(s) of each first training dataset may optionally have been captured using the same imaging modality. For example, each training dataset may comprise a single training image which is an X-ray image, e.g. relating to the same portion of the first individual’s body. Alternatively, each training dataset may comprise multiple training images, e.g. a CAT scan of the first individual comprising multiple images of different corresponding portions of the body of the first individual, and / or images of the same portion of the body of the first individual captured at different respective times (e.g. different times during a breathing cycle or cardiac cycle, and / or different times which are one or more days / weeks / months / years apart and show the progression of a medical condition during that period).

[0033] In a further possibility, some or all of the first training datasets may comprise multiple images captured from multiple corresponding imaging modalities, e.g. at least one image from each of at least two of the group consisting of an X-ray image, an magneticresonance imaging (MRI) image, a fluoroscopy image, an ultra-sound (ultrasonography) image, or a nuclear medicine image (e.g. an image obtained by positron emission tomography, PET). For example, one or more of the first training datasets may comprise one or more images captured by X-ray radiography and one or more images captured by ultrasound imaging.

[0034] Following the fine-tuning, when the trained (fine-tuned) neural network system is in use, it receives a network input comprising at least one medical image of an individual, and, optionally, a textual input. Again, the medical image(s) of the network input may include multiple medical images captured at different times and / or medical images of different respective portions of the body of the individual, and / or medical images captured with different imaging modalities.

[0035] From the network input, the trained (fine-tuned) neural network system may generate a predicted likelihood for each of a plurality of candidate text continuations of the textual input (or, if there is no textual input in the network model, an output based only on the at least one medical image). One of the candidate text continuations (e.g. a token from the vocabulary) may be selected, by an output unit of the trained neural network system based on the corresponding predicted likelihoods, as the network output of the neural network system. For example, the candidate text continuation for which the predicted likelihood is highest may be selected. Alternatively, a probability distribution over the candidate text continuations may be defined based on the predicted likelihoods, and a candidate text continuation may be selected from the probability distribution.

[0036] This process may be auto-regressive. That is, a new network input may then be constructed based on the previous network input and the selected candidate text continuation. By repeating this process (e.g. until a termination criterion is met, such as the VLM generating a text token of a certain type), a plurality of network outputs may be successively generated, and a medical report is generated based these network outputs, e.g. as a concatenation of the network outputs.

[0037] During the fine tuning of the neural network system, each likelihood value for a textual report of a training example may be formed by dividing the textual report into a sequence of portions, and calculating the likelihood value for the textual report given a training example by combining (e.g. adding) corresponding portion likelihood values for each portion which are conditioned on the at least one corresponding medical image and, for each portion except the first, on the earlier portions of the sequence. For example, denoting a given training example as (x,y) where x is a medical image (or multiple images) and y is a corresponding textual report, the portions may be denoted by y={yi} for 1=1, ...L, where / is an integer indexand Z is the number of portions into which y is divided. The portion likelihood value for portion yz (where / ’ is an integer in the range 1 to Z) is predicted by the VLM as a function of x w y<i (where y<i denotes {y } for l’=l, ...l- ). In one case, the portion likelihood value generated by the neural network system for yr may be logarithmic, i.e. log (p(yi \ y^, x)). Thus, the likelihood value generated by the neural network system, for textual report y conditioned on the medical image(s) x, may be given bylog (p(yi I yx- The corresponding cost value may simply be -- Xf=i log (p(yi I yx), although other cost values may be used instead.

[0038] In the prediction cost term, the cost values may be weighted based on whether the corresponding first individual exhibits a medical characteristic. The medical characteristic may be one which a proper subset of the first individuals exhibit. It may be simply that the first individuals are “abnormal”, e.g. the characteristic may be that a given first individual exhibits any one or more of a plurality of different (and logically -independent) ways of being abnormal (e.g. an enlargement of different respective organs of the body).

[0039] The first training database may contain an index value associated with each training dataset, indicating whether the corresponding first individual exhibits the medical characteristic (for example, that the individual exhibits one of more medical conditions), so that the weight values can be determined from the corresponding index values. In one example, the index value may be binary, for example the index value may be equal to 0 if the medical characteristic is not present in the first individual, and the index value may be equal to 1 if the medical characteristic is present in the first individual. Due to the weights, the cost function can be tailored based on the medical characteristic. In particular, in the training the neural network system can be penalised for composing inaccurate reports both for the first individuals exhibiting the medical characteristic (e.g. “abnormal” first individuals) and for the first individuals who do not (e.g. “normal” first individuals), even if one of these is a low proportion of the first individuals.

[0040] For example, the weight values for each of the first individuals exhibiting the medical characteristic may be lower if a higher proportion of the first individuals exhibit the medical characteristic, and conversely higher if a lower proportion of the first individuals exhibit the medical characteristic. Similarly, the weight value for first individuals not exhibiting the medical characteristic may be higher if a high proportion of the first individuals exhibit the medical characteristic, and conversely lower if a lower proportion of the first individuals exhibit the medical characteristic. In some examples, the weight valuescorresponding to the population of first individuals not exhibiting the medical characteristic may be inversely proportional to the proportion of the first individuals exhibiting the medical characteristic. Conversely, the weight values corresponding to the population of first individuals exhibiting the medical characteristic may be inversely proportional to the proportion of first individuals not exhibiting the medical characteristic. For example, if the proportion of first individuals exhibiting a medical characteristic is 10% of the total population of first individuals, the weight value for the first individuals exhibiting the medical characteristic may be ten times greater than the weight value applied to the first individuals not exhibiting the medical characteristic.

[0041] Thus, denoting the weight value for a first training dataset (x,y) as w(x,y), the prediction cost term may have the general form:where the expectation value is calculated over the first training database D . As discussed below in more detail, training a visual language model a cost function including a prediction cost term which is a weighted sum over the training datasets of a cost value, where the weight values depend upon whether the corresponding individuals exhibit a certain medical characteristic, constitutes a second independent aspect of the present disclosure.

[0042] A neural network system 200 which is suitable for use with either aspect of the invention is explained in detail below with reference to Fig. 2. The neural network system 200 may include one or more encoder networks for receiving images (“image encoders”) comprised in the network input, and generating encoding (embedding) from the images (that is, from pixel data of the image(s) defining one or more weights for each pixel of a multi-dimensional array of pixels). The encoding is representative of salient features in the pixel data. An encoder network may be one which has been pre-trained within a different neural network system. For example, Fig. la illustrates a neural network system including an encoder network 11 and an optional output layer 12. The encoder network 11 receives a data item comprising at least one image and generates an output from it called an encoded data item. The optional output layer 12 (if present) receives the output of the encoder network 11 and generates an output of the neural network system. The encoder network 11 may be trained using a database of data items and corresponding desired outputs, such that, upon a data item being input to the encoder network 11, the encoder network 11 and (if present) output layer 12 generate the corresponding desired output. During the training process, the output layer 12 (if present) may optionally be trained also. In the case that an output layer 12 is present, the encoder network may be trainedto extract features of a data item it receives, and the output layer 12 processes those features to generate the desired output (encoded data item). Following the training the output layer 12 (if any) may be discarded. The encoded data item may have a smaller number of components than the data item (e.g. image) from which it is generated. The encoder network 11 may comprise one or more stacked convolutional layers, and the encoded data item may be a (e.g. two- dimensional) feature map.

[0043] In the experiments reported below, the encoder network 11 employed was the “f6” model taken from the paper “High-Performance Large-Scale Image Recognition Without Normalization”, by A. Brock et al 2021. The encoder network 11 was pre-trained using a contrastive objective on datasets of image and text pairs, using the two-term contrastive loss from “Learning Transferable Visual Models From Natural Language Supervision”, by A. Radford et al, 2021. The output of the final stage, a 2D spatial grid of features, was “flattened” to form a ID sequence.

[0044] The neural network 200 system further includes a data-item-token processing model comprising token processing layers taken from a token processing model. The token processing model may be one which has been pre-trained. For example, Fig. lb illustrates a suitable token processing model comprising an integer number n of token processing layers 131, 132, ... 13n, arranged in a stack (sequence). The first token processing layer receives an input token string. Each token processing layer, except the first, receives the output of the preceding token processing layer.

[0045] The output of the n-th token processing layer, 13n, is one or more output tokens for the network input, e.g. one output token for each network input. Successive output tokens generated by the n-th token processing layer 13 for successive corresponding network inputs, form an output token string, y. As discussed above, the network inputs may be generated auto- regressively, including previously generated output tokens. The n-th token processing layer may generate, for each network input, a distribution over the vocabulary (i.e. respective values for each element of the vocabulary), and include an output unit (not shown) which selects a token based on the distribution. For example, the output unit may treat the distribution as a probability distribution, and select the output token according to the probability distribution (i.e. with each element of the vocabulary being selected with a probability proportional to the corresponding value of the distribution); or the output unit may select the output token as the element of the vocabulary for which the corresponding value of the distribution is highest.

[0046] The token processing layers may constitute a “language model” (e.g. a “large language model”), trained on a large database of data, e.g. natural language data, such that uponan input token string (sequence of tokens from the vocabulary) being input to the first token processing layer, the output token string is an appropriate response. For example, if the input token string is a question, the output token string may be an appropriate answer. In the experiments reported below, the token processing model was the Chinchilla model of “Training compute-optimal large language models”, J. Hoffmann et al, 2022. Alternatively, the token processing layers may be layers generated by another publicly-known language model training method (e.g. Brown et al, “Language models are few-shot learners”, in Conference on Neural Information Processing Systems, 2020).

[0047] Another visual language model which can be used as the pre-trained neural network system which is the subject of the fine-tuning operation is the Flamingo model described in “Flamingo: a Visual Learning Model for Few-Shot Learning”, J-B Alayrac et al (2022), https: / / arxiv.org / pdf / 2204.14198.pdf. This comprises at least one image encoder neural network, a trained language model comprising a sequence of token-processing layers, and a plurality of modification layers.

[0048] The modification layers are interleaved with the sequence of token processing layers of the trained language model, so as to form an interleaved sequence of layers (e.g. with each modification layer being placed between a corresponding pair of the token processing layers). The modification layers modify data flowing through the training language model, and the modification layers are controlled based on (e.g. different respective portions of) an embedding (“image encoding vector”) generated by the (or each) image encoder neural network based on at least one image the image encoder neural network receives. To put this another way, the VLM includes an interleaved sequence (stack of layers) including a number n of token processing layers interleaved by a number j of modification layers in the form of gated cross-attention layers. The number j of modification layers may be equal to the number n of token processing layers, but in variations n and j may be different from each other, e.g. such that there may be any number of token processing layers between any given pair of modification layers and vice versa.

[0049] If the medical image(s) of the training examples are all obtained using the same imaging modality, there may be one image encoder neural network. Alternatively, if at least one of the training examples contains medical images captured using multiple corresponding imaging modalities (e.g. a given training example includes one or more X-ray images and an MRI image), the VLM may comprise a different image encoder neural network for each imaging modality. For simplicity, the case that there is a single imaging modality and a single image encoder neural network is mainly assumed below.

[0050] During the training, or in subsequent use, the image encoder neural network receives (e.g. successively) the medical image(s) of a current network input, and generates from it / them an image encoding vector. The interleaved sequence of layers receives the textual part of the current network input. The image encoding vector (or corresponding portions of it) are fed to the modification layers. The modification layers modify the processing of any textual part of the current network input by the trained language model.

[0051] The (or each) image encoder neural network may comprise a pre-trained image encoder (a neural network), and a compressed representation generation system in the form of a respective “resampler” unit. As in J-B Alayrac et al (2022), the image encoder may be a pretrained and frozen Normalizer-Free Residual Network (NFNet) (A. Brock, et al, “High- performance large-scale image recognition without normalization”, arXiv: 2102.06171, 2021). The image encoder(s) may be trained using a contrastive objective of datasets of image and text pairs, using a two-term contrastive loss (Radford et al., “Learning Transferable visual models from natural language supervision”, arXiv:2103.00020. 2021). In the case of still images, the image encoder may generate a 2D spatial grid of features, and flatten it to a ID sequence. For video images, the image encoder may sample the frames and encode them independently to obtain a 3D spatio-temporal grid of features to which learned temporal embeddings are added. The features may then be flattened to ID.

[0052] As in J-B Alayrac et al (2022), the resampler(s) may be based on a “perceiver” as described in “Perceiver: General Perception with iterative attention”, A. Jaegle et al, 2021. The resampler(s) take as input a variable number of image or video features (spatial features or spatio-temporal features in the case that the input image(s) are video images) produced by the corresponding image encoder from input images (e.g. having different respective formats), and produce an image encoding vector as a compressed representation of the corresponding input image and having a fixed number of components (e.g. visual outputs).

[0053] The token processing layers may have been trained together (without the gated cross-attention layers) as a stack of layers which operates as a trained language model. That is, the token processing layers were trained as a stack, in the same order, such that when a first of the token processing layers receives a (text) input composed of text tokens from a vocabulary, each other token processing layer receives the output of the preceding token processing layer, and the last token processing layer generates an output composed of text tokens from a vocabulary (normally, the same vocabulary, but optionally the input could be text tokens from a first vocabulary, and the output could be text tokens from a different, second vocabulary). The token processing layers were trained, e.g. by any known large language model trainingalgorithms such as using a maximum-likelihood objective, on a large database of text, e.g. natural language text that is publically available from the Internet or another text corpus, such that upon an input token string (sequence of tokens from the vocabulary) being input to the first token processing layer, the output token string generated by the last token processing layer is an appropriate response. For example, if the input token string is a question, the output token string may be an appropriate answer. The trained language model may, for example, be the Chinchilla model of “Training compute-optimal large language models”, J. Hoffmann et al, 2022, or the trained language model of Brown et al, “Language models are few-shot learners”, in Conference on Neural Information Processing Systems, 2020.

[0054] Each modification layer (e.g. gated cross-attention layer) receives an image encoding vector (or a corresponding portion thereof) from the image encoder neural network. It further receives a vector (“layer output”) from the preceding layer of the VLM (i.e. the preceding layer of the interleaved sequence of layers), as a language input. If the gated crossattention layer is the first layer of the stack, the language input is typically the input text tokens, or a subset of them. Otherwise the language input to the modification layer is the output from the preceding layer of the VLM.

[0055] Each modification layer may apply an attention function over the inputs to the modification layer, such as a query -key -value (QKV) attention operation. For example, the image encoding vector(s) may be used to define key and value vectors, and a query vector is defined based on the language input. In general, an attention operation can be one that applies an attention mechanism to elements of an embedding to update each element of the embedding, e.g. where input embeddings are used to determine a query vector and a set of key -value vector pairs, and the updated embedding comprises a weighted sum of the values, weighted by a similarity function of the query to each respective key.

[0056] There are many different possible attention operations. Some examples of transformer blocks including attention operations, are described in Vaswani et al. “Attention is all you need”, 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA.

[0057] Some attention operation may map a query and a set of key -value pairs to an output, where the query, keys, and values are all vectors. The output is computed as a weighted sum of the values, where the weight assigned to each value is computed by a compatibility function, e.g. a dot product or scaled dot product, of the query with the corresponding key.

[0058] In some implementations the attention operation is configured to apply each of a query transformation e.g. defined by a query matrix WQ, a key transformation e.g. definedby a key matrix WK, and a value transformation e.g. defined by a value matrix VF1', to two attention layer inputs: a matrix X of input data (e.g. where each row may be the input data 33 for a respective one of the input text tokens of the input data element) and a matrix Y of input data (e.g. derived from the inputs 31, 32) to the attention layer, to derive a query matrix (formed of query vectors) Q = XWQ, a key matrix (form of key vectors) K = YWK, and value matrix (forms of value vectors) V =which are used to determine an attended sequence for the output. For example the attention operation may be a dot product attention operation applied by applying each query vector to each key vector to determine respective weights for each value vector, then combining the value vectors using the respective weights to determine the attention layer output for each element of the input sequence. The attention layer output may be scaled by a scaling factor e.g. by the square root of the dimensions of the queries and keys, to implement scaled dot product attention. Thus, for example, an output of the attention f Q operation may be determined as softmax ^--^=-JV where d is a dimension of the key vector and the query vector (and in some implementations the value vector). A summation over the value vectors included in V is assumed here, weighted by the respective values f QK^ \ softmax V. In another implementation the attention operation can comprise an “additive attention” mechanism that computes the compatibility function using a feed-forward network with a hidden layer. As previously mentioned, the output of the attention operation may be further processed by one or more fully-connected, feed forward neural network layers.

[0059] The attention operation may implement multi-head attention, that is, it may apply multiple different attention operations in parallel. The outputs of these may then be combined, e.g. concatenated, with a learned linear transformation applied to reduce to the original dimensionality if necessary.

[0060] The output of the attention function may be added to the language input by an addition unit. The result may then be input to a feed-forward (FFW) network, such as a multilayer perceptron. The output of the FFW network is a vector. This may be input to a tanh gating unit which applies a component- wise operation in which each component is multiplied by a gating parameter a, (and (a non-linear function such as a tanh function) may be applied to the result). The output of the FFW network (or the tanh gating unit if present) may be added to the input to the FFW network by an addition unit. This produces the output of the gated crossattention layer, which is used as a “layer input” to the succeeding layer of the interleaved sequence of layers.

[0061] Each token processing layer may also preferably include a self-attention layer. The attention layer operates on key vector, value vector and query vectors generated from the input to the token processing layer. The self-attention operation may apply the attention operation explained above, except that the matrix of data input X plays the role of matrix Y also. The output of the self-attention layer is added to its input by an addition unit, and the result is fed to a feedforward (e.g. multi-layer perceptron) layer. The output of the FFW network (optionally as modified by a non-linear function, such as a tanh function) is added to the input of the FFW layer by an addition unit. This produces the output of the token processing layer.

[0062] If the gating parameter a is near zero, the output of the modification layer (gated cross-attention layer) is very close to the language input, so the VLM performs almost the same function as the trained language model. However, as the VLM is been trained (during the pretraining which is before the fine-tuning described above) the value of the gating parameter a may be raised to be significantly higher than zero, so that the image encoding vectors affect the output of the modification layer, and therefore affect the subsequent layers of the VLM stack.

[0063] As noted, the image encoding vectors output from the resamplers are used as control inputs to the modification layers. For example, each image encoding vector may be, e.g. for each of a number patches of the image, the different (learned) latent vector, resampled by the corresponding resampler, to generate the control input for a corresponding one of the modification layers, or a single (e.g. larger) resampled latent vector may be used to generate the control inputs for all the modification layers at once.

[0064] The input text tokens of a network input to the VLM are used as a prompt input for the multi-modal neural network, and are supplied to the first processing layer of the stack, which may be the firstlstgated cross-attention layer. Data passes through the VLM, to produce a token output (output token string), which is the output of the VLM.

[0065] The VLM may also include a text embedder, which is an embedding function that independently transforms each token in the vocabulary of tokens to a respective text embedding that has the same dimensionality as the input embeddings. The text embeddings can be fixed, e.g., pre-determined, or can be learned during the training of the language model neural network. The text embedder may be configured to receive the input text tokens and process them before supplying them to the first layer of the VLM stack.

[0066] The training method proposed in J-B Alayrac et al (2022) is briefly as follows, and may be used for the pre-training of the VLM in the current disclosure. The language model,the image encoder models, and the perceiver resampler are trained separately, and then “frozen”, i.e. no further changes are made to these networks. The VLM is then constructed, including interleaving the modification layers with the token processing layers. A training procedure is then carried out in which (e.g. only) the modification layers are trained. This is done based on a database of multi-modal training examples composed of (a) sample input data elements which are pairs of sample images (and optionally, for some or all sample input data elements, sequences of input tokens), and (b) corresponding sample token outputs representing an (appropriate) text response to the corresponding sample input data element.

[0067] Following this training procedure the VLM may be used to generate a textual output iteratively. In a first step, a network input is formed which includes an image and optionally one or more text tokens (text input). This is input to the VLM, such that the last layer of the stack produces an output which is the output of the VLM. This output is then added to the previous network input, and this updated network is input to the VLM, such that the last layer of the stack produces a new output which is a new output of the VLM. Continuing this process for multiple steps, a succession of network outputs are produced, which are tokens forming a response to the image and text input.

[0068] This VLM which is the result of training process (that is, the VLM resulting from the training method proposed in J-B Alayrac et al (2022)) may be used as the neural network system which is the starting point for the fine-tuning process described here. During the fine-tuning process, numerical parameters defining VLM, and specifically the modification layer(s) and / or the image encoder neural network(s), may be iteratively updated, e.g. without varying the operation of the token processing layers. In this way, the sophisticated language processing ability of the trained language model is not reduced. In particular, iterative updating of the numerical parameters defining the image encoder neural network may comprise iterative updating of numerical parameters defining the operation of the perceiver(s), optionally without updating the parameters of the image encoder. The updating may be performed, for example, by back-propagating gradients of the cost function with respect to the ones of the numerical parameters defining the resampler unit and / or the image encoder through the trained language model without adjusting parameters of the trained language model (i.e. they were “frozen”). It has been found experimentally that training the parameters of the trained language model resulted in overfitting (e.g. due to the limited size of the training database(s)), and generation of textual reports which were evaluated as being less good.

[0069] Turning to Fig. 2, the structure of a VLM neural network system 200 employing the Flamingo model is illustrated. This is a neural network system suitable for use in the presentdisclosure. The neural network system 200 is configured to receive as an input a network input 201 which includes one or more non-text data items (typically multi-dimensional images) and a textual input in the form of an input token string. In one case, illustrated in Fig. 2, the network input 201 may contain two non-text data items 203, 205. However, it is to be understood that there may be any number of data item(s).

[0070] The neural network system 200 extracts the two data items 203, 205 from the network input 201, and processes each data item using a modality network. As illustrated in Fig. 2, the neural network system contains two modality networks 207, 209, but alternatively, in the case that all data items have the same modality, there may be only a single modality network which in turn receives the data items 203, 205, and generates a corresponding output from each.

[0071] Note that if there are multiple data items in the network input 201, and if they have different corresponding modalities, a modality network may be provided for each corresponding modality, and each modality network is configured to receive the data item(s) of the corresponding modality (e.g. sequentially if there are multiple data items of the corresponding modality in the network input). For example, the modality network 207 may be configured to receive data item(s) of the network input which are a single image (e.g. still image), and the modality network 209 may be configured to receive data item(s) of the network input which are audio data. In another example, the modality network 207 may be configured to receive data items of the network input which are a single image (e.g. a still image), and the modality network 209 may be configured to receive data items which are a video (i.e. sequence of images) optionally with a soundtrack of audio data. In another example, the modality network 207 may be configured to receive data items of the network input which are two- dimensional images, and the modality network 209 may be configured to receive data items which are three-dimensional images (e.g. a set of two-dimensional image representing respective planes within an imaged object). There may be any desired number of modality networks for each corresponding data item modality which the query processing system 200 is configured to process.

[0072] Each modality network 207, 209 is configured to generate, from each data item it receives, one or more corresponding compressed representations of the data item. Each modality network may comprise a pre-trained encoder network 11 (e.g. trained within the system of Fig. 1(a)), and a compressed representation generation system in the form of a respective resampler 210, 211. Each resampler 210, 211 may, for example, be a “perceiver” as described in “Perceiver: General Perception with iterative attention”, A. Jaegle et al, 2021.Alternatively, the resampler 210, 211 may be as described below with reference to Fig. 3 and Fig. 4.

[0073] The resamplers 210, 211 take as input a variable number of image, video or audio features produced by the corresponding encoder network 11 from a data item, and produce a compressed representation of the data item having a fixed number of outputs (e.g. visual outputs).

[0074] Fig. 3 shows the structure of a possible modality network 207 in the case that the data item 203 is an image, including a respective “perceiver” resampler 210. Note that the modality network 209 has the same structure if the data item 205 is an image (indeed, as mentioned above, when data items 203, 205 have the same modality, a single network may play the roles of both the modality networks 207, 209).

[0075] The resampler 210 receives the output of the encoder 11 (an encoded version of the data item 203), and generates from it flattened data denoted X_f. The perceiver resampler also receives a predefined number of latent vectors (i.e. sets of latent values), each denoted X, which are learnt (trained) during the training of the modality network 207 described below.

[0076] The resampler 210 also includes one or more “resampler layers” 300. For simplicity, the resampler 210 is shown in Fig. 3 as having a single resampler layer 300. The compressed representation generation system may alternatively comprise a plurality of resampler layers, each having the structure illustrated as 300 in Fig. 3. The resampler layers 300 may be arranged as a stack of resampler layers, in which each layer has the same form as the resampler layer 300 shown in Fig. 3. Each resampler layer 300 (except the first) receives the output of the preceding resampler layer of the stack.

[0077] The resampler layer 300 (or, if there are multiple resampler layers 300, each resampler layer of the stack) receives the output of the corresponding encoder network 11, e.g. following a transformation to the output of the encoder network 11, such as a flattening operation (i.e. conversion of an e.g. 2-dimensional array, to a 1-D sequence).

[0078] Each of the resampler layer(s) 300 may perform a transformer operation as defined above; that is, it includes one or more transformer blocks or self-attention layers. For this purpose it may comprise a unit 301 which uses the flattened data X] _f and, by turns, different ones of the latent vectors X, to generate a key vector K(Xf, X) and a value vector V(Xf, X). This process uses a key matrix and a value matrix as described below. A unit 302 uses the latent vector X and a value matrix to generate a query vector Q(X). An attention unit 303 combines K, V and Q by the attention operation described above. The result is added tothe latent vector X by an addition unit 305. The result is passed to an adaptive unit, such as a multi-layer feedforward (FFW) perceptron 307. The output of the FFW perceptron 307 is added to the output of the addition unit 305 by a further addition unit 311. The result is an output vector for each of the latent vectors X, providing different respective compressed representations of the data item. The output of each resampler layer 300 may have the same dimensionality as the dimensionality of the latent variables X. This dimensionality may be much lower than the dimensionality of the encoder network 11.

[0079] Fig. 4 shows a variation of the structure of Fig. 3 in the case that the input data item 203 is a plurality of images (e.g. a plurality of video frames, or multiple 2D images of a single object, e.g. medical images of different respective planes within a subject). The data item 203 is split into individual frames 41, 42, 43. Each of these frames is encoded sequentially by the encoder network 11, to generate a corresponding encoded frame. Respective data 44, 45, 46 is added to each of the encoded frames, specifying the position of the frame in data item 203. The resulting data is input to the resampler 210, which may have the same structure as explained above with reference to Fig. 3. Note that although the input to the resampler 210 is different from in Fig. 3, the output, for a given latent vector X, is the same number of output tokens.

[0080] The operation of a perceiver resampler may be represented by the following pseudocode:Let x_f represent the outputs of the corresponding encoder network(s) 11. This is an array with dimensionality [T, S, d], where T is an integer variable representing a number of times (i.e. the number of encoder networks 11, or the number of times a single encoder network is used), S is an integer representing a number of spatial positions, and d is a further integer indicating a number of feature values for each position.Let x represent R learned latent variables, so this is array with dimensionality [R, d]. The number of layers of the resampler is denoted num layers.First, the time embeddings are added and the result flattened: x_f = x_f + time embeddings x_f =flatten (x_f) This produces an array of dimension [T*S,d] Then, for each integer value i in the range 1 to num layers: x=x+attention_i(q=x, kv=concatenation([x_f,x])) x=x+ffw_i (x).Here attention ! represents the attention unit 303 performed by the i-th resampler layer 300, based on the corresponding key, value and query matrices, and ffw_i represents the operation performed by the corresponding perceptron 307 of the i-th resampler layer 300.

[0081] Thus, the resampler 210 maps a variable size grid of spatial visual features (spatio-temporal features in the case of Fig. 4), to a fixed number of output tokens, independently from the input image resolution of the number of input video frames. Each resampler 210 has a set of learned latent vectors as queries and the keys and values are a concatentation of the visual features with the learned latent vectors.

[0082] Returning to Fig. 2, the neural network system 200 further includes a data-item- token processing model 220. This comprises a stack of layers, which are the token processing layers 131, 132, ..., 13n of the token processing model of Fig. lb, interleaved by a number j of gated cross-attention layers 231, 232, ...., 23j. The number j of gated cross-attention layers may be equal to the number n of token processing layers, but in variations n and j may be different from each other, e.g. such that there may be any number of token processing layers between any given pair of gate cross-attention layers.

[0083] The outputs from the resamplers 210, 211 are used as control inputs to the gated cross-attention layers. Optionally, the same outputs from the resamplers 210, 211 can be used as control inputs for all the gated layers 231, 232, ...23j. This was implemented in some successful experiments, in which 64 latent vectors were used, and 6 resampler layers, and the corresponding 64 vectors output by the perceiver resamplers 210, 211 were used together as the control inputs for all the gated cross-attention layers. However, this is not the only possibility. For example, there may be corresponding different (learned) latent vectors to generate respective control inputs for corresponding ones of the gated cross-attention layers.

[0084] The input token string is used as a prompt input for the data -item -token processing model 220, and is supplied to the first processing layer of the stack, which may be the 1 st gated cross-attention layer 231. Data passes (upwardly in Fig. 2) through the data-item- token processing model 220, to produce an output token string, which is the output of the data- item -token processing model 220.

[0085] The gated cross-attention layers 231, 232, ..., 23j each receive at least one of the compressed representations from the modality networks 207, 209, and perform a gating operation based on the received compressed representations.

[0086] Fig. 5 shows the structure of two consecutive layers 523, 513 of the data-item- processing model 220 of Fig. 2. The layer 523 is a gated cross-attention layer which is one ofthe layers 231, 232, ..., 23j. The layer 513 is one of the token processing layers 131, 132, ..., 13n.

[0087] The gated cross-attention layer 523 receives vision inputs 51, 52 (compressed representations) from respective ones of the modality networks 207, 209. It further receives a vector which is a language input 53, and which may be denoted y. If the gated cross-attention layer 523 is the first layer of the stack, the language input 53 is typically the input token string (prompt input), or a part of it. Otherwise the language input 53 to the gated cross-attention layer 523 is the output from a preceding layer of data-item-token processing model 220.

[0088] The gated cross-attention layer 523 applies an attention function 501 as described above, in which the vision inputs 51, 52 may define key and value vectors, and a query vector is defined based on the language input 53. The output of the attention function 501 is added to the language input 53 by the addition unit 502. The result is then input to a FFW network 503, such as a multi-layer perceptron. The output of the FFW network 503 is a vector which may be denoted x. This is input to a tanh gating unit 504 which applies a component- wise operation in which each component is multiplied by a gating parameter a, and a tanh function is applied to the result. The output of the FFW network 504 is added to its input by the addition unit 505. This produces the output of the gated cross-attention layer 523, which may be written as y + tanh ax).

[0089] The token processing layer 513, as in some known systems, includes a self- attention layer 506. The attention layer operates on key vector, value vector and query vectors generated from the input to the token processing layer 513. The output of the self-attention layer 506 is added to its input by an addition unit 507, and the result is fed to a feedforward (e.g. multi-layer perceptron) layer 508. The output of the FFW network 508 is added to the input of the FFW network 508 by the addition unit 509. This produces the output of the token processing layer 513.

[0090] Note that if the gating parameter a is near zero, the output of the gated crossattention layer 523 is very close to the language input 53, so the neural network system 200 performs almost the same function as the trained token processing model shown in Fig. 1(b). However, in use (i.e. after the neural network system 200 has been trained during a pre-training process which is carried out before performing a fine-tuning operation as described below with reference to Fig. 6) the value of the gating parameter is significantly higher than zero, so that the compressed representations (e.g. vision inputs) 51, 52 affect the output of the token processing layer 513.

[0091] Fig. 6 shows a flow chart depicting the flow of processing in the first aspect of the invention. The process disclosed 600 is a fine-tuning operation to train a pre-existing trained VLM (such as the neural network system 200 of Fig. 2) to process medical images to generate textual reports related to the contents of the medical images. In some instances, the medical images may be X-rays or other such images of a human subject, and the textual reports may be radiology reports, such as generated by a radiologist reviewing said medical image. The process 600 starts with obtaining a pre-trained neural network system 602, such as trained pre-existing neural network system 200 described above with reference to Fig. 2. This may be pre-trained on a large database of medical or non-medical images prior to the beginning of the process 600.

[0092] Following step 602, the process passes to step 604, where the neural network system iteratively updates network parameters of the pre-trained neural network system to reduce a value of a cost function. Thus, the pre-trained neural network system is fine-tuned. The cost function includes a first prediction cost term based on a first training database. The first database contains a plurality of datasets (“training datasets”) with each dataset corresponding to an individual, and the datasets each comprising at least one medical image and an associated text report. These medical images and associated text reports may substantially be in the form of the aforementioned example X-ray images and radiology reports. The first prediction cost term includes a cost value for each individual related inversely to a likelihood value of the associated text report for the respective individual also generated by the neural network system. The likelihood value may be generated in a separate step of processing by the neural network system. The likelihood value is conditioned on the at least one medical image. As noted above, the modality network 207 of the neural network system 200 may be configured to receive data items of the network input which are two-dimensional images, and the modality network 209 may be configured to receive data items of the network input which are three-dimensional images. During step 602 the modality network 207 may receive two- dimensional images (e.g. single X-ray or MRI image(s) of the individual) from the training datasets, and / or the modality network 209 may receive three-dimensional data (e.g. a set of two-dimensional image representing respective planes within the individual) from the training datasets.

[0093] Fig. 7 shows an exemplary computer system that may be used for training of a neural network system. For example, the exemplary computer system may be used to perform training of the neural network system as disclosed in Fig. 6, such that the neural network system is fine-tuned. In other examples, the exemplary computer system may be used to pre-train the neural network system, prior to any fine-tuning of the neural network system.

[0094] Exemplary computer system 700 includes input / output interfaces 702 to receive inputs from other devices, for example an input device such as a keyboard, or a network port, and to transmit outputs to other devices, such as a display, or a network port. Exemplary computer system 700 further contains, and receives power from, a power supply 704, which may be a mains power system, a battery storage system or other power supply means. Exemplary computer system 700 also includes a memory 706 for the storage of data, instructions, and the like. In the example of the neural network system as described above, the memory may store network parameters related to the neural network system, and may further store instructions related to the methods for training the neural network system. The exemplary computer system 700 further includes a processor 708, which may operate to perform instructions stored in the memory 706 of the exemplary computer system 700, and to generate updated parameters to be stored in the memory 706, or transmitted to other devices via the input / output interfaces 702. Various elements of the exemplary computer system are interconnected by a bus 710.

[0095] As noted above, a method of training a visual language model based on a training database comprising a plurality of training datasets corresponding to a plurality of individuals, using a cost function including a prediction cost term which is a weighted sum over the training datasets of a cost value, where the weight values depend upon whether the corresponding individuals exhibit a certain medical characteristic, constitutes the second independent aspect of the present disclosure. The second aspect is freely combinable with the first aspect, but can also be used independently. For example, it may be used in a situation in which the training of the neural network system does not employ a pre-trained neural network system, but instead, for example, takes, as the starting point of the training based on the first training database, an untrained VLM model. Such an arrangement may be appropriate for example if the size of the first training database is sufficiently great.

[0096] In the case of either the first or second aspect of the invention, the procedure above may be generalised such that there are multiple medical characteristics, and a given first individual may exhibit any combination of them. For example, a combined classification may be performed on the first training dataset, for example classifying the first individual as “abnormal” if the index values of any of a predetermined number of medical characteristics are equal to a particular value, such as a non-zero value. Continuing this example, a classification of “normal” may be returned if index values for all of the predetermined number of medical characteristics are equal to a particular value, such as a zero value. There may be a corresponding weights value for each combination of the medical characteristics. This is usedin the first prediction cost term to weight the cost value for each corresponding first individual, and is a function of the combination of the medical conditions which the first individual exhibits.

[0097] Furthermore, in the case of either aspect of the disclosure, the first training database may not be the only training database used for the fine turning operation. For convenience, we will refer to the individuals for which training datasets are included in the first dataset as “first” individuals, and the corresponding training datasets as “first” training datasets. The fine tuning operation may include one or more further training datasets, such as a “second training dataset”, comprising a plurality of “second” training examples relating to “second” individuals. Like the first training examples, the second training examples comprise at least one medical image of the corresponding second individual (as discussed above, this may be multiple images captured at different times, or of different portions of the body of the second individual, and / or captured with different modalities) and an associated text report. The first and second training databases may have been compiled according to different protocols, e.g. by different organisations, or according to respective regulations of different health authorities of different countries or differing states within one country. For the fine-tuning using the cost function, the cost function may include a “second” prediction cost term. The second prediction cost term may be defined in the same way (as described above) as the first prediction cost term, but using network inputs based on the second training datasets from the second training database, instead of the first training database.

[0098] Due to the different protocols under which the two training databases were created, it is advantageous if the two prediction cost terms are weighted by respective data- specific coefficients, denoted A. The respective data-specific coefficients may be tuned to maximise the benefits of jointly training the neural network system on both datasets. Thus, the cost function may take the form:where the first and second training databases are denoted by f^and D2, and the corresponding data-specific coefficients are denotedand A2. The data-specific coefficientsand A2may be tuned, e.g. the fine tuning procedure may be carried out for each of multiple candidate realisations of theand A2, to identify values of and A2for which the fine-training resultsin a trained neural network system having desirable properties (e.g. generates the highest proportion of accurate textual reports, according to a human assessor).

[0099] Optionally, the cost function may further comprise an auxiliary classification loss term. This too may be based on a database comprising a plurality of training datasets corresponding to a plurality of individuals. For convenience, the database will be referred to as a “third” database, comprising a plurality of “third” training datasets corresponding to a plurality of “third” individuals, but it is to be understood that some or all of the third individuals may in fact be first individuals (or second individuals, if a second training database is used). Each third training dataset defines a corresponding third training example. Each of the third training examples comprise a medical image of the corresponding third individual and comprising one or more index values (“auxiliary index values”) indicative of whether the third individual exhibits one or more corresponding medical characteristics.[000100] The auxiliary classification loss term is calculated by the neural network system generating, for each of the one or more third training datasets, a corresponding textual output. For each of the one or more third training datasets, a corresponding auxiliary loss value is generated based on whether the textual output indicates an “auxiliary” characteristic of the third individual indicated by the auxiliary index value(s). The auxiliary classification loss term is based on the auxiliary loss values.[000101] Fig. 8 shows an example process 800 for generating an auxiliary classification loss term. The process 800 begins with generating in step 802, for each training dataset, a corresponding textual output. In some instances, this may include generating a plurality of textual outputs (e.g. reports, report sections, etc.) for a single training dataset. The process 800 then passes to a second step 804 wherein an auxiliary loss value is generated by the neural network system for each of the training datasets. The generation process of the auxiliary loss value takes as an input whether the corresponding textual output generated in the first step 802 indicates a characteristic of the individual, the characteristic being indicated by one or more auxiliary index values. For example, the auxiliary index values may correspond to a set of medical conditions, and the indication of the characteristic of the individual may be an indication that the medical condition is present in the individual, the textual output providing an indication of visible features of medical images processed by the neural network system. The auxiliary loss value may thus be generated by the neural network system by processing the textual output to determine the presence of particular medical conditions.[000102] Following the processing of the textual outputs to generate an auxiliary classification loss value for each of the training datasets, the process passes to a third step 806.The third step in the process comprises generating, on the basis of the auxiliary loss values generated for each of the training datasets, an auxiliary classification loss term. The auxiliary classification loss term may just be the sum of the auxiliary loss values.[000103] For example, a characteristic of the individual may be “normal” if the auxiliary index values for a particular training dataset indicate that none of the medical conditions from a list of medical conditions, each associated with a particular auxiliary index value, are present. Alternatively, the characteristic may be “abnormal” if the auxiliary index value corresponding to any of the medical conditions associated with a particular index value indicates the presence of the respective medical condition. However, other examples are also envisaged, such as the sub-division of a set formed of the one or more auxiliary index values and medical conditions corresponding to them, and correspondingly three or more characteristics may be used. In one example, the characteristic may correspond to a particular medical condition, and thus each characteristic may be associated with a particular respective auxiliary index value, and hence a particular underlying medical condition. Following this step 806, the process 800 ends.[000104] For example, the textual output for each third training dataset may just be a response to a network input to the neural network system including the medical image(s) of the third training dataset, and including an input text asking a question about the auxiliary characteristic of the corresponding third individual. Each training dataset in the third training database may be associated with an auxiliary index value indicating whether the corresponding third individual exhibits the auxiliary characteristic. The auxiliary characteristic may, for example, be whether the third individual is “abnormal” or “normal”. That is, the auxiliary characteristic may be the example of the medical characteristic given above (in this case, when the third training database overlaps with the first and / or second training databases, the index values for the first and second training databases may be used as the auxiliary index values for the corresponding third training datasets). The input text may ask whether the third individual is “normal” or “abnormal”. The textual output is the answer to this question (e.g. based on a plurality of successively generated network outputs, as described above). Note that the calculation of the auxiliary loss term does not rely on the neural network system generating a complete textual report based on the medical image(s) of the third individuals. In other words, during the training, the neural network system is asked to perform two different tasks: report generation (to obtain the prediction cost term(s))) and medical characteristic classification (to obtain the auxiliary loss term). The two tasks complement each other.[000105] Note that the auxiliary index values may alternatively be more specific than just “normal” / ” abnormal”. For example, the may be multiple auxiliary index values for each thirdindividual and each of the auxiliary index values may relate to whether the third individual exhibits any of a plurality of different respective medical characteristics (e.g. different medical conditions). Any one of these medical characteristics would be enough for the third individual to be classified as “abnormal”. However, it has been found experimentally that an auxiliary classification loss term which uses an auxiliary characteristic more specific than “normal / abnormal” (e.g. an auxiliary classification loss term which trains the neural network system to generate outputs specifying a particular one of multiple different medical characteristics, rather than a single category “abnormal”), gives inferior results compared to an auxiliary characteristic of normal / abnormal.[000106] The auxiliary index values for the third training database, and / or the index values for the first and second training databases, may be obtained by an automated procedure from the corresponding training datasets. For example, if the (e.g. third) training datasets include a textual report describing the medical image(s) of the (e.g. third) training dataset, an automatic process may be used to obtain the (e.g. auxiliary) index values from the textual reports, (e.g. using the CheXpert labeller described in J. Irvin, et al., “Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison”, in Proceedings of the AAAI conference on artificial intelligence, volume 33, pages 590-597, 2019.[000107] Fig. 9 shows a high-level representation of an alternative example neural network system 900 which may be used in the examples of the aspects of the disclosure. Broadly speaking, the neural network system is formed of a plurality (e.g. sequence) of layers interconnected with each other. For example, the output of each layer of the sequence (except the first) may be taken as the input to an adjacent layer (the next in the sequence). Seven layers are depicted in Fig. 9 for ease of reference, but it would be appreciated that a different number of layers may be used in practice.[000108] A first layer 902 is an image encoder neural network. This image encoder may comprise a pre-trained image encoder 902a and a resampler unit 902b. These may receive, as an input to the neural network, one or more medical image(s) and generate an image encoding vector. In other examples, a plurality of image encoders may be provided, for example if the neural network system is designed to receive medical images with a plurality of modalities. For simplicity of representation, a single image encoder 902a is presented here. For example, the image encoder 902a may generate as an intermediate result a grid of features, which may be two-dimensional, and subsequently transform the grid to output a sequence of features as a vector. Alternatively, the image encoder 902a may sample the input, for example if the input is a video, and similarly obtain a three-dimensional grid of features, comprising a 2-D grid offeatures for each sample. The 3-D grid of features may then be “flattened” to transform the 3- D grid into a vector of one dimension.[000109] A second portion of a stack of layers forming part of the neural network system may be formed of layers 904, 906, 908, 910, 912 and 914. These layers may comprise different types of layer, for example token processing layers, and indeed may perform different tasks. For example, token processing layers may be provided as one or more of the layers 904, 906, 908, 910, 912, or 914. These layers may process the tokens that the respective layer receives as an input from a previous layer. Layers of different types may be combined within the same stack of layers forming part of the neural network system. For example, layers of different types may be interleaved, and there may be at least one layer of a first type situated adjacent to two layers of a second type, such that the layers of different types are alternating layers. Alternatively, layers of multiple types may be interleaved in another fashion, such that any number of layers of a first type are situated between layers of a second type.[000110] The layers 904, 906, 908, 910, 912, 914 (i.e. all layers except the first layer 902) may receive inputs from previous layers in the direction of processing, as indicated by arrow 916. For example, an input to a layer such as 906 may be an image encoding vector as created by the image encoder neural network 902, as modified by the token processing layer 904. In this way, a representation of the image input to the image encoder neural network 902 may proceed through the neural network system 900, and be modified by the layers therein, until reaching a final layer 914, which generates an output of the neural network system 900.[000111] As for the n-th token processing layer, 13n, of the neural network system of Fig. 2, the final layer 914 of the neural network 902 may generate one or more output tokens (normally one) for each network input. Successive output tokens generated by the layer 914, for successive corresponding network inputs to the neural network system 900, form an output token string. As discussed above, the network inputs may be generated auto-regressively, including previously generated output tokens. The layer 914 may generate a distribution over the vocabulary (i.e. respective values for each element of the vocabulary), and include an output unit (not shown) which selects a token based on the distribution. For example, the output unit may treat the distribution as a probability distribution, and select the output token according to the probability distribution (i.e. with each element of the vocabulary being selected with a probability proportional to the corresponding value of the distribution); or the output unit may select the output token as the element of the vocabulary for which the corresponding value of the distribution is highest.[000112] The output of the neural network system 900 may thus be a portion of text, which may be related to the image input to the image encoder neural network 902a. For example, the portion of text may be a portion of a medical report, where the medical report indicates a description of a medical image input to the neural network system.[000113] Other layers may be provided in the neural network system as part of one or more layers, or as separate layers. For example, a token-processing layer may include a selfattention layer. Other layers that may be included are feedforward layers and / or convolutional layers.[000114] One visual language model which can be used as the pre-trained neural network system is the PaLi-3 visual language model described in “PALI-3 Vision Language Models: Smaller, Faster, Stronger”, Xi Chen et al. (2023), https: / / arxiv.org / pdf / 2310.09199.pdf. A vision transformer (ViT) encodes an image into tokens which, together with embedded text tokens generated by a text embedding transformer based on a text input (the question, prompt, instruction), are passed to an encoder-decoder transformer (Vaswani et al., 2017, “Attention is all you need”. CoRR, abs / 1706.03762, http: / / arxiv.org / abs / 1706.03762) that generates a text output. The ViT and a text embedding transformer are trained to separately embed images and texts, such that a binary classifier using the sigmoid cross-entropy of the dot product of image and text embeddings correctly classifies whether the respective image and text correspond to each other or not. The outputs of the ViT image encoder before pooling form the visual tokens, which are linearly projected and prepended to the embedded input text tokens. Together, these tokens are passed into a pre-trained UL2 encoder-decoder language model (Tay et al., 2023, “U12: Unifying language learning paradigms”), which generates text output. In one case, the training performed by the present disclosure may be performed by iteratively updating numerical parameters of any one or more of the ViT, the text embedding transformer, or the encoder-decoder language model.[000115] Once the neural network system has been fine-turned it may be used for generating medical reports relating to “current” individuals (e.g. current patients for whom at least one medical image has been obtained, and for whom a corresponding textual report is desired). This is based, as explained above, on generating successive network inputs each including the same one or more medical image(s) x relating to the current individual and optionally a text input, to generate corresponding network outputs. Each network input except the first may (e.g. in the case that the neural network is as shown in Fig. 2 or 9) include the network output(s) generated based on the preceding network input(s). Thus, the neural networksystem generates the medical report y as an output token string composed of successively generated output tokens.[000116] Fig. 10 depicts an example procedure 1000 of generating a medical report. The procedure begins, and advances to method stage 1002. In this stage, a trained neural network system is obtained, for example one which has been trained (fine-tuned) using the method of Fig. 6 as discussed above.[000117] Following this, the procedure passes to a second method stage 1004. Stage 1004 comprises the processing of a network input comprising for each of a sequence of steps and generating a corresponding network output for each step. The network input at one or more or each step may comprise at least one medical image x related to an individual (e.g., a human patient), and, in each step (except, optionally, the first), the network input may comprise the network output(s) at each previous step. Thus, network outputs may be generated auto- regressively. For example, for each of a sequence of steps, processing may be performed on the same medical image related to an individual, which is unchanged for different steps, and the network output to the previous step is dependent on the previous step. In this way, a portion of the network input may change for each step of processing method, whilst a portion of the network input may remain constant for each step of the processing method. Thus, an output string of output tokens is generated.[000118] In some example implementations, the network input for one or more of the sequence of steps may not comprise a medical image of the individual, but may comprise other inputs such as a system prompt, template, output example, or network output from a previous step. In some examples, the network input may, alternatively or additionally to a medical image, comprise a system prompt for the trained neural network system providing instructions to the system for how the system is to process and perform analysis on a medical image received in the same, or another, processing step. In some examples, the network input may comprise a medical report template, specifying a format for a medical report produced by the trained neural network in step 1006, such as a human-readable format or a machine-readable format. In some examples, the network input may comprise so-called few-shot examples related to a medical report to be generated by the trained neural network in step 1006, where few-shot can refer to, e.g. 1-20 examples or 1-10 examples. The few-shot examples may comprise second medical images of one or more individuals (which may or may not include the individual to which the medical image corresponds) and associated medical reports or subsections of medical reports. This may enable the trained neural network to output, in step 1006, a medical report similar to the few-shot examples.[000119] In some example implementations for some steps of the processing method, the network input may comprise reasoning and / or tool outputs forming part of one or more network outputs for one or more prior processing steps. For example, in one processing step the network input could comprise a medical image and the network output corresponding to the processing step could comprise reasoning regarding the medical image. The network output comprising reasoning may then form part of the network input for a subsequent processing step, and the network output of the subsequent step may comprise a medical report or part thereof. The reasoning may comprise so-called chain-of-thought reasoning, e.g. in which a sequence of logical steps, presented as network inputs, lead in a causal chain from a first network input that may, e.g. represent a query to a network output that represents a response to the first network input. As an example, at one step, the neural network system can receive a medical image and generate a chain-of-thought response that provides reasoning about the image to be used for the medical report, and at a subsequent step, this reasoning output can be used to generate the medical report.[000120] In the above-described examples the network output for a processing step is generally produced on the basis of the network input for that processing step. That is, the network output may be produced on the basis of the medical image, system prompt, template, output example, and / or network output from a previous step included in the network input for the relevant processing step. Following the method stage 1004, the procedure passes to a third method stage 1006, wherein a medical report is generated (and output) based on the outputs of the network in the preceding method stage.[000121] Stage 1006 may just comprise assembling the output string into a data file in a desired format and / or outputting it, e.g. to a user interface (screen or printer). Alternatively, it may comprise more complex processing. For example, stage 1004 may be performed repeatedly for different portions of an input medical image, and may generate for each portion of the input medical image a portion of a medical report related to the features visible in that portion of the input image. The method stage of generating the medical report may comprise assembling the portions of medical reports into a whole medical report. Alternatively, the method stages may correspond more abstractly to the input image, and greater processing may be applied to the assembly of the medical report, such that, for example, repetition of features may be avoided within the medical report, and such that the medical report may be accurate and concise. Following this method stage, the procedure ends.[000122] Optionally, there may be a process of providing the medical report(s) generated by the neural network system for current individual(s) to a human expert, and, upon receivingdata input from the human expert making a change to the medical report(s), updating the medical reports. This “human-in-the-loop” process, i.e. obtaining data input from a human expert to modify an automatically generated textual report (whether generated by the neural network systems described above, or otherwise), constitutes another independent aspect of the invention. Experimentally, it has been found to produce medical reports with higher accuracy (according to human reviewers) than medical reports produced from the medical images by the neural network system or the human expert alone.[000123] Accumulated evidence has shown that known methods for automatic report generation metrics fail to appropriately evaluate many nuanced issues of radiology reports (“Evaluating progress in automatic chest X-ray radiology report generation”, Yu et al., 2023 in medRxiv pp. 2022-08). To achieve a more fine-grained and realistic assessment of the clinical quality of radiology reports generated by an example of a process according to Fig. 10 (the one explained above with reference to Fig. 2), an expert evaluation for reports in both the MIMIC- CXR and INDI datasets was conducted. A group of 11 radiologists in the US and 16 certified radiologists in India were recruited as raters to perform the evaluation task, namely a pairwise preference test.[000124] In the pairwise preference evaluation task, the radiologist raters were provided with (i) a frontal view of a CXR (chest X-ray) image, (ii) a radiology report generated by an Al system as explained herein with reference to Fig. 2, and (iii) the original report written by a radiologist, and asked to assess the relative usefulness of the two reports for the given image. [000125] Fig. 11(a) shows a distribution of preference labels for both 300 cases (100 normal cases and 300 abnormal cases), across the datasets INDI and MIMIC-CXR, as assigned by four of the raters per case. The size of the areas 1001 indicates the proportion of cases for which none of the four human raters preferred the Al-generated report or was neural. The size of the areas 1002 indicates the proportion of cases for which one of the four human raters preferred the Al-generated report or was neural. The size of the areas 1003 indicates the proportion of cases for which two of the four human raters preferred the Al-generated report or were neural. The size of the areas 1004 indicates the proportion of cases for which three of the four human raters preferred the Al-generated report or were neural. The size of the areas 1005 indicates the proportion of cases for which all four of the four human raters preferred the Al-generated report or were neural (e.g. 36.7% in the case of the INDI dataset). Note that in over half of cases (in 77.7% of the INDI cases and 56.1% of the MIM-CXR cases), reports produced by the example of the disclosure were rated as equivalent or superior to a report by a clinician by at least half the raters.[000126] Fig. 11(b) breaks down the results for the dataset INDI as between the 200 abnormal reports, and the 100 normal reports, showing that for 94% of normal INDI cases, reports produced by an example of the present disclosure were rated as equivalent or preferred to reports written by clinicians by at least half of the radiologists in the panel.[000127] Fig. 12 provides an example procedure 1200 for generating a human-reviewed medical report. The procedure begins and passes to method stage 1202. Here, a pre-trained neural network system is obtained. This step may be akin to the correspondingly numbered step 1002 in Fig. 10, or may comprise obtaining an alternate neural network system, or one that has been trained in a different manner or extent. Following this stage 1202, the procedure passes to a second method stage 1204. The second stage comprises processing a corresponding network input and generating a network output. This is performed for each step in a sub-process of the processing stage. For example, the sub-process may comprise steps corresponding to the layers of the pre-trained neural network system, or may comprise steps corresponding to a subset of the layers of the pre-trained neural network system.[000128] Similarly to Fig. 10, as discussed above, a network input for one or more of the sequence of steps may contain a medical image of an individual such as a patient. In some example implementations the network input for one or more of the sequence of steps may not comprise a medical image of the individual. In some examples, the network input for one or more of the sequence of steps may, alternatively or additionally to a medical image, comprise a system prompt, a template, an output example, and / or a network output from a previous step. In some examples, the network input may, alternatively or additionally to a medical image, comprise a system prompt for the trained neural network system providing instructions to the system for how the system is to process and perform analysis on a medical image received in the same, or another, processing step. In some examples, the network input may comprise a medical report template, specifying a format for a medical report to be produced by the trained neural network in step 1206, such as a human-readable format or a machine-readable format. Other example formats for a medical report may include a Findings section of text and an Impression section of text. In some examples, the network input may comprise few-shot examples related to a medical report to be generated by the trained neural network in step 1206, and the few-shot examples may comprise second medical images of one or more individuals (which may or may not include the individual to which the medical image corresponds) and associated medical reports or subsections of medical reports. This may enable the trained neural network to output, in step 1206, a medical report similar to the few-shot examples.[000129] In some example implementations for some steps of the processing method, the network input may comprise reasoning and / or tool outputs forming part of one or more network outputs for one or more prior processing steps. For example, in one processing step the network input could comprise a medical image and the network output corresponding to the processing step could comprise reasoning regarding the medical image. The network output comprising reasoning may then form part of the network input for a subsequent processing step, and the network output of the subsequent step may comprise a medical report or part thereof. The reasoning may, e.g. comprise chain-of-thought reasoning as described above.[000130] In each of the above-described examples the network output for a processing step is produced on the basis of the network input for that processing step. That is, the network output may be produced on the basis of the medical image, system prompt, template, output example, and / or network output from a previous step included in the network input for the relevant processing step. Following the second method stage 1204, the procedure passes to a third stage 1206, wherein a medical report is generated on the basis of the network outputs previously generated. For example, this step may be similar to the corresponding method stage 1006 in Fig. 10, and may include assembling portions of a medical report generated by different steps of the procedure.[000131] Following this method stage 1206, the procedure passes to a fourth stage 1208, wherein the medical report is provided to a human expert. For example, this may comprise transmission of the medical report generated in the previous step to a human expert. Finally, the procedure passes to a fifth stage 1210, wherein data input by the human expert is used to modify the medical report. This may take any number of forms, such as corrections, modifications, additions, deletions, reordering of content, etc. Subsequently, the procedure ends, having produced a reviewed and modified medical report.[000132] The methods of Figs. 10 and 12 permit the generation of textual reports (medical report, such as radiology reports) with an accuracy which, according to experimental results, is comparable to, or even better than, textual reports generated by human experts (radiologists). There is a worldwide shortage of radiologists, which limits access to expert care and imposes heavy workloads on available radiologists, contributing to avoidable errors and delays in report delivery. These disadvantages can be addressed by generating textual reports according to the present disclosure, almost instantaneously and to high accuracy. Furthermore, if the training datasets of some or all of the training database(s) include images of different modalities, the textual reports can combine information from these different modalities, relieving possible shortages of radiologists expert in interpreting a set of images of multiple correspondingmodalities. Based on the textual reports, a clinician may devise a strategy for treatment of a current individual who is the subject of the text report (e.g. a current patient), and administer it to the patient.[000133] Furthermore, the training is achieved at reduced computational cost compared to some existing textual report generation systems, both as measured by the number of computational operations and the required size of the medical training database(s) employed, due to making use of a pre-trained VLM.[000134] By weighting the contributions to the cost function based on medical characteristics, during the VLM the performance of the VLM for individuals with abnormal characteristics can be maintained, even if the training database(s) used for the fine tuning only include a very low proportion of individuals exhibiting those characteristics (in some cases, the proportion of “normal” individuals in the training databases may be over 90%).[000135] Furthermore, although different available training databases may have been compiled with different characteristics (e.g. with textual reports produced by medical staff with different training or medical specialisms) and with different respective sizes, they can be used together to give superior training of the VLM than training based on one database alone, using the data-specific coefficients to prevent the VLM becoming over-trained in relation to one of the training databases, relative to another.[000136] While the explanation of the present concepts given above is based on the concept of training (e.g. fine-tuning) the neural network system based on the training database(s), an alternative or additional way of generating training reports using the medical database(s) is to use “multi-shot prompting”. That is, a prompt may be generated for a neural network system (either the pre-trained neural network system or a neural network system finetuned as described above) which comprises at least one example of performing the textual report generation text. For example, a (pre-trained or fine-tuned) neural network system, such as one of the ones describe above, may receive a network input including one or more “prompt examples”, where each prompt example is an image-text report pair, taken from one or more of the first (or if present) second training databases.[000137] Particularly if more than one of the prompt examples is used, they may optionally be weighted in the network input based in the manner described above, in a manner similar to that done for the training datasets in the fine-tuning example. For example, the weight value for a prompt example relating to a first individual exhibiting a medical characteristic may depend inversely on the proportion of the plurality of first individuals for which prompt examples are included in the network input who exhibit the medical characteristic. Similarly,the weight value for each first individual not exhibiting the medical characteristic depends positively on the proportion of the plurality of first individuals for which prompt examples are included in the network input who exhibit the medical characteristic.[000138] Optionally, an associated instruction (e.g. “generate a radiology report”) may also be included in the network input alongside the prompt examples (image-text report pairs). [000139] Subsequent to the prompt examples and associated instruction (if any), an image for a current individual (e.g. a patient - e.g. different from the first and second individuals - for whom a textual report is desired) may be given in the network input. The neural network system may then produce an initial portion of a text report for the current individual. This may be added to the network input and input again to the neural network system, to generate a new network output encoding a subsequent part of the text report for the current individual. This process is repeated to create the complete text report for the current individual. Thus, by incontext learning on the provided prompt examples in the network input, a textual report (radiology report) for the current individual may be produced. This initial textual report may be improved by the “human-in-the-loop” process described above, based on data input from a human expert.[000140] This specification uses the term “configured” in connection with systems and computer program components. For a system of one or more computers to be configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions.[000141] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a machine- readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, theprogram instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus.[000142] The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.[000143] A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a stand alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network.[000144] In this specification the term “engine” is used broadly to refer to a softwarebased system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers.[000145] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows canalso be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.[000146] Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.[000147] Computer readable media suitable for storing computer program instructions and data include all forms of non volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD ROM and DVD-ROM disks.[000148] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s device in response to requests received from the web browser. Also, a computer can interact with a user by sending text messages or other forms of message to a personal device, e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return.[000149] Data processing apparatus for implementing machine learning models can also include, for example, special-purpose hardware accelerator units for processing common and compute-intensive parts of machine learning training or production, i.e., inference, workloads. [000150] Machine learning models can be implemented and deployed using a machine learning framework, e.g., a TensorFlow framework.[000151] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.[000152] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the device, which acts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received at the server from the device.[000153] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.[000154] Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.[000155] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.

Claims

CLAIMS:

1. A computer-implemented method of training a neural network system for processing medical images to generate textual reports, the method comprising: obtaining a pre-trained neural network system defined by a plurality of network parameters and operative, upon receiving a network input comprising at least one image, to generate an output based on the at least one image; and repeatedly updating the network parameters to reduce the value of a cost function, the cost function including a first prediction cost term based on a first training database comprising a plurality of first training datasets corresponding to a plurality of first individuals, each first training dataset comprising at least one medical image of the corresponding first individual and an associated text report; the first prediction cost term comprising a corresponding cost value for each first individual which depends inversely on a likelihood value, generated using the neural network system, of the associated textual report, conditioned on the corresponding at least one medical image.

2. The method according to claim 1, wherein the network input includes a textual input.

3. The method according to claim 2, wherein the output generated is a value indicative of a predicted likelihood of one or more candidate text continuations of the textual input.

4. The method according to any one of claims 1 to 3, wherein the pre-trained neural network system was trained using a visual language model training database of training examples, each comprising at least one image and a text response to the image, wherein less than 20% of training examples comprise medical images.

5. The method according to any one of claims 1 to 4, in which each of the first training datasets is associated with a corresponding index value indicative of whether the first individual exhibits a medical characteristic, and the first prediction cost term is a sum over the cost values for the first individuals weighted by corresponding first weight values for the first individuals based on the corresponding index values.

6. A computer-implemented method of training a neural network system for processing medical images to generate textual reports, the method comprising:repeated updating the network parameters of a neural network system defined by a plurality of network parameters and operative, upon receiving a network input comprising an image, to generate an output based on the at least one image; the repeated updates each being to reduce the value of a cost function including a first prediction cost term based on a first training database comprising a plurality of first training datasets corresponding to a plurality of first individuals, each first training dataset comprising at least one medical image of the corresponding first individual and an associated text report, and each of the first training datasets being associated with a corresponding index value indicative of whether the first individual exhibits a medical characteristic; the first prediction cost term comprising a corresponding cost value for each first individual which depends inversely on a likelihood value, generated using the neural network, system of the associated textual report, conditioned on the corresponding at least one medical image, the first prediction cost term is a sum over the cost values for the first individuals weighted by corresponding first weight values for the first individuals based on the corresponding index values.

7. The method of claim 6, wherein the wherein the network input includes a textual input.

8. The method according to claim 7, wherein the output generated is a value indicative of a predicted likelihood of one or more candidate text continuations of the textual input9. The method of any one of claims 5 to 8, in which the weight value for each first individual exhibiting the medical characteristic depends inversely on the proportion of the plurality of first individuals exhibiting the medical characteristic; and the weight value for each first individual not exhibiting the medical characteristic depends positively on the proportion of the plurality of first individuals exhibiting the medical characteristic.

10. The method of any preceding claim, wherein each likelihood value is formed by dividing each textual report into a sequence of portions, and calculating the likelihood value by combining a corresponding portion likelihood value for each portion generated by the neural network system, conditioned on the corresponding at least one medical image and, for each portion except the first portion of the sequence, on the earlier portions of the sequence.

11. A method of any preceding claim in which the cost function further includes a second prediction cost term based on a second training database comprising a plurality of second training datasets corresponding to a plurality of second individuals who are different from the first individuals, each second training dataset comprising at least one medical image of the corresponding second individual and an associated text report; the second prediction cost term comprising a corresponding cost value for each second individual which depends inversely on a likelihood value, generated using the neural network, of the associated textual report, conditioned on the corresponding at least one medical image; the first and second prediction cost terms being weighted in the cost function by respective weighting values.

12. The method of claim 11, in which the first individuals are residents of a different country from the second individuals.

13. The method of any preceding claim, in which the cost function further includes an auxiliary classification loss term based on a third training database comprising a plurality of third training datasets corresponding to a plurality of third individuals, each of the third training examples comprising at least one medical image of the corresponding third individual and being associated with one or more auxiliary index values indicative of whether the third individual exhibits one or more corresponding medical characteristics, the auxiliary classification loss term being calculated by: generating, by the neural network system, for each of the one or more third training datasets, a corresponding textual output; generating, for each of the one or more third training datasets, a corresponding auxiliary loss value based on whether the corresponding textual output indicates a characteristic of the third individual indicated by the one or more auxiliary index values; and generating the auxiliary classification loss term based on the auxiliary loss values.

14. The method of claim 13, in which the characteristic is whether the individual exhibits any of a plurality of different respective medical characteristics associated with corresponding ones of the auxiliary index values, the auxiliary loss value being higher if the corresponding textual output does not indicate that the third individual exhibits the characteristic.

15. The method according to claim 13 or claim 14 in which the auxiliary index values are obtained by an automated procedure from the corresponding training datasets.

16. The method of claim 15 in which each third training dataset includes a textual report describing the corresponding at least one medical image, and the corresponding auxiliary index values are obtained based on the corresponding textual reports.

17. The method according to any of claims 13 to 16 in which the third individuals include one or more of the first individuals and / or one or more of the second individuals, the corresponding third training datasets being the corresponding first training datasets and / or the corresponding second datasets.

18. A method according to any preceding claim in which the neural network system comprises: a trained language model defined by a sequence of token processing layers, and operative to process a textual input comprising a plurality of text tokens selected from a vocabulary, to generate a textual output comprising a plurality of text tokens and which is a textual response to the text input; and at least one an image encoder neural network operative to receive at least one input image to generate at least one corresponding image encoding vector.

19. A method according to claim 18 in which the neural network system further comprises: a plurality of modification layers interleaved with the sequence of text processing layers to form an interleaved sequence of layers, each modification layer being arranged to receive a corresponding layer output of a corresponding preceding layer of the interleaved sequence of layers, and to transmit a corresponding layer input to the corresponding succeeding layer of the interleaved sequence of layers, and each modification layer being configured to modify the corresponding layer output, based on a corresponding portion of the at least one image encoding vector, to generate the corresponding layer input.

20. The method of claim 19, in which the modifying of the neural network system is performed by updates to numerical parameters defining the operation of the at least one modification layer and / or the image encoder neural network, without modifications to the trained language model.

21. The method of any of claims 18 to 20, in which the at least one image encoder neural network includes: an image encoder for generating visual features from the at least one image, and a resampler unit for processing the visual features into an image encoding vector of fixed length.

22. The method of claim 21, in which the modifying of the multi-modal language model neural network comprises back-propagating gradients of the cost function with respect to the ones of the numerical parameters defining the resampler unit and / or the image encoder through the trained language model without adjusting parameters of the trained language model.

23. A method of producing a medical report relating to an individual, the method comprising: obtaining a neural network system trained by a method according to any preceding claim, in each of a sequence of steps, the neural network system processing a corresponding network input for the step, to generate a corresponding network output for the step, one or more or each network input comprising at least one medical image of the individual and each network input except the first network input being based on a corresponding network output of the neural network in the preceding step, and generating the medical report based on the network outputs.

24. A method of producing a medical report relating to an individual, the method comprising: obtaining a neural network system for processing medical images to generate medical reports; in each of a sequence of steps, the neural network system processing a corresponding network input for the step, to generate a corresponding network output for the step, wherein one or more or each network input comprises at least one medical image of the individual and each network input, except the first network input, is based on corresponding network outputs of the neural network in the preceding steps; providing the medical report to a human expert based on the network outputs; and modifying the medical report based on data input from the human expert.

25. The method of claim 24, in which the neural network system was trained by the method of any of claims 1 to 22.

26. One or more computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform the operations of the respective method of any one of claims 1-25.

27. A system comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the respective method of any one of claims 1-25.

Citation Information

Patent Citations

  • Automatic diagnosis report preparation

    US20200411150A1

  • Generating reports from scanned images

    US20230102428A1

  • Method and system for automated generation of text captions from medical images

    US20230274420A1

Cited By

  • Multi-modal large model design method and system with capability of simultaneously processing multiple medical visual language tasks

    CN120597938A

  • Multi-modal semi-supervised image analysis and structured report generation method and application

    CN120783934A

  • A multimodal semi-supervised image analysis and structured report generation method and application

    CN120783934B

  • Training method and device for multi-modal medical pre-training model

    CN120851128A

  • Chest radiograph image report generation method and device based on retrieval enhancement generation

    CN121964040A