Methods and systems for mitigating effects of attacks of generative models

By using a boson sampler to generate latent vectors for training generative models while keeping configuration settings separate, the method secures the models against attacks and maintains performance, addressing resource-intensive training and vulnerability issues.

GB2635128APending Publication Date: 2025-05-07ORCA COMPUTING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
GB2023016443
Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-27
Publication Date
2025-05-07

AI Technical Summary

Technical Problem

Training generative models requires significant financial and computational resources, and they are vulnerable to attacks such as model extraction and model inference, where adversaries can degrade their performance by using different latent spaces.

Method used

Utilize a boson sampler to generate latent vectors for training generative models, keeping the configuration settings separate from the model, ensuring that only compatible latent vectors are used for inference, thereby securing the model and maintaining performance.

Benefits of technology

Enhances security by preventing adversaries from replicating the model's performance without access to the boson sampler and its configuration settings, mitigating attacks and maintaining optimal model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Methods and systems for training and using generative models. A method is provided for performance by a first system and a second system having access to a boson sampler which includes communicating,
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present disclosure relates to methods and systems or training a generative model or using a trained generative model. In particular but not exclusively methods and systems described herein utilise a boson sampler. Background

[0002] Trained machine learning models are increasingly becoming vital to the world economy and to the individual organisations that deploy them. Generative models, trained to generate a synthetic dataset (e.g. a fake image) similar to the collection of genuine datasets (e.g. genuine images) on which the model is trained, have proven to be particular useful. However, training such generative models may require sizeable collections of genuine datasets for training, may require large financial commitments (the financial costs of training high quality' generative models can run into millions of dollars), and may require extensive computational power to perform the training. Perhaps unsurprisingly, this lias led to increased attention in the security of such trained models and the genuine datasets on which they are trained. Summary

[0003] According to an aspect of the present disclosure, a method is provided. The method is suitable for performance by a first system (e.g. an electronic system, for example an electronic device) and a second system (e.g. a hybrid system or heterogeneous system comprising a set of one or more processors and a boson sampler). The method comprises communicating, from the first system to the second system, a request for a set of latent vectors for use in training a generative model to generate a synthetic dataset. The method further comprises, at the second system, based at least in part on the request, selecting configuration settings for the boson sampler. The method further comprises, at the second system, operating the boson sampler to produce a batch of samples, the boson sampler configured in accordance with the selected configuration settings. The method further comprises, at the second system, determining the set of latent vectors from the batch of samples. The method further comprises communicating, from the second system to the first system, the determined set of latent vectors. The method further comprises, at the first system, training the generative model to generate a synthetic dataset using the set of latent vectors.

[0004] Training a generative model in this way may inherently link the trained generative model at the first system with the latent space from which the set of latent vectors are derived, which in turn is inherently linked with the configuration settings of the boson sampler at the second system. When using the trained generative model for inference (in other words using the model to generate new synthetic datasets), the performance of the model (e.g. the quality of the output of the model) is degraded unless a latent vector from the same latent space is used. Accordingly, without access to a boson sampler or knowledge of the configuration settings of the boson sampler used to train the model, the performance of the trained model is suboptimal. Advantageously, this training method mitigates the effects of attacks on machine learning models including model extraction attacks and model inference attacks. Further advantageously, using this training method the configuration settings are not stored at the same system as tire trained model, thereby improving security.

[0005] The method may further comprise, at the second system, associating a model identifier with the selected configuration settings. The method may further comprise communicating the model identifier from the second system to the first system. The method may further comprise, at the first system, associating the model identifier with the trained generative model. Advantageously, use of a model identifier may enable the second system to assist in training or using multiple generative models at the first system or at multiple first systems.

[0006] The method may further comprise communicating, from the first system to the second system, a further request for a further set of latent vectors for use in further training a generative model to generate a synthetic dataset. The method may further comprise, at the second system, based on the further request, updating the configuration settings. The method may further comprise, at the second system, operating the boson sampler in accordance with tire updated configuration settings to produce a further batch of samples. The method may further comprise, at tire second system, determining the further set of latent vectors from the further batch of samples. The method may further comprise communicating, from the second system to the first system, the determined further set of latent vectors. The method may further comprise, at the first system, further training the generative model using the further set of latent vectors.

[0007] Using multiple sets of latent vectors to train the generative model and updating the configuration settings accordingly may enable the configuration settings of the boson sampler to also be trained, for example to make the latent space more suitable for the generative model. Advantageously, training the generative model in this way may mean that less training of the generative model is required for the same performance, or the generative model may be trained to provide a better performance compared to the case when the configuration settings are not updated. The further request for latent vectors may comprise information such as performance data indicating a performance of the trained model (e.g. one or more cost function values), that the second system may use to determine updated configuration settings. The further request may further include one or more latent vector identifiers identifying the specific latent vectors to which the performance data relates. The method may comprise, at the second system, associating a model identifier with the updated configuration settings.

[0008] The method may further comprise communicating, from the first system to the second system, a request for one or more latent vectors for use with the trained generative model (or further trained generative model as may be appropriate). The method may further comprise, at the second system, based at least in part on the request, identifying configuration settings for the boson sampler. The method may further comprise, at the second system, operating the boson sampler in accordance with the identified configuration settings to produce one or more samples. The method may further comprise determining, from the one or more samples, the one or more latent vectors. The method may further comprise communicating, from the second system to tire first system, the one or more latent vectors. The method may further comprise, at the first system, using at least one of the one or more latent vectors to generate a synthetic dataset. The request for one or more latent vectors may comprise a model identifier usable at the second system to identify the appropriate configuration settings. Advantageously, when using the trained model for inference, by using latent vectors derived from the same probability distribution as that used in training the model, the performance of the trained model is maintained.

[0009] According to an aspect of the present disclosure, a method is provided. The method is suitable for performance by a first system (e.g. an electronic system, for example an electronic device). The method comprises sending a request for a set of latent vectors for use in training a generative model to generate a synthetic dataset. The method further comprises receiving, in response to the request, the set of latent vectors, wherein the latent vectors have been determined from a batch of samples produced by a boson sampler. The method further comprises training a generative model to generate a synthetic dataset using the set of latent vectors.

[0010] The method may further comprise receiving a model identifier. The method may further comprise associating the model identifier with the generative model trained using the set of latent vectors. Associating tire model identifier with the generative model may comprise recording the association of the model identifier with the trained generative model.

[0011] The method may further comprise sending a further request for a further set of latent vectors for use in further training tire generative model. The method may further comprise receiving, in response to the further request, the further set of latent vectors, wherein the further set of latent vectors have been determined from a further batch of samples produced by a boson sampler. The method may further comprise using the further set of latent vectors, further training the generative model to generate a sy nthetic dataset. The further request may be indicative of a model identifier associated with the trained generative model. The further request may include performance data indicative of a performance of the generative model trained using the set of latent vectors. The performance data may comprise a function value of a function associated with the training of the generative model, and a latent vector identifier associated with a latent vector used in training the generative model.

[0012] The method may further comprise sending a request for one or more latent vectors for use with the trained (or further trained) generative model. The method may further comprise receiving, in response to the request, the one or more latent vectors. The method may further comprise using the one or more latent vectors with the trained (or further trained) generative model to generate a synthetic dataset.

[0013] According to an aspect of the present disclosure an electronic system is provided. The electronic system comprises one or more processors. The electronic system is configured to send a request for a set of latent vectors for use in training a generative model to generate a synthetic dataset. The electronic system is further configured to receive the set of latent vectors, wherein the latent vectors have been determined from a batch of samples produced by a boson sampler. The electronic system is further configured to use the set of latent vectors to train (or further train) a generative model to generate a synthelic dataset. For model inference, the electronic system may be configured to send a request for one or more latent vectors for use with the trained (or further trained) generative model, to receive the one or more latent vectors, and to use at least one of the one or more latent vectors to generate a synthetic dataset using the trained (or further trained) model.

[0014] According to an aspect of the present disclosure, a method is provided. The method is suitable for performance by a second system (e.g. a hybrid system or heterogeneous system comprising a set of one or more processors and a boson sampler). The method comprises receiving a request for a set of latent vectors for use with training a generative model to generate a synthetic dataset. The method further comprises based at least in part on the request, selecting configuration settings for a boson sampler. The method further comprises operating the boson sampler to produce a batch of samples, the boson sampler configured in accordance with the selected configuration settings. The method further comprises determining, from the batch of samples, the set of latent vectors. The method further comprises sending the set of latent vectors.

[0015] The method may further comprise associating a model identifier with the selected configuration settings. The method may further comprise sending the model identifier.

[0016] The method may further comprise receiving a further request for a further set of latent vectors for use with further training the generative model. The method may further comprise, based on the further request, identifying configuration settings for the boson sampler and updating the identified configuration settings. The method may further comprise operating the boson sampler to produce a further batch of samples, the boson sampler configured in accordance with the updated configuration settings. The method may further comprise determining, from the further batch of samples, the further set of latent vectors. The method may further comprise sending the determined further set of latent vectors, wherein the further set of latent vectors is usable to further train the generative model to generate a synthetic dataset. The further request may be indicative of a model identifier associated with the trained model. Tire further request may include performance data indicative of a performance of the generative model trained using the set of latent vectors. Updating the configuration settings may comprises determining, using the performance data, updated configuration settings. The performance data may comprise a function value of a function associated with the training of the generative model, and a latent vector identifier associated with a latent vector used in training the generative model.

[0017] The method may further comprise receiving a request for one or more latent vectors. The method may further comprise, based on the request, identifying configuration settings for the boson sampler. The method may further comprise operating the boson sampler to produce one or more samples, the boson sampler configured in accordance with the identified configuration settings. The method may further comprise determining, from the one or more samples, the one or more latent vectors. The method may further comprise sending the determined one or more latent vectors, wherein the one or more latent vectors are usable with the trained generative model to generate a synthetic dataset.

[0018] According to an aspect of the present disclosure a system is provided. The system comprises a boson sampler and a set of one or more processors. The system is configured to receive a request (or further request as may be appropriate) for a set of latent vectors for use with training a generative model to generate a synthetic dataset. The system is further configured to, based at least in part on the request, select configuration settings (or updated configuration settings as may be appropriate) for the boson sampler. The system is further configured to operate the boson sampler to produce a batch of samples (or further batch of samples as may be appropriate), the boson sampler configured in accordance with the selected configuration settings. Each sample may be representative of a measurement outcome of one or more photodetectors of the boson sampler. The system is further configured to determine, from the batch of samples, the set of latent vectors (or further set of latent vectors as may be appropriate). The system is further configured to send the set of latent vectors, for example to an electronic system for training a generative model. For model inference, the system may be configured to receive a request for one or more latent vectors, identify configuration settings for the boson sampler, operate the boson sampler in accordance with tire configuration settings to produce a one or more samples, determine one or more latent vectors from the one or more samples, and communicate the one or more latent vectors, wherein the one or more latent vectors are usable with the trained generative model to generate a sy nthetic dataset. The boson sampler may be a single-photon boson sampler or may be a Gaussian boson sampler.

[0019] According to an aspect of the present disclosure, a method is provided. The method is suitable for performance by a first system (e.g. an electronic system, for example an electronic device) and a second system (e.g. a hybrid system or heterogeneous sy stem comprising a set of one or more processors and a boson sampler). The method comprises, at the second system, selecting configuration settings for a boson sampler. The method further comprises, at the second system, operating the boson sampler to produce a batch of samples, the boson sampler configured in accordance with the selected configuration settings. The method further comprises, at the second system, determining, from the batch of samples, a set of latent vectors. The method further comprises, at the second system, training a generative model to generate a synthetic dataset using the set of latent vectors. The method further comprises outputting the trained generative mode, for example by communicating the trained model from the second system to the first system or providing the trained model for installation in tire first system. The method may further comprise locally erasing at least a part of the trained model.

[0020] In this example, one system (the second system) is responsible for training the generative model and then the generative model is made available to the first system. This method is suitable, for example, when the same model is to be deployed to multiple first systems, or where the computation resources for training the model are not powerful enough at the first system. For example, this method may be appropriate for e.g. mobile AI in which trained machine learning models are stored in user devices such as smartphones.

[0021] According to an aspect of the present disclosure, a method is provided. The method is suitable for performance by a hybrid s> stem or heterogeneous system comprising a boson sampler and a set of one or more processors. The method comprises selecting configuration settings for a boson sampler. The method further comprises operating the boson sampler to produce a batch of samples, the boson sampler configured in accordance with the selected configuration settings. The method further comprises determining, from the batch of samples, a set of latent vectors. The method further comprises using the set of latent vectors, training a generative model to generate a synthetic dataset. The method further comprises outputting the trained generative model.

[0022] According to an aspect of the present disclosure, a system is provided. The system comprises a boson sampler and a set of one or more processors. The system is configured to select configuration settings for a boson sampler. The system is further configured to operate the boson sampler to produce a batch of samples, the boson sampler configured in accordance with the selected configuration settings. The system is further configured to determine, from the batch of samples, a set of latent vectors. The system is further configured to use the set of latent vectors to train a generative model to generate a synthetic dataset. The system is further configmed to output tire trained generative model. Training the model may include training the configuration settings. For model inference, the system may be configmed to receive a request for one or more latent vectors, identify configuration settings for the boson sampler, operate the boson sampler in accordance with the configuration settings to produce a one or more samples, determine one or more latent vectors from the one or more samples, and communicate the one or more latent vectors, wherein the one or more latent vectors are usable with the trained generative model to generate a synthetic dataset. The boson sampler may be a single-photon boson sampler or may be a Gaussian boson sampler.

[0023] The methods and systems described herein are applicable to any suitable architecture for a generative model. For example, the generative model may comprise an artificial neural network (ANN), a diffusion model, or a variational autoencoder.

[0024] The output of the boson sampler is unrelated to tire choice of training data for training the generative model. Accordingly, the methods and systems described herein are applicable irrespective of the nature of the synthetic dataset that the generative model is trained to generate. For example, the generative model may be trained to generate a synthetic image. The descriptions herein are also applicable to generating other types of synthetic datasets, such as synthetic videos, synthetic 3D shapes, synthetic text, sy nthelic molecule geometries or formulae, and synthetic time series. Examples of synthetic time series include sound, financial time series, weather-related time series such as wind, clouds, or temperature, time series of energy production in a grid, or time series of sensor data (such as speed or acceleration in a vehicle). Other output datasets may comprise synthetic graphs such as social networks or transportation networks. An output dataset may also comprise conditional data, for example images conditioned on a text input, or images conditioned on other images. An example of an image conditioned on another image is a high-resolution image conditioned on a low-resolution image. An output dataset may also comprise any combination of the above, such as joint images and text caption.

[0025] The methods and systems described herein are applicable to any suitable training routine of the generative model. For example, training a generative model may comprise training an artificial neural network (ANN). For example, training the generative model may comprise training a generative adversarial network (GAN), the GAN comprising the ANN and a second ANN, wherein training the GAN may comprise: training the ANN, using the determined set of latent vectors and feedback from the second ANN to generate a synthetic dataset; training the second ANN, using a plurality of genuine datasets and a plurality of synthetic datasets generated by the ANN. to classify received datasets as synthetic datasets or genuine datasets, and to provide feedback to the ANN; and outputting the trained ANN configured to generate synthetic datasets. For example, training the generative model may comprise training an ANN as part of a hybrid neural network (HNN) comprising a boson sampling layer and the ANN.

[0026] The configuration settings of the boson sampler may comprise parameter values for parameters of a configurable interferometer of the boson sampler. For example, the coirfiguration settings may describe effective reflection coefficients of one or more reconfigurable beamsplitters of the interferometer, or phase shifts imparted by one or more phase shifters of the interferometer. In some examples, the configuration settings, may indicate an input multimodal photonic state provided to an interferometer of the boson sampler, for example an indication as to which input modes should contain a single photon of squeezed light input.

[0027] The methods and systems described herein are compatible with other security features. For example, each request for a set of latent vectors may be indicative of at least one of: a device identifier indicating a device from which the request is received; or a user identifier indicating a user of the device from which the request is received. This additional information may be used by the heterogeneous system to determine that a requesting device (or user of that device) is authorised to receive latent vectors.

[0028] According to an aspect of the present disclosure, a non-transitoiy computer readable medium is provided. The non-transitoiy computer-readable medium comprises stored instructions that, when executed by one or more processors coupled to a boson sampler, cause the one or more processors to perform a method for a heterogeneous system as described herein.

[0029] According to an aspect of tire present disclosure, a non-transitory computer readable medium is provided. The non-transitoiy computer-readable medium comprises stored instructions that, when executed by one or more processors of an electronic system, cause the one or more processors to perform a method for an electronic system as described herein. Brief description of the Drawings

[0030] Illustrative embodiments of the present disclosure will now be described by way of example only, with reference to the accompanying figures.

[0031] Fig. 1 shows an illustration of a trained artificial neural network according to an example.

[0032] Fig. 2 shows a block diagram of a heterogeneous system according to an example.

[0033] Fig. 3 shows an illustration of a boson sampler according to an example.

[0034] Fig. 4 shows an illustration of a boson sampler according to an example.

[0035] Fig. 5 shows a swim lane flowchart of a method for training a generative model according to an example.

[0036] Fig. 6 shows an illustration of a generative adversarial network.

[0037] Fig. 7 shows an illustration of a hybrid neural network comprising a boson sampling layer and an artificial neural network according to an example.

[0038] Fig. 8 shows a swim lane flow chart of a method for training configuration settings of a boson sampler and training a generative model according to an example.

[0039] Fig. 9 shows a swim lane flowchart of a method for training a generative model according to an example.

[0040] Fig. 10 shows a swim lane flowchart of an inference method for using a trained generative model according to an example.

[0041] Fig. 11 shows a block diagram of an electronic device according to an example.

[0042] Fig. 12 shows a graph indicating the performance of a trained generative model using a variety’ of latent spaces.

[0043] Fig. 13 shows a second graph indicating the performance of a trained generative model using a variety of latent spaces.

[0044] Throughout the description and the drawings, like reference numerals refer to like parts. Detailed Description

[0045] Embodiments of the disclosure are described with reference to the accompanying drawings. How'ever, it should be appreciated that the disclosure is not limited to the embodiments, and all changes and / or equivalents or replacements thereto also belong to the scope of the disclosure. The same or similar reference denotations may be used to refer to the same or similar elements throughout the specification and the drawings.

[0046] As used herein, the terms “have”, “may have”, “include”, or “may include” a feature (e.g. a number, function, operation, or a component such as a part) indicate the existence of the feature and do not exclude the existence of other features. Throughout the description and claims of this specification, the words “comprise” and “contain” and variations of them mean “including but not limited to”, and they are not intended to (and do not) exclude other components, integers or steps. Throughout the description and claims of this specification, the singular encompasses the plural unless the context otherwise requires. In particular, where the indefinite article is used, the specification is to be understood as contemplating plurality’ as well as singularity, unless the context requires otherwise.

[0047] As used herein, the terms “A or B”, “at least one of A and / or B”, or “one or more of A and / or B” may include all possible combinations of A and B. For example, “A or B”, “at least one of A or B”, “at least one of A and B” may indicate all of (1) including at least one A, (2) including at least one B, or (3) including at least one A and at least one B.

[0048] As used herein, the terms “first” and “second” may modify various components regardless of importance and do not limit the components. These terms are only used to distinguish one component from another. For example, reference to a first component and a second component may indicate different components from each other regardless of the order or importance of the components.

[0049] It will be understood that when an element (e.g. a first element) is referred to as being (physically, operatively or communicatively) “coupled with / to”, or “connected with / to” another element (e.g. a second element), it can be coupled with / to the other element directly or via a third element. In contrast, it will be understood that when an element (e.g. a first element) is referred to as being “directly coupled with / to” or “directly connected with / to” another element (e.g. a second element), no element (e.g. a third element) intervenes between the element and the other element.

[0050] The terms as used herein are provided merely to describe some embodiments thereof, but not to limit the scope of other embodiments of the disclosure. It is to be understood that the singular forms “a”, “an”, and “the” include plural references unless the context clearly dictates otherwise. All terms including technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the embodiments of the disclosure belong. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and will not be interpreted in an idealised or overly formal sense unless expressly so defined herein.

[0051] Generative models are machine learning models trained to generate synthetic datasets (output data) that are similar to the genuine datasets (training data) on which they are trained. At a high level, a generative model works by capturing the underlying patterns or distributions from a set of training data. Once these patterns are learned, the model can then generate new data that shares similar characteristics with tire original training data. For example a generative model may be trained to generate artificial images that are similar in style to the genuine images on which the model is trained.

[0052] For ease of explanation, the following description refers often to generative artificial neural networks, also referred to herein as generator networks, in particular as examples of generative models. However, this is merely for convenience. The methods and systems described herein are also applicable for other forms of generative model including but not limited to generative adversarial networks, variational autoencoders, or diffusion models.

[0053] In many of the examples provided herein, the description refers to the specific example of generating images. However, this is merely for convenience. The descriptions herein are also applicable to generating other types of synthetic datasets, such as synthetic videos, synthetic 3D shapes, synthetic text, synthetic molecule geometries or formulae, and synthetic time series. Examples of synthetic time series include sound, financial time series, weather-related time series such as wind, clouds, or temperature, time series of energy production in a grid, or time series of sensor data (such as speed or acceleration in a vehicle). Other output datasets may comprise synthetic graphs such as social networks or transportation networks. An output dataset may also comprise conditional data, for example images conditioned on a text input, or images conditioned on other images. An example of an image conditioned on another image is a high-resolution image conditioned on a low-resolution image. An output dataset may also comprise any combination of the above, such as joint images and text caption. The words “synthetic” and “artificial” are used interchangeably herein to indicate the generated output of the generative model.

[0054] As previously indicated, training a generative model may require vast amounts of training data, large financial resources, or great computational power. Accordingly, there is interest in securing such trained models and the training data on which they are trained. Most attention is usually given to securing the training data, which is typically not made available to outside parties. However, the task of protecting the models themselves is more difficult, as they may be exposed to at least an extent to outsiders via user interfaces or APIs.

[0055] Deployed machine learning models (that is, trained models with which outsiders may interact) may be subject to several attacks that are different to other types of cyber attacks. For example, machine learning models may be subject to so-called model inversion attacks, in which an adversaiy attempts to extract sensitive information about the training data used to train a model. In particular, a generative model may encode some information about the training data in its trained parameters and output behaviour. By queiying the trained model and analysing behaviour of the model’s outputs, an adversary may be able to infer information about training data or even reconstruct some training data. Such attacks may be mitigated to some extent using distinctive private training, but this is difficult to implement in practice and may reduce performance of the model or lead to unstable training.

[0056] Deployed machine learning models may also be subject to so-called model extraction attacks, in which an adversary seeks to produce their own copy of the model. In blackbox attacks, in which the adversaiy’ does not have full knowledge of the trained model, the adversary may attempt to use queries against the deployed model to generate outputs and then infer the model’s architecture and trained parameters. It may be possible for an adversary to reproduce the deployed model using less data than was required to train the model in the first place. In whitebox attacks, the adversary is assumed to have full knowledge of the model architecture and the trained parameters and so can trivially recreate the trained model. Blackbox model extraction attacks are typically mitigated to an extent by embedding a signature or watermark into the model parameters which may be revealed by inspecting the model or queiying its parameters. A whitebox attacker may be able to remove such signatures or watermarks.

[0057] When a trained model is deployed via e.g. cloud services over the internet, the owner of the model may have some success in securing the model by storing the model (e.g. descriptions of the model architecture and the parameters thereof) in secure servers, making the model accessible to outsiders only via an API or other interface. However, the model may still be vulnerable to data leaks or theft. Furthermore, in recent years there has been a trend towards mobile AI, in which trained machine learning models are stored and operated on user electronic devices, leading to lower latencies and a better user experience. Mobile AI places trained models on devices in the hands or pockets of users, which may include adversaries.

[0058] The methods and systems described in the present disclosure may mitigate the effects of attacks such as model inference attacks and model extraction attacks on a generative model, even in the event that an adversaiy' lias full knowledge of the architecture and trained parameters of tire trained generative model. Generative models rely on sampling from probability distributions to produce synthetic datasets. These samples provide the generative model with the randomness that may be required for the generative model to function correctly, for example to produce a synthetic dataset that is different from any of the genuine datasets on which the model was trained. The inventors have recognised the importance of the latent space (e.g. the range of all possible random noisy samples) that a generative model uses to generate synthetic datasets in helping to secure trained models, and have recognised that even with full knowledge of the architecture and model parameters of a trained generative model, if the adversary’ is unable to produce latent vectors from the latent space with which the generative model was trained, then a cloned model will give a much degraded performance in comparison to the original trained model. This is discussed further below in relation to Fig. 12 and Fig. 13.

[0059] Traditionally, each latent vector component value input to a generative model is independently selected from a probability distribution such as a normalised Gaussian distribution or a uniform distribution. For example, in the traditional approach each input to a generative model may be independently sampled from a Gaussian distribution having a known mean and covariance. However, such independent latent spaces are often suboptimal when generating realistic synthetic datasets (e.g. images). Furthermore, an adversary' with knowledge of the trained model architecture and all trained parameters may be able to guess or closely approximate the Gaussian distribution with which the model was trained, and accordingly when using the cloned model for inference may produce synthetic datasets with a high level of performance.

[0060] The inventors have recognised that a non-universal quantum processor known as a boson sampler can help with generating latent vectors from a structured latent space that cannot be replicated without access to tire boson sampler and knowledge of the configuration settings associated with the latent space with which the model was trained. As will be discussed further below in relation to Fig. 12 and Fig. 13, attempts to use a clone of a trained model with a different latent space result in a degraded performance - in other words the synthetic datasets generated will not reliably share characteristics with the training data on which the model was trained.

[0061] A boson sampler is a non-universal quantum computer that relies on the interference of identical photons to generate its output. More particularly, a boson sampler comprises a network of optical components or elements (an interferometer) in which photons interfere with one another. As photons are quantum objects, the output of tins network is described by a quantum superposition of all the possible outcomes. When a measurement is performed at the output of this network using one or more photodetectors, a single measurement outcome is realised from thi s superposition. For example, if photon number resolving (PNR) detectors are used, then each sample or measurement outcome may be described by an array or string or sequence of integers indicating how many photons were found in each output mode of the output state; if threshold detectors are used, then each measurement outcome may be described by an array or string or sequence of integers, for example a binary sequence, indicating whether photons were present or absent in each output mode of the output state. By repeatedly sampling measurement outcomes, one can build up knowledge of the probability' distribution governing the quantum superposition.

[0062] The photonic superposition states output from an interferometer of a boson sampler can be highly entangled. Accordingly, the output probability distributions generated by a boson sampler may have a complex structure and simulating this sampling task is understood to be intractable classically. Modern supercomputers fail to simulate boson sampler distributions generated from more than a few tens of modes.

[0063] Advantageously, the training and inference methods described herein ensure that a trained generative model accessible to users is separated from the configuration settings and boson sampler used to produce the latent vectors with which the model is trained. Even if an adversary is able to obtain full knowledge of tire architecture and trained parameters of the trained generative model, they are unable to get the same level of performance out of the cloned model. Accordingly, the effect of a model extraction attack is mitigated. Furthermore, due to tire degraded performance of the cloned model, the difficulty of probing the cloned model to infer information about the training data is greatly magnified. Accordingly, the effect of a model inversion attack is mitigated. Furthermore, due to inherent complexity in a boson sampler distribution, even if an adversary successfully obtains the configuration settings of the boson sampler, they would still be unable to recreate or accurately simulate the latent space with which the trained model is trained without access to an appropriate boson sampler.

[0064] An example of a generative model will now be described with reference to Fig. 1, which illustrates an artificial neural network according to an example.

[0065] Generally speaking, an artificial neural network (ANN) is a method of function approximation loosely modelled on an animal brain, and comprising a plurality7 of nodes known as neurons, a plurality of connections between the nodes, and a plurality7 of weights and biases associated with the neurons and neuron-to-neuron connections therebetween. Each neuron is configured to receive one or more inputs and to provide those one or more inputs as weighted argument(s) to a non-linear transfer function that provides the neuron’s output. The transfer function is sometimes known as an activation function. The weightings of the inputs of the activation function are defined by the weights and biases associated with that neuron and its connections. The activation function may be. for example, the sigmoid activation function, the tanh activation function, or the rectified linear (ReLu) activation function.

[0066] The neurons are ty pically arranged in layers, such as visible layers including input and output layers, and hidden layers. The outputs of neurons in the input layer and each hidden layer are provided as inputs to a subsequent layer or layers, and the output layer produces the output of the network. Accordingly, the ANN receives a plurality7 of input values and converts them to a plurality of output values / results.

[0067] Some ANNs can be used to generate synthetic data (sometimes referred to as artificial data) and may be referred to in this document as generator networks. The input values received by the generator network may comprise values selected at random from some probability distribution, and the generator network may produce synthetic data as output values that mimic the genuine datasets on which the ANN is trained. In the terminology of generator networks, an instance of input values may be referred to as a latent vector, and the probability distribution from which a latent vector is selected may be referred to as a latent space.

[0068] A neural network needs to be trained in order to perform a task correctly and may be trained in many different ways. During a training process an ANN learns (or is “trained”) by processing data of a collection of representative examples (genuine datasets) according to a prescribed training routine, forming a probability -weighted distribution between the input values and output values of the ANN. For example, when training a generator network, synthetic data generated by the network may be compared in some way with representative examples from the training data, ty pically by use of a cost function, and the weights and biases of the generator network may be iteratively updated according to a learning rule. Successive adjustments will cause the artificial neural network to produce synthetic data that is increasingly similar to the target output data. After a sufficient number of these adjustments the training can be terminated based upon certain criteria. Once trained, the trained model (for example, the trained weights and biases of the neural network) can be stored for future use.

[0069] An illustration of a trained ANN 100 for generating an image according to an example is shown in Fig. 1. The trained generator network 100 comprises an input layer 102, a hidden layer 104, and an output layer 106. The input layer 102 comprises a first plurality7 of neurons 102-1 to 102-r, the hidden layer 104 comprises a second plurality of neurons 104-1 to 104-v, and the output layer 106 comprises a third plurality of neurons 106-1 to 106-w. While only a single hidden layer is shown in Fig. 1. the skilled person would appreciate that an ANN may have several hidden layers between the input layer 102 and output layer 106. The number of neurons in each layer may be the same or different. Furthermore, while every7 neuron in a layer is connected to every neuron in the next layer in the ANN 100, the skilled person will appreciate that each neuron may be connected to fewer neurons of the next layer, or may be connected to a neuron in the same layer or a preceding layer. The activation function implemented by each neuron may be, for example, the sigmoid activation function, the tanh activation function, or the rectified linear (ReLu) activation function.

[0070] The trained generator network 100 is configured to receive as input a latent vector 108, denoted z in the figure. The latent vector z 108 is from a latent space derived from a boson sampler. The neurons 102-1 to 102-r of the input layer 102 are each configured to take tire magnitude of a component (zx to zr) of the latent vector z 108 as input, to provide that magnitude as an argument to the neuron’s activation function, and to output the result to neurons of the hidden layer 104. The neurons 104-1 to 104-v of the hidden layer 104 are each configured to receive inputs from the preceding layer (in this example the input layer 102), to provide those inputs as weighted arguments to the neuron’s activation function, and to output the result to neurons of the output layer 106. The neurons 106-1 to 106-w of the output layer are each configured to receive inputs from the preceding layer (in this example the hidden layer 104). to provide those inputs as weighted arguments to the neuron’s activation function, and to output the result.

[0071] The output values yr to yw together comprise a synthetic dataset. In the example shown in Fig. 1 .the synthetic dataset comprises a synthetic image 110. For example, the output values y± to yw of the neurons of the output layer may comprise pixel values. In some examples, the image 110 may be a grayscale image and each output value yj of the output layer 106 may correspond to a respective pixel of the image 110. In some examples, the image 110 may be a colour image and, for example, each pixel of the image 110 may be represented by three output values of the neural network 100, one for each of the red, green, and blue channels of the pixel. Accordingly, the ANN 100 is configured to receive as input a latent vector z (108) and to generate an artificial / synthetic image (110).

[0072] The skilled person will appreciate that neural network architectures different to that of Fig. 1 are compatible with the methods and systems described herein. For example an ANN may comprise one or more convolution layers, one or more max-pool layers, and / or a soft-max layer, and may include skip connections. An ANN may take other additional input besides latent vectors in order to condition the output of the ANN. For example, an ANN for modifying images may take as input a latent vector and an image for modification and output the modified image.

[0073] The ANN 100 may have been trained by attempting to optimise a cost function indicative of the error between synthetic images (which are examples of synthetic datasets) generated by the ANN to a training set of genuine images (which are examples of genuine datasets). For example, training may comprise minimising a cost function such as a quadratic cost function, a cross-entropy cross function, a log-likelihood cost function. The minimisation may be performed for example by gradient descent, stochastic gradient descent or variations thereof, using backpropagation to adjust weights and biases within the neural network accordingly. Training may involve the use of further techniques known to the skilled person, such as regularization. Mini-batch sizes, learning rates, numbers of epochs and other hyperparameters may be selected and fine-tuned during training. The ANN 100 may be trained as part of a generative adversarial network for example.

[0074] Fig. 2 depicts a block diagram of a heterogeneous system 200 in which illustrative embodiments may be implemented. The heterogeneous system 200 comprises both classical processing apparatus and quantum processing apparatus. Other architectures to that shown in Fig. 2 may be used as will be appreciated by the skilled person. For example, system 200 may be distributed across multiple interconnected devices.

[0075] System 200 is an example of a specialised computing apparatus, in which computer usable program code or instructions implementing the processes may be located. In this example, system 200 includes communications fabric 202, which provides communications between a processor unit 204, memory unit 206, input / output unit 208, communications module 210, display 212, and boson sampler 214, the boson sampler comprising a state generation unit 216, an interferometer 218, a state detection unit 220 and a dedicated controller unit 222.

[0076] The system 200 may be implemented in any of a number of ways. For example, the system 200 may be provided as a number of hardware modules suitable for installation in a server / computer rack (for example a conventional 19-inch server rack). For example, the processor unit 204, memory unit 206, input / output unit 208, and communications module 210 may be provided in a first rack-mounted hardware module, the controller 222 may be implemented in a second rack-mounted hardware module and electronically coupled to the first hardware module, the state generation unit 216 may be implemented in a third rack-mounted hardware module electronically coupled to the controller 222. the interferometer 218 may be implemented in a fourth rack-mounted hardware module electronically coupled to the controller 222 and optical fibre-connected to the state generation unit 216, and the photodetectors of the state detection unit 220 may be provided in another hardware module electronically coupled to the controller 222 and optical fibre-connected to the interferometer module and, optionally, to the state generation unit 216. In other examples, the system 200 may be implemented using one or more separate devices communicatively coupled (at least in part) over a network such as the internet.

[0077] The processor unit 204 is configured to execute instructions for software that may be loaded into the memory unit 206. Processor unit 204 may be a set of one or more processors or may be a multi-processor core, depending on the particular implementation. Furthermore, processor unit 204 may be implemented using one or more heterogeneous processor systems in which a main processor is present with secondary processors on a single chip. The processor unit 204 may comprise one or more central processing units (CPUs), one or more graphics processing units (GPUs) or any combination thereof. If the processor unit 240 comprises multiple processors, the multiple processors may operate individually or collectively.

[0078] The memory unit 206 may comprise any piece of hardware that is capable of storing information, such as, for example, data, program code in functional form, and / or other suitable information on a temporary basis and / or a permanent basis. The memory unit 206 may include, for example, a random-access memory or any other suitable volatile or non-volatile storage device. The memory unit 206 may include a form of persistent storage, for example a hard drive, a flash memory, a rewritable optical disk, a rewritable magnetic tape, or some combination thereof. The media used for persistent storage may also be removable. For example, the memory unit 206 may include a removable hard drive.

[0079] Input / Output unit 208 enables the input and output of data with other devices that may be in communication with the system 200. For example, input / output unit 208 may provide a connection for user input through a keyboard, a mouse, and / or other suitable devices. The input / output unit 208 may provide outputs to, for example, a printer.

[0080] Communications module 210 enables communications with other data processing systems or devices. The communications module 210 may provide communications through the use of either or both physical and wireless communications links. For example, the communications module 210 may be configured to communicate with other data processing systems or devices via a wired local area network connection, via WiFi or over a wide area network such as the internet.

[0081] Instructions for the applications and / or programs may be located in the memory unit 206. which is in communication with the processor unit 204 through communications fabric 202. Computer-implementable instructions may be in a functional form on persistent storage in the memory unit 206 and may be performed by processor unit 204. These instructions may sometimes be referred to as program code, computer usable program code, or computer-readable program code that may be read and executed by a processor in processor unit 204. The program code in the different embodiments may be embodied on different physical or tangible computer-readable media.

[0082] The program code may contain instructions which, when processed by the processor unit 204. cause the processor unit 204 to communicate with the boson sampler to sample a bosonic probability distribution.

[0083] In Fig. 2, computer-readable instructions 226 are located in a functional form on computer-readable storage medium 224 that is selectively removable and may be loaded onto or transferred to system 200 for execution by processor unit 204. Alternatively, computer-readable instructions 226 may be transferred to system 200 from computer-readable storage medium 224 through a communications link to communications module 210 and / or through a connection to input / output unit 208. The communications link and / or the connection may be physical or wireless.

[0084] A computer-readable storage medium may be, for example but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, device, or any suitable combination thereof. More specific examples of the computer-readable medium include the following: a portable computer diskette, a hard disk, a random-access memory (RAM), a read-only memory’ (ROM), an erasable programmable read-only memory’ (EPROM or Flash memory'), a portable compact disc read-only memory (CDROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer-readable storage medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0085] In some illustrative embodiments, computer-implementable instructions 226 may be downloaded over a network to the memory' unit 206 from a remote device for use with system 200. For instance, computer-implementable instructions stored in a remote server may be downloaded over a network from the server to the system 200.

[0086] The boson sampler 214 comprises a state generation unit 216, a linear interferometer 218, a state detection unit 220 and a dedicated (classical) control unit 222. Example boson sampler architectures are described further below in relation to Fig. 3 and Fig. 4.

[0087] The state generation unit 216 is configured to generate an input multimodal photonic state [WIN) comprising a plurality’ of input modes. The input multimodal photonic state is a product state (in other words, there is no quantum entanglement between input modes) comprising a plurality of N non-vacuum optical inputs distributed across a plurality of M input modes.

[0088] In some examples, the state generation unit 216 may be configured to generate an input multimodal photonic state \VIN) comprising N photons distributed across the plurality of M modes. When the number of photons N is less than or equal to the number of input modes M, and one photon is provided in each of the populated input modes (in which case the boson sampler may be referred to as a single-photon boson sampler), the input state can without loss of generality’ be expressed as \VIN)sp = |11F 12, -, lw, 0w+1, -, 0M) = aj .-¼ -, 0M> (EQ. 1) where ais the bosonic creation operator in the fcth mode. The skilled person will appreciate that the methods and systems described herein are also applicable when one or more input modes comprise more than one photon.

[0089] In other examples, the state generation unit 216 may be configured to generate an input multimodal photonic state | WIN) comprising a Gaussian photonic input in each of the N input modes (in which case the boson sampler may be referred to as a Gaussian boson sampler). For example, a single mode squeezed state (SMSS), also referred to as a squeezed coherent state, may be input into each input mode.

[0090] In some examples, the state generation unit 216 may be configured to generate an input multimodal photonic state | WIN) that comprises single photons in some modes and squeezed coherent states in other modes.

[0091] The state generation module 216 may comprise one or more light sources. For example, in a singlephoton boson sampler, the state generation module 216 may comprise a non-linear photonic material (such as periodically-poled lithium niobate (PPLN) or potassium titanyl phosphate (KTP)) configured to receive a pump beam from a pump laser and to probabilistically generate pairs of entangled photons, and may further comprise a photodetector configured to detect a photon of the entangled pair, thereby heralding the presence of the other photon of the pair. For example, in a Gaussian boson sampler, the state generation module 216 may comprise PPLN waveguides configured to generate two entangled modes of light, and a 50:50 beamsplitter for interfering the two modes of light, thereby generating two independent single mode squeezed Gaussian states.

[0092] The interferometer 218 comprises a plurality of optical elements arranged to interfere the modes of the input multimodal photonic state, thereby transforming the input multimodal photonic state to produce an output multimodal photonic state. The interferometer 218 is configured to receive the input multimodal photonic state, to transform the input multimodal photonic state to an output multimodal state, and to output the output multimodal photonic state to the state detection unit 220. The transformation is dependent on the values {0} of a set of parameters 6 One or more of the parameters 0 may characterise a single mode operation. For example, a parameter may characterise the phase shift imparted by a phase shifter of a passive linear interferometer. One or more of the parameters 0 may characterise a multimodal operation. For example, a parameter may characterise a transmission (or equivalently, a reflection) coefficient of a reconfigurable beam splitter in a passive linear interferometer. If the values {0} of one or more of the parameters 0 may be reconfigured, then the boson sampler is said to be a reconfigurable boson sampler.

[0093] The interferometer 218 may be designed and manufactured in any suitable and desired way e.g. depending on the modes of electromagnetic radiation to be transformed by the interferometer 218. Thus, for example, when the electromagnetic radiation lias an optical or infrared wavelength (e.g. between 400nm and 700nm or between 700nm and 1600nm). the optical paths through the interferometer 218 may be implemented at least partially using optical fibres. In some examples, the interferometer 218 may be implemented in bulk optics. However, in other examples, the interferometer 218 may comprise a photonic integrated circuit. In the photonic integrated circuit, the optical paths may be implemented with, for example, a plurality of etched waveguides and plurality of coupling locations arranged in the photonic integrated circuit. At each coupling location, tuneable elements may be arranged (e.g. EOM phase shifters) that are configured to control the coupling interaction between the waveguides. The photonic integrated circuit may be implemented in silicon nitride (Si3N4) or any other suitable material.

[0094] Due to interference between photons in different temporal inodes, in operation the boson sampler 214 transforms the input multimodal photonic state into an output multimodal photonic state that may be expressed as a superposition of the different possible configurations of the photons in the output modes as \VOUT(d)) = 2> c where C is a configuration. is the number of bosons in the / th output mode in configuration C, and ac is the probability amplitude associated with configuration C. The skilled person would appreciate that while the state of (EQ. 2) is expressed as a pure state, this is for illustrative purposes only - photon loss may, for example, mean that the output state can be expressed only as a mixed state. By tuning the parameter values {0}, the probability amplitudes associated with each configuration may be changed. Accordingly, a measurement of the number of photons in each output mode can yield a measurement outcome representable as a string of integers corresponding to a configuration C. By operating the boson sampler 214 a number of times to produce a batch of samples S, it is possible to establish an empirical probability' distribution of the bosonic configurations of the output state. One can expect that with many samples, the probability pc of obtaining a measurement outcome corresponding to configuration C is approximately given by pc = | ac |2.

[0095] The state detection unit 220 comprises an arrangement of one or more photodetectors configured to detect photons output from the interferometer 218 and produce corresponding detection event signals. In some examples, the photodetectors may comprise photon number resolving (PNR) detectors, capable of determining how many photons are received. For example, the detectors may comprise superconducting nanowire detectors that generate an output signal intensity proportional to the (discrete) number of photons that strike a detector. The PNR detectors may comprise transition edge sensors (TESs). In other examples, the photodetectors may comprise threshold detectors, also known as on / off detectors. Threshold detectors are not capable of determining how many photons are received but are capable of determining the presence / absence of photons in an output mode.

[0096] The controller 222 is communicatively coupled to the processor unit 204, the state generation unit 216. the interferometer 218 and the state detection unit 220. The controller 222 may be any suitable classical computing resource for controlling the operation of the boson sampler 214. In some examples, the controller 222 is implemented in a dedicated, application-specific processing unit. For example, the controller 222 may comprise an application-specific integrated circuit (ASIC) or an application-specific standard product (ASSP) or another domain-specific architecture (DSA). Alternatively, the controller 222 may be implemented in adaptive computing hardware (in other words, hardware comprising configurable hardware blocks / configurable logic blocks) that has been configured to perform the functions, for example in a configured field programmable gate array (FPGA). The controller may have a dedicated random-access memory or other memory element for temporarily logging data. In some examples, the functionality of the controller 222 may be incorporated into the functionality of the processor unit 204.

[0097] The controller 222 is configured to receive instructions from the processor unit 204. More particularly, tire controller is configured to. if so directed by the processor unit 204, configure the interferometer 218 according to a set of parameter values {0} and thereby control the transformation of the input multimodal photonic state that is implemented by the interferometer 218. For example, the controller 222 may directly send control signals that tune the reflectivity / transmittance of a reconfigurable beam splitter or the phase imparted by a phase shifter. The controller 222 is further configured to, if so directed by the processor unit 204, generate one or more control signals to cause the state generation unit 216 to produce an input multimodal photonic state. The controller 222 may optionally be able to control which input multimodal photonic state is input into the boson sampler, for example by generating one or more control signals to control a number of photons in each input mode. For example, in a photonic boson sampler in which the state generation module comprises a plurality of single photon sources, the controller 222 may be able to generate one or more control signals to cause a selected number of photons to be emitted at a particular time point.

[0098] The controller 222 is further configured to receive a response from the state detection unit 220. More particularly, the controller 222 is configured to receive samples or measurement outcomes from the photodetectors of the state detection unit 220. In other words, the controller is configured to sample from the output distribution of the configured boson sampler. For example, the measurement outcomes may comprise an electrical signal from each photodetector at which a detection event occurs.

[0099] In examples wherein the photodetectors are PNR detectors, the electrical signals may be indicative of the number of photons received. The controller 222 may interpret the number of photons detected in each output mode of the output multimodal photonic state as a sequence or string of integers and provide this integer sequence to the processor unit 204. For example, a sample may comprise an integer string in which each integer corresponds to a detected number of photons in a corresponding output mode.

[0100] In examples in which the photodetectors are threshold detectors, the controller 222 may interpret samples as binary strings, with each element indicative of the presence or absence of a detected photon in an output mode of the output multimodal photonic state. For example, an element of a binary’ sequence may have a value of one if one or more photons are detected in the corresponding output mode, while an element of the binary’ sequence may have a value of zero if no photons are detected in the corresponding output mode.

[0101] The controller 222 is further configured to communicate the response from the state detection unit 220 to the processor unit 204.

[0102] In operation, the boson sampler 214 is configured to receive parameter values from the processor unit 204; to produce a batch (in other words, a plurality) of samples by configuring the interferometer 218 according to the received parameter values, generating an input photonic state and measuring an output photonic state; and to communicate the batch of samples to the processor unit 204.

[0103] The processor unit 204 is further configured to determine, from the batch of samples, a set of latent vectors. The processor 204 may determine the set of latent vectors from the batch of samples by performing one or more post-processing operations on the integer sequences. As an example, the processor unit 204 may truncate a received integer sequence to a size compatible with the design of the generative model to be trained. As an example, the processor unit 204 may add or remove an offset value from each element of an integer sequence, for example to ensure that the sequence has a predefined average value (such as zero) which may be preferable for training some ANNs. As another example, the processor unit 204 may convert an integer sequence to a binary sequence, for example by assigning all non-zero elements of the integer sequence the value 1, or for example by assigning all even integers the value 0 and all odd integers the value 1 (or vice versa). In some examples, each sample may itself be considered as a latent vector and the processor unit 204 may perform no post-processing of the sample.

[0104] The skilled person would appreciate that the architecture described above in relation to Fig. 2 is not intended to provide limitations on the computing devices with which the methods described herein may be implemented. Instead, the skilled person would appreciate that other architectures may be applicable. For example, tire computing device may include more or fewer components.

[0105] The boson sampler 214 of Fig. 2 may comprise a spatial mode interferometer, such as the single-photon boson sampler 214a illustrated in Fig. 3. In the boson sampler 214a of Fig. 3. the modes of the input multimodal photonic state are spatial modes - that is, the state is defined by the number of photons in each of a plurality of spatially distinct paths. The boson sampler 214a may be implemented, at least in part, in a photonic integrated circuit.

[0106] The state generation unit 216a of Fig. 3 comprises a plurality of single-photon sources 310 configured to produce single photons. One suitable photon source technology is spontaneous parametric down-conversion (SPDC). In SPDC, a non-linear crystal is pumped with a laser and, probabilistically, entangled photons are emitted (the “signal” and the “idler”). A photodetector (not shown in Fig. 3) is arranged to detect the presence of the idler photon which, due to the entanglement, heralds the presence of a photon in the signal mode. Other photon sources may also be used, for example solid state photon sources and quantum dots.

[0107] The number of single photon sources 310 may be greater than the number M of input modes of the input multimodal photonic state l^IN) in order to account for the fact that single photons may be generated only probabilistically. The state generation unit 216a of Fig. 3 comprises a multiplexer 320 to route successfully generated single photons to N input ports of the M input ports of the interferometer 218a. In the example shown in Fig. 3, the number of single photons N is equal to the number of input ports M of the interferometer 218a.

[0108] The interferometer 218a comprises M input ports, M output ports, and a plurality of waveguides arranged to pass through the interferometer 218a to connect the M input ports to the M output ports. The plurality of waveguides are arranged to provide a plurality of coupling locations between pairs of the plurality of waveguides. The interferometer 218a may be designed and manufactured in any suitable and desired way e.g. depending on the modes of electromagnetic radiation to be transformed by the interferometer. Thus, for example, when the electromagnetic radiation has an optical or infrared wavelength (e.g. between 400nm and 700nm or between 700nm and 1600mn). the waveguides may comprise optical fibres. In some examples, the interferometer may be implemented in bulk optics. However, in other examples the interferometer comprises a photonic integrated circuit, with the plurality of waveguides and plurality of coupling locations arranged in the photonic integrated circuit. The photonic integrated circuit may be implemented in silicon nitride (Si3N4) or any other suitable material, for example thin-fihn lithium niobate.

[0109] A reconfigurable beam splitter 330 is arranged at each of the coupling locations such that at each coupling location the two modes of electromagnetic radiation carried by tire two respective waveguides are capable of coupling with each other with a reconfigurable reflection coefficient (transmission coefficient). The reflection (transmission) coefficient of each reconfigurable beam splitter is denoted with a theta in the figure.

[0110] A parametrised / reconfigurable beam splitter is understood to mean any tuneable element or device or tuneable collection of elements / devices capable of coupling two modes of electromagnetic radiation with each other with a rcconfigurablc reflection / transmission coefficient and optionally a reconfigurable phase shift coefficient (not indicated in Fig. 3). The parametrised beam splitters may be implemented in any suitable way -for example a parametrised beam splitter may comprise a Mach-Zehnder ty pe interferometer containing a variable phase shifter in one internal path for controlling the effective beam splitter reflection coefficient of the Mach-Zehnder interferometer. The Mach-Zehnder interferometer may further comprise an external phase shifter on one external path of the Mach-Zehnder interferometer to control the relative phases of the two mode acted upon. For example, when the interferometer 218a is implemented in an integrated circuit, a reconfigurable beamsplitter may comprise a first waveguide coupling region for coupling the electromagnetic radiation modes in each waveguide, an electro-optical phase shifting element for adjusting the phase in one of the outgoing waveguides from that coupling region, and a second waveguide coupling region for recoupling the two electromagnetic modes output from the first waveguide coupler. For example, when implemented in bulk or fibre optics, a reconfigurable beamsplitter may comprise two 50 / 50 beamsplitters and a phase shifter element arranged therebetween.

[0111] The interferometer 218a may further comprise reflective elements (e.g. mirrors) and other passive photonic elements (not shown). Accordingly, the interferometer 218a couples tire single photons received at the M input ports to the plurality of M output ports based on operations defined by a set of parameter values.

[0112] The interferometer 218a of Fig. 3 is suitable for transforming an input multimodal photonic state comprising M input spatial modes to an output multimodal photonic state comprising M output spatial modes. The skilled person will appreciate that other architectures for the interferometer 218a may be utilised. Of course, while in the illustration the number of input and output modes is M = 4, an interferometer 218a may be provided to operate on a greater number of spatial modes. Of course, the interferometer 218a may comprise any number of reconfigurable / parametrised elements and in any configuration that leads to interference between spatial modes.

[0113] The state detection unit 220a comprises a plurality of photon number resolving (PNR) photodetectors 340, each arranged to receive any photons output from a corresponding output port of the interferometer 218a. The state detection unit 220a comprises one PNR detector for each of the M output modes and accordingly the measurement outcomes are representative of the number of photons measured in all output modes of the output multimodal photonic state. The PNR detectors may comprise nanowire photodetectors.

[0114] The controller 222a is coupled to each of the state generation unit 216a, the interferometer 218a and the state detection unit 220a. The controller 222a is further communicatively coupled to the processor units 204. The controller 222a may receive a set of parameter values from the processor unit 204, and may generate control signals to configure the tuneable elements 330 of the interferometer 218a in accordance with those parameter values. For example, each reconfigurable beam splitter 330 may comprise a Mach-Zehnder interferometer comprising two 50 / 50 beam splitters and a phase shifter located in each of one or both of its internal optical paths. The phase shifter may be implemented using an electro-optical modulator. The control signals may comprise an electric field for controlling the phase shift imparted by the internal phase shifters and therefor the coupling strength of the reconfigurable beam splitter. The controller 222a may further generate a control signal to cause the single photon sources 310 to begin generating single photons, for example the control signal may cause a pump laser to pump light into the non-linear material of the single-photon sources 310. The controller 222a may further receive signals from each of the PNR detectors 340 indicative of the number of photons detected at each of the PNR detectors 340, which may be interpreted as a sample of the output probability distribution produced by the boson sampler 214a. The controller 222a may then communicate the sample information to the processor unit 204.

[0115] The boson sampler 214 of Fig. 2 may comprise a temporal mode interferometer, such as the singlephoton boson sampler 214b illustrated in Fig. 4. In the boson sampler 214b of Fig. 4. the modes of the input multimodal photonic state are temporal modes, which means that the state is defined by the number of photons in each of a plurality of temporal modes or time bins.

[0116] The state generation unit 216b of Fig. 4 comprises a single-photon source 410 operable to produce a single photon in each of a plurality of time bins, so that each photon enters the time-bin interferometer 218b separated from the next by a duration t. As in the boson sampler 214a of Fig. 3, the state generation unit 216a may comprise further single photon sources and a multiplexer in order to reliably ensure that a single photon is generated in each time period t.

[0117] The interferometer 218b comprises a temporal mode coupling device. In particular, in Fig. 4, a temporal mode coupling device comprises a reconfigurable beam splitter 420 and a delay line 430. The delay line 430 is arranged to connect one input port of the reconfigurable beam splitter 420 with one output port of tire reconfigurable beam splitter 420. The delay line may comprise, for example, optical fibre. The delay line 430 Iras a length ct where c is the speed of light in the fibre. In this way, the field of the photon in one temporal mode may be coupled, at least partially, into the delay line 430 so as to interfere with photons in the next temporal mode on the parametrised beam splitter 420. The time-bin interferometer may comprise further optical components including further optical switches.

[0118] The controller 222b is configured to tune the parameter value (e.g. transmittance) of the parametrised beam splitter 420 for each time interval. For example, for four input modes, the temporal mode coupling device can be used to implement the equivalent operations of the three beam splitters defined by parameters 91,92 and 03 shown in Fig. 3. For example, the controller 222b may, as a first photon is emitted from the photon source 410, configure the reconfigurable beamsplitter 420 to couple the first photon into the delay line 430. The controller 222b may then, as a second photon is emitted from the photon source 410, configure the reconfigurable beamsplitter 420 using parameter value to cause interference between the first temporal mode and second temporal mode (e.g. the first and second photon) of the input state |T71V). The controller 222b may then, as a third photon is emitted from the photon source 410, configure the reconfigurable beamsplitter 420 using parameter value 02 to cause interference between the second temporal mode and third temporal mode. The controller 222b may then, as a fourth photon is emitted from the photon source 410, configure the reconfigurable beamsplitter 420 using parameter value 03 to cause interference between the third temporal mode and fourth temporal mode. This may continue until a predetermined transformation has been performed on the input photon sequence of M time bins.

[0119] The state detection unit 220b comprises a photon number resolving (PNR) photodetector 440 configured to detect the number of photons in each temporal mode. By measuring the number of photons in each of M time bins output from the interferometer 218b, the boson sampler 214b takes a sample of the output distribution.

[0120] The skilled person would appreciate that the architecture of tire temporal mode boson sampler 214b of Fig. 4 may be varied in several ways. For example, the boson sampler 214b may comprise further reconfigurable beamsplitters 420 and further delay lines 430 in order to generate more complicated interference between temporal modes. The skilled person would further appreciate that delay lines of different lengths may be used to vary which temporal modes are interfered with one another. In other temporal mode boson samplers, the temporal mode coupling device may comprise a quantum memory’ that may be controlled to selectively interfere photons in different temporal modes.

[0121] The skilled person would appreciate that the spatial mode boson sampler of Fig. 3 and tire temporal mode boson sampler of Fig. 4 may further be operable as Gaussian boson samplers with a suitable substitution of the state generation unit 216a / 216b. Furthermore, the PNR detector(s) of the state detection modules 220a / 220b may be replaced with threshold detectors, in which case the measurement outcomes output from the state detection unit are indicative of the presence or absence of photons in each output mode but not tire number of photons in output modes. The PNR detector(s) may be replaced with pseudo-threshold detectors.

[0122] Fig. 5 shows a swim lane flowchart of a method of training a generative model according to an example. The figure shows actions that may be performed by an electronic system 500 and a hybrid quantum-classical system 550. The electronic system 500 may comprise an electronic device, for example a personal computer, a server, a laptop computer or other such machine. An example electronic device is presented further below in relation to Fig 11. The electronic system 500 may alternatively comprise several coimected devices (for example in a distributed computing environment). The hybrid quantum-classical system 550 comprises classical computing apparatus and boson sampling apparatus. For example, the hybrid quantum-classical system 550 may comprise the heterogeneous system 200 described above in relation to Fig. 2. and the actions performed by the hybrid system 550 may be coordinated by the processor unit 204 of the hybrid system 550..

[0123] The electronic system 500 is configured to securely communicate with the hybrid system 550. For example, the electronic system 500 may be configured to communicate with the hybrid system 550 over a fixed or wired connection such as a LAN cable or similar, or alternatively over a secure wireless connection. For example, the electronic system 500 may be configured to initiate and establish a TLS communication with the hybrid system 550. Communications between the electronic system 500 and hybrid quantum-classical system 550 may be encrypted.

[0124] At 502, the electronic system 500 sends a request for latent vectors to the hybrid system 550 over the communication channel. The latent vectors are to be used by the electronic system 500 in training a generative model to generate a synthetic dataset.

[0125] The request comprises an indication that the electronic system 500 would like to obtain latent vectors. The request may include further information. For example, the request may include a device identifier that identifies the electronic system 500 on which the generative model is to be trained. For example, the request may include a user identifier that identifies a user (pseudonymously or otherwise) of the device on which the generative model is trained. The device identifier or user identifier may be used by the hybrid system 550 in verifying that the requesting electronic system 500 is authorised to receive latent vectors. The request may comprise further information relating to tire requested latent vectors or the generative model to be trained. For example, the request may indicate a number of latent vectors to be sent to the electronic system 500 for training.

[0126] The hybrid system 550 may be configured to analyse the received request to determine that the requesting electronic system 500 or a user thereof is authorised to receive latent vectors for training. For example, the hybrid system 550 may analyse a device identifier or user identifier provided in the request and consult a locally stored record of permitted device identifiers or user identifiers to verify that the electronic system (or user thereof) is authorised.

[0127] Based at least in part on the received request, at 504 the hybrid system 550 selects configuration settings for the boson sampler. The hybrid system 550 may be prompted to select configuration settings based on the fact of the request being received alone. For example, the configuration settings may be selected at random in response to the hybrid system 550 registering that a request has been received. Alternatively the hybrid system 550 may be prompted to select configuration settings based on particular information provided in the request. For example, the hybrid system 550 may be configured to use at least a part of the information provided in the received request to select the configuration settings. For example, if the request comprises a device identifier or user identifier then such identifiers may be used to seed a selection process for selecting the configuration settings.

[0128] The configuration settings describe parameter values {0} of a set of parameters 0 of the interferometer of the boson sampler of the hybrid device 550. For example, the configuration settings may describe the effective transmission or reflection coefficients of reconfigurable beamsplitters of the interferometer. For example, the configuration settings may describe phase shifts to be imparted on light modes in the interferometer.

[0129] The configuration settings may further comprise a selection of an input multimodal photonic state to be generated in the boson sampler. For example, in a single photon boson sampler the configuration settings may indicate in which input modes a single photon is to be provided. For example, in a Gaussian boson sampler, the configuration settings may comprise squeezing parameters for the input state.

[0130] In some embodiments, the parameter values or input multimodal photonic state are selected such that the boson sampler creates an entangled quantum state. Additionally, the parameter values {0} or input multimodal photonic state may be selected such that the statistical properties of the samples output from the boson sampler match or are similar to those of the output data generated by the generative model.

[0131] At 506, the hybrid system operates the boson sampler to produce a batch of samples (measurement outcomes).

[0132] At 508, the method comprises determining, from the batch of samples, a set of latent vectors. Note that a single sample output from the boson sampler may be used to determine a single latent vector.

[0133] At 510, the hybrid system 550 securely communicates the set of latent vectors to the electronic system 500 over a secure channel.

[0134] At 512, the electronic system 500 uses the received latent vectors to train a generative model to generate synthetic datasets. Using the received latent vectors in training may include further processing of the received latent vectors prior to undertaking the training.

[0135] The generative model may be, for example, an artificial neural network, a diffusion model, a variational autoencoder, or any other suitable generative model. The electronic system 500 may use any suitable training data (genuine datasets) at its disposal to train the generative model to produce a suitable synthetic dataset.

[0136] As an example, tire electronic system 550 may train an artificial neural network (ANN) to generate images by training a generative adversarial network (GAN). GANS are a class of deep learning architectures whereby two networks train simultaneously, with a first ANN focused on data generation (the generator network) and a second ANN focused on data discrimination (known as the discriminator or critic). With reference to Fig. 6, the generator network 604 and the discriminator network 610 ‘compete’ against each other. The generator network 604 is trained using a set 602 of latent vectors, each latent vector determined from a sample produced by the boson sampler of the hybrid system 550. The generator network 604 uses latent vectors to generate corresponding artiricial / synlhctic / gcncratcd images 606. The discriminator network 610 receives either a genuine image from a training set 608 of genuine images, or an artificial image from the set 606 generated by the generator network 604, and must distinguish between the two (indicated at 612 in the figure). The generator network 604 is trained to fool the discriminator 610. Feedback from the discriminator 610 is used to train the discriminator 610 until it achieves acceptable accuracy. Feedback from the discriminator network 610 is also used to train the generator network 604 based on whether it fools the discriminator 610. Formally, in this example the game between the generator network 604 and the discriminator 610 may be expressed as one of optimising the minimax objective: minmax E [loeD(x)l + E [loe£>(G(z))] G d xepr(x) ° zepz(z) ° v J where Pr(x) is tire data distribution of the genuine image set x and Pz(z) is the distribution of the latent vectors z produced using the boson sampler of the hybrid system 550. The functions G(z) and D(x) refer respectively to the output of the generator network 604 and discriminator network 610. Further GAN-training techniques may be used to improve the quality" of tire images generated by the generator network 604. Such techniques may include, for example, feature matching, minibatch discrimination, historical averaging, one-sided label smoothing, and virtual batch normalisation.

[0137] Advantageously, in the method of Fig. 5, the training of the generative model is performed on or at the electronic system 500 while the boson sampler and the configuration settings for the boson sampler are at the hybrid system. The hybrid system 550 is not required to possess details of the architecture or trained parameters of the trained generative model. The electronic system 500 is not required to have a boson sampler or the configuration settings for the boson sampler. In the event that an adversary appropriates the trained model, then they would only be able to operate the trained model with degraded or suboptimal performance without further communication with the hybrid system 550.

[0138] In the method of Fig. 5 described above, the configuration settings of the boson sampler are selected once (at 504) and then are not changed. In other examples, the configuration settings, and in particular the parameters of the boson sampler on the hybrid system 550, may be trained as well as the generative model itself on the electronic sy stem 500.

[0139] As an example, Fig. 7 illustrates a hybrid neural network (HNN) 700 comprising a boson sampling layer 710 and a generative model, in this example a classical artificial neural network (ANN) 720. Similar to the ANN of Fig. 1, the ANN 720 comprises an input layer 730, a hidden layer 740. and an output layer 750. The input layer 730 comprises a first plurality of neurons 730-1 to 730-r, the hidden layer 740 comprises a second plurality of neurons 740-1 to 740-v, and the output layer 750 comprises a third plurality of neurons 750-1 to 750-w. While only a single hidden layer is shown in Fig. 7, the skilled person would appreciate that an ANN may have several hidden layers between the input layer 730 and output layer 750. The number of neurons in each layer may be the same or different. Furthermore, while every neuron in a layer is connected to every neuron in the next layer in the ANN 720, the skilled person will appreciate that each neuron may be comiected to fewer neurons of the next layer, or may be connected to a neuron in the same layer or a preceding layer. The activation function implemented by each neuron may be, for example, the sigmoid activation function, the tanh activation function, or the rectified linear (ReLu) activation function.

[0140] The boson sampling layer 710 may be implemented by the boson sampler of the hybrid system 550, while the generator network 720 may be implemented in logic on the electronic system 500. The ANN 720 is configured to receive as input a latent vector (indicated by the values z± to zr in the figure) determined from a sample output from the boson sampling layer 710. However, the parameters of the boson sampling layer 710 may be trained also during training of the ANN. That is, training the ANN 720 may comprise training the ANN as part of a hybrid neural network 700 that also comprises a boson sampling layer.

[0141] Each sample of the batch of samples may be used to determine a latent vector. Each latent vector may be provided as input to the ANN 720 which produces synthetic data output values yl to yw. The output values produced for all of the samples of the batch of samples may be compared in some way with representative examples from the training data, typically by use of a cost function. The parameters of the boson sampling layer 710 and the weights and biases of the ANN 720 may be iteratively updated according to a learning rule. Successive adjustments will cause the HNN 700 to produce synthetic data that is increasingly similar to the target output data. After a sufficient number of these adjustments the training can be terminated based upon certain criteria.

[0142] The HNN 700 may be trained by attempting to optimise a cost function indicative of the error between synthetic data generated by die HNN and a training set of genuine data. For example, training may comprise reducing a cost function such as a quadratic cost function, a cross-entropy cross function, or a log-likelihood cost function. The reduction may be performed for example by gradient descent, stochastic gradient descent or variations thereof, using backpropagation to adjust the parameters of die boson sampling layer 710 and the weights and biases within the ANN 720 accordingly. A method of using gradient descent to train parameters of the boson sampling layer 710 of a HNN 700 is described in, for example, UK patent application number GB2309657.1, “Boson Sampler Parameter Configuration”, filed on 27 June 2023, which is hereby incorporated by reference in its entirety. Training may involve the use of further techniques known to the skilled person, such as regularization. Mini-batch sizes, learning rates, numbers of epochs and other hyperparameters may be selected and fine-tuned during training.

[0143] The skilled person would appreciate that the boson sampling layer 710 on the hybrid system 550 and the ANN 720 on the electronic system 500 may be trained in any suitable manner. For example, the ANN 720 may be trained as part of a GAN.

[0144] Fig. 8 shows a swim lane flowchart of a method of training a generative model according to an example. In this example, the generative model on the electronic system 500 is trained as part of a hybrid generative model (such as the hybrid neural network 700 shown in Fig. 7) including a boson sampling layer. In other words, the configuration settings of the boson sampler at the hybrid system 550 are trained in an iterative process while training the generative model at the electronic system 500. A first iteration (steps 802 to 814) and a second iteration (steps 816 to 828) are shown in Fig. 8, but further iterations may be used to further train the configuration settings and generative model.

[0145] At 802, the electronic system 500 sends a request to the hybrid system, and at 804, the hybrid system 550 selects configuration settings for the boson sampler. These processes may take place in a similar manner to steps 502, 504 described above in relation to Fig. 5.

[0146] At 806, the hybrid system 550 associates a model identifier (“model ID”) with the selected configuration settings. Associating a model identifier with the selected configuration settings may comprise, for example, recording the model identifier and selected configuration settings in a database in memory at the hybrid system 550. The model identifier may comprise, for example, a name, a number or an alphanumeric string. In some examples, for example when the hybrid system 550 is being used with multiple electronic systems to train multiple generative models, the use of model identifiers may enable each of those generative models to be distinct. In some examples, for example when the hybrid system 550 is only communicating with one electronic system to train one generative model, step 806 may be skipped.

[0147] At 808, the hybrid system operates the boson sampler in accordance with the selected configuration settings to produce a batch of samples, and at 810, the hybrid system 550 determines a set of latent vectors from tire batch of samples. These processes may take place in a similar manner to steps 506, 508 described above in relation to Fig. 5.

[0148] The hybrid system 550 may associate each determined latent vector with a latent vector identifier. A latent vector identifier may be any suitable identifier of a specific latent vector. For example, a latent vector identifier may comprise a hash of a latent vector or an index number allocated to a particular latent vector.

[0149] At 812, the hybrid system 550 communicates tire determined latent vectors to the electronic system 500. The hybrid system 550 further communicates the model identifier to the electronic system 500. Further information, for example one or more latent vector identifiers, may also be communicated to the electronic system 500.

[0150] At 814, the electronic system 500 uses the received latent vectors to train a generative model to produce a sy nthetic dataset. The skilled person will appreciate that the electronic system 500 may further post-process the latent vectors prior to training of the generative model.

[0151] The electronic system 500 may associate the model identifier with the trained (first iteration) generative model. For example, the electronic system 500 may record in a local database the model identifier and a hash of the trained generative model (e.g. a hash of a computer-readable file documenting the architecture or structure of the model and the associated model parameters).

[0152] At 816, tire electronic system 500 sends a further request for latent vectors to the hybrid system 550. The request is similar to that sent at 802, but further comprises the model identifier and performance data. As mentioned above, in some examples such as when the hybrid system 550 is only involved in the training of a single model, the model identifier may be omitted. The performance data is indicative of a performance of the (first iteration) generative model trained using the set of latent vectors communicated at 812. For example, the performance data may comprise a function value (e.g. a cost function value) of a (cost) function associated with the training of the (first iteration) generative model. The performance data may further comprise a latent vector identifier associated with a latent vector received at 812 that was used in training the generative model at 814.

[0153] At 818, the hybrid system 550 identifies, based on the request at 816, the configuration settings for the boson sampler that are associated with the first iteration of the generative model. In particular, the hybrid system determines the model identifier from tire request, and identifies the configuration settings using the received model identifier. The hybrid system 550 also updates the configuration settings of the boson sampler based at least in part on the request. For example, the hybrid system 550 may analyse the performance data provided in the request and then select updated parameter values for the interferometer of the boson sampler.

[0154] At 820, the hybrid system associates the model identifier with the updated configuration settings. At 822 tire hybrid system operates the boson sampler in accordance with the updated configuration settings to produce a further batch of samples. At 824. the hybrid system determines a further set of latent vectors from the further batch of samples. At 826, the latent vectors are sent to the electronic system 500, and at 828 the electronic system uses the received latent vectors to further train the generative model.

[0155] Steps 816 to 828 may be repeated one or more times to further train the generative model. The steps 816 to 828 may be repeated until a stopping condition is reached. For example, the stopping condition may be that the generative model reaches a particular performance threshold e.g. the cost function values of samples are indicative that tire synthetic datasets generated by the model are convincing. For example, the stopping condition may be that a predetermined number of iterations has been performed.

[0156] Advantageously, the method of Fig. 8 can lead to an improved generative model at the electronic system 500. By training the parameters of the boson sampler at the hybrid system 550, the latent vectors produced may be better suited to the generation task of the generative model.

[0157] Fig. 9 shows a swim lane flowchart of an alternative method of training a generative model according to an example. In this example, latent vectors are not communicated from the hybrid system 550 to the electronic device 500 during training. Instead, the hybrid system selects configuration settings (902), operates the boson sampler in accordance with the selected configuration settings to produce a batch of samples (904), determines latent vectors from the batch of samples (906) and uses the latent vectors to train a generative model (908). The skilled person will appreciate that the steps 902 to 908 may be varied. For example, in an iterative process the configuration settings of the boson sampler may be trained while training the generative model.

[0158] In contrast to, for example, the method of Fig. 5, in the method of Fig. 9 all training is performed on the hybrid system 550. At 910, the trained model (for example, a computer-readable file detailing the architecture or structure and trained parameters of the trained model) is communicated to the electronic system 500. In some examples, the trained model may be enciypted.

[0159] At 912, the hybrid system 550 locally erases the trained model. By locally erasing the model, the trained model and configuration settings are not stored at the same system. The skilled person will appreciate that in some examples, the trained model may not be erased.

[0160] Using any of the methods of Figs. 5, 8 and 9 described above, a trained generative model is stored at a separate system to that at which the configuration settings are stored. By training a generative model according to any of the methods described herein, the performance of the trained model is inherently linked to the boson sampler and configuration settings associated with the latent space with which the model is trained. Accordingly, further communication and authentication between the electronic system 500 and hybrid system 550 may be required for inference.

[0161] Fig. 10 shows a swim lane flowchart of an inference method according to an example. For inference, it is assumed that the electronic system 500 lias full knowledge and control of the trained generative model. For example, the electronic system 500 may possess a data file detailing the architecture and trained parameters of the trained model. The inference method of Fig. 10 is compatible with any of the training methods described herein.

[0162] At 1002, the electronic system 500 sends a request to the hybrid system 550 for one or more latent vectors to use with a trained generative model. The request may include, for example, a user identifier, a device identifier, or a model identifier.

[0163] At 1004, the hybrid system 550 identifies configuration settings for the boson sampler based on the request. For example, the hybrid system 550 may verify from information provided in the request (for example a device identifier or user identifier) that a request is valid. For example, the hybrid system 550 may consult a locally stored database of model identifiers and configuration settings and may use a model identifier provided in the request to look up the configuration settings in the database.

[0164] At 1006. the hybrid system 550 operates the boson sampler in accordance with the identified configuration settings to produce one or more samples. At 1008, the hybrid system 550 determines, from the one or more samples, one or more latent vectors. At 1010, the hybrid systems 550 securely communicates the one or more latent vectors to the requesting electronic system 500.

[0165] At 1012, the electronic system 550 uses the received one or more latent vectors with the trained generative model to generate a synthetic dataset.

[0166] Variations of the inference method are envisaged. As an example, at 1010 the hybrid system may communicate enough latent vectors for multiple uses of the trained generative model. For example, the electronic system 500 may be able to locally store enough latent vectors for multiple uses of the trained model. Advantageously, this would reduce the number of times that the electronic system 500 needs to communicate with tire hybrid system 550, thereby reducing delays and improving the user experience while still ensuring that tire performance of the model degrades after a predetermined number of model uses unless the electronic system 500 reestablishes connection with the hybrid system 550.

[0167] Fig. 11 shows a block diagram of an electronic device 1100 according to an example. The electronic device 1100 may be a personal computer or laptop, or may be a server or mainframe, or any other computing device. The electronic device is suitable for performing the actions of the electronic system 500 in any of the methods described herein. The electronic device 1100 comprises a processor 1110 (for example a CPU) and a memory 1120. The processor 1110 is configured to execute computer-readable instructions stored in the memoiy 1120 to perform the actions of the electronic system 500 in one or more of the methods described above. For example, the processor 1110 may execute computer-readable instructions in memory 1120 that cause the processor to initiate a request for latent vectors. The processor 1110 may execute computer-readable instructions in memory 1120 that cause the processor 1110 to train a generative model. For example, the memory may store a data file detailing an architecture of an artificial neural network, and the processor 1110 may use received latent vectors to train parameters of that neural network. The processors 1110 may be further configured to store the trained model in memory 1120. For inference, the processor 110 may be configured to initiate a request for latent vectors, and to use the requested latent vectors in conjunction with a trained model stored in me mon 1120 to generate a synthetic dataset.

[0168] Fig. 12 shows a graph that demonstrates that information about the probability distribution used in training, and therefore information about the configuration settings of a boson sampler used in training, is useful for inference. In particular, a GAN was trained using the MNIST database (Modified National Institute of Standards and Technology database) to generate synthetic images. All latent vectors used during training were derived from a probability' distribution corresponding to possible outcomes from a boson sampler similar to that illustrated in Fig. 3 witli fixed beamsplitter reflection / transmission coefficients (in other words, tire configuration settings are selected and then not updated during training). After each training iteration of the GAN, the trained GAN was used for inference and the performance of the trained GAN was quantitatively evaluated using the Inception Score (more details on how the Inception Score for a GAN is calculated can be found in Salimans et al., Improved Techniques For Training GANs, International conference on intelligent, secure, and dependable systems in distributed and closed environments, pages 127-138, Springer, 2017). In particular, for inference the latent vectors used were (i) derived from the same probability distribution as used in training, (ii) derived from a second probability' distribution corresponding to possible outcomes from a similar boson sampler witli randomly selected beamsplitter coefficients different to those used in training, or (iii) derived from an independent Bernoulli distribution (i.e. each component of the latent vector was independently selected from a random Bernoulli distribution). In Fig. 12, the Inception Score is plotted against the number of training iterations of the GAN when inference is performed with latent vectors from each of the three latent vector distributions. The solid curve shows tire Inception Score when the latent vectors are derived from the same probability distribution as that used for training. The dotted curve shows the Inception Score when the latent vectors are derived from the second probability distribution (different to that used in training). The dashed cun e shows the Inception Score when each latent vector component is derived from an independent Bernoulli distribution. The trained GAN performs better when the latent vectors for inference are derived from the same probability distribution as the latent vectors used in training (the higher the Inception Score the better).

[0169] Fig. 13 shows a second graph that demonstrates drat information about the probability distribution used in training, and therefore information about the configuration settings of a boson sampler used in training is important for inference. In particular, a StyleGAN2 model was trained using the Flickr-Faces-HQ Dataset (FFHQ), which is a high-quality image dataset of human faces. All latent vectors used during training were derived from a probability distribution corresponding to possible outcomes from a boson sampler similar to that illustrated in Fig- 3 with fixed beamsplitter reflection / transmission coefficients (in other words, the configuration settings were selected and then not updated during training). After each training iteration, the performance of the trained model was evaluated using the Frechet Inception Distance (FID) (see Hensel et al.. GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium, arXiv: 1706.08500. 12 Jan 2018 for a definition of this measure). In particular, for inference tire latent vectors used were (i) derived from the same probability distribution as used in training, or (ii) derived from a “spoof’ distribution. The spoof distribution is a classical probability distribution attempting to emulate the output of the probability distribution used in training. In particular, the spoof distribution is a classical probability7 distribution that has the same probabilities p(i,j) as the probability distribution used in training, wherein the probabilities p(i,j) represent tire probabilities that that there are i photons measured in output mode j. The adversary7 may be able to produce the spoof distribution by. for example, obtaining several hundred of the latent vectors used in training. The spoof method accordingly enables the adversary to approximate the probability distribution used in training to first order. Fig. 13 shows the FID score against number of training iterations of the model when inference is performed using latent vectors derived from the same probability distribution as used in training (solid line) and when the classical “spoofer” distribution is used (dashed line). As can be seen in the figure, the trained model performs better when the latent vectors used in inference are from the same distribution as those used in training (the lower the FID the better).

[0170] Producing a spoofer distribution such as that used in Fig. 13 represents a sophisticated attack, but the performance of the trained model with latent vectors derived from the spoofer distribution is noticeably different to the performance of the trained model with latent vectors derived from the same probability distribution as used in training. Accordingly, even if the an adversary is able to obtain some or even all latent vectors used in training to try to reproduce a produce a spoofer distribution, they will be unable to achieve the same level of performance with the trained model indefinitely. Accordingly, if the probability distribution used in training is that produced by a boson sampler, then an adversary requires both an appropriate boson sampler and knowledge of the configuration settings of the boson sampler to reach good performance with the trained model.

[0171] Variations of the described embodiments are envisaged.

[0172] In some embodiments, the block diagram of Fig. 2 is part of a cloud computing system where boson computing is provided as a shared service to separate users. For example, a cloud computing sen ice provider operates the boson sampler 214 and allows users to use the boson sampler 214. For example, a user using a computing apparatus remote from system 200, generates control instructions and transmits the control instructions to the system 200.

[0173] As will be appreciated by one skilled in the art, the present disclosure may be embodied as a system, method, or computer program product. Accordingly, aspects of the present disclosure may take tire form of an entirely hardware embodiment an entirely software embodiment (including firmware, resident software, microcode, etc.) or an embodiment combining software and hardware aspects. Furthermore, aspects of the present disclosure may take the fonn of a computer program product embodied in any one or more computer-readable medium / media having computer usable program code embodied thereon.

[0174] The flowchart and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that, in some alternative implementations, the functions noted in the block may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.

[0175] Each feature disclosed in this specification (including any accompanying claims, abstract or drawings), may be replaced by alternative features serving the same, equivalent or similar purpose, unless expressly stated otherwise. Thus, unless expressly stated otherwise, each feature disclosed is one example only of a generic series of equivalent or similar features. The disclosure is not restricted to the details of any foregoing embodiments. The disclosure extends to any novel one, or any novel combination, of the features disclosed in this specification (including any accompanying claims, abstract and drawings), or to any novel one, or any novel combination, of the steps of any method or process so disclosed. The claims should not be construed to cover merely the foregoing embodiments, but also any embodiments which fall within the scope of the claims.

Claims

1. A method for performance by an electronic system, the method comprising:sending, to a heterogeneous system comprising a boson sampler, a request for a set of latent vectors for use in training a generative model to generate a synthetic dataset;receiving from the heterogeneous system, in response to the request, the set of latent vectors, wherein the latent vectors have been determined from a batch of samples produced by the boson sampler; andusing the set of latent vectors, training a generative model to generate a synthetic dataset.

2. The method of claim 1, further comprising:receiving from the heterogeneous system a model identifier; andassociating the model identifier with the generative model trained using the set of latent vectors.

3. The method of claim 2, wherein associating the model identifier with the generative model comprises recording the association of the model identifier with the trained generative model.

4. The method of any preceding claim, further comprising:sending to the heterogeneous system a further request for a further set of latent vectors for use in further training the generative model;receiving from the heterogeneous system, in response to the further request, the further set of latent vectors, wherein the further set of latent vectors have been determined from a further batch of samples produced by the boson sampler; andusing the further set of latent vectors, further training the generative model to generate a synthetic dataset.

5. The method according to claim 4, wherein the further request is indicative of a model identifier associated with the trained generative model.

6. The method of claim 4 or claim 5, wherein the further request includes performance data indicative of a performance of the generative model trained using the set of latent vectors.

7. The method of claim 6, wherein the performance data comprises a function value of a function associated with the training of the generative model, and a latent vector identifier associated with a latent vector used in training the generative model.

8. The method of any preceding claim, wherein the generative model comprises an artificial neural network (ANN).

9. The method of claim 8, wherein training the generative model comprises training the ANN as part of a hybrid neural network (HNN) comprising a boson sampling layer and the ANN.

10. A method according any of claims 8 or 9, wherein training the generative model comprises training a generative adversarial network (GAN), the GAN comprising:the ANN; anda second ANN; andwherein training the GAN comprises:training the ANN, using the determined set of latent vectors and feedback from the second ANN to generate a synthetic dataset;training the second ANN, using a plurality of genuine datasets and a plurality of synthetic datasets generated by the ANN, to classify received datasets as synthetic datasets or genuine datasets, and to provide feedback to the ANN; andoutputting the trained ANN configured to generate synthetic datasets.

11. The method of any of claims 1 to 10, wherein training a generative model comprises training a diffusion model.

12. The method of any preceding claim, further comprising:sending to the heterogeneous system a request for one or more latent vectors for use with the trained generative model,05 07 24receiving from the heterogeneous system, in response to the request, the one or more latent vectors; andusing the one or more latent vectors with the trained generative model to generate a synthetic dataset.

13. A method for performance by a heterogeneous system comprising a boson sampler, the method comprising:receiving, from an electronic system, a request for a set of latent vectors for use with training a generative model to generate a synthetic dataset;based at least in part on the request, selecting configuration settings for the boson sampler;operating the boson sampler to produce a batch of samples, the boson sampler configured in accordance with the selected configuration settings;determining, from the batch of samples, the set of latent vectors; andsending the set of latent vectors to the electronic system.

14. The method of claim 13, further comprising:associating a model identifier with the selected configuration settings; andsending the model identifier to the electronic system.

15. The method of claim 13 or claim 14 further comprising:receiving from the electronic system a further request for a further set of latent vectors for use with further training the generative model;based on the further request:identifying configuration settings for the boson sampler; andupdating the identified configuration settings;operating the boson sampler to produce a further batch of samples, the boson sampler configured in accordance with the updated configuration settings;determining, from the further batch of samples, the further set of latent vectors; andsending the determined further set of latent vectors to the electronic system, wherein the further set of latent vectors is usable to further train the generative model to generate a synthetic dataset.

16. The method of claim 15, wherein the further request is indicative of a model identifier associated with the trained model.

17. The method of claim 15 or claim 16, wherein the further request includes performance data indicative of a performance of the generative model trained using the set of latent vectors.

18. The method of claim 17, wherein updating the configuration settings comprises:determining, using the performance data, updated configuration settings.

19. The method of claim 17 or claim 18, wherein the performance data comprises a function value of a function associated with the training of the generative model, and a latent vector identifier associated with a latent vector used in training the generative model.

20. The method of any of claims 13 to 19, wherein a request for a set of latent vectors is indicative of at least one of:a device identifier indicating a device from which the request is received; ora user identifier indicating a user of the device from which the request is received.

21. A method for performance by a heterogeneous system comprising a boson sampler, the method comprising:selecting configuration settings for a boson sampler;operating the boson sampler to produce a batch of samples, the boson sampler configured in accordance with the selected configuration settings;determining, from the batch of samples, a set of latent vectors;using the set of latent vectors, training a generative model to generate a synthetic dataset; and communicating the trained generative model to an electronic system.05 07 2422. The method of claim 21, further comprising locally erasing at least a part of the trained generative model.

23. The method of claim 21 or 22,wherein training the generative model comprises training an artificial neural network (ANN) as part of a hybrid neural network (HNN) comprising a boson sampling layer and the ANN; andwherein selecting configuration settings for the boson sampler comprises selecting configuration settings for the boson sampler while training the HNN.

24. The method of any of claims 13 to 23, further comprising:receiving from the electronic system a request for one or more latent vectors;based on the request, identifying configuration settings for the boson sampler;operating the boson sampler to produce one or more samples, the boson sampler configured in accordance with the identified configuration settings;determining, from the one or more samples, the one or more latent vectors; andsending the determined one or more latent vectors to the electronic system, wherein the one or more latent vectors are usable with the trained generative model to generate a synthetic dataset.

25. The method of any of claims 13 to 24, wherein the configuration settings comprise parameter values for parameters of a configurable interferometer of the boson sampler.

26. The method of any of claims 13 to 25, wherein the configuration settings indicate an input multimodal photonic state provided to an interferometer of the boson sampler.

27. The method of any of claims 13 to 26, wherein each sample is representative of a measurement outcome of one or more photodetectors of the boson sampler.

28. The method of any preceding claim, wherein training the generative model to generate a synthetic dataset comprises training the generative model to generate a synthetic image.

29. A system comprising:a boson sampler; anda set of one or more processors, the set of one or more processors configured to perform a method according to any of claims 13 to 29.

30. The system of claim 29, wherein the boson sampler comprises a configurable interferometer, and wherein the configuration settings comprise parameter values for configurable parameters of the configurable interferometer.

31. The system according to claim 29 or claim 30, wherein the boson sampler comprises one or more photon sources, and wherein the configuration settings indicate an input multimodal photonic state provided to an interferometer of the boson sampler.

32. The system according to any of claims 29 to 31, wherein the boson sampler is:a single-photon boson sampler; ora Gaussian boson sampler.

33. An electronic system comprising a set of one or more processors, the set of one or more processors configured to perform a method according to any of claims 1 to 12.

34. A non-transitory computer readable medium comprising stored instructions that, when executed by one or more processors coupled to a boson sampler, cause the one or more processors to perform the method of any of claims 13 to 28.

35. A non-transitory computer-readable medium comprising stored instructions that, when executed by one or more processors of an electronic system, cause the one or more processors to perform the method of any of claims 1 to 12.

Citation Information

Patent Citations

  • Differentiable generative modelling using a hybrid computer including a quantum processor

    EP4224378A1

  • Image generation system and method

    GB2619368A