Quantum machine learning

A hybrid quantum-classical machine learning model trains on distributed devices, enabling efficient training on resource-constrained systems by using quantum computing advantages and addressing data privacy issues, despite the lack of fault-tolerant quantum computers.

WO2026030790A1PCT designated stage Publication Date: 2026-02-12COMMONWEALTH SCI & IND RES ORG +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
PCT/AU2025/050837
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-09
Filing Date
2025-08-06
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Resource-constrained classical devices lack the computational power to train machine learning models to desired accuracy, and fault-tolerant quantum computers are years away, making it difficult to leverage quantum computing for effective training.

Method used

A hybrid quantum-classical machine learning model is trained by distributing the model across classical and quantum devices, where the classical sub-model is trained locally and the quantum sub-model is trained on a quantum server, using a quantum circuit configured by classical outputs to minimize loss and update both models collaboratively.

Benefits of technology

This approach enables efficient training of machine learning models on resource-constrained devices, leveraging quantum computing advantages without sharing raw data, while addressing data privacy leakage through noise-based defense mechanisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure AU2025050837_12022026_PF_FP_ABST
    Figure AU2025050837_12022026_PF_FP_ABST
Patent Text Reader

Abstract

The disclosure relates to machine learning, more particular to training a machine learning model comprising a classical sub-model and a quantum sub-model. A classical processor, receives, from a classical device, an intermediate classical output from a classical sub-model of the machine learning model and configures quantum gates of a quantum circuit based on the intermediate classical output. A quantum processor executes the quantum circuit using the quantum gates to determine a quantum circuit output, the quantum circuit being configured to represent a quantum sub-model of the machine learning model. The classical processor adapts the quantum circuit output to determine a further classical output; updates the quantum sub-model based on minimising a loss involving the further classical output; and transmits a loss propagation value to the classical device, to cause the classical device to update the classical sub-model based on the loss propagation value, thereby training the machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

"Quantum machine learning" Cross-Reference to Related Applications

[0001] The present application claims priority from Australian Provisional Patent Application No 2024902467 filed on 9 August 2024, the contents of which are incorporated herein by reference in their entirety. Technical Field

[0002] The disclosure relates to machine learning. More particularly but not necessarily exclusively, this disclosure relates to training a machine learning model comprising a classical sub-model and a quantum sub-model. Background

[0003] Machine learning has emerged as a prominent tool in many different applications. For example, machine learning can be used to classify images or other types of data into one or more categories or classes. However, to achieve sufficient accuracy, machine learning models need to be trained using a large amount of training data. In some cases, this is difficult to achieve with standard computer hardware, such as a personal computer, as these devices are typically resource-constrained, meaning that they do not have the computational power to feasibly train machine learning models to the desired accuracy.

[0004] Quantum computing may have potential benefits to training machine learning models. Quantum computers utilise quantum mechanical phenomena (such as entanglement and superposition) to perform calculations that would be unfeasible for Boolean algebra-based classical computing. However, adapting machine learning models to utilise quantum computing is not straightforward, especially given that fault-tolerant quantum computers are still many years away.

[0005] Any discussion of documents, acts, materials, devices, articles or the like which has been included in the present specification is not to be taken as an admission that any or all of these matters form part of the prior art base or were common general knowledge in the field relevant to the present disclosure as it existed before the priority date of each of the appended claims.

[0006] Throughout this specification the word “comprise”, or variations such as “comprises” or “comprising”, will be understood to imply the inclusion of a stated element, integer or step, or group of elements, integers or steps, but not the exclusion of any other element, integer or step, or group of elements, integers or steps. Summary

[0007] Disclosed herein are methods and systems for machine learning. More specifically, disclosed herein is a method and system for training a machine learning model which utilises techniques in classical computing and quantum computing. Also, disclosed herein is a method and system for classifying test data using a trained machine learning model which also utilises techniques in classical computing and quantum computing.

[0008] According to the present disclosure, there is provided a method for training a machine learning model, the method comprising: receiving, from a classical device, an intermediate classical output from a classical sub-model of the machine learning model, the intermediate classical output representing an evaluation of the classical sub-model on training data; configuring quantum gates of a quantum circuit based on the intermediate classical output; executing the quantum circuit using the quantum gates to determine a quantum circuit output, the quantum circuit being configured to represent a quantum sub-model of the machine learning model; adapting the quantum circuit output to determine a further classical output; updating the quantum sub-model based on minimising a loss involving the further classical output; and transmitting a loss propagation value based on the updated quantum sub-model to the classical device, to cause the classical device to update the classical sub-model based on the loss propagation value, to thereby train the machine learning model.

[0009] It may be an advantage to train the machine learning model by minimising the loss to update the quantum sub-model as the training is not solely performed on the classical device. This is particularly advantageous for resource-constraint classical devices, which lack the computational power to individually train the machine learning model. Further, the classical sub- model and the quantum sub-model of the machine learning model can be distributed to different devices, enabling users of classical device to utilise quantum computing advantages.

[0010] In some embodiments, the machine learning model comprises a further classical sub- model and adapting the quantum circuit output comprises applying the further classical sub- model to the quantum circuit output to determine the further classical output.

[0011] In some embodiments, training the machine learning model comprises: updating the further classical sub-model based on minimising the loss involving the further classical output; and determining a further loss propagation value based on the updated further classical sub- model; and updating the quantum sub-model comprises updating the quantum sub-model based on the further loss propagation value.

[0012] In some embodiments, configuring the quantum gates of the quantum circuit comprises embedding the intermediate classical output into one or more qubits representing the quantum circuit to reduce a dimensionality of the intermediate classical output.

[0013] In some embodiments, the quantum gates of the quantum circuit comprise rotation gates; and configuring the quantum gates of the quantum circuit comprises configuring the rotation gates based on the intermediate classical output, wherein the rotation gates are parameterised by the intermediate classical output.

[0014] In some embodiments, the quantum gates of the quantum circuit comprise entangling gates, and the rotation gates and the entangling gates of the quantum circuit are alternating in the quantum circuit.

[0015] In some embodiments, the classical device is one of multiple classical devices, each of the multiple classical devices comprising the classical sub-model; and the method further comprises, for each of the multiple classical devices: repeating the steps of receiving the intermediate classical output from the classical device, configuring the quantum gates, executing the quantum circuit, adapting the quantum circuit output, and updating the quantum sub-model; and transmitting the loss propagation value based on the updated quantum sub-model to the each of the multiple classical devices to cause each of the multiple classical devices to update the classical sub-model on the respective classical device based on the received loss propagation values.

[0016] In some embodiments, each of the multiple classical devices stores respective training data and performing the method of claim 8 trains the machine learning model on the training data of each of the multiple classical devices while isolating the respective training data from other ones of the multiple classical devices.

[0017] In some embodiments, the respective training data of each of the multiple classical devices is confidential training data.

[0018] In some embodiments, the classical sub-model is a neural network, and the quantum sub-model is a quantum neural network.

[0019] In some embodiments, the machine learning model is a classifier.

[0020] In some embodiments, the method further comprises receiving, from the classical device, a label being associated with the training data; and updating the quantum sub-model is based on minimising the loss involving the further classical output and the label.

[0021] In some embodiments, the quantum circuit is implemented on a quantum device comprising at least two qubits.

[0022] In some embodiments, the quantum sub-model and the further classical sub-model are implemented on a hybrid quantum server comprising a quantum processor and a classical processor, the hybrid quantum server being configured to perform the method of according to the previous embodiments.

[0023] According to the present disclosure, there is provided a method for classifying test data using a trained machine learning model as performed by a classical device, the trained machine learning model comprising a classical sub-model and a quantum sub-model, the method comprising: evaluating the classical sub-model on the test data to determine a classical output; applying a noise layer to the classical output to determine a noisy output, the noise layer being configured to add noise to an output of the classical sub-model; and sending the noisy output to a server to cause the server to: configure quantum gates of a quantum circuit based on the noisy output; execute the quantum circuit using the quantum gates to determine a quantum circuit output, the quantum circuit being configured to represent the quantum sub-model; adapt the quantum circuit output to determine a classification output representing a classification of the test data; and send the classification output to the classical device to classify the test data.

[0024] In some embodiments, the noise layer is configured to add Laplacian noise to the output of the classical sub-model.

[0025] In some embodiments, wherein the Laplacian noise is based on a Laplace distributionwith a mean of about 4^ and a scale parameter between about 0.01 and 1.

[0026] According to the present disclosure, there is provided software that, when installed on a classical processor and executed by the classical processor, causes the classical processor to perform the method of any of the previous embodiments.

[0027] According to the present disclosure, there is provided a hybrid quantum server comprising: a classical processor configured to: receive, from a classical device, an intermediate classical output from a classical sub-model of the machine learning model, the intermediate classical output representing an evaluation of the classical sub-model on training data; and configure quantum gates of a quantum circuit based on the intermediate classical output; a quantum processor configured to execute the quantum circuit using the quantum gates to determine a quantum circuit output, the quantum circuit being configured to represent a quantum sub-model of the machine learning model; and wherein the classical processor is further configured to: adapt the quantum circuit output to determine a further classical output; update the quantum sub-model based on minimising a loss involving the further classical output; and transmit a loss propagation value based on the updated quantum sub-model to the classical device, to cause the classical device to update the classical sub-model based on the loss propagation value, to thereby train the machine learning model.

[0028] Optional features provided in relation to the first method, equally apply as optional features to the second method, the software and the hybrid quantum server. Brief Description of Drawings

[0029] An example will be described with reference to the following drawings:

[0030] Fig.1 illustrates a system for machine learning.

[0031] Fig.2 illustrates a method for training a machine learning model.

[0032] Fig.3 illustrates another system for machine learning.

[0033] Fig.4 illustrates another method for training a machine learning model.

[0034] Fig.5 illustrates a method for classifying test data using a trained machine learning model.

[0035] Fig.6 shows the Hybrid Quantum Split Learning (HQSL) model used in theexperiments described herein with N clients and 1 hybrid quantum server.

[0036] Fig.7a shows the qubit-efficient data loading scheme used in the experiments described herein.

[0037] Fig.7b shows the optimal quantum circuit utilised in the experiments described herein.

[0038] Fig. 8 shows a representation of a 2 qubit circuit (Circuit 7) with a single layer, L ,without entanglement consisting of the arrangement U ( X )− U ( ^ ) , where U is decomposed herein RZ− RY − RZ gates, for each qubit.

[0039] Fig.9 shows a representation of a reorganization of Circuit 7 of Fig.8 to introduce entanglement between each qubit (Circuit 8).

[0040] Fig.10 presents Table 1, which shows the results comparing Mean Accuracy and F1- score of HQSL models consisting of Circuits 7, 8, 6, 9, 10 during classification of the FMNIST dataset.

[0041] Fig.11 shows the effect of the number of qubits and circuit depth on the mean test accuracy and 5-fold training time of HQSL on FMNIST dataset.

[0042] Fig.12a shows HQSL variant 1, which was used for binary classification on multivariate datasets.

[0043] Fig. 12b shows HQSL variant 2, which was used for n -class classification on single-channel image datasets.

[0044] Fig.13a shows the performance comparison in terms of test accuracy of split learning (white) versus HQSL (black) with 5-fold cross-validation.

[0045] Fig.13b shows the performance comparison in terms of F1-score of split learning (white) versus HQSL (black) with 5-fold cross-validation.

[0046] Fig.14a shows the test accuracy of HQSL compared with split learning on all 5 datasetsfor up to K =100 clients.

[0047] Fig.14b shows the F1-score performances of HQSL compared with split learning on all5 datasets for up to K =100 clients.

[0048] Fig.15 illustrates the setup of a reconstruction attack.

[0049] Fig. 16a shows the probability density function of Laplacian noise centred at ^ withvarying scale parameter, b .

[0050] Fig.16b shows Rotation of qubit (red dot) around the x-axis of the Bloch sphere afteradding Laplacian noise, ^ , with mean ^ = 4 ^ and small scale, b .

[0051] Fig.17 illustrates the defence setup deployed at inference time after training the HQSL / split learning models.

[0052] Fig.18a shows the results for comparing the difference between original and reconstructed images under HQSL and split learning settings at different noise levels for the FMNIST.

[0053] Fig.18b shows the results for comparing the difference between original and reconstructed images under HQSL and split learning settings at different noise levels for the MNIST.

[0054] Fig.18c shows the results for comparing the difference between original and reconstructed images under HQSL and split learning settings at different noise levels for the Speech Commands dataset.

[0055] Fig. 19a shows the 5-fold test accuracy and F1-score versus scale parameter, b , withmean fixed atthe FMNIST dataset.

[0056] Fig. 19b shows the 5-fold test accuracy and F1-score versus scale parameter, b , withmean fixed at ^ = 2 ^ for the MNIST dataset.

[0057] Fig. 19c shows the 5-fold test accuracy and F1-score versus scale parameter, b , withmean fixed at ^ = 2 ^ for the Spectrograms for Speech Commands dataset.

[0058] Fig. 20a shows the 5-fold test accuracy and F1-score versus scale parameter, b , withmean fixed at ^ = 4 ^ for the FMNIST dataset.

[0059] Fig. 20b shows the 5-fold test accuracy and F1-score versus scale parameter, b , withmean fixed at ^ = 4 ^ for the MNIST dataset.

[0060] Fig. 20c shows the 5-fold test accuracy and F1-score versus scale parameter, b , withmean fixed at ^ = 4 ^ for the Spectrograms for Speech Commands dataset.

[0061] Fig.21a shows the original and reconstructed images in classical and hybrid settingsrespectively, with Laplacian noise layer with parameters ^ = 4 ^ , b = 0.01 for the FMNISTdataset.

[0062] Fig.21b shows the original and reconstructed images in classical and hybrid settingsrespectively, with Laplacian noise layer with parameters ^ = 4 ^ , b = 0.01 for the MNISTdataset.

[0063] Fig.22 presents Table 3, which shows the quantum circuits trialled during the experiments described herein.

[0064] Fig.23 presents Table 4, which shows a summary of accuracy and F1-score for HQSL model using each quantum circuit from Table 3 of Fig.22 compared against their equivalent split learning models. Description of Embodiments

[0065] In this disclosure, the unique properties of qubit systems, such as superposition and entanglement, are leveraged to improve machine learning tasks using quantum computing, which may be referred to as quantum machine learning. More specifically, a machine learning model is present which comprises a classical sub-model and a quantum sub-model, where the methods described herein relate to training and application of this machine learning model. As the machine learning model comprises a classical sub-model and a quantum sub-model, the machine learning model may be referred to as a hybrid quantum machine learning model, given that it is neither a fully classical model nor a fully quantum model.

[0066] While fault-tolerant quantum computers are still years away, the disclosed machine learning model and the methods for training thereof provide a feasible route for addressing real- world problems in the Noisy-Intermediate Scale Quantum (NISQ) era. In particular, this disclosure also presents the architecture of such a machine learning model, which accounts for the current limitation of quantum computing hardware. Moreover, throughout experiments discussed later in this disclosure, an optimal quantum circuit for such a machine learning model was determined, which has a limited circuit depth and width and hence, can be easily implemented on current NISQ devices.

[0067] Inputting a relatively large dimension of classical data into a quantum circuit can be achieved using a qubit-efficient data loading technique described herein. This approach is particularly advantageous for current-generation quantum computers which are constrained by quantum circuit size and depth limitations. Further, the utilisation of quantum computing enables advantages that are unfeasible in the classical domain. More specifically, qubits of a quantum computer have more degrees of freedom compared to a classical bit. For example, for classification tasks, the extra degrees of freedom enable higher dimensional clusters to be determined to enable better classification accuracy.

[0068] However, a client may have to run the entire machine learning model on their personal classical device, which can be too computationally intensive for resource-constrained clients. As such, this disclosure presents an application of a (hybrid quantum) machine learning model on a distributed computing environment. Preferably, the machine learning model disclosed herein may be divided into two parts: the client-side model, which has only the classical sub-model and is implemented on the client’s classical device, and the server-side model, which comprises a quantum sub-model. In other words, a hybrid quantum machine learning model is leveraged in a decentralized setup involving resource-constrained classical clients, such as Internet of Things (IoT) and mobile devices, which interact with an edge server possessing quantum computing capabilities.

[0069] By offloading the computationally intensive parts of the model, distributing the machine learning model enables more efficient resource allocation on the client devices, benefiting resource-constrained clients. For example, the machine learning model can be split into client and server components, where the client trains the initial few layers on local data, and then the server continues training the remaining layers. This enables collaborative training without sharing raw data. This is indicative of “split learning”. As such, this enables clients (with no quantum computing capabilities) to compute in the classical domain in resource-constrained NISQ-era environments. The methods of training the machine learning model may be considered to be split learning of a hybrid quantum machine learning model. This may be referred to as “hybrid quantum split learning”.

[0070] However, distributed computing environment can suffer from data privacy leakage as a result of performing distributed machine learning. Data privacy leakage in split learning is the risk of private information being revealed or reconstructed from exchanged information during collaborative training, despite not sharing raw data directly. However, while some techniques prevent the sharing of clients’ raw data with the server, information leakage still occurs from the data communicated across the distributed computing environment. This is a concerning security issue because the data may contain latent information about the raw input data and can be used to stage an input data reconstruction attack at the server side. There is also a risk of reconstruction attacks on the server side if the server is honest but curious. To tackle the data privacy leakage, this disclosure presents a noise-based defence mechanism including a noise layer based on noise in the client-side model, which is tuned considering the rotational properties of single-qubit rotation gates. System

[0071] Fig.1 illustrates system 100 for machine learning. Fig.1 is one example of a configuration of system 100. However, system 100 is not strictly limited to this configuration and this may be one possible embodiment of system 100. It is noted that system 100 of Fig.1 is only meant to illustrate an example and a preferred system which is capable of performing the disclosed methods.

[0072] System 100 comprises classical device 110, which may be smartphone, computer, tablet, a server device, or any other similar classical computing device. Classical device 110 may also be a field-programmable gate array (FPGA), an application specific integrated circuits (ASIC), or one or more single board computers, such as a Raspberry Pi or an Arduino. In some embodiments, classical device 110 is resource-constrained, which may mean classical device 110 has limited processing power, which affects how fast the device can execute tasks. In particular, it may be unfeasible for classical device 110 to train a machine learning model (such as a deep neural network, for example) with a large set of training data.

[0073] Classical device 110 comprises classical processor 111, which may be configured to perform the methods described in this disclosure or parts thereof. Classical device 110 comprises memory 112, which comprises non-volatile memory 113 and / or volatile memory 114. Classical processor 111 may communicate with memory 112 by communicating with non-volatile memory 113 and / or volatile memory 114. Non-volatile memory 113 is a non-transitory computer readable medium and may be an optical disk drive, hard disk drive, solid-state drive, flash memory, storage server, cloud storage or another equivalent type of memory. Volatile memory 114 may be cache, RAM, or another equivalent type of memory.

[0074] Memory 112 may store data to be retrieved for later use. For example, memory 112 may store training data, such as images in some examples. Memory 112 may also store outputs from the sub-models, including one or more classical outputs and a quantum circuit output. Memory 112 may also store loss values and loss propagation values. In essence, memory 112 may store any data used by classical processor 111 when performing the methods described herein or parts thereof. The data thereof may be stored in memory 112 in the form of a JSON format file, XML format file or another equivalent data format file. The images which may form the training data may be stored in memory 112 in a Joint Photographic Experts Group (JPEG) format, RAW data format, or another equivalent data format. It is noted that the reference to “image” in this disclosure refers to “image data”, which may be two-dimensional image data, three-dimensional image data or multi-dimensional image data (in the case of multi-spectral or hyperspectral images).

[0075] The methods described herein comprise training or applying one or more machine learning models. These machine learning models may be stored on memory 112 by storing the weights that form the respective models, for example. Memory 112 may also store any output values calculated by classical processor 111 applying the one or more machine learning models, or any other variable or data necessary to perform such methods described herein. For a hybrid quantum machine learning model, memory 112 may store the parameters that represent the classical sub-model, for example. The quantum sub-model may also be represented by one or more parameters (similar to the classical sub-model), which may be stored on memory 112 in a similar manner.

[0076] Software, that is, an executable program stored on non-volatile memory 113 causes classical processor 111 to perform methods for training a machine learning model. While the singular of “processor” is used herein, it is meant to also encompass multiple processors that are individually or together configured (e.g., programmed) to perform the methods disclosed herein. As such, classical processor 111 may refer to multiple central processing units (CPUs) and / or graphical processing units (GPUs) that are configured to collectively perform the methods disclosed herein.

[0077] Once executed, the software may cause classical processor 111 to apply a classical sub- model to training data to determine an intermediate classical output, transmit the intermediate classical output, receive a loss propagation value, and update the classical sub-model based on the received loss propagation value.

[0078] Software may provide a user interface (such as a graphical user interface) presented to the user on classical device 110. The user interface may be configured to accept input (via buttons or text fields etc) from the user, via a touch screen or a device attached to classical device 110 such as a keyboard or computer mouse. These devices may also include a touchpad, an externally connected touchscreen, a joystick, a button, and a dial. In an example, the user interface may display multiple training data sets and a user may choose one of the multiple training data sets by interacting with the user interface. The user interaction may then cause classical processor 111 to perform a method of training a machine learning model using the chosen training data set.

[0079] System 100 also comprises server 150, which may be a remote server, a quantum device or an equivalent device. As depicted in Fig.1, server 150 is a hybrid quantum device, meaning that it comprises classical part 160 and quantum part 170. In other words, server 150 may comprise classical processor 161 (which may be equivalent or may perform similar functionalities to classical processor 111) and quantum processor 171. However, in someembodiments, server 150 may simply comprise a quantum part and hence, server 150 may be quantum computing server. Server 150 may provide computing services to a user of classical device 110. As such, the user may be referred to a client and classical device 110 may be referred to as a client device.

[0080] Classical part 160 may be similar to, equivalent to, or have similar capabilities as classical device 110. For example, classical part 160 may be a computer, a server device, or any other similar classical device. Server 150 comprises classical processor 161, which may be configured to perform the methods described in this disclosure or parts thereof. Classical part 160 may also have memory capabilities similar to classical device 110. While classical part 160 and quantum part 170 are depicted as the same device in Fig.1, it is noted that classical part 160 and quantum part 170 may be separate and distinct parts (e.g., computers) that may be in communication (such as through a wired connection, like an optical fibre).

[0081] In system 100, quantum processor 171 comprises multiple qubits 172. Quantum processor 171 implements one or more operations on multiple qubits 172. The structure or arrangement of the multiple qubits 172 in quantum processor 171 in Fig.1 is a simplified illustration to assist understanding and may or may not be related to the topology or architecture of quantum processor 171. In the example shown in Fig.1, the multiple qubits 172 are illustrated in the form of a 5-by-3 array of qubits. In some embodiments, quantum part 170 is a NISQ device. For example, quantum processor 171 may only comprises a small number of multiple qubits 172, such as two qubits. Quantum processor 171 may also comprise more than two qubits. In some examples, a quantum circuit that represents the quantum sub-model of the machine learning model may use less qubits than the total number of available qubits of quantum processor 171.

[0082] It is understood that quantum processor 171 comprising multiple qubits 172 is implemented as a quantum annealer or quantum gate model computer. Example of quantum computers which may benefit from the disclosed method include D-Wave’s quantum annealers for quantum annealing, Google’s, IBM’s, or Rigetti’s quantum gate model computers, and / or NEC’s quantum inspired vector annealers, or the like. Operations implemented by quantum processor 101 may include a universal gate set to perform “circuit model” quantum computing, entangling operations for measurement-based quantum computing, or adiabatic evolutions for adiabatic quantum computing.

[0083] In some examples, classical processor 161 may control the operation of quantum processor 171. For example, software, that is, an executable program stored on non-volatile memory causes classical processor 161 and / or quantum processor 171 to perform methodsdisclosed herein or part thereof. Once executed, the software may cause classical processor 161 to receive an intermediate classical output from classical device 110, configure quantum gates of a quantum circuit implemented on quantum processor 171, execute the quantum circuit, adapt the quantum circuit output, update the quantum sub-model and transmit a loss propagation value. Quantum part 170 (or more specifically, quantum processor 171) may also be implemented with embedded code that causes quantum processor 171 to perform the methods described herein or parts thereof. It is also noted that classical processor 111 of classical device 110 may also control the operation of quantum processor 171.

[0084] Classical device 110 may further comprise an input-output (I / O) port 115, which enables classical device 110 to establish a communication with server 150. Server 150 may similarly comprise I / O port 155, configured similarly to I / O port 115. In some examples, classical device 110 may communicate with server 150 via I / O port 115 by using a Wi-Fi network according to IEEE 802.11. The Wi-Fi network may be a decentralised ad-hoc network, such that no dedicated management infrastructure, such as a router, is required or a centralised network with a router or access point managing the network. Classical device 110 and server 150 may also communicate via I / O port 115 using a wired connection, such as Ethernet. System 100 may further be implemented within a cloud computing environment, such as a managed group of interconnected servers hosting a dynamic number of virtual machines. In some examples, classical device 110 may communicate with server 150 via an interface device, or the like. Method for training Training by server 150

[0085] Fig.2 illustrates method 200 for training a machine learning model. Fig.2 is to be understood as a blueprint for a software program and may be implemented step-by-step, such that each step in Fig.2 is represented by a function in a programming language, such as, but not limited to, Python, C++, or Java. The resulting source code is then compiled and stored as computer-executable instructions on non-volatile memory, which causes a processor (or multiple processors or a distributed computing architecture) to perform method 200. For explanatory purposes, some of the steps of method 200 will be explained with reference to classical processor 161. However, it is noted that the steps of method 200 may be performed by other components of system 100, such as classical processor 111, for example.

[0086] The machine learning model (being trained) may perform a machine learning task, such as classification, or the like. The machine learning model comprises a classical sub-model and a quantum sub-model. However, the machine learning model may comprise more than one classical sub-models and more than one quantum sub-models, which may be arranged in anymanner (such as alternating classical or quantum sub-models, for example). Preferably, the classical sub-model is stored and operated on classical device 110 and the quantum sub-model is stored and operated on server 150. This may represent “split learning”, where the whole machine learning model is split into parts that are distributed among different devices or parties. Preferably, the machine learning model is ‘split’ at the connection between the classical sub- model and the quantum sub-model (this may be referred to as the “split point”). In other words, preferably, the first layer implemented on server 150 is the “quantum layer” of the machine learning model (i.e., the quantum sub-model). However, as server 150 has both classical and quantum computing capabilities, the split point may be within the classical sub-model. In other words, the split point may be a classical layer of the machine learning model.

[0087] In split learning, classical processor 111 may train the initial few layers on local data, and then a server may continue training the remaining layers. However, the local data of the client never leaves the respective classical device, which is advantageous especially if the training data is confident or sensitive training data. Split learning enables collaborative training without sharing raw data, for which a machine learning model may be trained using multiple sets of training data without sharing the training data among multiple clients. As such, split learning addresses a challenge of balancing data privacy during collaborative machine learning.

[0088] Classical processor 161 receives 201 an intermediate classical output for classical device 110. More specifically, classical processor 161 receives 201 an intermediate classical output from a classical sub-model of the machine learning model. The intermediate classical output represents an evaluation of the classical sub-model on training data. In other words, classical processor 111 of classical device 110 may apply the classical sub-model on the training data to determine the intermediate classical output, which is then transmitted and received by server 150 (specifically, classical processor 161). The training data may an image (such as a two- dimensional, three-dimensional or multi-dimensional image). In some examples, the training data may be text, such as plain English text, or the like.

[0089] As method 200 is indicative of “split learning”, the intermediate classical output may be considered to be “smashed data”. In the context of split learning “smashed data” refers to the intermediate results or hidden representations that are generated at a certain layer of a machine learning model, also known as the split point. “Intermediate classical output” and “smashed data” may be used interchangeably in this disclosure. In some cases, the smashed data can be as large as and as similar as the input training data.

[0090] In some examples, memory 112 of classical device 110 may store the intermediate classical output rather than explicitly applying the classical sub-model to determine theintermediate classical output. In other examples, the intermediate classical output may be stored on memory of classical part 160. Hence, classical processor 161 receives the intermediate classical output indirectly from classical device 110.

[0091] The machine learning models described in this disclosure (such as the classical sub- model, further classical sub-model and, to an extent, the quantum sub-model) are understood to be models, such as mathematical models or mathematical functions, which receive input and generate an output based on the input. The machine learning models may be of an architecture, such as, but not limited to, a neural network, for example. In general, machine learning models are ‘trained’ to learn and recognise patterns in an input and provide an output that is a prediction based on the training it has undergone. Training involves updating weights or parameters of the machine learning model, which represent the machine learning models, to minimise a loss value, thereby creating a trained machine learning model (in other words, a machine learning model trained to generate an output). This may involve a gradient descent and backpropagation method.

[0092] The machine learning models recited in this disclosure may be stored on memory 112 storing the weights that represent the model. As such, each of the machine learning models may be referred to as a “memory model”, given that it is represented by parameters (i.e., the weights) which can be stored on computer memory. In some embodiments, the machine learning models may be programmed on an integrated circuit, such as a field-programmable gate array (FPGA) or an NVIDIA processing unit. In such an embodiment, the handling modules (or processor(s) that may perform the method or part thereof) may not retrieve the parameters from memory 112. Instead, an input may be communicated to the integrated circuit and the integrated circuit may apply the machine learning model to the input and generate an output, which is then communicated to the handling modules (or processor(s) that may perform the method or part thereof).

[0093] Integrated circuits, such as FPGAs, can be used where flexibility, speed, and parallel processing capabilities are desired. In such an embodiment, the integrated circuit may be part of classical device 110 or classical part 160 of system 100 and may be considered as a “processor” or “processing unit”. Other implementations, such as application specific integrated circuits (ASIC) or neuromorphic architectures are equally useable.

[0094] Applying, executing, or evaluating the machine learning models may involve calling an application programming interface (API) routine to send the input to a server and the server then performs the calculations according to the machine learning model and returns the results. In other examples, applying, executing, or evaluating may involve issuing a command to local hardware, such as a local chip, device, machine learning accelerator (e.g., a USB device designto efficiently perform machine learning tasks or NVIDIA’s Deep Learning Accelerator (DLA)), etc., that has the machine learning model stored thereon and provides a command interface to interact with the model. It is also possible to have a local copy of the machine learning model available so that the calculations are performed by the main processor of the local machine. Other local, remote, or distributed implementations (such as cloud computing environments) are equally useable.

[0095] In some examples, the machine learning models described herein may be a neural network. In further examples, these machine learning models may be a neural network comprising one or more convolutional layers. As such, the machine learning models may perform the methods described herein by creating feature maps using convolutional filters. Such a machine learning model is known as a convolutional neural network (CNN). A CNN is ideal for applications involving images and the image as it accounts for the positioning and shape of objects captures in the image.

[0096] However, other types of machine learning models are equally applicable here, such as K nearest neighbour, decision tree, support vector machines, regression models and other artificial neural networks, such as long short-term memory (LSTM) networks or deep neural networks. The machine learning models described herein may also be a collective of different models. The machine learning models may also be based on a transformer models, which comprises encoder and / or decoder blocks and predictions by ‘tokenising’ the input (such as an image). As such, the machine learning models may also comprise self-attention. The machine learning models may output (and hence, the multiple machine learning outputs may be) a numerical value, a generated output (such as a generated image) or some other type of output (such as a word or sentence).

[0097] In some embodiments, the machine learning models are a multimodal machine learning model, in which multiple inputs of different modalities (e.g., text, image data and audio data) are used to provide one or more generated outputs. An example of a multimodal machine learning model is an object detection model, which detects the location of a specific object (specified by input text, for example) in an image. This example model may generate output text that describes the location of the specified object in the image. Although the multimodal machine learning model can be evaluated on multiple input of different modalities, the multimodal machine learning model can also be evaluated on a single input and still generate an output based on the single input.

[0098] The quantum sub-model described in this disclosure may also be considered to be a machine learning model and the aspects of (classical) machine learning models described above may be equally applicable to quantum machine learning models. For example, a quantummachine learning model may be represented by one or more adjustable parameters (i.e., weights) that are modified during the training process. These weights may be represented as parameterised gates, which are applied to one or more of multiple qubits 172 in the quantum circuit. In some examples, the quantum machine learning models described herein are quantum neural networks, which are similar to (classical) neural networks. In essence, one or more or multiple qubits 172 may be equated to the ‘neurons’ of a (classical) neural network.

[0099] Classical processor 161 then configures 202 quantum gates of a quantum circuit based on the intermediate classical output. The quantum gates may include, but are not limited to, memory (I), bit flip (X), phase flip (Z), rotation gates, T, S, CNOT, CPHASE, or entangling gates. The rotation gates may perform a rotation of a qubit state around the Bloch sphere. For example, one rotation gate may rotation a qubit state round the X-axis, Y-axis or Z-axis of the Bloch sphere. The rotation gates may be indicative of Pauli gates. The quantum gates may also be encoding gates, which are used to encode classical data (such as the intermediate classical output) into quantum data. Encoding gates may be rotation gates, in some examples. In one example, the quantum gates of the quantum may be parameterized, and the intermediate classical output may correspond to the parameters of the quantum gates. Classical processor 161 may apply an algorithm or mathematical model to the intermediate classical output to determine one or more parameters to configure the quantum gates.

[0100] In one example, the quantum gates may be rotation gates which are parameterised by a rotational angle and the classical processor 161 configures 202 the rotation gates by using the intermediate classical output as the rotational angle. More specifically, in some embodiments, the quantum gates of the quantum circuit comprise rotation gates and classical processor 161 configures 202 the quantum gates of the quantum circuit by configuring the rotation gates based on the intermediate classical output, wherein the rotation gates are parameterised by the intermediate classical output.

[0101] In some embodiments, classical processor 161 configures 202 the quantum gates of the quantum circuit by embedding the intermediate classical output into one or more qubits representing the quantum circuit. This may reduce a dimensionality of the intermediate classical output. For example, the intermediate classical output may correspond to the values of an outputlayer of the classical model, which may be represented as an output vector of dimension M . Thenumber of multiple qubits 172 may be N , which is less than M . As such, classical processor161 may embed the intermediate classical output onto multiple qubits 172 to reduce the dimensionality. However, it is noted that, in some examples, the dimension of the intermediateclassical output may be the same as the number of multiple qubits 172 (i.e., N may equal M ).Hence, in some embodiments, classical processor 161 may not need to embed the intermediate classical output to reduce its dimensionality.

[0102] As described later in the disclosure, a preferred quantum circuit was found through experimentation. However, it is noted that the methods and systems disclosed herein are not limited to this quantum circuit. This quantum circuit provided high performance based on the chosen metrics, while being able to be implemented on current NISQ hardware, given the small circuit depth and width. As will be discussed later in this disclosure, the quantum circuit comprises entangling gates and rotation gates. In particular, the quantum circuit comprises rotation gates and entangling gates, which are alternating in the quantum circuit. For example, the quantum circuit may comprise a layer, corresponding to gates being applied to the qubits of the quantum circuit at the same time. The quantum circuit may comprise alternating layers. For example, the quantum circuit may comprise a layer of encoding gates, followed by a layer of rotation gates, followed by a layer of encoding gates and so on.

[0103] Quantum processor 171 then executes 203 the quantum circuit using the quantum gates to determine a quantum circuit output. The quantum circuit is configured to represent a quantum sub-model of the machine learning model. In some examples, classical processor 161 causes quantum processor 171 to execute 203 the quantum circuit. Executing 203 the quantum circuit may comprise applying electromagnetic field to directly control the qubit states of multiple qubits 172. More specifically, applying the configured quantum gates may comprise applying electromagnetic field to directly control the qubit states of multiple qubits 172. As such, configuring 202 the quantum gates may comprise configuring the electromagnetic field which is applied to the multiple qubits 172. For example, configuring the electromagnetic field may involve determining the frequency of the electromagnetic field based on the type of quantum gates, as well as the input parameters of the quantum gate (which may be the intermediate classical output, in some examples).

[0104] The quantum circuit output may be the result or outcome of the execution of the quantum circuit. In particular, the quantum circuit output may be indicative of measurement of each of the multiple qubits 172 after the execution. The measurement of each of the multiple qubits 172 may be a measurement of the qubit state after the execution, for example. The measurement may be indicative of an expectation value of the multiple qubits 172 after the execution, for example.

[0105] Classical processor 161 then adapts 204 the quantum circuit output to determine a further classical output. More specifically, classical processor 161 adapts 204 the quantum circuit output into a classical output, so the quantum circuit output can be used to train theclassical sub-model on classical device 110, or be used as input into a further (classical) machine learning model. Classical processor 161 may adapt 204 the quantum circuit output by applying an algorithm or mathematical operation to the quantum circuit output. In some examples, the further classical output can be considered to be the output of the machine learning model being trained. In other words, the further classical output may be the final output.

[0106] In some embodiments, classical processor 111 may adapt 204 the quantum circuit output, rather than classical processor 161. For example, classical processor 161 may transmit the quantum circuit output to classical device 110, which in then received by classical processor 111. Classical processor 111 may then adapt 204 the quantum circuit output to determine a further classical output, rather than classical processor 161 adapting the quantum circuit output. Classical processor 111 may then transmit the further classical output to server 150 to be received by classical processor 161.

[0107] In some embodiments, the machine learning model comprises a further classical sub- model and classical processor 161 adapts the quantum circuit output by applying the further classical sub-model to the quantum circuit output to determine the further classical output. As such, the machine learning model being trained in method 200 may comprise two classical sub- models and a quantum sub-model. The further classical sub-model may be at least one layer. In the example of neural networks, the further classical sub-model may be a fully connected layer. The classical sub-model of the classical device 110 may be a similar architecture to the further classical sub-model. For example, both the classical sub-model and the further classical sub- model may have a similar neural network architecture. Preferably, the further classical sub- model is stored in the memory of classical part 160 of server 150.

[0108] Classical processor 161 then updates 205 the quantum sub-model based on minimising a loss involving the further classical output. Classical processor 161 may update 205 the quantum sub-model by updating or adjusting the weights that represent the quantum sub-model based on the loss. Classical processor 161 may determine the loss (which may be represented by a loss value) by applying an algorithm or mathematical operation to the further classical output. In some embodiments, classical processor 161 may receive a label being associated with the training data from classical device 110. The loss may be based on the further classical output and the label received from classical device 110. In other words, classical processor 161 may update 205 the quantum sub-model based on minimising the loss involving the further classical output and the label. For example, classical processor 161 may determine the loss based on a difference between the label and the further classical output.

[0109] The training data may comprise label data, such as image data that has been labelled with a class, where the label for each item of training data may be manually provided by a user. Training data may also or otherwise be obtained from a database containing already labelled data, for which manually labelling the data is not needed. Training the machine learning model in this way is known as “supervised learning”, in the sense that, the outcome of the machine learning model is validated, and the weights of the machine learning model are adjusted accordingly. In “supervised learning”, the goal is to find weights that minimise an objective function, which is some function of the label from the training data and the output of the machine learning model. In some examples, training the machine learning model, multiple machine learning models and / or the segmentation model comprises minimising a cross-entropy loss function. However, other types of loss functions are equally applicable here.

[0110] It should be noted that “supervised learning” is not the only method that can be used to train the machine learning models. For example, “semi-supervised learning” can be used in situations where only a relatively small amount of training data can be obtained. In some examples, a training set may be augmented to produce more training data, which is advantageous if a large training data is difficult to obtain. A training set may be augmented by randomly resizing, cropping, flipping, or rotating the image data in the training set, for example.

[0111] In another example, “unsupervised learning” may be used to train the machine learning models, which involves training the machine learning model using unlabelled training data. “Unsupervised learning” thereby enables the machine learning model to recognise patterns within the training data, then classify new image data into one of the recognised patterns during inference. Memory 112 or the memory of server 150 may store the machine learning model by storing the weights that have been optimised.

[0112] Finally, classical processor 161 transmits 206 a loss propagation value based on the updated quantum sub-model to classical device 110, to cause classical device 110 to update the classical sub-model based on the loss propagation value, to thereby train the machine learning model. In other words, classical device may be configured to update the classical sub-model based on the loss propagation value received from classical processor 161. The loss propagation value may be indicative of one or more gradients calculated during the process of updating 205 the quantum sub-model. Training of the machine learning model is indicative of backpropagation, in the sense that a loss is determined from the output of the machine learning model (i.e., the further classical output) which is then used to update the machine learning model in a backwards manner.

[0113] In some embodiments, classical processor 111 may perform the entire training of the classical sub-model and the quantum sub-model. For example, classical processor 111 may determine the further classical output and may calculate a loss based on the further classical output. As such, classical processor 111 may not transmit a label to server 150. Classical processor 111 may also understand the architecture of the quantum circuit representing the quantum sub-model. Hence, classical processor 111 may determine the updated weights (parameters) of the quantum sub-model to update 205 the quantum sub-model. This may be achieved through simulation of the quantum circuit by classical processor 111. Furthermore, classical processor 111 may calculate the loss propagation value from the updated quantum sub- model to update the classical sub-model.

[0114] In embodiments where the machine learning model comprises a further classical sub- model, classical processor 161 may update the further classical sub-model based on minimising the loss involving the further classical output. Classical processor 161 may then determine a further loss propagation value based on the updated further classical sub-model. Classical processor 161 may then update the quantum sub-model comprises updating the quantum sub- model based on the further loss propagation value. Training of the machine learning model is this way is also indicative of backpropagation.

[0115] In some embodiments, classical processor 161 and quantum processor 171 may communicate using encrypted communication, such as through homomorphic encryption or a symmetric-key algorithm. As such, preferably, only the quantum processor 171 can decrypt communication (such as the classical output) received from classical processor 161 and vice versa. In some examples, data that is communicated between classical processor 161 and quantum processor 171 may be encrypted in such a way that the encrypted data does not need to be decrypted by the receiving processor. More specifically, quantum processor 171 may transmit encrypt the loss propagation value in such a way that classical processor 161 update (i.e., train) the classical sub-model without decrypting the loss propagation value. Such encryption may be used for other values of the disclosed method. Multiple client devices

[0116] Similar techniques to those described above may be applicable to training a machine learning model using multiple sets of training data. In some embodiments, the disclosed training method can be used to train a machine learning model on multiple sets of training data provided by multiple clients. In particular, each of the multiple clients has a set of training data, but it is desired to keep the training data private and not share the training data among other clients.

[0117] Fig.3 illustrates system 300 for machine learning. Fig.3 is one example of a configuration of system 300. However, system 300 is not strictly limited to this configuration and this may be one possible embodiment of system 300. It is noted that system 300 of Fig.3 is only meant to illustrate an example and a preferred system which is capable of performing the disclosed methods or embodiments thereof. System 300 depicts server 150 from Fig.1, which may provide computing services to multiple clients. System 300 also comprises multiple classical devices 310, 320, 330, each associated with one or multiple clients, which may be similar to equivalent to classical device 110 of Fig.1.

[0118] In some embodiments of method 200, there may be multiple classical devices, 310, 320, 330, where each of multiple classical devices 310, 320, 330 comprises a classical sub-model. As such, classical processor 161 of server 150 may, for each of multiple classical devices 310, 320, 330, repeat the steps of: receiving 201 the intermediate classical output from the classical device, configuring 202 the quantum gates, executing 203 the quantum circuit, adapting 204 the quantum circuit output, and updating 205 the quantum sub-model. Classical processor 161 may then transmit 206 the loss propagation value based on the updated quantum sub-model to the each of multiple classical devices to cause multiple classical devices 310, 320, 330 to update of the classical sub-model on each of multiple classical devices 310, 320, 330 based on the loss propagation values. The classical sub-model of each of multiple classical devices 310, 320, 330 may be the same classical sub-model. However, classical processor 161 may only transmit the respective loss propagation value to the corresponding classical device. In other words, system 300 may be used for collaborative and distributed learning, and may also be used for individual distributed learning.

[0119] In some embodiments, each of the multiple classical devices 310, 320, 330 stores respective training data. In particular, the training data may never leave the respective classical device. Hence, in the above embodiment, the machine learning model may be trained on the training data of each of the multiple classical devices 310, 320, 330 while isolating the respective training data from other ones of the multiple classical devices 310, 320, 330. For example, the respective training data of each of the multiple classical devices 310, 320, 330 may be confidential (or sensitive) training data. Therefore, the embodiments of method 200 involving multiple classical devices 310, 320, 330 enable a machine learning model to be trained on multiple (isolated) sets of training data, while utilising the advantages provided by quantum computing. Such embodiments also enable the multiple classical devices 310, 320, 330 to be resource-constrained devices, as previously described. Training by classical device 110

[0120] Fig.4 illustrates method 400 for training a machine learning model. Fig.4 is to be understood as a blueprint for a software program and may be implemented step-by-step, such that each step in Fig.4 is represented by a function in a programming language, such as, but not limited to, Python, C++, or Java. The resulting source code is then compiled and stored as computer-executable instructions on non-volatile memory, which causes a processor (or multiple processors or a distributed computing architecture) to perform method 400. Similar embodiments and examples described above in reference to method 200 are equally applicable to method 400.

[0121] Classical processor 111 applies 401 a classical sub-model to training data to determine an intermediate classical sub-model. The intermediate classical output represents an evaluation of the classical sub-model on training data. Classical processor 111 then transmits 402 the intermediate classical output to server 150, which may be received by classical processor 161.

[0122] As previously mentioned, classical device 110 may be configured to configure quantum gates of a quantum circuit based on the intermediate classical output, execute the quantum circuit using the quantum gates to determine a quantum circuit output, adapt the quantum circuit output to determine a further classical output; update the quantum sub-model based on minimising a loss involving the further classical output; and transmit a loss propagation value based on the updated quantum sub-model to classical device 110. As such, classical processor 111 of classical device 110 may cause server 150 to perform method 200 by transmitting the intermediate classical output.

[0123] Classical processor 111 then receives 403 the loss propagation value from server 150 and updates 404 the classical sub-model based on the loss propagation value, to thereby train the machine learning model. In particular, receiving the loss propagation value from server 150 may cause classical processor 111 to update 404 the classical sub-model based on the loss propagation value, to thereby train the machine learning model. Method for classifying

[0124] Fig.5 illustrates method 500 for classifying test data using a trained machine learning model. Method 500 is performed by a classical device, such as classical device 110 of Fig.1. As such, for explanatory purposes, some aspects of method 500 will be explained with reference to classical processor 111 of classical device 110. However, it is noted that the aspects of method 500 may be performed by other components of system 100, such as classical part 160 of server 150, for example.

[0125] In method 500, the trained machine learning model comprises a classical sub-model and a quantum sub-model. As such, the trained machine learning model be considered to be a hybrid quantum machine learning model. The trained machine learning model of method 500 may betrained using method 200, 400, for example. As such, similar embodiments and examples described above in reference to method 200, 400 are equally applicable to method 500.

[0126] Fig.5 is to be understood as a blueprint for a software program and may be implemented step-by-step, such that each step in Fig.5 is represented by a function in a programming language, such as, but not limited to, Python, C++, or Java. The resulting source code is then compiled and stored as computer-executable instructions on non-volatile memory, which causes a processor (or multiple processors or a distributed computing architecture) to perform method 500.

[0127] Classical processor 111 evaluates 501 the classical sub-model on the test data to determine a classical output. The test data may be any type of data that may be classified. For example, the test data may an image (such as a two-dimensional, three-dimensional or multi- dimensional image). In some examples, the test data may be text, such as plain English text.

[0128] Classical processor 111 then applies 502 a noise layer to the classical output to determine a noisy output. The noise layer is configured to add noise to an output of the classical sub-model. The noise added by the noise layer may be random, pseudo-random or based on an algorithm, mathematical operation or mathematical model. Although classical processor 161 of server 150 may be configured to apply a noise layer to the classical output of classical sub- model, preferably classical processor 111 applies 502 a noise layer to the classical output to determine a noisy output. Applying the noise layer to the classical output before transmitting the output to server 150 may prevent an active attacker who intercepts the output from reconstructing the original input.

[0129] In some embodiments, the noise layer is configured to add Gaussian noise to the output of the classical sub-model. Gaussian noise is a type of random noise whose amplitude follows a Gaussian (normal) distribution. Moreover, in some embodiments, the noise layer is configured to add Laplacian noise to the output of the classical sub-model. Laplacian noise is a type of random noise characterized by its probability density function (PDF), which may be referred to as a Laplace distribution, which is similar in shape to the Gaussian distribution but with heavier tails.

[0130] The noise (such as the Gaussian or Laplacian noise) may be parameterised. Forexample, the Laplacian noise may be parameterised by two values: ^ and b , which correspondto the mean and a scale parameter (related to the spread or width) of the Laplacian distribution. The noise may be parameterised randomly. However, as will be discussed later in this disclosure, there may be preferred parameters. Hence, in some embodiments where the noise added by thenoise layer is Laplacian noise, ^ may be within the range of about 0 to 4^ , and b may bewithin the range of about 0.01 to 1. In particular, it was found that a mean of about ^ = 4 ^maximizes the difference between the original and reconstructed images during the experiments described herein.

[0131] Classical processor 111 then sends 503 the noisy output to server 150. Server 150 (more specifically, classical processor 161) may then configure 504 quantum gates of a quantum circuit based on the noisy output, execute 505 the quantum circuit using the quantum gates to determine a quantum circuit output, the quantum circuit being configured to represent the quantum sub- model, adapt 506 the quantum circuit output to determine a classification output representing a classification of the test data, and send 507 the classification output to classical device 110 to classify the test data.504-507 of method 500 may be similar to 202-204, 206 of method 200.

[0132] As will be discussed later in this disclosure, the machine learning model may still perform the classification of the test data, despite the noisy output. In other words, the machine learning model classifies the test data as if there were no noise layer. This is a result of the quantum sub-model. In particular, when using noise, such as Gaussian or Laplacian noise, the quantum state encoded in the quantum circuit from the noisy output is expected to be close to the quantum state corresponding to its clean version (i.e., the classical output), which would not be the case for a classical sub-model. As such, server 150 (more specifically, classical processor 161) may not need to determine the original classical output from the received noisy output. However, in some examples, classical processor 161 may determine the original classical output from the received noisy output by applying an algorithm, mathematical operation or machine learning model to the received noisy output.

[0133] The classification output may be a numerical value, indicative of a classification of the test data, such as a class. For example, the trained machine learning model may be a classifier trained to predict a number (between 0-9) from an image of a handwritten number. Hence, the trained machine learning model may output a numerical value between 0-9 as the classification output, which represents the predicted class. For example, the classification output may be 7.23, which indicates that the trained machine learning model has classified the test data (in this case, an image of a handwritten number) as the number 7. Experimental setup

[0134] Experiments to test the performance of the disclosed methods will now be discussed. The following sections discuss the architecture of the machine learning models used during the experiments described herein. It is noted that the following descriptions of the disclosed methods are only embodiments. Hybrid Quantum Split Learning (HQSL) Architecture Hybrid Quantum Split Learning Model

[0135] In the setup proposed herein, the labels and smashed data are transmitted to the server- side model at the split layer, while the raw data remain on the client side. During backpropagation, the gradients are transmitted from the server side to the client side across the split layer. The quantum layer is introduced as the first layer of the server-side model of split learning. The quantum layer consists of a quantum circuit with classical features as input and output. The classical inputs are encoded to quantum states and processed by the quantum circuit. Expectation measurements of the resultant quantum states for each qubit lead to classical features output from the quantum layer. These are then fed to the next classical layer. The quantum circuit is designed to be of practical importance in the near-term quantum computing era. The general structure of HQSL is shown in Fig.6.

[0136] Fig.6 shows the Hybrid Quantum Split Learning (HQSL) model used in theexperiments described herein with N clients and 1 hybrid quantum server. The clients have aclassical model (assumed to be the same for all clients), and the server model has a quantum layer as its first layer.

[0137] Construction and Selection of Quantum Circuit

[0138] To construct the quantum circuit used in HQSL’s quantum layer, the limitations of current-generation quantum computers were considered (see Appendix A for details). Specifically, the circuit width (number of qubits) and depth (number of gates per qubit) were considerations as it is desirable to keep these as small as possible. The presence or absence of entangling layers is also another consideration. Different quantum circuit configurations were trialled by adjusting the number of qubits, entangling gates, and encoding type. Each configuration was tested by incorporating the circuit as a quantum layer in the proposed HQSL model. The quantum circuit was selected based on the comparison of accuracy and F1-score between HQSL and its classical analogue. The circuit that consistently outperforms split learning in all experiments described herein was chosen and this is detailed in the Appendix C. This circuit is shows in Fig.7b, which is referred to as Circuit 6 later is this disclosure.

[0139] Fig.7a shows the qubit-efficient data loading scheme used in the experiments described herein, which comprises layers of encoding and parameterized (Param) gates. Fig.7b shows the quantum circuit utilized in the experiments described herein, which comprises RX gates that serve as data-loading points (i.e., the encoding layer as seen in Fig.7a). The quantum circuit shown in Fig.7b was designed to make efficient use of qubits while loading data to the quantum circuit. The quantum circuit of Fig.7b may be referred to as the optimal quantum circuit in this disclosure. However, this does not imply that this is the only quantum circuit that can be used inthe machine learning model architecture. This simply means this quantum circuit was found to be the best performing circuit from the circuits considered in the experiments described herein.

[0140] The quantum circuit with the optimal performance is shown in Fig.7b consists of 2qubits with 3-dimensional classical input features (x 1 , x 2 , x 3 ) each encoded at 3 different pointsalong each qubit using RX-gates. As can be seen in Fig.7a, each data loading point follows a parameterized rotation gate. More specifically, the RX-gates of Fig.7b provide the encoding layer shows in Fig.7a, while the RY-gates and RZ-gates of Fig.7b provide the parameterised layer of Fig.7a. As can be seen in Fig.7a, the rotation gates and the entangling gates of the quantum circuit are alternating in the quantum circuit.

[0141] CZ-gates provide entanglement between qubits. This qubit-efficient data-loading circuit was developed while keeping the circuit depth low. The construction of this circuit is discussed below. It is noted that this proposed circuit provides better classification performance than deeper circuits (three layers) with other data re-uploading techniques (see Table 3 in Appendix C). At the measurement stage, the expectation value is computed at the output of the quantum circuit using the Pauli-Y operator on each qubit.

[0142] The proposed qubit-efficient data loading methodology for quantum circuit design can be extended beyond the scope of HQSL, presenting a versatile framework with broad applicability. In addition, given that it is only a 2-qubit low-depth circuit, using this circuit in HQSL gave a relatively low runtime compared to the other circuits trailed. This circuit is anticipated to be feasible for practical implementation on error-corrected quantum devices in the near future (such as low qubit fault tolerant devices), from its low-qubit and low circuit depth requirements. For simulation purposes using the PennyLane library (https: / / pennylane.ai / ), this quantum circuit is integrated as a quantum layer for constructing HQSL.

[0143] Construction and Selection of Quantum Circuit

[0144] The data re-uploading mechanism used in this disclosure consists of multiplerepetitions of a layer, L i , consisting of the generic single-qubit rotation gate, U^ SU (2) , for theencoding layer as U( X ) and for the parameterized layer as U(^ ), where X = ( x 1 , x 2 , x 3 ) and^ i = ( ^ i1 , ^ i2 , ^ i3 ) , where ^ i represents the parameters for layer L i . Thus, each layer can berepresented as Li = U ( X ) U ( ^ i ) . L i , for i = 0,..., N , is repeated giving rise to the data re-uploading mechanism, and the single-qubit classifier is introduced, given enough repeats, N , areperformed.

[0145] The U-gate consists of parameters ^ =(^ 1 , ^ 2 , ^ 3 ) and can be expressed asU(^ 1 , ^ 2 , ^ 3 ) ^ SU (2) . The U-gate may be given the following decomposition:

[0146] The above decomposition of this layer, L i , may be adopted while taking intoconsideration that the width and depth of a circuit should be kept as small as possible, given the limitations due to current quantum devices. Thus, a circuit consisting of only a single layer, with 2 qubits, and without any entangling gates was initially considered. The circuit is represented as Circuit 7 in Fig.8.

[0147] However, entangling gates may improve the classification performance of the circuit.To introduce entanglement within the circuit consisting of a single layer,, the U -gate isdecomposed into RZ− RY − RZ gates and the circuit was reorganised such that, for each qubit, afeature, x k is uploaded for k^ 1,2,3 , corresponding to a 3-dimensional input, using rotationgates and followed immediately by a single-qubit parameterized rotation gate of the same type, and a CZ entangling gate between the 2 qubits. This is repeated until all the features have been uploaded. The CZ-gate at the end of the circuit is omitted. This circuit is given as Circuit 8 and present in Fig.9.

[0148] The performance of both these circuits were test when utilized in the quantum layer of HQSL. Their performance was compared for the classification of FMNIST dataset using the accuracy and F1-score metrics. The result is shown in Table 1 presented as Fig.10. The experiment shows that this modification improves the accuracy and F1-score during classification of the FMNIST dataset.

[0149] A final modification to the current circuit (i.e., Circuit 8 shown is Fig.9) to develop Circuit 6, which is the circuit shown in Fig.7b. Specifically, the embedding gates for loading(x 1 , x 2 , x 3 ) to the circuit are replaced with RX-gates, as shown in Fig. 7b. This change in theembedding gates further improves the classification performance as shown in Table 1.

[0150] The last two rows of Table 1 show the performance of when the data re-uploadingtechnique is adopted in a 2-qubit circuit with 3 layers, L , with (Circuit 9) and withoutentanglement (Circuit 10) respectively. Here, CZ-gates are used after each layer, except the last layer. It is shown that the qubit-efficient data loading circuit (Circuit 6) disclosed herein still provides a better accuracy and F1-score than a 3-layer deep circuit using the data re-uploading scheme.

[0151] Hence, the proposed circuit (i.e., Circuit 6) keeps the training time during simulations within reasonable limits, is more likely to be feasible on near-term devices, and provides a better classification performance than a 3-layer deep circuit that uses the data reuploading technique.

[0152] Given the chosen quantum circuit 6, further experiments were performed to compare the accuracy of HQSL in classifying the FMNIST dataset when the number of qubits and the depth of the circuit is varied. Increasing the number of qubits is simply increasing the width of the circuits such that the expectation measurement output has a dimension equal to the number of qubits. Entanglement via CZ-gates is provided between pairs of qubits such that every qubit in the circuit is entangled with another qubit only once. This is alternated as next layer is switched. An illustration of this is given by comparing Circuits 1 and 2 in Table 1 presented in Fig.10.

[0153] The proposed circuit (i.e., Circuit 6) is ‘qubit-efficient’ as, in theory, it enables one to load infinitely large dimensional data to each qubit (this comes at the cost of an infinitely deep quantum circuit and, hence, an unbounded simulation runtime). However, here, the circuit depth is restricted by only considering 3-dimensional data being loaded onto a 2-qubit quantum circuit. By imposing this restriction, the simulations are kept within reasonable runtimes, while still providing high performance during classification.

[0154] To demonstrate this, experiments with Circuit 6 were carried out by increasing the number of qubits and circuit depth. The resulting accuracy, and simulation time taken for 5-fold training can be seen in Fig 11. In Fig.11, ‘shallow’ refers to a low-depth circuit with 3 re- uploading points, while ‘deep’ corresponds to a circuit twice as deep as the shallow one, i.e., 6 data-loading points. Hence, a shallow circuit would require 3-dimensional classical features, while a deep circuit would require a dimension of 6.

[0155] Fig.11 shows the effect of the number of qubits and circuit depth on the mean test accuracy and 5-fold training time of HQSL on Fashion-MNIST dataset. Accuracy improves only slightly as the number of qubits and circuit depth increase at the expense of rapidly increasing training time.

[0156] It is highlighted here that although as the number of qubits and circuit depth increased, accuracy increased marginally, the runtime during simulation increases very rapidly (comparing the shallow circuit with 2 qubits and the deep circuit with 8 qubits, an improvement of only around 1.5% causes the time taken for training to increase by more than 500 hours). It is also highlighted that the deeper (number of layers) and wider (number of qubits) a circuit is, the more difficult it is to implement on currently available quantum computers. For these reasons, it is desirable to keep the quantum circuit shallow and narrow (3 data-loading points and 2 qubits), while still obtaining high accuracy. Hence, the proposed quantum circuit keeps the training time during simulations within reasonable limits, is more likely to be feasible on near-term devices, and provides a better classification performance than a 3-layer deep circuit that uses the data reuploading technique.

[0157] Classical Benchmarking of HQSL

[0158] To benchmark the performance of HQSL, it was desirable to create an equivalent classical model. To achieve this, the classical counterpart of the selected quantum circuit (Circuit 6) was constructed as follows: The classical counterpart consisted of a dense layer with the same number of output neurons as the quantum layer – two output neurons with ReLU activation units (the best-performing activation unit based on the experiments described herein). The number of input neurons in this layer corresponded to the dimension of the classical input (3 input features). The aim of this benchmarking was to compare the compact quantum layer to a simple classical dense layer, enabling a fair comparison, taking into consideration the size limitations of quantum circuits. In Appendix C, a comparative analysis of the performances of different quantum circuits is provided, in comparison to their corresponding classical counterparts when introduced as a layer in the split learning models. This experimental analysis enabled the selection the quantum circuit to be used in HQSL, as described earlier.

[0159] HQSL Model Variants

[0160] Two variants of HQSL were used in the experiments described herein. These variants were designed for multivariate binary and multi-class single-channel image dataset classification. These variants were designed by constructing a hybrid quantum neural network (HQNN) and subsequently dividing it at the split point (see Figs.12a and 12b) so that the initial part of the HQNN represents the client-side (classical neural network layers) and the remaining segment functions as the server-side (hybrid quantum neural network layers). However, it is noted that other machine learning model architectures are possible with the disclosed methods.

[0161] Fig.12 shows the HQSL variants used in the experiments described herein. The split point splits the model into the client-side portion (left) and server-side portion (right). Fig.12a shows HQSL variant 1, which was used for binary classification on multivariate datasets. Fig.12b shows HQSL variant 2, which was used for n -class classification on single-channel imagedatasets. The number of output nodes, n , can be set depending on the number of classes in thedataset.

[0162] HQSL variant 1: The client-side model consists of 4 fully connected dense layers with ReLU activation functions with a 7-dimensional input and outputs smashed data of dimension 3. For the server-side model, the quantum layer was constructed using the optimally chosen quantum circuit shown in Fig.7b, which has an input dimension of 3 and an output dimension of 2.4 additional fully connected classical dense layers were concatenated with the ReLU activation function (except the last layer, which uses Sigmoid) with an output dimension of 1 (for binary classification). Dropout layers were applied on the first two classical layers on theserver-side model to improve regularization. For the classical counterpart of this variant, the quantum layer was simply replaced with the classical dense layer as described earlier.

[0163] HQSL variant 2: The client-side model consists of two Convolutional Neural Network filters, each with ReLU activation function and a MaxPool2d Layer, with a stride of 2 and kernel size of 2, followed by a flatten layer, and 2 fully connected layers to bring the feature dimensions down to 3. This is the client-side portion of the model. For the server side, the selected quantum circuit shown in Fig.7a was used as a quantum layer and it was concatenated with 3 more fullyconnected classical dense layers with an output dimension of n (for n -class classification).ReLU activation function was then applied on the first 2 classical layers on the server side. For the classical version, the quantum layer was simply replaced with its equivalent classical dense layer, acting as a benchmark for this HQSL variant.

[0164] Training the Hybrid Quantum Split Learning Model

[0165] For the classical layers of HQSL, gradients are computed from the loss function, L ,using the classical backpropagation algorithm via the chain rule. However, it becomes non-trivial to do the same for the quantum layer. An algorithm, which permits backpropagation through the quantum layer, is discussed in the following. This technique may be referred to as a parameter- shifting algorithm.

[0166] Consider a quantum circuit, fqwhere f q is parameterized bya set of k free parameters, ^={^ 1 , ^ 2 ,..., ^k } , and maps input, Z = ( z 1 , z 2 ,..., zM ) to themeasurement outputs, E = ( e 1 , e 2 ,..., eN ) . In HQSL, the quantum layer can be considered as beingsandwiched between two sets of classical layers: the client-side model, fc : X Z(f in Mc :R → R ) , and the server-side classical layers, fs : E Y ˆ, where Ŷ isthe output of HQSL.

[0167] The derivatives of the expectation value of the measurement, E , of the quantum circuit,f q , with respect to (w.r.t.) the gate parameters, ^ , are computed by evolving the circuit twiceper parameter, with a (^ ) shift in that parameter. Representing the partial derivative off (^ , Z ) = E , w.r.t ^ ^ {1,2,...,k }^ f q q. parameter i for i as, the parameter-shift rule can beexpressed as:where r is a shift constant. This gives us the exact gradient w.r.t. each free parameter in thequantum circuit. Similarly, the derivative w.r.t. input,z j can be computed for j^ {1,2,..., M } ,to the quantum circuit:

[0168] In summary, for the quantum layer, during forward propagation, the input classical features are fed to the quantum layer, and the quantum circuit is evolved for a default of 1000 shots. At the end of the quantum circuit evolutions, an expectation value for the measurement is calculated as per Eq.10 in Appendix A. These calculated values are then used as classical features for the next classical layer. During backpropagation, to compute the derivatives w.r.t.parameter ^ 1 , ^ 2 ,..., ^ k } and inputs zi^{ z 1 , z 2 ,..., z M } , two evaluations of the circuit arecomputed with a ^ shift in the parameter ^ i . These two computations are then added accordingto Eq. 2 and Eq. 3. This is repeated until the derivatives w.r.t. all ^ i and z j are computed. Using^ fthe computed derivatives, the gate parameters, ^ , are updated accordingly. The derivativesq ^ z j represent the derivatives w.r.t. to the input to the quantum layer and are transmitted across thesplit layer to the client side model, to continue the backpropagation of the loss function, L , tothe client-side classical layers. The HQSL training is summarised in Algorithm 1. In the ^(^ algorithm,)represents the Jacobian matrix that stores the element-wise gradients of the vector in the numerator with respect to the vector in the denominator. Algorithm 1: Hybrid Quantum Split Learning Algorithm with a single classical client and hybrid quantum serverInput: Number of training rounds n ^ {1,...,N } , learning rate ^Output: Optimal weights, W = ( W C , W S ) , for classical layers and optimal parameters, ^ , for quantumlayerNotations: Client’s inputs: X , client’s labels: Y , predicted labels: Ŷ , client’s smashed data: Z , quantumlayer’s expectation measurement output: E , loss function: LInitialization: Initialize client-side layers ( f c ) weights, W C ; server-side’s quantum layer ( f q )parameters, ^ , and classical layers ( f s ) weights, W S .START OF TRAININGforn= 1, ^ , N Training Epochs doClient side: / / Forward Pass Z= f ( Cc X , W n )Send Z and Y to the server……………………………………………………………………………………….Server side: / / Forward Pass, Compute Less, Backward Pass A, ,ComputeCompute gradient w.r.t. toWnSas^ WS= ˆ ^S / / Chain rule n^Y ^ W n^ˆ Compute gradient w.r.t. to E as L ^ L ^Y ^E= ^ / / Chain rul ^ ˆ^ Ee Y Update server classical layers weightsW S as: WS S^ L n+1^ W n −^ ^ WSnCompute gradient w.r.t. n as =^ ^^ / / Chain rule n^E^ ^ nCompute gradient w.r.t. to Z as^ L^ Z = ^ L ^E ^E^ ^Z / / Chain rule date Quantum Layer Parameters as:n + 1 ^Up^ n −^ ^ ^ n Send ^ L to client side to continue backward pass ^ Z ………………………………………………………………………………………. Client side / / Backward Pass ComputeCompute gradient w.r.t.WnCas ^ L= ^ L ^Z ^WC ^ Z^ ^C / / Chain rule nW ndel weightsW C as: WCn +1^ W CUpdate client mon −^ WCn End of Epoch n END OF TRAININGreturn W = ( WC , W S ), ^Multiple client setup and training of Hybrid Quantum Split Learning

[0169] It is desirable to scaling HQSL to accommodate a larger number of clients as it would enable multiple clients to leverage the quantum resources at the central server. The method usedto scale HQSL with K clients in the experiments described herein follows the basic round-robinprotocol and is as follows:1. Initialize the client, W C , and the server side (^ ,WS ) model parameters.2. Randomly choose a client k and train it in collaboration with the server. This consists ofone round of forward and backward propagation and makes 1 local epoch for client k .3. The updated client k model weights W C^are then sent to the next client (k ' ) thatupdates its model weights before another round of forward and backward propagation with the server. This marks the end of client k + 1 ’s local epoch.4. After all K clients have been served in that 1 global epoch, the protocol then moves tothe next global epoch and resume training client k .

[0170] The differences due to the presence of the quantum layer in HQSL are as follows: A forward and backward propagation round consists of encoding classical smashed data within the quantum layer. At the measurement stage, a default of 1000 shots were used to sample the expectation value of the measurements due to the probabilistic nature of quantum circuits. Finally, the parameter-shifting algorithm described earlier for the backpropagation algorithm within the quantum layer. The method for including multiple clients is summarised in Algorithm 2.Algorithm 2: Hybrid Quantum Split Learning Algorithm with K Classical Clients and Hybrid Server(Initialization Phase)Input: Number of clients K , training rounds n^ {1,..., N } , learning rate ^Output: Optimal weights, W = ( WC , W S ) , for classical layers and optimal parameters, ^ , for quantumlayerNotations: Client k ’s inputs: X k , labels: Y k , predicted labels: Yˆk , smashed data: Z k , quantum layer’soutputs when client k is being served: E k , loss function: LInitialization: Initialize client-side layers ( f c ) weights, W C ; server-side’s quantum layer ( f q )parameters, ^ , and classical layers ( f s ) weights, W S .Algorithm 2: Hybrid Quantum Split Learning Algorithm with K Classical Clients and Hybrid Server(Training Phase) START OF TRAININGforn= 1, ^ , N Training Epochs doStart of global epoch nfor each client k , k^ {1,...,− 1} (mod k )Start of client k ’s local epochClient k side: / / Forward PassZC k= f c ( X k , W n , k ) ; Send Z k and Y k to server………………………………………………………………………………………. Server side: / / Forward Pass, Compute Loss, Backward Pass Quantum Layer:Ek = f q ( Z k , ^ n , k ) / / Compute E k as per Eq. 10Classical Layers:Y ˆk = f s ( E k , WSn,k )Compute Loss, Lk = L (Y k , Y ˆ k ) , and gradient^ Lk ^Yˆ k^Y ˆ^fS s( E k , W n k ) ^Y ˆ^fS s( E k , W n,k )Computek, ank= ^WS=Sd n, k ^ W n , k ^ E k ^ E kdient^ L ^ L ^ˆ pute grak kY Comk^W S= ^S / / Chain rule n, k ^ Y k ^ W n , kient^ L ^ L ^Y ˆ Compute gradk k k^ E= ^ / / Chain k^Yˆk ^ Erule k2.3 Compute gradient / / Chain rule Compute gradientChain rule Update^ as: ^ n ,Sendto client k to continue backward pass^Zk………………………………………………………………………………………. Client k side: / / Backward PassCompute^Zk^ WC= n, kCompute gradient Chain ruleUpdate client k model weights W C as:End of client k ’s local epochSendW C^ =W Cn , k+1 to client k ' and start local epoch for client k 'End of global epoch nStart global epoch n+ 1 with client k with updated weights,, WS n+ 1, k , and ^n+ 1, kEND OF TRAININGreturn W = ( W C , W S ) , ^Experiments and Results

[0171] In this section, the feasibility and scalability of HQSL was empirically assessed by comparing its performance in classification tasks to that of their corresponding split learningmodels. All programs were written using Python 3.11.3 and PyTorch 2.1.0 libraries. The quantum part of the HQSL models were simulated using the Pennylane library with its PyTorch backend. The experiments were conducted on an NVIDIA GeForce RTX 2080 Ti GPU machine system.

[0172] Five publicly available datasets were used to test and validate the HQSL architectures. These datasets are summarised in Table 2. Detailed descriptions, including dataset processing methods, can be found in Appendix B. Setting a fixed seed, the datasets were split into five non- overlapping folds, in preparation for five-fold cross-validation for the experiments described herein. Specifically, four of the folds were used for collectively training the machine learning models, and the remaining fold was used to test the trained machine learning models. For the multi-class image datasets (MNIST and FMNIST), the datasets were split in a stratified manner to ensure an even and balanced distribution of classes. Dataset HQSL model Training Testing# Features / # Classesvariant Samples Samples Image Size Botnet DGA 1 800 200 7 2 Breast Cancer 1 455 114 7 2 MNIST 2 4800 1200 28 x 28 10 FMNIST 2 4800 1200 28 x 28 10 Speech Commands 2 6388 1597 28 x 28 2 Table 2: Datasets used for the experiments. Hybrid Quantum Split Learning versus Classical Split Learning: Single Client

[0173] Experiment

[0174] The quantum circuit shown in Fig.7b was built using PennyLane’s default-qubit device, which was converted to the quantum layer of the machine learning model. HQSL variant 1 shown in Fig.12a was constructed for training on Botnet DGA and Breast Cancer datasets and HQSL variant 2 shown in Fig.12b was constructed on MNIST, FMNIST and Speech Commandspectrograms datasets. In HQSL variant 2, the number of output nodes was set to n =10 (10classes) for the MNIST and FMNIST datasets and n = 2 (2 classes) for the Speech Commandsspectrograms dataset. Each of these 5 experiments were paired with their classical counterparts to benchmark the performance of the HQSL model variants disclosed herein.

[0175] For the hyperparameters, the Adam optimizer was utilised with a learning rate of 10− 3across all experiments, except for the Speech Commands dataset, where a learning rate of 10− 4provided better convergence. A binary cross entropy loss function was used for classifying the 2 multi-variate datasets and the cross-entropy loss function for the 3 image classification tasks. The models were trained as per Algorithm 1 (i.e., method 200) for 100 epochs for Botnet DGAand Breast Cancer datasets, and 50 epochs for MNIST, FMNIST and Speech Commands spectrograms datasets. Across all pairs of experiments, the random seed was set to be a constant and fixed the training and testing dataset for each of the 5 folds. These tests were also performed on the centralized (unsplit) versions of the split learning models, and the results are reported next.

[0176] Results and Discussions

[0177] Fig.13a shows the performance comparison in terms of test accuracy of split learning (white) versus HQSL (black) with 5-fold cross-validation. Fig.13b shows the performance comparison in terms of F1-score of split learning (white) versus HQSL (black) with 5-fold cross- validation. The horizontal dashed lines spanning the width of each box represent the mean test results. In every case, HQSL outperforms or is nearly on par with split learning. On the FMNIST and Speech Commands datasets, a mean improvement of approximately 3% and 1.5% respectively due to HQSL is observed. The test results for the split and centralized HQSL versions are identical.

[0178] Although marginal, improvements in both mean accuracies and F1-scores of HQSLwere obtained on the Botnet DGA (mean accuracy improvement: ^ 0.1% , mean F1-scoreimprovement: ^ 0.4% ), Breast Cancer (^ 0.6% , ^ 0.7% ) and MNIST (^ 0.4% , ^ 0.4% )datasets. Significant outperformance in terms of mean accuracy and F1-score were obtained onthe FMNIST (^ 3% , ^ 4% ) and Speech Commands (^ 1.5% , ^ 1.5% ) datasets. For all of thefive datasets used in the experiments described herein, HQSL proved feasible as it at least matched the testing performance of split learning in terms of accuracies and F1-scores. As demonstrated by the results on the FMNIST and Speech Commands spectrograms datasets, HQSL has the potential to achieve higher accuracy and F1-score compared to its classical counterpart. It is noteworthy to highlight that with just the introduction of 1 quantum layer consisting of a small 2-qubit quantum circuit, HQSL can slightly outperform classical split learning, showcasing the potential power of quantum computation beyond the NISQ-era.

[0179] It is also noteworthy that simulating a quantum computer using a classical computer is computationally intensive, despite having only 1 quantum layer consisting of a simple quantum circuit on the server side of HQSL. The training time for HQSL was significantly longer than that of split learning, depending on the size of the dataset. For instance, 5-fold training for 100 epochs each with 569 data points from the Breast Cancer dataset split in the train: test ratio of 4:1, took approximately 1 hour to run on HQSL. On the other hand, 5-fold cross-validation with 6000 image-label pairs from the MNIST dataset took around 9 hours to complete 50 trainingepochs of HQSL. In contrast, their classical analogues only took a few minutes to train, which is significantly faster than HQSL.

[0180] Key takeaway: In the single client case, HQSL with a quantum layer consisting of only a small quantum circuit is feasible and offers tangible improvements in classification accuracy and F1-score compared to split learning. Hybrid Quantum Split Learning versus Classical Split Learning: Multiple Clients

[0181] Experiment

[0182] To assess the impact of the number of clients K on model accuracy and F1-score, thetraining set was divided into K independent and identically distributed (IID) subsets forK = 2,3,4,5,10,20,50,100. Each subset represents an IID distribution of the dataset assigned toeach client, ensuring a balanced portion for each client.

[0183] The training process follows Algorithm 2. The hyperparameters used are the same as in the case of a single client in the previous section. At the end of each global epoch, the model was evaluated on the testing set. This training process was repeated for 100 global epochs on the multivariate datasets (Botnet DGA and Breast Cancer) and 50 global epochs on the image datasets (MNIST, FMNIST, and Speech Commands spectrograms). The test results are then compared to those of their classical equivalent. The reported results include the test accuracy and F1-score at the end of training comparing HQSL against split learning. These tests are conducted for only one-fold in this set of experiments and the test accuracies and F1-scores are reported to evaluate the scalability of HQSL compared to split learning.

[0184] Results and Discussions

[0185] Fig.14a shows the test accuracy of HQSL compared with split learning on all 5 datasetsfor up to K =100 clients. Fig. 14b shows the F1-score performances of HQSL compared withsplit learning on all 5 datasets for up to K =100 clients. As can be seen from Figs. 14a and 14b,HQSL maintains a high performance like split learning while accommodating multiple clients.

[0186] For both HQSL and split learning, the accuracy and F1-score do not show a significantdrop in performance across all datasets as the number of clients, K , are increased. The resultsuggests that HQSL can match, and in some cases potentially surpass, the performance of split learning as the client count increases. This finding underscores the viability of HQSL. The comparable performance between HQSL and split learning across increasing client numbers is a significant finding, opening up new avenues for quantum-enhanced distributed learning in real- world applications.

[0187] An independent and identically distributed (IID) dataset distribution was used amongthe K clients to ensure an even and balanced distribution of the dataset for each client.Moreover, only the round-robin communication protocol was considered for training each client with a centralized server. However, other method may be used, such as asynchronous training. In addition, increasing the number of clients increases the training latency for a particular client in the network, which is further exacerbated by the round-robin communication protocol used.

[0188] Key takeaway: HQSL accommodates multiple clients without losing its performance significantly. The performance trend is similar to that of the split learning counterpart. Strengthening the Security of HQSL against Reconstruction Attacks

[0189] Considering the rotational properties of the encoding gates in the quantum layer, a noise-based defence mechanism is proposed in the form of a Laplacian noise layer. This layer adds randomness to the smashed data before they are transmitted to the server-side model, making it more difficult for the reconstruction attack models to recreate private input raw data. Next, experiments are conducted under different noise parameters to extract the best-performing noise layer whereby HQSL has a clear advantage over split learning in two key aspects: (i) impairing the performance of the reconstruction attack models and (ii) maintaining a high classification performance despite the presence of the noise defence layer. In these experiments, reconstruction attacks are invested considering the image datasets previously described, specifically, MNIST, FMNIST, and the spectrograms from the Speech Commands datasets. Reconstruction Attack from Smashed Data

[0190] Threat Model

[0191] In the experiment described herein, it is assumed that a server that collaborates with multiple clients or data owners according to the split learning protocol. The clients do not share their private dataset,with the server or other clients in the network. it is also assumed thatthe server is honest-but-curious (semi-honest), i.e., the server does not deviate from the specified protocol instructions but attempts to infer information about the client’s model and data. Further,it is assumed that the server has access to a publicly available auxiliary dataset, D aux , that has asimilar distribution as, butand D aux are non-overlapping disjoint datasets. Theadversary can have access to a ‘shadow’ model, f shadow , to generate outputs, Z aux , in the samefeature space as the smashed data, Z , received from a client. The shadow model is trained on theauxiliary dataset, D aux . The server then uses the pair (Zaux = f shadow ( X aux ), X aux ) , whereX aux ^ D aux to train a reconstruction attack model, f rec , that recreates X aux .

[0192] With the trained reconstruction attack model, f rec and the smashed data, Z , generatedfrom the private datasets,D priv , the honest-but-curious server reconstructs the private input ofthe clients. It is assumed that the reconstruction attack happens during deployment / inference time, i.e., the adversary attempts to recreate the data clients provides during inference time.

[0193] Reconstruction Attack Models

[0194] 3 reconstruction model architectures are described that use the smashed data, Z , toreconstruct the input raw data,X ^D priv . These are trained on the publicly available auxiliarydatasets, D aux . Since the dimension of the smashed data, dim ( Z ) = 3 and the dimension of theclient’s input data, dim ( X ) = 28 x 28 for the image datasets (MNIST, FMNIST, and SpeechCommands spectrograms), the only adjustments are to the input and output layer dimensions of the reconstruction models listed. During training, the input to these models is the smashed data,Z aux , and the output is the reconstructed image, X aux . The attack models investigated aredescribed as follows: a. Reconstruction Model 1: This model performs feature inversion to reconstruct images from low dimensional intermediate features. The model is a neural network with transpose convolutions. It consists of skip connections such as those used in ResNet models, with ReLU activation and batch normalization. The mean absolute error (MAE / L1Loss) loss function and a stochastic gradient descent (SGD) optimizer with a learning rate of 0.001 and momentum of 0.9 was used to train the model. b. Reconstruction Model 2: This is a fully convolutional autoencoder for image feature extraction. The decoder part of the auto-encoder is adopted to reconstruct the input images. Upsampling layers are used to recover the feature maps along with deconvolutional layers with ReLU activation function. Batch normalization is used after each deconvolutional layer except for the last layer. The MSELoss loss function is used to compute the difference between the original and reconstructed images. A learning rate of 0.001 and Adam optimizer is used to train Reconstruction Model 2. c. Reconstruction Model 3: This is a fully connected model consisting of 2 hidden layers with a width of 1000 and the ReLU activation function is used as the decoder part of the autoencoder. The loss function used here is the mean absolute error (MAE / L1Loss) + mean square error (MSELoss). A learning rate of 0.001 with the RMSProp optimizer is used to train Reconstruction Model 3.

[0195] After training, the reconstruction models were deployed to recreate a client’s private input. Fig.15 illustrates the setup of a reconstruction attack, on the private data from a client, ,using their smashed data,, to successfully recreate their private input data,represented by,X priv . The reconstruction attack model, f rec is trained, using the auxiliarydataset, D aux , as described in the previous section.

[0196] The smashed data,X priv , transmitted to the server model consists of latent informationabout the client’s private input data,X priv , which can be reconstructed by the setup shown in Fig.15. In the following section, the baseline results for the image comparison metrics demonstrate the similarity between the original (X priv ) and reconstructed images (X priv ). Hence, to maintainthe privacy of the client’s input data, the smashed data communicated to the server model should be desensitized to reduce the possibility of reconstruction attacks. Thus, a defence mechanism may be utilised to mitigate the risks of reconstruction of private input raw images from smashed data in HQSL. Encoding Gate-based Noise Layer Defence Mechanism

[0197] To strengthen the security of HQSL against the risk of reconstruction attacks, a Laplacian noise layer defence mechanism is proposed at the end of the client-side model. However, this noise defence comes with a privacy-utility trade-off. Here, the rotational properties of quantum encoding gates were studied to design the noise layer to (i) effectively hinder the reconstruction attack, and (ii) maintain a high classification performance by the server during inference. Hence, this addresses the privacy-utility trade-off associated with noise layer defence mechanism.

[0198] Noise sampled from a Laplacian distribution, Laplace (^ , b ) , may be systematicallyadded to the smashed data, Z , to provide controlled randomness, while enabling the calculationof the level of perturbation introduced. This may be used for maintaining privacy while retainingdata utility. The location or mean parameter, ^ , provides a translational property to the data,while the scale parameter, b , indicates the variance or perturbation level caused by the noise.The scale parameter also determines the sharpness of the Laplacian distribution about its mean value as shown in Fig.12a.

[0199] Noise Layer Design based on Rotational Properties of RX-gates

[0200] Adding noise sampled from a Laplacian distribution, with appropriately chosen mean and scale parameters, may benefit the quantum circuit on the server-side in terms of maintaining a high classification performance and impeding reconstruction of raw data due to the jitter added to the data. In the following presented experimental results, it is demonstrated how this is not possible for a server without a quantum layer. In other words, the quantum circuit in a machine learning model enables a noise layer to be added without other modification, which is notpossible with a purely classical machine learning model. The following was considered to develop a better understanding of the selection of Laplacian noise parameters.

[0201] The Bloch sphere gives a geometrical representation of the pure state space of a qubit. Quantum gates, e.g., single-qubit rotation gates cause a qubit’s state vector to move around the surface of the Bloch sphere as shown in Fig.16b. In the proposed quantum circuit for HQSL (Fig.7b), RX-gates are responsible for encoding classical data to quantum states. RX-gates mayhave a period of rotation of 4^ . These gates may rotate qubits around the x-axis of the Blochsphere by an angle, ^ , equal to the numerical value of the classical feature, z . Thus, the RX-gates can be written as: RX (^ ) = RX ( z ) . An additive noise to the classical smashed data may beequivalent to a noise, ^ , that may cause rotation of the qubit by a perturbed angle, ^ = z+ ^about the x-axis of the Bloch Sphere. Hence, RX (^ ) = RX ( z+ ^ ) . Owing to the rotationalproperty of qubits, introducing a Laplacian noise with a mean of 4k^ , for k ^ Z , and a smallscale, b , may result in a quantum state similar to that before the noise is applied. This isformalised below.

[0202] In the hybrid setting, for ^ = 4k ^ , k ^Z ,^ : Laplace (^ , b ) . Equivalently, due tothe periodicity of the RX-gate, ^ : Laplace (0, b ) , and hence

[0203] Next, considering a mean of ^ = (4k + 2) ^ , where k^ , the resulting quantum stateencoded by the RX-gate corresponds to a global phase factor of −1 to the noise-free quantumstate. Hence, for ^ = (4k + 2) ^ , Eq. 4 can be updated as:

[0204] A global phase factor of -1 may not affect the physical state. This means that RX ( z+ ^ ), where ^ Laplace ( ^ , b ) and ^ = (4k + 2) ^ , is equivalent to RX(z) up to a global phase. Thismay not affect the further processing along the circuit or measurement outcomes, as it remains a global factor throughout the entire circuit evolution

[0205] Hence, in general, for the quantum layer on the server-side, given ^ = 2k ^ , wherek^ and a small enough scale parameter b , the quantum state encoded in the quantum circuitfrom the noisy smashed data, Z , is expected to be close to the quantum state corresponding to itsclean version, Z . Hence, this may ensure that the performance of the server model in HQSL ismaintained. In contrast, due to the classical nature of the reconstruction model and the server model in split learning, the additive noise can adversely impact their performances. This is demonstrated later in this disclosure.

[0206] Maintaining a small-scale parameter, b , may limit the noise to a minimal perturbation.This ensures a low variability in the resulting quantum state, enabling a consistent encoding that closely aligns with the clean version of the quantum state. Thus, the primary objective forkeeping the scale parameter, b , small is to preserve the accuracy of the hybrid model.

[0207] Fig.17 illustrates the defence setup deployed at inference time after training the HQSL / split learning models. The defence mechanism involves using the Laplacian noise layer to obscure the clean smashed data,Z inf , thereby making it more challenging for the attack modelsto reconstruct the input data.Z inf represents the noisy version of Z inf . As shown in Fig. 13, Z infis communicated to the HQSL or split learning server to carry out the classification task with outputsY inf as intended or leaked to an adversary that uses a trained reconstruction model toreconstruct the private input data,X inf .Experiments on Encoding gate-based noise defence

[0208] In this section, experiments to test the defence mechanism are described. The experiments are conducted with varying Laplacian noise layer parameters to (i) compare the differences between reconstructed and original images in hybrid and classical settings and (ii) compare the performances of HQSL and split learning at inference time, in the presence of the noise layer. From these experiments, the optimal noise parameters are devised that give HQSL a distinct advantage over split learning.

[0209] Each dataset is split in the Xtrain : X test 4:1 ratio. The client and server models weretrained in both HQSL and split learning settings using the training dataset, X train . The testdataset, X test , is further split in the ratio Xrec : X inf of 4 : 1. X rec represents a subset of theauxiliary dataset, X rec ^ D aux , as described in the threat model section, to train the reconstructionmodel. Here, each dataset is split into three subsets such that they are disjoint from each other.X rec is passed to the trained client model to generate the smashed data, Z rec . This step assumesthat the client model represents the shadow model, f shadow , described in the threat modeldescribed herein. The smashed data, Z rec , is generated from the auxiliary dataset and use thedata pair (Z rec , X rec ) to train the adversary’s reconstruction model, f rec .

[0210] The stochastic gradient descent, SGD, optimizer, is used with a learning rate of 10− 3and momentum of 0.9, to train the reconstruction model for 200 epochs using the MAE / L1Loss loss function. After training the client and server models, a Laplacian noise layer is added, withmean ^^ {0, ^ ,2 ^ ,3 ^ ,4 ^ } and scale b^ {0.01,0.1,0.5,1} , at the end of the client-side model.This range of values for the mean ^ is considered because they represent the location of thequbits at different points in a full rotation around the x-axis (RX-gates encode classical data to quantum states rotated around the x-axis) of the Bloch Sphere.

[0211] X inf represents a portion of the private data of a client at inference time, X inf ^D priv ,that is aimed to be reconstruct using the trained reconstruction model. However, due to the noise layer, the client model now outputs noisy smashed data,Z inf . Z inf represents the input to thereconstruction model during the attack phase to recreateX inf as Xˆinf .

[0212] The difference between the original (X inf ) and reconstructed images ( Xˆinf ) arecomputed from the FMNIST and MNIST datasets using three image comparison metrics: cosine distance, mean square error (MSE), and structural dissimilarity (DSSIM). For the spectrograms from the Speech Commands dataset, the log spectral distance (LSD) instead of DSSIM is employed, which is more suitable for spectrograms. To ensure a fair comparison, a mask is applied to the original and reconstructed images. This masking function takes 2 images of the same dimensions and extracts the non-zero pixel values from both images at positions where at least one image has a non-zero pixel. These processed images are then used for further analysis and comparison using the image comparison metrics. The means of these metrics are taken over a batch of testing data. Further detail about these metrics are presented in Appendix D. Results and Discussions

[0213] Comparing the difference between original and reconstructed images under HQSL and split learning settings at different noise levels

[0214] To evaluate the performance of the reconstruction model in classical and hybridsettings in the presence of Laplacian noise, the impact of noise means, ^ , and scale, b , wasinvestigated in recreating FMNIST, MNIST and spectrogram images. This comparative analysis focused on 4 key metrics: Mean Cosine Distance, Mean MSE, Mean Structural Dissimilarity Index (Mean DSSIM), and Mean Log Spectral Distance (Mean LSD). The results for the FMNIST, MNIST and Speech Commands datasets are depicted in Fig.18a, 18b and 18c,respectively, which demonstrates the effect of varying noise parameters ^ and b on thedatasets.

[0215] The baseline values (represented by the horizontal dashed lines) refer to the noise-free implementation of the reconstruction attack, i.e., the smashed data transmitted to the reconstruction model are free from Laplacian noise. In both the classical and hybrid settings, the baseline results are close to 0, showing the reconstruction attacks are successful in recreating the private inputs.

[0216] The general trends from Fig. 18a and 18b were that as the mean parameter, ^ , isincreased, the metric values were well above the baseline, irrespective of the noise scaleparameter, b . Notably, on the MNIST dataset, the deviations from the baseline in the hybridsetting were more significant than in the classical setting for all 3 metrics. This shows that there was a larger difference between original and reconstructed images in the hybrid setting than in the classical setting. In the classical setting, the minimal deviations from the baseline indicate that a reconstruction attack can still be successful for the MNIST dataset, regardless of the noise parameters. This could reflect a model characteristic where the noise parameters do not perturbthe foundational performance of the reconstruction model. It is also highlighted that at ^ = 4 ^ ,irrespective of the scale parameter value, b , the largest deviations were obtained, particularly inthe hybrid setting, indicating that at this ^ value, the additive noise was successfully hinderingthe reconstruction performance of the adversary’s model. It is also noted that significantdeviations were still obtained at ^ = 2 ^ . These findings underscore that the mean values ^ = 2 ^and 4^ caused the performance of the adversary’s reconstruction model to drop in the hybridsetting.

[0217] The results for the Speech Commands spectrograms shown in Fig.18c indicated similar findings. The baseline results were very close to 0, showing that in the absence of the noise layer, the reconstructed spectrograms were very similar to the original spectrograms. As themean, ^ , was increased up to 4^ , the deviations increased. For this dataset, the mean cosinedistance and mean LSD metrics showed that in the classical setting, the deviations from the baselines were more significant than in the hybrid setting. Still, large deviations above the baseline were obtained in the hybrid setting, signalling the increased difficulty of the adversary’smodel to recreate the original spectrograms. For all scale parameters, b , at ^ = 2 ^ and 4^ ,significant deviations were obtained in both the hybrid and classical settings.

[0218] Overall, these results demonstrate that the introduction of the Laplacian noise layer, with appropriately configured noise parameters, can obfuscate the smashed data thereby hindering the performance of the reconstruction model, and in some cases, to a much greaterextent in the hybrid setting. From these results, it was established that mean values of ^ = 2 ^and 4^ effectively increase the difference between the original and reconstructed images, hencesupporting the analysis provided earlier in the disclosure.

[0219] Key takeaway: By tuning the Laplacian noise layer parameters considering the rotational properties of encoding gates, the reconstruction attack model’s performance can behindered in the hybrid setting. In the classical setup, the reconstruction attack can still be successful.

[0220] Inference time performance analysis of HQSL and split learning at different noise levels

[0221] The effect of varying b on the inference performance of HQSL and split learning wasinvestigated given mean valueand ^ = 4 ^ . The performance was measured bycomparing the accuracy and F1-score in classifying FMNIST, MNIST and Speech Commandsspectrogram images with varying scale parameter, b . These results are presented using the boxplots shown in Fig.19 and Fig.20, representing the five-fold test accuracy and F1-score in the presence and absence of noise.

[0222] Fig. 19 shows the effect of noise with fixed mean ^ = 2 ^ and varying scale b oninference performance in classifying FMNIST, MNIST and spectrograms from Speech Commands datasets. Fig.19a shows the 5-fold test accuracy and F1-score versus scaleparameter, b , with mean fixed at ^ = 2 ^ for the FMNIST dataset. Fig. 19b shows the 5-fold testaccuracy and F1-score versus scale parameter, b , with mean fixed at ^ = 2 ^ for the MNISTdataset. Fig. 19c shows the 5-fold test accuracy and F1-score versus scale parameter, b , withmean fixed at ^ = 2 ^ for the Spectrograms for Speech Commands dataset. HQSL (black) canmaintain a higher accuracy and F1-score and lower performance variance than split learning(white) in the presence of Laplacian noise with ^ = 2 ^ and b = 0.01,0.1. The baselineperformances are represented by the noise-free boxplots.

[0223] Fig. 20 shows the effect of noise with fixed mean ^ = 4 ^ and varying scale b oninference performance in classifying FMNIST, MNIST and spectrograms from Speech Commands datasets. Fig.20a shows the 5-fold test accuracy and F1-score versus scaleparameter, b , with mean fixed at ^ = 4 ^ for the FMNIST dataset. Fig. 20b shows the 5-fold testaccuracy and F1-score versus scale parameter, b , with mean fixed at ^ = 4 ^ for the MNISTdataset. Fig. 20c shows the 5-fold test accuracy and F1-score versus scale parameter, b , withmean fixed at ^ = 4 ^ for the Spectrograms for Speech Commands dataset. HQSL (black) canmaintain a higher accuracy and F1-score and lower performance variance than split learning(white) in the presence of Laplacian noise with ^ = 4 ^ and b = 0.01,0.1. The baselineperformances are represented by the noise-free boxplots.

[0224] Based on the results in Fig.19 and Fig.20, the following deductions can be made. The noise-free box plots represent the baseline performance showing a high accuracy and F1-score with low fluctuations across all folds in both HQSL and split learning. By varying the Laplaciannoise scale and keeping the mean fixed at ^ = 2 ^ (Fig. 19), and ^ = 4 ^ (Fig. 20), it isdemonstrated that a scale small enough, e.g., b = 0.01 , caused the performance of split learningat inference time to show high fluctuations in accuracy and F1-score. On the other hand, HQSL could maintain a high performance with less fluctuations – very similar to the noise-free performance. The fluctuations in performance are quantified in terms of the variance of the accuracies and F1-scores at inference time, across the five folds of the dataset. The heights of thebox plots represented this. At b = 0.1 , HQSL still maintained a superior performance to splitlearning but was slightly worse than when b = 0.01. This showed that when the level ofperturbation increased, the performance of HQSL started to drop. As b increased above 0.1,HQSL under-performed split learning in terms of both accuracy and F1-score. Hence, thisexperiment indicates that the scale value b should be kept close to 0, i.e., b = 0.01. Therefore,b = 0.01 may be a desirable scale parameter value for the noise layer as it allows HQSL tomaintain a high performance, similar to its noise-free baseline.

[0225] Key takeaway: By tuning the Laplacian noise layer parameters considering the rotational properties of encoding gates, HQSL’s classification performance is as high as its baseline performance, showing robustness to the noise layer. However, the performance of split learning is less robust to the noise layer.

[0226] Reconstruction performance when HQSL has an advantage.

[0227] Using the Laplacian Noise layer with parameters ^ = 4 ^ and b = 0.01 , some sampleinput images are reconstructed to observe the effectiveness of the noise layer in hindering the performance of the reconstruction attack model. This is shown in Fig.21.

[0228] Fig.21a shows the original and reconstructed Images in classical and hybrid settingsrespectively, with Laplacian noise layer with parameters ^ = 4 ^ , b = 0.01 for the FMNISTdatabase. Fig.21b shows the original and reconstructed Images in classical and hybrid settingsrespectively, with Laplacian noise layer with parameters ^ = 4 ^ , b = 0.01 for the MNISTdatabase. For both FMNIST and MNIST images, reconstruction is almost impossible in the hybrid setting, whereas some visual similarities are observed between reconstructed and original images in the classical setting.

[0229] Fig.21 illustrates that the reconstructed images show more significant differences from the original images in the hybrid setting than in the classical setting. Hence, it can be inferred that the designed noise layer effectively impedes the reconstruction model in recreating the raw images. This is also depicted in the results shown in Fig.18. Furthermore, in HQSL, the noise layer causes the smashed data to be handled properly by the quantum layer on the server-side model, hence maintaining a similar classification performance to the noise-free performance asshown in Fig. 19 and 20. This is due to the configured mean values of ^ = 2 ^ and 4^ , whichcauses the quantum state encoded by the RX-gates in the quantum layer to be nearly invariant tothe additive Laplacian noise, given the scale parameter is small (b = 0.01) . This means that, forthe quantum layer, the obfuscated classical intermediate features correspond to quantum states that are close to those encoded from clean classical smashed data.

[0230] By leveraging the rotational properties of the encoding gates in the quantum layer of HQSL, the noise layer parameters can be tuned to obtain clear advantages in terms of enhancing security in split learning against reconstruction attacks from data privacy leakage. The advantages of the proposed noise layer in HQSL are that it (i) impedes the reconstruction attack model’s performance in reconstructing raw data, and (ii) enables HQSL to maintain a high performance despite its presence. HQSL addresses the privacy-utility trade-off associated with the use of noise layers to enhance security against data privacy leakage in split learning. Conclusion

[0231] This disclosure addresses the problem of applying split learning concepts to pure Quantum Machine Learning (QML) models for resource-constrained clients lacking quantum computing resources. This disclosure also addresses the issue of reconstruction attacks due to data privacy leakage in split learning. This disclosure proposes the application of split learning in Hybrid QML instead. Specifically, a Hybrid Quantum Split Learning, HQSL, architecture is introduced, where split learning methods are applied to a machine learning model comprising a classical sub-model and a quantum sub-model. This HQSL architecture enables classical clients to train their machine learning models collaboratively with a hybrid quantum server. This disclosure also presents a Laplacian noise defence mechanism based on the rotational properties of encoding gates to obfuscate the intermediate smashed data transferred to the server model, to enhance the security against reconstruction attacks.

[0232] The proposed HQSL consists of a HQNN split into two parts: the client-side consisting of a classical sub-model and the server-side comprising a quantum sub-model model where only the first layer is a quantum layer or node. The quantum node consists of only a 2-qubit quantum circuit equipped with a qubit-efficient data loading technique to demonstrate the feasibility of HQSL within the NISQ-era. Despite the small quantum circuit in its quantum layer, HQSL shows promises for enhancing model performance in classification. Moreover, the noise defence mechanism strengthens its security against data privacy leakage and risks of reconstruction attacks.

[0233] The experiments described herein showed that HQSL is feasible, and can also improve classification accuracy and F1-score compared to its classical counterpart. By expanding the experiments to include multiple clients, it is also shown that HQSL is scalable, being able tomaintain a high classification performance as the number of clients increases. This enables multiple clients to benefit from the quantum resources provided by the hybrid quantum server. In the presence of the proposed noise defence mechanism disclosed herein, HQSL addresses the privacy-utility trade-off associated with the use of noise layer to enhance security against reconstruction attacks in split learning.

[0234] As fault tolerance in quantum computers becomes feasible, HQSL could be scaled with larger quantum circuits and potentially involve more than one quantum layer in the server-side model. This would enable resource-constrained clients to further benefit from the potential advantages of quantum computing. Appendix A: Quantum Computing Background Basics and current state of Quantum Computing

[0235] Qubits. These are the basic unit of information of quantum computing. A qubit can berepresented mathematically using Dirac notation as |^^ = ^ 0 | 0 ^ + ^ 1 |1 ^ , where |^ ^ is thequantum state, ^ 0, ^ 1 ^C represent the complex probability amplitudes, and | 0^ and |1^ denotethe computational basis states representing a quantum state. This superposition allows for thesimultaneous representation of both | 0^ and |1^ states until measured, following thenormalization condition^ 20 + ^ 21 =1.Superposition and entanglement.

[0236] A unique property in quantum mechanics is superposition. Superposition allows a quantum state to exist simultaneously in multiple basis states. In quantum computing, this ofteninvolves a qubit being in a combination of | 0^ and |1^ states, but superposition can includemany more states in more complex systems. This property allows calculations to be performed on multiple states at the same time. Entanglement is another property of qubits that occurs when two or more particles become correlated in such a way that their states are no longer independent and cannot be described individually. Instead, their states are intrinsically connected, regardless of the spatial separation between them. Noisy Intermediate-Scale Quantum (NISQ) Computers.

[0237] Currently, quantum computing hardware exists in the NISQ era, where quantum hardware has evolved from a few qubits to modular technology boasting 3 x 133 qubits. The disclosed methods and systems can be implemented on such quantum hardware and on future hardware with more qubits. Gate-Based Quantum Circuits

[0238] In the quantum circuit model, wires represent qubits, and gates represent quantum operations acting on qubits. An example is shown in Fig.7b. Quantum gates are elementary operations that transform the state of one or more qubits. These gates are represented mathematically by unitary matrices that operate on a quantum state vector, changing itsamplitude and phase. The RX-gate, RX ( ^ ) , is an example of a single qubit gate, that performsrotations around the x-axis by an angle of ^ within the Bloch Sphere. The Bloch Sphere gives ageometrical representation of the pure state space of a qubit. Variational Quantum Circuit (VQC)

[0239] An example of the use of gate-based quantum circuits is in VQCs. They are the functional building blocks of near-term QML algorithms. The quantum gates within a VQC are represented using tuneable parameters, which is why VQCs are also called parameterized quantum circuits (PQCs). As a quantum model or as part of a hybrid quantum-classical machine learning model, the parameters of the VQC are determined by an optimization process. This is described earlier in this disclosure.

[0240] Constructing a VQC entails the following 3 steps:a. Encoding Classical Features, X , to Quantum States, ^ (X )A classical data point can be represented as a vector of M features: X = ( x 1 , x 2 , , xM ) .The M -dimensional classical vector is encoded into a quantum state, ^ (X ) , using anencoding process that can be described generically as follows: ^(X ) = U ^Qe ( X ) | 0^ , (6)where X is the classical data vector,the initial quantum state of Q qubits,typically all initialized to | 0^ , and Ue ( X ) is a unitary transformation that encodes theclassical vector X into a quantum state. The unitary transformation Ue ( X ) can beconstructed using various encoding strategies such as basis, amplitude, or phase encoding. b. The quantum circuit Represented by the unitary, Up ( ^ ) , the quantum circuit is parameterized by a set of k freeparameters ^ = {^ 1 , ^ 2 ,..., ^k } . The quantum state at the end of the parameterized circuit,^ (X , ^ ) , can be represented as the product of the unitary Up ( ^ ) and input quantum stateUe ( X ) :c. Quantum MeasurementsA certain observable, Â , is measured after each evolution of a quantum circuit and find theexpectation value, E, of the measurement as: E=^A ˆ ^ = ^^ | A ˆ | ^^ (8)An example of an observable is the Pauli-Y observable, Â = ^ y , where a measurement ateach qubit would generate ^ 1 , which are the eigenvalues for the ^ y operator, for the qubitin the states respectively. The experimentalimplementation of measurements of a particular quantum circuit consists of evolving the quantum algorithm M times and computing the average for each qubit,Q } , as anapproximation of Eq. 8, decomposed for each q as,(9)whereM− 1,qandM+ 1,qare the number of times -1 and +1 measurements are obtained for qubit q respectively. This average serves as an approximation of the probability that a givenqubit is in a particular quantum state. For a Q-qubit system, Eq.9 is computed for each q^ {1,..., Q } , and obtain a feature vector, E , consisting of the expectation measurementoutput at each qubit q as,E= ( e 1 , e 2 ,..., eQ ) (10)Appendix B: Datasets and data processing details

[0241] 5 publicly available datasets were used to test and validate the proposed HQSL architectures: • Botnet DGA Dataset: Botnets consist of compromised devices controlled by a central entity known as the botmaster, and they engage in malicious activities such as launching Distributed Denial-of-Service (DDoS) attacks or spreading malware. DDoS is an example of unauthorized impairment of electronic communication. The domain generation algorithm (DGA) is a technique employed by botnets to evade detection and maintain communication with their command-and-control (C&C) infrastructure. For the experiments described herein, the benchmarked dataset was used, which consisted of 7 features and 1 binary label (malicious or non-malicious). • Breast Cancer Wisconsin (Diagnostic) Dataset: This dataset consists of 569 instances of breast cancer samples with 30 attributes. Each sample in this dataset was classified as benign or malignant. Principal Component Analysis (PCA) was then carried out to reduce thenumber of attributes or features from 30 to only 7 to prevent a high number of trainable parameters. • MNIST Dataset: This dataset consists of 70,00028 x 28 grayscale images of handwritten digits (0, ..., 9) and their respective labels. For the experiments described herein, 6000 normalized random samples were utilised using a set random seed. • Fashion-MNIST (FMNIST) Dataset: This is another image dataset consisting of 70,000 examples of clothing items and 10 labels. Each sample is a 28 x 28 grayscale image associated with a label from the 10 classes.6,000 random samples were extracted using the same set random seed as in the MNIST case, and normalized the images. • Speech Commands Dataset: This is an audio dataset consisting of 105,829 utterances of 35 commonly spoken words from 2168 speakers. Each sample is stored as a .wav file with a length of a maximum of one second. The utterances corresponding to the words “yes" and “no" were extracted and their respective spectrograms were generated and resized into single channel 28 x 28 images. Hence, the classification of the Speech Commands dataset is a binary image classification task, as the images are of the same shape as MNIST and FMNIST datasets. Appendix C: Using a Heuristic method to select Quantum Circuit

[0242] In this section, the heuristic approach employed to select the quantum circuit utilised in the experimental section is described This approach is adopted to find the quantum circuit by making informed decisions and reaching a solution that satisfies some metric. The chosen metric is as follows: Which quantum circuit in from the proposed set of quantum circuits can give rise to an HQSL model that outperforms split learning in terms of test accuracy and F1-score for all datasets?

[0243] For the classical counterpart of each of these circuits, the same rule introduced earlier is followed, i.e., a hidden layer is used where the number of input nodes is equal to the size of the input data dimension, the number of output nodes is the same as the number of qubits (or measurement outputs) in the circuit and the ReLU activation at the end of the dense layer.

[0244] Two considerations for potential quantum circuits are the circuit depth and width (number of qubits). It is desirable to have the smallest circuit that satisfies the above-mentioned metric. The compact size of the circuit is advantageous to ensure a reasonable runtime for the simulations using the default-qubit simulator by PennyLane. Moreover, this makes it more likely to be implemented on real devices in the near-term. In short, the larger the number of qubits and the deeper the circuit, the longer the runtime of HQSL and the harder it is to implement on actual quantum computers.

[0245] The quantum circuits trialled in these experiments are given in Table 3, which is presented as Fig.22. Firstly, 2-dimensional inputs to a quantum circuit with 2 qubits and entanglement were investigated, where each feature was uploaded 3 times per qubit. This is Circuit 1. In Circuit 2, the number of qubits was scaled to 4, to accommodate 4-dimensional inputs and updated the entanglement strategy accordingly. In Circuit 3, the 4 qubits layout was kept but a variational quantum circuit (VQC) configuration was adopted with phase-encoding to convert 4-dimensional classical data to quantum states, which were operated on by 3 parameterized rotation gates each and entangling gates. Circuit 4 was constructed similarly to Circuit 1 but lacking entanglement, to assess the performance without entangling gates. In Circuit 5, the depth of the circuit was increase.

[0246] For all the circuits described above (1-5) when used in HQSL, none of these circuits were able to obtain an accuracy and F1-score that outperformed their equivalent split learning model on all datasets. Hence, the above-mentioned metric was not satisfied for these circuits. Circuit 6 was then developed according to the description presented earlier in the disclosure. The experiments showed that Circuit 6 satisfies the metric and hence, concludes this heuristic approach.

[0247] The HQSL was trained and test on all datasets described previously, and the testing accuracy and F1-score was compared against their classical counterparts. The results of these tests are present in Table 4, which is presented as Fig.23. Only HQSL with Circuit 6 in its quantum layer outperforms split learning on all datasets considered. The blacked-out areas demonstrate that these experiments were not carried out as the circuits they corresponded to were already rejected as they failed the set metric. The dashed boxes indicate the best result between the classical and hybrid approach for each circuit.

[0248] Circuit 6 is the one that makes efficient use of the 2 qubits in the circuit, whereby each of the 3 classical input features were loaded onto each qubit at 3 different RX-gates. Hence, this quantum circuit was chosen due to its superior performance compared to its classical analogue and, it is also easy to simulate without significant computational complexity. Appendix D: Metrics to compare original and reconstructed images

[0249] The metrics used to compare the difference between the client’s original input and reconstructed images are now described.

[0250] Metric 1 Cosine Distance, D c : D c is a metric used to complement cosine similarity.The latter is a measure of similarity between two non-zero vectors and is represented as the cosine of the angle between the vectors. In this disclosure, the cosine distance is used to indicatethe difference between two vectors instead of their similarity. Hence, the larger the cosinedistance is, the larger the difference between two images, A and B .CosineSimilarity,CosineDistance,Dc ( A , B ) =1− S c ( A , B ),0 ^ D c ^ 2 (11)

[0251] Metric 2 Mean Square Error, MSE: In this disclosure, MSE is used to compare the difference between the reconstructed images and the original images. The greater the MSE is, thegreater the difference between the 2 images. For two images A and B of size M×N, the MSE iscalculated as:where A ( i , j ) and B ( i , j ) are the pixel values of the original and reconstructed imagesrespectively at position (i , j ) , and M and N are the dimensions of the images.

[0252] Metric 3 Structural Dissimilarity Index, DSSIM: DSSIM can be considered as the complement of the structural similarity index (SSIM). The SSIM index is a real number between -1 and 1, where 1 indicates perfect similarity, 0 indicates no similarity and -1 indicates perfect anti-correlation. In this disclosure, values less than 0 are treated as showing no similarity withthe original image. Hence, SSIM is redefined as SSIM = max (SSIM,0) such that now,0 < SSIM <1. The DSSIM used in this disclosure is given as:

[0253] Given this redefinition of SSIM, similar to the previous two metrics, the larger the DSSIM value is, i.e., the closer it is to 0.5, the larger the difference between the reconstructed and original images.

[0254] Metric 4 Log Spectral Distance, LSD: LSD can be used to compare the original and reconstructed spectrograms. This metric is applied to the original and reconstructed spectrogramsfrom the speech commands dataset. The LSD between 2 images A and B of size M×N can becomputed asAppendix E: Additional experiments with Reconstruction attack models 2 and 3

[0255] Once the Laplacian noise parameters (^ = 2 ^ ,4 ^ ,b = 0.01 ) are chosen, it can bechecked whether HQSL shows robustness to reconstruction attacks using other reconstruction models, which were discussed earlier in this disclosure. To demonstrate the reconstruction fidelity in the hybrid and classical settings, the metrics described above are used. The results of these additional experiments are presented in Table 5 and Table 6 below. Table 5 presents the results for the FMNIST and MNIST reconstructions and Table 6 presents the spectrogram reconstruction.Table 5: Comparison of Mean Cosine Distance, MSE and mean DSSIM (MDSSIM) values between raw FMNIST and MNIST images and reconstructed images for reconstruction models 2 and 3, with noise parameters ^ = 4 ^ ,b = 0.01. The larger the metric, the worse thereconstruction fidelity. The values in the bracket represent the baseline results, corresponding to noise-free reconstruction performance.Table 6: Comparison of Mean Cosine Distance, MSE and Mean Log Spectral Distance (Mean LSD) values between original and reconstructed spectrogram images from Speech Commandsdataset for reconstruction models 2 and 3, with noise parameters ^ = 4 ^ ,b = 0.01. The larger themetric, the worse the reconstruction fidelity. The values in the bracket represent the baseline results, corresponding to noise-free performance.

[0256] In general, from these results, the reconstruction fidelity is poorer in the hybrid settings,showing reconstruction becomes harder with the introduction of Laplacian noise with ^ = 4 ^and b = 0.01. When coupled with the inference performance comparison from Figs. 20a-c, thisshows an advantage of HQSL compared to split learning. More precisely, reconstruction attacks become more difficult in the hybrid setting than in the classical setting, and HQSL provides a more robust performance than split learning in terms of accuracy and F1-score of the model in the presence of noise.

[0257] It will be appreciated by persons skilled in the art that numerous variations and / or modifications may be made to the above-described embodiments, without departing from the broad general scope of the present disclosure. The present embodiments are, therefore, to be considered in all respects as illustrative and not restrictive.

Claims

1. CLAIMS:

1. A method for training a machine learning model, the method comprising: receiving, from a classical device, an intermediate classical output from a classical sub- model of the machine learning model, the intermediate classical output representing an evaluation of the classical sub-model on training data; configuring quantum gates of a quantum circuit based on the intermediate classical output; executing the quantum circuit using the quantum gates to determine a quantum circuit output, the quantum circuit being configured to represent a quantum sub-model of the machine learning model; adapting the quantum circuit output to determine a further classical output; updating the quantum sub-model based on minimising a loss involving the further classical output; and transmitting a loss propagation value based on the updated quantum sub-model to the classical device, to cause the classical device to update the classical sub-model based on the loss propagation value, to thereby train the machine learning model.

2. The method of claim 1, wherein the machine learning model comprises a further classical sub-model and adapting the quantum circuit output comprises applying the further classical sub-model to the quantum circuit output to determine the further classical output.

3. The method of claim 2, wherein training the machine learning model comprises: updating the further classical sub-model based on minimising the loss involving the further classical output; and determining a further loss propagation value based on the updated further classical sub-model; and updating the quantum sub-model comprises updating the quantum sub-model based on the further loss propagation value.

4. The method of any one of the preceding claims, wherein configuring the quantum gates of the quantum circuit comprises embedding the intermediate classical output into one or more qubits representing the quantum circuit to reduce a dimensionality of the intermediate classical output.

5. The method of claim 4, wherein the quantum gates of the quantum circuit comprise rotation gates; and configuring the quantum gates of the quantum circuit comprises configuring the rotation gates based on the intermediate classical output, wherein the rotation gates are parameterised by the intermediate classical output.

6. The method of claim 5, wherein the quantum gates of the quantum circuit comprise entangling gates, and the rotation gates and the entangling gates of the quantum circuit are alternating in the quantum circuit.

7. The method of any one of the preceding claims, wherein the classical device is one of multiple classical devices, each of the multiple classical devices comprising the classical sub-model; and the method further comprises, for each of the multiple classical devices: repeating the steps of receiving the intermediate classical output from the classical device, configuring the quantum gates, executing the quantum circuit, adapting the quantum circuit output, and updating the quantum sub-model; and transmitting the loss propagation value based on the updated quantum sub-model to the each of the multiple classical devices to cause each of the multiple classical devices to update the classical sub-model on the respective classical device based on the received loss propagation values.

8. The method of claim 7, wherein each of the multiple classical devices stores respective training data and performing the method of claim 7 trains the machine learning model on the training data of each of the multiple classical devices while isolating the respective training data from other ones of the multiple classical devices.

9. The method of claim 8, wherein the respective training data of each of the multiple classical devices is confidential training data.

10. The method of any one of the preceding claims, wherein the classical sub-model is a neural network, and the quantum sub-model is a quantum neural network.

11. The method of any one of the preceding claims, wherein the machine learning model is a classifier.

12. The method of any one of the preceding claims, wherein the method further comprises receiving, from the classical device, a label being associated with the training data; and updating the quantum sub-model is based on minimising the loss involving the further classical output and the label.

13. The method of claim 12, wherein the quantum circuit is implemented on a quantum device comprising at least two qubits.

14. The method of any one of claims 2 to 13, wherein the quantum sub-model and the further classical sub-model are implemented on a hybrid quantum server comprising a quantum processor and a classical processor, the hybrid quantum server being configured to perform the method of any one of claims 2 to 13.

15. A method for classifying test data using a trained machine learning model as performed by a classical device, the trained machine learning model comprising a classical sub-model and a quantum sub-model, the method comprising: evaluating the classical sub-model on the test data to determine a classical output; applying a noise layer to the classical output to determine a noisy output, the noise layer being configured to add noise to an output of the classical sub-model; and sending the noisy output to a server to cause the server to: configure quantum gates of a quantum circuit based on the noisy output; execute the quantum circuit using the quantum gates to determine a quantum circuit output, the quantum circuit being configured to represent the quantum sub-model; adapt the quantum circuit output to determine a classification output representing a classification of the test data; and send the classification output to the classical device to classify the test data.

16. The method of claim 15, wherein the noise layer is configured to add Laplacian noise to the output of the classical sub-model.

17. The method of claim 16, wherein the Laplacian noise is based on a Laplace distributionwith a mean of about 4^ and a scale parameter between about 0.01 and 1.

18. Software that, when installed on a classical processor and executed by the classical processor, causes the classical processor to perform the method of any one of claims 1 to 14 or part thereof.

19. Software that, when installed on a classical processor and executed by the classical processor, causes the classical processor to perform the method of any one of claims 15 to 17 or part thereof.

20. A hybrid quantum server comprising: a classical processor configured to: receive, from a classical device, an intermediate classical output from a classical sub-model of a machine learning model, the intermediate classical output representing an evaluation of the classical sub-model on training data; and configure quantum gates of a quantum circuit based on the intermediate classical output; a quantum processor configured to execute the quantum circuit using the quantum gates to determine a quantum circuit output, the quantum circuit being configured to represent a quantum sub-model of the machine learning model; and wherein the classical processor is further configured to: adapt the quantum circuit output to determine a further classical output; update the quantum sub-model based on minimising a loss involving the further classical output; and transmit a loss propagation value based on the updated quantum sub-model to the classical device, to cause the classical device to update the classical sub-model based on the loss propagation value, to thereby train the machine learning model.

Citation Information

Cited By

  • Quantum-inspired method and system for super-resolution reconstruction of plant microtubule images

    CN122335552A