Quantum Machine Learning Devices and Methods

A quantum ML device using quantum dots and gates enhances feature generation and machine learning performance by converting input data into non-linear mappings, addressing the limitations of classical computers and current quantum machines.

JP2025523534APending Publication Date: 2025-07-23SILICON QUANTUM COMPUTING PTY LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024575835
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-07-05
Filing Date
2023-07-05
Publication Date
2025-07-23

AI Technical Summary

Technical Problem

Classical computers using binary bits are limited by their binary nature, which slows down calculations and requires multiple bits for simple equations, while quantum computers using qubits can perform calculations faster and with fewer qubits, but current quantum machines are in the noisy intermediate-scale quantum era and lack fault-tolerant algorithms, limiting their performance in machine learning tasks.

Method used

A quantum ML device comprising quantum dots, source gates, drain gates, and control gates is used to convert input data into voltages, measure signals, and analyze parameters as non-linear mappings for improved feature generation in machine learning models, employing techniques like quantum extreme learning machines, kernel learning machines, and random kitchen sinks.

Benefits of technology

The quantum ML device significantly outperforms classical methods in feature generation and machine learning tasks, particularly in high-dimensional datasets, demonstrating improved computational efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025523534000001_ABST
    Figure 2025523534000001_ABST
Patent Text Reader

Abstract

Disclosed are a method and a device for generating quantum features for a machine learning model. The method includes providing a quantum ML device (QMLD) comprising one or more quantum dots, one or more source gates, one or more drain gates, and one or more control gates. The method further includes converting input data of a machine learning model into a first voltage, applying the first voltage to one or more control gates and / or source gates and / or drain gates, applying a second voltage to one or more of the one or more source gates, measuring a signal at one or more of the one or more drain gates, analyzing the measured signal to determine values of one or more parameters, and interpreting the values of the one or more parameters as a non-linear mapping of the input data used by the machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Aspects of the present disclosure relate to quantum processing devices, and more particularly, to methods and devices for implementing machine learning techniques using such quantum processing devices.

Background Art

[0002] The developments described in this section are known to the inventors. However, unless otherwise specified, none of the developments described in this section should be considered to be admitted as prior art solely for the reason that they are included in this section, nor should they be considered to be known to those skilled in the art.

[0003] Machine learning (ML) has had a great impact on our daily lives, from the progress of computational methods in materials design and chemical processes to pattern recognition for autonomous vehicle transportation and cell classification for cancer cell detection. Along with the progress of computing power that doubles approximately every year according to Moore's law, ML algorithms have also advanced rapidly.

[0004] To date, in ML, classical computers that perform calculations using binary bits that can take on one of two different states, 0 or 1, have been used. The binary nature of classical computing bits can slow down the speed, and usually multiple bits are required to complete the simplest equations on classical computers. On the other hand, quantum computers perform calculations using qubits, or quantum bits, which can exist in multiple states, unlike classical bits. A qubit can take on the state of 0, 1, or a superposition of the two states (referred to as a quantum state). Therefore, quantum computers can complete algorithms much faster and may require fewer qubits to perform operations. Due to this advantage, quantum computers are touted as being able to solve ML problems that may be difficult to solve with classical computing.

[0005] With this in mind, recently, efforts have been underway to solve ML problems using quantum computers, giving rise to the subfield of quantum machine learning, and various quantum algorithms (which are partially or fully executed on quantum computers) for ML tasks have been shown to outperform classical algorithms (which are executed on classical computers). However, the performance of such quantum machine learning algorithms can be further improved. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM

[0006] According to a first aspect of the present disclosure, a method for generating quantum features for a machine learning model is provided. The method includes providing a quantum ML device (QMLD) comprising one or more quantum dots, one or more source gates, one or more drain gates, and one or more control gates. The method further includes converting input data of the machine learning model into a first voltage, applying the first voltage to one or more control gates and / or source gates and / or drain gates, applying a second voltage to one or more of the one or more source gates, measuring a signal at one or more of the one or more drain gates, analyzing the measured signal to determine values of one or more parameters, and interpreting the values of the one or more parameters as a non-linear mapping of the input data used in the machine learning model.

[0007] Converting the input data to a voltage may include performing a random transformation of the input data and converting the randomly transformed input data to a voltage. In other embodiments, converting the input data to a voltage includes directly mapping the input data to a voltage. In such embodiments, interpreting the value of one or more parameters may include combining the values of one or more parameters as features for a machine learning model.

[0008] In some embodiments, converting the input data to a voltage includes pairing data points of the input data and converting the paired data points to a combined voltage, and applying a voltage to one or more control gates includes applying a combined voltage to one or more control gates. In such cases, interpreting the value of one or more parameters includes determining a distance metric or similarity score between the values of one or more parameters.

[0009] In some examples, a quantum ML device includes a plurality of source gates, a plurality of drain gates, and a plurality of control gates, and the quantum ML device is used as a quantum random kitchen sink device.

[0010] In other examples, a quantum ML device includes one source gate, a number of drain gates that matches the desired feature dimension, and a number of control gates that matches the dimension of the input data, and the quantum ML device is used as a quantum extreme learning machine.

[0011] In yet other examples, a quantum ML device includes one source gate, one drain gate, and a number of control gates that is twice the dimension of the input data, and the quantum ML device is used as a quantum kernel learning machine.

[0012] This method may further include steps of manufacturing a quantum ML device. The manufacturing steps include fabricating a bulk layer of a semiconductor substrate, fabricating a second semiconductor layer, exposing a clean crystal surface of the second semiconductor layer to dopant molecules to generate an array of dopant dots on the exposed surface, annealing the arrayed surface to incorporate dopant atoms into the second semiconductor layer, and forming one or more gates, one or more source leads, and one or more drain leads.

[0013] One or more control gates may be formed in the same plane as the dopant dots. In other examples, a dielectric material may be deposited on the annealed second semiconductor layer, and one or more control gates may be formed on the dielectric material.

[0014] In some examples, the dopant dots are phosphorus dots, the second semiconductor layer is silicon-28, and the quantum ML device includes 10 quantum dots.

[0015] In other aspects of the present disclosure, a quantum ML device (QMLD) is provided. The QMLD includes one or more quantum dots, one or more source gates, one or more drain gates, and one or more control gates. The quantum ML device is used to generate quantum features for a machine learning model by applying a first voltage corresponding to input data of the machine learning model to one or more control gates and / or source gates and / or drain gates, applying a second voltage to one or more of the one or more source gates, measuring a signal at one or more of the one or more drain gates, analyzing the measured signal to determine values of one or more parameters, and interpreting the values of the one or more parameters as a non-linear mapping of the input data used in the machine learning model.

[0016] Further aspects of the present disclosure and embodiments of the aspects summarized in the immediately preceding paragraphs will become apparent from the following detailed description and the accompanying drawings.

Brief Description of the Drawings

[0017]

Figure 1A

Figure 1B

Figure 2AB

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12A

Figure 12B

Figure 12C

Figure 12D

Figure 13A

Figure 14A

Figure 14B

Figure 14C

Figure 14D

Figure 14E

Figure 14F

Figure 14G

Figure 15A

Figure 15B

Figure 16A

Figure 16B

Figure 16C

Figure 16D

DETAILED DESCRIPTION OF THE INVENTION

[0018] Summary As described above, ML algorithms and models are currently used in almost all technical fields to assist in data classification, result prediction, or prescription of solutions. For example, ML algorithms can be used to automatically classify emails as spam or not, predict weather patterns, or formulate action plans based on given input conditions. To achieve these goals, first an appropriate ML model (e.g., a binary classification model or a regression model) is selected and then trained with training data. For example, to classify emails, a binary classification model can be used, and the training data can be emails. To predict weather patterns, a regression model can be selected, and the training data can be various types of weather phenomena and past weather data, etc.

[0019] At the macro level, an ML model takes an input data vector

[0020]

number

[0021] Depending on how well the model is trained,

[0022]

number

[0023] For example, an ML model may be trained to take an image as input and determine whether the image contains a cat. In such an example, an ML model is first trained using a set of images; some images contain cats and some do not. Further, the training data may be labeled such that the model knows which images contain cats and which do not. Once the ML model is trained with enough data, it can classify unlabeled images as either images of cats or not. The accuracy of such an ML model depends on several factors, including but not limited to the amount of training data used, the model itself, and how well the model was trained (i.e., the quality and amount of training).

[0024] One way to improve the accuracy of an ML model is to use feature engineering. In machine learning, a feature is an individual measurable property or characteristic of a phenomenon. For example, in a spam detection algorithm, features can include the presence or absence of certain email headers, the structure of the email, the language, the frequency of certain terms, and the grammatical correctness of the text. The model can use these features to assist with classification, prediction, or prescription. Feature engineering refers to the process of selecting, manipulating, and transforming raw data to extract features that can be used for training an ML model. In feature engineering, data from the training dataset is typically utilized to create new features. This new set of features can then be used to train an ML model with the aim of simplifying and accelerating the overall computation.

[0025] One exemplary technique for engineering features is called the kernel method or kernel trick. This method generates features for an algorithm based only on the inner product between pairs of input data points.

[0026] [Number]

[0027] For any positive semi - definite function K(x i , x j ), the inner product in the transformed space is defined, that is, K(x i , x j ) = <φ(x i ), φ(x j )> based on the observation that here, φ(x i ) and φ(x j ) are the transformations of the data points x i and x j respectively. Another method for engineering features is the random kitchen sink method, which, as its name suggests, randomly selects a subset of features from a feature set and uses these features to train the corresponding ML model.

[0028] As described above, classical computers quickly reach a point where quantum mechanical effects prevent further development and a quantum computer is needed to perform calculations that cannot be executed on classical computers. However, a caveat of quantum computing is that currently constructed physical devices are still in the so-called noisy intermediate-scale quantum (NISQ) era, and sufficiently fault-tolerant quantum algorithms cannot be implemented. Instead, these NISQ systems can be used to solve certain problems that are important in practice. Such quantum systems that are constructed (or hard-coded) for a specific purpose to execute one or more specific problems are called analog quantum computers or analog quantum processors.

[0029] Such analog quantum computers or processors have recently been constructed to simulate the Fermi-Hubbard model, magnetism, and topological phases.

[0030] The present disclosure introduces a quantum ML device for performing quantum machine learning using classical inputs and outputs, specifically utilizing semiconductor quantum dots to simulate the Hamiltonian of the Fermi-Hubbard model. With minimal modifications to the quantum ML device, this device can be used as a quantum extreme learning machine, a quantum kernel learning machine, and a quantum random kitchen sink. It has been found that the quantum ML device and related quantum ML techniques for engineering the features of the present disclosure exhibit significantly better performance than corresponding classical computing techniques.

[0031] Exemplary System Figure 1A shows an exemplary semiconductor quantum dot device 100 that can be implemented in the quantum ML device of the present disclosure. As shown in this figure, the quantum dot device 100 includes a semiconductor substrate 102 and a dielectric 104. In this example, the substrate is isotopically purified silicon (silicon 28), and the dielectric is silicon dioxide. In other examples, the substrate can be silicon (Si). An interface 106 is formed at the portion where the substrate 102 and the dielectric 104 are in contact. In this example, this is the Si / SiO2 interface. To form the quantum dot, donor atoms 108 are disposed within the substrate 102. The quantum dot is defined by the Coulomb potential of the donor atoms. The donor atoms 108 can be introduced into the substrate using nanofabrication techniques such as hydrogen lithography provided by a scanning tunneling microscope or industry-standard ion implantation techniques. In some examples, the donor atoms 108 can be phosphorus atoms within the silicon substrate, and the quantum dot can be called a Si:P quantum dot.

[0032] In the example shown in Figure 1A, the quantum dot includes a single donor atom 108 embedded in a silicon 28 crystal. In other examples, the quantum dot can include multiple donor atoms embedded in proximity to each other.

[0033] Gates 112 and 114 can be used to adjust the filling of electrons into the quantum dot 100. For example, electrons 110 can be loaded into the quantum dot by a gate electrode, such as 112. The physical state of the electrons 110 is described by a wave function 116, which is defined as the probability amplitude of finding the electrons at a specific location. The donor qubit in silicon relies on binding the electron spin using a potential well naturally formed by the donor atomic nucleus.

[0034] Figure 1B shows another exemplary semiconductor quantum dot device 150 that can be implemented in the quantum ML device of the present disclosure. This device is similar to the quantum dot device 100 shown in Figure 1A, but the difference lies in the gate arrangement. In Figure 1A, the gates 112, 114 were arranged on the dielectric 104. In this example, the gate 152 is located within the semiconductor substrate 102. In some embodiments, the gate 152 is arranged in the same plane as the donor dot 108. Such an in-plane gate can be connected to the surface of the substrate via a metal via (not shown). A voltage can be applied to the gate electrode 152 to confine one or more electrons 110 within the Coulomb potential of the donor atom 108. In some examples, the quantum dot device can include a gate arranged within the semiconductor substrate 102 and a gate arranged on the dielectric 104.

[0035] Figure 2A shows an exemplary quantum ML device (QMLD) 200 according to an aspect of the present disclosure. The QMLD 200 includes a donor quantum dot array 202, a source 204, a drain 206, and a plurality of input gate electrodes G1 - G8. In one example, the quantum dot array 202 is composed of an array of donor quantum dots 208, and each quantum dot 208 is similar to those shown in Figure 1A and / or Figure 1B. The inset 210 in Figure 2A shows an enlarged view of a 5×5 section of the quantum dot array 202. As seen in the inset, the quantum dots 208 are arranged in a square lattice. Each circle in the inset represents a quantum dot, and the arrow within the quantum dot represents the spin of the electron bound to the donor atom of the quantum dot.

[0036] It will be understood that the quantum dots 208 within the array need not be arranged in a square lattice. Instead, the quantum dots 208 can be arranged in an array of any shape without departing from the scope of this disclosure. In some examples, the quantum dots 208 can be arranged in a random manner that has no symmetry available within the array. This randomness in the array design can improve the ML prediction. Further, the number of quantum dots 208 present within the array 202 can vary depending on the particular implementation. Generally, the larger the array, the better or more accurate the results. However, it should be noted that beyond about 50 quantum dots, it becomes impossible to simulate the device on a classical computer.

[0037] Although a single source 204 and a single drain 206 are shown in this example, this need not always be the case. In other embodiments, the QMLD 200 can include multiple drains and / or source leads. Further, the number of gates used in the QMLD 200 can vary depending on the feature generation method used.

[0038] Furthermore, the quantum dot array 202 can be fabricated in 2D where the input gate electrodes G1 - G8 are in the same plane as the array 202 (see Figure 1B), and / or can be fabricated in 3D such that after covering the quantum dot array layer with epitaxial silicon, the gates G1 - G8 can be patterned on a second layer (see Figure 1A). Since the gate density of the Si:P quantum dots 208 is very low, it becomes possible to fabricate a large-scale quantum dot array with a small number of control electrodes. In embodiments of this disclosure, a low gate density of 1 gate per hundreds of quantum dots is possible. However, more gates may be required to operate the array. Thus, in some embodiments, there may be about 10 gates to control about 75 quantum dots. However, this is not critical and this ratio can be varied to suit various implementations without departing from the scope of this disclosure.

[0039] To measure the electron transport through the quantum dot array 202, the quantum dot array 202 is weakly coupled to a source 204 and a drain 206. Further, the quantum dot array 202 is capacitively coupled to a plurality of control gates G1 - G8. The control gate 208 can be used to adjust the electron filling, the inter-dot coupling, and the single-particle energy levels of the quantum dot array 202.

[0040] The Hamiltonian (i.e., the underlying mathematical description) that describes the behavior of electrons in the quantum dot array 202 is a 2D extended Hubbard model with long-range electron-electron interactions (V), on-site Coulomb interactions (U), and nearest-neighbor electron transport (t). The Hamiltonian describing the system is as follows.

[0041]

Equation

[0042] Here, the electron hopping term (t) is related to the tunneling probability of electrons between nearest-neighbor donor sites. The on-site Coulomb interaction (U) is the energy required to add a second electron to a site. The inter-site Coulomb interaction (V) is the energy required to add an electron to a neighboring site, which can include interactions over all pairs of sites i and j. The parameters U, V, and t are determined by the configuration and the distance between the dots. Finally, ε i is the single-particle energy level of the i-th site.

[0043] The problem of the ground state of the Hubbard model has been shown to be Quantum Merlin Arthur (QMA:Quantum Merlin Arthur) complete (a complexity type in computational complexity theory), which means that the ability to control the ground state of the Hubbard Hamiltonian offers the potential of large-scale computational resources. The techniques described in the following sections utilize the computational power of this Hubbard model for machine learning in a way that is independent of downstream tasks.

[0044] Feature Generation for ML The QMLD200 can be used to generate features for the ML model. For example, classical input data

[0045] [Number]

[0046] can be converted into a voltage applied to one or more of the gate electrodes. In some examples, different voltages can be applied to each gate. In other examples, the same voltage can be applied to two or more of the gate electrodes. What voltage is applied to the gate electrodes can ultimately depend on the particular ML algorithm used. The applied voltage is used to change the Hubbard parameter (ε i ) of the ML. Then, the source voltage can be swept to measure the current curve at one or more drain leads 206 of the QMLD200. Then, this measured current curve can be analyzed to find the charge excitation gap (CEG) or the current at a particular source voltage. Then, the CEG, current, or conductance can be used to output a non-linear function mapping that can be used in a classical ML model.

[0047] FIG. 3 is a flowchart showing an exemplary method 300 for generating features for quantum ML according to an aspect of the present disclosure.

[0048] Method 300 begins at step 302, where input data is applied to one or more of the control gates (e.g., one or more of control gates G1 - G8). In some examples, the raw input data is applied as a voltage to the control gates. In other examples, the input data is converted before being applied as a voltage to the control gates. In some embodiments, the input data is represented as a vector. For example, if the input data is an image of a cat, the input data vector could be the color scale of each pixel in the image. This input data vector is converted to a voltage and used as an input to the ML model.

[0049] Next, at step 304, a constant voltage or a voltage sweep is applied to the source lead 202. The sweep voltage can be in the range of a few millivolts. A direct current (DC) or radio frequency (RF) sweep can be used. In an RF sweep, each drain lead 206 is connected to a resonant circuit (not shown) of a different frequency. Using RF enables parallel operation by multiplexing and measuring several drain leads 206, further shortening the processing time of the machine learning method. Since each drain lead 206 is connected to a resonator of a different frequency, this multiplexing can be achieved by an RF sweep. In this way, voltage sweeps of different frequencies can be applied corresponding to the various drain leads.

[0050] Next, in step 306, a current is measured as a function of the source voltage at the drain lead 206. When the measured current data is plotted against the source voltage, a current curve can be obtained. FIG. 2B shows an example of these current curves measured at an independent drain lead 206. Specifically, FIG. 2B shows three plots 220, 222, 224, where the x-axis of each plot is the source voltage and the y-axis is the current. Plot 220 shows the current measured as a function of the source voltage for a first data point x1 applied as a voltage to the gate 208. Plot 222 shows the current measured as a function of the source voltage for a second data point x2 applied as a voltage to the control gate 208. Finally, plot 224 shows the current measured as a function of the source voltage for a third data point x3 applied as a voltage to the control gate 208.

[0051] In step 308, the current data obtained in step 306 is analyzed to determine one or more parameters. In some examples, the parameter can be a charge excitation gap (CEG), and the measured current curve can be analyzed to find the CEG. The CEG is the energy of the gap corresponding to the amount of energy required to add electrons to the quantum dot array. The CEG is determined by finding the point at which the current rises sharply. At that time, the voltage that increases the current most rapidly is regarded as the CEG. In other examples, the measured current curve can be analyzed to determine the current at a specific source voltage. In other examples, the parameter determined from the measured current curve is conductance.

[0052] Next, in step 310, the values of one or more parameters are interpreted as a non-linear mapping of the input data used in the machine learning model. For example, if the parameter is CEG and a quantum random kitchen sink is used, one or more values of CEG are interpreted as features in an extended space ready to be used in the ML model. Alternatively, if a quantum kernel learning machine is used, interpreting the values of one or more parameters may involve determining a distance metric or similarity score between the values of one or more parameters.

[0053] Importantly, these various features can be generated from a single voltage sweep of the source voltage at a specific set of input voltages to the gate. Alternatively, since current and conductance can be generated by a single current measurement, the processing time of the QMLD200 is significantly reduced. These quantum extended features are expected to perform better than classical features because they increase the computational complexity compared to classical ones.

[0054] Method 300 can be adapted according to the ML techniques used for feature generation. For example, method 300 describes that the voltage corresponding to the input data is applied to the control gate, but this is not necessary in all embodiments. Instead, in some other embodiments, the voltage corresponding to the input data can be applied to one or more source gates and / or one or more drain gates in addition to or instead of the control gate.

[0055] Furthermore, depending on the method of use, the way of applying data as voltage and the way of interpreting the output can vary. In the following sections, some exemplary methods and variations will be described. For example, when using the Quantum Random Kitchen Sink (QRKS) method, the input voltage is a random transformation of the input data, and the output parameter is an additional feature dimension. In other examples, using the Quantum Extreme Learning Machine (QELM), the input voltage is directly mapped to the input data, and the output parameter is interpreted as an additional feature dimension. In still other examples, using the Quantum Kernel Learning Machine (QKLM) method, the input data is applied as a pair of voltages. For example, data point x i and data point x2 are combined and applied as voltage to gate 208. And the output parameter is interpreted as the distance metric or similarity score between them.

[0056] Also, in step 306, a current signal is measured as a function of the source voltage at drain lead 206, but this is not necessary in all embodiments. In some embodiments, other types of signals such as voltage, capacitance, conductance, inductance, etc. can be measured together with or instead of the current signal. Then, when the measured signal data is plotted against the source voltage, a signal curve can be obtained.

[0057] Figures 2C and 2D show examples of these other signal curves measured at independent drain lead 206. Specifically, Figure 2C shows a setup for measuring the RF transmission of QMLD by applying an oscillating voltage (V RF ) to source 204 using a bias tee (not shown). Drain 206 is connected to an LC resonance circuit (L) and parasitic capacitance, which converts the output current signal into a voltage signal. Then, this voltage signal is amplified and the original V RFIt is demodulated by a signal to obtain changes in the amplitude and phase of the voltage signal that has passed through the array. FIG. 2D shows three plots 230, 232, 234, where the x-axis of each plot is the source voltage and the y-axis is another signal (e.g., voltage output). Plot 230 shows the voltage amplitude or phase measured as a function of the source voltage for the first data point x1 applied as a voltage to gate 208. Plot 232 shows the voltage amplitude or phase measured as a function of the source voltage for the second data point x2 applied as a voltage to control gate 208. Finally, plot 234 shows the voltage amplitude or phase measured as a function of the source voltage for the third data point x3 applied as a voltage to control gate 208. Then, the amplitude and phase of the voltage signal can be used to perform machine learning, similar to the current signal.

[0058] Although QMLD200 has been described using donor-based quantum dots, it will be understood that the QMLD system 200 and the ML method can also function with gate-defined quantum dots.

[0059] In the following sections, several different quantum ML techniques that can be performed using the QMLD200 described herein will be described.

[0060] Quantum Extreme Learning Machine (QELM) In the classical machine learning literature, an extreme learning machine is a type of feedforward neural network architecture that implements a non-linear projection by fixing and randomizing the network parameters and structure. Then, a simple model such as linear regression is trained in this projected feature space. Extreme learning machines have been shown to perform universal approximation even under very loose conditions and have been shown to outperform various other techniques including support vector machines.

[0061] Figure 4 shows a schematic diagram 400 of a hardware-based QELM process according to an aspect of the present invention. The QELM process utilizes QMLD200 and includes a quantum dot array 202, a single source lead 204, one or more drain leads 206, and one or more gates G1 to G6. Each drain lead 206 may correspond to a dimensional feature. In this exemplary figure, a 3×3 quantum dot array, three drain leads 206A, 206B, 206C, and six control gates G1 to G6 are shown, capable of generating three data dimensions. If more dimensions are required, the number of drain leads 206 can be increased.

[0062] For each input data point x i , by mapping each input dimension to a control gate and applying voltages to control gates G1 to G6, a non-linear projection is realized. The input data point is converted into a voltage and used to generate additional features. Then, the output current is measured at the drain lead 206 and used as the extended feature space f(x i ). The dimension of the resulting feature map is equal to the number of drain leads 206, whereby the ability of this method can be subject to experimental limitations. The extended feature space obtained from this technique can then be used in downstream ML tasks. For example, the dataset f(x i ) is used in an ML model, and some property y i regarding the input can be classified or predicted. In the QELM process, all training data points and test data points need to be converted. Therefore, the number of measurements is the product of the number of measurements per conversion and the number of data points.

[0063] Quantum Kernel Learning Machine (QKLM) Another way to generate features is to use QMLD200 as a quantum kernel learning machine.

[0064] Figure 5 shows a schematic diagram 500 of the QKLM process according to an aspect of the present disclosure. In the QKLM process, a QMLD 200 having a quantum dot array 202, a single source lead 204, a single drain lead 206, and a plurality of control gates G1 to G6 is used. In the QKLM method, the size of the data is limited by the number of control gates. This is different from the above-described QELM method in which data is mapped to different gate voltages. In the QKLM method for d-dimensional data, 2d control gates can be used so that two data points can be mapped simultaneously.

[0065] The kernel trick is a technique used in ML where the dot product between vectors is replaced by a kernel function. Mercer's theorem states that if the kernel function is symmetric and continuous and its evaluation between all pairs of data points forms a positive semi-definite matrix, it can be expressed as the inner product in the transformed space (feature space), that is, for a certain transformation φ, K(x i ,x j ) = <φ(x i ), φ(x j )>. This means that it is possible to calculate the similarity measurement values in those spaces without calculating the vector representations in a high-dimensional and complex vector space, thereby improving efficiency or enabling the use of transformations that were originally impossible. For example, a common kernel used in a support vector machine (SVM) is a radial basis function kernel

[0066]

Number

[0067] . The feature space representation of this kernel is infinite-dimensional.

[0068] The QKLM process starts with two data points x i and x j being symmetrically applied to the control gate 208. Here, symmetric means f(xi , x j ) = f(x j , x i ), that is, it means that the result does not change even if the order of the inputs is swapped. In one example, this can be achieved by applying data point 1 (x i ) to the upper control gates (G1~G3), applying data point 2 (x j ) to the lower control gates (G4~G6), then applying data point 2 (x j ) to the upper control gates (G1~G3), applying data point 1 (x i ) to the lower control gates (G4~G6), and summing the outputs. As a result, even if data points 1 and 2 (x i and x j ) are swapped, the result does not change. Thus, an appropriate distance metric is maintained. For example, the two questions "How far is it from Melbourne to Sydney?" and "How far is it from Sydney to Melbourne?" produce the same output.

[0069] Then, based on the measured value of the current in the resulting drain lead 206, a kernel function that satisfies the kernel criterion can be defined, and this kernel is used in the ML model. As described above, in this technique, two control gates 208 per data dimension and a single drain lead 206 are required. Using the QKLM process, the kernel is calculated between all pairs of training data points, as well as between all test data points and training data points. Therefore, in the case of the QKLM process, the number of measurements can be expressed as follows. Number of measurements per kernel calculation × (Number of training data points 2 + (Number of training data points × Number of test data points))

[0070] Quantum Random Kitchen Sink (QRKS) The Random Kitchen Sink is a technique in which a feature space is generated by randomly transforming input data points. It will be understood that a function can be approximated under certain conditions by this method. For example, in this method, a cost function such as an L-Lipschitz function whose weight decays rapidly from a given sample distribution when the maximum size of any projection point is 1 or less can be approximated. QRKS has been empirically shown to perform equivalently to other techniques under such conditions.

[0071] Figure 6 shows a schematic diagram 600 of the QRKS process according to an aspect of the present disclosure. The QRKS process can also be executed with the QMLD200 described above. In this example, the QMLD200 includes a quantum dot array 202, a single source lead 204, a single drain lead 206, and a plurality of control gates (six gates, G1 to G6 in this example) for each data dimension of the randomly linearized input data. Some of these control gates can facilitate access to or modification of the Hamiltonian, which in some cases leads to higher-dimensional outputs. However, the QRKS process can be used with any number of gates.

[0072] Similar to QELM, a random non-linear projection into a new feature space is achieved by the Quantum Random Kitchen Sink (QRKS). However, unlike QELM, the dimensionality of the feature space can be of any size. That is, the dimension is not limited by the number of drain leads.

[0073] For each desired feature dimension, a random linear transformation is applied to all data points within the dataset

[0074]

Number

[0075] is applied. Here, w is an n×m matrix having elements sampled from a Gaussian distribution, n is the dimension of the data, and m is the number of control gates 208. And,

[0076] [Number]

[0077] is a vector of the same length as the number of input gates 208, having elements sampled from a uniform distribution. This transformed dataset is directly applied to the input gates 208. Current is measured at the drain lead 206 and used as the extended feature space f(x i ). This is repeated for each desired feature dimension to realize the desired dimensions for use in downstream ML processes.

[0078] As shown in FIG. 6, the data point x i is randomly transformed multiple times. Each random transformation corresponds to a new feature dimension, and each feature dimension has the same linear transformation for all data points. For each new feature of each data point, the transformed data is applied as a gate voltage to the control gates G1 to G6, and the output current is measured by the drain lead 206. Each transformation and measurement adds one dimension to the extended feature space representation of that data point.

[0079] Note that although sampled randomly, the corresponding transformations must be the same across different data points. In this model, since more transformations can be sampled, the size of the extended feature space is arbitrary. In addition, when multiple drain leads are used, multiple dimensions can be added to the extended feature space in each measurement. In this model, all training points and test points must be transformed, that is, the number of measurements is expressed as follows. Number of measurements per transformation × Number of data points × Number of feature dimensions

[0080] As described above, the various ML algorithms (QELM, QKLM, and QRKS) have slightly different QMLD200 requirements. The requirements for these source, drain, and control gates are summarized in Table A below.

[0081]

Table 1

[0082] Here, the parameters in Table A are defined as n s : the number of source gates, n d : the number of drain gates, n c : the number of control gates, d f : the feature dimension, d d : the input data dimension, and n x : the number of data points.

[0083] Results - 1 In the first experimental set, for three datasets of hypersphere, ad hoc dataset, and polynomial separation, the performance of QRKS was evaluated using a 10-quantum dot device and compared with classical random kitchen sink.

[0084] Figure 7 shows the STM lithography of QMLD700 used to evaluate the performance of QRKS. This device has 10 donor-based quantum dots 208 in the quantum dot array 202, a source 204, a drain 206, and six control gates G1 - G6. The 10 donor quantum dots 208 are phosphorus quantum dots embedded in a silicon substrate. Two data acquisition units 702A, 702B were used to measure the device 700, and a 1:50 voltage divider 704 and a current amplifier 706 were used to amplify the signal.

[0085] Ad hoc dataset The ad-hoc dataset (schematically shown in FIG. 12A) classifies points using a quantum circuit that is presumably difficult to simulate classically. This is based on a low-depth quantum circuit and is designed to be fully separable by variational quantum methods. The dimension of this dataset corresponds to adding qubits to the quantum circuit. The coordinates of each data point within this dataset essentially parameterize the quantum circuit.

[0086] FIG. 8 shows plot 800 comparing the performance of QRKS in 4K and mK measurements when completing an ad-hoc classification task with a classical random kitchen sink. Specifically, this plot shows the performance of QRKS and a classical random kitchen sink when completing this task in various data dimensions of 1D, 2D, and 3D.

[0087] The x-axis of plot 800 indicates the number of features generated, and the y-axis indicates the classification error. The model was trained with approximately 1000 data points and tested with approximately 270 data points.

[0088] The performance of QRKS in 4K and mK measurements is approximately the same in all three dimensions. However, for the 2D dataset, as the number of features generated by the model increases, the performance of cRKS improves compared to QRKS.

[0089] In this dataset, as can be seen from plot 800, the features generated using quantum functions exhibit performance similar to classical features for 1D and 3D data.

[0090] Hypersphere The hypersphere dataset is the n-dimensional version of the circular dataset commonly used in simple ML models. FIG. 12B shows a schematic diagram of the hypersphere dataset.

[0091] n-dimensional coordinates x ∈ [-1, 1] nWhen given, x 2 If ≤ r (r is the radius of the hypersphere), it is considered to be "inside" the hypersphere; otherwise, it is considered to be "outside". The task of this model is to classify points as "inside" or "outside". Since a linear support vector machine (SVM) can only separate points with a hyperplane, in this task, only up to 50% accuracy can be expected.

[0092] Figure 9 shows Plot 900 comparing the performance of QRKS in 4K and mK measurements when completing this classification task with that of a classical random kitchen sink. Specifically, this plot shows the performance of QRKS and the classical random kitchen sink when completing this task in various data dimensions of 1D, 3D, 10D, and 60D. The x-axis is the number of features generated, and the y-axis is the classification error.

[0093] In this dataset, as can be seen from Plot 900, the features generated using quantum functions are significantly more performant than classical features, at least for 60D data. Specifically, as the number of features increases, the classification error decreases. For 60D, the quantum method achieves an error rate of approximately 20%, while classical features have an error rate of 30% - 40%. In lower dimensions, as the number of features increases, the performance of both CRKS and QRKS is comparable, but when the number of features is small, QRKS performs better than CRKS.

[0094] Polynomial root separation dataset The polynomial-separated dataset (see Figure 12C) is composed of the coefficients of a univariate polynomial of degree n - 1. The minimum interval between the two roots of each polynomial is calculated, and values greater than the threshold are marked. A larger-dimensional sample space corresponds to a higher-degree polynomial. In 2D, the coordinates are (x, y). The polynomial is defined using these coordinates, for example, (xa + y), where the variable in this case is a. The roots of this polynomial (xa + y = 0) are used to define the dataset. In 3D, the coordinates are (x, y, z), and the polynomial can be defined as (xa 2 + ya + z = 0), and in 4D, the coordinates can be (x1, x2, x3, x4), and the polynomial can be (x1a 3 + x2a 2 + x3a + x4 = 0). The degree of the polynomial is the highest power of a, which is d - 1, where d is the number of dimensions.

[0095] Figure 10 shows a plot 900 comparing the performance of quantum RKS with classical RKS in 4K and mK measurements. Specifically, this plot shows the performance of QRKS and classical random kitchen sink when completing this task in various data dimensions such as 3D, 5D, 7D, and 10D. The x-axis is the number of features generated, and the y-axis is the model error.

[0096] Quantum features, as in this case, have equivalent performance to classical features in 4K and mK measurements for all data dimensions.

[0097] Results - 2 In the second experiment, the performance of QRKS was evaluated using different quantum dot devices for various datasets and compared with classical random kitchen sink.

[0098] Figure 11 shows a STM micrograph of QMLD1100 used in the second experiment to evaluate the performance of QRKS. This device 1100 includes 75 donor-based quantum dots 208 within a quantum dot array 202, a source 204, a drain 206, and ten control gates G1 to G10. The 75 donor quantum dots 208 are phosphorus quantum dots embedded in a silicon substrate.

[0099] As described above, the Hubbard parameters U, V, and t are determined by the configuration and the distance between the dots 208 within the quantum dot array 202. QMLD1100 is designed to randomize these parameters (U, V, T) to provide a richer dynamics to the reservoir. This was achieved by constructing the quantum dot array 202 on a triangular lattice. To introduce randomness and enhance the potential effectiveness of the device, the positions were randomly jittered in the x coordinate and / or y coordinate, such that each moves slightly by a random amount in a random direction. As a result, the final array seen in Figure 11 was obtained.

[0100] The voltage applied to the control gates G1 to G10 reconfigures the on-site energy level ε i of the site, and the resulting charge transport through the quantum dot array 202 is measured as the current through the leads of the source 204 and drain 206. This current varies according to the quantum state of the quantum dot array 202, which, due to the strong coupling strength between adjacent quantum dots and the low measurement temperature (about 30 mK), is in a strong quantum regime where quantum effects are dominant. The parameter range for which this holds is when the thermal energy is small compared to all other energy scales in the system.

[0101] In this exemplary QMLD1100, the dimension of x e ’ is chosen to be n = 10 to correspond to the number of control gates G1 to G10 within the device. Each element

[0102]

Number

[0103] is directly applied as a voltage to gate Gi, and when measuring the current flowing through the device, the result of the non-linear transformation is returned.

[0104] Each matrix w is generated by randomly sampling each matrix element from a Gaussian distribution with mean μ = 0 and standard deviation σ, and each b is generated by sampling each vector element IID from a uniform distribution in the interval [w, b]. When σ is changed, the volume of the gate space accessible to the algorithm changes, and features with different complexities are obtained.

[0105] To prevent the silicon substrate 102 from being damaged at high voltages and leakage current from occurring between the gates, the voltage applied to the gates must be within a maximum range of [-0.5V, 0.5V]. The voltage generated in

[0110] is calculated, and for features having a range exceeding the allowable voltage range, the corresponding rows of the transformation matrix are resampled until the range becomes sufficiently small or the threshold number of trials is reached. Then, the range of the uniform offset is defined by the range of the voltage of each feature. In all experiments, the source-drain bias was set to 4 mV.

[0106] Dataset QMLD1100 was tested using three synthetic datasets: hypersphere (shown in Fig. 12B), polynomial separation (shown in Fig. 12C), and ad hoc (shown in Fig. 12A). Figs. 12A to 12C show visualizations of the synthetic datasets. Each dataset is composed of points separated into two classes represented by different colors. The role of QRKS is to learn a method of separating points into two classes when only the coordinates of each point are given.

[0107] Furthermore, in each plot, the bright gray and dark gray regions represent the regions where the points of each class exist. Points within the dark region are classified as dark, and points within the bright gray region are classified as bright. The white region is the region where points are not sampled. The bright gray and dark gray points in these exemplary schematic diagrams are the plotted exemplary points and are exemplary points classified into the colors corresponding to the regions where they exist.

[0108] The ad-hoc dataset (shown in Figure 12A) is based on a low-depth quantum circuit and is designed to be fully separable by variational quantum methods. The dimension of this dataset corresponds to adding qubits to the quantum circuit. The coordinates of each data point essentially parameterize the quantum circuit. The results of the quantum circuit are used to classify the data points, and using higher-dimensional data points (corresponding to more values within the vector) is equivalent to operating a similar circuit with more qubits.

[0109] The hypersphere dataset (see Figure 12B) consists of points randomly sampled within an m-dimensional unit hypercube. Then, the points inside the hypersphere are marked, and the model must identify which points are marked. The differently colored dots in Figure 12B represent different classification classes, i.e., the dots classified as inside the hypersphere are one color, and the dots classified as outside the hypersphere are another color.

[0110] The polynomial separation dataset (see Figure 12C) is composed of the coefficients of a univariate polynomial of degree n - 1. The minimum distance between the two roots of each polynomial is calculated, and values greater than the threshold are marked. The larger-dimensional sample space corresponds to higher-degree polynomials. In 2D, the coordinates are (x, y). The polynomial is defined using these coordinates, for example, (xa + y), where the variable in this case is a. The roots of this polynomial (xa + y = 0) are used to define the dataset. In 3D, the coordinates are (x, y, z), and the polynomial is (xa2 It can be defined as +ya + z = 0), and in 4D, the coordinates can be (x1, x2, x3, x4), and the polynomial can be (x1a 3 + x2a 2 + x3a + x4 = 0). The degree of the polynomial is the highest power of a, which is d - 1, where d is the number of dimensions.

[0111] For these three datasets, the output threshold was chosen such that 50% of the input data points were marked, and points that were too close to this threshold were discarded to ensure class separation.

[0112] These three synthetic datasets were chosen to test the QRKS model by presenting various levels of difficulty. The hypersphere is considered an easy dataset because the separation boundary is simple, while the polynomial separation and ad - hoc datasets have more complex separation boundaries. In particular, the ad - hoc dataset is conjectured to be classically difficult to compute, which means that ad - hoc datasets based on a large number of qubits appear completely random to classical computers.

[0113] In addition to selecting datasets of various difficulties, each dataset was tested in various dimensions. This has two effects. First, the sampling density of points decreases exponentially in proportion to the number of dimensions, and since the data points are relatively scarce, the difficulty of the dataset increases. Second, in the case of polynomial separation and ad - hoc, the computational complexity of the function defining the separation boundary increases.

[0114] Since all synthetic data points are based on randomly sampled points, by simply selecting different subsets of the data and redefining the class assignment of each point, the same measurements can be reused to evaluate the performance of the model in all three cases. After measuring a total of 3000 points and discarding the points close to the separation boundary, approximately 2700, 2400, and 1850 points were obtained for the hypersphere, polynomial separation, and ad-hoc datasets, respectively. There was some variation in the number of points in different dimensions. For each dataset, 70% of the points were used as the training set and the remaining 30% were used as the test set.

[0115] QMLD1100 was also tested using the Modified National Institute of Standards and Technology dataset (MNIST), which is an actual dataset. MNIST is a dataset composed of 28×28 pixel images of 70,000 handwritten digits (60,000 training examples and 10,000 test examples) and is a standard benchmark dataset in the field of ML. Figure 12D shows some examples of these images.

[0116] In this experiment, the model was trained on both the complete dataset and a subset of MNIST that contains only the digits 3 and 5, which are two digits that are judged to be the most difficult for a linear classifier to separate. This subset contains 11,551 training examples and 1903 test examples. For the purpose of this experiment, principal component analysis (PCA) decomposition was used to reduce the dimension of the MNIST input data from 784 dimensions to 10 dimensions.

[0117] Model The QRKS model on QMLD1100 was compared with three models: the linear support vector machine (LSVM), the SVM using the radial basis function (RBF) kernel, and the classical version of the random kitchen sink method using cosine linearity (CRKS).

[0118] The LSVM was used to classify the features generated by QMLD1100. Note that in the random kitchen sink process, no linearity other than quantum mapping was introduced. That is, the improvement in performance for the LSVM itself directly results from device 1100 and is not a secondary effect of preprocessing or postprocessing.

[0119] Each model has a scale parameter "gamma" and a regularization parameter "C", and the optimal values of each depend on both the model and the dataset. The scale parameter of CRKS and QRKS represents the width of the Gaussian distribution from which the transformation matrix w is sampled, and the scale parameter of the RBF SVM defines the influence region of each support vector. The regularization parameter defines how much each model is penalized for extreme values within the weight matrix. In other words, regularization is a way to prevent overfitting, which is when a machine learning model memorizes the data within the dataset rather than learning the general trends and patterns that would allow it to generalize to unseen data. An overfitted model tends to perform very poorly when used to classify data points outside the training set. If the parameters learned by the model are extremely large, this may indicate that the model is overfitted. Regularization mitigates this phenomenon by introducing a cost for having large model parameters.

[0120] Figure 13 is a grid of nine subplots showing the performance of each model (RKS (left three subplots), RBF (middle three subplots), and QRKS (right three subplots)) in each dataset (the upper three subplots of the hypersphere, the middle three subplots of polynomial separation, and the lower three subplots of ad hoc) as a function of the hyperparameters used to perform the optimization. The color gradation represents the accuracy of the model, black is the highest error rate (0.5), and white is the lowest error rate (0). The axes of each subplot represent the values of the hyperparameters C (x-axis) and gamma (y-axis). The spots on each tile represent the optimal hyperparameters for each model in each dataset.

[0121] Training and Testing Procedures Before training, each dataset was rescaled to the range [-1, 1], and the output features of QMLD1100 were also scaled back to this range before being input to LSVM. The hyperparameters of scale and regularization were optimized by grid search for each dataset, and each model was trained on a subset of 500 data points for each combination of hyperparameter values and tested on the validation set. Then, the best value leading to the model with the highest performance was used during training on the entire dataset.

[0122] For each synthetic dataset, a total of 1000 features were generated for CRKS and QRKS, and 10,000 features were generated for the MNIST dataset.

[0123] To construct the statistics of the model performance, 300 random training / test splits were generated, and the models randomly initialized in each of these splits were trained and tested. In other words, error bars were created by training many models with random initial conditions and contrasting their accuracies to obtain a more accurate estimate of the mean and standard deviation of the accuracy of this technique.

[0124] Result Figure 14A is a plot 1400 comparing the performance of the QRKS models (4K 1402 and mK 1404) in a polynomial separation dataset with three classical models RBF SVM1405, LSVM1406, and CRKS1408. The x-axis is the dimension of the generated feature space, and the y-axis is the classification error.

[0125] Figure 14B is a plot 1410 comparing the performance of the QRKS models (4K 1412 and mK 1404) in an ad-hoc dataset with three classical models RBF SVM1415, LSVM1416, and CRKS1418. The x-axis is the dimension of the generated feature space, and the y-axis is the classification error.

[0126] Figure 14C is a plot 1420 comparing the performance of the QRKS models (4K 1422 and mK 1424) in a hypersphere dataset with three classical models RBF SVM1425, LSVM1426, and CRKS1428. The x-axis is the dimension of the generated feature space, and the y-axis is the classification error.

[0127] QRKS is consistently significantly more performant than LSVM in all datasets, indicating that the non-linear transformation provided by QMLD1100 is useful and versatile. In addition to this, QRKS performs equivalently compared to the RBF SVM and CRKS methods, despite this technology being far from mature and implemented in noisy hardware. The inherent noise robustness of the Reservoir / Random Kitchen Sink algorithm is evident when comparing the performance of QRKS at 30mK to that at 4K. The performance drops only slightly as the measurement noise increases.

[0128] The plots shown in FIGS. 14A to 14C also show the differences in the difficulty levels of each dataset. The hypersphere dataset can be easily separated by all models (except linear) up to a very large dimension. However, in the case of ad-hoc, there is no model that performs better than randomly guessing for dimensions higher than 3. In the polynomial separation dataset, the performance of the model gradually decreases according to the dimension, so it becomes a dataset suitable for judging the robustness of the model against complexity. QRKS is still inferior in performance to RBF SVM and CRKS, but since it does not converge quickly to random guessing, this technique is general-purpose and supports the fact that it can learn complex separation boundaries.

[0129] FIG. 14D is a table comparing the performance of the QRKS model 1440 in 3-5 subsets of the MNIST dataset with three classical models: RBF SVM 1442, LSVM 1444, and CRKS 1446.

[0130] FIG. 14E is a plot 1450 of the performance of each model as a function of the number of data points (x-axis) in a 22D hypersphere dataset. Specifically, plot 1456 is for the LSVM model, plot 1454 corresponds to the CRKS model, plot 1452 corresponds to the RBF model, plot 1458 corresponds to the mK QRKS model, and plot 1459 corresponds to the 4K QRKS model. What can be seen from FIG. 14E is that at a small number of data points, no model can learn more effectively than the other models. However, as the number of data points increases, the error rates (y-axis) of the CRKS, RBF, and QRKS models decrease, while those of the linear models remain the same.

[0131] FIG. 14F is a plot 1460 showing the performance of each model as a function of the number of data points in a 5D polynomial root separation dataset. The x-axis is the number of generated features, and the y-axis is the classification error.

[0132] Specifically, plot 1466 corresponds to the LSVM model, plot 1454 corresponds to the CRKS model, plot 1462 corresponds to the RBF model, plot 1468 corresponds to the mK QRKS model, and plot 1469 corresponds to the 4K QRKS model. In this plot 14F, it has been revealed that all models except the linear model exhibit relatively similar performance for high-dimensional polynomial root separation datasets.

[0133] Figure 14G is plot 1470 showing the performance of each model as a function of the number of data points in a 2D ad-hoc dataset. The x-axis is the number of features generated, and the y-axis is the classification error.

[0134] Specifically, plot 1476 corresponds to the LSVM model, plot 1474 corresponds to the CRKS model, plot 1472 corresponds to the RBF model, plot 1478 corresponds to the mK QRKS model, and plot 1479 corresponds to the 4K QRKS model. In this plot 14G, it has been revealed that for the 2D ad-hoc dataset, the QRKS model performs better than the linear model, but not as well as the RBF model and the CRKS model.

[0135] Other Variants Another way to operate a quantum ML device is to operate it with a reservoir. In this operating regime, the temporal dynamics of the device are utilized, and the input signal is applied faster than the device can settle. In the case of random kitchen sink, this will cause the output to depend on the order of the data points in the dataset, which is not desirable for datasets such as the aforementioned hypersphere and ad-hoc datasets because each data point in these datasets is independent. However, for other types of datasets such as time series datasets, the data points are not independent and can be ordered.

[0136] In this operating regime, random transformations are still used. Specifically, several random transformations are generated and applied to the input data (e.g., time series), and then provided as various input voltages to the QMLD. The random transformations define paths through the voltage-gate space. The paths within the voltage-gate space are between the first voltage and the second voltage of the input voltages. New features are measured for each random transformation. These features are then measured and can be used for prediction in a machine learning model.

[0137] An exemplary response of the QMLD1100 operating as a reservoir to random binary inputs is shown in FIG. 15. Specifically, FIG. 15A shows the response at 4K, and FIG. 15B shows the response at approximately 30mK. In FIGS. 15A and 15B, the input to the quantum ML device is shown by a dashed line, and the generated features are shown by a solid line. As an example of how individual features change over time, five random features are highlighted in a darker color.

[0138] When the input signal is provided faster than the settling time of the QMLD, the history of the input signal is encoded into an instantaneous quantum state, i.e., the measured value at that point encodes information about both the current input and the past inputs. This ability to extract information about past inputs is called memory, and the distance into the past that the QMLD can remember is called the "memory capacity". In addition to this, this device can non-linearly interact the current input with past inputs. The ability of the QMLD to perform complex non-linear transformations based on points in memory is called "non-linear processing ability", and the ability to perform linear transformations based on points in memory is called "linear processing ability".

[0139] These two properties were measured using the binary inputs of QMLD1100. The results are shown in FIGS. 16A to 16D. Regarding the memory capacity, the accuracy of correctly remembering the t-i-th input (where t is the current time step and i is swept from 1 to 10) is calculated and summed to obtain a metric. Regarding the processing ability, the accuracy of predicting the parity of the inputs from t-i to t is calculated and again summed to obtain another metric. These metrics mainly vary with two hyperparameters, namely gamma and the ramp length. Similar to the random kitchen sink, the gamma variable controls the size of the random transformation and defines the amount of the voltage gate space accessible to the model. QMLD1100 was measured at an input rate of 500 kHz, but by changing the number of points required for the ramp between consecutive inputs along the path of the gate space, the rate at which data points are measured can be changed. This variable is called the ramp length.

[0140] FIGS. 16A to 16D show the memory capacity and processing ability of the device at 4K and about 30 mK according to both hyperparameters. Specifically, FIGS. 16A and 16B show the memory capacity at 4K and 30 mK, respectively. The scale represents the memory capacity, that is, the number of past data points that the device can remember. As shown in the figure, the memory capacity of the device can vary depending on the gamma and / or ramp length selected at both temperatures. Furthermore, the maximum memory capacity of this device (i.e., QMLD1100) is about 6 data points at 4K and 30 mK. It will be understood that the memory capacity of the device can be increased by increasing the time it takes for the device to stabilize.

[0141] Figures 16C and 16D show the processing capabilities of the device at 4K and 30mK, respectively. The scale represents the processing capabilities of the device based on the number of data points in the memory. In this example, the device was programmed to identify the number of 1s at the data points in the memory. As can be seen from the figure, the processing capabilities of the device vary with gamma and lamp length at both temperatures. Further, the device can perform processing on up to six data points from the memory.

[0142] References to prior art in this specification do not admit, nor do they suggest, that this prior art forms part of the common general knowledge in any jurisdiction, that this prior art is understood and relevant by a person skilled in the art, and / or that it is reasonably foreseeable that this prior art could be combined with other prior art.

[0143] As used in this specification, unless the context requires otherwise, the term "comprise" and variations of the term, such as "comprising", "comprises" and "comprised", are not intended to exclude further additions, components, integers or steps.

Claims

1. A method for generating quantum features for a machine learning model, comprising: providing a quantum ML device comprising one or more quantum dots, one or more source gates, one or more drain gates, and one or more control gates; converting input data of the machine learning model into a first voltage; applying the first voltage to the one or more control gates and / or source gates and / or drain gates; applying a second voltage to one or more of the one or more source gates; measuring a signal at one or more of the one or more drain gates; analyzing the measured signal to determine values of one or more parameters; interpreting the values of the one or more parameters as a non-linear mapping of the input data used in the machine learning model; A method comprising the above.

2. The converting the input data into the first voltage comprises: performing a random transformation of the input data; converting the randomly transformed input data into the first voltage; The method according to claim 1, comprising the above.

3. The converting the input data into the first voltage comprises directly mapping the input data to the first voltage. The method according to claim 1, comprising the above.

4. The interpreting the values of the one or more parameters comprises combining the values of the one or more parameters as features for the machine learning model. The method according to claim 2 or 3, comprising the above.

5. The converting the input data into the first voltage comprises combining data points of the input data in pairs and converting the combined data points into a combined voltage. The applying the first voltage to the one or more control gates comprises applying the combined voltage to the one or more control gates. The method according to claim 1, comprising the above.

6. The interpreting the values of the one or more parameters comprises determining a distance metric or similarity score between the values of the one or more parameters. The method according to claim 5, comprising the above.

7. The quantum ML device includes a plurality of source gates, a plurality of drain gates, and a plurality of control gates, and the quantum ML device is used as a quantum random kitchen sink device. The method according to any one of claims 1 to 6.

8. The quantum ML device includes one source gate, a number of drain gates that matches the desired feature dimension, and a number of control gates that matches the dimension of the input data, and the quantum ML device is used as a quantum extreme learning machine. The method according to any one of claims 1 to 6.

9. The quantum ML device includes one source gate, one drain gate, and a number of control gates that is twice the dimension of the input data, and the quantum ML device is used as a quantum kernel learning machine. The method according to any one of claims 1 to 6.

10. Further including manufacturing the quantum ML device, and manufacturing the quantum ML device includes fabricating a bulk layer of a semiconductor substrate, fabricating a second semiconductor layer, exposing a clean crystal surface of the second semiconductor layer to dopant molecules to generate an array of dopant dots on the exposed surface, annealing the arrayed surface to incorporate dopant atoms of the dopant molecules into the second semiconductor layer, forming the one or more gates, the one or more source leads, and the one or more drain leads, The method according to any one of claims 1 to 9.

11. The method according to claim 10, wherein the one or more control gates are formed in the same plane as the dopant dots.

12. Further including depositing a dielectric material on the second semiconductor layer, and the one or more control gates are formed on the dielectric material. The method according to claim 10.

13. The method according to any one of claims 10 to 12, wherein the dopant dots are phosphorus dots.

14. The method according to any one of claims 10 to 12, wherein the second semiconductor layer is silicon 28.

15. One or more quantum dots, One or more source gates, One or more drain gates, A quantum ML device comprising one or more control gates, wherein The quantum ML device is Applying a first voltage corresponding to the input data of the machine learning model to the one or more control gates and / or source gates and / or drain gates; Applying a second voltage to one or more of the one or more source gates; Measuring a signal at one or more of the one or more drain gates; Analyzing the measured signal to determine values of one or more parameters; Interpreting the values of the one or more parameters as a non-linear mapping of the input data used in the machine learning model; A quantum ML device used to generate quantum features for the machine learning model by the above. **Claim 16** The quantum ML device according to claim 15, comprising a plurality of source gates, a plurality of drain gates, and a plurality of control gates, wherein the quantum ML device is used as a quantum random kitchen sink device. **Claim 17** The quantum ML device according to claim 15, comprising one source gate, a number of drain gates that matches the desired feature dimension, and a number of control gates that matches the dimension of the input data, wherein the quantum ML device is used as a quantum extreme learning machine. **Claim 18** The quantum ML device according to claim 15, comprising one source gate, one drain gate, and a number of control gates that is twice the dimension of the input data, wherein the quantum ML device is used as a quantum kernel learning machine. **Claim 19** The quantum ML device according to any one of claims 15 to 18, wherein the one or more control gates are formed in the same plane as the quantum dot. **Claim 20** The quantum ML device according to any one of claims 15 to 19, wherein the quantum dot is a silicon dot.