Managing feature vectors for classification procedures
Patent Information
- Application Number
- PCT/US2024/052000
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-24
- Filing Date
- 2024-10-18
- Publication Date
- 2025-12-26
AI Technical Summary
Existing quantum computers face challenges in efficiently solving graph classification and similarity problems, particularly in determining if one graph is a vertex minor of another, due to the NP-Complete nature of the decision problem and the lack of effective performance evaluations of quantum-enhanced random samplers like Gaussian Boson Sampling (GBS) compared to classical algorithms.
A hybrid quantum-classical algorithm using Gaussian Boson Sampling (GBS) to generate feature vectors from linear-optical interferometers, combined with machine learning models like SVM, to classify graph pairs based on vertex-minor equivalence, allowing for a trade-off between input squeezing and classification accuracy through repeated trials.
The hybrid algorithm achieves quantum acceleration in solving graph classification problems, outperforming classical algorithms in runtime by up to a factor of 3 for 12-node vertex-minor instances, despite exponential time complexity, using a readily-realizable GBS device with reduced squeezing and repeated trials.
Smart Images

Figure US2024052000_26122025_PF_FP_ABST
Abstract
Description
Attorney Docket No. UAZ2-208-A-WO / UA23-240MANAGING FEATURE VECTORS FOR CLASSIFICATION PROCEDURESCROSS-REFERENCE TO RELATED APPLICATION(S)
[0001] This application claims priority to and the benefit of U.S. Provisional Patent ApplicationSerial No. 63 / 545,418, entitled “MANAGING FEATURE VECTORS FOR CLASSIFICATIONPROCEDURES,” filed October 24, 2023, which is incorporated herein by reference.STATEMENT AS TO FEDERALLY SPONSORED RESEARCH
[0002] This invention was made with government support under Grant No. 1842559, awardedby the National Science Foundation (NSF), and under grant No. DE-AC05-00OR22725, awardedby the Department of Energy (DOE). The government has certain rights in the invention.TECHNICAL FIELD
[0003] This disclosure relates to managing feature vectors for classification procedures.BACKGROUND
[0004] In recent years, Noisy Intermediate-Scale Quantum (NISQ) processors have been usedto generate random samples from complex probability distributions that are hard to sample fromusing a classical computer. For example, Gaussian Boson Sampling (GBS) generates randomsamples of photon-click patterns from a class of probability distributions that may be difficult fora classical computer to sample from, especially as the GBS size grows. Despite demonstrations ofquantum supremacy using GBS, Boson Sampling, and instantaneous quantum polynomial (IQP),systematic performance evaluations of the power of these quantum-enhanced random samplerswhen applied to real problems, and performance comparisons with classical algorithms have beenlacking. SUMMARY
[0005] In one aspect, in general, a method comprises: generating a plurality of sets of vectorsbased on respective output signals from at least M+N detectors configured to detect differentrespective output modes emitted from one or more linear-optical interferometers, the plurality ofsets of vectors comprising: a first set of vectors, where each vector in the first set of vectorscorresponds to an M-mode quantum state emitted from at least one of the linear-opticalinterferometers, and a second set of vectors, where each vector in the second set of vectorscorresponds to an N-mode quantum state emitted from at least one of the linear-opticalinterferometers; and training a classifier configured to classify each of a plurality of pairs ofX-dimensional input vectors as a member of one of two or more different classes based on an X-1dimensional hyperplane that is determined at least in part by the training, where X is at least aslarge as M and is at least as large as N, the training comprising: receiving training datacomprising at least one vector from the first set of vectors and at least one vector from the secondset of vectors, and updating the X-1 dimensional hyperplane based at least in part on the receivedtraining data.
[0006] Aspects can include one or more of the following features.
[0007] The method further comprises: encoding a representation of a first graph onto inputmodes coupled into the one or more linear-optical interferometers for at least one vector in thefirst set of vectors, and encoding a representation of a second graph onto input modes coupled intothe one or more linear-optical interferometers for at least one vector in the second set of vectors.
[0008] The two or more different classes comprise a first class in which the two vectors in thepair of X-dimensional input vectors correspond to respective graphs that are vertex-minorequivalent to each other, and a second class in which the two vectors in the pair of X-dimensionalinput vectors correspond to respective graphs that are not vertex-minor equivalent to each other.
[0009] In another aspect, in general, a method comprises: coupling input modes into one ormore linear-optical interferometers; generating a first vector and a second vector based onrespective output signals from two or more of a plurality of detectors configured to detectdifferent respective output modes emitted from the one or more linear-optical interferometers;providing input to a machine learning model based at least in part on information associated withthe first vector and information associated with the second vector.
[0010] Aspects can include one or more of the following features.
[0011] The method further comprises: classifying a pair of graphs corresponding to the firstand second vectors, respectively, as a member of one of two or more different classes based onperforming multiple executions of the machine learning model and selecting a classification basedon a majority of the multiple executions that produce an identical classification.
[0012] The input modes are generated by nonlinear spontaneous four-wave mixing of at least afirst optical wave and a second optical wave.
[0013] In another aspect, in general, a method comprises: generating at least one set of vectorsbased on respective output signals from a plurality of detectors configured to detect differentrespective output modes emitted from one or more linear-optical interferometers, the generatingcomprising: coupling input modes into the one or more linear-optical interferometers for each ofa plurality of iterations, and configuring a plurality of phase shifts and a plurality of beamsplitting ratios imposed by the linear-optical interferometers to reduce an amount of squeezingassociated with the output modes for each of the plurality of iterations relative to a maximumamount of squeezing attainable based at least in part on a selected constant factor applied to amatrix that is based on an N × N adjacency matrix of a graph; and training a machine learningmodel based at least in part on training data comprising N-dimensional vectors based onrespective output signals from the plurality of detectors.
[0014] Aspects can include one or more of the following features.
[0015] The machine learning model comprises at least one of: a support vector machinemodel, a neural network model, a random decision forest model, a random forest model, or adecision tree model.
[0016] The method further comprising encoding a representation of a graph onto input modescoupled into the one or more linear-optical interferometers.
[0017] In another aspect, in general, a method comprises: generating at least one set of vectorsbased on respective output signals from a first set of detectors configured to detect differentrespective output modes emitted from a first independent linear-optical interferometer, thegenerating comprising: coupling input modes into the first independent linear-opticalinterferometer for each of a plurality of iterations, and configuring a plurality of phase shifts and aplurality of beam splitting ratios imposed by the first independent linear-optical interferometer;generating at least one set of vectors based on respective output signals from a second set ofdetectors configured to detect different respective output modes emitted from a secondindependent linear-optical interferometer independent from the output modes emitted from thefirst independent linear-optical interferometer, the generating comprising: coupling input modesinto the second independent linear-optical interferometer for each of a plurality of iterations, andconfiguring a plurality of phase shifts and a plurality of beam splitting ratios imposed by thesecond independent linear-optical interferometer; and training a machine learning model based atleast in part on training data comprising M-dimensional vectors based on respective outputsignals from the first set of detectors, and N-dimensional vectors based on respective outputsignals from the second set of detectors.
[0018] Aspects can include one or more of the following features.
[0019] The first independent linear-optical interferometer and the second independentlinear-optical interferometer comprise independent portions of a common linear-opticalinterferometer.
[0020] The plurality of iterations is determined based at least in part on a probability of errorassociated with an iteration of the plurality of iterations.
[0021] In another aspect, in general, an apparatus comprises: a Gaussian Boson Sampling(GBS) module configured to generate a plurality of sets of vectors based on respective outputsignals from at least M+N detectors configured to detect different respective output modes emittedfrom one or more linear-optical interferometers, the plurality of sets of vectors comprising: a firstset of vectors, where each vector in the first set of vectors corresponds to an M-mode quantumstate emitted from at least one of the linear-optical interferometers, and a second set of vectors,where each vector in the second set of vectors corresponds to an N-mode quantum state emittedfrom at least one of the linear-optical interferometers; and at least one processor configured totrain a classifier configured to classify each of a plurality of pairs of X-dimensional input vectorsas a member of one of two or more different classes based on an X-1 dimensional hyperplane thatis determined at least in part by the training, where X is at least as large as M and is at least aslarge as N, the training comprising: receiving training data comprising at least one vector fromthe first set of vectors and at least one vector from the second set of vectors, and updating the X-1dimensional hyperplane based at least in part on the received training data.
[0022] Aspects can include one or more of the following features.
[0023] Each linear-optical interferometer comprises at least one beam-splitter that isconfigured to receive a first optical mode and generate, based at least in part on the received firstoptical mode, a second optical mode and a third optical mode.
[0024] At least one of the linear-optical interferometers is configured to generate the M-modequantum state based at least in part on a decomposition of an adjacency matrix of a first graph.
[0025] At least one of the linear-optical interferometers is configured to generate the N-modequantum state based at least in part on a decomposition of an adjacency matrix of a second graph.
[0026] In another aspect, in general, a method comprises: generating a plurality of sets ofvectors, the plurality of sets of vectors comprising: a first set of vectors, where each vector in thefirst set of vectors corresponds to a set of M eigenvalues of a Laplacian of an M × M adjacencymatrix of a graph, and a second set of vectors, where each vector in the second set of vectorscorresponds to a set of N eigenvalues of a Laplacian of an N × N adjacency matrix of a graph; andtraining a classifier configured to classify each of a plurality of pairs of X-dimensional inputvectors as a member of one of two or more different classes based on an X-1 dimensionalhyperplane that is determined at least in part by the training, where X is at least as large as M andis at least as large as N, the training comprising: receiving training data comprising at least onevector from the first set of vectors and at least one vector from the second set of vectors, andupdating the X-1 dimensional hyperplane based at least in part on the received training data.
[0027] Aspects can include one or more of the following features.
[0028] The two or more different classes comprise a first class in which the two vectors in thepair of X-dimensional input vectors correspond to respective graphs that are vertex-minorequivalent to each other, and a second class in which the two vectors in the pair of X-dimensionalinput vectors correspond to respective graphs that are not vertex-minor equivalent to each other.
[0029] The first set of vectors is generated based on respective output signals from M detectorsconfigured to detect different respective output modes emitted from one or more linear-opticalinterferometers and the second set of vectors is generated based on respective output signals fromN detectors configured to detect different respective output modes emitted from one or morelinear-optical interferometers.
[0030] In another aspect, in general, a method comprises: generating a first set of vectorsbased at least in part on a first graph, where each vector in the first set of vectors corresponds to arespective set of eigenvalues of a respective Laplacian of a respective adjacency matrix of arespective graph determined at least in part on the first graph; and classifying a plurality of pairsof input vectors as a member of one of two or more different classes, where at least one of theplurality of pairs of input vectors comprises a first vector from the first set of vectors and a secondvector corresponding to a set of eigenvalues of a Laplacian of an adjacency matrix of a secondgraph.
[0031] Aspects can include one or more of the following features.
[0032] A vector in the first set of vectors corresponds to a third graph that is related to the firstgraph by at least one of (1) local complementation of one or more nodes of the first graph or (2)deletion of one or more nodes of the first graph.
[0033] The first graph and third graph are vertex-minor equivalent to each other.
[0034] The at least one of (1) local complementation of one or more nodes of the first graph or(2) deletion of one or more nodes of the first graph is based at least in part on a random orpseudo-random algorithm.
[0035] The two or more different classes comprise a first class in which the two vectors in thepair of input vectors correspond to respective graphs that are vertex-minor equivalent to eachother, and a second class in which the two vectors in the pair of input vectors correspond torespective graphs that are not vertex-minor equivalent to each other.
[0036] The method further comprises: classifying each of a set of pairs of input vectors ascorresponding to a first or second class of the two or more different classes, where a first numberof the set of pairs of input vectors are classified as the first class and a second number of the set ofpairs of input vectors are classified as the second class.
[0037] The method further comprises: classifying the first graph and the second graph asvertex-minor equivalent to each other if the first number is greater than the second number, andclassifying the first graph and the second graph as not vertex-minor equivalent to each other if thefirst number is less than the second number.
[0038] Aspects can have one or more of the following advantages.
[0039] Some of the techniques disclosed herein may be applied to challenges facing quantumcomputers. For example, the decision problem to determine if G′ is a vertex minor of G is animportant problem underlying Clifford manipulation of Stabilizer quantum states, but is known tobe NP-Complete. Some of the techniques disclosed allow for reduced computational complexityfor computing such a decision problem.
[0040] Some of the techniques disclosed herein map a graph onto a Gaussian Boson Samplingstate by more efficient means than other methods, and the GBS’ purported quantum-enhancedcomputational prowess suggests that the GBS might provide quantum acceleration in solvingproblems defined on graphs, such as classification, optimization, and search problems.
[0041] The disclosed graph embedding techniques allows for a lower squeezing amount at theGBS input, a hard-to-produce quantum optical resource, at the expense of a lower one-shotclassification accuracy, which in turn results in a larger number of repeated trails to obtain adesired (high) target accuracy. We also disclose a classical spectral algorithm for the vertex minorproblem that uses the graph Laplacians’ eigenvalues and a support vector machine (SVM)classifier, which we show outperforms various state-of-the-art classical graph classificationalgorithms in prophetic simulations. We show in prophetic examples that with a readily-realizableon-chip GBS, our hybrid special-purpose processor can outperform the aforesaid classicalalgorithm when executed on the most powerful MacBook Pro.
[0042] Other features and advantages will become apparent from the following description,and from the figures and claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The disclosure is best understood from the following detailed description when read inconjunction with the accompanying drawing. It is emphasized that, according to commonpractice, the various features of the drawing are not to-scale. On the contrary, the dimensions ofthe various features are arbitrarily expanded or reduced for clarity.
[0044] FIGS. 1A-1F shows schematic diagrams of example methods for implementing aquantum-assisted algorithm.
[0045] FIG. 2A shows a flowchart of an example procedure for obtaining a vertex minor of agraph.
[0046] FIG. 2B shows a schematic diagram of an example framework for the implementationof a quantum-assisted algorithm.
[0047] FIG. 3 shows a prophetic plot of the accuracy of an example quantum-assistedalgorithm as a function of squeezing, loss, and n(N).
[0048] FIG. 4 shows a prophetic plot of accuracy, as a function of sample length, for aclassical algorithm
[0049] FIG. 5 shows a prophetic plot of accuracy, as a function of sample length, for aquantum-assisted algorithm.
[0050] FIG. 6 shows prophetic plots of error, as a function of nodes in the graph, for variousclassical algorithms.
[0051] FIG. 7 shows prophetic plots of time taken (i.e., execution time), as a function of nodesin the graph, for various quantum and classical algorithms.
[0052] FIG. 8 shows all k ∈ {3,4,5} connected non-isomorphic graphlets we used for ourgraphlet sampling kernel algorithm.
[0053] FIG. 9 shows an example of the Weisfeiler-Lehman algorithm.
[0054] FIG. 10 shows example network classification using adjacency matrix embeddings forour vertex minor graphs.
[0055] FIGS. 11A and 11B show coupling loss terms together in an input mode of amultimode interferometer.DETAILED DESCRIPTION
[0056] Herein, examples of quantum-assisted algorithms are described. As used herein, aquantum-assisted algorithm refers to an algorithm that uses one or more modules capable ofquantum processing, including quantum processing of optical signals, such as in a GBS device,and may use the output of such quantum processing for configuring other modules, includingclassical modules. For example, one type of quantum-assisted algorithm is a hybridquantum-classical algorithm to solve the NP-complete problem of determining if two givengraphs are a vertex minor of one another. The presented pair of graphs are encoded in GBSdevices and the generated samples serve as feature vectors in a support vector machine (SVM)classifier. A method for mapping a graph into a GBS that allows trading between the one-shotclassification accuracy and the amount of input squeezing needed, a hard-to-producequantum-optical resource, followed by repeated trials and a majority vote decision to reach anoverall desired accuracy is disclosed. Further disclosed is a classical algorithm based on graphspectra, which is shown to outperform several graph-similarity algorithms for the vertex-minorproblem. We compare the performance of our hybrid quantum-classical algorithm with the aboveclassical algorithm and analyze their time versus problem-size scaling, to yield a desiredclassification accuracy, using repeated trials. Prophetic simulation results suggest that with anear-term realizable GBS device—5 dB pulsed squeezers, 12-mode unitary, and reasonableassumptions on coupling efficiency, on-chip losses, and detection efficiency of photon numberresolving detectors—the GBS processor can solve 12-node vertex-minor instances with about 3fold lower time compared to a powerful desktop computer.
[0057] In some implementations, training a SVM classifier can comprise constructing ahyperplane, or set of hyperplanes based on training data. In some examples, a SVM classifier canfind a boundary, or hyperplane, in N-dimensional space that can classify data points in the trainingdata. A SVM classifier can then classify input data as a member of one of two or more differentclasses based at least in part on the hyperplane. For X-dimensional input vectors of training data,a SVM classifier can construct an X-1 dimensional hyperplane such that the classifier can classifyeach of a plurality of pairs of X-dimensional vectors as a member of one or two or more differentclasses. In some implementations, a hyperplane can be updated based on training data.
[0058] In GBS, N independent squeezed-light pulses are sent through an N-port passivelinear-optical interferometer, and each output mode is detected by a photon number resolving(PNR) detector. The random click pattern that appears at the N PNR detectors can be samplesfrom a distribution that can be classically hard to sample from. There are multiple practicalapplications of GBS, which include: graph similarity, point processes, graphs problems such asthe Maximum Clique problem, the Dense Subgraph problem and various machine-learninginspired primitives.
[0059] An N-node graph can be embedded into an N-mode GBS circuit by mapping thegraph’s adjacency matrix into the covariance matrix of the N-mode Gaussian quantum state at theoutput of the linear-optical interferometer, such that the probability mass function of the outputPNR click pattern is a function of the Hafnian of the graph’s adjacency matrix. This mapping of agraph onto a GBS, and the GBS’ purported quantum-enhanced computational prowess, suggeststhat the GBS might provide quantum acceleration in solving problems defined on graphs, such asclassification, optimization, and search problems.
[0060] Herein we consider a GBS-based solution to the NP-Complete decision problem ofdetermining if one graph is a vertex minor of another. A graph G′ is a vertex minor of graph G ifG′ can be obtained via a series of local complementations and vertex deletions starting with G.Local complementation applied on a vertex inverts the neighborhood of that vertex. For instance,for every pair of neighboring vertices, if there is an edge, it is deleted; and if there is no edge, oneis created).
[0061] FIG. 1A shows an example method 100A for implementing a quantum-assistedalgorithm. The method 100A utilizes linear-optical interferometers 102A-102B. Eachlinear-optical interferometer 102A-102B emits different respective output modes 103A-103D.Detectors 104A-104D are configured to detect the different respective output modes 103A-103Demitted from the linear-optical interferometers 102A-102B. In some examples, the output modes103A-103B can be an M-mode quantum state and at least M detectors 104A-104B can beconfigured to detect the optical modes 103A-103B. In some examples, the output modes103C-103D can be an N-mode quantum state, and at least N detectors 104C-104D can beconfigured to detect the optical modes 103C-103D. A plurality of sets of vectors 107A-107B aregenerated based on respective output signals 106A-106D from detectors 104A-104D. Theplurality of sets of vectors 107A-107B comprises a first set of vectors 107A, where each vector108A-108N in the first set of vectors 107A corresponds to an M-mode quantum state emitted fromthe linear-optical interferometer 102A. The plurality of sets of vectors 107A-107B also comprisesa second set of vectors 107B, where each vector 109A-109N in the second set of vectors 107Bcorresponds to an N-mode quantum state emitted from the linear-optical interferometer 102B.The plurality of sets of vectors 107A-107B are utilized to train a classifier 110 configured toclassify each of a plurality of pairs of X-dimensional input vectors as a member of one of two ormore different classes based on an X-1 dimensional hyperplane that is determined at least in partby the training, where X is at least as large as M and is at least as large as N. In some examples,the training comprises receiving training data comprising at least one vector 108N from the firstset of vectors 107A and at least one vector 109N from the second set of vectors 107B, andupdating the X-1 dimensional hyperplane based at least in part on the received training data.
[0062] In some implementations, an apparatus that can perform the example method 100A forimplementing a quantum-assisted algorithm can comprise a Gaussian Boson Sampling modulethat is configured generate the sets of vectors 107A-107B as described above. In someimplementations, the classifier 110 can be trained by a processor configured to perform theprocesses described above.
[0063] FIG. 1B shows an example method 100B for implementing a quantum-assistedalgorithm. The method 100B comprises coupling input modes 111 into one or more linear-opticalinterferometers 112. The linear-optical interferometer 112 emits output modes 113A-113N. Aplurality of detectors 114A-114N is configured to detect different respective output modes113A-113N emitted from the linear-optical interferometer 112. Each detector 114A-114N isassociated with a respective output signal 115A-115N. A first vector 116A is generated based onthe output signal 115A from the detector 114A and a second vector 116B is generated based onthe output signal 115B from the detector 114B. Input 118 is provided to a machine learningmodel 119 based at least in part on information 117A associated with the first vector 116A andinformation 117B associated with the second vector 116B.
[0064] FIG. 1C depicts an example method 100C for implementing a quantum-assistedalgorithm. The method 100C comprises coupling input modes 120A-120N into a linear-opticalinterferometer 121 for each of a plurality of iterations. The linear-optical interferometer emitsoutput modes 122A-122N. The linear-optical interferometer 121 can be configured to impose aplurality of phase shifts and a plurality of beam splitting ratios to reduce an amount of squeezingassociated with the output modes 122A-122N for each of the plurality of iterations relative to amaximum amount of squeezing attainable based at least in part on a selected constant factorapplied to a matrix that is based on an N × N adjacency matrix of a graph (not shown). A pluralityof detectors 123A-123N is configured to detect different respective output modes 122A-122Nemitted from the linear-optical interferometer 121. Each detector 123A-123N has a respectiveoutput signal 124A-124N. The output signals 124A-124N are utilized to generate a set of vectors125 comprising vectors 126A-126N. In some examples, vectors 126A-126N are N-dimensional.The set of vectors 125 is used to train a machine learning module 127.
[0065] FIG. 1D depicts an example method 100D for implementing a quantum-assistedalgorithm. The method 100D utilizes a first independent linear-optical interferometer 128A and asecond independent linear-optical interferometer 128B. The first independent linear-opticalinterferometer 128A can be configured to impose a plurality of phase shifts and a plurality ofbeam splitting ratios on input modes 129A-129N coupled into the first independent linear-opticalinterferometer 128A. The input modes 129A-129N are coupled into the first independentlinear-optical interferometer 128A for each of a plurality of iterations. The second independentlinear-optical interferometer 128B can be configured to impose a plurality of phase shifts and aplurality of beam splitting ratios on input modes 130A-130N coupled into the second independentlinear-optical interferometer 128B. The input modes 130A-130N are coupled into the secondindependent linear-optical interferometer 128B for each of a plurality of iterations. The firstindependent linear-optical interferometer 128A emits optical modes 131A-131B. Detectors132A-132B are configured to detect the different respective optical modes 131A-131B. Thedetectors 132A-132B form a first set of detectors 133A. The second independent linear opticalinterferometer 128B emits optical modes 131C-131D. Detectors 132C-132D are configured todetect the different respective optical modes 131C-131D. The detectors 132C-132D form asecond set of detectors 133B. Each of the detectors 132A-132D has respective output signals134A-134D. A first set of vectors 135A is generated based on respective output signals134A-134B from the first set of detectors 133A. A second set of vectors 135B is generated basedon respective output signals 134C-134D from the second set of detectors 133B. A machinelearning model 138 is trained based at least in part on training data comprising M-dimensionalvectors 136A-136N based on respective output signals 134A-134B from the first set of detectors133A, and N-dimensional vectors 137A-137N based on respective output signals 134C-134Dfrom the second set of detectors 133B.
[0066] FIG. 1E depicts an example method 100E for implementing a quantum-assistedalgorithm. The method 100E comprises generating a plurality of sets of vectors 142A-142B. Theplurality of sets of vectors 142A-142B comprises a first set of vectors 142A. Each vector143A-143N in the first set of vectors 142A corresponds to a set of M eigenvalues of a Laplacianof an M × M adjacency matrix 141A of a graph 140A. The plurality of sets of vectors 142A-142Balso comprises a second set of vectors 142B. Each vector 144A-144N in the second set of vectors142B corresponds to a set of M eigenvalues of a Laplacian of an N × N adjacency matrix 141B ofa graph 140B. A classifier 145 is trained by receiving training data comprising at least one vector143N from the first set of vectors 142A and at least one vector 144N from the second set ofvectors 142B. The classifier 145 is configured to classify each of a plurality of pairs ofX-dimensional input vectors as a member of one of two or more different classes based on an X-1dimensional hyperplane that is determined at least in part by the training, where X is at least aslarge as M and is at least as large as N. Training the classifier 145 also comprises updating theX-1 dimensional hyperplane based at least in part on the received training data.
[0067] FIG. 1F depicts an example method 100F for implementing a quantum-assistedalgorithm. The method 100F comprises generating a first set of vectors 152A based at least in parton a first graph 150A. Each vector 153A-153N in the first set of vectors 152A corresponds to arespective set of eigenvalues of a respective Laplacian of a respective adjacency matrix 151A of arespective graph 150B determined at least in part based on the first graph 150A. A second set ofvectors 152B is also generated from a second graph 150C. Each vector 154A-154N in the secondset of vectors 152B corresponds to a respective set of eigenvalues of a Laplacian of an adjacencymatrix 151B of the second graph 150C. A plurality of pairs of input vectors 155A-155B areclassified at a member of one of two or more different classes. At least one of the plurality ofpairs of input vectors 155A comprises a first vector 153A from the first set of vectors 152A and asecond vector 154A from the second set of vectors 152B.
[0068] FIG. 2A shows an example procedure for obtaining a vertex minor of a graph. In thisexample, a 4-node graph 206 that is a vertex minor of a 5-node graph 202 is determined. A firstgraph 202 undergoes local complementation at vertex 5, which results in a second graph 204. Athird graph 206 is obtained by vertex deletion of node 4 of the second graph 204. The third graph206 is a vertex minor of the first graph 202, i.e., a graph that is obtained via localcomplementations applied at a collection of vertices and a deletion of a collection of vertices, onthe original graph.
[0069] FIG. 2B shows an example framework for the implementation of a quantum-assistedalgorithm. The algorithm takes two graphs, denoted as G1 and G2, as inputs. Adjacency matrices208A and 208B are generated for the respective graphs G1 and G2. Autonne-Takagidecomposition values 210A, 210B of the adjacency matrices 208A, 208B are computed andutilized to program a respective GBS 212A, 212B. Samples are extracted from each GBS 212A,212B. The sampling algorithm is executed multiple times (denoted as n(N)) and results areaggregated using majority voting. For ML training, a dataset with 500 vertex-minor graph pairsand 500 not vertex-minor graph pairs can be created. Each pair is processed through the samearchitecture, producing a single-shot sample. These samples serve as training data for the MLalgorithm, which is labeled “output”.
[0070] FIG. 3 shows a prophetic plot of the accuracy of an example quantum-assistedalgorithm as a function of squeezing, loss, and n(N), including the value of n(N) required toachieve a 98% accuracy. For each value of squeezing and loss, we average over 7 graph pair’sn(N) value. We observe that as the value of maximum squeezing per mode increases, the value ofn(N) decreases and as loss increases, the value of n(N) increases.
[0071] Herein we propose a randomized classical algorithm that utilizes the spectral propertiesof the graphs, i.e., the eigenvalues of their Laplacians, and uses them in a Support Vector Machine(SVM) to decide whether one is a Vertex Minor of another. We show this algorithm outperformsseveral classical algorithms for graph classification. We also propose a hybrid quantum-classicalalgorithm that encodes the two graphs in two GBS devices and uses the photon-click-patternsamples as feature vectors for an SVM classifier. We disclose a graph embedding algorithm thatallows for a trade-off between the one-shot sampling accuracy from one run of the GBS with themaximum amount of per-mode squeezing used at the input. Thereafter, several runs of the GBSand a majority-vote decision at the end to declare the final verdict, may be used to meet thedesired overall classification accuracy. We discuss a runtime versus problem size complexity forour hybrid quantum-classical computer, and lend evidence towards its time-scaling superiorityover using a classical computer, even though both ultimately can have an exponential timecomplexity. Also disclosed is that using 5 dB pulsed squeezing per mode at the input of a12-mode on-chip GBS generated using a nonlinear spontaneous four-wave mixing (sFWM)process with a mode-locked fiber laser pump of 10 MHz repetition rate, 80% coupling efficiency(squeezing chip to photonic integrated circuit (PIC) realizing the interferometer), 0.25 dB / cmon-chip losses, and 95% PNR detector efficiency, with an appropriate number of repeated trials toget the final classification accuracy up to a desired value, prophetic results show that thequantum-classical algorithm outperforms the aforesaid classical spectral algorithm (also providedwith an appropriate number of repeated trials to get to the same desired accuracy) for up to 12node vertex-minor problems, by anywhere between a factor of 103 to a factor of 102 in runtime.
[0072] Consider a graph G(V,E) defined over a vertex set V and edge set E, with |V | = Nnodes. Let us say its adjacency matrix is A, which is a N ×N symmetric matrix. Define:^ ^A ^ (1)where c is a constant chosen such that 0 < c < 1 / smax, where smax is the maximum singular valueof A. Let us consider an N-mode pure Gaussian state defined uniquely by a 2N ×2N covariancematrix σ, with ^ ^σ = Q− I / 2 with Q = (I −XÃ)−1,X = ^0 I^ ^^ , (2)where I is the 2N ×2Non each of the N modes of this state, we would get a random vector of photon number outcomes(henceforth called a GBS sample) n= [n1, ....,nM], where ni is the number of photons detected inthe i-th mode, the probability mass functionof which is given by:p(n) = √1Haf2(An), (3)where n! = n !n ! ...n !,12 N and Ancolumns of A. For example, if the sample is n j = 1, 1 ≤ j ≤ N, then An = A. If some of the n jvalues are zero (and rest are 1), An would be obtained by removing from A the j-th rows and j-thcolumns corresponding to the zero n j values. If some of the n j values of a sample n are > 0, wewould obtain An by repeating those ( j-th) rows and columns of A, respectively n j numbers oftimes. The hafnian of a N ×N symmetric matrix C:Haf(C) =∑Π(u,v)∈πCu,v, (4)π∈PN{2}where P{2}N is the set of all ways to partition the index set {1,2, ... ,N} into N / 2 unordered pairsthat each index only appears in one pair. If C were the adjacency matrix of a graphG, the set P{2}N contains the edge sets of all possible perfect matchings on G. Therefore Haf(C)counts all matchings of G, a #P-complete problem.
[0073] Therefore, if σ is interpreted as the covariance matrix of the output of an N-mode GBS,the GBS sample n can be drawn from a p.m.f. that is proportional to Haf2(An) per Eq. (3), whichin turn is a function n and the adjacency matrix A of the graph G encoded into the GBS.
[0074] Any pure N-mode Gaussian state (described by a covariance matrix σ) can be obtainedby a GBS, i.e., a linear-optical (passive) N-mode unitary U acting on the N-mode tensor-productquantum state |0;r1^|0;r2^ ... |0;rN^, where |0;ri^ is a squeezed vacuum state with squeezingparameter ri in decibels (dBs) is 10log 2r10(e ). To obtain U , andr1, ... ,rN , that would result in the state σ, a Takagi A can be used:A=Udiag(λ , .. T1 ..,λM)U , (5)where the squeezing parameters are given by r −1i = tanh (cλi),1 ≤ i ≤ N. The U is the unitarythat characterizes the interferometer, which in turn can be used to calculate the N(N −1)single-mode phases in a universal programmable linear optical interferometer comprisingN(N −1) / 2 Mach-Zehnder interferometers (MZIs). Setting the scaling constant c close to 1 / smaxcould result in unrealistically-high single-mode squeezing values. One can clamp the maximumper-mode squeezing to x dB, i.e., a maximum squeezing parameter r per mode, withx = 10log 2r10(e ), by picking(6)
[0075] A graph G′ is a vertex minor of a graph G if there exists a sequence of localcomplementations and vertex deletions, applied to G, which yields G′. Local complementation ona vertex v of a graph G (denoted by τv(G)) is a graph operation that complements the edgesbetween the vertices that are in the neighbourhood of v. Deletion of a vertex v of graph G(denoted by χv(G)) is another graph operation that deletes the vertex v and all its adjacent edges inthe graph G. A vertex minor G′ of G can be obtained by having vertices (v1, ....,v j) of G undergolocal complementations, i.e., H = τv j ◦ τv j−1...◦ τv1(G), followed by vertex deletions on vertices(u1, ....,uk) of H, leading to: G′ = χuk ◦χuk−1...◦χu1(H), where k and j are arbitrary integers.
[0076] Both of the abovegraph G is associated with a quantumgraph state (or cluster state)—correspond to local single-qubit operations. Localcomplementation is a single-qubit Clifford unitary, and vertex deletion can be realized by a PauliZ measurement on a vertex. Therefore, a sequence of local complementations and vertexdeletions correspond to a sequence of single-qubit Clifford unitaries and measurements. Thedecision problem to determine if G′ is a vertex minor of G is an important problem underlyingClifford manipulation of Stabilizer quantum states, but is known to be nondeterministicpolynomial-time complete (NP-Complete). This is one motivation behind our study of thisproblem as a candidate for quantum acceleration by a NISQ processor.
[0077] To compare a quantum-assisted algorithm with the classical algorithms, severalclassical algorithms for graph similarity were explored and implemented for the vertex-minorproblem. The accuracies and runtimes of the classical algorithms were compared relative to thequantum-assisted algorithm. Brief synopses and further details of the algorithms implemented areprovided herein.
[0078] 1. Graphlet kernel — Graphlet kernel is an algorithm that exploits a similarity measure forgraphs that captures local structural information. A graphlet kernel algorithm can computethe frequency of each possible graphlet (or small subgraph) within two graphs andcompares their distributions. The kernel value is high if the two graphs share similargraphlet distributions, and low otherwise.2. Shortest path kernel — This algorithm is based on a graph kernel that measures thesimilarity between two graphs based on their shortest-path distances. The shotest pathkernel computes a kernel value that is high if the two graphs have similar shortest-pathdistance distributions and is low otherwise.3. Weisfeiler-Lehman Kernel — The Weisfeiler-Lehman kernel is a graph kernel thatcompares the structural similarity of two graphs by counting the number of commonsubstructures in their Weisfeiler-Lehman subtree patterns. The Weisfeiler-Lehman kerneliteratively labels the nodes of the graphs with the local patterns of their neighbours andaggregates these labels to form a global representation of the graph.4. Image classification on ordered adjacency matrices — This is a method for classifyinggraphs based on their adjacency matrices. Image classification on ordered adjacencymatrices involves comparing the similarity between the matrices by first ordering the nodesand adjusting their size, then converting them into images and using a neural network toclassify them. In essence, this method does image classification onappropriately-conditioned adjacency matrices as input ‘images.’5. Spectral method—This method is a quantum-assisted algorithm that is proposed andevaluated. This method can perform much better than all the above four algorithms forsmall instances of the vertex minor problem. This method calculates the eigenvalues of theLaplacian matrices of the two graphs and concatenates them to construct a feature vector.An SVM classifier is utilized thereafter.
[0079] To train the SVM back end, both for the classical (spectral) method as well as the GBSmethod, a data set comprising 400 vertex-minor graph pairs and 400 non-vertex-minor graphpairs was generated by brute force. In some examples, an exact algorithm (i.e., zero misseddetections or false alarms) is used to determine if two graphs are related to each other by a seriesof local complementations and vertex deletions. Since the vertex minor is an NP-Completeproblem, we do not have a so-called “efficient” algorithm and hence the algorithm we undertake,described below, has runtime that grows exponentially in N, the number of nodes in the larger ofthe two graphs presented to the classifier.
[0080] Let G and G′ be the two graphs presented. Without loss of generality, let us assume Ghas the larger (or equal) number of nodes of the two, which we will call the parent graph, and G′the child graph. We then calculate z, the difference in the number of nodes between the twographs. We then make G go through all possible sequences of local complementations (LCs). Acrucial point to note here is that in local complementation, we do not repeat vertices sinceτv ◦ τv(G) = G. Hence, the total number of such sequences is 2N . For the graph produced by eachLC sequence, we delete all possible combinations of z nodes. The total number of ways deleting zdistinct vertices (without paying heed to any graph symmetries) is 2z. Finally, we check if thegraph produced for each LC sequence and each vertex deletion set is isomorphic to G′ or not. Wesubtract the adjacency matrices of the graph obtained and G′ (modulo all possible nodepermutations). If we obtain an all-zero matrix for any of those trials, we declare that G′ is avertex-minor of G. Otherwise, G′ is not a vertex-minor of G. The time complexity of thisalgorithm can be written:O(N ×22N ×2−N′ )(7)where N′ = |G′| is the number of nodes in G′. If N′ = N, there is a polynomial-time algorithmwhich checks if G and G′ are vertex minors or not. This is because the problem of deciding if twographs are LC-equivalent admits a polynomial-time algorithm. Eq. 7 clearly shows that ouralgorithm is suboptimal, since its time scaling does not become poly(N) when N′ = N. But, sincewe used the brute force algorithm only to generate our training data (once) for N ranging from 6to 12, we did not look for a more efficient algorithm.
[0081] For our classical algorithm’s feature vector, we use the spectral method. Let us picktwo graphs G (parent graph) and Gprime (child graph) from our training data set such that Gprime isa vertex minor of G. Let the Laplacian matrices be L and L′ and the corresponding eigenvaluearrays of these Laplacian be e and e′, respectively. Since the number of vertices of G is greaterthan or equal to that of G′ and we want to keep the length of the feature vectors constant for ourML algorithms, e′ can be padded with 0s until the length of e′ matches that of e. Both of thesearrays can be concatenated and a label can be added at the end called vertex-minor. For featurevectors of graphs where one of them is not a vertex minor of the other, the brute force method canbe used to first check if the graphs are vertex minor of each other or not.
[0082] Let the graph that is not a vertex minor of G be Gnvm. A Gnvm whose number of verticesis less than or equal to G but not higher can be chosen. We then check if Gnvm is a vertex minor ofG using our brute-force method. If it is not a vertex minor, we construct its Laplacian matrix Lnvmand the corresponding array of eigenvalues called envm. Since we want to keep the length of thefeature vector the same, we pad envm with 0 until the dimension of envm matches the dimension ofe. We then concatenate both of these arrays and add a label at the end called “not vertex-minor.”
[0083] In some examples, for a quantum-assisted algorithm’s feature vector, efficient samplingfrom the output distribution of a Gaussian boson sampling device can be performed. The methodcan involve introducing conditioning auxiliary variables, using virtual measurements, anditeratively replacing heterodyne outcomes with photon number measurements. This process cancreate a chain of conditional probabilities, where only the probabilities of a pure Gaussian stateneed to be calculated at each step. This technique simplifies the sampling process, allowing forfaster and more accurate computation. The same technique can be done for the Gvm graph andGnvm graph and the samples can be labeled as Snm and Snvm respectively. The arrays Svm and Snvmcan be padded with 0s until the length of the array is equal to that of S so that the length of thefeature vector remains constant. The arrays S and Svm can then be concatenated and labeled as“vertex-minor.” The arrays S and Snvm can also be concatenated and labeled as “notvertex-minor”. This feature vector is the same for a quantum-assisted algorithm simulated on aclassical computer and that using a Gaussian Boson Sampler.
[0084] In some examples, in order to increase the accuracy of a machine learning algorithm toa desired level of 99%, the following method can be employed. It involves running the simulationmultiple times (denoted by the variable n that may preferably be an odd number) and selecting themajority result as the final output, and may use the assumption that each trial is independent. Forinstance, if a binary classifier outputs 0 and 1 with probabilities 0.25 and 0.75 respectively, andthe desired output is 1, we can run the simulation three times and select outputs that have amajority of 1. By doing so, the probability of obtaining the correct output increases to 0.84375.This technique can be generalized such that if the probability of error in a single trial is e, then theprobability of error after running the classifier n times can be expressed as a function of e and n as:⌊ n2 ⌋( ) P(n,enn−k kerror ) =∑e (1− e) (8)kWe know we can do a normalgiven as:() n en−k(1− e)k =√1e−(k−µ)2 / (2σ2)(9) k σ 2πwhere µ = n(1− e) and σ2pretty low and the relation of n with e in that region. Let δ be the error. Then the number of stepsrequired for that is given by the formula:n(e) = c(δ) / ε2 (10)where ε = (0.5− e) if δ = 0.01 then the number of steps required for that is given by the formula:n(e) = 0.6 / ε2 (11)
[0085] A derivation of the above equation is disclosed herein. Now, ε is dependent on N, thenumber of vertices in a graph. Since we are dealing with a small number of vertices it may bechallenging to conclude anything about the dependence of error on N. Let E(N) be the one-shoterror corresponding to N for a particular algorithm. However, the E(N) function differs fordifferent algorithms used. We denote Ec(N) as the one-shot error corresponding to N for theclassical algorithm and Eq(N) as the one-shot error corresponding to N for the quantum-assistedalgorithm. Hence:n(N) =0.6 2 (12) (0.5−E(N))The above equation has a lower bound of O(eN). This means that if y(N) = (0.5−E(N))2 theny(N) scales a O(e−N) for large values of N and E(N) scales as O(e−N / 2). However, due topositive or negative correlations among trials, it may not be true that the scaling happens as percalculated in Eq.12 (as shown, for example, in FIGS. 4 and 5).
[0086] In order to determine the duration of the GBS-based quantum-assisted algorithm, weutilize an N-mode multiport interferometer design. Next, we aim to estimate the time that isrequired for a single photon to travel through one of the modes present in the interferometer. Thenumber of beam splitters in a single mode is equivalent to the total number of modes, which isdependent on the number of vertices (N). In this example, each beam splitter has a length of 10µm. The velocity of light is approximately 3×108 m / s, and the refractive index of silicon in thewaveguides is 1.44. Therefore, the time it takes for a photon to traverse one mode (in seconds), ofan N-mode multiport interferometer can be calculated as follows:1.44×N ×10−13to= (13)
[0087] The time taken is of the order O(N).rate of the pulses generated by a mode-locked fibre laser. To produce a squeezed vacuum, amode-locked fibre laser may be used that generates pulses at a repetition rate (R), followed by itsfrequency being doubled by a periodically-poled lithium niobate crystal followed by pumping aperiodically-poled potassium titanyl phosphate (ppKTP) waveguide which produces degeneratetwo-mode squeezed vacuum through type-II spontaneous parametric down-conversion. The timetaken for the next pulse to enter the device is 1 / R. Interestingly, this value does not depend on N.Time because of repetition rate (in seconds) is given by:tr = 1 / R (14)
[0088] Three of the main sources of loss in a GBS are due to coupling efficiency, on-chiplosses and detector transmissivity. Let the transmissivity by coupling efficiency be given by ηc,transmissivity by on-chip loss be given by ηo and transmissivity by the detector efficiency begiven by ηd . The resulting covariance matrix after loss is:σ Nloss = ηcηN oηdσpure +(1−ηcηo ηd)I (15)The entire derivation of the
[0089] Herein we disclose a randomized classical algorithm to increase the accuracy of ourML classification using repeated trials. Given the graphs G1 and G2, where G2 has a lesser orequal number of nodes than G1, we construct another graph G3 which we obtain by performinglocal complementation on a random list of vertices of graph G1 hence making it locally equivalentto G1. Whether G2 is a vertex-minor of G1 or not, it would share the same relation with G3. Foreach trial, we calculate a new G3, construct its classical feature vector and compare it with thefeature vector of G2. We can repeat the trials and take the maximum vote until, for example, theaccuracy starts to plateau at a fixed accuracy.
[0090] Herein we disclose a quantum-assisted algorithm to increase the accuracy of our MLclassification using repeated trials, Given the graphs G1 and G2, where G2 has a lesser or equalnumber of nodes than G1, we calculate the sample array for both, concatenate them and give it asan input in our ML algorithm. We repeat our sampling algorithm n(N) a number of times until wereach an accuracy close to 99% and conclude our classification. As shown in FIG. 5, because ofcorrelations, it is not exactly functioning like the Bernoulli approximation. However, it is prettyclose.
[0091] Finally, we calculate the time complexity each of our algorithms takes to solve theproblem.
[0092] 1) Classical-spectral algorithm- An example first step of our classical algorithm is finda locally equivalent graph of the larger graph, construct its Laplacian Matrix and find itscorresponding eigenvalues followed by the method of repeated trials to increase our accuracy.The time complexity of deriving a random local equivalent graph is O(N2). The time complexityto construct a Laplacian of an N node graph is O(N2 +M) where M is the number of edges in thegraph. Following this, the time complexity of calculating the eigenvalues of a N ×N matrix isO(N3) since we are using the QR decomposition for our algorithm. The time complexity ofpredicting data using a linear SVM is O(N). The complexity of repeated trials has already beenderived in the above section and we use nc(N) as the corresponding length of repeated trials to geta 97% accuracy for the given Ec(N). Hence, our final time-complexity for our classical algorithmis given as:Tc = nc(N)× (O(N3)+O(N2)+O(N)) (16)
[0093] 2) Classical realisation of the GBS- An example first step of our quantum-assistedalgorithm is to embed our two graphs in the GBS and generate samples. The sampling algorithmwe use has a time complexity of O(mN32N / 2) where m is the number of modes and N is thenumber of photons which is proportional to the N ×N matrix embedded in the GBS, which for usis the adjacency matrix and hence, N is the number of vertices. However, the algorithm samplesfrom the covariance matrix, not the adjacency matrix, so we convert it to its following covariancematrix using Eq (8). This has a time complexity of O(N3). For this example device, the number ofmodes is equal to the number of vertices in the graph hence the time complexity of the samplingalgorithm becomes O(N42N / 2) . We then concatenate these two samples and pass them throughthe linear SVM with a time complexity of O(N). This is then followed by the method of repeatedtrials to increase our accuracy. The complexity of repeated trials has already been derived in theabove section and we use nc(N) as the corresponding length of repeated trials to get a 97%accuracy for the given Ec(N). Hence, our final time-complexity for our algorithm is given as:T 3( ) cgbs = O(N )+nq(N)×O(N42N / 2)+O(N) (17)
[0094] 3) Hybrid quantum-classical realisation- An example first step of our quantum-assistedalgorithm is to embed our two graphs in the GBS and generate samples. Embedding the graph inour GBS can be done by doing the Takagi decomposition of the adjacency matrix to give us thesqueezing parameters and the unitary. The time complexity for the Takagi decomposition isO(N2) . Here, our sampling algorithm is described in the optical propagation section and hence ithas a time complexity of O(N) where N is the number of vertices in the graph. We thenconcatenate these two samples and pass them through the linear SVM which has a timecomplexity of O(N). This may then be followed by repeated trials to increase our accuracy. Thecomplexity of repeated trials has already been derived in the above section. Hence, our finaltime-complexity for our algorithm is given as:Tqgbs = O(N2)+nq(N)×( ) (18)
[0095] Here we show that the classical algorithm based on the spectra of graphs performsbetter (in accuracy and time) than the other graph classification method. Secondly, we see that ourlossy GBS in reasonable limits of loss and squeezing parameter performs better than the classicalspectra method. The details of the classical system where we ran all our simulations are disclosedherein.
[0096] Before we generate prophetic plots, we make our dataset of graphs. For each number ofthe node of the graph, we generate 500 pairs of “vertex-minor” graphs and 500 pairs of “notvertex-minor” graphs using the brute force method. Then for each pair, we calculate the featurevector for all the algorithms. For Graphlet, Shortest path and WL Kernel, we use a C-supportedSupport Vector Machine (SVM) by passing through the precomputed kernels. We use 25% of ourdataset for test data and the remaining for training our SVM. The C-parameter of the SVMcontrols the penalty for misclassifying training examples and we use a C value in the range 5 to10. The best model is then used to get the accuracy of the test set. The hyperparameters used forNetwork Classification are disclosed herein. For the spectral method, we use a linear SVM. Weuse 25% of our dataset for test data and the remaining for training our SVM. The C-parameter ofthe SVM is in the range of 5 to 10. The best model is then used to get the accuracy of the test set.
[0097] FIG. 6 shows prophetic plots of error, as a function of nodes in the graph, for variousclassical algorithms. The accuracy is averaged over 800 datasets and the error plotted here is(1-accuracy). We can conclude that the spectral method performs better than the other graphsimilarity algorithms. The slope of our method is less than other algorithms, which may help forhigh node graphs.
[0098] FIG. 7 shows prophetic plots of time taken (i.e., execution time), as a function of nodesin the graph, for various quantum and classical algorithms. The photon loss value used here isaround 0.24 (1.2 dB). The lossy GBS at a maximum 5 dB squeezing per mode performs betterthan the classical algorithm with a magnitude of 103 in time. Time vs N for classical vs allpossible quantum simulations. We observe the increment in improvement in performance due tohigher squeezing per mode. We observe that the time taken by the classical realisation of the GBSis indeed exponential as compared to our hybrid-classical and spectral methods. We observe thatthe slope for our GBS algorithms is less than the slope of our classical algorithm, suggestingbetter performance with larger-sized graphs.
[0099] After plotting the error vs the nodes in the graph (as shown in FIG. 6), the error being(100-accuracy of the test dataset) for each classical algorithm, we see that our spectral methodcan be an optimal working classical algorithm. Hence we can conclude that for our problem, onepossible approximate classical algorithm is the spectral method.
[0100] FIG. 4 shows a prophetic plot of accuracy, as a function of sample length, for aclassical algorithm. The accuracy is computed by averaging over 7 graph pairs for each particularn(N). In a scenario where each trial is independent of the others, the accuracy would follow theBernoulli approximation, resulting in increased accuracy with higher values of n(N). However,due to the negative correlations present in the algorithm, it does not perform optimally andapproaches a near-perfect accuracy but falls short of reaching 100%.
[0101] FIG. 5 shows a prophetic plot of accuracy, as a function of sample length, for aquantum-assisted algorithm. The accuracy is computed by averaging over 7 graph pairs for eachparticular n(N). In a scenario where each trial is independent of the others, the accuracy wouldfollow the Bernoulli approximation, resulting in increased accuracy with higher values of n(N).However, we observe that there are some positive correlations and then the Bernoulli takes overdue to the negative correlations present in the algorithm and approaches a near-perfect accuracybut falls short of reaching 100%.
[0102] Here we demonstrate that on fair assumptions of losses, maximum squeezing per modeand repetition rate of the mode-locked fiber laser, our quantum-assisted algorithm performs up to103 times better than the classical algorithm for the same set of graph pairs (as shown in FIG. 7).For our GBS simulation, we assumed the loss to be 80% coupling efficiency, 95 % detectorefficiency and 0.25 dB / cm. Maximum squeezing per mode is 5 dB and the mode-locked fiberlaser is assumed to have a repetition rate of 10 MHz. The feature vector for both classical andquantum-assisted algorithms is generated exactly as it was mentioned in the previous sections.
[0103] An explanation for this might be that Takagi decomposition, which is the dominatingterm of the GBS simulation, is faster than calculating the eigenvalues of the Laplacian of thegraphs, hence performing better even though the GBS requires more trials than a classicalalgorithm. We multiply n(N) with the time complexity of the QR decomposition in the classicalalgorithm whereas in quantum, we multiply it with an O(N) factor and not the Takagi which is thedominating factor, hence giving us time advantage. The optical propagation is in the order of10−7 hence not contributing much to the overall time calculation. Another interesting thing tonotice is that the slope of the GBS’s time vs N is less than the slope of the classical algorithm’stime vs N hence if we increase the number of nodes, it is more likely that the GBS tends toperform better than classical.
[0104] As disclosed herein, we describe a hybrid quantum-classical approach using aGaussian Boson Sampler (GBS) to the NP-Complete decision problem of determining if twographs are a vertex-minor pair (or not). The premise behind our approach is the assertion that bymapping a graph’s adjacency matrix into the quantum state of a GBS, theclassically-hard-to-sample photon-click-pattern samples produced by the GBS powerfully encodefeatures of the graph embedded in it that assists with the classification. We design a graphembedding that lets us pick a lower squeezing amount at the GBS input (a hard-to-producequantum optical resource) at the expense of a lower one-shot classification accuracy, which inturn results in a larger number of repeated trials to obtain a desired (high) target accuracy. We alsopropose a classical spectral algorithm for the vertex minor problem that uses the graphLaplacians’ eigenvalues and a support vector machine (SVM) classifier, which we showoutperforms various state-of-the-art classical graph classification algorithms. We show that with areadily-realizable on-chip GBS, our hybrid special-purpose processor can outperform theaforesaid classical algorithm when executed on the most powerful MacBook Pro.
[0105] Other implementations can includes a variety of other features, including the possiblybetter embedding of this problem to a GBS (or another NISQ quantum processor), findingpossibly better classical algorithms, or using a different machine learning (ML) classifierback-end that might give higher accuracy. Other NISQ processors can also be used for variouspractical problems in optimization, search, estimation and classification (e.g., of chemicalvibronic spectra), and other graph problems, and performance comparisons can be performedunder realistic assumptions with the respective classical algorithmic approaches.
[0106] Without intending to be bound by theory, the following is an example of a theoreticalmodel for illustrating features. Here we focus on deriving how our number of trials n is related tothe accuracy of our ML algorithms. Suppose the probability of error in a single trial is e, then theprobability of error after running the classifier n times can be expressed as a function of e and n as:⌊ n2 ⌋( ) nkWe know we can do a normal distribution approximation to the Perror which is a binomial. It isgiven as:() nn−k k = 1−(k−µ)2 / (2σ2)where µ = n(1− e) and σ2.pretty low and the relation of n with e in that region. Let δ be the error. Hence:⌊ n2 ⌋( ) δ(n,e) =n ∑en−k(1− e)k (21)kSubstituting y = x / n, we∫1n⌊n 2⌋2 2δ(n,e) =n √ e−(ny−µ) / (2σ ) dy (23)which is the samecalculations more accessible, let us divide the first integral into 2 parts.δ(n,e) =1∫ µ√ e−(x−µ)2 / (2σ2)dxWe know that the first 2. y = x−µand swapping the limits, we1 1∫ (⌊ n2⌋−µ) 2 / ( 2δ =+ √−(y)2σ )Now we introduce the definition of erf(z) function, given by:er f =2∫ z√ dtNow we can write the integral δ in terms of er f (z):δ(n,e) = 1− er f (k)(28) 2where k is given by:(1− n√ e)n−⌊2⌋k(n,e) =(29) 2ne(1− e)We know that erfc(z)=1-erf(z),erfc(k) = 2δ (30)− 2 ek√= 2δ (31)Taking logs on both sides:−k2 − lnk−0.5724 = ln2δ (32)We can see that:k = c(δ) (33)Hence, using Eq.29 ,n = c(δ) / ε2 (34)where ε = (0.5− e). Using Newton-Raphson method and substituting δ = 0.01 in Eq.32, we get:k(n,e) = 1.09629 (35)Substituting k(n,e) back in the above equation, we get:n(e) = 0.6 / ε2 (36)Now, ε depends on N, the number of vertices in a graph. Hence the n depends on N.
[0107] We are assuming G and G′ as our two graphs of comparison where the number ofnodes of G′ are less than or equal to G. The graphlet sampling kernel is a popularmachine-learning-based method for graph classification and clustering tasks. In this method, ourgraphs are decomposed into graphlets (small induced subgraphs with k nodes wherek ∈ {3,4,5..}.
[0108] FIG. 8 shows all k ∈ {3,4,5}-node connected non-isomorphic graphlets that can beused for a graphlet sampling kernel algorithm.
[0109] The feature vector of each graph is then calculated in the following way: LetS = {graphlet1,graphlet2...graphletr} be a set of k-size graphlets. We then define a vector fG oflength r such that:fG,i = #(graphleti ⊑ G) (37)We even normalise the counts to frequency vectors because we mostly deal with graphs ofdifferent sizes.1 NG= #allfG(38)We then calculate the kernel given by thekGS(G,G′) = NTGNG′ (39)
[0110] We then use this precomputed kernel and pass it through our SVM classifier labellingit “vertex-minor” or “not vertex-minor” depending on the pair of graphs chosen. For our problem,we choose all possible non-isomorphic graphlets of size k ∈ {3,4,5} giving us a total of 29graphlets (shown in FIG. 8). We then calculate the normalised feature vector and flattened kernelmatrix for G and Gvm and label it “vertex minor” and do the same for G and Gnvm and label it “notvertex minor”. Each row of our data set is 29 × 29 = 841.
[0111] The Weisfeiler-Lehman algorithm (e.g., FIG. 9) is a technique used to label eachvertex of a graph with a multiset label. This label consists of the original vertex label and thesorted set of labels of its neighbouring vertices. The multiset label is then compressed into a new,shorter label. This relabeling process is repeated for several iterations. It is important to note thatthis procedure is applied simultaneously to all input graphs. If two vertices from different graphshave identical multiset labels, they will receive the same new label.
[0112] FIG. 9 shows an example implementation of the Weisfeiler-Lehman algorithm ongraphs 902A, 902B. The vertex labels of graphs 902A, 902B are updated by concatenating thevertex number with the adjacent vertices to generate graphs 904A, 904B. The label is againupdated by a pre-defined hash table 906 to generate graphs 908A, 908B. In the end, the vertex(defined in the hash table) frequencies of graphs 908A, 908B become a feature vector ψ(G),ψ(G′) and the kernel K(G,G′) is defined as an inner product of the two vectors.
[0113] The shortest-path kernel is a graph analysis technique that involves breaking downcomplex graphs into their constituent shortest paths and then comparing these paths based ontheir lengths and the labels of their start and end points. Firstly, we would want to compute theall-pairs-shortest path for G and G′ using the Floyd-Warshall algorithm. The Floyd-Warshallalgorithm is a dynamic programming algorithm for finding the shortest path between all pairs ofvertices in a graph. It achieves this by computing a matrix of shortest path distances, where eachentry in the matrix represents the shortest path distance between two vertices. The shortest-pathkernel is then given by:k′SP(G,G ) =k(d(vi,v j),d(vk′,vl′)) (40)where d(u,v) gives us the shortest path between nodes u and v and k is the base kernel. For ourproblem, we used the linear kernel k(d(vi,v j),d(vk′,vl′)) = d(vi,v j)∗d(vk′,vl′). This is aprecomputed kernel so we pass itminor” or ”notvertex-minor” depending on the pair of graphs chosen.
[0114] FIG. 10 shows example network classification using adjacency matrix embeddings forour vertex minor graphs. As we can see from the image transformation of the graphs, thevertex-minor pair looks similar compared to the not vertex-minor pair hence helping the neuralnetwork to predict the correct label.
[0115] Formally, we assign the same label, l = l0 to all the vertices of our graphs G and G′. Inthe next iteration, each label is updated with a sequence of its own label followed by labels fromthe neighbouring vertices. We rename this sequence of labels for each vertex with a new labell1 = h(l0), from a defined hash table, h and repeat this process for a certain number of iterations.The relabeled graphs can be called as Gi where i is the number of iterations G and then we havean array of graphs {G0,G1,G2....Gh} where h is the total number of iterations. We then define ourkernel as:kWL(G,G′) = k(G0,G′0)+ k(G1,G′ 1)+ ...+ k(Gh,G′ h) (41)where k is our basedefined in the hash table, make a vector for each graph and then take an inner product of the twovectors. This is a precomputed kernel so we pass it through our SVM classifier labelling it“vertex-minor” or “not vertex-minor” depending on the pair of graphs chosen. An example of thiskernel calculation is given in FIG. 9.
[0116] Network classification using adjacency matrix embeddings and deep learning is analgorithm where the adjacency matrix of the graph is embedded into an image and we performimage classification using a neural network. The procedure for embedding the adjacency matrixof a graph into an image for the neural network is given below:
[0117] 1. Ordering of Nodes- To make the adjacency matrix informative, an ordering algorithm isused. This algorithm starts by selecting the node with the highest degree, and if there is atie, it chooses the node with the largest k-neighbourhood for increasing values of k. Thealgorithm then proceeds to order the rest of the nodes by their shortest path distances to thealready ordered nodes. If a tie exists, it prioritizes the node with the highest degree or picksrandomly. This ordering algorithm prioritizes nodes that are well-connected and closer tothe previously ordered nodes, while also considering the degree and size of each node’sneighbourhood. 2. Rescale Algorithm- Since we are working with graphs of different sizes, we need to rescaleall the graphs to one particular dimension of the image. There are two ways of doing that,image resizing and padding. In image resizing, the “pixels” of the matrix are mappedlinearly to the corresponding pixels in the final resized image. This mapping ensures thatthe information in the adjacency matrix is preserved even after resizing, allowing foraccurate analysis and comparison of graphs of different sizes. In padding, zeros are addedas padding to the matrix, with the original matrix remaining in the top-left corner. Thispadding ensures that the matrix is resized uniformly without changing the relative positionsof the nodes or altering the overall connectivity of the graph. For resizing, since the sizes ofthe graphs vary from 3 to the number of nodes of the parent graph, we take the leastcommon multiple. However, that number would be too big and it would be time-consumingto map all the adjacency matrices to that particular size linearly. Hence for our algorithm,we stick to padding.
[0118] After ordering the nodes and padding to a particular size, we flatten the pixel matricesand concatenate the flattened arrays of the pair of graphs depending if they are vertex minor ornot. We then pass these arrays into our neural network with three layers. The first layer has 64neurons and uses the rectified linear unit (ReLU) activation function. The second layer has 240neurons and also uses the ReLU activation function. The third layer has 20 neurons and uses therectified linear unit (ReLU) activation function. The final layer has a number of neurons equal tothe number of classes in the dataset and uses the softmax activation function, which normalizesthe output of the network to represent the probabilities of each class. The autoencoder uses alearning rate of 0.001 during pre-training and a learning rate of 0.1 during fine-tuning. Todetermine when to stop the training process, 10% of the training data is used as a validation set.Pre-training uses a maximum of 10 epochs, while fine-tuning uses a maximum of 400 epochs.Both pre-training and fine-tuning use cross-entropy cost. As we can see from FIG. 10, thevertex-minor pair of graphs look similar to the not-vertex minor pair.
[0119] Here we focus on deriving the loss function for our GBS that undergoes photon lossfrom coupling efficiency, on-chip losses and detector efficiency. Losses can be added to thedevice with the help of a vacuum state coupled with the help of a beam splitter whosetransmissivity is the transmissivity of the photon loss. It can be proven that if there is an equalloss on the output modes of an MZ interferometer, it is equivalent to having the same equal loss inthe input modes as shown in FIG. 11A. With the help of this, we can push back all the loss termsfrom the on-chip and detector efficiency to the input modes and get a system where all the lossesare in the input as shown in FIG. 11B.
[0120] FIGS. 11A and 11B show coupling the loss terms together in the input mode of themultimode interferometer. FIG. 11A depicts a MZ interferometer 1102A with losses 1104A,1104B in the output modes. In some implementations, the interferometer 1102A can beequivalent to the MZ interferometer 1102B having losses 1106A, 1106B in the input modes if thelosses 1104A, 1104B are equal to 1106A, 1106B. FIG. 11B shows an example of how the lossterms can move move back in an interferometer 1110A. Each step 1110A1-1110A4 depictsequivalent representations of interferometer 1110A wherein the losses, which are represented bydashes, are moved from output ports at each step to input ports. Each loss represented by a dash issummed with another loss represented by a dash. For instance, a double dash signifies twice theloss because the length of the fiber there is double that of the single dash.
[0121] Without intending to be bound by theory, the following is an example of a theoreticalmodel for illustrating features. Given G1 and G2, where G2 is a smaller graph and not a vertexminor of G1. Let G3 be a graph which is locally equivalent to G1, that is, we obtained G3 by aseries of local complementations on G1 denoted like this:G3 = τvi ◦ τvi−1...◦ τv1(G1) (42)where k is an arbitrary integer. We prove that “G2 is not a vertex-minor of G3” iff “G2 is not avertex-minor of G1”. To prove this, we need to prove the following two premises:
[0122] Premise A: If G2 is not a vertex-minor of G3 then G2 is not a vertex-minor of G1.
[0123] We use proof by contradiction to prove this. Let G2 be a vertex-minor of G1. Then,G2 = χuk ◦χuk−1...◦χu1 ◦ τv j ◦ τv j−1...◦ τv1(G1) (43)Since local we canG1 = τv1 ◦ τv2...◦ τvi(G3) (44)Substituting Eq.44 in Eq.43, we get,G2 = χuk ◦χuk−1...◦χu1 ◦ τv j ◦ τv j−1...◦ τv1(45)
[0124] Hence fromcontradiction. Hence, G2 can never be a vertex minor of G1.
[0125] Premise B: If G2 is not a vertex-minor of G1 then G2 is not a vertex-minor of G3.
[0126] We use proof by contradiction to prove this. Let G2 be a vertex-minor of G3. Then,G2 = χuk ◦χuk−1...◦χu1 ◦ τv j ◦ τv j−1...◦ τv1(G3) (46)Substituting Eq.42 in Eq.46, we get:G2 = χuk ◦χuk−1...◦χu1 ◦ τv j ◦ τv j−1...◦ τv1
[0127] Hence from Eq.47, we can conclude that G2 is a vertex-minor of G1. But this is acontradiction. Hence, G2 can never be a vertex minor of G3. Hence we can conclude that “G2 isnot a vertex-minor of G3” iff “G2 is not a vertex-minor of G1”.
[0128] While the disclosure has been described in connection with certain embodiments, it isto be understood that the disclosure is not to be limited to the disclosed embodiments but, on thecontrary, is intended to cover various modifications and equivalent arrangements included withinthe scope of the appended claims, which scope is to be accorded the broadest interpretation so asto encompass all such modifications and equivalent structures as is permitted under the law.
Claims
. - , - , , - , - - - , - , - ,. ,. , - -. . , - - , --- . -. - -. , . , , . . - ,. . - , , - - ---. . , , , . , ,. , . - . - , , - - - - , , - - --- , -. . -, -. - ,. . . , - - , , - , - - - , - , - , . --. - , - , ,. , . - -. . -, -. . , , , , - , - , - ,. ,. - - , --- . -. , -. - . , , . ., . . . , -. , . ,- . , - - . --. , ,. , . , - . - --
Citation Information
Patent Citations
Parallel network flow classification method
CN104702465A
Systems And Methods For Training Matrix-Based Differentiable Programs
US20190354894A1
Apparatus and methods for gaussian boson sampling
US20220051124A1
Systems and methods for analog computing using a linear photonic processor
US20230353252A1
Apparatus and methods for generating non-gaussian states from gaussian states
US20240053615A1