Method and apparatus for distributed quantum computing

By partitioning optimization problems into factorizable Markov random fields and mapping them onto trainable quantum circuits, the method enhances distributed quantum computing's efficiency in solving large-scale telecommunications network optimization tasks, overcoming noise and dynamic network challenges.

WO2025216675A1PCT designated stage Publication Date: 2025-10-16TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/SE2024/050341
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-10
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

Existing distributed quantum computing systems face challenges in efficiently solving large-scale optimization problems in telecommunications networks due to noise, decoherence, and dynamic network conditions, which hinder achieving quantum supremacy, particularly in tasks like energy optimization and beamforming.

Method used

A method for distributed quantum computing that utilizes an inherent graph structure in system optimization by representing optimization problems as factorizable Markov random fields (MRF) and partitioning the graph structure into partitions, mapping these partitions onto trainable quantum circuits across multiple quantum processing units (QPUs), and training them sequentially or in parallel based on available resources and cost functions.

Benefits of technology

This approach significantly reduces computational complexity and improves training speed and accuracy in solving optimization problems, demonstrating superior performance compared to traditional quantum circuits like Ising Born machines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SE2024050341_16102025_PF_FP_ABST
    Figure SE2024050341_16102025_PF_FP_ABST
Patent Text Reader

Abstract

A computer-implemented method (100) for system optimization using distributed quantum computing performed by a network node The optimization problem can be represented as a factorizable Markov random field, MRF, model. The method comprises partitioning (S101) a graph structure of the MRF model into two or more partitions based on a cost function of the optimization problem. The method comprises mapping (S105) the optimization problem onto a plurality of trainable quantum circuits, each trainable quantum circuit corresponding to a partition of the two or more partitions. The method comprises initiating training (S107) each of the trainable quantum circuits on a separate quantum processing unit of the distributed quantum computing device. Further disclosed are a related network node, quantum computing device, computer program, and computer program product.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHOD AND APPARATUS FOR DISTRIBUTED QUANTUM COMPUTING

[0002] TECHNICAL FIELD

[0003] The disclosure relates generally to methods for quantum computing, and more specifically relates to methods for distributed quantum processing units. Further disclosed are a related network node, a system comprising quantum processing circuitry, a computer program, and a computer program product.

[0004] BACKGROUND

[0005] A quantum computer is a computer which exploits quantum mechanical phenomena, for example by leveraging quantum superposition and quantum entanglement. A quantum computer typically comprises hardware capable of preparing particles in specified quantum states (e. g. qubits, or more generally qudits), a circuit of quantum gates which manipulate the quantum states, and measurement circuitry to measure output of the circuit and obtain a final result of a calculation. Theoretical results suggest a potential for quantum computers to be much faster at certain types of computations than a classical computer, referred to as quantum supremacy. In particular, there may be an advantage for quantum computers when performing certain types of calculation-heavy machine learning tasks, such as optimization tasks in communication networks. However, considerable obstacles remain before the theorized quantum supremacy can be achieved in practice.

[0006] The present state of quantum computing is sometimes referred to as the noisy intermediate-scale quantum, NISQ, era. The NISQ era is defined by intermediatesize quantum processors typically containing from 50 up to a few hundred qubits, with considerable noise and short decoherence limits, making calculations requiring many quantum gates or calculations requiring a large number of input qubits difficult or impossible.

[0007] One solution to mitigate the lack of large, fault-tolerant quantum computers is using a distributed quantum processing unit, wherein a large problem such as an optimization problem in a communication network is broken down into smaller components which are executed in parallel on a plurality of quantum processing units or sequentially using one or more quantum processing units. In the case of parallel execution / distributed quantum computing, the distributed quantum computers may take advantage of quantum entanglement to further improve the performance of the quantum circuit through quantum correlations.

[0008] A number of issues remain to be solved to improve distributed quantum computing systems, including engineering problems with communication between the quantum processing units and theoretical computer science problem with how to distribute a given problem across a plurality of quantum processing units in an efficient way.

[0009] Large scale optimization problems in telecommunication networks are a class of problems for which quantum computing may be useful. Optimization problems in telecommunications networks may refer to, for example, energy optimization problems, capacity / load balance problems, beamforming problems, or similar problems which arise in high capacity / high performance networks. These optimization problems may involve a large number of nodes and a large amount of data, which means that a computer tasked with learning a probability distribution representing some aspect of the optimization problem has to perform a considerable number of calculations. Moreover, a telecommunications network is often highly dynamic, as people move around carrying their communication devices, sometimes at high speeds such as in high-speed trains. Other telecommunications networks may comprise base stations in the form of unmanned aerial vehicles to supplement fixed base stations, further adding to the dynamic condition. Solutions which reduce exponential complexity in runtimes are therefore desirable. One class of quantum circuits which shows promise for exhibiting quantum supremacy is the class called Bom machines, and in particular a subclass defined by Ising Hamiltonians, as shown in “The Bom Supremacy: quantum advantage and training of an Ising Bom machine” by Brian Coyle, Daniel Mills, Vincent Danos, and Elham Kashefi, published in npj Quantum Information, 6(1 ):60, 2020.

[0010] The paper “Simple explanation of the no-free-lunch theorem and its implications" by Yu-Chi Ho and David L. Pepyne, Journal of Optimization Theory and Applications 115 (2002), pp 549-570 discusses the limitations of general-purpose optimization methods. In particular, the paper discusses the fact that any general-purpose optimization method will have the same expected outcome on a randomly selected optimization problem. Developing methods which are customized to the precise type of optimization problem is therefore of great interest.

[0011] SUMMARY

[0012] It is an object of the present disclosure to provide a computer-implemented method for efficient distributed quantum computing. In particular, it is an object of the present disclosure to provide a method for distributed quantum computing utilizing an inherent graph structure in system optimization.

[0013] According to a first aspect of the disclosure, there is a computer-implemented method for system optimization by distributed quantum computing, the computer- implemented method performed by a network node. The method comprises obtaining an optimization problem, the optimization problem representable as a factorizable Markov random field, MRF, model. The method further comprises partitioning a graph structure of the MRF model into two or more partitions based on a cost function of the optimization problem. The method further comprises mapping the optimization problem onto a plurality of trainable quantum circuits, each trainable quantum circuit corresponding to a partition of the two or more partitions. The method further comprises initiating training each of the trainable quantum circuits on a separate quantum processing unit of the distributed quantum computing device.

[0014] According to an embodiment of the first aspect, the MRF model is a binary MRF model.

[0015] According to an embodiment of the first aspect, the partitioning separates the partitions along edges of the MRF model with lowest total cost.

[0016] According to an embodiment of the first aspect, the cost function is based on a key performance indicator, KPI, of a telecommunications network.

[0017] According to an embodiment of the first aspect, the KPI is a power consumption metric and the MRF model is associated to a power consumption optimization problem.

[0018] According to an embodiment of the first aspect, a clique of order n of the graph partition is mapped onto a quantum circuit comprising 2n- 1 quantum gates, each quantum gate corresponding to a sub-clique of the clique. According to an embodiment of the first aspect, mapping the partitions of the optimization problem onto trainable quantum circuits depends on one or more of a size of a register of each of the quantum processing units of the distributed quantum computing device and an available set of quantum gates.

[0019] According to an embodiment of the first aspect, the method further comprises obtaining a plurality of graph partitionings, each graph partitioning associated with a cost relative to the cost function, and selecting the graph partitioning of the plurality of graph partitionings with lowest associated cost which can be mapped onto quantum circuits which can be implemented by available quantum computing resources.

[0020] According to an embodiment of the first aspect, selecting a graph partitioning comprises selecting a graph partitioning wherein the partitions are split along edges that minimize the cost along split edges.

[0021] According to an embodiment of the first aspect, mapping a graph partition onto a quantum circuit comprises selecting a unitary operator for state initialization for the quantum circuit encoding the input states of the quantum circuit with initial knowledge of a solution to the optimization problem.

[0022] According to an embodiment of the first aspect, vertices belonging to split edges are trained with the partition to which the cost associated to the vertices has the least impact.

[0023] According to an embodiment of the first aspect, when the number of quantum processing units is insufficient for the number of partitions in the selected partitioning, the method further comprises initiating training of each of the trainable quantum circuits sequentially on the available quantum processing units.

[0024] According to an embodiment of the first aspect, an order to train the trainable quantum circuits is selected based on the cost associated to each partition.

[0025] According to an embodiment of the first aspect, initiating training of a quantum circuit comprises initiating a training process associated to a machine learning algorithm on a quantum processing unit to obtain an approximation of optimal parameters for the optimization problem. According to a second aspect of the disclosure, there is a network node comprising a memory and processing circuitry, the network node configured for system optimization using distributed quantum computing. The network node is configured to obtain an optimization problem, wherein the optimization problem can be represented as a factorizable Markov random field, MRF, model. The network node is further configured to partition a graph structure of the MRF model into two or more partitions based on a cost function of the optimization problem. The network node is further configured to map the optimization problem onto a plurality of trainable quantum circuits, each trainable quantum circuit corresponding to a partition of the two or more partitions. The network node is further configured to initiate training each of the trainable quantum circuits on a separate quantum processing unit of the distributed quantum computing device.

[0026] According to embodiments of the second aspect, the network node is further configured to perform a method according to any embodiment of the first aspect.

[0027] According to an embodiment of the second aspect, the network node is a radio access node in a telecommunication system.

[0028] According to a third aspect of the disclosure, there is a quantum computing device comprising quantum processing circuitry, the quantum computing device configured to solve an optimization problem by implementing quantum circuits obtained according to any embodiment of the first aspect.

[0029] According to a fourth aspect of the disclosure, there is a computer program comprising computer readable instructions such that executing the program on processing circuitry of a network node causes the network node to perform a method according to any embodiment of the first aspect.

[0030] According to a fifth aspect of the disclosure, there is a computer program product comprising a computer readable storage medium on which a computer program according to the fourth aspect is stored.

[0031] BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Fig. 1 depicts a flowchart of a method according to embodiments presented herein. Fig. 2 depicts a communication system according to embodiments presented herein.

[0033] Fig. 3 depicts an apparatus according to embodiments presented herein.

[0034] Fig. 4 depicts a quantum computing device according to embodiments presented herein.

[0035] Fig. 5 depicts an example of a graph, associated factor graphs, and corresponding quantum circuits according to methods presented herein.

[0036] Fig. 6 depicts a quantum circuit according to an embodiment presented herein.

[0037] Fig. 7 depicts a partitioned graph according to an embodiment presented herein.

[0038] Fig. 8a is a scatterplot of the training speed of a machine learning algorithm according to an embodiment presented herein.

[0039] Fig. 8b is a line graph of a number of training steps needed to train a machine learning model according to an embodiment presented herein.

[0040] Fig. 8c is a line graph of a number of training steps needed to train a machine learning model according to an embodiment presented herein.

[0041] DETAILED DESCRIPTION OF THE DRAWINGS

[0042] Fig. 1 depict a method 100 according to presented embodiments. The method 100 is a computer-implemented method for system optimization by distributed quantum computing. A computer implementing the method 100 may be a classical computing device. For the purpose of the present disclosure, the term classical refers to computing units comprising a processor and a memory, such that the processor performs operations on bits. A classical computing unit may, in embodiments, comprise a virtual computing unit. The word classical in reference to methods, or more specifically algorithms, refers to methods and algorithms to be performed by a classical computing unit.

[0043] System optimization problems refer to a broad class of problems which may be modelled mathematically by a probability distribution, often referred to as a target distribution or target model, such that finding an optimal set of parameters for the target distribution corresponds to finding an optimal solution to an optimization problem. In some embodiments, solving the optimization problem comprises finding or approximating a local optimum of a cost function, whereas other embodiments may comprise finding or approximating a global optimum of the cost function. In some embodiments, the optimization problem comprises finding an optimal set of parameters for a mathematical model. In such embodiments, solving the optimization problem may correspond to a machine learning algorithm learning optimal parameters for a model. The learning process may be supervised or unsupervised. The optimization problem may correspond to obtaining parameters for a target distribution which arises in a communications network. In some embodiments, the distribution may model a resource allocation problem in a communications network. In embodiments presented herein, solving the optimization problem, or, equivalently, obtaining approximations to optimal parameters for the target distribution, is performed by a distributed quantum computer.

[0044] Methods presented herein may be performed in any communications system, such as an example of a communication system 200 of Fig. 2. The communication system of Fig. 2 comprises a host 201 in communication with a telecommunication network 202. The telecommunication network comprises a core network 206 comprising one or more core network nodes 208, and an access network 204 comprising one or more radio nodes 210A, 210B. The radio nodes 210A, 210B may communicate with one or more user equipments, UEs, 212A, 212B. Alternatively or in addition, the radio nodes 210A, 210B may communicate with one or more hubs 214, the hub communicating with one or more UEs 212C, 212D.

[0045] In particular, the method may be performed in a communication system such as a communication system implementing wireless local area networking protocols such as Wi-Fi 4, Wi-Fi 5, Wi-Fi 6, and Wi-Fi 7 as defined by the Institute of Electrical and Electronics engineers, or a wireless telecommunications network such as Universal Mobile Telecommunications System, UMTS, Long-Term Evolution, LTE, New Radio, NR, as defined by the 3rdGeneration Partnership Program, 3GPP, or a network according to any other or future standard as defined by a standardizing body. Alternatively or in addition, methods as presented herein may be implemented in a hybrid network incorporating aspects of several different types of networks. Networks implementing methods as presented herein may be private networks or public networks.

[0046] Embodiments of the method 100 may be performed by a network node. A network node in a communication network may comprise a physical or virtual unit or device of the network capable of creating, receiving, and / or transmitting information over a communication channel of the network. A physical network node may for example comprise a modem, a router, or a wireless access point in a Wi-Fi network, or it may comprise a radio access node such as an eNodeB or a gNB, or a core network node in a 3GPP system. A virtual network node may perform some or all the tasks of a physical node, but distributed over a series of client computing devices. In some embodiments of the method, the network node may be an Open Radio Access Network, O-RAN node. An O-RAN node is a network node that supports an O-RAN specification, such as a specification published by the O-RAN Alliance or a similar organization.

[0047] An example of a network node 300 is depicted in Fig. 3. The network node comprises processing circuitry 305 and a memory 304. The processing circuitry may comprise an optimization apparatus / device 324, or the network node may comprise the optimization apparatus with additional circuitry 326 and memory 328 adapted to carrying methods according to the disclosure.

[0048] In some embodiments, the network node of Fig. 3 comprises further optional features. These include a communication interface 306. The communication interface 306 comprising an antenna 310, a port / terminal 316, and radio front-end circuitry 318. The radio front-end circuitry 318 comprises a filter 320 and an amplifier 322. The network node 300 further comprises a power source 308 and the processing circuitry 305 comprises radio front-end circuitry 312 and baseband circuitry 314. The network node 300 further comprises an optimization apparatus / device 324, the optimization apparatus configured to perform methods as presented herein. The optimization apparatus comprises processing circuitry 326 and a memory 328. The memory may further comprise a computer program 330 comprising computer-readable code which, when executed by the processor, cause the optimization apparatus to perform methods as presented herein. The memory may further comprise a computer program product 332, comprising the computer program 330.

[0049] In embodiments, a network node performing methods as presented herein may have access to a quantum computing device 400 (Fig 4). The quantum computing device may comprise a single quantum processing unit, QPU, 402 or it may be a distributed quantum computing unit comprising a plurality of QPUs. The quantum processing unit 402 may comprise state preparation circuitry 404, quantum processing circuitry 406 capable of executing a selection of quantum gates, and measurement circuitry 408. In some embodiments, the quantum computing device may be comprised in the network node. In other embodiments, the network node performing the method may be separate from the network node comprising the quantum computing device. A QPU, as presented herein, comprises a state preparation circuit, a unitary circuit, and a measurement circuit. The unitary circuit may be adapted to implement a subset of possible quantum gates by acting on qubits or, more generally, on qudits. The state preparation circuit may be selected based on, for example, a coherence time of the QPU. Each quantum gate may act on two or more qubits or qudits.

[0050] Distributed quantum computing is a family of techniques which enable simulation of quantum computers comprising more qubits than available. Broadly speaking, distributed quantum computing comprises partitioning a theoretical quantum circuit into a plurality of partitions such that a minimal number of gates act on qubits from distinct partitions. Such partitions can be executed on distinct quantum processing units, QPUs. Gates which act on qubits from distinct partitions can be handled in several different ways. One alternative is to enforce a strict partitioning, such that no two qubits in distinct partitions are acted on by the same gate. Alternatively or in addition, two distinct QPUs may communicate through a non-quantum, classical, channel, where the classical channel is either one-way or two-way, so that the output of the quantum circuit in a first QPU simulating a first partition is sampled and communicated to a second QPU, where the result is encoded and used as input in a quantum circuit in the second QPU. Alternatively or in addition, two qubits intended for distinct QPUs may be entangled by the state preparation circuit. However, implementation involving communication between the QPUs and / or entanglement are exponentially more complex to implement and therefore add considerable overhead to the distributed quantum computing device. Implementations of distributed quantum computing with strictly local operations or which use one-way communication channels may be implemented by a single QPU by implementing each circuit? in sequence on the QPU. One-way communication would then correspond to measuring the output of a first circuit and using the output of the first circuit as input when preparing the second circuit.

[0051] Optimization problems arising from learning distributions of communication networks often involve problems which cannot be mapped to a NISQ device, and hence benefit from distributed quantum computing methods.

[0052] The optimization problem can be represented as learning the parameters of a factorizable Markov random field problem. Markov random field, MRF, problems are a class of problems which may be represented as a combinatorial undirected graph G = (y,E) where V is a set of vertices and E is a set of edges, each edge comprising k vertices, the edges describing relations between the vertices. The vertices correspond to a set of random variables X = indexed by the vertices. The set of random variables form an MRF with respect to the graph G if they satisfy at least one of the following three Markov properties:

[0053] 1 . Pairwise Markov property, whereby any two non-adjacent variables are conditionally independent given all other variables: Xu

[0054] 2. Local Markov property, whereby a variable is conditionally independent of all other variables given its neighbors: Xv1 X{^\v[v]} I^Gv(v), where N[v] = v u N(v).

[0055] 3. Global Markov property, whereby any two subsets of variables are conditionally independent given a separating subset: XA1 XB\XS, where every path from a node in A to a node in B passes through S.

[0056] For most real-world applications representable as pairwise MRFs, the probability distributions associated to the vertices are such that the three properties are equivalent.

[0057] Solving the MRF problem comprises determining a joint distribution of the random variables relative a cost function which assigns values to both vertices and edges. Broadly speaking, the cost function may be such that it ranks the vertices and edges by their relative importance or relative influence on the MRF problem. Optimization problems which may be represented as learning the parameters of an MRF problem are a subclass of graphical problems, where the random variables whose joint distribution is to be learned have dependencies between them which may be represented as an undirected (hyper)graph. Undirected here means that the dependencies between the vertices encoded by the edges are symmetrical.

[0058] The subclass of MRF models referred to as pairwise MRF models are represented as simple graphs, where each edge encodes a dependency between exactly two vertices. The class of non-pairwise MRF models are represented as fc-uniform hypergraphs, where each edge expresses a dependency between k vertices. MRF models are particularly suited for modelling situations where there may be a cyclical dependency between vertices.

[0059] A clique is a fully connected graph, wherein the set of edges consists of every possible edge given the vertex set. In the case of simple graphs, a clique is a graph on n vertices, where the edge set consists of all (”) size two subsets of the vertex set. Similarly, a fc-uniform hypergraph on n vertices is a clique if the edge set consists of all subsets of size k of the vertex set.

[0060] The class of clique factorizable MRFs is of interest when analyzing a system optimization problem as an MRF. The Markov properties of an arbitrary collection of random variables relative an arbitrary graph may be difficult to establish, so analysis of MRFs is typically restricted to MRFs corresponding to a set of random variables X = where P(X = x) is the probability of a particular configuration of the random variables and P(X = x) can be understood as a joint distribution of the set of random variables, such that the joint density can be factorized over cliques of the graph:

[0061] In Eq. 1 , cZ(G) denotes a set of cliques C of G and the functions may be referred to as clique potentials. The set of cliques and associated clique potentials is referred to as a clique factorization of the MRF model. The clique factorization is maximal if the number of cliques is minimized over all clique factorizations of the graph. Note that there may be several clique covers of the graph structure, but a unique maximal clique factorization in a given graph.

[0062] With reference to Fig. 5a-c, Fig. 5a illustrates an example of a graph on 4 vertices labelled A, B, C, and D. Fig. 5b illustrates a factor graph of a maximal clique factorization of Fig. 5a of two factors, factor containing the vertices A, B, and C, and factor f2containing the vertices C and D. Fig. 5c illustrates a factor graph of a non-maximal clique factorization of Fig. 5a of 4 factors - A of 2 vertices each, so factor contains vertices A and B, factor y contains vertices B and C, factor f3contains vertices A and C, and factor y contains vertices C and D.

[0063] A clique factorization of a graph G may be represented as a factor graph. A factor of an MRF is a function which, broadly speaking, describes compatibilities between different values of the random variables of the corresponding clique. A factorized joint distribution of an MRF as in Eq. 1 can be thought of as a normalized product of the factors of the corresponding clique factorization. The clique potentials are therefore normalized factors of the MRF model.

[0064] Returning to Fig. 1 , the method further comprises S101 partitioning the graph structure based on the importance metric. Partitioning the graph structure comprises creating a plurality of (possibly overlapping) partitions of the graph in the form of induced subgraphs, each partition comprising at least one clique of the graph. The graph structure is associated with a system optimization problem, and associated to the system optimization problem is an importance metric. The importance metric assigns a relative impact to each vertex and each edge of the graph. The relative impact is a measure of how central the node or vertex is to the system, or, equivalently, the impact of the node or vertex on the solution to the problem. Broadly, in a telecommunications network, comparatively small proportional energy savings in a high-traffic RAN node may have a much larger impact on overall energy consumption than proportionally larger energy savings in a low-traffic RAN node. The high-traffic RAN node would then be assigned a higher relative importance by the importance metric in the problem description than the low-traffic RAN node.

[0065] With reference to Fig. 7, a graph partitioning process is depicted. Step A of the graph partitioning process comprises identifying a maximal clique cover of the graph, in this case resulting in 3 cliques 1-3. The graph comprises a split edge, clique 2. A partitioning of the graph may be obtained by splitting the graph along clique 2 to obtain cluster 1 and cluster 2 as shown in step B. Splitting the graph along clique 2 results in the partitioning of step C, comprising partitions 1 and 2, where partition 1 comprises clique 1 and clique 2 and partition 2 comprises clique 2 and clique 3.

[0066] Partitioning the graph may optionally comprise obtaining S102 a plurality of graph partitionings. The plurality of graph partitionings may all be based on the same maximal clique factorization of the underlaying graph of the MRF model. In some embodiments, obtaining a plurality of graph partitionings may comprise obtaining a pre-determined number of graph partitionings obtained by assigning the cliques of the maximal clique factorization to different parts of the partitioning.

[0067] The method may optionally comprise 105 selecting, from the plurality of graph partitionings, a graph partitioning which can be mapped onto available quantum resources. In particular, some graph partitionings from the plurality of graph partitionings may comprise partitions which are larger than what may be executed on available quantum resources. Selecting a graph partitioning which can be mapped onto available quantum resources then comprises selecting a partitioning which fits on available QPUs. If several of the graph partitionings can be mapped onto available quantum resources, selecting may comprise selecting a partitioning at random.

[0068] Additionally, selecting a graph partitioning may comprise selecting S104 a graph partitioning with lowest associated cost. The associated cost of a graph partitioning is the cost of the edges and vertices of the graph partitioning which overlap between two partitions of the graph partitioning. With reference to Fig. 7, the edge connecting the vertices labelled 2 and 3, as well as vertices 2 and 3 are overlapping between the partitions. The associated cost of the partitioning may be obtained by summing all the costs of the overlapping edges and vertices of the partitioning.

[0069] The method further comprises mapping S105 the MRF model to a plurality of quantum circuits. The selected mapping is such that there is a one-to-one correspondence between the plurality of quantum circuits and the partitions of the selected partitioning. The quantum circuits may comprise quantum circuit Bom machines, QCBMs. The basis for a QCBM Ansatz is a quantum circuit Ising Born machine implementing time evolution under a Pauli-Markov Hamiltonian as Uz(a) = exp(— iH(a)) (Eq.2) where H(a) denotes a simplified Pauli-Markov Hamiltonian defined as where Zvis the Pauli-Z matrix acting on qubit v and a is a vector of parameters for the gates of the resulting quantum circuit. The vector a is the parameters of the distribution that is to be learned by solving the optimization problem.

[0070] Using the graph illustrated in Fig. 5a as an example, Fig. 5b shows a corresponding factor graph with a maximal clique factorization. The clique factorization produces the following Hamiltonian: + c |Z3(Eq.4) which generates the quantum circuit in Fig. 5d. In particular, the gate parametrized by ajlacting on n qubits corresponds to exp (icj ®nZn).

[0071] By contrast, using a pairwise factorization of Fig. 5c, without regard to the larger graph structures of the MRF model as in Fig. 5b, yields the Hamiltonian: which generates the quantum circuit in Fig. 5e.

[0072] In particular, the quantum circuit of Fig. 5e does not comprise the 3-qubit gate corresponding to the clique of order 3 in the graph which is in the quantum circuit Fig. 5d.

[0073] Mapping the problem onto quantum circuits may optionally comprise selecting S106 an Ansatz for each quantum circuit. Selecting the Ansatz for the quantum circuit may, in embodiments, comprise selecting an Ansatz which encodes initially known information into the quantum circuit. The initially known information may be partial solutions to the optimization problem obtained from already executed quantum circuits, in the case where the quantum circuits are executed sequentially.

[0074] Learning in MRF models may be described mathematically as follows. Assume there is some underlying distribution P* induced by an MRF . To learn the parameters of the distribution, we obtain a set of training data D comprising n samples from Learning the distribution then comprises learning parameters of an MRF such that the distribution defined by , PM, is an accurate approximation of the original distribution P*. In implementations, the accuracy of the distribution PMis determined by comparing predictions made by PMto a validation set of data from the distribution P*. A criterion and / or threshold value for evaluating the accuracy of the distribution may be set by the person skilled in the art depending on the implementation. Setting a threshold value may comprise determining the accuracy needed for the evaluation and weighing the precision against the available computing resources. Weighing the precision against available computing resources may comprise setting a multicondition criterion for stopping, such as stopping when a rolling average accuracy on the validation set is above a threshold or stopping when a set maximum number of training iterations have been performed, whichever condition is reached first.

[0075] Initiating training the MRF model therefore comprises initiating the distributed quantum computer using a suitable machine learning algorithm for quantum machine learning to learn parameters of the QPU gates, so that the learned parameters perform accurately on a validation data set. Any state-of-the-art quantum machine learning algorithm may be used with the methods of the disclosure to perform the training, for example SGD, Adam, RMSprop optimizers, Genetic search algorithms, or SPSA stochastic approximators.

[0076] In embodiments presented herein, initiating training the quantum circuits may optionally comprise initiating S108 training overlapping vertices and / or edges with the lowest-impact partition. The lowest-impact partition is the partition of the graph partitioning in which the cost of the overlapping edge or overlapping vertex has the lowest relative impact. By way of a simplified example, if a graph partitioning comprises two partitions, A and B, where the cost associated to partition A is 2 and the cost associated to partition B is 3. A and B share an overlapping vertex, v, with associated cost 1. The relative impact of vertex v is larger in clique A than in partition B, and therefore vertex v is trained with partition B. Training vertex v with partition B comprises training the quantum gates associated with vertex v in the quantum circuit associated to partition B. When the quantum circuit associated with partition A is trained, the gate parameters associated to vertex v are set as fixed and not retrained. In embodiments of the method, initiating the training of the quantum circuits optionally comprises initiating S110 training the quantum circuits sequentially. In some embodiments, there are more quantum circuits to train than there are available quantum processing units and some or all of the quantum circuits have to be trained separately. Initiating a sequential training of the quantum circuits may optionally comprise selecting S109 an order to train the quantum circuits. In embodiments, the order may be selected based on considerations related to overlapping vertices. Alternatively or in addition, the order may be selected based on times when specific quantum computing resources are available.

[0077] Fig. 6 depicts an example of a quantum circuit according to an embodiment of the method, where the quantum circuit is obtained from mapping the problem of learning a distribution of a cell key performance indicator, KPI, in a time window for, e.g. generating realistic synthetic datasets, replacing missing data, or identifying anomalies in the network. The quantum circuit is split into parts 601 and 602 for the drawing to fit on the page, not due to technical considerations. The graph is constructed to describe the state of a mobile network at a given time window, where nodes correspond to transmission points and edges correspond to handovers between transmission points. The target distribution to learn is a discrete binary state describing the cell KPI in a time window T = 1 hour. The binary distribution is obtained by a threshold rule: setting the state to 1 if the energy consumption of the node is higher than the threshold value for at least one time interval longer than some time threshold At, here set to 1 minute, and 0 otherwise. The cost associated to each node is a function of the monetary cost of local power consumption.

[0078] The costs of the edges are, in this example, proportional to the aggregated handovers between the corresponding transmission points. The edge cost is determined by the fact that power consumption of a transmission point is directly dependent on the number of UEs served by the transmission point.

[0079] The target distribution to be learned by the quantum circuit of Fig 6 is derived from timeseries data of the node states, collected as joint states of the nodes per time window over a given period of time.

[0080] The input in the method is the graph topology with the determined clique structure, metrics giving the node weights and the edge costs, the target distribution to be learned, and QPU information about available QPUs, in this example the input size and available gate set for each available QPU.

[0081] Partitioning the graph structure may, in embodiments, comprise deriving a community structure by identifying strongly connected subgraphs formed by subsets of cliques of the graph. Identifying strongly connected subgraphs may, in embodiments, be performed by creating a clique-overlap matrix and identifying closely connected subgraphs forming communities. In embodiments where the QPU size is such that the largest found community size may be executed by an available QPU and the available gate set of available QPUs is compatible with the structure of the communities found, the communities may be used as partitions, the set of communities forming the partitioning of the graph.

[0082] Alternatively or in addition, the partitioning may be further refined by using the importance metric of the edges connecting found communities, and partitioning based on separating least important edges. Alternatively or in addition, if the largest partition is larger than the largest available QPU, a partitioning algorithm may be run again, using the largest partition as input to obtain at least two additional partitions.

[0083] In embodiments where the distributed quantum computing unit comprises k modules, a minimum fc-cut partitioning may be used to create k partitions in polynomial time.

[0084] Once the graph has been partitioned so that the partitions match the available quantum processing unit, node importance per partition is aggregated to determine an order in which to train the partitions. The partitions may then be mapped to the available QPUs and training of the quantum circuits in the determined order initiated. Figs. 8a-c depicts the result when performing a method presented herein to a random graph with 10 nodes simulating a communication network according to the embodiment of Fig 6. Fig. 8a depicts a scatterplot of the total variational distance and the training speed of implementation of applying methods as disclosed herein to the embodiment of Fig. 6 labelled quantum circuit Markov random field, QCMRF, compared to attempts to learn the same distribution using a quantum circuit Ising Bom machine, QCIBM, and a quantum circuit Born machine, QCBM. Fig. 8b depicts a line graph of total variational distance as a function of the number of training steps for the same three methods. Fig 8c depicts a negative log-likelihood loss function as a function of the number of training steps for each of the three methods. As can be seen, the methods presented herein show superior performance compared to the state-of-the-art on this type of problem.

Claims

CLAIMS1. A computer-implemented method (100) for system optimization using distributed quantum computing, the computer-implemented method performed by a network node, wherein the system optimization problem can be represented as a factorizable Markov random field, MRF, model, the method comprising: partitioning (S101 ) a graph structure of the MRF model into two or more partitions based on a cost function of the optimization problem; mapping (S105) the optimization problem onto a plurality of trainable quantum circuits, each trainable quantum circuit corresponding to a partition of the two or more partitions; initiating (S107) training each of the trainable quantum circuits on a separate quantum processing unit of the distributed quantum computing device.

2. The method (100) according to claim 1 , wherein the MRF model is a binary MRF model.

3. The method (100) according to claims 1 or 2, wherein the partitioning separates the partitions along edges of the MRF model with lowest total cost.

4. The method (100) according to any one of claims 1-3, wherein the cost function is based on a key performance indicator, KPI, of a telecommunications network.

5. The method (100) according to claim 4, wherein the KPI is a power consumption metric and the MRF model is associated to a power consumption optimization problem.

6. The method (100) according to any one of claims 1-5, wherein a clique of order n of the graph partitioning is mapped onto a quantum circuit comprising 2n- 1 quantum gates, each quantum gate corresponding to a sub-clique of the clique.

7. The method (100) according to any one of claims 1-6, wherein mapping (S105) the partitions of the optimization problem onto trainable quantum circuits depends on one or more of:- a size of a register of each of the quantum processing units of the distributed quantum computing device; and- an available set of quantum gates.

8. The method (100) according to any one of claims 1 -7, further comprising obtaining (S102) a plurality of graph partitionings, each graph partitioning associated with a cost relative to the cost function, and selecting (S103) the graph partitioning of the plurality of graph partitionings with lowest associated cost which can be mapped onto quantum circuits which can be implemented by available quantum computing resources.

9. The method (100) according to claim 8, wherein selecting (S103) a graph partitioning comprises selecting (S104) a graph partitioning wherein the partitions are split along edges that minimize the cost along split edges.

10. The method (100) according to any one of claims 1-9, wherein mapping (S105) a graph partition onto a quantum circuit comprises selecting (S106) a unitary operator for state initialization for the quantum circuit encoding the input states of the quantum circuit with initial knowledge of a solution to the optimization problem.11 . The method (100) according to any one of claims 1 -10, wherein when vertices belonging to split edges are trained (S108) with the partition to which the cost associated to the vertices has the least impact.

12. The method (100) according to any of claims 1-11 , wherein the number of quantum processing units is insufficient for the number of partitions in the chosen partitioning, further comprising: initiating training (S110) of each of the trainable quantum circuits sequentially on the available quantum processing units.

13. The method (100) according to claim 12, wherein an order to train the trainable quantum circuits is selected (S109) based on the cost associated to each partition.

14. The method (100) according to any one of claims 1-13, wherein initiating training (S107) of a quantum circuit comprises initiating a training process associatedto a machine learning algorithm on a quantum processing unit to obtain an approximation of optimal parameters for the system optimization problem.

15. A network node (300) comprising a memory (328) and processing circuitry (326), the network node configured for system optimization using distributed quantum computing device, wherein the system optimization problem can be represented as a factorizable Markov random field, MRF, model, the network node configured to: partition (S101 ) a graph structure of the MRF model into two or more partitions based on a cost function of the optimization problem; map (S105) the optimization problem onto a plurality of trainable quantum circuits, each trainable quantum circuit corresponding to a partition of the two or more partitions; initiate (S107) training each of the trainable quantum circuits on a separate quantum processing unit of the distributed quantum computing device.

16. The network node (300) according to claim 15, further configured to perform a method according to any one of claims 2-14.

17. The network node (300) according to claim 15 or 16, wherein in the network node is a radio access node (210A, 210B) in a telecommunication system (200).

18. A quantum computing device (400) comprising quantum processing circuitry (404, 406, 408), the quantum computing device configured to solve an optimization problem by implementing quantum circuits obtained according to any one of claims 1-14.

19. A computer program (330) comprising computer readable instructions such that executing the program on processing circuitry (326) of a network node (300) causes the network node to perform a method according to any one of claims 1-14.

20. A computer program product (332) comprising a computer readable storage medium on which a computer program (330) according to claim 19 is stored.