Parallelization of Controlled-NOT Gates in Quantum Computing Simulation
By simulating a controlled NAG parallelization system during the qubit reordering process, the inefficient memory access and thread localization problems in existing quantum computing simulation technologies are solved, and the simulation efficiency is improved.
Patent Information
- Application Number
- CN201980080089.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-12-06
- Filing Date
- 2019-11-22
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2039-11-22
AI Technical Summary
Existing quantum computing simulation technologies have inefficient execution time and efficient storage space requirements, resulting in inefficient simulation efficiency.
By simulating controlled NAG gates during qubit reordering, the system is parallelized to mitigate and/or reduce fragment access and thread locality of quantum memory.
It realizes more efficient quantum memory access and reduces thread locality, improving the efficiency of quantum computing simulation.
Smart Images

Figure CN113168582B_ABST
Abstract
Description
BACKGROUND OF THE INVENTION
[0001] Due to the size of quantum computing systems, directly simulating quantum computing is challenging even with powerful computers. For example, Smelyanskiy1 et al. discussed "implementing a quantum emulator on a classical computer, which can simulate general single-qubit gates and two-qubit controlled gates". "See Smelyanskiy1 et al.," qHiPSTER: Quantum High-Performance Software Testing Environment ", 2016, arXiv: 1601.07195v2 [Quant-ph], abstract. Additionally, Smelyanskiy1 et al. discussed the performance of "multiple single-node and multi-node optimizations, including vectorization, multithreading, cache blocking, and overlapping computation with communication". See ibid.
[0002] However, these simulations can be inefficient because they involve execution time (e.g., the time required to perform the computations for the simulation). These simulations can also be inefficient due to the amount of storage space required. Thus, there is an opportunity to improve quantum computing simulation. SUMMARY OF THE INVENTION
[0003] An overview is given below to provide a basic understanding of one or more embodiments of the present invention. This overview is not intended to identify key or important elements, or to delineate any scope of a particular embodiment of the present invention or any scope of any claims. Its sole purpose is to present concepts in a simplified form as a prelude to the more detailed description that is presented later. In embodiments of the present invention described herein, systems, computer-implemented methods, apparatuses, and / or computer program products are provided that facilitate parallelization of controlled NOT gates in quantum computing simulation.
[0004] According to one embodiment of the present invention, a system may include a memory storing computer-executable components and a processor executing the computer-executable components stored in the memory. The computer-executable components may include a replication component that simulates a controlled NOT gate during a qubit reordering process. The computer-executable components may further include an analysis component that performs memory access balancing based on the controlled NOT gates simulated by the replication component during qubit reordering. Thus, the advantage of reducing and / or alleviating fragmented access to quantum memory can be provided. Additionally, another advantage may be that inefficiencies in thread locality can be reduced and / or alleviated.
[0005] According to another embodiment of the present invention, a computer-implemented method may include simulating the controlled NOT gate during qubit reordering by a system operatively coupled to a processor. The computer-implemented method may further include performing memory access balancing by the system based on simulating the controlled NOT gate during the qubit reordering. Thus, the benefits of mitigating and / or reducing inefficient thread locality and / or fragmented access to quantum memory may be provided.
[0006] According to another embodiment of the present invention, there is provided a computer program product facilitating quantum computing simulation of a controlled NOT gate. The computer program product may include a computer-readable storage medium having program instructions stored therein. The program instructions may be executed by a processor to cause the processor to simulate the controlled NOT gate during qubit reordering. The program instructions may further cause the processor to perform memory access balancing based on the controlled NOT gate being simulated during the qubit reordering. Thus, the advantages of providing mitigation and / or reduction of fragmented access to quantum memory may be realized. Additionally, the advantages of providing mitigation and / or reduction of inefficient thread locality may be realized.
[0007] Another embodiment of the present invention relates to a method that may include selecting a first qubit and a second qubit by a system operatively coupled to a processor, where the first qubit is a control qubit. The method may further include reordering the first qubit with the second qubit by the system. The controlled NOT gate may be simulated during the reordering. Advantages of such a method include migrating and / or reducing fragmented access to quantum memory and / or mitigating and / or reducing inefficient thread locality.
[0008] Another embodiment of the present invention relates to a computer program product that facilitates improving quantum computing simulation of a controlled NOT gate while avoiding unbalanced memory access to a control qubit. The computer program product includes a computer-readable storage medium having program instructions stored therein, and the program instructions are executed by a processor to cause the processor to select the control qubit and a non-control qubit, where the non-control qubit and the control qubit are different qubits. The program instructions may further cause the processor to reorder the control qubit with the non-control qubit and simulate the controlled NOT gate when reordering the control qubit with the non-control qubit. An advantage of such a computer program product is that it can mitigate and / or reduce inefficient thread locality and / or fragmented access to quantum memory. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1A is a non-uniform memory access architecture and cache including balanced memory access
[0010] Figure 1B is a schematic representation of a non-uniform memory access architecture and cache lines including unbalanced memory access due to inefficient thread locality.
[0011] Figure 1C is a schematic representation of a non-uniform memory access architecture and cache lines including unbalanced memory access due to segmented access.
[0012] Figure 2 It is a schematic representation of the simulation results of a controlled NOT gate.
[0013] Figure 3 is a bit reordering example according to one embodiment of the present invention.
[0014] Figure 4 is a block diagram of a system that facilitates parallelization of controlled NOT gates in quantum computing simulations according to one embodiment of the present invention.
[0015] Figure 5 is a block diagram of a system implementing selection of qubits for memory access balancing according to one embodiment of the present invention.
[0016] Figure 6 is a flow chart of a computer-implemented method for facilitating parallelization of controlled-NOT gates in quantum computing simulations according to one embodiment of the present invention.
[0017] Figure 7 is a flow chart of a computer-implemented method that facilitates selecting one or more bits to assist in qubit reordering according to one embodiment of the invention.
[0018] Figure 8 is a flow chart of a computer-implemented method that facilitates evaluation of memory access balancing attempts according to one embodiment of the present invention.
[0019] Figure 9 is a flow chart of a computer-implemented method for facilitating parallelization of controlled-NOT gates in quantum computing simulations according to one embodiment of the present invention.
[0020] Figure 10 is a block diagram of an operating environment in which embodiments of the present invention are facilitated. DETAILED DESCRIPTION
[0021] The following detailed description is illustrative only and is not intended to limit the embodiments of the invention and / or its application or use. In addition, it is not intended to be bound by any explicit or implicit information presented in the previous background technology or summary or detailed description.
[0022] Reference is now made to the accompanying drawings to describe embodiments of the present invention, where the same reference numerals are always used to denote the same elements. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a more thorough understanding of the embodiments of the present invention. However, in various instances, it is apparent that the embodiments of the present invention may be practiced without these specific details.
[0023] In quantum computing, the controlled NOT gate (CNOT or C-NOT) is a quantum gate that can be simulated to various degrees of precision using a combination of CNOT gates and single-qubit rotations. A quantum gate (or quantum logic gate) is a basic quantum circuit that operates on a small number of qubits. Quantum gates are reversible gates and are the building blocks of quantum circuits.
[0024] The state of the q qubits of a quantum computer is represented by a matrix of size 2 q × 2. The indices of the matrix are represented in binary format, and the probability of 0 for the i-th qubit is the calculated complex value, where the i-th bit of the index (in order from right to left) is 0. For example, the state of four qubits is represented by a matrix of 16 complex values, and the probability of 0 for the second qubit is calculated using the complex values at indices 0000, 0001, 0100, 0101, 1000, 1001, 1100, and 1101 in the matrix (all second binary digits in these indices are 0).
[0025] Quantum gates can be represented by unitary matrices. Various quantum gates can operate on the space of one or two qubits. A quantum gate can be described by a unitary matrix of size 2 n × 2 n × 2. The variables (e.g., quantum states) on which the gate acts are vectors in 2 n complex dimensions, where n is the number of qubits on which the gate acts (e.g., the number of qubits of the variable). A controlled gate acts on two qubits, where one qubit is used as the control for the operation.
[0026] From the perspective of non-uniform memory access (NUMA) architectures and cache lines, the simulation of CNOT may result in unbalanced memory accesses. For example, unbalanced memory accesses may be based on inefficient thread locality and / or fragmented accesses.
[0027] More specifically, CNOT can modify half of the states in a quantum computing simulation. The control bit q c can determine the modification region. The teleportation bit q t can determine what states are swapped. By knowing the NUMA architecture, individual threads can frequently access the states allocated in their NUMA nodes. This can be expressed as the thread having affinity for the NUMA node. Excessively high q c and / or low q cmay lead to inefficient memory access in its simulation. For example, if q c is high, the memory access may become unbalanced. In another example, if q c is low, the memory access may become fragmented.
[0028] Figure 1A FIG. 100 shows a schematic representation of a NUMA architecture and cache lines including balanced memory access, where each thread accesses local memory. In this example, the memory allocation is between NUMA0 and NUMA1. For the example shown, there can be 32 states, and this is a 5-qubit example. To simulate a quantum computer, matrices of complex number types can be utilized, and for the example shown, there are 32 quantum complex values for representing a 5-qubit quantum computer. As shown, the first set of boxes 102 (e.g., the first 16 states shown) represents NUMA0, while the second set of boxes 104 (e.g., the last 16 states shown) represents NUMA1. The boxes can include corresponding complex values. For example, the first box has a complex value of 00001.
[0029] The thread affinity of NUMA0 is represented by the first set of arrows 106, and the thread affinity of NUMA1 is represented by the second set of arrows 108. In addition, the cache lines are represented by arrows 110.
[0030] The CNOT operation can utilize two qubits. For the CNOT operation, for example, half of the states can be swapped. The following formula can be used to swap qubits, where CX is the CNOT gate and q is the qubit.
[0031] CX q[a]q[b]: If the a-th bit is 1, then swap the b-th bit
[0032] Formula 1
[0033] Using the above Figure 1A formula gives CX q[3]q[2]. Therefore, Figure 1A one or more threads of access local memory, which includes balanced memory access.
[0034] Figure 1B FIG. 112 shows a schematic representation of a NUMA architecture and cache lines that include unbalanced memory access due to inefficient thread locality. Using Formula 1 gives Figure 1B CX q[4]q[2], which results in a problem of inefficient thread locality. This result is due to q cToo high (e.g., 4 in this example). In this case, the threads of NUMA0 (e.g., NUMA0 hardware) always access the state of NUMA1, as indicated by the first set of arrows 114. As shown, accessing the memory value of the remote memory node takes longer to access the memory than in the case of performing a local memory access (as Figure 1A shown). Therefore, Figure 1B has slower memory access. A solution to the various aspects discussed herein is qubit reordering to minimize and / or reduce the inefficiency of thread locality, resulting in more efficient thread locality.
[0035] Figure 1C Shows a schematic representation 116 of a NUMA architecture and cache lines, which include unbalanced memory accesses due to segmented access. In Figure 1C the case of, CX q[0]q[2] is obtained using Equation 1, as shown by the four sets of arrows 118 - 124. In this case, q c is too low (e.g., 0 in this example), which results in the problem of segmented access. A solution to the various aspects discussed herein is qubit reordering to minimize and / or reduce the fragmented access of quantum memory.
[0036] Figure 2 Shows a schematic representation 200 of the results of a controlled NOT gate simulation. The target bits 202 (e.g., 30 target bits numbered 0 to 29) are represented on the X-axis, the control bits 204 (e.g., 30 control bits numbered 0 to 29) are represented on the Z-axis, and the elapsed time 206 (in seconds) is represented on the Y-axis. The schematic representation is a 30-qubit simulation with 4 NUMA nodes and 128-byte cache lines.
[0037] As discussed in reference Figure 1B , inefficient thread locality is shown within the box 208. Based on 30 control bits being less than or equal to log2(number of NUMA) or 30 control bits <= log2(number of NUMA), this problem may occur. The present invention solves the problem of inefficient thread locality by implementing qubit reordering for quantum computing to avoid unbalanced memory access while simulating the CNOT gate and at the same time mitigate and / or reduce inefficient thread locality, resulting in more efficient thread locality.
[0038] As regarding Figure 1CThe issue of segmented access being discussed is shown within block 210. This can be caused by the control bit being less than log2(cache line size) or the control bit being less than log2(cache line size). The present invention solves the problem of fragmented access by implementing qubit reordering for quantum computing to avoid unbalanced memory access, while simulating CNOT gates and at the same time alleviating and / or reducing fragmented access to quantum memory.
[0039] The various aspects discussed herein can change quantum register allocation to minimize fragmented access and inefficient thread locality. For example, Figure 3 An example of bit reordering according to an embodiment of the present invention is shown.
[0040] Shown at 300 is a first quantum register allocation. Applied to the first qubit (e.g., bit index 0) can be a first gate, which is a first X gate 302; a second gate, which is a CNOT gate 304; a third gate, which is a second X gate 306, and a first measurement 308. A second measurement 310 can be on the second qubit (e.g., bit index 1). Shown at 312 is a second quantum register allocation, which is a reallocation (or reordering) of the first quantum register allocation. For example, the first qubit and the second qubit can be swapped.
[0041] As an example, the following simple algorithm can be used for bit reordering. Qubits can be randomly allocated in a register, and the number of inefficient memory states can be calculated. The minimum number of inefficient memory accesses can be determined, and qubit reordering can be performed. The most efficient qubit reordering can be determined. Note that although an example algorithm has been described, other algorithms can be utilized, and the disclosed aspects are not limited to this example.
[0042] In quantum computing, when a quantum register is measured, its state (0 or 1) is passed to a classical register. Even if the runtime changes the allocation of qubits in a quantum circuit, by keeping the allocation of the classical register, the result of the circuit is equivalent. The embodiments described herein include systems, computer-implemented methods, and computer program products that facilitate CNOT parallelization. For example, as discussed herein, bit reordering can be performed to minimize the number of control qubits that are equal to or higher than (the number of qubits) - log2(the number of NUMA nodes). Bit reordering discussed herein can additionally or alternatively be performed to minimize the number of control qubits that are less than log2(the number of NUMA nodes).
[0043] Figure 4FIG. 400 is a block diagram of a system 400 that facilitates controlled-NOT gate parallelization in quantum computing simulation. Aspects of the systems (e.g., system 400, etc.), apparatus, or processes explained in this disclosure may constitute machine-executable components contained within one or more machines, e.g., contained within one or more computer-readable media (or media) associated with one or more machines. When executed by one or more machines (e.g., computers, computing devices, virtual machines, etc.), such components may cause the machines to perform the described operations.
[0044] System 400 may be a system that includes a processor and / or any type of component, machine, device, facility, apparatus, and / or instrument that can be capable of communicating effectively and / or operatively with a wired and / or wireless network. The components, machines, devices, equipment, facilities, and / or instruments that may constitute system 400 may include tablet computing devices, handheld devices, server-class computer machines and / or databases, laptop computers, notebook computers, desktop computers, cellular phones, smart phones, consumer appliances and / or instruments, industrial and / or commercial devices, handheld devices, digital assistants, multimedia Internet-enabled phones, multimedia players, etc.
[0045] System 400 may be a quantum computing system associated with a variety of technologies, such as but not limited to quantum circuit technology, quantum processor technology, quantum computing technology, artificial intelligence technology, medical and materials technology, supply chain and logistics technology, financial services technology, and / or other digital technologies. System 400 may employ hardware and / or software to solve problems that are inherently highly technical, non-abstract, and cannot be performed as a set of mental acts by a human. Some of the processes performed may be executed by one or more dedicated computers (e.g., one or more dedicated processing units, dedicated computers with quantum computing components, etc.) to perform defined tasks related to machine learning.
[0046] System 400 and / or the components of system 400 may be used to solve new problems arising due to advancements in the above-mentioned technologies, computer architectures, etc. System 400 may provide technical improvements to quantum computing systems, quantum circuit systems, quantum processor systems, artificial intelligence systems, and / or other systems. System 400 may also provide technical improvements to a quantum processor (e.g., a superconducting quantum processor) by improving the processing performance, processing efficiency, processing characteristics, timing characteristics, and / or power efficiency of the quantum processor.
[0047] System 400 may include a replication component 402, a parallelization component 404, a processing component 406, a memory 408, and / or a storage 410. The memory 408 may store computer-executable components and instructions. The processing component 406 (e.g., a processor) may facilitate the execution of instructions (e.g., computer-executable components and corresponding instructions) by the replication component 402, the parallelization component 404, and / or other system components. One or more of the replication component 402, the parallelization component 404, the processing component 406, the memory 408, and / or the storage 410 may be electrically coupled, communicatively coupled, and / or operatively coupled to each other to perform one or more functions of the system 400.
[0048] The replication component 402 may receive quantum circuit data 412 as input data. For example, the quantum circuit data 412 may be a machine-readable description of a quantum circuit. A quantum circuit may be a model for one or more quantum computations associated with a sequence of quantum gates. In one example, the quantum circuit data may include text data indicating a text format language (e.g., QASM text format language) that describes the quantum circuit. For example, the text data may describe, in text, one or more qubit gates of the quantum circuit associated with one or more qubits.
[0049] At least in part based on the input data (e.g., quantum circuit data 412), the replication component 402 may simulate a controlled NOT gate (e.g., CNOT gate 304) during qubit reordering. Additionally, the parallelization component 404 may perform memory access balancing based on the controlled NOT gates simulated by the replication component 402 during qubit reordering. The parallelization component 404 may output the result of the memory access balancing as output data 414. The parallelization component 404 may perform memory access balancing to reduce and / or minimize fragmented access to quantum memory (as discussed with respect to Figure 1C and Figure 2 ). Additionally, or alternatively, in an implementation, the parallelization component 404 may perform memory access balancing to reduce and / or minimize inefficient thread locality (as referenced with respect to Figure 1B and Figure 2 ).
[0050] The parallelization component 404 can perform memory access balancing and can generate output data 414 based on classification, correlation, inference, and / or expressions associated with the principles of artificial intelligence. For example, the parallelization component 404 and other system components can employ an automatic classification system and / or an automatic classification process to determine which qubit to select as a control qubit, which qubit to select as a target qubit, when to select one or more other qubits as control qubits and / or target qubits, whether the memory access balancing is successful, and so on. In one example, the parallelization component 404 can employ probability- and / or statistics-based analysis (e.g., decomposed into analysis utility and cost) to learn and / or generate inferences regarding the selection of one or more qubits and the corresponding balancing that should be applied to one or more qubits. In one aspect, the parallelization component 404 can include an inference component (not shown) that can partially utilize an inference-based scheme to facilitate learning and / or generating inferences associated with the selection of qubits and / or the results of memory access balancing in order to achieve reducing and / or minimizing fragmented access to quantum memory and / or reducing and / or minimizing thread locality, thereby further enhancing the automated aspects of the parallelization component 404.
[0051] The parallelization component 404 can employ any suitable machine learning-based techniques, statistics-based techniques, and / or probability-based techniques. For example, the parallelization component 404 can employ expert systems, fuzzy logic, SVM, hidden Markov models (HMMs), greedy search algorithms, rule-based systems, Bayesian models (e.g., Bayesian networks), neural networks, other non-linear training techniques, data fusion, utility-based analysis systems, systems employing Bayesian models, etc. In another aspect, the parallelization component 404 can perform a set of machine learning computations associated with qubit selection and / or memory access balancing. For example, the parallelization component 404 can perform a set of clustering machine learning computations, a set of logistic regression machine learning computations, a set of decision tree machine learning computations, a set of random forest machine learning computations, a set of regression tree machine learning computations, a set of least squares machine learning computations, a set of instance-based machine learning computations, a set of regression machine learning computations, a set of support vector regression machine learning computations, a set of k-means machine learning computations, a set of spectral clustering machine learning computations, a set of rule learning machine learning computations, a set of Bayesian machine learning computations, a set of deep Boltzmann machine computations, a set of deep belief network computations, and / or a set of different machine learning computations to determine the manner of memory access balancing and / or the results of memory access balancing.
[0052] It should be recognized that system 400 (e.g., replication component 402 and / or parallelization component 404, and other system components) performs qubit selection, memory access balancing (or rebalancing), and / or generates results of memory access balancing substantially simultaneously while replication component 402 emulates a controlled NOT gate during qubit reordering, which cannot be performed by a human (e.g., greater than the ability of a single human mind). For example, the amount of data processed by system 400 (e.g., replication component 402 and / or parallelization component 404) over a certain period of time, the speed at which the data is processed, and / or the type of data processed can be greater, faster, and different than the amount, speed, and data type processed by a single human mind over the same period of time. System 400 (e.g., replication component 402 and / or parallelization component 404) can also be fully operable to perform one or more other functions (e.g., fully powered on, fully executed, etc.) while also performing the above-mentioned quantum circuit analysis and / or pulse signal generation processes. Additionally, the output data 414 generated and coordinated by system 400 (e.g., replication component 402 and / or parallelization component 404) can include information that cannot be obtained manually by a user. For example, the type of information included in the quantum circuit data 412 that is used to facilitate memory access balancing and / or generate and output data 414, the various information associated with the quantum circuit data 412, and / or the optimization of the quantum circuit data 412 can be more complex than the information that can be obtained manually and processed by a user.
[0053] Figure 5 FIG. is a block diagram of a system 500 that implements selection of qubits for memory access balancing according to an embodiment of the present invention.
[0054] System 500 can include one or more components and / or functions of system 400, and vice versa. The various aspects discussed herein can be used to improve the quantum computing emulation of CNOT gates to avoid (e.g., mitigate and / or reduce) c (control qubit qbit or qubit qubit) unbalanced memory access, and improve cache hits for q t (target qubit) in a computer system with N NUMA nodes.
[0055] System 500 may include a selector component 502 that may select a first bit and a second bit for qubit reordering. The first bit may be a control bit. In an instance, the selector component 502 may select the first bit based on determining that the first bit is a bit that is equal to or higher than the sum of the binary logarithm (log2) of the number of qubits minus the number of non-uniform memory access nodes. In another instance, the selector component 502 may select the first bit based on determining that the first bit is a bit that is less than the binary logarithm (log2) of the number of non-uniform memory access nodes. In yet another instance, the selector component 502 may select the second bit based on determining that the second bit is different from the first bit and is not a target bit, where the first bit is a control bit. The selector component 502 may also select other qubits.
[0056] System 500 may also include an evaluation component 504 that may determine whether the memory access balance performed by the parallelization component 404 is successful or unsuccessful. Thus, the evaluation component 504 may determine whether there is an improvement in memory access or whether there is a lack of improvement in memory access. For example, the evaluation component 504 may analyze the memory access balance to determine whether a reduction and / or minimization of control qubits has been achieved, where the control qubits are equal to or higher than the sum of the first binary logarithm (log2) of the number of qubits minus the number of non-uniform memory access nodes. In another instance, the evaluation component 504 may analyze the memory access balance to determine whether a reduction and / or minimization of control qubits has been achieved, where the control qubits are less than the second binary logarithm (log2) of the number of non-uniform memory access nodes.
[0057] System 500 may also include an arrangement component 506 that may implement qubit reordering for quantum computing, where the arrangement component 506 may reorder the first bit using the second bit. Additionally, based on the evaluation component 504 determining that the qubit reordering and / or the memory access balance is successful (e.g., the above results have been achieved), the arrangement component 506 may reorder the first bit and / or the second bit using a third bit (and / or subsequent bits). For example, the evaluation component 504 (or another system component) may determine that there are more qubits that should be reordered, and thus, as discussed herein, the third bit (and / or subsequent bits) may be reordered until there are no additional qubits that should be reordered, as determined by the evaluation component 504 (or another system component).
[0058] Additionally, based on the evaluation component 504 determining that there is an improvement in memory access, the selector component 502 may select another qubit as a control qubit or as a target qubit. The parallelization component 404 may perform another memory access balance based on the CNOT gates emulated by the replication component 402 during the second (or subsequent) qubit reordering.
[0059] If the determination of the evaluation component 504 is that qubit reordering and / or memory access balancing is unsuccessful (e.g., the above results are not achieved), the restoration component 508 can restore the (most recent) qubit reordering.
[0060] In addition, based on the evaluation component 504 determining a lack of memory access improvement, the selector component 502 can select another qubit as the control qubit or as the target qubit. The parallelization component 404 can perform another memory access balance based on the CNOT gates emulated by the replication component 402 during the second (or subsequent) qubit reordering. As described above, the evaluation component 504 (or another system component) can determine that there are more qubits that should be reordered, and thus, as discussed herein, the third (and / or subsequent bits) can be reordered until, as determined by the evaluation component 504 (or another system component), there are no additional qubits that should be reordered.
[0061] Figure 6 is a flowchart of a computer-implemented method 600 according to an embodiment of the present invention that facilitates controlled-NOT gate parallelization in quantum computing simulation.
[0062] At 602 of method 600, a system operatively coupled to a processor can emulate a controlled-NOT gate during qubit reordering (e.g., via the replication component 402). In addition, at 604 of method 600, the system can perform memory access balancing based on the emulation of the controlled-NOT gate during qubit reordering (e.g., via the parallelization component 404). Memory access balancing can minimize fragmented access to quantum memory. Additionally or alternatively, memory access balancing can minimize thread locality.
[0063] Figure 7 is a flowchart of a computer-implemented method 700 according to an embodiment of the present invention that facilitates selecting one or more bits to assist qubit reordering.
[0064] Method 700 begins at 702, where a system including a processor can select a first bit and at least a second bit (e.g., via the selector component 502). According to various implementations, the first bit and at least the second bit can be selected for qubit reordering, and the first bit can be a control bit.
[0065] In an example, the selection of the first bit can be based on determining that the first bit is equal to or higher than the sum of the binary logarithm (log2) of the number of qubits minus the number of non-uniform memory access nodes. In another example, the first bit can be selected based on determining that the first bit is less than the binary logarithm (log2) of the number of non-uniform memory access nodes. In yet another example, the selection of the first bit can be based on determining that the second bit is different from the first bit and is not the target bit, where the first bit serves as a control bit.
[0066] At 704, the method 700 continues when the system can simulate a controlled-NOT gate during qubit reordering (e.g., via the copy component 402). Additionally, at 706 of the method 700, the system can perform memory access balancing (e.g., via the parallelization component 404). During memory access balancing, a controlled-NOT gate can be simulated.
[0067] Figure 8 It is a flowchart of a computer-implemented method 800 that facilitates the evaluation of a memory access balancing attempt according to an embodiment of the present invention.
[0068] When the system operatively coupled to the processor simulates a controlled-NOT gate during a first qubit reordering (e.g., via the copy component 402), the method 800 starts at 802. Additionally, at 804, the system can perform a first memory access balancing (e.g., via the parallelization component 404) based on the simulation of the controlled-NOT gate during the first qubit reordering.
[0069] At 806 of the method 800, it can be determined whether the first memory access balancing is successful (e.g., via the evaluation component 504). If it is determined that the first memory access balancing is not successful ("no"), then at 808, the system can reverse the first qubit reordering (e.g., via the reversal component 508). If it is determined that the first memory access balancing is successful ("yes"), or after the reversal at 806, the method 800 can continue at 810, and the system can simulate a controlled-NOT gate during a second qubit reordering (e.g., via the parallelization component 404).
[0070] Method 800 may continue at 812 to determine whether a second memory access balance is successful (e.g., via evaluation component 504). If not successful ("no"), method 800 may return to 806 and may restore the second qubit reordering (e.g., via restoration component 508). If successful ("yes"), or after restoration of the second qubit reordering, and based on the existence of more qubits to reorder, method 800 may continue at 810 and may emulate the controlled NOT gate during a subsequent qubit reordering (e.g., via parallelization component 404). The determination at 812, the restoration at 806, and / or the emulation at 810 may continue until a determination is made that there are no additional qubits to reorder according to the various implementations.
[0071] Figure 9 is a flowchart of a computer-implemented method 900 that facilitates controlled NOT gate parallelization in quantum computing emulation according to an embodiment of the present invention.
[0072] Method 900 begins at 902, where a system including a processor may select a first bit, which may be a control bit q c (e.g., via selector component 502). The selection of the first bit may be based on a determination that one of the plurality of bits includes a value equal to or higher than [(the number of qubits) - log2(the number of NUMA nodes)]. Alternatively, the selection of the first bit may be based on a determination that one of the plurality of bits includes a value lower than log2(the number of NUMA nodes).
[0073] At 904 of method 900, the system may select a second bit, which may be bit q (e.g., via selector component 502). The selection of the second bit may be based on a determination that one of the plurality of bits is not the first bit (e.g., not control bit q c , is a bit different from control bit q c and the bit has not been designated as a target bit, where q c is the control bit.
[0074] At 906 of method 900, the system may attempt to reorder the first bit with the second bit, e.g., (reorder control bit q with bit q c )(e.g., via parallelization component 404). Additionally, at 908 of method 900, it may be determined whether the memory access has been improved (e.g., via evaluation component 504). For example, the determination at 908 may be whether an inefficient memory access has been improved (e.g., the inefficiency has been improved).
[0075] If it is determined that the inefficiency has not been improved ("No"), then method 900 may continue at 910, and the last reordering may be restored (e.g., via restoration component 508). For example, the last reordering may be restored to the previous reordering. After the restoration of the last reordering, or if the determination at 908 is that the inefficiency has been improved ("Yes"), then method 900 may continue at 912 to determine whether additional reordering should be implemented (e.g., via evaluation component 504). For example, the determination of performing additional reordering may be based on whether there are additional bits that can be reordered.
[0076] If no additional reordering is performed ("No"), then method 900 may stop at 914. However, if additional reordering should be performed ("Yes"), then the method may continue at 902, where the first bit is selected (e.g., the control bit may be reused). However, in some implementations, the selection at 902 may be another bit that is available as the control bit. It should be understood that the determination of performing another (or subsequent) reordering at 912 may be recursive. For example, after selecting the control bit (e.g., the first bit) and another bit (e.g., the second bit or subsequent bits), reordering may be performed, and it may be determined whether another reordering should be performed.
[0077] As discussed herein, a system, computer-implemented method, computer program product, or other embodiments of the present invention may be provided that may facilitate quantum computing simulation of a controlled NOT gate. For example, a computer program product may include a computer-readable storage medium having program instructions included therein that are executable by a processor to cause the processor to simulate the controlled NOT gate during qubit reordering and perform memory access balancing based on the controlled NOT gate simulated during the qubit reordering.
[0078] In an example, the program instructions may cause the processor to select the first bit based on determining that the first bit is equal to or higher than the bit that is the sum of the second binary logarithm (log2) of the number of qubits minus the number of non-uniform memory access nodes, where the first bit is the control bit. Alternatively, the program instructions may cause the processor to select the first bit based on determining that the first bit is less than the second binary logarithm (log2) of the number of the non-uniform memory access nodes. According to some embodiments, the program instructions may cause the processor to select the second bit based on the determination that the second bit is different from the first bit and is not the target bit, where the first bit is the control bit.
[0079] In addition, as discussed herein, a system, computer-implemented method, computer program product, or other embodiments can be provided that can help improve the simulation of a controlled NOT gate in quantum computing while avoiding unbalanced memory access to a control qubit. For example, a computer program product can include a computer-readable storage medium having program instructions contained therein that are executable by a processor to cause the processor to select a control qubit and a non-control qubit, where the non-control qubit and the control qubit are different qubits, reorder the control qubit with the non-control qubit, and simulate a controlled NOT gate when reordering the control qubit with the non-control qubit.
[0080] In an exemplary embodiment, the program instructions can cause the processor to reduce the control qubits that are equal to or higher than the first binary logarithm (log2) of the first quantity of qubits minus the second quantity of non-uniform memory access nodes. According to another example implementation, the program instructions can cause the processor to reduce the control qubits that are lower than the second binary logarithm (log2) of the quantity of non-uniform memory access nodes.
[0081] For simplicity of explanation, the computer-implemented method is depicted and described as a series of acts. It will be understood and appreciated that the present invention is not limited by the acts and / or order of acts shown, e.g., the acts can occur in various orders and / or concurrently, and can occur in conjunction with other acts not presented and described herein. Further, not all acts shown are necessary to implement the computer-implemented method according to the disclosed subject matter. Additionally, those skilled in the art will understand and appreciate that the computer-implemented method can alternatively be represented as a series of interrelated states via a state diagram or events. Additionally, it should also be understood that the computer-implemented methods disclosed hereinafter and throughout this specification can be stored on an article of manufacture to facilitate the transfer and conveyance of these computer-implemented methods to a computer. As used herein, the term article of manufacture is intended to encompass a computer program accessible from any computer-readable device or storage medium.
[0082] To provide context for the various aspects of the present invention, Figure 10 and the following discussion is intended to provide a general description of a suitable environment in which the various aspects of the present invention can be implemented. Figure 10 A block diagram of an operating environment 1000 in which embodiments of the present invention can be facilitated is shown. Refer Figure 10, the operating environment 1000 may also include a computer 1012. The computer 1012 may also include a processing unit 1014, a system memory 1016, and a system bus 1018. The system bus 1018 couples system components including, but not limited to, the system memory 1016 to the processing unit 1014. The processing unit 1014 can be any of a variety of available processors. Dual microprocessors and other multiprocessor architectures can also be used as the processing unit 1014. The system bus 1018 can be any of several types of bus structures, including a memory bus or memory controller, a peripheral bus or external bus, and / or a local bus using any of a variety of available bus architectures, including but not limited to Industry Standard Architecture (ISA), Micro Channel Architecture (MSA), Extended ISA (EISA), Intelligent Drive Electronics (IDE), Video Electronics Standards Association (VESA), Local Bus (VLB), Peripheral Component Interconnect (PCI), Card Bus, Universal Serial Bus (USB), Advanced Graphics Port (AGP), FireWire (IEEE 1394), and Small Computer System Interface (SCSI). The system memory 1016 may also include volatile memory 1020 and non-volatile memory 1022. The Basic Input / Output System (BIOS) contains basic routines such as transferring information between elements within the computer 1012 during startup, and it is stored in the non-volatile memory 1022. By way of illustration and not limitation, the non-volatile memory 1022 may include Read-Only Memory (ROM), Programmable ROM (PROM), Electrically Programmable ROM (EPROM), Electrically Erasable Programmable ROM (EEPROM), flash memory, or non-volatile Random Access Memory (RAM) (e.g., Ferroelectric RAM (FeRAM)). The volatile memory 1020 may also include RAM, which acts as an external cache. By way of illustration and not limitation, RAM can be obtained in many forms, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Direct Rambus RAM (DRRAM), Direct Rambus Dynamic RAM (DRDRAM), and Rambus Dynamic RAM.
[0083] The computer 1012 may also include removable / non-removable, volatile / non-volatile computer storage media. For example, Figure 10Disk storage 1024 is shown. Disk storage 1024 may also include, but is not limited to, devices such as disk drives, floppy disk drives, tape drives, Jaz drives, Zip drives, LS-100 drives, flash cards, or memory sticks. Disk storage 1024 may also include separate storage media or storage media in combination with other storage media, and other storage media includes, but is not limited to, optical disk drives such as compact disk ROM devices (CD-ROM), CD recordable drives (CD-R drives), CD rewritable drives (CD-RW drives), or digital versatile disk ROM drives (DVD-ROM). To facilitate connecting disk storage 1024 to system bus 1018, a removable or non-removable interface such as interface 1026 is typically used. Figure 10 Software that acts as an intermediary between the user and the basic computer resources described in a suitable operating environment 1000 is also depicted. Such software may also include, for example, operating system 1028. Operating system 1028, which may be stored on disk storage 1024, is used to control and allocate the resources of computer 1012. System applications 1030 utilize the management of resources by operating system 1028 through program modules 1032 and program data 1034 stored, for example, in system memory 1016 or disk storage 1024. It should be understood that the present disclosure may be implemented with various operating systems or combinations of operating systems. The user inputs commands or information into computer 1012 through input device 1036. Input device 1036 includes, but is not limited to, pointing devices such as a mouse, trackball, stylus, touchpad, keyboard, microphone, joystick, gamepad, satellite dish, scanner, TV tuner card, digital camera, digital video camera, web camera, etc. These and other input devices are connected to processing unit 1014 via interface port 1038 through system bus 1018. Interface port 1038 includes, for example, serial ports, parallel ports, game ports, and universal serial bus (USB). Some of the (one or more) output devices 1040 use the same type of ports as the (one or more) input devices 1036. Thus, for example, a USB port may be used to provide input to computer 1012 and output information from computer 1012 to output device 1040. Output adapter 1042 is provided to account for the presence of certain output devices 1040, such as monitors, speakers, and printers, and other output devices 1040 that require a dedicated adapter. By way of example and not limitation, output adapter 1042 includes video cards and sound cards that provide a means of connection between output device 1040 and system bus 1018. It should be noted that other devices and / or systems of devices provide input and output capabilities, such as remote computer 1044.
[0084] Computer 1012 can operate in a networked environment using a logical connection to one or more remote computers, such as remote computer 1044. Remote computer 1044 can be a computer, server, router, network PC, workstation, microprocessor-based appliance, peer device, or other common network node, etc., and generally may also include many or all of the elements described relative to computer 1012. For simplicity, only memory storage device 1046 is shown with remote computer 1044. Remote computer 1044 is logically connected to computer 1012 via network interface 1048 and then physically connected via communication link 1050. Network interface 1048 includes wired and / or wireless communication networks, such as local area network (LAN), wide area network (WAN), cellular network, etc. LAN technologies include Fiber Distributed Data Interface (FDDI), Copper Distributed Data Interface (CDDI), Ethernet, Token Ring, etc. WAN technologies include, but are not limited to, point-to-point links, circuit-switched networks like Integrated Services Digital Network (ISDN) and its variants, packet-switched networks, and Digital Subscriber Line (DSL). Communication link 1050 refers to the hardware / software for connecting network interface 1048 to system bus 1018. Although shown inside computer 1012 for clarity, it can also be outside computer 1012. For illustrative purposes only, the hardware / software for connecting to network interface 1048 can also include internal and external technologies, such as modems including conventional telephone-grade modems, cable modems, and DSL modems, ISDN adapters, and Ethernet cards.
[0085] The present invention can be a system, method, and / or computer program product at any possible technical detail integration level. The computer program product can include a computer-readable storage medium (or media) having thereon computer-readable program instructions for causing a processor to implement various aspects of the present invention. The computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device.
[0086] The computer-readable storage medium can be, by way of example and not limitation, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing.
[0087] A non-exhaustive list of computer-readable storage media includes: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disks (DVDs), memory sticks, floppy disks, mechanically encoded devices such as punch cards or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage media as used herein is not to be construed as transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., optical pulses through an optical fiber cable), or electrical signals transmitted through wires.
[0088] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a respective computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device. The computer program instructions for performing the operations of the present invention can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through an Internet service provider via the Internet). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present invention.
[0089] Aspects of the present invention are described herein with reference to the flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions. These computer-readable program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a method for implementing the functions / acts specified in one or more blocks of the flowchart illustrations and / or block diagrams. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium in which the instructions are stored comprises an article of manufacture including instructions which implement aspects of the functions / acts specified in one or more blocks of the flowchart illustrations and / or block diagrams. The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other devices to produce a computer-implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other devices implement the functions / acts specified in one or more blocks of the flowchart illustrations and / or block diagrams.
[0090] The flowchart and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagrams may represent a module, a segment of code, or a portion of an instruction, which comprises one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.
[0091] Although the subject matter has been described above in the general context of computer-executable instructions of a computer program product running on one and / or more computers, those skilled in the art will recognize that the present disclosure may also be implemented in conjunction with other program modules. Generally, program modules include routines, programs, components, data structures, etc. that perform tasks and / or implement specific abstract data types. In addition, those skilled in the art will appreciate that the computer-implemented methods of the present invention may be implemented with other computer system configurations, including single-processor or multi-processor computer systems, small computing devices, mainframe computers, and computers, handheld computing devices (e.g., PDAs, telephones), microprocessor-based or programmable consumer or industrial electronic products, etc. The aspects shown may also be practiced in a distributed computing environment where tasks are performed by remote processing devices linked through a communication network. However, some aspects of the present disclosure, if not all, may be practiced on a stand-alone computer. In a distributed computing environment, program modules may be in local and remote memory storage devices.
[0092] As used in this application, the terms "component", "system", "platform", "interface", etc. may refer to and / or may include computer-related entities or entities related to an operating machine having one or more specific functions. The entities disclosed herein may be hardware, a combination of hardware and software, software, or software in execution. For example, a component may be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. By way of illustration, an application running on a server and the server can both be components. One or more components may reside within a process and / or an execution thread, and a component may be located on one computer and / or distributed between two or more computers. In another example, corresponding components may execute from various computer-readable media on which various data structures are stored. These components may communicate via local and / or remote processes, such as in accordance with a signal having one or more data packets (e.g., data from one component that interacts with another component in a local system, a distributed system, and / or interacts with other systems via a network such as the Internet). As another example, a component may be a device having a specific function provided by a mechanical component operated by an electrical or electronic circuit, which is operated by a software or firmware application executed by a processor. In such a case, the processor may be internal or external to the device and may execute at least a portion of the software or firmware application. In another example, a component may be a device that provides a specific function through an electronic component rather than a mechanical component, where the electronic component may include a processor or other means to execute software or firmware that at least partially imparts the function of the electronic component. In one aspect, a component may emulate an electronic component via a virtual machine, such as within a cloud computing system.
[0093] In addition, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified or clear from the context, "X employs A or B" is intended to mean any natural inclusive permutation. That is, if X uses A; X employs B; or X employs both A and B, then "X employs A or B" is satisfied in any of the foregoing instances. In addition, unless otherwise specified or clear from the context to refer to the singular form, the articles "a" and "an" as used in this specification and the drawings shall generally be construed to mean "one or more". As used herein, the terms "example" and / or "exemplary" are used to mean serving as an example, instance, or illustration. To avoid doubt, the subject matter disclosed herein is not limited by these examples. In addition, any aspect or design described herein as "example" and / or "exemplary" is not necessarily to be construed as preferred or advantageous over other aspects or designs, nor does it imply exclusion of equivalent exemplary structures and techniques known to those of ordinary skill in the art.
[0094] As used in this specification, the term "processor" can refer to substantially any computing processing unit or device, including but not limited to a single-core processor; a single processor with software multithreading execution capabilities; a multi-core processor; a multi-core processor with software multithreading execution capabilities; a multi-core processor with hardware multithreading technology; a parallel platform; and a parallel platform with distributed shared memory. Additionally, a processor can refer to an integrated circuit, an application specific integrated circuit (ASIC), a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic controller (PLC), a complex programmable logic device (CPLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof that is designed to perform the functions described herein. Furthermore, a processor can employ a nanoscale architecture, such as but not limited to transistors, switches, and gates based on molecules and quantum dots, in order to optimize space usage or enhance the performance of a user device. A processor can also be implemented as a combination of computing processing units. In the present disclosure, terms such as "storage", "database", and substantially any other information storage component related to the operation and function of a component are used to refer to "memory components", entities embodied in "memory", or components that include memory. It should be understood that the memory and / or memory components described herein can be volatile memory or non-volatile memory, or can include both volatile and non-volatile memory. By way of illustration and not limitation, non-volatile memory can include ROM, PROM, EPROM, EEPROM, flash memory, or non-volatile RAM (e.g., FeRAM), and volatile memory can include RAM, for example, which can serve as an external cache memory. Additionally, the memory components of the systems disclosed herein or computer-implemented methods are intended to include but not limited to including these and any other suitable types of memory.
[0095] The foregoing description only includes examples of systems and computer-implemented methods. Of course, it is not possible to describe every conceivable combination of components or computer-implemented methods for the purpose of describing the present disclosure, but one of ordinary skill in the art will recognize that many further combinations and permutations of the present disclosure are possible. In addition, insofar as the terms "including", "has", "possesses", etc. are used in the detailed description, claims, appendices, and drawings, these terms are intended to be inclusive in a manner similar to the way the term "comprising" is interpreted when used as a transitional word in the claims. The description of the various embodiments has been presented for purposes of illustration, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to one of ordinary skill in the art without departing from the scope of the invention. The terms used herein were chosen to best explain the principles of the embodiments, the practical application, or the improvement of technologies found in the marketplace, or to enable other ordinary skilled artisans in the art to understand the embodiments of the invention disclosed herein.
Claims
1. A system, comprising: a memory that stores computer-executable components; and a processor that executes the computer-executable components stored in the memory, wherein the computer-executable components include: a replication component that emulates a controlled NOT gate during qubit reordering; a parallelization component that performs memory access balancing based on the controlled NOT gate emulated by the replication component during the qubit reordering; and a selector component that selects a first bit and a second bit for the qubit reordering, wherein the first bit is a control bit, and wherein the selector component selects the first bit based on: determining that the first bit is a bit that is equal to or higher than the sum of the binary logarithm log2 of the number of qubits minus the number of non-uniform memory access nodes; or determining that the first bit is a bit that is less than the binary logarithm log2 of the number of non-uniform memory access nodes; or determining that the second bit is different from the first bit and is not a target bit.
2. The system of claim 1, wherein the computer-executable components further include: an arrangement component that implements the qubit reordering for quantum computing, wherein the arrangement component reorders the first bit with the second bit.
3. The system of claim 1, wherein the computer-executable components further include: a restoration component that restores the qubit reordering based on an evaluation component that determines a lack of memory access improvement.
4. The system of claim 1, wherein the qubit reordering minimizes fragmented access to quantum memory.
5. The system of claim 1, wherein the qubit reordering minimizes thread locality.
6. The system of claim 1, wherein the qubit reordering is a first qubit reordering, and wherein the replication component emulates the controlled NOT gate during a second qubit reordering based on a determination of success of the memory access balancing made by an evaluation component.
7. A computer-implemented method, comprising: emulating a controlled NOT gate during qubit reordering by a system operatively coupled to a processor; and performing, by the system, memory access balancing based on the emulation of the controlled NOT gate during the qubit reordering; and selecting, by the system, a first bit and a second bit for the qubit reordering, wherein the first bit is a control bit, and wherein the selection includes: selecting the first bit based on determining that the first bit is a bit that is equal to or higher than the sum of the binary logarithm log2 of the number of qubits minus the number of non-uniform memory access nodes; or selecting the first bit based on determining that the first bit is a bit that is less than the binary logarithm log2 of the number of non-uniform memory access nodes; or selecting the second bit based on determining that the second bit is different from the first bit and is not a target bit.
8. The computer-implemented method of claim 7, further comprising restoring, by the system, the qubit reordering based on determining a lack of memory access improvement.
9. The computer-implemented method according to claim 7, wherein the qubit reordering is a first qubit reordering, and wherein the computer-implemented method further comprises: determining, by the system, that the memory access balance is successful; and simulating, by the system, the controlled NOT gate during a second qubit reordering based on the determination.
10. A method comprising: selecting, by a system operatively coupled to a processor, a first qubit and a second qubit, wherein the first qubit is a control qubit; and reordering, by the system, the first qubit with the second qubit, wherein a controlled NOT gate is simulated during the reordering; wherein reordering the first qubit with the second qubit comprises minimizing, by the system, the occurrence of one or more control qubits that are equal to or higher than the binary logarithm log2 of the first number of qubits minus the second number of non-uniform memory access nodes; or reordering the first qubit with the second qubit comprises minimizing, by the system, the occurrence of one or more control qubits that are lower than the binary logarithm log2 of the number of non-uniform memory access nodes.
11. A computer program product that facilitates quantum computing simulation of a controlled NOT gate, the computer program product comprising a computer-readable storage medium having program instructions embodied therein, the program instructions executable by a processor to cause the processor to perform the steps of the method according to any one of claims 7 to 10.