Noise quantum channel error mitigation
Patent Information
- Application Number
- CN202610310408.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2026-02-12
- Filing Date
- 2026-03-13
- Publication Date
- 2026-09-15
Smart Images

Figure CN122764367A_ABST
Abstract
Description
[0001] Cross-references to related applications This application is a continuation-in-part of U.S. Patent Application No. 19 / 295,997, filed August 11, 2025, which claims the benefit of U.S. Provisional Patent Application No. 63 / 771,789, filed March 14, 2025, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates generally to quantum computing, and more specifically to the mitigation (e.g., error correction) of noisy quantum channel errors. Background Technology
[0003] Quantum metrology and computing networks can utilize the laws of quantum mechanics, such as superposition and entanglement, to perform measurements on physical systems, process the results of those measurements, and execute computations. In quantum computing, quantum circuits can include hardware and / or software for quantum computing, where quantum computing is performed using a series of quantum (logic) gates, quantum measurements, and so on. Attached Figure Description
[0004] Figure 1A-1C The illustration shows a schematic diagram of an example system for mitigating noise quantum channel errors, according to at least one embodiment. Figure 2 The illustration is a schematic diagram of an example method for mitigating errors in a noisy quantum channel using weak measurements, according to at least one embodiment. Figure 3 The illustration shows an example machine learning model architecture that can be used to achieve noise quantum channel error mitigation, according to at least one embodiment. Figures 4A-4B This is an example algorithm according to at least one embodiment that can be used to implement a forward diffusion process, which can be used to train a machine learning model to achieve noise quantum data error mitigation; Figure 5 A graph illustrating the relationship between average fidelity and the number of training instances, according to at least one embodiment; Figures 6A-6C The flowchart illustrates an example method for training a machine learning model to achieve noise quantum channel error mitigation, according to at least one embodiment. Figure 7 The flowchart illustrates an example method for mitigating noisy quantum channel errors using a trained machine learning model, according to at least one embodiment. Figure 8 The figure illustrates the test set fidelity distribution according to at least one embodiment; Figure 9AThe illustration shows a tensor network representation of a separable quantum register according to at least one embodiment, wherein the state of independently parameterized subsystems is propagated through local tensor operations without coupling between subsystems; Figure 9B The illustration shows a tensor network architecture according to at least one embodiment, wherein an initially separable local tensor state is transformed into a non-separable tensor by applying a randomly selected controlled NOT gate operation between local tensor quantum system pairs; Figure 10 The illustration shows an example visual transformer architecture that can be used to implement noise quantum data error mitigation according to at least one embodiment; Figure 11A-11B The diagram illustrates a network architecture according to at least one embodiment; Figure 12 An example data center according to at least one embodiment is illustrated; and Figures 13A-13B An example data center architecture according to at least one embodiment is described. Detailed Implementation
[0005] The embodiments described herein relate to techniques for mitigating (e.g., error correction) errors in noisy quantum channels. Quantum computers have entered the era of Noisy Medium-Scale Quantum (NISQ), characterized by the use of a finite number of qubits in quantum devices. Variational quantum algorithms (VQA) have become the dominant paradigm for demonstrating quantum advantage on NISQ devices. VQA operates within a hybrid quantum-classical framework, synergistically combining the advantages of both classical and quantum computing paradigms. The core structure of VQA involves encoding classical data into a parameterized cost function, which is then evaluated using a quantum system (e.g., a quantum simulator or actual hardware). Subsequently, a classical optimizer iteratively refines the parameters of the quantum circuit, aiming to minimize the cost function. This iterative feedback loop between the quantum and classical components improves model accuracy and mitigates the effects of inherent noise in NISQ devices. Hybrid quantum-classical frameworks, such as VQA, cleverly combine the advantages of classical computing (e.g., its robustness in optimization, control, and error mitigation) with the unique computational power of quantum systems, making them particularly suitable for specific subroutines that are difficult to handle using classical methods. This pragmatic approach acknowledges the limitations of current quantum hardware, such as noise and the finite number of qubits.
[0006] Quantum circuits can be designed such that the horizontal axis represents time (usually from left to right), a single horizontal line represents a qubit, and two horizontal lines represent classical bits. The objects connected by these lines define the operations performed on the qubits (such as measurement or gate operations). A quantum gate is a quantum circuit that operates on a certain number of qubits. A quantum gate is the basic building block of a quantum circuit, similar to a classical logic gate in a classical circuit. A quantum gate is an operator that can be represented by a matrix corresponding to an orthogonal basis. For example, a quantum gate can be a unitary operator and can be represented by a unitary matrix. Examples of quantum gates include identity gates, Pauli gates (e.g., Pauli-X, Pauli-Y, and Pauli-Z gates), controlled gates, Hadamard gates, phase-shift gates, commutation gates, Tovey gates, and so on. Quantum circuits can be based on a set of parameter values (e.g., a parameter value vector). To construct, thereby generating wave functions (e.g., For example, this set of parameter values can include a set of rotation angles used to define the rotation of a qubit about an axis (e.g., the X, Y, and Z axes). More specifically, a set of parameter values can define the parameters of a quantum gate used to construct a parameterized gate for a quantum circuit. For example, a n The wave function of a quantum bit can be expressed as ,in Represents the "0" state. It is a unitary operator, indicating that by Parameterized quantum circuits or quantum gate sequences, symbols This represents the tensor product. The other state can be a "1" state, represented by... express.
[0007] In classical information theory, a "classical communication channel" (or "classical channel") refers to a medium that transmits classical information from a sender to a receiver. In quantum information theory, a "quantum communication channel" (or simply "quantum channel") refers to a communication path capable of transmitting quantum information, which is essentially encoded in the quantum states of a quantum system (e.g., in qubits). The evolution of quantum states in a quantum channel can be described by mapping from one quantum state to another. Conceptually, a quantum channel can be viewed as a path designed specifically for reliable transmission between locations. For example, a quantum channel can be used to transmit encoded information about the quantum states of a quantum system (e.g., via qubits). More precisely, a quantum channel serves as a framework for describing how the quantum states of a system evolve over time, including the effects of noise and decoherence, which pose challenges to practical quantum computing. The behavior and properties of quantum channels can be determined by principles of quantum mechanics, such as the superposition of quantum states and quantum entanglement between multiple quantum systems. These principles enable quantum channels to achieve unique and powerful functions that are impossible using classical communication channels or computational methods. Quantum channels can be used to implement quantum communication protocols. For example, quantum channels can support quantum teleportation, enabling the transfer of any quantum state from one physical location to another, even over vast distances, without the need for a physical transfer of the particle or qubit carrying that state.
[0008] In a general sense, the discrete-time changes of the state of any quantum system that conforms to the fundamental principles of quantum mechanics can be fully described by a quantum channel. Specifically, a quantum channel can be mathematically represented as a linear transformation from the input density matrix representing at least one quantum state of the quantum system to the output density matrix representing the evolution of that at least one quantum state. For example, a quantum channel can be represented using a combination of linear operators in an operator summation framework, also known as Kraus operators. The minimum number of Kraus operators used to represent a quantum channel is defined as its Kraus rank. A quantum channel that can be represented by only one Kraus operator (i.e., a quantum channel with a Kraus rank of 1) is called a pure quantum channel. In this case, the evolution of the quantum state can be described by a single deterministic operation (e.g., a unitary transformation) without any probabilistic mixing or noise. The operator summation framework can be used to capture the complex effects of noise, energy dissipation, and other interactions that a quantum state may experience in a quantum channel. To maintain consistency with the principles of quantum mechanics, a quantum channel can be defined using a perfectly positive (CP) and trace-preserving (TP) linear mapping (CPTP linear mapping). The CP property means that the transformation not only maps any valid quantum state to another valid quantum state, but also ensures that this mapping holds even if the quantum system is considered as part of a larger system (e.g., a subsystem of a larger quantum entangled system). The TP property ensures that the total probability associated with any quantum state remains unchanged after the transformation (i.e., the probability is always 1). Therefore, the CP property guarantees that the output of the channel is always a physically possible quantum state, while the TP property guarantees that the probability is neither lost nor increased as the quantum state evolves through the quantum channel.
[0009] An example of a quantum channel is the unitary quantum channel, which reflects the process of applying quantum gates to a quantum state. Mathematically, if a quantum state is represented by a density matrix... ρ This indicates that quantum gates use unitary operators. U This means that the quantum state after passing through the unitary quantum channel is determined by... ρ ' = UρU † This process is reversible and represents the noiseless evolution of a quantum state without any interaction with its environment. A unitary quantum channel characterizes the time evolution of a closed quantum system where there is no exchange of energy or information with its surroundings and serves as an ideal building block for quantum computing. For example, a perfect quantum channel capable of transmitting any input quantum state without alteration can be represented by a unitary channel, where the unitary operator... U It is the unit operator ( I ), making ρ ' = IρI † = ρ .
[0010] Quantum channels can be used as a transmission medium in quantum computing systems, such as connecting different quantum computing devices, different components of a quantum system, and so on. Quantum channels can transmit quantum information (such as qubits) as well as classical information from one location to another.
[0011] Certain quantum channels associated with the quantum circuitry of a quantum computer may be noisy. For example, data transmitted through a quantum channel may experience dissipation or other types of noise, making the quantum channel a so-called "noisy" quantum channel. Noisy quantum channels are typically modeled as non-unitary channels, meaning they cannot maintain the original quantum state data (e.g., density matrix) transmitted from sender to receiver. Instead, noise alters the quantum state during transmission, leading to errors. Noise refers to any interaction with the environment that causes a quantum state to deviate from its expected evolution. This includes factors such as decoherence (loss of quantum information), imperfect quantum gate control, or interactions with other quantum systems. Noise in a quantum channel can originate from a variety of sources, including imperfect control operations, environmental disturbances, and hardware limitations, all of which can significantly reduce the accuracy of quantum computing. For example, noise can cause errors in measurements of observables (e.g., energy levels) of a quantum system, determined by the quantum circuitry running on the quantum computer. Even small environmental disturbances—such as stray magnetic or electric fields, thermal fluctuations, or unexpected interactions with other particles or qubits—can significantly perturb the quantum state.
[0012] Noisy quantum channels can exert asymmetric effects on quantum systems, for example, through bit-flip noise, phase-flip noise, amplitude-damped noise, and other asymmetric noise processes. However, a series of such random asymmetric noise processes can accumulate to produce an effective symmetric noise model—the depolarization channel. The depolarization channel can be considered universal because it erases quantum information regardless of the input, causing the quantum state to tend towards a maximally mixed state. Due to this symmetric decoherence dynamics, the resulting process can be modeled as a quantum diffusion model in some embodiments. In at least one embodiment, this symmetric noise model also appears in quantum state reconstruction, including quantum tomography and classical shadowing. In these cases, measurement processes (e.g., positive operator value measure (POVM)) and device defects can lead to decoherence. Furthermore, random measurements and rotation operations often approximate depolarization of the effective noise. In at least one embodiment, it becomes challenging to identify it accurately enough for inversion when the effective noise is only approximately depolarized—i.e., when the real channel deviates from the ideal depolarization mapping. In this case, channel-independent or robust mitigation methods may outperform exact inversion.
[0013] When noise is present, it can lead to information loss, erroneous manipulation of qubits during quantum logic operations (e.g., accidental rotation or incorrect gate application), and other operational errors. As the scale and complexity of quantum computers increase, so too does the likelihood and severity of these errors, due to the increased number of qubits and operations involved. For example, noise can gradually degrade or erase quantum information (e.g., information encoded in qubits), resulting in a loss of quantum state coherence and fidelity. This corruption of quantum data leads to the accumulation of errors in quantum computing and communication. Over time, these accumulated errors can cause the transmitted quantum state to deviate completely from its original state (e.g., through decoherence). This presents a significant challenge to achieving reliable quantum computing and highlights the importance of effective error mitigation (e.g., error correction) techniques to address errors caused by transmission through noisy quantum channels.
[0014] The embodiments described herein can map noisy quantum state information to the original noise-free quantum state by training a machine learning model, thereby mitigating noisy quantum channel errors and addressing these and other defects. The embodiments described herein also relate to long-distance transmission of optical qubits, i.e., purifying the quantum state before it enters the next quantum computing system.
[0015] The physical implementation of quantum channel 115A may involve different ambient temperature ranges, each characterized by different noise profiles. While optical qubits can operate in room-temperature links, qubits at cryogenic temperatures face unique challenges when directly coupled to system interconnects such as NVIDIA® NVLink™. Cryogenic quantum systems typically operate at temperatures near absolute zero (e.g., millikelvin ranges for superconducting qubits), and direct coupling to room-temperature interconnects introduces significant thermal gradients, leading to decoherence and information loss. The interface between cryogenic quantum hardware and room-temperature interconnect infrastructure represents a critical junction where quantum state degradation can occur due to thermal noise injection, impedance mismatch, and electromagnetic interference from warmer environments. Transferring from a cryogenic source through a room-temperature medium can inject additional thermal noise, thus requiring more repetition and intensive cleanup in machine learning models to recover the original quantum information. In at least one embodiment, the interconnect (e.g., NVLink) serves as both the transmission medium and the location for error mitigation operations. Instead of acting as a passive channel, the interconnect infrastructure can integrate trained machine learning models to identify the types of noise introduced during the cryogenic-to-room-temperature transition and simulate the inverse quantum channel in situ. This integration enables the interconnect hardware to autonomously restore the integrity of the transmitted quantum message before it reaches the computation layer of the receiving processing node, thus effectively performing quantum state cleaning at the interconnect layer. Therefore, NVLink or similar high-bandwidth interconnects can serve a dual purpose: facilitating high-speed data transmission between quantum and classical processing components while simultaneously implementing real-time error mitigation protocols to compensate for thermal noise accumulated during computation performed by and from cryogenic quantum systems.
[0016] This technique utilizes weak measurements (e.g., weak measurements in the middle of a circuit) to simulate noisy quantum channels, thereby generating training data for training a machine learning model to perform noisy quantum channel error mitigation. A “weak” measurement is one that only slightly perturbs the quantum system, as opposed to a “strong” measurement that completely perturbs the quantum system (e.g., wavefunction collapse). Weak measurements can be used to create any type of noisy quantum channel. This can be achieved by performing multiple weak measurements to create an artificial noisy quantum channel. The measurement intensity can be adjusted, allowing the measurement of any type of property and the creation of any type of noisy channel. The trained machine learning model can then be used to perform error mitigation. The machine learning model described herein can approximate an inverse quantum channel by recovering the original quantum state by removing the noise introduced by the noisy quantum channel. The machine learning model described herein can be used to identify at least one type of noise added through at least one non-unitary quantum channel. In at least one embodiment, the machine learning model is a diffusion model. References will follow below. Figure 1A-13BFurther details on achieving error mitigation in noisy quantum channels are described.
[0017] The trained machine learning model can be deployed in practical quantum computing systems or experimental setups to execute error mitigation protocols (e.g., in-situ front error mitigation protocols), thereby providing more accurate and reliable data for model training. In at least one embodiment, the training process utilizes both simulated and non-simulated (experimental) data simultaneously, enabling the machine learning model to learn from real-world noise characteristics and operational defects. This approach can produce more robust and effective error mitigation protocols, better suited to practical quantum computing environments.
[0018] The embodiments described herein can be used in a variety of quantum computing applications. Examples of quantum computing applications include error mitigation, secure quantum communication, quantum teleportation, quantum key distribution, quantum sensing and metrology, quantum simulation, and distributed quantum computing. The output of the machine learning models described herein (including output data indicating the original quantum state recovered from noisy input data) can be used in each of the above applications to improve performance and reliability.
[0019] For example, in error mitigation applications, the output of a machine learning model (representing the denoised quantum state) can be used to improve tolerance to various types of noise introduced by non-unitary quantum channels, such as depolarization noise, amplitude damping noise, and phase decay noise. The recovered quantum state output can replace the damaged quantum state in subsequent quantum computations, thereby reducing the accumulation of errors during the execution of quantum algorithms.
[0020] In secure quantum communication applications, the output of machine learning models can be used to reconstruct the original quantum state of the transmitted qubits, thereby improving the fidelity of the transmitted quantum state between the communicating parties and enhancing the reliability of quantum cryptography protocols. The denoised quantum state output, by correcting channel-induced errors before key extraction, enables more robust long-distance quantum key distribution.
[0021] In quantum teleportation applications, the output of a machine learning model can be used to correct the received quantum state after teleportation, improving the accuracy of quantum state transfer between spatially separated quantum systems by eliminating noise accumulated during the teleportation process. The corrected quantum state output can then be used for subsequent quantum operations at the receiving node.
[0022] In quantum sensing and metrology applications, the output of machine learning models can be used to eliminate errors caused by noise in the output of quantum sensors. The denoised quantum state output provides a more accurate representation of the measured physical quantity.
[0023] In quantum simulation applications, the output of machine learning models can be used to correct the quantum states of the simulated system. By correcting hardware-induced errors, the accuracy of the simulated quantum system is improved, enabling more reliable modeling of complex molecular structures, material properties, and chemical reactions. The error-mitigated quantum state output can be used to extract more accurate expected values and observables from the simulation. In distributed quantum computing applications, the output of machine learning models can be used to recover the original quantum states transmitted between quantum processing nodes connected by noisy quantum channels, thereby enabling reliable communication and allowing the construction of larger-scale quantum computing systems using smaller interconnected quantum processors. The denoised quantum states output at each receiving node ensure that subsequent distributed quantum operations can be performed using high-fidelity quantum data.
[0024] The systems and methods described herein can be used for a wide range of purposes, including, but not limited to, machine control, machine motion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, safety and supervision, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or role simulation and / or digital twins, data center processing, conversational AI, optical transmission simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation of 3D assets, generative AI operations using (e.g., large-scale) language models, cloud computing, and / or any other suitable applications. The systems and methods described herein can also be used in systems for compiling quantum circuits, systems for executing quantum circuits, systems for measuring quantum states, and / or systems for measuring the states of one or more qubits.
[0025] The disclosed embodiments may be included in a variety of different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, marine systems, smart area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems containing one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in a data center, systems for performing conversational AI operations, systems for performing generative AI operations using (large) language models, systems for performing optical transmission simulations, systems for performing collaborative content creation of 3D assets, systems implemented at least partially using cloud computing resources, and / or other types of systems.
[0026] In at least one embodiment, the resulting machine learning model can be integrated into a hybrid quantum-classical workflow. For example, the machine learning model can be integrated with a quantum computing system via a communication interface (e.g., a QLink-type interface). In addition to error mitigation, the machine learning model can also play a role in state preparation (e.g., generating a target quantum state from noise). Applications can include embedding these models more efficiently into quantum algorithms, such as variational quantum linear solvers (VQLS). In at least one embodiment, the machine learning model is configured to efficiently manage resources in the presence of measurements in the circuit. Integrating the error mitigation model into the hybrid quantum-classical workflow enables more reliable quantum computing by correcting errors introduced during the execution of the quantum algorithm. In at least one embodiment, classical computing resources are post-processed to stabilize the output of the quantum algorithm, thereby supporting the quantum algorithm. Although the target evolution in the quantum algorithm is unitary, noise effectively turns it into a general noise channel. The machine learning model described herein can help mitigate this noise, enabling hybrid quantum-classical systems to achieve higher fidelity results.
[0027] In at least one embodiment, the improved error mitigation described herein reduces the number of sampling runs required by the quantum computer. Since sampling requirements depend on the purity of the quantum circuit, noisier circuits require more runs to obtain statistically accurate results. The performance of error mitigation can be improved by training machine learning models using the protocols described herein (e.g., segmenting communication lines into small segments where the noise in each segment resembles a weak measurement). Therefore, the number of processor cycles required by the quantum computer can be reduced, which can translate to lower power consumption, higher bandwidth, or enabling the qubits to be used for other tasks. Furthermore, effective error mitigation measures can reduce or eliminate the need for other error correction protocols that require additional qubits, which may be unavailable in resource-constrained quantum systems.
[0028] In at least one embodiment, the techniques described herein offer several technical advantages. In at least one embodiment, training a machine learning model to simulate the inverse of a non-unitary quantum channel can improve error mitigation performance by enabling the recovery of noisy quantum states without requiring precise knowledge of the noisy channel. Improved error mitigation can reduce the number of samples required by a quantum computer, thereby reducing processor cycles, lowering power consumption, and increasing bandwidth, or enabling qubits to be used for other tasks. Effective error mitigation can also reduce or eliminate the need for additional error correction protocols that require additional qubits, which may be unavailable in resource-constrained quantum systems. In at least one embodiment, segmenting the communication line into small segments such that the noise in each segment approximates a weak measurement can improve error mitigation performance. In at least one embodiment, when the effective noise only approximately matches the depolarized channel, channel-independent mitigation methods can be used, thus avoiding the need for accurate channel characterization for direct inversion. In at least one embodiment, attention mechanisms can capture non-local quantum correlations, such as entanglement between distant qubits, which cannot be adequately captured by pure convolution operations with spatially limited receptive fields. In at least one embodiment, applying the attention mechanism to the bottleneck of reduced resolution rather than at full input resolution can reduce the secondary computational cost of attention while still enabling the model to learn global dependencies. In at least one embodiment, the techniques described herein can be extended to different quantum register types, including separable quantum registers where features scale linearly with the number of qubits and entangled quantum registers where features scale exponentially with the number of qubits, wherein attention-based models are more efficient at handling entangled quantum registers than recurrent neural network models. In at least one embodiment, the machine learning model can be trained using only the local density matrix to reconstruct the global quantum state, achieving reconstruction from limited input information. In at least one embodiment, average pooling can preserve phase information in the quantum state representation by retaining the aggregated signal energy of all features (rather than selectively retaining only the maximum activation value). In at least one embodiment, blind denoising can be performed without explicit noise level conditions, thereby achieving generalization to different noise levels. In at least one embodiment, training stability can be improved through teacher coercion (gradually transitioning to autoregressive generation), mixed-precision computation, exponential moving average of model weights, and gradient pruning. In at least one embodiment, the trained machine learning model can be used for error mitigation and state preparation, thus providing flexibility for quantum computing applications. In at least one embodiment, the techniques described herein can be integrated with hybrid quantum-classical workflows, including variable quantum algorithms such as VQLS, and can efficiently manage resources in the presence of intermediate circuit measurements. In at least one embodiment, tensor network representations enable efficient implementations of graphics processing units (GPUs), thereby rapidly generating training data.
[0029] In the following and foregoing description, numerous specific details will be set forth in order to provide a more complete understanding of at least one embodiment. However, those skilled in the art will understand that the inventive concept can be practiced without the foregoing one or more specific details. The following figures illustrate example systems and methods for collective communication protocols for quantum metrology, but are not limited thereto.
[0030] Figure 1A This is a block diagram of a system 100A according to at least one embodiment. (e.g.) Figure 1A As shown, system 100A may include quantum computing device 110A which is communicatively coupled to quantum computing device 120A via quantum channel 115A.
[0031] In at least one embodiment, quantum channel 115A corresponds to an interconnect. For example, the interconnect can enable communication between one or more processing units of quantum computing device 110A and one or more processing units of quantum computing device 120A. The interconnect can be physically reconfigured (e.g., rewiring) to apply transformations without active measurement. In at least one embodiment, the interconnect or quantum channel 115A can be physically reconfigured to manipulate quantum state trajectories. For example, the trajectory of one or more poles in a quantum state can be transformed by physically modifying the interconnect (e.g., “rewiring the cable”) without active measurement. This hardware-level reconfiguration allows the application of specific transformations or noise suppression protocols by physically altering the transmission path. Interconnects (e.g., NVLink) and network interface cards (NICs) play a primary functional role in error correction protocols; they are both the medium for quantum state transmission and the site where “cleaning” occurs. Instead of acting as passive channels, NICs and system-level interconnects can integrate trained machine learning models to identify the types of noise added during transmission and simulate in-situ inverse quantum channels. This integration enables the interconnect hardware to autonomously restore the integrity of the transmitted quantum message before it reaches the computing layer of the receiving processing node.
[0032] Quantum computing device 110A can transmit the quantum state 112A of a quantum system to quantum computing device 120A via quantum channel 115A. More specifically, quantum state 112A can represent the primitive quantum state of the quantum system. In at least one embodiment, quantum channel 115A is a noisy quantum channel that introduces noise into quantum state 112A. More specifically, the noisy quantum channel can be any suitable non-unitary quantum channel in the embodiments described herein. An example of a non-unitary quantum channel is a depolarization channel. Depolarization channels can be used to simulate depolarization noise. For example, quantum information units (e.g., qubits) are probabilistically... p Maintain its quantum state and with probability 1- pAll quantum information is lost. In the latter case, quantum state 112A is transformed into a maximally mixed state, which is determined by the unit operator. I Divide by the dimension of the Hilbert space associated with the quantum system (e.g., for a qubit, ). I / 2).
[0033] Another example of a non-unitary quantum channel is the amplitude-damped channel. Amplitude-damped channels can be used to simulate the process by which quantum information units (such as qubits) lose energy to their environment with a certain amplitude-damped probability, which may lead to transitions from excited states to the ground state.
[0034] Another example of a non-unitary quantum channel is the phase-damped channel. Phase-damped channels can be used to simulate the loss of quantum information due to random phase fluctuations without any energy exchange with the environment.
[0035] Another example of a non-unitary quantum channel is the bit-flipping channel. A bit-flipping channel can be used to... p The state of the qubit from Flip to Conversely, the same applies. The following text will refer to... Figure 1B The error mitigation system that achieves noise quantum channel error mitigation can be used to recover the initial or original quantum state affected by noise caused by quantum channel 115A.
[0036] Figure 1B A block diagram of a system 100B for mitigating noise quantum channel errors is illustrated according to at least one embodiment. Figure 1B As shown, in at least one embodiment, system 100B includes computing device 110B, which includes an error mitigation system 120B for mitigating noisy quantum channel errors. In at least one embodiment, computing device 120B is a classical computing device. In at least one embodiment, computing device 120B is a quantum computing device. In other embodiments, at least one component of error mitigation system 120B is located in a separate computing device or system.
[0037] The error mitigation system 120B can perform noisy quantum channel error mitigation. Specifically, as described herein, noisy quantum channel error mitigation can be achieved by training a machine learning model to map noisy quantum information to its corresponding ideal (noise-free) quantum state.
[0038] For example, error mitigation system 120B may include training component 122B. Training component 122B can train a machine learning model during the training phase to approximate the inverse of a noisy quantum channel, thereby eliminating the noise introduced by the noisy quantum channel. In at least one embodiment, the training process utilizes both simulated and non-simulated (experimental) data, enabling the machine learning model to learn from real-world noise characteristics and operational defects. This approach can produce a more robust and effective error mitigation protocol, better suited to practical quantum computing environments. The trained machine learning model can be deployed in practical quantum computing systems, hybrid computing systems including both quantum computing components and conventional computing components, and / or experimental setups to execute the error mitigation protocol in situ. Alternatively or additionally, the error mitigation system may also include inference component 124B. Inference component 124B can use the trained machine learning model during the inference phase to recover quantum state 112A. Therefore, error mitigation system 120B can receive input data generated by processing one or more physical signals (e.g., optical signals, electrical signals, and / or microwave signals).
[0039] Input data can indicate quantum state 112A (sent by quantum computing device 110), which is modified by the amount of noise added by quantum channel 115A. Error mitigation system 120B can provide the input data to a machine learning model trained by training component 122B, which predicts quantum state 112A by simulating the inverse of quantum channel 115A, and obtain output data indicating quantum state A from the machine learning model via inference component 124B.
[0040] In at least one embodiment, the machine learning model is a generative model trained to predict and reconstruct the original quantum state (e.g., quantum state 112A) based on input data, thereby generating the original quantum state as output. For example, the machine learning model could be a diffusion model, a generative model designed to iteratively denoise the input data. Typically, diffusion models are trained to generate output data by learning to remove noise systematically added to the input data. This makes diffusion models particularly suitable for recovering damaged or noisy data, such as quantum state data affected by noise due to transmission through noisy quantum channels.
[0041] If the machine learning model is a diffusion model, training the machine learning model using training component 122B during the training phase typically involves two main processes: a forward diffusion process and a backward diffusion process. In the forward diffusion process, random noise is progressively added to the data over a series of discrete time steps, causing the data to transition from its original state to a state of almost pure noise. Each step in the forward diffusion process depends on the previous step. For example, the forward diffusion process can be mathematically modeled as a Markov chain, where the probability distribution of the data at each step depends only on the immediately preceding step. The diffusion model then learns the backward diffusion process during training. In this process, the model is trained to progressively reverse the effects of the forward diffusion process by predicting and removing the noise added at each step. The goal of the backward diffusion process is to enable the model to reconstruct noise-free data from noisy input. Using the trained machine learning model during the inference phase using inference component 124B can include: receiving input data representing noisy or damaged quantum states (even pure random noise), and iteratively applying the trained backward diffusion process to generate output data that approximates the original (denoised) quantum state. Therefore, the output quantum state data generated by the trained diffusion model corresponds to a predicted, noise-corrected version of the input quantum state data, thus enabling a practical implementation scheme to effectively mitigate quantum data errors.
[0042] In at least one embodiment, the machine learning model (e.g., a diffusion model) is implemented using a recurrent neural network (RNN)-based architecture. An RNN is a type of neural network designed to process sequential data by maintaining internal hidden states that capture information from previous time steps. This makes them well-suited for tasks involving time or sequence dependencies, such as time series prediction, speech recognition, and video analysis. In the context of quantum systems, RNNs are particularly advantageous in modeling noisy quantum channels that can be represented as a series of weak measurements, where each measurement provides partial information about the quantum state.
[0043] Long Short-Term Memory (LSTM) networks are a type of RNN that can be used. LSTMs are designed to capture short-term and long-term dependencies in sequential data through a system of gates (called input gates, forget gates, and output gates) and unit states that act as memory. These gates regulate the flow of information, allowing the network to selectively retain or discard information over time. LSTMs can consist of a main layer and hidden layers: the main layer is responsible for capturing long-term dependencies (e.g., relationships between initial quantum states in a sequence), while the hidden layers capture short-term dependencies (e.g., changes between consecutive time steps or measurements). In at least one embodiment, the LSTM architecture employs a multi-layer encoder followed by a residual feedforward network, where the hidden layers propagate through time steps to capture temporal dependencies, thereby enabling autoregressive predictions. Within this framework, the short-term dependencies modeled by the hidden layers can be used to represent the dissipation or evolution of quantum states over a series of weak measurements. The long-term dependencies captured by the main layers can be used to correlate these sequences with a reference scenario, such as the evolution of quantum states towards a maximally mixed state (i.e., a state in which the quantum system is a completely random mixture of all possible ground states). The maximally mixed state has maximum entropy and retains no information about the initial quantum state. This two-layer modeling allows LSTM to simultaneously learn the direct effects of a single weak measurement and the overall trajectory of the quantum state, thus providing a comprehensive approach to reconstructing or denoising quantum states from weak measurement data. The following will refer to... Figure 3 Describe an example RNN (e.g., LSTM) that can be used to implement a machine learning model for mitigating errors in noisy quantum channels.
[0044] One technical challenge faced by the error mitigation system 120B when using machine learning models (e.g., diffusion models) to mitigate noisy quantum channel errors is the potential lack of sufficient and representative training data required for effective training. To address this, the error mitigation system 120B can utilize a synthetic training data generation process specifically tailored for training machine learning models to perform noisy quantum channel error mitigation. This process may include generating multiple random pure quantum states, for example, by applying a random unitary matrix to a fixed reference state (e.g., the ground state). Each generated pure state can then be transmitted through a randomly selected quantum channel, characterized by a specific type of quantum channel (e.g., a depolarization channel, amplitude-damped channel, or phase-damped channel). The transmitted output (or the generated noisy quantum state) is recorded along with the corresponding original pure state. This pair of data (including the noisy output and the known original input) serves as training instances. By repeating this process with a sufficiently large number of random pure states and random quantum channels, a comprehensive training dataset can be constructed. The dataset can then be used to train machine learning models (e.g., diffusion models) to learn the mapping from noisy quantum states to corresponding denoised quantum states, thereby enabling the model to effectively mitigate quantum data errors.
[0045] In at least one embodiment, generating training data includes performing a “weak” measurement on a quantum computing device 110 using a quantum measurement device 130. A weak measurement is a quantum measurement that produces minimal perturbation to a quantum state 112A, such as partially collapsing a wavefunction. This contrasts with a “strong” measurement, which is a quantum measurement that completely collapses a wavefunction and completely perturbs a quantum state 112A. In a weak measurement, the quantum state 112A is only partially projected onto the quantum measurement device 130, meaning that the measurement provides limited information about the measured observable. Therefore, the quantum state 112A after a weak measurement may be partially affected by noise, the degree of which depends on the specific observable being measured and the strength of the measurement interaction. The measured property can include magnetic fields under different basis vectors (X, Y, Z). The introduced noise depends on the measured basis vectors. The effect of the measurement on the quantum state depends on the quantum state itself. For example, a spin Z measurement on a quantum state with spin Z does not introduce noise, while a spin Z measurement on a quantum state with spin X will lose information. Furthermore, weak measurements can produce so-called "outliers," meaning that the measurement result may fall outside the expected eigenvalue range of the measured observable. Additionally, by performing multiple weak measurements consecutively, a quantum state can gradually deviate from its pristine state in an observable manner.
[0046] Due to these properties of weak measurements, they can be used as training data to simulate the effects of noisy quantum channels by introducing controllable incremental noise into quantum computing device 110, while still retaining at least some information about quantum state 112A. Specifically, training component 122B can train a machine learning model (e.g., a diffusion model) to perform error mitigation on noisy quantum channels by removing noise from the data generated by one or more weak measurements. Since a single weak measurement typically provides only limited information about quantum state 112A, multiple weak measurements that may be performed on a single quantum system or a quantum system including the homomorphic preparation of quantum computing device 110 can be combined (e.g., through statistical averaging or data fusion techniques) to extract more accurate and meaningful information about quantum state 112A. Furthermore, weak measurements can be used in conjunction with pre-selection (preparing the quantum system in a known initial state) and post-selection (selecting a specific final state after a weak measurement). Inference component 124B can use such a trained machine learning model to remove noise from the data generated by one or more weak measurements that model noise in the quantum system. For example, a single weak measurement or a series of weak measurements can be fed as input to a trained machine learning model, which then predicts quantum state 112A based on the input and reconstructs quantum state 112A as the output.
[0047] As an illustrative example of training via training component 122B, the machine learning model described herein (e.g., an LSTM model) can be trained with appropriate training data to predict the inverse evolution of each density matrix already used to simulate a noisy quantum channel. By predicting the inverse evolution of each density matrix, the machine learning model can identify the corresponding inverse quantum channel, thereby effectively eliminating the noise introduced by the noisy quantum channel. In the initial phase of training, the machine learning model can be exposed to an unknown quantum channel constructed using a set of random elements (e.g., random positive operator value metric (POVM) elements) applied as weak projections to a randomly selected quantum state. This initial quantum state is designated as a prime example. The control system can repeatedly apply these measurements to the prime example with constant interaction strength, thereby driving its evolution until the quantum state approaches a maximally mixed state. Once the maximally mixed state is reached or very close to is reached, the control system can stop the measurements and record the complete sequence of density matrices representing the state evolution (assuming the density matrices have been reconstructed at each step, e.g., by performing quantum state tomography or other suitable techniques). Following this initial phase, other pure quantum states (correlated with the typical instance via unitary rotations) can be input into the same unknown quantum channel. The evolution of their corresponding density matrices can be monitored and recorded in a similar manner. Unlike the typical instance, these additional quantum states may not necessarily reach a maximally mixed state, as their final state depends on the interaction strength and the number of measurements applied. Machine learning models can be trained to analyze the evolution of these different examples, thereby learning the properties and structure of the underlying quantum channel. As an illustrative example of inference component 124B, a trained machine learning model can be used to identify an inverse quantum channel that can transform the maximally mixed state of the typical instance back to its original pure state. This approach enables the reconstruction or denoising of noisy quantum states. In at least one embodiment, at least two quantum channels are sequentially combined to form a composite quantum channel. This framework can be further generalized to higher-dimensional Hilbert spaces, enabling machine learning models to handle quantum systems with larger state spaces and more complex noisy processes.
[0048] In at least one embodiment, the forward diffusion process uses a Pauli-6 POVM as a set of quantum instruments to create a noisy quantum channel. In a Pauli-6 POVM, each projection has the same measurement strength, but its effect on each density matrix may differ depending on the quantum state. The explicit form of the Pauli-6 POVM can be defined using operator summation. In at least one embodiment, other POVM sets can be used, such as SIC (symmetric information complete) POVM or any other suitable set of measurement operators. The choice of POVM set may affect the characteristics of the resulting noisy quantum channel and the training data generated for machine learning models.
[0049] In at least one embodiment, the machine learning model (e.g., a diffusion model) is implemented using an encoder-decoder architecture with a multi-head attention mechanism. Unlike convolutional neural networks (CNNs), which primarily capture local spatial dependencies through limited receptive fields, or LSTMs, which excel at sequence dependencies but may not efficiently model spatial correlations, attention mechanisms can identify dependencies between arbitrary, distant spatial locations. This capability is particularly advantageous for quantum state representations, as quantum states can exhibit non-local correlations (e.g., entanglement between distant qubits) that cannot be adequately captured by pure convolution operations with spatially limited receptive fields. In at least one embodiment, a multi-head self-attention mechanism is integrated into the bottleneck of the encoder-decoder architecture to capture non-local quantum correlations. This architecture may include three main components: an encoder path that progressively reduces the spatial dimension while increasing the feature channels; an attention-based bottleneck for capturing long-range spatial dependencies; and a decoder path that reconstructs the denoised state at the original resolution. References will follow below. Figure 10 Describe an example visual transformer architecture that can be used to mitigate errors in noisy quantum data.
[0050] In at least one embodiment, the machine learning model is implemented using the U-Net architecture. The encoder may include residual blocks operating at progressively decreasing spatial resolution, with the channel dimension of each stage of the residual block gradually increasing. In at least one embodiment, the downsampling operation may employ average pooling. Average pooling is chosen over max pooling because it preserves the aggregated signal energy of all features in each pooling window, while max pooling may selectively preserve only the maximum activation value. For quantum state representations, phase information can be encoded across multiple spatial features, so preserving all signal contributions may be beneficial for maintaining the coherent structure of the quantum state. In at least one embodiment, the residual blocks employ a pre-activation design using group normalization and the SiLU (Swish) activation function, followed by reflection-filled convolutions. The pre-activation residual design (normalization → activation → convolution) facilitates gradient flow in deep networks by ensuring that gradients propagate along clean additive paths without being modulated by the activation function. At the lowest spatial resolution (bottleneck), a multi-head self-attention mechanism can be integrated, enabling the model to capture global spatial correlations in the quantum state. The decoder may mirror the encoder structure, progressively reconstructing the spatial resolution through upsampling stages.
[0051] Skip connections concatenate features from the corresponding encoder layer with decoder features at each resolution level, preserving high-resolution spatial details and phase information that might otherwise be lost during dimensionality reduction in the encoding process. The concatenated features can be processed by residual blocks that learn to selectively combine information from multiple scales: coarse-grained semantic features from the decoder path and fine-grained details from the skip connections. In at least one embodiment, a hierarchical multi-scale architecture (which simultaneously processes inputs at multiple spatial resolutions) combined with residual connections enables the model to learn adaptive denoising strategies. This adaptability facilitates generalization across different noise levels without explicitly conditionalizing the noise level as input, enabling blind denoising when the noise level is unknown at inference time.
[0052] The complexity of attention operations may be quadratic with respect to spatial resolution, expressed as O(...). N 2 ),in N The number of spatial locations. By applying attention only at the bottleneck of reduced resolution (e.g., 32×32 resolution, 1024 locations) rather than at the full input resolution (e.g., 128×128=16384 locations), the quadratic computation cost of attention can be reduced by 16x, while still enabling the model to learn global dependencies.
[0053] In at least one embodiment, training the machine learning model includes using a teacher-coercive policy that gradually transitions to autoregressive generation. In the early stages of training, the model can receive the ground-truth target from the previous time step as input at each point in the sequence. This can accelerate convergence by exposing the model to ideal local inputs. However, relying on the ground truth during training can lead to a mismatch at inference time, as only the model's own predictions are available at this point. To bridge this gap, linear platform scheduling can be used to reduce the teacher-coercive probability. During training, the model increasingly relies on the outputs of its previous steps, adapting to the final inference scenario. This balance between teacher coercion and autonomous prediction improves the robustness and stability of the trained model.
[0054] In at least one embodiment, the machine learning model is trained by being fed an inverse sequence to perform error mitigation at each time step. The training process may include inverting the simulated sequence to ensure that the ground truth corresponds to a previous (less noisy) state rather than a subsequent (more noisy) state, thereby framing the task as a regression problem, where the model progressively predicts the denoised quantum state in an autoregressive manner.
[0055] In at least one embodiment, training the machine learning model utilizes mixed-precision computation (automatic mixed precision, AMP) to improve computational efficiency, employs exponential moving averages (EMA) of model weights to improve generalization at test time, and utilizes gradient pruning to prevent instability when the model encounters high-amplitude noise inputs in the early training phase. In at least one embodiment, an optimizer with cosine annealing learning rate scheduling (e.g., an Adam optimizer with weight decay or AdamW) is used to optimize the model.
[0056] In at least one embodiment, training the machine learning model includes using a cost function based on mean squared error (MSE) with additional constraints. This cost function may include a trace normalization constraint that enforces Tr( ρ The cost function is set to 1, thus forcing the output density matrix to normalize. This cost function may also include a penalty for non-Hermitian density matrices to account for the physical properties of the generated quantum system. In at least one embodiment, model performance is evaluated using a geometric fidelity metric, which measures the degree of overlap between the predicted and target density matrices. To ensure that the predicted outputs retain valid density matrices, post-processing can be applied to map them back to Hermitian form according to the constraints of the CPTP mapping.
[0057] In at least one embodiment, in addition to error mitigation, the trained machine learning model can also be used for state preparation (e.g., generating a target quantum state from noise). More specifically, after generating sufficient training data through a forward diffusion process, the trained model can serve as either an error mitigation protocol or a state preparation protocol. In state preparation applications, the model can be used to generate a target quantum state by providing appropriate input and allowing the model to generate a target output state through learned inverse dynamics.
[0058] In at least one embodiment, the forward diffusion process uses a Pauli-6 POVM as a set of quantum instruments to create a noisy quantum channel. In a Pauli-6 POVM, each projection has the same measurement strength, but its effect on each density matrix may differ depending on the quantum state. The explicit form of the Pauli-6 POVM can be defined using operator summation. In at least one embodiment, other POVM sets can be used, such as SIC (symmetric information complete) POVM or any other suitable set of measurement operators. The choice of POVM set may affect the characteristics of the resulting noisy quantum channel and the training data generated for machine learning models.
[0059] In at least one embodiment, Bloch vector representation is used to train a machine learning model to mitigate single-qubit errors. In Bloch sphere representation, a single-qubit quantum state can be represented by a Bloch vector with three real-valued components. The shape of the Bloch vector dataset can be [N , S [3], of which N This refers to the batch size (number of instances). S is the sequence length (number of noise steps), and 3 corresponds to the three Bloch vector components, all of which are real numbers. In the Bloch vector representation, the greater the channel noise, the smaller the norm of the Bloch vector; when the quantum state approaches the maximally mixed state, its norm approaches zero. In at least one embodiment, each iteration of the forward diffusion process causes the Bloch vector to undergo an affine transformation, thereby changing the Bloch vector components. The inverse quantum channel can also be represented as an affine map that reverses this transformation.
[0060] Since the forward process is non-unitary, the inverse quantum channel may not have physical meaning because entropy is not reduced. To make the process physically meaningful and unitary in at least one embodiment, an auxiliary quantum system is needed, according to the Stinespring dilation theorem. The Stinespring dilation theorem states that any CPTP mapping can be represented as a unitary operation over a larger Hilbert space including the original system and the auxiliary quantum system, followed by partial trace operations on the auxiliary degrees of freedom. Here, the forward quantum channel can be understood as heating the system according to the second law of thermodynamics, while the inverse process cools the system, similar to a purification process. This quantum channel has a non-physical inverse channel, which may require dilution of the existing subsystem, typically achieved by using auxiliary quantum states (e.g., auxiliary qubits), thus making the inverse process unitary and physically realizable. These auxiliary quantum states are discarded after the process ends. In the context of the above Bloch vector representation, the inverse quantum channel can be represented as an affine mapping. ,in It is an inverse linear transformation matrix. However, such an inverse transformation may amplify the Bloch vector beyond the unit sphere, corresponding to a non-physical quantum state. In at least one embodiment, a machine learning model learns to approximate this non-physical inverse channel through classical post-processing, variational circuitry, or measurement conditional mapping. The inverse process is approximate and can be implemented by: classical post-processing that operates on the measurement results; variational quantum circuitry that parameterizes a series of quantum operations and optimizes the parameters to approximate the inverse dynamics; or measurement conditional mapping that applies different corrections based on intermediate measurement results. This approach progressively converts noisy mixed states into estimates of their pre-noisy pure states without requiring an exact physical realization of the inverse channel, which would otherwise violate the second law of thermodynamics because it would reduce entropy without coupling to an external system. Therefore, in at least one embodiment, at least one processing device couples the quantum system to an auxiliary quantum system to perform unitary operations approximating the inverse of a non-unitary quantum channel and performs partial trace operations on the auxiliary quantum system.
[0061] In at least one embodiment, the training dataset has a specific shape representation that depends on whether the quantum registers are separable or entangled. For separable quantum registers, the dataset shape can be [ N , S , D [,2, 2], where N This refers to the number of instances (batch size). S It is the sequence length (number of noise steps). D is the number of qubits, and the 2×2 dimension corresponds to the real and imaginary parts of the density matrix for each single qubit, respectively. For entangled quantum registers, the dataset shape can be [ N , S , 2 D , 2 D This reflects the fact that the Hilbert space dimension scales exponentially with the number of qubits, resulting in a much larger matrix and including more parameters. In at least one embodiment, the input shape can be [...] when training the local density matrix of the entanglement register. N , S , D [, 2, 2, 2], where the model treats each qubit independently, without explicit correlation, but the initial correlated states of all qubits must be considered together as a single 2 D × 2 D The matrix output demonstrates that the model performs well even with limited input resources.
[0062] In at least one embodiment, quantum states and quantum operations are represented using tensor networks. Tensor networks provide a compact graphical representation of quantum states, where nodes represent tensors and edges represent index contractions, thus explicitly expressing entanglement structures. Tensor network representations enable efficient many-body simulation methods. In at least one embodiment, the input quantum state enters as a tensor and undergoes a series of CPTP mappings, where each mapping is represented as a tensor block acting at a labeled time step. Tensor network representations align well with GPU-based computation and quantum mechanics in many-body systems, enabling efficient parallel processing of quantum state data. Tensor network representations can serve as computational blueprints for GPU implementations, where the tensor network defines the structure of quantum operations and noise channels, while the GPU supports parallel computation across multiple training instances. Essentially, there are two blueprint types corresponding to two types of datasets: separable quantum registers and entangled quantum registers. Parameters can be randomly generated within this blueprint structure, and once the structure is defined, training instances can be generated very efficiently because all operations are integrated onto the GPU via the tensor network. This approach can rapidly generate large training datasets for machine learning models. This article will describe in detail the additional details about tensor networks.
[0063] One phenomenon that can occur with continuous weak measurements is the generation of many mixed components that are projected onto different variants of the measurement projection operator. Therefore, in practice, reaching the maximum mixing state may be impossible. Furthermore, the process can become increasingly difficult as fidelity approaches the maximum mixing state, meaning that the required sequence length becomes very large when low fidelity criteria are required. Therefore, in practical implementations, the fidelity threshold may be limited to a feasible range. The fidelity requirement can be adjusted. For example, depending on the algorithm, the fidelity of the result may be considered pure when it reaches approximately 80%, approximately 90%, or approximately 100%, respectively. For a global state consisting of multiple qubits, the system cannot arbitrarily approach the global maximum mixing state. This practical limitation can affect the design of the training data generation process and the forward diffusion process.
[0064] Figure 1C An example computing system 100C according to at least one embodiment is illustrated, which is capable of supporting noisy quantum channel error mitigation operations. In at least one embodiment, the computing system 100C may include... Figure 1A The quantum computing device 110A and Figure 1B The computing device 120B.
[0065] The 252C quantum bit circuit system may include multiple qubits (three-qubit, quantum digital, etc.), which exist in two states (three-state, three ... d The qubit circuit system 252C can be implemented in the form of physical systems such as states. It may include Josephson junction arrays, trapped ions / atoms, quantum dots, electron spins, nuclear spins, diamond vacancies, coherent optical states, topological electronic states, magnetic flux states, and / or other quantum systems with a target Hilbert space dimension, for example, a two-state Hilbert space for qubits (or a 2-state Hilbert space for quantum numbers). d (Hilbert space). In the qubit circuit system 252C, qubits can physically interact with each other in a controllable manner, thereby implementing quantum gates and realizing quantum computing. Although in Figure 1CIn the schematic diagram of the qubit circuit system 252C shown, each qubit is controllably coupled to four adjacent qubits (three for edge qubits and two for corner qubits). However, the qubit circuit system 252C may also have other topologies, including non-rectangular topologies, non-planar (e.g., three-dimensional) topologies, and so on. The specific physical implementation of the qubit circuit system 252C and the specific interactions (couplings) between the qubits can depend on the type of quantum computing device 250C or the type of quantum computing that the quantum computing device 250C is performing. The interactions between qubits can be controlled by voltage (current) signals, microwave (or radio frequency) signals, controlled propagation in a medium (e.g., for optical qubits), controlled exposure to a magnetic field (causing nuclear spin or electron spin precession), and so on.
[0066] The quantum bit circuit system 252C may further include a readout circuit system that uses a readout signal to detect the state of the quantum bit, for example, using microwave, radio frequency, optical signals and / or other signals depending on the specific physical type of the quantum computing device 110A.
[0067] The state control and readout of the qubit circuit system 252C can be performed by a control system 254C. The control system 254C may include various components of a classical computer system that can be used to initiate the state of the qubits, implement quantum gates, trigger readout measurements and implement error correction codes, store quantum computation results, and / or perform other operations as part of the execution of the quantum circuit or algorithm. The control system 254C may include: one or more classical processors, such as CPUs, GPUs, FPGAs, DPUs, ASICs, etc.; one or more memory devices; and input / output (I / O) units connected via one or more data buses. The control system 254C may also include digital-to-analog (DA) converters, analog-to-digital (AD) converters, electromagnetic signal generators, transmitters, receivers, filters, and / or other digital and / or analog devices. In one example, the control system 254C can be programmed to generate a series of control signals for the qubit circuit system 252C, such as initializing the correct number of qubits to implement the desired quantum circuit, applying a selected set of quantum gates to the initialized qubits, generating readout signals, performing error correction operations, etc.
[0068] Operations of software / firmware running on the (classic) computing device 120B can be performed using one or more GPUs 110C, one or more CPUs 130C, one or more parallel processing units (PPUs) or accelerators (e.g., deep learning accelerators, data processing units (DPUs), etc.). In at least one embodiment, the GPU 110C includes multiple cores 211C, each capable of executing multiple threads 212C. Each core can run multiple threads 212C concurrently (e.g., in parallel). In at least one embodiment, threads 212C can access registers 213C. Registers 213C can be thread-specific registers, with access restricted to the respective thread. Furthermore, one or more (e.g., all) threads of a core can access shared registers 214C. In at least one embodiment, each core 211C may include a scheduler 215C for distributing computational tasks and processes among different threads 212C of the core 211C. A dispatch unit 216C can implement the scheduled tasks on the appropriate threads using the correct private registers 213C and shared registers 214C. The computing system 100C may include one or more input / output components (I / O) 217C to facilitate the exchange of information with one or more users or developers.
[0069] In at least one embodiment, GPU 110C may have a (high-speed) cache 218C, to which multiple cores 211C may share access. Furthermore, computing system 100C may include GPU memory 219C, in which GPU 110C may store intermediate and / or final results (outputs) of various computations performed by GPU 110C. After completing a specific task, GPU 110C (or CPU 130C) may move the output to (main) memory 104C. In at least one embodiment, CPU 130C may execute processes involving serial computation tasks, while GPU 110C may execute tasks suitable for parallel processing (e.g., multiplying the input of a neural node by weights and adding a bias).
[0070] Further details on implementing noise channel error mitigation, including training machine learning models to implement noise channel error mitigation, will be referenced below. Figures 2 to 13B Describe it.
[0071] Figure 2 This is a schematic diagram 200 illustrating a method for mitigating errors in a noisy quantum channel using weak measurements, according to at least one embodiment. As shown, the density matrix " ρ The quantum state 210, represented by ", enters and is associated with " Â The base shown is composed of " εThe weak measurement channel 220, as shown, is related to the measurement strength. The weak measurement channel 220 can generate quantum states 210 on the basis "Â" to measure strength. ε Observable measurements 230, denoted as The following is a formal mathematical description of weak measurements (which can be used as training data to train machine learning models for mitigating noisy quantum channel errors) and machine learning models (e.g., diffusion models) for achieving noisy quantum channel error mitigation. In at least one embodiment, the complex density matrix is represented as a real-valued vector by splicing and flattening its real and imaginary parts, resulting in an input sequence composed of these vectors. This representation enables machine learning models to process quantum state data using neural network operations designed specifically for real-valued inputs.
[0072] For any two-state quantum system , can be Relative to observable  Perform weak measurements. From quantum states. The quantum system represented can be weakly coupled to an auxiliary quantum system, which is called a "pointer state" or "pointer". Pointers are typically modeled as continuous-variable quantum systems composed of quantum states. It means that among them Quantum systems With pointers The interaction between them can be controlled by the von Neumann projection measurement scheme, which is based on... and The joint state is applied using unitary operators. For example, von Neumann projectives can be performed as follows: (1) in It is a correspondence to the interaction strength ( 1) Small coupling parameter, It is a momentum operator. This interaction causes the pointer to deflect according to the eigenvalues of  and the coupling parameter γ. After the interaction, the tracking of the pointer's degrees of freedom defines a monitoring transformation. It can describe the initial quantum state (by the density matrix). The evolution of (represented by) under the influence of weak measurement is shown below: (2) in The measurement intensity parameter is used to quantify the degree of interference to the system. =0 corresponds to no measurement. =1 corresponds to a fully projective measurement. Indicates action on Specific types of noise, errors, or operations (e.g., depolarization, decoherence, bit flipping, phase flipping, amplitude damping, etc.). The monitoring transformation described by equation (2) decomposes the initial predictor density matrix into two components: the original (unmeasured) quantum state component and the projected quantum state component. The projected quantum state component corresponds to the quantum state projected onto the eigenbase of the measured observable and represents the result of the projected measurement. Therefore, the measured intensity parameter The relative weights of each component are determined. In at least one embodiment, successive weak measurements are performed to generate simulated training data. More specifically, multiple weak projections can be applied consecutively. These multiple weak projection measurements can collectively approximate the effect of a single strong (projective) measurement, but each step only partially collapses the quantum state. This method allows for the simultaneous extraction of partial or approximate information about several observables, particularly suitable for observables that do not satisfy commutation relations. However, it should be noted that the information about observables obtained from weaker measurements is less accurate and the results are more susceptible to statistical fluctuations. Furthermore, successive weak measurements of observables that do not satisfy commutation relations can be more destructive to the quantum state than a single projection measurement, because each measurement may partially erase the information obtained from previous measurements. When the same measurement strength parameter is used for the same observable... When performing continuous weak measurements, the cumulative effect can be represented by a series of... n This is described by a secondary monitoring transformation. This leads to quantum states. It evolved gradually, as shown in the following formula: (3) It can be seen that quantum state It can be gradually transformed into projected components on an observable basis. For multiple different observables... By performing continuous weak measurements, a transformation can be obtained, represented by the following nested summation equation: (4) Therefore, the evolution of quantum states can depend on the measurement of intensity parameters. In some cases, this dependency can lead to changes in the density matrix. The quantum state is uniformly decomposed over all measurable observables, thus effectively dispersing quantum information. When a series of weak measurements are performed, the quantum state undergoes an asymptotic evolution, taking incremental steps toward a maximally mixed state (i.e., completely losing information about the original quantum state). This process is caused by cumulative decoherence introduced by successive weak measurements. However, if the quantum state is weakly projected onto a large set of noncommutative observables, the system can resist complete decoherence, making it almost impossible to reach a maximally mixed state. This is because the information about the initial quantum state is distributed over incompatible measurement bases, which cannot be completely erased by any single measurement. In this case, the density matrix is typically relied upon. The chosen two-state vector form (TSVF) may not be suitable for describing the evolution of a system. Instead, quantum states can naturally approach a maximally mixed state through the dissipative effects of weak measurements, and coupling with an auxiliary system or environment can further facilitate this process. In some cases, machine learning techniques such as generative adversarial networks (GANs) can be used to assist in quantum state denoising, improving the fidelity of quantum information processing in the presence of decoherence caused by weak measurements by learning to distinguish and correct errors caused by noise.
[0073] The implementation of machine learning models (particularly the diffusion model in at least one embodiment) for error mitigation in noisy quantum channels is described in more detail below. As mentioned above, the diffusion model can be trained using a forward diffusion process, which progressively adds noise to the original (“denoised” or “clean”) quantum state data across multiple steps, ultimately resulting in “fully noisy” quantum state data (i.e., the quantum state after the last noise-adding step). For example, the original quantum state data could represent a pure quantum state. The diffusion model can then learn a reverse diffusion process, which aims to reconstruct the original quantum state data from the fully noisy data by progressively removing noise sequentially. This approach allows the diffusion model to approximate the inverse of the noise added in the forward diffusion process, thus aiding in the recovery of the original quantum state. The forward and reverse diffusion processes can be parameterized as Markov chains, and the diffusion model can be trained to minimize the difference between the reconstructed quantum state and the original quantum state, typically using a loss function such as mean squared error or quantum fidelity.
[0074] For example, data distribution The forward diffusion process can be written as: (5) In formula (5), Indicates that during the forward diffusion process t Quantum state data after adding noise (e.g., It is the original quantum state data. (It is completely noisy quantum state data). This represents the probability distribution of the original quantum state data. Indicates the steps of the forward diffusion process t The probability distribution of quantum state data, and Indicates from arrive The joint probability distribution of the entire diffusion trajectory (i.e., the probability of observing this particular noisy quantum state sequence from start to finish). (Term) Indicates that in a given In this case, steps tThe conditional probability of quantum state data (i.e., the probability of transitioning from one noisy quantum state to the next noisy quantum state). In at least one embodiment, ,in It is a Markov nucleus. is the diffusion rate, used to control the amount of noise added at each step. Therefore, equation (5) essentially shows that the joint probability of the entire diffusion trajectory can be expressed as the product of the probability of the original (denoised) quantum state data multiplied by the conditional probability of transitioning from one noisy quantum state to the next noisy quantum state.
[0075] The reverse diffusion process can be written as: (6) In equation (6), Steps representing the forward diffusion process t The probability distribution of time-space quantum state data This represents the probability distribution of the original quantum state data, while This represents the probability distribution of fully noisy quantum state data. (Item) Indicates that in a given In this case, steps t The conditional probability of the data at -1 (i.e., the probability of denoising the quantum state data in one step). Therefore, Equation (6) essentially shows that the probability of the original quantum state data can be expressed as the product of the probability of the fully noisy quantum state data and the inversion of the conditional probability of each denoising step. Accordingly, Equation (6) can govern the process of training the diffusion model to implement the backdiffusion process to generate predictions of the original quantum state based on arbitrary noisy quantum states.
[0076] Acting on the density matrix The noisy quantum channel representing the input quantum state can be expressed as: ,in ,in This represents an inverse quantum channel. An inverse quantum channel is not necessarily a physical quantum channel. As an illustrative example, a noisy quantum channel can be implemented as a depolarized channel. However, other types of noisy quantum channels are also considered, which may have more complex inverse quantum channels. The depolarized channel model can be used to represent a type of noise in a quantum system that loses information by projecting the state onto a statistical mixture of all possible states. A depolarized channel can be modeled as... ,in This represents the output quantum state after passing through the depolarized channel. (Term) Denotes the maximum mixture state in the associated Hilbert space, where It is the trace of the density matrix representing the quantum state (e.g., for an effective quantum state, ), nIt is the number of qubits that define the dimension of Hilbert space. It is the identity matrix in Hilbert space (e.g., in a single-qubit implementation). Therefore, in a single-qubit implementation scheme, the maximum mixture of effective quantum states can be represented by... Indicates. Parameters p Indicates by The probability that a quantum state passes through a depolarized channel undisturbed is represented by (1- p ) represents the probability that a quantum state is depolarized by a depolarization channel. For example, if p =1, then the depolarization channel is ideal, and For example, if p If the depolarization channel equals 0, then the depolarization channel will depolarize the quantum state to the maximum extent, thereby transforming the quantum state into a maximally mixed state. Therefore, it can be deduced that the reverse depolarization channel can be modeled using the following formula: .
[0077] In at least one embodiment, a quantum channel (and similar inverse quantum channels) can be represented by a global channel, defined as the product of local channels. For example, ,in It corresponds to the first i Local quantum states The i A local quantum channel, and As mentioned above, quantum channels can act on quantum states. T Next, its expression is: (7) The inverse quantum channel can be represented as: (8) The product of small inverse quantum channels can serve as an error mitigation protocol to eliminate noise in noisy quantum channels. It can also help identify simpler and more compact structures in quantum computer algorithms. For example, for operations acting on arbitrary states... Quantum circuit algorithms may be able to denoise single-qubit quantum states, among which... Corresponding to the initial pure state Departure, passing through Hamiltonian Time t Pure quantum states of time evolution U It is an operator that represents reversible quantum operations or evolution. yes U Hermitian conjugate (e.g., if) U If it is a unitary arithmetic operator, then... The Hamiltonian of the entire system (including pointers). Trotter approximation can be used Separated into n The steps, among which Using the Trotter approximation can remove noise from gate-dependent quantum algorithms and allow noise to be combined with auxiliary quantum data (such as qubits) to generate quantum algorithms.
[0078] As described above, the goal of this paper's noisy quantum channel error mitigation is to determine a suitable inverse quantum channel that can transition from an approximately maximally mixed state to a pure quantum state. As mentioned above, LSTM is an architecture that can be used to implement machine learning models (such as diffusion models) to determine the inverse quantum channel, provided that LSTM can effectively process time-series data. The short-term correlation of LSTM can be determined at each time step. t The calculation shows that each time step t Each iteration involves unknown monitoring channels. ( ρ ) density matrix ρ (t) Transform into ρ (t+1) The input to the LSTM is ρ (T) ,in ρ (T) Indicates the final time step T The completely noisy quantum state before reaching a certain predetermined threshold; the output of the LSTM is from ρ (T) arrive ρ (0) The approximate inverse sequence, where ρ (0) This represents the original denoised quantum state. Therefore, LSTM can remove... ρ The noise introduced above. LSTM can also transform one density matrix into another using the following sequence: .
[0079] LSTM can transform pure quantum states into mixed states (e.g., maximally mixed states). Under the first weak measurement sequence in an analog noisy quantum channel, this is achieved by the first density matrix. The evolution of the first quantum state can be represented as .in, Represents the first random pure quantum state. This represents the maximum mixed state (e.g., can be interpreted as white noise). It can be approximated as equal to ,in d It is the dimension of Hilbert space (e.g., for a single qubit, d(2). This is the initial quantum state of the noisy quantum channel, meaning the noisy quantum channel is designed to operate according to this transformation. This set of weak measurements may be unknown, but they can serve as a representation of the noisy quantum channel, on average, if... 1. Then the noisy quantum channel can be represented as a depolarized quantum channel. In In the limiting case approaching zero, the model simulates an average behavior describing a noncausal quantum process. In the first quantum state... After evolution, the input state of the noisy quantum channel can be transformed into a second random pure quantum state. .For example, It can be rotation (e.g., Therefore, under the second weak measurement sequence of the simulated noisy quantum channel, the second density matrix... The evolution of the second quantum state can be represented as Training datasets used to train machine learning models D It can include n One instance. For example, .
[0080] In at least one embodiment, the input features of the machine learning model are defined as follows: [ N , S , 2, D [, 2, 2], where N This represents the number of instances (batch size). S The sequence length is 2, where 2 represents the real and imaginary parts. D For the number of qubits, 2,2 defines a 2×2 density matrix representing the quantum state. The data loader class converts the input features into transformed features. N , S 8 D The transformed features are then fed into a machine learning model. For each sequence, the machine learning model can predict the output corresponding to the original quantum state and / or the inverse quantum channel that transforms the final completely noisy quantum state back to the original quantum state. The output can be determined by the weights of the machine learning model.
[0081] Consider a simple single-qubit case, where D =1. The state space of a single qubit can be geometrically represented using a Bloch sphere (unit sphere). Specifically, each pure state of a qubit can be mapped to a point on the surface of the Bloch sphere, while each mixed state of a qubit can be represented by any point inside the Bloch sphere (i.e., not on the sphere). For a single qubit, the density matrix representing the state of the qubit can be written as... + + ,in It is the Bloch vector (i.e., in pure state) In the mixed state , This represents the Pauli matrix. The Pauli matrix is a Hermitian unitary matrix, representing an operator that can be used to describe particles with spin 1 / 2 (such as qubits). Specifically, each Pauli matrix corresponds to a rotation about a corresponding axis on a Bloch sphere. For example, , , .
[0082] Any change or shift in the quantum state of a qubit (e.g., due to weak interactions, small rotations, or transient evolution under the action of the Hamiltonian) can be described as a Bloch vector. Changes or displacements This change or displacement can occur from near the maximum mixing state ( The quantum state begins.
[0083] The effect of a noisy quantum channel on the Bloch radius can be described by an affine transformation. To describe, among which It is a time step t The Bloch radius, The time step after adding noise t +1 Bloch radius It represents one or more linear transformations that are affine transformations (e.g., for...). Linear transformation matrices for scaling, rotation, and / or shearing. is the translation vector representing the translation transformation of the affine transformation. Therefore, the inverse quantum channel used to reverse the noise introduced by the quantum channel to recover the original quantum state of the qubit can be expressed as: ,in It is the inverse linear transformation matrix. Since the quantum channel is weakly affected by perturbations of the quantum state, the affine transformation representing the quantum channel can also be written considering the measurement intensity parameter. The prior knowledge is as follows: .
[0084] and It can be defined according to the type of noisy quantum channel represented. In the depolarized channel shown in the example above, it can map pure states to maximally mixed states, where It can be defined as the scaled identity matrix of a reduced Bloch sphere, while It can be the zero vector.
[0085] The forward diffusion process will now be described in detail. As mentioned above, the quantum channel can be represented by formula (7) as follows: The forward diffusion process can be further performed based on a finite set of POVM elements, the simplest of which is the Pauli-6 POVM. Using the forward diffusion algorithm, the simulated quantum channel can be randomized through weak measurements, where the measurement strength can be the same for each random measurement. The first quantum state can be weakly measured until it is mapped to a completely noisy quantum state (e.g., transformed into a maximally mixed state or near-maximally mixed state). At this point, the quantum channel is well-defined and can be applied to other instances. More details on how to implement the forward diffusion process will be referenced below. Figures 4A to 4B Describe it.
[0086] Figure 3 This is a schematic diagram of a machine learning model architecture (“architecture”) 300 according to at least one embodiment. Architecture 300 can be used to implement a machine learning model capable of performing noise quantum channel error mitigation (e.g., a quantum diffusion model process). More specifically, architecture 300 represents an LSTM.
[0087] like Figure 3 As shown, LSTM at the current time step t Receive input data. In at least one embodiment, the input data includes three input data items (e.g., a vector). The three input data items include: the LSTM at the previous time step... t -1 generates the cell state ("C") t-1 》302, Previous time step t -1 hidden state ("H" t-1 》304 and at the current time step t External data items provided to the LSTM ("X") t 》306.C t-1 302 retains information from multiple time steps as the network's long-term memory, while H... t-1 304 captures short-term dependencies by capturing information relevant to the current time step. X t 306 can represent the current time step. t Related new data or features, such as measurement results or encoded quantum information.
[0088] LSTM can generate at the current time step using input data. t Unit state during the process ("C") t 》308 and at the current time step t Hidden states during the process ("H") t 》310.H t310 can be used for downstream tasks or passed to subsequent layers. The performance of the LSTM can be evaluated based on this output, as it reflects the LSTM's ability to capture and utilize relevant temporal data dependencies.
[0089] To generate C t 308 and H t 310, Architecture 300 can further include a set of gates. This set of gates enables LSTM to pass H t-1 304 indicates recent historical information and X t 306 represents combining new data or features to learn temporal patterns and dependencies in sequence data, thereby updating C. t-1 302 represents long-term memory. Specifically, this set of gates determines how information is retained, updated, or discarded, and sets C... t-1 302 and H t-1 The value of 304 is transformed into C t 308 and H t 310. Architecturally, each gate in this set of gates can be implemented as a dedicated neural network layer, which includes an activation function to regulate the information flow. In at least one embodiment, the activation function of at least one gate is a sigmoid function, whose output is a value between 0 and 1 to control the degree of information transmission. These gates are used to generate updated unit states and hidden states by selectively filtering and combining relevant information from previous states and current inputs. For example, as... Figure 3 As shown, this group of gates may include forget gates (F... t 312. Input Gate (I) t )314 and output gate (O t 316. Each of gates 312-316 can be provided by H t-1 304 and X t 306 are combined to form the same gate input (e.g., concatenating them together to form a gate input vector).
[0090] Forget gate 312 evaluates the gate input to determine whether the previous cell state C should be retained or discarded. t-1 Which parts of 302, and subsequently generate from C t-1 The updated cell state 313 of 302. Specifically, forget gate 312 generates forget gate selector data items, typically a selector vector with values ranging from 0 to 1. Each element of this selector vector corresponds to C. t-1 The corresponding element in 302 is used as a multiplication mask. For example, the closer the value is to 1, the more likely the corresponding information is to be retained; the closer the value is to 0, the more likely the information is to be discarded. The updated cell state 313 is achieved by combining the forget gate selector vector with C. t-1The result is obtained by element-wise multiplication using 302. This allows LSTM to selectively retain or remove C. t-1 The information in 302 can help with the retention or forgetting of long-term memories (depending on the situation).
[0091] Input gate 314 determines the input from X t To what extent should new information from 306 be incorporated into C? t 308. To achieve this, architecture 300 may include a candidate memory component (CM) 318, which may be implemented as a separate neural network layer. CM 318 may generate candidate memory data items, typically represented as candidate memory vectors. In at least one embodiment, the output neurons of CM 318 apply a candidate memory value function to ensure that all elements of the candidate memory vector are constrained to a given range. For example, the given range may be -1 to 1. The hyperbolic tangent function (tanh) is an example of a candidate memory value function that can be used to ensure that all elements of the candidate memory vector are constrained to the range of -1 to 1.
[0092] Input gate 314 generates input gate selector data items, typically in the form of a selector vector with values between 0 and 1. This selector vector is used to modulate the candidate memory vector through element-wise multiplication, thereby generating cell state modification data items (“data items”) 319. This operation allows the network to control the degree of influence of each element in the candidate memory vector on the cell state update.
[0093] Then, the generated data item 319 can be combined with the updated cell state 313 at combiner 320 to produce a new cell state C. t 308. In at least one embodiment, combiner 320 performs element-wise addition on data item 319 and the updated cell state 313 to generate C. t 308. This process enables LSTM to retain C t-1 While addressing relevant aspects of 302, selectively integrating new information can support both the learning of new patterns and the maintenance of long-term dependencies.
[0094] Output gate 316 can determine H t The value of 310, this output can be used as the next time step. t +1 input. Therefore, the output gate 316 can be based on H. t 304 and X t 306 generates output gate selector data items (e.g., selector vectors). Architecture 300 may include a cell state normalizer 322, which applies a nonlinear activation function to the cell states C generated by combiner 320. t 308. This normalization ensures Ct The value of 308 is constrained to a range, typically between -1 and 1 (e.g., using the tanh function), which helps stabilize training and prevents the unit states from growing indefinitely. The output gate selector data items and the normalized unit states are then combined by a combiner 324, which typically performs element-wise multiplication to generate H. t 310. This mechanism allows LSTM to control how much cell state information is exposed at each time step, thereby enabling selective output of relevant information.
[0095] The LSTM represented by architecture 300 can be trained using any suitable technique. In at least one embodiment, the LSTM is trained using a backpropagation algorithm, specifically, a backpropagation time step (BPTT) algorithm. In BPTT, gradients are propagated backward along the sequence of time steps, enabling the model to learn temporal dependencies. The effectiveness of BPTT depends on the cell states, as the gradient at each time step relative to the cell states affects how information is retained or forgotten. The gradient calculation for backpropagation time step can be at least partially accomplished by C... t 308 and C t-1 The rate of change between 302 (e.g., The ratio is determined by the combined effect of the forget gate, input gate, and output gates 312-316, as well as CM 318. By adjusting these components during training, LSTM can learn to control the information flow and manage long-term dependencies, thereby optimizing the performance of the task of recovering original quantum state data from noisy quantum state data.
[0096] The following provides an exemplary, non-restrictive training example. Figure 3 The illustrated LSTM process is designed to approximate the inverse of any non-unitary quantum channel, thereby mapping noisy quantum state data back to the corresponding noise-free (original) quantum state data. Unless otherwise stated, all variables, symbols, and reference numerals correspond to those introduced elsewhere in this specification.
[0097] In at least one embodiment, a supervised learning task is used to train the LSTM. This supervised learning task is formalized as training the LSTM to learn a function. ,in Represents a non-unitary (noise) quantum channel M interaction T The density matrix of the completely noisy quantum states after the second step. This represents the density matrix representing the LSTM's prediction of the original quantum state (e.g., a pure quantum state). It is a parameter that includes all trainable weights and biases of the LSTM.
[0098] To train an LSTM, a random pure quantum state can be initialized. More specifically, initializing this random pure quantum state may include: setting the polar angle parameter... Initialize to a randomly selected value between 0 and π (inclusive), and set the azimuth parameter Initialize to a randomly selected value between 0 and 2π (inclusive), and initialize a representation of a primitive (e.g., pure) quantum state. The primitive (e.g., pure) quantum state vector, which is denoted as .
[0099] Then, a forward diffusion process can be performed to add noise to the original quantum state. The output of the forward diffusion process can be recorded as training instances. For example, Figure 4B It is a forward diffusion process algorithm (“Algorithm”) 400 according to at least one embodiment. For example... Figure 4B As shown, random pure quantum states were initialized. More specifically, as referenced above... Figure 3 The initialization of the random pure quantum state includes: setting the polar angle parameter Initialize to a randomly selected value between 0 and π (inclusive); set the azimuth parameter Initialize to a randomly selected value between 0 and 2π (inclusive); set the pure quantum state vector Initialize to ; and the density matrix Initialize to .
[0100] After initializing the parameter set, a while loop is started. The while loop will continue to execute as long as the threshold quantum state purity condition is met. In this example, the quantum state purity is determined by... trace Tr( ) definition. For a pure state, Tr( ) = 1; For the maximum mixture state, Tr( ) = ,in d It is the dimension of the relevant Hilbert space. In the example of a single qubit (e.g.) Figure 4B As shown), the associated Hilbert space has two dimensions ( d =2), therefore, the Tr( ) of a single qubit in the maximum mixing state is 2. )for As long as Tr( Greater than or equal to The while loop will continue to run. That is, in this example, the quantum state represents a qubit. However, in other embodiments, when the quantum state does not correspond to a single qubit, It can be promoted as .
[0101] During the while loop, a random basis measure is applied to the density matrix, thereby changing the density matrix. For example, the Pauli-6 POVM set can be used as the measure. Of course, any suitable POVM set can also be used. For example, values can be randomly selected. c It can be 0, 1, or 2. Without loss of generality, c==0 corresponds to a measurement using the z-basis, c==1 corresponds to a measurement using the x-basis, and c==2 corresponds to a measurement using the y-basis, such as... Figure 4B As shown (similar to formula (2) above). After applying a randomly selected base measurement, Updated to If Tr( Still greater than or equal to If the initial quantum state is successfully transformed into a maximally mixed state or a near-maximally mixed state, then another random basis measurement is applied. Otherwise, the process ends. Therefore, the number of iterations in the while loop can correspond to the number of weak measurements required to transform the primordial quantum state into a maximally mixed state or near-maximally mixed state, determined by... The size is controlled. After each measurement, the updated density matrix of the quantum state can be stored and normalized as follows: (9) in It is in the while loop. t The observables measured in each iteration.
[0102] review Figure 3 It can perform real-valued tensor encoding. In at least one embodiment, for the product state, the input features of the machine learning model are defined as follows: [ N , S , 2, D [, 2, 2], where N This represents the number of instances (batch size). S Let be the sequence length, and 2 define the real and imaginary parts. D Where is the number of qubits, and 2,2 defines the 2×2 size of the density matrix representing the quantum states. The data loader class can convert input features into transformed features. N , S , 8 D In at least one embodiment, for non-separable states, the input features of the machine learning model are defined as follows: N , S , twenty two D , 2 D ], and the input features can be converted into [ N , S , 2 2D+1 Further details regarding separable and non-separable states will be described below.
[0103] Any suitable parameters can be initialized according to the embodiments described herein. For example, the forget gate bias can be set to 1 to encourage retention of initial memories. As another example, an appropriate batch size N can be chosen to leverage parallel processing without memory overflow (e.g., N =64).
[0104] A loss function can be defined for use during training. This loss function can be used to calculate a distance metric. In at least one embodiment, the loss function includes an MSE loss term. Its properties are similar to those of the Euclidean norm described above. The loss function may also include other loss terms. In at least one embodiment, the loss function further includes a fidelity loss term. This term corresponds to a similarity measure between the true density matrix and the predicted density matrix. For example, fidelity could correspond to Uhlmann fidelity. Therefore, in at least one embodiment, the loss function is defined as... + .
[0105] The input quantum state data for machine learning models can be represented by a density matrix with trace 1. This is represented as follows: If a machine learning model is trained on multiple qubits, the input quantum state data of the machine learning model can be represented as... ,in Because machine learning models can operate in the real number domain. Non-complex fields Therefore, the input quantum state data should be represented as real numbers. For example, the density matrix. The values of the diagonal can be real numbers, and the density matrix It can be a Hermitian matrix. Therefore, for a two-state quantum system, the input quantum state data for each iteration can be represented by a 1×8 dimensional vector, which has the form: The sequence, and the output of the machine learning model can be the next approximate state. One or more features of the vector can be zero. This effectively creates a pure quantum state from the mixed quantum state by amplification and predetermined rotation (i.e., the parameters of the inverse quantum channel, and effectively acts as a purification protocol).
[0106] Regarding the training process, LSTM can be trained using any suitable number of quantum state input instances, which pass through a random non-unitary quantum channel created by adding noise to the initial quantum state.
[0107] In at least one embodiment, the training process includes randomly sampling from the training set. NA batch assembly process is performed using a sample. Input training data (e.g., input tensors) comprising an inverse time sequence can be constructed such that... t 0 corresponds to , T Corresponding to The input training data can be fed into an LSTM to determine the characteristics of each sample in the input training data. ,sample i of Then, the loss can be calculated using the loss function described above. After calculating the loss, BPTT can be performed to update the settings. In updating At this time, the parameters should keep the evolution of the quantum state within an acceptable quantum mechanical range. The hyperparameters controlling the noisy quantum channel simulation can be varied across training samples to enhance robustness to unknown noise levels. As long as the causal constraints are satisfied, Architecture 300 can be augmented with additional components, such as attention mechanisms, bidirectional LSTM layers, etc. Regularization, such as through weight decay and / or early stopping, can be used to prevent overfitting.
[0108] To monitor the training process, the average fidelity can be calculated on the validation set after each training epoch. For example, Figure 5 A graph 500, according to at least one embodiment, illustrates the relationship between average fidelity and the number of training instances. Specifically, graph 500 includes an x-axis 510 corresponding to the number of training instances, ranging from 100 to 10,000, and a y-axis 520 corresponding to the average fidelity. The machine learning model used includes an LSTM with 250 hidden layers and 3 main layers, and the sequence length of the instances is 17 steps. In this illustrative example, =0.3, the lower limit of purity is Tri(ρ) 2 = 0.7. As shown in Table 500, the average fidelity of the machine learning model plateaus at around 0.975. This is at least partly due to the small number of parameters in the machine learning model, resulting in minimal variation in each iteration. Therefore, different models may achieve better performance (e.g., an average fidelity of at least approximately 0.999).
[0109] Figure 6A This is a flowchart of a method 600 for training a machine learning model to achieve noise quantum channel error mitigation, according to at least one embodiment. Method 600 can be executed by processing logic, which may include hardware (e.g., circuit systems, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions running on a processing device to perform hardware simulation), or a combination thereof. In at least one embodiment, method 600 is as described above. Figure 1BThe training component 122B of the computing device 110B is executed.
[0110] In operation 610, the processing logic causes a measurement sequence of the quantum states of the quantum system to be obtained. In at least one embodiment, the measurement sequence of the quantum states is a weak measurement sequence of the quantum state, which does not completely collapse the wavefunction corresponding to the quantum state. In at least one embodiment, the measurement sequence of the quantum states includes multiple quantum state measurements captured in time sequence. In at least one embodiment, each measurement in the measurement sequence simulates the corresponding amount of noise added to the original quantum state by a non-unitary quantum channel at the corresponding time step. In at least one embodiment, the weak measurements are performed using a quantum measurement device, such as the one referenced above. Figure 1A The quantum measurement device 130 is described above. In at least one embodiment, weak measurement is performed by weakly coupling the quantum system with an auxiliary quantum system called a pointer state, as referenced above. Figure 2 As described in Equation (1). In at least one embodiment, the interaction between the quantum system and the pointer is controlled by a von Neumann projection measurement scheme, which applies a unitary operator to the joint state based on the quantum state and the pointer state. In at least one embodiment, the measurement intensity parameter ε determines the degree of perturbation of the quantum system by the measurement, with ε=0 corresponding to no measurement and ε=1 corresponding to a fully projective measurement. In at least one embodiment, the measurement sequence includes measurements performed on different observables, such as measurements in x-basis, y-basis, and z-basis. In at least one embodiment, the measurement sequence is used to simulate the effects of various noisy quantum channels, including depolarization channels, amplitude-damped channels, phase-damped channels, and bit-flipping channels.
[0111] In operation 620, the processing logic generates a training dataset comprising quantum state measurement sequences. In at least one embodiment, the training dataset comprises multiple measurement sequences corresponding to multiple different initial quantum states. In at least one embodiment, the training dataset comprises pairs of data, wherein each pair comprises a noisy quantum state (or a series of noisy quantum states) and its corresponding pristine pure quantum state. In at least one embodiment, generating the training dataset comprises generating multiple random pure quantum states by applying a random unitary matrix to a fixed reference state (e.g., a ground state). In at least one embodiment, each generated pure state is transmitted through a randomly selected quantum channel, which may be characterized by a specific type of quantum channel, such as a depolarization channel, an amplitude-damped channel, or a phase-damped channel. In at least one embodiment, the transmitted output noisy quantum state is recorded together with the corresponding pristine pure state to form training instances. In at least one embodiment, the training dataset comprises a density matrix representation of the quantum states at each time step during the forward diffusion process. In at least one embodiment, the training dataset is generated using both simulated data and experimental data obtained from actual quantum hardware, which enables the machine learning model to learn from real-world noise characteristics and operational defects.
[0112] In step 630, the processing logic trains a machine learning model using a training dataset to predict the primitive quantum state by simulating the inverse of a non-unitary quantum channel. More specifically, the machine learning model is trained to predict the primitive quantum state, thereby performing error mitigation. In at least one embodiment, the primitive state is a pure quantum state. In at least one embodiment, the machine learning model is trained to map noisy quantum state information to its corresponding ideal (noise-free) quantum state. In at least one embodiment, training the machine learning model includes minimizing a cost function that measures the difference between the predicted quantum state and the primitive quantum state. In at least one embodiment, the cost function includes an MSE term to measure the difference between the predicted density matrix and the target density matrix. In at least one embodiment, the cost function includes a trace normalization constraint that forces Tr(ρ) = 1, thereby forcing the output density matrix to be normalized. In at least one embodiment, the cost function includes a penalty on the non-Hermitian density matrix to account for the physical properties of the generated quantum system. In at least one embodiment, the model performance is evaluated using a geometric fidelity metric that measures the degree of overlap between the predicted density matrix and the target density matrix. In at least one embodiment, post-processing is applied to map the predicted output back to Hermitian form based on the constraints of the CPTP mapping.
[0113] In at least one embodiment, the machine learning model is a diffusion model. In at least one embodiment, the machine learning model includes a recurrent neural network. In at least one embodiment, the recurrent neural network includes an LSTM, as referenced above. Figure 3As described above. In at least one embodiment, the machine learning model is implemented using an encoder-decoder architecture with a multi-head attention mechanism (e.g., a visual transformer architecture that employs a self-attention mechanism to directly model the relationship between arbitrary pairs of elements in the input), as referenced. Figure 10 As described above. In at least one embodiment, the machine learning model is implemented using a U-Net architecture, where the encoder includes residual blocks running at progressively decreasing spatial resolution, and the channel dimension expands at each stage. In at least one embodiment, quantum states and quantum operations are represented using tensor networks, as described in reference [reference missing]. Figures 9A to 9B As stated above.
[0114] In at least one embodiment, the LSTM model can achieve high fidelity for separable quantum registers with multiple qubits (e.g., more than 10 qubits), where noise acts similarly to a local filter with low complexity. However, for entangled quantum registers, the LSTM model may become slow and fidelity may decrease as the number of qubits increases (e.g., more than 6 qubits). Attention-based models can process entangled quantum registers more efficiently and faster than LSTM models. In at least one embodiment, for entangled quantum registers, each density matrix requires 8 qubits proportionally. n Features (of which) n (This refers to the number of qubits), and for training local density matrices, each density matrix also requires 8 qubits proportionally. n Each feature is used to reconstruct the initial density matrix in both cases.
[0115] In at least one embodiment, the model may tend to overfit when trained on the local density matrix of the entangled register, which may necessitate the implementation of regularization methods. This overfitting behavior may be due to insufficient parameters as the model size increases, since the larger the global state, the greater the decoherence representation.
[0116] In at least one embodiment, training a machine learning model to predict the original quantum state by simulating the inverse of a non-unitary quantum channel includes: performing a reverse diffusion process to predict the reverse evolution of noise added to the original quantum state based on the forward diffusion process.
[0117] In at least one embodiment, training the machine learning model includes using a teacher-coercive policy that gradually transitions to autoregressive generation. In at least one embodiment, in the early stages of training, the model receives the ground truth target from the previous time step as input at each point in the sequence, which accelerates convergence by exposing the model to ideal local inputs. In at least one embodiment, linear platform scheduling is used to gradually reduce the teacher-coercive probability, such that during training, the model becomes increasingly dependent on the output of its previous steps, thus adapting to the final inference time scenario. In at least one embodiment, the trained machine learning model, in addition to being used for error mitigation, can also be used for state preparation, generating a target quantum state from noise by providing appropriate input, and allowing the model to produce a target output state through learned inverse dynamics.
[0118] More details about operating 610-630 have been referenced above. Figures 1A to 5 The following will be described and referenced. Figures 6B to 6C Describe it.
[0119] Figure 6B The flowchart, based on at least one embodiment, illustrates the execution... Figure 6A The example method in operation 630 is used to train a machine learning model.
[0120] In operation 640, the processing logic performs a forward diffusion process to obtain input data indicating a noisy quantum state. More specifically, the noisy quantum state is a pure quantum state modified with a certain amount of noise. In at least one embodiment, the noisy quantum state is a completely noisy quantum state. In at least one embodiment, the noisy quantum state is a maximally mixed quantum state. In at least one embodiment, the forward diffusion process progressively adds noise to the original quantum state data through multiple steps, eventually obtaining completely noisy quantum state data. In at least one embodiment, the forward diffusion process is parameterized as a Markov chain, where the probability distribution of the data at each step depends only on the immediately preceding step. In at least one embodiment, the forward diffusion process of the data distribution is described according to formula (5). In at least one embodiment, the forward diffusion process is implemented using the algorithm described above (see above). Figure 4A and Figure 4B ), where the density matrix representing the quantum state is iteratively modified by applying a monitoring transformation until the purity threshold condition is met.
[0121] In at least one embodiment, the forward diffusion process includes one or more iterations. In at least one embodiment, training a machine learning model to predict the original quantum state by simulating the inverse of a non-unitary quantum channel includes: performing each iteration of one or more iterations of the forward diffusion process by initializing a first quantum state represented by a first density matrix associated with Hilbert space; determining, based on the density matrix, whether the purity of the first quantum state satisfies a threshold condition defined based on the maximum mixing quantum state in Hilbert space; and modifying the first density matrix to obtain a second density matrix representing the second quantum state in response to determining that the purity of the first quantum state satisfies the threshold condition. In at least one embodiment, determining whether the purity of the first quantum state satisfies the threshold condition includes determining whether the purity determined by trace analysis using the first density matrix is greater than or equal to a threshold (e.g., in extreme cases, the threshold may be defined based on the purity determined by trace analysis using a density matrix representing the maximum mixing quantum state). In at least one embodiment, the threshold condition is determined by checking Tr[ρ 2 The evaluation is based on whether it is greater than the threshold δ, as mentioned above. Figure 4A As described above. In at least one embodiment, modifying the first density matrix to obtain the second density matrix includes applying randomly selected weak measurements to the first density matrix. In at least one embodiment, the randomly selected weak measurements are selected from measurements under the x-basis, y-basis, or z-basis, as referenced above. Figure 4B As described above. In at least one embodiment, a monitoring transformation is applied to the density matrix, as described above. Figure 4B An example implementation of the forward diffusion process is described below, which will be referenced in the following text. Figure 6C Describe it.
[0122] In operation 650, the processing logic trains a diffusion model to perform a backdiffusion process, thereby recovering the original quantum state of the quantum system from the input data. In at least one embodiment, the backdiffusion process is learned by the diffusion model during training, wherein the model is trained to progressively reverse the effects of the forward diffusion process by predicting and removing noise added at each step. In at least one embodiment, the backdiffusion process can be represented according to formula (6). In at least one embodiment, the goal of the backdiffusion process is to enable the model to reconstruct noise-free data from noisy input. In at least one embodiment, the diffusion model is trained to minimize the difference between the reconstructed quantum state and the original quantum state using a loss function such as MSE or quantum fidelity metric. In at least one embodiment, the trained diffusion model can approximate the inverse of the depolarization channel as described above. In at least one embodiment, the trained diffusion model can approximate the inverse of other types of noisy quantum channels, including amplitude-damped channels and phase-damped channels. Further details regarding operations 640-650 have been referenced above. Figures 1A to 6A To describe, now refer to Figure 6C Describe it.
[0123] Figure 6C The flowchart, based on at least one embodiment, illustrates the execution... Figure 6B The example method in operation 640 is used to perform the forward diffusion process.
[0124] In operation 642, the processing logic initializes a quantum state represented by a density matrix associated with the Hilbert space. In at least one embodiment, initializing the quantum state includes initializing a set of parameters. In at least one embodiment, initializing the quantum state includes initializing a set of angles used to define the Bloch sphere representation of the quantum state. In at least one embodiment, the set of angles includes polar angles and azimuth angles. In at least one embodiment, initializing the set of angles includes randomly selecting values for polar angles from the interval [0, π]. In at least one embodiment, initializing the set of angles includes randomly selecting values for azimuth angles from the interval [0, 2π]. In at least one embodiment, initializing the set of parameters includes initializing (e.g., constructing) the quantum state based on the set of angles. In at least one embodiment, initializing the set of parameters includes initializing the density matrix ρ based on the quantum state, for example by calculating ρ = In at least one embodiment, the initialized quantum state is of purity Tr[ρ]. 2 A pure quantum state where ] = 1. In at least one embodiment, for a multi-qubit system, the quantum state is initialized as a separable state, wherein each qubit is independently parameterized, as referenced below. Figure 9A As described above. In at least one embodiment, for a multi-qubit system, the quantum state is initialized as an inseparable (entangled) state (e.g., by applying a controlled NOT gate operation between qubit pairs), as referenced. Figure 9B As stated above.
[0125] In operation 644, the processing logic determines whether the purity of the quantum state satisfies a threshold condition based on the definition of a maximum mixed quantum state in Hilbert space. In at least one embodiment, determining whether the purity of the first quantum state satisfies the threshold condition includes: determining whether the trace determined by the density matrix is greater than or equal to a threshold defined based on the trace determined for a maximum mixed quantum state (e.g., the trace of a maximum mixed quantum state for a single qubit is...). In at least one embodiment, the threshold is defined as the sum of the trace determined for the maximally mixed quantum state and an error value greater than or equal to zero. In at least one embodiment, the purity of the quantum state is calculated as Tr[ρ 2 ], where ρ is the density matrix representing the quantum state. In at least one embodiment, the threshold condition is Tr[ρ 2] ≥ δ, where δ is a threshold greater than or equal to the purity of the maximum mixing state. In at least one embodiment, for a single-qubit system, the purity of the maximum mixing state is Tr[ρ 2 The threshold δ can be set to 1 / 2 + ε, where ε is a small positive value. In at least one embodiment, this threshold condition ensures that the forward diffusion process continues until the quantum state accumulates enough noise to approach the maximum mixing state.
[0126] If the purity of the quantum state does not meet the threshold condition, it means that sufficient noise has been added to the quantum state, and the forward diffusion process ends. In at least one embodiment, when the forward diffusion process ends, a sequence of density matrices representing the evolution of the quantum state from an initial pure state to a final noisy state is recorded and used as training data for a machine learning model. Otherwise, in operation 646, the processing logic modifies the density matrix to obtain a new density matrix representing the new quantum state. In at least one embodiment, modifying the density matrix to obtain the new density matrix includes applying randomly selected basis measurements to the density matrix. In at least one embodiment, the randomly selected basis measurements correspond to the x-axis, and the density matrix is updated according to ρ' = (1-ε)ρ + εΦ_X(ρ), where Φ_X(ρ) represents the projection onto the x-basis. In at least one embodiment, the randomly selected basis measurements correspond to the y-axis, and the density matrix is updated according to ρ' = (1-ε)ρ + εΦ_Y(ρ), where Φ_Y(ρ) represents the projection onto the y-basis. In at least one embodiment, the randomly selected basis measurement corresponds to the z-axis, and the density matrix is updated according to ρ' = (1-ε)ρ + εΦ_Z(ρ), where Φ_Z(ρ) represents the projection onto the z-basis. In at least one embodiment, after applying the monitoring transformation, the density matrix is updated by dividing by Tr[ρ']. 2 Normalization is performed to ensure proper normalization. In at least one embodiment, the measured intensity parameter ε is a constant that determines the rate at which noise is added to the vector quantum state at each step.
[0127] After modifying the density matrix to obtain a new density matrix, the forward diffusion process returns to operation 644 to determine whether the purity of the new quantum state satisfies a threshold condition. In at least one embodiment, this iterative loop continues until the purity of the quantum state falls below a threshold, indicating that the quantum state has accumulated sufficient noise. In at least one embodiment, the number of iterations required to reach the threshold condition depends on the measured intensity parameter ε and the initial quantum state. In at least one embodiment, the sequence of density matrices generated during the forward diffusion process represents the gradual evolution of the quantum state from a pure state to a noisy state, which can be used to train a machine learning model to learn the reverse diffusion process. In at least one embodiment, the forward diffusion process uses a POVM, such as the Pauli-6 POVM or SIC POVM as described above. In at least one embodiment, the quantum communication channel is divided into multiple segments, where the noise in each of the multiple segments approximates the weak measurement as described above. Further details regarding operations 642-646 have been referenced above. Figures 4A-4B and Figure 6B It has been described.
[0128] Figure 7 This is a flowchart illustrating an example method for mitigating noisy quantum channel errors using a trained machine learning model, according to at least one embodiment. Method 700 can be executed by processing logic, which may include hardware (e.g., circuit systems, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions running on a processing device to perform hardware emulation), or a combination thereof. In at least one embodiment, method 700 is as described above. Figure 1B The inference component 124B of the computing device 110B is executed.
[0129] In step 710, the processing logic receives input data indicating the original quantum state of the quantum system, which is modified by the amount of noise added by the non-unitary quantum channel. In at least one embodiment, receiving the input data includes processing one or more physical signals (e.g., optical signals, electrical signals, and / or microwave signals) received from the non-unitary quantum channel. For example, the original quantum state may be transmitted from a first quantum computing system to a second quantum computing system via the quantum channel. In at least one embodiment, the non-unitary quantum channel is a depolarization channel that transforms the quantum state according to the probability that the quantum state passes through the channel undisturbed. In at least one embodiment, the non-unitary quantum channel is an amplitude-damped channel that models the process by which a qubit loses energy to its environment. In at least one embodiment, the non-unitary quantum channel is a phase-damped channel that models the loss of quantum information due to random phase fluctuations. In at least one embodiment, the non-unitary quantum channel is a bit-flipping channel that flips the state of a qubit from |0> to |1> or vice versa with a certain probability.
[0130] In at least one embodiment, the input data includes a plurality of local density matrices, each of which corresponds to a corresponding qubit of an entangled quantum register, and is obtained by performing a partial trace operation on the global density matrix as described above.
[0131] In at least one embodiment, the quantum communication channel is divided into multiple segments, wherein the noise in each of the multiple segments approximates the weak measurement as described above.
[0132] In operation 720, the processing logic provides input data to a machine learning model that predicts the primitive quantum state by simulating the inverse of a non-unitary quantum channel, thereby reversing one or more errors caused by the non-unitary quantum communication channel. More specifically, the machine learning model is trained to predict the primitive quantum state, thereby achieving error mitigation. In at least one embodiment, the machine learning model is a diffusion model. In at least one embodiment, the machine learning model includes a recurrent neural network, such as the one referenced above. Figure 3 The LSTM described above. In at least one embodiment, the machine learning model includes the encoder-decoder architecture (e.g., a visual transformer architecture) with a multi-head attention mechanism as described above. In at least one embodiment, the machine learning model includes a U-Net architecture with residual blocks and skip connections as described above. In at least one embodiment, providing input data to the machine learning model includes encoding the density matrix as a real-valued tensor by separating the real and imaginary parts, as described above.
[0133] In at least one embodiment, the machine learning model is trained by simulating one or more non-unitary quantum channels based on one or more measurement sequences of one or more quantum states, wherein each measurement sequence of the quantum state comprises multiple measurements of that quantum state captured in time sequence. In at least one embodiment, each measurement sequence of the quantum state is a weak measurement sequence of that quantum state, which does not completely collapse the wavefunction corresponding to that quantum state. In at least one embodiment, the machine learning model uses the above reference... Figure 4A , Figure 4B and Figure 6CThe forward diffusion process is trained by applying successive weak measurements to the random pure quantum state until a threshold purity condition is met. In at least one embodiment, the machine learning model is trained using both simulated data and experimental data obtained from actual quantum hardware. In at least one embodiment, the machine learning model is trained using training data comprising multiple local density matrices, each corresponding to a corresponding qubit of an entangled quantum register, and obtained as described above by performing partial trace operations on the global density matrix. In at least one embodiment, the machine learning model is trained to reconstruct the global quantum state from the multiple local density matrices.
[0134] In at least one embodiment, the machine learning model is trained to perform a backdiffusion process to predict the reverse evolution of noise added to the original quantum state via a forward diffusion process. In at least one embodiment, the forward diffusion process includes one or more iterations, wherein each iteration of the one or more iterations of the forward diffusion process is performed by: initializing a first quantum state represented by a first density matrix associated with a Hilbert space; determining, based on the density matrix, whether the purity of the first quantum state satisfies a threshold condition defined by a maximum mixing quantum state based on the Hilbert space; and in response to determining that the purity of the first quantum state satisfies the threshold condition, applying a random basis measurement to the first density matrix to obtain a second quantum state associated with a second density matrix. In at least one embodiment, the backdiffusion process iteratively applies learned denoising operations to the input data to reconstruct the original quantum state from the noisy input. In at least one embodiment, the number of backdiffusion steps corresponds to the number of forward diffusion steps used during training. In at least one embodiment, the machine learning model is trained using a teacher-coercive strategy that progressively transitions to autoregressive generation and uses linear platform scheduling to progressively reduce the teacher-coercive probability, as described above. In at least one embodiment, the machine learning model is trained using a cost function that includes trace normalization constraints and penalties for the non-Hermitian density matrix, as described above.
[0135] In operation 730, the processing logic obtains output data indicating the primordial quantum state from a machine learning model. More specifically, the primordial quantum state refers to the quantum state of the quantum system before transmission through a noisy quantum channel. For example, the output data may indicate the primordial state of a qubit transmitted through a noisy quantum channel. In at least one embodiment, the output data includes a predicted density matrix corresponding to the primordial quantum state before noise contamination. In at least one embodiment, post-processing is applied to the output data to ensure that the predicted density matrix satisfies the physical constraints of a valid quantum state, including Hermitianness, positive semi-definiteness, and unit trace. In at least one embodiment, the output data is validated using a fidelity metric as described above, which measures the overlap between the predicted density matrix and the target density matrix. In at least one embodiment, the output data indicates a global quantum state reconstructed from multiple local density matrices, as described above. In at least one embodiment, the output data is used for state preparation of a quantum computing system as described above.
[0136] In at least one embodiment, the input data includes a modified message corresponding to the original message, adjusted for noise levels, and the output data includes a prediction of the original message. In at least one embodiment, the input data corresponds to one or more optical qubits obtained after processing an optical signal received through a non-unitary quantum channel. For example, the one or more optical qubits can be transmitted over long distances between computing systems, and a machine learning model can be used to clean up noisy quantum states before they enter the next quantum computing system. In at least one embodiment, the machine learning model identifies at least one type of noise introduced by the non-unitary quantum channel. In at least one embodiment, the machine learning model is trained to handle multiple types of noise simultaneously, thereby enabling error mitigation for quantum channels with complex noise characteristics. In at least one embodiment, the machine learning model is able to identify, based on the characteristics of the input data, whether the noise is primarily depolarization noise, amplitude-damped noise, phase decay noise, or a combination thereof.
[0137] In operation 740, the processing logic performs at least one quantum computing operation using the output data. In at least one embodiment, the at least one quantum computing operation includes sending data to the quantum computing system indicating a primitive quantum state predicted by a machine learning model.
[0138] In at least one embodiment, the at least one quantum computing operation includes performing quantum sensing and / or metrology to improve measurement accuracy. For example, a denoised quantum state output can be used to eliminate errors caused by noise in the quantum sensor output, thereby improving measurement accuracy.
[0139] In at least one embodiment, the at least one quantum computing operation includes: performing a quantum simulation to generate one or more models with higher accuracy. For example, the error-mitigated quantum state output can be used to extract more accurate expected values and observables from quantum simulations of complex molecular structures, material properties, and chemical reactions.
[0140] In at least one embodiment, the at least one quantum computing operation includes performing at least one quantum cryptographic operation. In at least one embodiment, performing the at least one quantum cryptographic operation includes performing quantum key distribution with higher precision. For example, the denoised quantum state output can achieve more robust long-distance quantum key distribution by correcting channel-induced errors before key extraction.
[0141] In at least one embodiment, the at least one quantum computing operation includes: performing quantum teleportation, wherein the output of a machine learning model is used to correct the received quantum state after teleportation, thereby improving the accuracy of quantum state transfer between spatially separated quantum systems. The corrected quantum state output can then be used for subsequent quantum operations at the receiving node.
[0142] In at least one embodiment, the at least one quantum computing operation includes facilitating distributed quantum computing, wherein the denoised quantum state output at each receiving node ensures that subsequent distributed quantum operations are performed using high-fidelity quantum data.
[0143] In at least one embodiment, the at least one quantum computing operation includes executing a quantum algorithm. In at least one embodiment, the at least one quantum computing operation includes executing a quantum circuit. In at least one embodiment, the method is integrated with a variational quantum algorithm (e.g., a variational quantum linear solver (VQLS) as described above).
[0144] Figure 8The diagram illustrates the test set fidelity distribution according to at least one embodiment. The test set fidelity distribution 800 depicts a histogram showing the distribution of fidelity values across the entire test dataset, where the horizontal axis 802 represents the fidelity value range of approximately 0.2 to approximately 1.0, and the vertical axis 804 represents the count value range of 0 to approximately 1200. The histogram bars indicate that most test samples achieve high fidelity values concentrated around 1.0, with the highest bar reaching nearly 1200 counts in the highest fidelity range. The dashed vertical line 810 represents the average fidelity value of approximately 0.9365, and the dashed vertical line 820 represents the median fidelity value of approximately 0.9757. The test set fidelity distribution 800 shows a clear bias towards high fidelity values, while instances with fidelity values below 0.8 are relatively few, indicating that the trained machine learning model is able to successfully reconstruct quantum states with high accuracy on the test dataset. In at least one embodiment, the test set fidelity distribution 800 corresponds to a validation dataset, which includes instances not encountered by the machine learning model during training, thereby demonstrating the trained model's generalization ability to unseen quantum state data. The fidelity metric used to generate the test set fidelity distribution 800 may correspond to geometric fidelity, defined as... , It measures the overlap between the predicted density matrix and the target density matrix. In at least one embodiment, the test set fidelity distribution 800 is generated using a dataset comprising 10,000 instances, of which 7,000 instances are used for training and 3,000 instances are reserved for validation, measuring the intensity parameter. =0.15, the stop parameter corresponds to the fidelity threshold. =0.7. The concentration of fidelity values close to 1.0 in the test set fidelity distribution 800 indicates that the machine learning model effectively learns the inverse dynamics of the noisy quantum channel, thus enabling accurate reconstruction of the original quantum state from noisy input data. In at least one embodiment, the difference between the average fidelity value and the median fidelity value in the test set fidelity distribution 800 reflects the presence of a small number of outlier instances with lower fidelity values, which may correspond to quantum states that are difficult to reconstruct due to proximity to a maximally mixed state or other factors affecting the learning process.
[0145] Figure 9AThe diagram illustrates a tensor network representation of a separable quantum register tensor network 900A according to at least one embodiment, wherein independently parameterized subsystem states are propagated through local tensor operations, with no coupling between the subsystems. The separable quantum register tensor network 900A comprises two parallel single-qubit tensor networks 910A and 920A, which operate independently and are not coupled to each other. This architecture reflects the mathematical structure of separable quantum states, where each local subsystem evolves in its own quantum channel, unaffected by other subsystems.
[0146] The single-qubit tensor network 910A begins with the locally pure state 912A, which is determined by the angular parameters. θ 1 and Characterization, based on The qubit states on the Bloch sphere are parameterized. A local pure state 912A is fed into a tensor operation module 914A, which receives local feedback 915A from an external input representing a weak measurement operation applied to the qubit. The tensor operation module 914A implements the aforementioned monitoring transformation. The local feedback 915A can specify the measurement basis and intensity for a specific time step. The output of the tensor operation module 914A then flows into the tensor operation module 916A, which can also receive the local feedback 917A representing subsequent weak measurement operations. The output of the tensor operation module 916A enters the tensor operation module 918A, which receives the local feedback and generates the first output connection 919A at the bottom of the network. Each subsequent tensor operation module represents an additional time step in the forward diffusion process, and the cumulative effect of multiple weak measurements gradually transforms the pure state into a mixed state.
[0147] The 920A single-qubit tensor network can operate in parallel with the 910A single-qubit tensor network, and in terms of angular parameters... θ 2 and The local pure state 922A is represented first. Local pure state 922A is fed into tensor operation module 924A, which receives local feedback 925A. The output flows to tensor operation module 926A, which receives local feedback 927A. The output of tensor operation module 926A enters tensor operation module 928A, which receives local feedback and generates a second output connection 929A at the bottom of the network. The parallel structure of the two single-qubit tensor networks 910A and 920A reflects the independence of the local quantum channel, i.e., a weak measurement applied to one qubit does not affect the evolution of the other qubit.
[0148] The separable quantum register tensor network 900A demonstrates a structure where each qubit evolves locally within its respective sequence of tensor operation modules, with no cross-interactions between the two single-qubit tensor networks 910A and 920A. Local feedback connections can represent the inputs to each tensor operation module, which influence the transformations applied at each stage, corresponding to a randomly chosen measurement basis (e.g., the x, y, or z basis from the Pauli-6 POVM) applied in each iteration of the forward diffusion process. This architecture allows the quantum register to be written as a tensor product of the individual qubit states, thus enabling the number of features to grow linearly with the number of qubits, rather than exponentially. Specifically, for a separable quantum register with d qubits, the input features of the machine learning model have [ N , S , d The shape of [, 2, 2, 2] can be converted into [ N , S , 8 d ] features, among which N It refers to the batch size. S It is the sequence length, 8 d This represents the real and imaginary parts of each 2×2 density matrix for all d qubits. In at least one embodiment, this separable representation is used when the quantum state being processed does not exhibit entanglement between subsystems, such as in quantum computing applications where the qubits are independently prepared and manipulated before any entanglement operations are applied.
[0149] Figure 9B The diagram illustrates a tensor network architecture in which initially separable local tensor states are transformed into non-separable tensors according to at least one embodiment. The non-separable quantum register tensor network 900B comprises three single-qubit tensor networks 910B, 920B, and 930B, coupled via a multi-qubit entanglement operator layer 905B that introduces entanglement between previously independent subsystems. This architecture represents a more general case where quantum correlations exist between different qubits in the register, where the machine learning model processes the complete global density matrix rather than the individual local density matrices.
[0150] 910B Single-Qubit Tensor Network and Angular Parameters θ 1 and The system is associated with and includes a locally pure state 912B, which is fed into a multi-qubit entanglement operator layer 905B. Below the multi-qubit entanglement operator layer 905B, a single-qubit tensor network 910B includes: a tensor operation module 914B that receives local feedback 915B; and a tensor operation module 916B located below the tensor operation module 914B that receives local feedback 917B. The tensor operation modules 914B and 916B are connected in series, with the output flowing downwards from the tensor operation module 916B. Although the noise channels implemented by the tensor operation modules act locally on each qubit, the global quantum state still maintains the correlation introduced by the entanglement operation, meaning that the local evolution of one qubit will affect the reduced density matrix of other qubits through shared entanglement.
[0151] 920B Single-Qubit Tensor Network and Angular Parameters θ 2 and Associated with and including a local pure state 922B, which is fed into a multi-qubit entangled operator layer 905B. Below the multi-qubit entangled operator layer 905B, a single-qubit tensor network 920B includes a tensor operation module 924B that receives local feedback 925B, and a tensor operation module 926B located below the tensor operation module 924B that receives local feedback 927B.
[0152] 930B Single-Qubit Tensor Network and Angular Parameters θ 3 and Associated with and including a local pure state 932B, which is fed into a multi-qubit entangled operator layer 905B. Below the multi-qubit entangled operator layer 905B, a single-qubit tensor network 930B includes a tensor operation module 934B that receives local feedback 935B, and a tensor operation module 936B located below the tensor operation module 934B that receives local feedback 937B.
[0153] A multi-qubit entanglement operator layer 905B serves as a coupling mechanism, connecting three locally pure states 912B, 922B, and 932B, thereby enabling the generation of entangled quantum states by applying controlled NOT gate operations between randomly selected qubit pairs. In at least one embodiment, the multi-qubit entanglement operator layer 905B applies controlled NOT gate operations between randomly selected qubit pairs. m There are n controlled NOT gates, where each controlled NOT gate can be represented using tensor notation as follows: The COPY operator copies the state of the qubits, while the XOR operator performs a conditional bit flip on the target qubit. After entanglement is established through the 905B multi-qubit entanglement operator layer, a local noise channel implemented by the tensor operation module processes each qubit independently while maintaining the global correlation introduced by the entanglement operation. Local feedback connections represent the iterative application of weak measurement operations to each subsystem during the forward diffusion process used to generate training data for a machine learning model that performs quantum error mitigation. For qubits with... d The inseparable quantum states of qubits, the input features of the machine learning model have [ N , S , twenty two d , 2 d The shape of ] can be converted into [ N , S , 2 2d+1 This characteristic reflects the exponential growth of the Hilbert space dimension with the number of qubits. In at least one embodiment, when the quantum states being processed are entangled between subsystems, an inseparable representation is used, a situation that may occur in quantum computing applications involving multi-qubit gates or quantum communication protocols. In at least one embodiment, the machine learning model is trained based on a local density matrix obtained by performing a partial trace operation on the global density matrix, denoted as... This can reduce the number of input features, but due to the loss of relevant information between qubits, it may lead to more challenging learning tasks, resulting in overfitting.
[0154] In at least one embodiment, a machine learning model is trained to reconstruct the global quantum state using only local density matrices. More specifically, at each time step, each local density matrix can be computed by performing partial trace operations on the global density matrix, thereby performing trace operations on all qubits except the one of interest. The model can then be trained based on the local evolution of these individual qubit density matrices. This approach is more practical but also more challenging because the local density matrices lack information about the correlations between qubits. Despite the reduced input information, the model can still be trained to predict the initial global density matrix based on the evolution of the local density matrices.
[0155] Figure 10A visual transformer architecture 1000 according to at least one embodiment is illustrated, which can be used to mitigate errors in noisy quantum data. The visual transformer architecture 1000 can serve as an alternative to recurrent neural network architectures for processing quantum state representations, particularly where capturing nonlocal correlations and long-range dependencies in the data may benefit the error mitigation process. Unlike recurrent neural networks (e.g., LSTMs) that process sequential data by iteratively updating hidden states, the visual transformer employs a self-attention mechanism, which can directly model the relationships between arbitrary pairs of elements in the input sequence, thereby enabling more efficient capture of complex correlations present in quantum state data. The visual transformer architecture 1000 may be particularly advantageous when processing quantum states in high-dimensional Hilbert spaces, where the number of parameters grows exponentially with the number of qubits, whereas convolutional or recurrent methods may require very deep networks to achieve a similar receptive field size.
[0156] The vision transformer architecture 1000 receives an input tensor 1010 representing quantum state data to be processed. In at least one embodiment, the input tensor 1010 may represent a series of density matrices corresponding to the evolution of the quantum state through a noisy quantum channel, wherein each density matrix is encoded as a real-valued tensor by separating its real and imaginary parts. For example, for a quantum state with… d A quantum register with qubits in an inseparable configuration, the shape of the input tensor 1010 can be [ N , S , twenty two d , 2 d ], where N is the batch size, S This corresponds to the sequence length of the time steps in the forward diffusion process; the other dimensions are encoded as 2. d × 2 d The real and imaginary parts of the density matrix. The input tensor 1010 is provided to the block embedding layer 1020, which divides the input data into blocks and embeds them into a representation suitable for transformer processing. In the context of quantum state data, the block embedding layer 1020 can divide the density matrix representation into smaller components for processing by the transformer architecture, similar to how image patches are extracted in standard vision transformer applications. In at least one embodiment, the block embedding layer 1020 extracts non-overlapping blocks from the density matrix representation, where each block corresponds to a subset of matrix elements encoding correlations between specific subsets of qubits.
[0157] A block embedding layer 1020 is connected to a linear projection layer 1030, which projects the embedded blocks into a higher-dimensional feature space suitable for the transformer encoder. The linear projection layer 1030 can implement a learned linear transformation that maps each block to a fixed-dimensional embedding vector, enabling the transformer to process blocks of different sizes uniformly. The linear projection layer 1030 is fed into a transformer encoder module 1040, which processes the projected features through an attention mechanism. A position encoding module 1050 adds positional information to the features, enabling the transformer to understand spatial or sequential relationships in the data. In the context of quantum error mitigation, the position encoding module 1050 can encode temporal information corresponding to the time steps of a weak measurement sequence or a forward diffusion process, enabling the model to distinguish density matrices at different stages of noise evolution. In at least one embodiment, the position encoding module 1050 implements learnable positional embeddings that are added to the block embeddings, allowing the model to learn the optimal positional representation of the quantum state data during training.
[0158] The position encoding module 1050 combines the output of the transformer encoder module 1040 and feeds it into the transformer module 1060. The transformer module 1060 applies a multi-head self-attention mechanism and a feedforward operation to capture dependencies between the input data. The multi-head self-attention mechanism enables the visual transformer architecture 1000 to recognize correlations between different components in the quantum state representation, which is very useful for reconstructing the original quantum state from noisy observations. Specifically, the self-attention mechanism computes attention weights to determine how much attention each element in the input sequence should give to each other element, thus enabling the model to capture both local and global dependencies without being constrained by the sequential processing of a recurrent architecture. The attention operation can be represented as... ,in Q , K and V These represent the query matrix, key matrix, and value matrix derived from the input features, respectively. It is the dimension of the bond vector. This capability may be particularly advantageous for handling entangled quantum states, where the correlation between distant qubits is crucial for accurate error mitigation. For example, Figure 9B The self-attention mechanism, as shown, can learn to recognize and maintain entangled structures when reconstructing the original quantum state from noisy observations by applying an inseparable quantum state generated by a controlled NOT gate. In at least one embodiment, the transformer module 1060 includes a multi-layer multi-head self-attention network followed by a feedforward network, and applies residual connections and layer normalization after each sub-layer to facilitate gradient flow during training.
[0159] Following the transformer module 1060 is a layer normalization module 1070, which normalizes the activation values to stabilize training and improve convergence. The layer normalization module 1070 applies normalization independently to the feature dimension at each location, which may be more suitable for variable-length sequences than batch normalization. In at least one alternative embodiment, batch normalization is used instead of layer normalization.
[0160] A layer normalization module 1070 is connected to a projection head module 1080, which maps processed features to a target output dimension for error mitigation. In at least one embodiment, the projection head module 1080 includes one or more fully connected layers that transform the transformer's output to match the dimension of the target density matrix representation. The projection head module 1080 generates an output tensor 1090 representing the denoised or error-mitigated quantum state data. In at least one embodiment, the output tensor 1090 may represent a predicted density matrix corresponding to the original quantum state before noise contamination, where the real and imaginary parts are reconstructed from the model output and combined to form an effective Hermitian density matrix. The output tensor 1090 may be post-processed to ensure that the predicted density matrix satisfies the physical constraints of the effective quantum state, including Hermitianness, positive semi-definiteness, and unit trace.
[0161] In at least one embodiment, the visual transformer architecture 1000 is trained using the same training data and loss function as described above for the LSTM architecture, including a mean squared error loss term. And fidelity loss item The latter is used to measure the similarity between the predicted density matrix and the target density matrix. Training data can use reference... Figures 4A to 4B The forward diffusion process described in [the document] generates [the data], in which continuous weak measurements are applied to random pure quantum states until a threshold purity condition is met. The Visual Transformer Architecture 1000 has advantages in processing quantum state data with complex correlations because the self-attention mechanism can capture long-range dependencies without being limited by sequential processing like recurrent architectures. The computational complexity of self-attention is quadratic with the sequence length, expressed as [equation]. ,in NThis refers to the number of spatial locations. However, this computational cost can be significantly reduced by applying attention at a lower spatial resolution (e.g., at the bottleneck layer), while still enabling the model to learn global dependencies. Furthermore, the Visual Transformer Architecture 1000 is better suited for parallel computation during training and inference, potentially enabling faster processing of large datasets or high-dimensional quantum state representations. The parallelization of attention computation, compared to the inherent sequentiality of recurrent architectures, can help utilize GPU resources more effectively during training. In some cases, the Visual Transformer Architecture 1000 is combined with recurrent components to form a hybrid architecture that leverages the advantages of both approaches, using attention mechanisms to capture global correlations while employing recurrent connections to model the temporal dynamics of noise evolution. In at least one embodiment, the Visual Transformer Architecture 1000 is combined with an encoder-decoder structure similar to the U-Net architecture, where the encoder progressively reduces the spatial dimension and increases the feature channels, the bottleneck applies multi-head self-attention to capture long-range spatial dependencies, and the decoder uses skip connections to reconstruct the denoised quantum state at the original resolution, thus preserving high-resolution details and phase information.
[0162] Figures 11A to 11B This is a schematic diagram of a network architecture 1100 that can be used to implement a quantum communication protocol according to at least one embodiment. For example, network architecture 1100 may include a data center 1102, a communication network 1104, and one or more network devices 1106. Network architecture 1100 may demonstrate a general computing architecture within which more specific systems and / or subsystems can operate.
[0163] For example, data center 1102 can be a centralized facility designed to house computing resources and related components. Data center 1102 can operate to support the infrastructure required for advanced computing tasks, enabling efficient, secure, and reliable operation. Data center 1102 may include building and structural components, including power, cooling systems, fire protection systems, and physical security measures designed to maintain optimal operating conditions and / or protect equipment from environmental hazards and unauthorized access. Data center 1102 may include high-performance servers or computing nodes, typically arranged in racks as shown in network architecture 1100 and connected via the high-speed network described herein. These servers may include processors (also referred to herein as processing devices, such as central processing units (CPUs), graphics processing units (GPUs), data processing units (DPUs), etc.), quantum processing units (QPUs), multiple parallel processing units (PPUs), and application-specific integrated circuits (ASICs), memory (e.g., RAM), and storage solutions (e.g., hard disk drives (HDDs), solid-state drives (SSDs)). Hardware configurations can be designed for parallel processing and high throughput to meet the demands of high-performance computing (HPC) applications. A QPU is configured to perform one or more operations related to a quantum algorithm. In some embodiments, each of the one or more QPUs may include multiple qubits, and the one or more QPUs may communicate with each other via a quantum channel. In some embodiments, each of the multiple qubits may include local qubits, global qubits, and / or synchronization qubits. In some embodiments, the local qubits of each QPU may be configured to perform one or more operations related to a quantum algorithm on the QPU associated with the local qubits.
[0164] Data center 1102 may include high-speed network devices, such as network switches, routers, and firewalls, to facilitate fast and secure data transfer within data center 1102 (e.g., between servers or compute nodes) and between external networks. Data center 1102 can facilitate communication between servers or compute nodes through a network topology that ensures efficient data exchange, minimizes latency, and maximizes bandwidth. The network topology determines how various network devices (e.g., switches and routers) interconnect to stream data. By implementing an effective network topology, data center 1102 can support high-performance computing tasks. Examples of various network topologies may include hierarchical networking topologies, such as fat-tree topologies, Slim Fly topologies, and Dragonfly topologies.
[0165] Communication network 1104 can communicatively couple data center 1102 to one or more network devices 1106 and other external devices to enable data exchange and connectivity. Examples of communication network 1104 may include Internet Protocol (IP) networks, Ethernet, InfiniBand (IB) networks, Fibre Channel networks, the Internet, cellular communication networks, wireless communication networks, combinations thereof (e.g., Fibre Channel over Ethernet), variations thereof, and so on. The ability of communication network 1104 to integrate multiple network types and configurations enables data center 1102 to adapt to a variety of application needs, from general data communication to dedicated HPC tasks. As described herein, communication network 1104 can utilize various optical components to establish communication links (e.g., communicatively coupled) between components in architecture 1100. Therefore, communication network 1104 may include various optics, transceivers, modules, and so on, configured to generate optical signals (e.g., provide optical transmitter functionality) and / or receive optical signals (e.g., provide optical receiver functionality).
[0166] One or more network devices 1106 may include a variety of computing devices capable of sending and receiving signals via communication network 1104. The range of network devices 1106 can be from personal computing devices to complex server configurations. Examples include personal computers (PCs), laptops, tablets, smartphones, and servers. One or more network devices 1106 can facilitate user interaction with data center 1102, allowing data input, retrieval, and processing from remote locations. In addition to individual computing devices, one or more network devices 1106 may also include server clusters or additional data centers. For example, these data centers may be other data centers similar to or identical to data center 1102. This interconnection can allow the formation of a distributed computing environment, thereby improving redundancy, load balancing, and disaster recovery capabilities. By linking multiple data centers, network architecture 1100 can leverage geographically dispersed resources, optimize performance, and ensure high availability.
[0167] As described herein, data center 1102 and / or one or more network devices 1106 may include storage devices and processing circuitry systems for performing computational tasks, such as controlling the flow of data internally and through communication network 1104. The processing circuitry systems may include software, hardware, or a combination thereof. For example, the processing circuitry systems may include memory containing executable instructions and a processor (e.g., a microprocessor) for executing those instructions. The memory may correspond to any suitable memory device type or a set of memory devices configured to store instructions. Non-limiting examples of suitable memory devices include flash memory, random access memory (RAM), read-only memory (ROM), variations thereof, combinations thereof, or similar technologies. In certain embodiments, the memory and processor may be integrated into the same device, such as a microprocessor with integrated memory. Additionally or alternatively, the processing circuitry systems may also include hardware components such as application-specific integrated circuits (ASICs). Other non-limiting examples of processing circuitry systems include integrated circuit (IC) chips, CPUs, GPUs, microprocessors, field-programmable gate arrays (FPGAs), collections of logic gates or transistors, resistors, capacitors, inductors, and diodes. Part or all of the processing circuitry can be housed on a printed circuit board (PCB) or an assembly of PCBs. It should be understood that any suitable type of electrical component or assembly of electrical components can be appropriately included in the processing circuitry.
[0168] Each network device 1106 may include or be connected to a power distribution unit (PDU) that powers the power-consuming components of the network device 1106 (e.g., individual network switches within each network device). Related PDUs may take the form of multiple power interfaces (e.g., AC and / or DC power outlets) integrated on a single power strip, which requires cumbersome power cable wiring and complex assembly / disassembly. Furthermore, these power strips are not easily replaceable. For example, replacing or repairing one or more outlets or other components of a PDU may require shutting down the entire strip, disconnecting each power cable, installing a new power strip, and then reconnecting the power cables to each outlet. This process is time-consuming and may result in unnecessary replacement of normally functioning components of the PDU integrated with malfunctioning components. To address these and other shortcomings of the prior art, exemplary embodiments provide PDUs capable of blind-mating or direct connection to power-consuming components within the network device 1106. Furthermore, the PDUs in exemplary embodiments may employ a modular design and work collaboratively to power devices, so that replacing or repairing a malfunctioning PDU module does not interrupt the power supply to devices connected to other normally functioning PDU modules in the network device.
[0169] Furthermore, although not explicitly shown in the figures, this disclosure envisions that data center 1102 and one or more network devices 1106 may include one or more communication interfaces to facilitate wired and / or wireless communication between them and with other unshown elements in network architecture 1100. These communication interfaces may include a variety of technologies, including but not limited to Ethernet ports, fiber optic connections, Wi-Fi® transceivers, Bluetooth® modules, and cellular communication modules, to enable integration and interoperability among the various components in network architecture 1100.
[0170] Furthermore, this disclosure envisions that network architecture 1100 may include other components and functions. For example, the network architecture may include, but is not limited to, additional processing units, dedicated accelerators (e.g., tensor processing units or TPUs), enhanced security modules, and redundant power supplies. The inclusion of these elements is intended to ensure that network architecture 1100 is robust, scalable, and capable of meeting a variety of operational requirements. Any variations, modifications, or adaptations of the elements falling within the spirit and scope of this disclosure are considered to be included in this disclosure. This includes any combination, sub-combination, or enhancement of the various elements to improve the performance, reliability, and efficiency of network architecture 1100.
[0171] Figure 12 A data center 1200 is illustrated according to at least one embodiment. In at least one embodiment, the data center 1200 includes, but is not limited to, a data center infrastructure layer 1210, a framework layer 1220, a software layer 1230, and an application layer 1240.
[0172] In at least one embodiment, the data center infrastructure layer 1210 may include a resource orchestrator 1212, packet computing resources 1214, and node computing resources (“nodes CR”) 1216(1)-1216(N), where “N” represents any positive integer. In at least one embodiment, the nodes CR 1216(1)-1216(N) may include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including accelerators, field-programmable gate arrays (“FPGAs”), graphics processors, etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state drives or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules, and cooling modules, etc. In at least one embodiment, one or more of the nodes CR 1216(1)-1216(N) may be servers having one or more of the aforementioned computing resources.
[0173] In at least one embodiment, the packet computing resource 1214 may include multiple node CR packets housed within one or more racks (not shown), or multiple racks housed within data centers (not shown) in different geographical locations. Individual node CR packets within the packet computing resource 1214 may include packet computing, networking, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, several node CRs, including CPUs or processors, may be grouped within one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, the one or more racks may also include any number and any combination of power modules, cooling modules, and network switches.
[0174] In at least one embodiment, resource orchestrator 1212 may be configured or otherwise control one or more nodes CR 1216(1)-1216(N) and / or grouped computing resources 1214. In at least one embodiment, resource orchestrator 1212 may include a Software Design Infrastructure (“SDI”) management entity for data center 1200. In at least one embodiment, resource orchestrator 1212 may include hardware, software, or some combination thereof.
[0175] In at least one embodiment, the framework layer 1220 includes, but is not limited to, a job scheduler 1232, a configuration manager 1234, a resource manager 1236, and a distributed file system 1238. In at least one embodiment, the framework layer 1220 may include a framework of software 1252 supporting the software layer 1230 and / or one or more applications 1242 supporting the application layer 1240. In at least one embodiment, software 1252 or one or more applications 1242 may respectively include web-based service software or applications, such as software or applications provided by a cloud service provider. In at least one embodiment, the framework layer 1220 may be, but is not limited to, a free and open-source software web application framework, such as Apache Spark™ (hereinafter “Spark”), which can utilize the distributed file system 1238 for large-scale data processing (e.g., “big data”). In at least one embodiment, the job scheduler 1232 may include a Spark driver to facilitate the scheduling of workloads supported by the various layers of the data center 1200. In at least one embodiment, configuration manager 1234 may be able to configure different layers, such as software layer 1230 and framework layer 1220, including Spark and distributed file system 1238, to support large-scale data processing. In at least one embodiment, resource manager 1236 may be able to manage cluster or group computing resources mapped to or allocated to support distributed file system 1238 and job scheduler 1232. In at least one embodiment, cluster or group computing resources may include group computing resources 1214 at data center infrastructure layer 1210. In at least one embodiment, resource manager 1236 may coordinate with resource orchestrator 1212 to manage these mapped or allocated computing resources.
[0176] In at least one embodiment, the software 1252 included in the software layer 1230 may include the software used by at least a portion of the nodes CR 1216(1)-1216(N) of the framework layer 1220, the grouped computing resources 1214, and / or the distributed file system 1238. One or more types of software may include, but are not limited to, internet web search software, email virus scanning software, database software, and streaming video content software.
[0177] In at least one embodiment, the application layer 1240 may include one or more applications 1242 that are available for use by at least a portion of the nodes CR 1216(1)-1216(N) of the framework layer 1220, the group computing resources 1214, and / or the distributed file system 1238. The one or more application types may include, but are not limited to, CUDA applications, 5G network applications, artificial intelligence applications, data center applications, and / or variations thereof.
[0178] In at least one embodiment, any of the configuration manager 1234, resource manager 1236, and resource orchestrator 1212 can implement any number and type of self-modification operations based on any amount and type of data obtained in any technically feasible manner. In at least one embodiment, the self-modification operations can enable the data center operator of data center 1200 to avoid making potentially erroneous configuration decisions and may avoid underutilization and / or use of poorly performing portions of the data center.
[0179] Figures 13A-13B A data center architecture (“architecture”) 1300 is depicted according to at least one embodiment. Architecture 1300 may represent a cluster including servers and corresponding graphics processing units (GPUs). For example, architecture 1300 may represent an OVX L40S GPU cluster provided by NVIDIA®. By default, this cluster may include 128 servers, totaling 512 GPUs. However, architecture 1300 is scalable, capable of scaling up to 2048 GPUs with minimal impact on performance expansion. Architecture 1300 is organized into 16 server racks (e.g., OVX L40S GPU server racks), one leaf-spine switch rack, and two management and storage racks. Each server rack can accommodate 8 compute server nodes, each equipped with 4 L40S GPUs, thus providing high-density GPU resources per rack. The cluster configuration may include 12 control server nodes and 12 storage server nodes, ensuring robust management and data processing capabilities. The default rack layout uses a single leaf-spine switch rack and assumes power redundancy within each rack. Alternatively, an alternative rack layout can be considered to maintain system reliability. The number of servers per rack can be determined based on available rack power supplies, ensuring secure and efficient operation. Each server can be configured with 2x200Gbps of single-node bandwidth, and the network design is flexible and scalable to meet ever-increasing capacity and bandwidth demands.
[0180] For example, architecture 1300 may include servers interconnected via compute architecture 1310 (e.g., a Spectrum-X 400G compute architecture). Compute architecture 1310 employs a leaf-spine network topology designed for high performance, utilizing orbit-optimized non-blocking 128-port leaf switches 1312-1 to 1312-4 and spine switches 1314-1 and 1314-2. Leaf switches 1312-1 to 1312-4 provide 64 downlinks and 64 uplinks to scalable units (SUs) 1330-1 to 1330-4, where each SU may include 4 racks and 32 GPU nodes. Spine switches 1314-1 and 1314-2 are the core backbone switches connecting leaf switches 1312-1 to 1312-4. Leaf switches 1312-1 to 1312-4 provide aggregated connections from GPU nodes to spine switches 1314-1 and 1314-2. This architecture supports high-throughput and low-latency communication between nodes. The computing architecture 1310 can employ a fully non-blocking fat-tree topology, which is well-suited for high-performance applications and is based on Remote Direct Memory Access (RDMA) technology to minimize latency and maximize bandwidth.
[0181] The Compute Architecture 1310 employs a leaf-ridge network topology, combined with Track-Optimized Data Processing Units (DPUs). These DPUs enable SmartNICs to provide hardware acceleration for network and storage functions. The combination of Compute Architecture 1310 and the Track-Optimized DPUs can be designed to provide the shortest and most efficient network paths for inter-GPU communication, whether within a single node or across multiple nodes. This architecture is highly effective for multi-GPU workloads and supports both Top-of-Rack (ToR) and End-of-Row (EoR) switch configurations, enabling flexible deployment options to adapt to various data center layouts. Compute Architecture 1310 provides a scalable and robust network infrastructure capable of supporting clusters of all sizes, from small to very large, while maintaining predictable and consistent performance. By leveraging high-speed, low-latency switching and optimized routing algorithms, Compute Architecture 1310 can be designed to maximize aggregate bandwidth and minimize network latency for GPU interconnects, both within servers and across tracks.
[0182] As another example, in terms of management and storage, architecture 1300 can include management and storage fabric 1320 (e.g., a 200G Ethernet in-band management and storage fabric). Management and storage fabric 1320 can utilize two 128-port leaf switches 1322-1 and 1322-2 to support high availability (HA). All nodes are redundantly connected to these two leaf switches 1322-1 and 1322-2, ensuring uninterrupted operation even in the event of switch failure. Management and storage fabric 1320 is architecturally independent of compute fabric 1310, which allows for optimization of storage and application performance. For example, in addition to connecting to SU 1330-1 through 1330-4, leaf switches 1322-1 and 1322-2 can also connect to control nodes and storage nodes 1340. In a typical configuration, there may be 12 control nodes and 12 storage nodes. The configuration and orchestration of the cluster are managed by the Basic Command Manager (BCM), which is responsible for configuring all server nodes and network devices. BCM supports automatic deployment of Kubernetes or SLURM (Simple Linux Resource Management Utility) clusters, enabling dynamic resource management, workload scheduling, and rapid scaling of cluster resources.
[0183] For example, such as Figure 13B As shown, SU 1330-1 may include DPUs 1332-1 and 1332-2 for its GPU network, DPU 1334-1 for its CPU network, and GPU server 1336-1. SU 1330-2 may include DPUs 1332-3 and 1332-4 for its GPU network, DPU 1334-2 for its CPU network, and GPU server 1336-2. SU 1330-3 may include DPUs 1332-5 and 1332-6 for its GPU network, DPU 1334-3 for its CPU network, and GPU server 1336-3. SU 1330-4 may include DPUs 1332-7 and 1332-8 for its GPU network, DPU 1334-4 for its CPU network, and GPU server 1336-4. Each DPU may implement a SmartNIC to provide hardware acceleration for networking and storage functions. For computing architecture 1310, leaf switch 1312-1 can connect to DPUs 1332-1 and 1332-3, leaf switch 1312-2 can connect to DPUs 1332-2 and 1332-4, leaf switch 1312-3 can connect to DPUs 1332-5 and 1332-7, and leaf switch 1312-4 can connect to DPUs 1332-6 and 1332-8. For management and storage architecture 1320, leaf switches 1322-1 and 1322-2 can connect to DPUs 1334-1 through 1334-4, respectively. Although Figure 13BNot shown, but each of SU 1330-1 through 1330-4 may include its own Intelligent Platform Management Interface (IPMI), and each of leaf switches 1322-1 and 1322-2 may connect to each IPMI. Although Figures 13A-13B Not shown, but control and storage node 1340 may include storage node, control node and IPMI, and each of leaf switches 1322-1 and 1322-2 may be connected to storage node, control node and IPMI.
[0184] For inference workloads, east / west (E / W) networks are typically unnecessary because most pre-trained models can fit in the memory of a single GPU. However, models with over approximately 13 billion parameters may require model parallelization and the use of multiple GPUs, in which case E / W network connections may be necessary. For typical inference applications, E / W networks can be omitted, simplifying network design and reducing cost and complexity. It's important to note that while tensor parallelism allows large models to be distributed across multiple GPUs, its performance may be lower than running the model entirely on a single GPU due to increased communication overhead. For optimal performance, models with up to 13 billion parameters are typically supported, although even larger models can be supported with appropriate parallelization strategies and network design.
[0185] Inference and prediction tasks within a cluster are highly dynamic, typically requiring automatic scaling of the compute engine and inference server instances. Fine-tuning can optimize which models are loaded onto which inference servers, ensuring efficient resource utilization and rapid response to changing workload demands. Other technical considerations may include monitoring GPU utilization, managing model versions, and implementing load balancing strategies to further improve inference performance and reliability.
[0186] Other variations also fall within the spirit and scope of this disclosure. Therefore, while the disclosed technology can be modified and alternatively constructed in various ways, certain exemplary embodiments have been shown in the accompanying drawings and described in detail above. However, it should be understood that this disclosure is not intended to limit its scope to the specific forms or multiple specific forms disclosed, but rather to cover all modifications, alternative constructions, and equivalents that fall within the spirit and scope of this disclosure as defined in the appended claims.
[0187] In describing the disclosed embodiments (particularly in the following claims), the terms “a,” “an,” “the,” and similar pronouns should be interpreted to cover both singular and plural forms unless otherwise stated herein or the context clearly indicates otherwise, and should not be considered as limiting the terms. The terms “comprising,” “having,” “including,” and “containing” should be interpreted as open-ended terms (meaning “including but not limited to”) unless otherwise stated. When the word “connection” is unmodified and refers to a physical connection, it should be interpreted as partially or wholly included, attached to, or linked together, even if intermediates are present. The enumeration of numerical ranges herein is intended only as a convenient method to individually refer to each individual value falling within a range, unless otherwise stated herein, and each individual value is incorporated into the specification as if it had been individually enumerated herein. In at least one embodiment, unless otherwise stated herein or the context clearly indicates otherwise, the terms “set” (e.g., “a group of items”) or “subset” should be interpreted as a non-empty set comprising one or more members. Furthermore, unless otherwise stated or the context indicates otherwise, a “subset” of a set does not necessarily mean a proper subset of the set, but rather a subset and the set can be equal.
[0188] Conjunctive phrases such as “at least one of A, B, and C” or “at least one of A, B, and C” are generally understood, depending on the context, to indicate that an item, term, etc., can be A, B, or C, or any non-empty subset of the set A, B, and C, unless explicitly stated otherwise or contradicted by the context. For example, in an illustrative example of a set comprising three members, the conjunctive phrases “at least one of A, B, and C” and “at least one of A, B, and C” refer to any of the following sets: {A}, {B}, {C}, {A,B}, {A,C}, {B,C}, {A,B,C}. Therefore, such conjunctive phrases are generally not intended to imply that some embodiments require at least one A, at least one B, and at least one C simultaneously. Furthermore, unless explicitly stated otherwise or contradicted by the context, the term “multiple” indicates a plural state (e.g., “multiple items” means multiple items). In at least one embodiment, the number of multiple items is at least two, but the number can be more if explicitly stated or determined by the context. Furthermore, unless otherwise stated or the context clearly indicates otherwise, the phrase “based on” means “at least partially based on” rather than “completely based on”.
[0189] The operations of the processes described herein can be performed in any suitable order unless otherwise stated herein or there is a clear contradiction in the context. In at least one embodiment, processes such as those described herein (or variations and / or combinations thereof) are executed under the control of one or more computer systems configured with executable instructions and implemented in the form of code (e.g., executable instructions, one or more computer programs, or one or more application programs) that execute cooperatively on one or more processors, or implemented by hardware or a combination thereof. In at least one embodiment, the code is stored on a computer-readable storage medium, for example, in the form of a computer program comprising multiple instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-volatile computer-readable storage medium that excludes transient signals (e.g., propagating transient electrical or electromagnetic transmissions) but includes non-volatile data storage circuitry systems (e.g., buffers, caches, and queues) within transient signal transceivers. In at least one embodiment, code (e.g., executable code or source code) is stored on a collection of one or more non-volatile computer-readable storage media that store executable instructions (or other memory for storing executable instructions) that, when executed by one or more processors of a computer system (i.e., as a result of execution), cause the computer system to perform the operations described herein. In at least one embodiment, the collection of non-volatile computer-readable storage media includes multiple non-volatile computer-readable storage media, and one or more individual non-volatile storage media do not contain all the code, while the multiple non-volatile computer-readable storage media collectively store all the code. In at least one embodiment, the executable instructions are executed in such a manner that different instructions are executed by different processors.
[0190] Therefore, in at least one embodiment, the computer system is configured to implement one or more services that individually or collectively perform the operations of the processes described herein, and the computer system is configured with corresponding hardware and / or software to enable the execution of these operations. Furthermore, the computer system implementing at least one embodiment of this disclosure may be a single device; in another embodiment, it may be a distributed computer system comprising multiple devices operating in different ways to cause the distributed computer system to perform the operations described herein, and such that a single device does not perform all operations.
[0191] Any and all examples or exemplary language provided herein (e.g., "for example") are used only to better illustrate embodiments of this disclosure and, unless otherwise stated, do not constitute a limitation on the scope of this disclosure. No language in the specification should be construed as indicating that any unstated element is essential to the practice of this disclosure.
[0192] In the specification and claims, the terms “coupled” and “connected” and their derivatives may be used. It should be understood that these terms are not synonymous with each other. More specifically, in some examples, “connected” or “coupled” may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. “Coupled” may also indicate that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.
[0193] Unless otherwise expressly stated, it is understood that throughout this specification, terms such as “processing,” “calculation,” “operation,” and “determine” refer to the operations and / or processes of a computer or computing system or similar electronic computing device that manipulate and / or transform data represented as physical quantities (e.g., electronic quantities) in the registers and / or memory of the computing system to generate data also represented as physical quantities, which are stored in the memory, registers, or other such information storage, transmission, or display devices of the computing system.
[0194] Similarly, the term "processor" can refer to any device or part of a device that processes electronic data from registers and / or memory and transforms that electronic data into other electronic data that can be stored in registers and / or memory. A "computing platform" can include one or more processors. As used herein, a "software" process can include, for example, software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Furthermore, each process can refer to multiple processes for executing instructions sequentially or in parallel, continuously or intermittently. In at least one embodiment, the terms "system" and "method" are used interchangeably herein, as a system can contain one or more methods, and a method can be considered a system.
[0195] This document may refer to obtaining, acquiring, receiving analog or digital data, or inputting analog or digital data into a subsystem, computer system, or computer-implemented machine. In at least one embodiment, the process of obtaining, acquiring, receiving, or inputting analog and digital data can be implemented in various ways, for example, by receiving data as a parameter of a function call or application programming interface (API) call. In at least one embodiment, the process of obtaining, acquiring, receiving, or inputting analog or digital data can be implemented by transmitting data through a serial or parallel interface. In at least one embodiment, the process of obtaining, acquiring, receiving, or inputting analog or digital data can be implemented by transmitting data from a providing entity to a receiving entity via a computer network. In at least one embodiment, the provisioning, outputting, transmitting, sending, or presenting analog or digital data may also be mentioned. In various examples, the process of providing, outputting, transmitting, sending, or presenting analog or digital data can be implemented by transmitting data as an input or output parameter of a function call, a parameter of an application programming interface (API), or a parameter of an inter-process communication mechanism.
[0196] Although exemplary embodiments of the technology described herein are presented in this document, other architectures may be used to implement the described functionality, and all such architectures are within the scope of this disclosure. Furthermore, while specific assignments of responsibilities may have been defined above for ease of description, various functions and responsibilities may be assigned and divided in different ways depending on the specific circumstances.
[0197] Furthermore, although this document has described the subject matter using language specific to structural features and / or method steps, it should be understood that the subject matter claimed in the appended claims is not necessarily limited to the specific features or steps described. Rather, the specific features and steps described are disclosed only as exemplary forms for implementing the claims.
[0198] This application provides the following: 1) A system comprising: Memory; and At least one processing device, operatively coupled to the memory, is used for: - Receive input data indicating the original quantum state of the quantum system, which is modified according to the amount of noise added by the non-unitary quantum channel; - Provide input data to a machine learning model, which is trained to predict the original quantum state, thereby mitigating errors by simulating the inverse of the non-unitary quantum channel; and - Obtain output data indicating the original quantum state from the machine learning model.
[0199] 2) According to the system described in 1), wherein the machine learning model is a diffusion model.
[0200] 3) According to the system described in 2), wherein the machine learning model includes a recurrent neural network, wherein the machine learning model is trained by simulating one or more non-unitary quantum channels, including the non-unitary quantum channel, based on one or more weak measurement sequences of one or more quantum states; and each weak measurement sequence of the quantum state includes multiple weak measurements of the quantum state acquired in chronological order, wherein the multiple weak measurements do not cause the wave function corresponding to the quantum state to completely collapse.
[0201] 4) The system according to 2), wherein the machine learning model is trained to perform a back diffusion process to predict the back evolution of noise added to the original quantum state according to the forward diffusion process.
[0202] 5) The system according to 4), wherein the forward diffusion process includes one or more iterations, and wherein each of the one or more iterations of the forward diffusion process is performed by: Initialize the first quantum state represented by the first density matrix associated with the Hilbert space; Based on the density matrix, determine whether the purity of the first quantum state satisfies the threshold condition defined by the maximum mixing quantum state based on the Hilbert space; and In response to determining that the purity of the first quantum state satisfies the threshold condition, the first density matrix is modified to obtain a second density matrix representing the second quantum state.
[0203] 6) The system according to 1), wherein the at least one processing device is further configured to perform at least one quantum computing operation using the output data.
[0204] 7) The system according to 6), wherein the at least one quantum computing operation includes: sending data to the quantum computing system indicating the original quantum state predicted by the machine learning model.
[0205] 8) The system according to 1), wherein the input data includes a modified message corresponding to the original message after modification by the noise amount, and wherein the output data includes a prediction of the original message.
[0206] 9) The system according to 1), wherein the machine learning model is used to identify the type of noise added through the non-unitary quantum channel.
[0207] 10) The system according to 1), wherein the input data corresponds to one or more optical qubits.
[0208] 11) The system according to 1) further includes: The first quantum computing system; and A second quantum computing system associated with the one or more processing devices; The original quantum state is transmitted by the first quantum computing system to one or more quantum computing systems, including the second quantum computing system, through one or more non-unitary quantum channels, including the non-unitary quantum channel.
[0209] 12) A system comprising: Memory; and At least one processing device, operatively coupled to the memory, is used for: -A measurement sequence that enables the acquisition of the quantum state of a quantum system, wherein the measurement sequence comprises multiple measurements captured in time order, and wherein each measurement in the measurement sequence simulates the corresponding amount of noise added to the original quantum state at the corresponding time step via a non-unitary quantum channel; - Generate a training dataset including the measurement sequence of the quantum state; and - A machine learning model is trained using the training dataset to predict the original quantum state, thereby mitigating errors by simulating the inverse of the non-unitary quantum channel.
[0210] 13) The system according to 12), wherein each measurement in the measurement sequence of the quantum state is a weak measurement of the quantum state, the weak measurement not completely collapses the wave function corresponding to the quantum state.
[0211] 14) The system according to 12), wherein the machine learning model is a diffusion model.
[0212] 15) The system according to 14), wherein the machine learning model comprises a recurrent neural network.
[0213] 16) The system according to 14), wherein, in order to train the machine learning model to predict the original quantum state by simulating the inverse of the non-unitary quantum channel, the at least one processing device is used to train the machine learning model to perform a back-diffusion process to predict the back evolution of noise added to the original quantum state according to the forward diffusion process.
[0214] 17) The system according to 16), wherein: The forward diffusion process includes one or more iterations; and To train the machine learning model to predict the original quantum state by simulating the inverse of the non-unitary quantum channel, the at least one processing device is further configured to perform each of the one or more iterations of the forward diffusion process by performing the following operations: Initialize the first quantum state represented by the first density matrix associated with the Hilbert space; Based on the density matrix, determine whether the purity of the first quantum state satisfies the threshold condition defined by the maximum mixing quantum state based on the Hilbert space; and In response to determining that the purity of the first quantum state satisfies the threshold condition, the first density matrix is modified to obtain a second density matrix representing the second quantum state.
[0215] 18) The system according to 17), wherein: Determining whether the purity of the first quantum state satisfies the threshold condition includes: determining whether the trace determined by the first density matrix is greater than or equal to a threshold, the threshold being defined based on the trace determined for the maximally mixed quantum state; and Modifying the first density matrix to obtain the second density matrix includes applying a randomly selected basis measurement to the first density matrix.
[0216] 19) A system comprising: The first quantum computing system, used to measure the primitive quantum state of a quantum system; A second quantum computing system, which is communicatively coupled to the first quantum computing system via an interconnect, wherein the second quantum computing system is configured to: - Receive input data from the first quantum computing system, the input data indicating the original quantum state of the quantum system, the original quantum state being modified by the amount of noise added by the non-unitary quantum communication channel; - Providing input data to a machine learning model trained to predict the primitive quantum state, thereby mitigating errors by simulating the inverse of the non-unitary quantum communication channel; and - Obtain output data indicating the original quantum state from the machine learning model.
[0217] 20) The system according to 19), wherein the machine learning model includes a recurrent neural network, wherein the machine learning model is trained by simulating one or more non-unitary quantum channels, including the non-unitary quantum channel, based on one or more weak measurement sequences of one or more quantum states; and each weak measurement sequence of the quantum state includes multiple weak measurements of the quantum state acquired in chronological order, the multiple weak measurements not causing the wave function corresponding to the quantum state to completely collapse.
[0218] 21) A method comprising: - Receive input data indicating the original quantum state of the quantum system, which is modified according to the amount of noise added by the non-unitary quantum channel; - Provide input data to a machine learning model, which is trained to predict the original quantum state, thereby mitigating errors by simulating the inverse of the non-unitary quantum channel; and - Obtain output data indicating the original quantum state from the machine learning model.
[0219] 22) The system according to 21), wherein the machine learning model is a diffusion model, the machine learning model includes a recurrent neural network, wherein the machine learning model is trained by simulating one or more non-unitary quantum channels, including the non-unitary quantum channel, based on a sequence of one or more weak measurements of one or more quantum states, wherein the one or more weak measurements do not cause the wavefunction corresponding to the quantum state to completely collapse; and Each weak measurement sequence of a quantum state comprises multiple weak measurements of the quantum state acquired in chronological order.
Claims
1. A system comprising: Memory; as well as At least one processing device, said at least one processing device being operatively coupled to the memory, is used for: Input data is provided to a machine learning model, the input data indicating the original quantum state of the quantum system, the original quantum state being modified by the amount of noise added by a non-unitary quantum communication channel, wherein the machine learning model is trained to predict the original quantum state, thereby performing error mitigation by simulating the inverse of the non-unitary quantum communication channel to reverse one or more errors caused by the non-unitary quantum communication channel; Obtain output data indicating the original quantum state from the machine learning model; and The quantum computing system performs at least one quantum computing operation based on the output data, the at least one quantum computing operation involving at least one of error mitigation or quantum state preparation.
2. The system of claim 1, wherein, The machine learning model is a diffusion model.
3. The system of claim 1, wherein, The machine learning model includes a recurrent neural network or an encoder-decoder network.
4. The system as claimed in claim 1, wherein, The machine learning model is trained to perform a reverse diffusion process to predict the reverse evolution of noise added to the original quantum state according to the forward diffusion process.
5. The system as described in claim 4, wherein, The forward diffusion process includes one or more iterations, wherein each iteration of the one or more iterations of the forward diffusion process is performed by the following operation: Initialize the first quantum state represented by the first density matrix associated with the Hilbert space; Based on the density matrix, determine whether the purity of the first quantum state satisfies the threshold condition defined by the maximum mixing quantum state based on the Hilbert space; and In response to determining that the purity of the first quantum state satisfies the threshold condition, the first density matrix is modified to obtain a second density matrix representing the second quantum state.
6. The system as claimed in claim 1, wherein, The machine learning model includes a visual transformer architecture that divides the density matrix representation of the quantum state into multiple blocks, wherein each of the multiple blocks is embedded in a feature representation to perform spatial correlation learning on the density matrix.
7. The system as claimed in claim 1, wherein, The at least one quantum computing operation includes at least one of the following: executing a quantum algorithm or initializing the quantum computing system using the original quantum state.
8. The system of claim 1, wherein, The at least one processing device is further configured to couple the quantum system to an auxiliary quantum system to achieve a unitary operation that approximates the inverse of the non-unitary quantum communication channel, wherein a partial trace operation is performed on the auxiliary quantum system.
9. The system as claimed in claim 1, wherein, The input data includes a modified message corresponding to the original message, modified by the noise level, and the output data includes a prediction of the original message.
10. The system of claim 1, wherein, The machine learning model is used to identify the types of noise added through the non-unitary quantum communication channel.
11. The system of claim 1, wherein, The input data corresponds to one or more optical qubits.
12. The system of claim 1, wherein, The at least one processing device is further configured to receive the input data from the quantum computing system via the non-unitary quantum communication channel by processing one or more physical signals.
13. The system of claim 1, wherein: The input data includes multiple local density matrices, each of which corresponds to a corresponding qubit in an entangled quantum register and is obtained by performing a partial trace operation on the global density matrix. The machine learning model is trained to reconstruct the global quantum state from the plurality of local density matrices by simulating the inverse of the non-unitary quantum communication channel; and The output data indicates the global quantum state.
14. A system comprising: Memory; as well as At least one processing device, operatively coupled to the memory, for: This allows for the acquisition of a measurement sequence of the quantum state of a quantum system, wherein the measurement sequence comprises multiple measurements captured in time order, wherein each measurement in the measurement sequence corresponds to a corresponding segment of a communication line, and wherein each measurement in the measurement sequence simulates a corresponding amount of noise added to the original quantum state at a corresponding time step through a non-unitary quantum communication channel. Generate a training dataset that includes the measurement sequence of the quantum state; as well as The training dataset is used to train a machine learning model to predict the original quantum state, thereby performing at least one of error mitigation or quantum state preparation by simulating the inverse of the non-unitary quantum communication channel to reverse one or more errors caused by the non-unitary quantum communication channel.
15. The system of claim 14, wherein, Each measurement in the measurement sequence of the quantum state is a weak measurement of the quantum state, and the weak measurement does not completely collapse the wave function corresponding to the quantum state.
16. The system of claim 14, wherein, The machine learning model includes at least one of the following: diffusion model, recurrent neural network, encoder-decoder network, or visual transformer architecture.
17. The system of claim 14, wherein, In order to train the machine learning model to predict the original quantum state by simulating the inverse of the non-unitary quantum communication channel, the at least one processing device is used to train the machine learning model to perform a reverse diffusion process to predict the reverse evolution of noise added to the original quantum state according to the forward diffusion process.
18. The system of claim 17, wherein: The forward diffusion process includes one or more iterations; and To train the machine learning model to predict the original quantum state by simulating the inverse of the non-unitary quantum communication channel, the at least one processing device is further configured to perform each of the one or more iterations of the forward diffusion process by performing the following operations: Initialize the first quantum state represented by the first density matrix associated with the Hilbert space; Based on the density matrix, determine whether the purity of the first quantum state satisfies the threshold condition defined by the maximum mixed quantum state based on the Hilbert space; as well as In response to determining that the purity of the first quantum state satisfies the threshold condition, the first density matrix is modified to obtain a second density matrix representing the second quantum state.
19. The system of claim 18, wherein: Determining whether the purity of the first quantum state satisfies the threshold condition includes: determining whether the trace determined by the first density matrix is greater than or equal to a threshold, the threshold being defined based on the trace determined for the maximally mixed quantum state; and Modifying the first density matrix to obtain the second density matrix includes applying a randomly selected basis measurement to the first density matrix.
20. A method comprising: Input data is provided to a machine learning model, the input data indicating the original quantum state of the quantum system, the original quantum state being modified by the amount of noise added by a non-unitary quantum communication channel, wherein the machine learning model is trained to predict the original quantum state, thereby performing error mitigation by simulating the inverse of the non-unitary quantum communication channel to reverse one or more errors caused by the non-unitary quantum communication channel; Obtain output data indicating the original quantum state from the machine learning model; and The quantum computing system performs at least one quantum computing operation based on the output data, the at least one quantum computing operation involving at least one of error mitigation or quantum state preparation.