Image diffusion method and device, electronic equipment and medium
By constructing a quantum classic hybrid diffusion model, combining downsampling network, upsampling network and quantum network layer, the problem of low image diffusion accuracy in quantum devices is solved, and efficient image diffusion generation in quantum devices is achieved.
Patent Information
- Application Number
- CN202510392554.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-18
AI Technical Summary
The existing quantum diffusion model is in the NISQ period and cannot achieve large-scale computing in quantum devices, resulting in low image diffusion accuracy.
The quantum classical hybrid diffusion model is adopted to build a quantum classical hybrid diffusion model through the combination of downsampling network, upsampling network and quantum network layer. The downsampling network is used to reduce the dimensions of input image data. The quantum network layer performs linear processing in the quantum field and restores the dimensions through the upsampling network to achieve image diffusion.
The accuracy of image diffusion in quantum devices is improved, and the mixing advantages of classical computing and quantum computing is used to improve the rate and accuracy of image diffusion.
Smart Images

Figure CN120339104A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of image diffusion, and specifically relates to a method and device for image diffusion, an electronic device, and a medium. Background Art
[0002] Currently, the diffusion model for images has become a gradually popular generative model. This diffusion model is widely applied in various fields such as computer vision, natural language processing, and audio generation. In recent years, with the development of deep learning, the diffusion model has achieved remarkable results in content generation tasks due to its advantages in generation quality and stability. The core idea of the diffusion model is to simulate the process of gradually adding noise to the data distribution and generate new data from random noise by learning the inverse operation of this process. Compared with traditional generative adversarial networks (GANs), the diffusion model avoids the instability and mode collapse problems brought by adversarial training and generates high-quality and detail-rich images and data through the way of gradually denoising.
[0003] The working process of the diffusion model is usually divided into two processes: forward diffusion and reverse diffusion. In the forward diffusion process, the model starts from the original data and gradually adds noise to the data until the data is completely transformed into pure noise close to the Gaussian distribution. In the reverse diffusion process, the model learns how to gradually remove the interference from the noise and reconstruct samples similar to the original data distribution. Through this gradually restored process, the diffusion model can generate data that conforms to the natural distribution, and the generation process is smoother and more natural.
[0004] With the development of quantum technology, those skilled in the art are more inclined to construct a quantum diffusion model to simulate the process of gradual evolution of data through a quantum diffusion process, so as to be able to capture the complex distribution of the data. Compared with the traditional diffusion model, the quantum diffusion model can better simulate the non-equilibrium state of the data. It has higher generation efficiency and diversity control ability in the fields of computer vision, natural language processing, and quantum simulation.
[0005] However, currently quantum computing is in the NISQ (Noisy Intermediate-Scale Quantum) era. The quantum devices in this era cannot achieve large-scale computing, while the diffusion of images requires large-scale computing, resulting in that the pure quantum diffusion model cannot be implemented in quantum devices, leading to a low accuracy rate of image diffusion. Summary of the Invention
[0006] To solve the above technical problems, embodiments of the present application provide a method and apparatus for image diffusion, an electronic device, and a storage medium to improve the accuracy of image diffusion.
[0007] According to one aspect of the embodiments of the present application, there is provided a method for image diffusion, including: obtaining first image data corresponding to a first noisy image to be diffused; inputting the first image data into a preset quantum-classical hybrid diffusion model to obtain a first predicted noise corresponding to the first noisy image; the quantum-classical hybrid diffusion model includes a downsampling network, an upsampling network, and a quantum network layer; the downsampling network and the upsampling network are connected through the quantum network; and generating an image according to the first image data and the first predicted noise.
[0008] In some embodiments, the obtaining first image data corresponding to a first noisy image to be diffused includes: obtaining the first noisy image to be diffused; and obtaining the first image data corresponding to the first noisy image.
[0009] In some embodiments, the quantum-classical hybrid diffusion model is obtained through the following steps: constructing an initial diffusion model; the initial diffusion model includes an initial downsampling network, an initial upsampling network, and an initial quantum network layer; the initial downsampling network and the initial upsampling network are connected through the initial quantum network layer; obtaining a raw image dataset; the raw image dataset includes image data of multiple images to be trained; adding noise to the image data in the raw image dataset to obtain a training dataset; and inputting the training dataset into the initial diffusion model for training to obtain the quantum-classical hybrid diffusion model.
[0010] In some embodiments, the downsampling network is used to reduce the dimension of the input image data; the quantum network layer is used to perform linear processing on the image data with reduced dimensions in the quantum domain; and the upsampling network is used to restore the dimension of the linearly processed image data.
[0011] In some embodiments, the quantum network layer is used to encode the image data with reduced dimensions into a quantum state to obtain a quantum encoding, then perform linear processing on the quantum encoding using a preset parameterized quantum circuit, and then measure the linearly processed quantum encoding to obtain the linearly processed image data.
[0012] In some embodiments, the encoding the image data with reduced dimensions into a quantum state to obtain a quantum encoding includes: padding the image data with reduced dimensions; normalizing the padded image data; and encoding the normalized image data into the amplitude of quantum bits to obtain the quantum encoding.
[0013] In some embodiments, the image generation based on the first image data and the first predicted noise includes: obtaining second image data corresponding to a second denoised image according to the first image data and the first predicted noise; the second denoised image is an image obtained by removing the first predicted noise from the first denoised image; inputting the second image data into a preset quantum-classical hybrid diffusion model to obtain a second predicted noise corresponding to the second denoised image; and performing image generation according to the second image data and the second predicted noise.
[0014] According to one aspect of the embodiments of the present application, there is provided an apparatus for image diffusion, including: an acquisition module configured to acquire first image data corresponding to a first denoised image to be diffused; an input module configured to input the first image data into a preset quantum-classical hybrid diffusion model to obtain a first predicted noise corresponding to the first denoised image; the quantum-classical hybrid diffusion model includes a downsampling network, an upsampling network, and a quantum network layer; the downsampling network and the upsampling network are connected through the quantum network; and a generation module configured to perform image generation according to the first image data and the first predicted noise.
[0015] According to one aspect of the embodiments of the present application, there is provided an electronic device, including: one or more processors; a storage device for storing one or more programs, which when executed by the one or more processors, cause the electronic device to implement the above-mentioned method for image diffusion.
[0016] According to one aspect of the embodiments of the present application, there is provided a computer-readable storage medium, on which computer-readable instructions are stored, which when executed by a processor of a computer, cause the computer to execute the above-mentioned method for image diffusion.
[0017] In the technical solution provided by the embodiments of the present application, first image data corresponding to a first noisy image to be diffused is obtained; the first image data is input into a preset quantum-classical hybrid diffusion model to obtain first predicted noise corresponding to the first noisy image; the quantum-classical hybrid diffusion model includes a downsampling network, an upsampling network, and a quantum network layer; the downsampling network and the upsampling network are connected by a quantum network; image generation is performed according to the first image data and the first predicted noise. In this way, the quantum-classical hybrid diffusion model composed of the downsampling network, the upsampling network, and the quantum network layer, which has both a classical network and a quantum network, is used to predict the noise in the first noisy image to obtain the first predicted noise, and then image generation is performed according to the first image data and the first predicted noise, realizing the removal of the noise in the first noisy image by using the quantum network. Compared with simply using a quantum model for noise prediction, the downsampling network can reduce the scale calculated by the quantum network layer, so that image diffusion generation can be performed in the quantum domain on existing quantum devices, thereby improving the accuracy of image diffusion.
[0018] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the accompanying drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts. In the drawings:
[0020] Figure 1 is a flowchart of a method for image diffusion shown in an exemplary embodiment of the present application;
[0021] Figure 2 is a schematic diagram of a quantum-classical hybrid diffusion model shown in an exemplary embodiment of the present application;
[0022] Figure 3 is a schematic diagram of a parameterized quantum circuit shown in an exemplary embodiment of the present application;
[0023] Figure 4 is a flowchart of step S120 in an exemplary embodiment shown in an exemplary embodiment of the present application;
[0024] Figure 5 is a schematic diagram of adding noise to image data shown in an exemplary embodiment of the present application;
[0025] Figure 6It is a flowchart of step S130 in an exemplary embodiment of the present application in an exemplary embodiment;
[0026] Figure 7 It is a block diagram of a device for image diffusion shown in an exemplary embodiment of the present application;
[0027] Figure 8 It shows a schematic structural diagram of an electronic device suitable for implementing the embodiments of the present application. Detailed implementation manners
[0028] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0029] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0030] The flowcharts shown in the drawings are only exemplary descriptions, and do not necessarily include all contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined. Therefore, the actual execution order may be changed according to the actual situation.
[0031] In the present application, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0032] See Figure 1 , Figure 1 It is a flowchart of a method for image diffusion shown in an exemplary embodiment of the present application.
[0033] As Figure 1 shown, in an exemplary embodiment, the method for image diffusion at least includes steps S110 to S130, which are introduced in detail as follows:
[0034] Step S110, obtaining first image data corresponding to a first noisy image to be diffused.
[0035] It should be noted that the first noisy image to be diffused is an image with random Gaussian noise.
[0036] Furthermore, obtaining the first image data corresponding to the first noisy image to be diffused includes: obtaining the first noisy image to be diffused; obtaining the first image data corresponding to the first noisy image. In this way, by obtaining the first noisy image to be diffused and then obtaining the first image data corresponding to the first noisy image, it is convenient to input the first image data into a quantum-classical hybrid diffusion model that has both a classical network and a quantum network, so as to be able to perform diffusion generation on the image in the quantum domain in existing quantum devices, thereby improving the accuracy of image diffusion.
[0037] It should be noted that the first image data is feature data that can characterize the first noisy image. It can be one or more of the feature data such as the gray-scale feature data of the first noisy image, the color feature data of the first noisy image, and the texture feature data of the first noisy image. Its specific content is not limited here.
[0038] Step S120, input the first image data into a preset quantum-classical hybrid diffusion model to obtain the first predicted noise corresponding to the first noisy image. The quantum-classical hybrid diffusion model includes a downsampling network, an upsampling network, and a quantum network layer; the downsampling network and the upsampling network are connected by a quantum network.
[0039] It should be noted that the preset quantum-classical hybrid diffusion model is used to predict the noise in the input image data. For example: the preset quantum-classical hybrid diffusion model can predict the noise in the first image data to obtain the first predicted noise corresponding to the first noisy image.
[0040] Please refer to Figure 2 , Figure 2 which is a schematic diagram of the quantum-classical hybrid diffusion model. As Figure 2 shown, the quantum-classical hybrid diffusion model includes a downsampling network, an upsampling network, and a quantum network layer; the downsampling network and the upsampling network are connected by a quantum network.
[0041] It should be noted that the downsampling network is used to reduce the dimension of the input image data; the quantum network layer is used to linearly process the image data with reduced dimension in the quantum domain; the upsampling network is used to restore the dimension of the linearly processed image data. In this way, the dimension of the input image data is reduced by the downsampling network, then the linearly processed image data with reduced dimension is processed linearly in the quantum domain by the quantum network layer, and finally the dimension of the linearly processed image data is restored by the upsampling network. The number of qubits used in the quantum network layer is reduced, making the quantum-classical hybrid diffusion model easier to run on devices in the NISQ era. At the same time, by adopting the idea of hybrid classical and quantum computing, the deterministic advantage of classical computing and the parallel advantage of quantum computing are exploited, improving the rate of image diffusion.
[0042] As Figure 2 shown, the quantum-classical hybrid diffusion model 201 adds a quantum network between the decoder and encoder of the classical UNet model, and it has a U-shaped structure. Among them, the decoder of the classical UNet model is the downsampling network. The encoder of the classical UNet model is the upsampling network.
[0043] The following Figure 2 arrows will be explained. As Figure 2 shown, in the downsampling network 202, the downward solid arrow represents the pooling layer, which is used to downsample the image data to reduce the dimension of the image data. For example, the pooling layer can reduce the horizontal and vertical dimensions of the image data by half.
[0044] The horizontal short arrows are convolutional layers, which are used to perform convolution on the image data. The convolutional layer consists of a 3×3 convolutional kernel. During the convolution process, padding needs to be performed on the convolved image to make the size of the image before and after convolution consistent. Two convolutional layers are combined into a convolutional kernel block, such as the first convolutional kernel block 203. The convolutional kernel block is used to extract the features of the image data and increase the number of channels. In this first convolutional kernel block 203, the image data with a dimension of 28×28 passes through the first convolution to obtain image data with 10 channels and a dimension of 28×28; then the convolved image data is convolved again to obtain image data with 10 channels and a dimension of 28×28.
[0045] The downsampling network 202 is composed of multiple convolutional kernel blocks as shown in the first convolutional kernel block 203.
[0046] In the upsampling network 204, the upward solid arrow represents the upsampling layer, which is used to upsample the image data to increase the dimension of the image data. For example, the pooling layer can double the horizontal and vertical dimensions of the image data.
[0047] Similar to the downsampling network, the horizontal short arrows are convolutional layers for convolving image data. The convolutional layer consists of a 3×3 convolutional kernel. During convolution, padding is required for the convolved image to keep the image size consistent before and after convolution. Two convolutional layers are combined into a convolutional kernel block, such as the second convolutional kernel block 205. The convolutional kernel blocks of the upsampling network 204 are used to extract features of the image data and reduce the number of channels. In this second convolutional kernel block 205, the image data with a dimension of 28×28 and 10 + 10 channels undergoes the first convolution to obtain image data with 10 channels and a dimension of 28×28; then the convolved image data is convolved again to obtain image data with 1 channel and a dimension of 28×28.
[0048] The upsampling network 204 is composed of multiple convolutional kernel blocks as shown in the second convolutional kernel block 205.
[0049] It should be noted that the convolutional kernel blocks with the same number of layers in the upsampling network 204 and the downsampling network 202 are correspondingly connected.
[0050] As Figure 2 shown, the horizontal hollow long arrows are replication layers. The replication layer copies the data before pooling into the upsampled image, keeping the data dimensions before pooling and after upsampling consistent. In the upsampling network 204, the hollow long strip shape represents the replicated image data.
[0051] The convolutional kernel block with the lowest number of layers in the upsampling network 204 and the downsampling network 202 is connected through the quantum network layer 206, and the convolutional kernel blocks of other layers are connected through the replication layer.
[0052] It should be noted that the quantum network layer is used to encode the image data with reduced dimensions into a quantum state to obtain a quantum encoding, then linearly process the quantum encoding using a preset parameterized quantum circuit, and then measure the linearly processed quantum encoding to obtain the linearly processed image data. In this way, the quantum network layer can perform linear operations on image data in the quantum domain, thereby improving the rate of linear operations.
[0053] Specifically, as Figure 2 shown, the quantum network layer 206 includes an encoding layer, a measurement layer, and a preset parameterized quantum circuit 207.
[0054] The downsampling network 202 and the parameterized quantum circuit 207 are connected through an encoding layer. The encoding layer is used to encode the image data output by the downsampling network 202, that is, the encoding layer is used to encode the image data with reduced dimensions into a quantum state to obtain a quantum encoding.
[0055] Further, encoding the image data with reduced dimensions into a quantum state to obtain a quantum encoding includes: padding the image data with reduced dimensions; normalizing the padded image data; and encoding the normalized image data into the amplitude of a qubit to obtain a quantum encoding. In this way, by padding the image data with reduced dimensions; normalizing the padded image data; and encoding the normalized image data into the amplitude of a qubit to obtain a quantum encoding, it is convenient to perform linear processing on the quantum encoding using a preset parameterized quantum circuit, realizing linear operations on image data in the quantum domain, thereby improving the rate of linear operations.
[0056] In this embodiment, the image data with reduced dimensions is the image data output by the downsampling network. The number of bits of this image is 3×3 = 9-dimensional data.
[0057] It should be noted that padding the image data with reduced dimensions means padding the image data with reduced dimensions with 0 to obtain data with a preset dimension. The preset dimension can be 16.
[0058] In some embodiments, the values in the normalized image data are between 0 and 1.
[0059] Encoding the normalized image data into the amplitude of a qubit to obtain a quantum encoding means encoding the normalized image data into the amplitude of a quantum state using a preset quantum amplitude encoding method, and adjusting the amplitude of the qubit through quantum gate operations so that the amplitude distribution of the quantum state corresponds to the normalized image data, obtaining the quantum encoding corresponding to the normalized image data.
[0060] In some embodiments, 16-dimensional image data can be encoded into 4-bit quantum encoding through quantum amplitude encoding.
[0061] The upsampling network 204 and the parameterized quantum circuit 207 are connected through a measurement layer, which is used to encode the quantum data output by the parameterized quantum circuit, that is, to measure the linearly processed quantum encoding to obtain the linearly processed image data.
[0062] It should be noted that the dimension of the image data obtained after measurement is 16. In some embodiments, only the first 9 dimensions of the 16-dimensional image data are input into the upsampling network.
[0063] The preset parameterized quantum circuit is used to perform linear processing on the quantum encoding.
[0064] Specifically, please refer to Figure 3 , Figure 3 which is a schematic diagram of a parameterized quantum circuit.
[0065] As Figure 3 shown, the parameterized quantum circuit consists of two layers of HardwareEfficientAnsatz (hardware-efficient trial solution).
[0066] As Figure 3 shown, the parameterized quantum circuit 207 includes a first layer of HardwareEfficientAnsatz 301 and a second layer of HardwareEfficientAnsatz 302.
[0067] Among them, each layer of HardwareEfficientAnsatz includes multiple single-qubit rotation gates, such as: RY gates (Y-axis rotation gates) and multiple CNOT gates (controlled-NOT gates).
[0068] The RY gate is used to rotate the single-bit quantum data in the quantum encoding. The rotation angles of each RY gate can be the same or different, and can be set according to specific circumstances, which will not be elaborated here.
[0069] The CNOT gate is a two-bit entanglement gate. One of the bits is the control bit, and the other bit is the target bit. If the control bit is |1>, then the X gate is applied to the target bit, that is, the target bit is flipped; if the control bit is |0>, then the target bit remains in its original state.
[0070] As Figure 3 shown, as Figure 3As shown, the first layer HardwareEfficientAnsatz301 includes the first RY gate 303, the second RY gate 304, the third RY gate 305, the fourth RY gate 306, the first CNOT gate 307, the second CNOT gate 308, and the third CNOT gate 309. The second layer HardwareEfficientAnsatz302 includes the fifth RY gate 310, the sixth RY gate 311, the seventh RY gate 312, the eighth RY gate 313, the fourth CNOT gate 314, the fifth CNOT gate 315, and the sixth CNOT gate 316. Specifically, the first RY gate 303 is respectively connected to the input end of the fifth RY gate 310 and one input end of the first CNOT gate 307; the second RY gate 304 is connected to the other input end of the first CNOT gate 307; the output end of the first CNOT gate 307 is respectively connected to the input end of the sixth RY gate 311 and one input end of the second CNOT gate 308; the third RY gate 305 is connected to the other input end of the second CNOT gate 308; the output end of the second CNOT gate 307 is respectively connected to the input end of the seventh RY gate 312 and one input end of the third CNOT gate 309; the fourth RY gate 306 is connected to the other input end of the third CNOT gate 309; the output end of the third CNOT gate 309 is connected to the input end of the eighth RY gate 313; the output end of the fifth RY gate 310 is connected to one input end of the fourth CNOT gate 314; the output end of the sixth RY gate 311 is connected to the other input end of the fourth CNOT gate 314; the output end of the fourth CNOT gate 314 is connected to one input end of the fifth CNOT gate 315; the output end of the seventh RY gate 312 is connected to the other input end of the fifth CNOT gate 315; the output end of the fifth CNOT gate 315 is connected to one input end of the sixth CNOT gate 316; the output end of the eighth RY gate 313 is connected to the other input end of the sixth CNOT gate 316.
[0071] Please refer to Figure 4 , Figure 4 is Figure 1 the flowchart of step S120 in the exemplary embodiment shown in. As Figure 4 shown, the process of obtaining the quantum-classical hybrid diffusion model may include steps S410 to S440, which are introduced in detail as follows:
[0072] Step S410, constructing an initial diffusion model; the initial diffusion model includes an initial downsampling network, an initial upsampling network, and an initial quantum network layer; the initial downsampling network and the initial upsampling network are connected through the initial quantum network layer.
[0073] It should be noted that the structure and function of the initial diffusion model are the same as those of the quantum-classical hybrid diffusion model, which will not be elaborated here.
[0074] Step S420: Obtain the original image dataset; the original image dataset includes the image data of multiple images to be trained.
[0075] In some embodiments, the original image dataset includes the image data of the images in the preset Mnist image dataset and the image data of the images in the preset FashionMnist image dataset. Among them, the Mnist image dataset is a handwritten digit image dataset, which contains handwritten digit images of 10 categories from 0 to 9; the FashionMnist image dataset contains grayscale images of various clothing items such as T-shirts / tops, trousers, pullovers, dresses, coats, sandals, shirts, sports shoes, bags, and ankle boots.
[0076] Step S430: Add noise to the image data in the original image dataset to obtain the dataset to be trained.
[0077] In some embodiments, when adding noise to the image data in the original image dataset, multi-step noise addition can be performed on the image data corresponding to the original image to obtain the first image data. The number of steps of noise addition is random.
[0078] Among them, the added noise is Gaussian noise, that is, noise with an independent Gaussian distribution.
[0079] In other embodiments, due to the additivity of noise with an independent Gaussian distribution, that is, the superposition of any number of noises can be replaced by another noise with an independent Gaussian distribution.
[0080] Specifically, for the first image data obtained by adding noise in the t-th step, where x t is the image data obtained after adding noise in the t-th step to the image data corresponding to the original image; α t is the noise addition parameter in the t-th step of noise addition; x t-1 is the image data obtained after adding noise in the (t - 1)-th step to the image data corresponding to the original image; ε t-1 is the noise data in the t-th step of noise addition; α t-1 is the noise addition parameter in the (t - 1)-th step of noise addition; α t-1 >α t ; ε t-2 is the noise data in the (t - 1)-th step of noise addition. Due to the additivity of noise with an independent Gaussian distribution, then where is the equivalent noise of ε t-1 and ε t-2 ; then Among them, x t-2 is the image data obtained after noise addition in the (t - 2)-th step of the image data corresponding to the original image; α1 is the noise addition parameter during the first-step noise addition; x0 is the image data corresponding to the original image; is ε t-1 ε t-2 …… the equivalent noise of ε0; ε0 is the noise data during the first-step noise addition; is the corresponding noise addition parameter when the t-step noise addition is changed to one-step noise addition; α1 > α2 > …… > α t-1 > α t .
[0081] Specifically, please refer to Figure 5 , Figure 5 which is a schematic diagram of adding noise to the image data.
[0082] As Figure 5 shown, Figure 5 the 7 attached figures in
[0083] are respectively the non-noisy image 500, the noisy image 501-1 during the first-step noise addition, the image 502-1 after the non-noisy image passes through the first-step noise addition, the noisy image 501-2 during the second-step noise addition, the image 502-2 after the image after the first-step noise addition passes through the second-step noise addition, ……, the noisy image 501-t during the t-step noise addition, and the image 502-t after the image after the (t - 1)-step noise addition passes through the t-step noise addition.
[0084] Among them, the image 502-1 after the non-noisy image passes through the first-step noise addition, the image 502-2 after the image after the first-step noise addition passes through the second-step noise addition, ……, the image 502-t after the image after the (t - 1)-step noise addition passes through the t-step noise addition can all be obtained by adding noise to the non-noisy image 500 once.
[0084] Specifically, the image data of the non-noisy image 500 is x = x0. The image data of the noisy image 501-1 during the first-step noise addition is ε0, then the image data of the image 502-1 after the first-step noise addition obtained by adding noise to the image data of the noisy image 501-1 during the first-step noise addition is
[0085] The image data of the noisy image 501-2 during the second-step noise addition is ε1, then the image data of the image 502-2 after the second-step noise addition obtained by adding noise to the image data of the noisy image 501-2 during the second-step noise addition is
[0086] The image data of the noisy image 501-t during the t-step noise addition is ε t-1, then the image data of the image 502-t after the t-th step of noise addition, which is obtained by adding noise to the image data of the noise image 501-t at the t-th step of noise addition, is
[0087]
[0088] In some embodiments, the number of noise addition steps corresponding to the first image data can be input into a preset quantum-classical hybrid diffusion model together with the first image data, so that the preset quantum-classical hybrid diffusion model can predict the noise corresponding to the number of noise addition steps in the first image data to obtain a first predicted noise.
[0089] Step S440, input the training dataset into the initial diffusion model for training to obtain a quantum-classical hybrid diffusion model.
[0090] In this embodiment, by constructing an initial diffusion model; the initial diffusion model includes an initial downsampling network, an initial upsampling network, and an initial quantum network layer; the initial downsampling network and the initial upsampling network are connected through the initial quantum network layer; obtaining an original image dataset; the original image dataset includes the image data of multiple images to be trained; adding noise to the image data in the original image dataset to obtain a training dataset; inputting the training dataset into the initial diffusion model for training to obtain a quantum-classical hybrid diffusion model. In this way, the initial diffusion model including the initial downsampling network, the initial upsampling network, and the initial quantum network layer is trained using the image data in the initial image dataset after adding noise, and a quantum-classical hybrid diffusion model is obtained. The downsampling network in the trained quantum-classical hybrid diffusion model can reduce the scale calculated by the quantum network layer, and thus can perform diffusion generation on images in the quantum domain in existing quantum devices, thereby improving the accuracy of image diffusion.
[0091] In some embodiments, the initial diffusion model can be trained multiple times using the training dataset to obtain a quantum-classical hybrid diffusion model. And / or, the initial diffusion model is trained multiple times using the training dataset, and the loss function corresponding to the initial diffusion model is obtained until the loss function tends to be stable, and a quantum-classical hybrid diffusion model is obtained. Among them, the number of training rounds is greater than 30 and less than 150.
[0092] In some embodiments, the initial diffusion model can be trained 30 times using the training dataset to obtain a loss function that tends to be stable. Finally, when the number of training rounds reaches 100, the training is ended to obtain a quantum-classical hybrid diffusion model.
[0093] Step S130, perform image generation according to the first image data and the first predicted noise.
[0094] In this embodiment, first image data corresponding to a first noisy image to be diffused is obtained; the first image data is input into a preset quantum-classical hybrid diffusion model to obtain first predicted noise corresponding to the first noisy image; the quantum-classical hybrid diffusion model includes a downsampling network, an upsampling network, and a quantum network layer; the downsampling network and the upsampling network are connected by a quantum network; image generation is performed according to the first image data and the first predicted noise. In this way, the quantum-classical hybrid diffusion model composed of the downsampling network, the upsampling network, and the quantum network layer, which has both a classical network and a quantum network, is used to predict the noise in the first noisy image to obtain the first predicted noise, and then image generation is performed according to the first image data and the first predicted noise, realizing the removal of the noise in the first noisy image by using the quantum network. Compared with simply using a quantum model for noise prediction, the downsampling network can reduce the scale calculated by the quantum network layer, so that image diffusion generation can be performed in the quantum domain on existing quantum devices, thereby improving the accuracy of image diffusion.
[0095] It should be noted that performing image generation according to the first image data and the first predicted noise means using the first image data and the first predicted noise to remove the noise in the first image data and the first predicted noise to obtain a new image.
[0096] Specifically, please refer to Figure 6 , Figure 6 is Figure 1 a flowchart of step S130 in the exemplary embodiment shown in. As Figure 6 shown, the process of performing image generation according to the first image data and the first predicted noise may include steps S610 to S630, which are introduced in detail as follows:
[0097] Step S610, obtaining second image data corresponding to a second noisy image according to the first image data and the first predicted noise; the second noisy image is the image obtained by removing the first predicted noise from the first noisy image.
[0098] It should be noted that the second image data is feature data capable of characterizing the second noisy image. It may be one or more of the feature data such as the grayscale feature data of the second noisy image, the color feature data of the second noisy image, and the texture feature data of the second noisy image. Its specific content is not limited here.
[0099] Step S620, inputting the second image data into the preset quantum-classical hybrid diffusion model to obtain second predicted noise corresponding to the second noisy image.
[0100] It should be noted that the number of noise addition steps corresponding to the second image data can be input into a preset quantum-classical hybrid diffusion model together with the second image data, so that the preset quantum-classical hybrid diffusion model can predict the noise corresponding to the number of noise addition steps in the second image data to obtain a second predicted noise.
[0101] Step S630, perform image generation according to the second image data and the second predicted noise.
[0102] In this embodiment, the second image data corresponding to the second denoised image is obtained according to the first image data and the first predicted noise; the second denoised image is the image after the first denoised image removes the first predicted noise; the second image data is input into a preset quantum-classical hybrid diffusion model to obtain the second predicted noise corresponding to the second denoised image; and image generation is performed according to the second image data and the second predicted noise. In this way, the quantum-classical hybrid diffusion model composed of a downsampling network, an upsampling network, and a quantum network layer, which has both a classical network and a quantum network, is used to predict the noise in the second denoised image to obtain the second predicted noise, and then image generation is performed according to the second image data and the second predicted noise, realizing the removal of the noise in the second denoised image by using the quantum network. Compared with simply using a quantum model for noise prediction, the downsampling network can reduce the scale calculated by the quantum network layer, so that image diffusion generation can be performed in the quantum domain on existing quantum devices, thereby improving the accuracy of image diffusion.
[0103] It should be noted that when the number of noise addition steps corresponding to the second image data is a preset number of steps, the target image data obtained by removing the second predicted noise from the second image data can be obtained, and then the target image can be obtained according to the target image data, that is, the image after all the noise in the first denoised image is removed. The preset number of steps can be 1 or. The number of noise addition steps corresponding to the second image data is the number of noise addition steps corresponding to the first image data - 1.
[0104] When the number of noise addition steps corresponding to the second image data is not the preset number of steps, the third image data obtained by removing the second predicted noise from the second image data is obtained and re-input into the preset quantum-classical hybrid diffusion model to obtain the third predicted noise corresponding to the third denoised image. That is, each time, the noise of the image data input into the preset quantum-classical hybrid diffusion model is removed according to the noise predicted by the preset quantum-classical hybrid diffusion model to obtain new image data, and the noise of the new image data is predicted by using the quantum-classical hybrid diffusion model. After each new image data is obtained, the number of noise addition steps is -1 until the number of noise addition steps is the preset number of steps, and the new image data obtained is the target image data, and then the target image with all the noise removed is obtained.
[0105] In some embodiments, please refer to Figure 7 , Figure 7It is a block diagram of a device for image diffusion shown in an exemplary embodiment of the present application.
[0106] As Figure 7 shown, the exemplary device for image diffusion includes:
[0107] An acquisition module 701, configured to acquire first image data corresponding to a first noisy image to be diffused;
[0108] An input module 702, configured to input the first image data into a preset quantum-classical hybrid diffusion model to obtain a first predicted noise corresponding to the first noisy image; the quantum-classical hybrid diffusion model includes a downsampling network, an upsampling network, and a quantum network layer; the downsampling network and the upsampling network are connected through the quantum network;
[0109] A generation module 703, configured to perform image generation according to the first image data and the first predicted noise.
[0110] In an exemplary embodiment, the acquisition module 701 includes:
[0111] A first noisy image acquisition sub-module, configured to acquire a first noisy image to be diffused;
[0112] A first image data acquisition sub-module, configured to acquire first image data corresponding to the first noisy image.
[0113] In an exemplary embodiment, the device for image diffusion further includes:
[0114] A construction sub-module, configured to construct an initial diffusion model; the initial diffusion model includes an initial downsampling network, an initial upsampling network, and an initial quantum network layer; the initial downsampling network and the initial upsampling network are connected through the initial quantum network layer;
[0115] A dataset acquisition sub-module, configured to acquire an original image dataset; the original image dataset includes image data of multiple images to be trained;
[0116] A noise addition sub-module, configured to add noise to the image data in the original image dataset to obtain a dataset to be trained;
[0117] A training sub-module, configured to input the dataset to be trained into the initial diffusion model for training to obtain a quantum-classical hybrid diffusion model.
[0118] In an exemplary embodiment, the downsampling network is used to reduce the dimension of the input image data; the quantum network layer is used to perform linear processing on the image data with reduced dimension in the quantum domain; the upsampling network is used to restore the dimension of the linearly processed image data.
[0119] In an exemplary embodiment, the quantum network layer is used to encode the image data with reduced dimensions into a quantum state to obtain a quantum encoding, and then a preset parameterized quantum circuit is used to linearly process the quantum encoding, and then the linearly processed quantum encoding is measured to obtain the linearly processed image data.
[0120] In an exemplary embodiment, the quantum network layer is configured to pad the image data with reduced dimensions; normalize the padded image data; and encode the normalized image data into the amplitudes of quantum bits to obtain a quantum encoding.
[0121] In an exemplary embodiment, the generation module 703 includes:
[0122] A second image data acquisition sub-module, configured to acquire second image data corresponding to a second noisy image according to first image data and first predicted noise; the second noisy image is an image obtained by removing the first predicted noise from the first noisy image;
[0123] An input sub-module, configured to input the second image data into a preset quantum-classical hybrid diffusion model to obtain second predicted noise corresponding to the second noisy image;
[0124] A generation sub-module, configured to perform image generation according to the second image data and the second predicted noise.
[0125] It should be noted that the apparatus for image diffusion provided in the above embodiments and the method for image diffusion provided in the above embodiments belong to the same concept. The specific manners in which each module and unit perform operations have been described in detail in the method embodiments and will not be elaborated here. In practical applications, the apparatus for image diffusion provided in the above embodiments may, according to needs, allocate the above functions to different functional modules, that is, divide the internal structure of the apparatus into different functional modules to complete all or part of the functions described above, and this is not limited here either.
[0126] An embodiment of the present application further provides an electronic device, including: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device implements the method for image diffusion provided in each of the above embodiments.
[0127] Figure 8 The structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown. It should be noted that Figure 8 The computer system 800 of the electronic device shown is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present application.
[0128] AsFigure 8 As shown, computer system 800 includes a Central Processing Unit (CPU) 801, which can perform various appropriate actions and processes according to a program stored in a Read-Only Memory (ROM) 802 or a program loaded from a storage section 808 into a Random Access Memory (RAM) 803, such as executing the method described in the above embodiments. In the RAM 803, various programs and data required for system operations are also stored. The CPU 801, ROM 802, and RAM 803 are connected to each other via a bus 804. An Input / Output (I / O) interface 805 is also connected to the bus 804.
[0129] The following components are connected to the I / O interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including, for example, a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), etc. and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed so that a computer program read from it can be installed into the storage section 808 as needed.
[0130] Specifically, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 809, and / or installed from the removable medium 811. When the computer program is executed by a Central Processing Unit (CPU) 801, various functions defined in the system of the present application are executed.
[0131] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0132] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0133] The units involved in the embodiments described in this application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not constitute a limitation to the unit itself in some cases.
[0134] Another aspect of this application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for image diffusion as described above. The computer-readable storage medium can be included in the electronic device described in the above embodiments, or can exist alone without being assembled into the electronic device.
[0135] Another aspect of this application also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method for image diffusion provided in the above various embodiments.
[0136] The above content is only a preferred exemplary embodiment of this application and is not used to limit the implementation of this application. Those of ordinary skill in the art can easily make corresponding adaptations or modifications according to the main concept and spirit of this application. Therefore, the protection scope of this application should be subject to the protection scope required by the claims.
Claims
1. A method for image diffusion, characterized in that, Including: Obtain first image data corresponding to a first noisy image to be diffused; Input the first image data into a preset quantum-classical hybrid diffusion model to obtain a first predicted noise corresponding to the first noisy image; The quantum-classical hybrid diffusion model includes a downsampling network, an upsampling network, and a quantum network layer; the downsampling network and the upsampling network are connected through the quantum network; Perform image generation based on the first image data and the first predicted noise.
2. The method according to claim 1, characterized in that, The obtaining of the first image data corresponding to the first noisy image to be diffused includes: Obtain a first noisy image to be diffused; Obtain first image data corresponding to the first noisy image.
3. The method according to claim 1, wherein The quantum-classical hybrid diffusion model is obtained through the following steps: Construct an initial diffusion model; the initial diffusion model includes an initial downsampling network, an initial upsampling network, and an initial quantum network layer; the initial downsampling network and the initial upsampling network are connected through the initial quantum network layer; Obtain an original image dataset; the original image dataset includes image data of multiple images to be trained; Add noise to the image data in the original image dataset to obtain a training dataset; Input the training dataset into the initial diffusion model for training to obtain the quantum-classical hybrid diffusion model.
4. The method according to claim 1, wherein The downsampling network is used to reduce the dimension of the input image data; the quantum network layer is used to perform linear processing on the image data with reduced dimensions in the quantum domain; the upsampling network is used to restore the dimension of the linearly processed image data.
5. The method according to claim 4, characterized in that, The quantum network layer is used to encode the image data with reduced dimensions into a quantum state to obtain a quantum encoding, then perform linear processing on the quantum encoding using a preset parameterized quantum circuit, and then measure the linearly processed quantum encoding to obtain linearly processed image data.
6. The method according to claim 5, wherein The encoding of the image data with reduced dimensions into a quantum state to obtain a quantum encoding includes: Pad the image data with reduced dimensions; Normalize the padded image data; Encode the normalized image data into the amplitude of quantum bits to obtain the quantum encoding.
7. The method according to claim 1, characterized in that, The performing of image generation based on the first image data and the first predicted noise includes: Obtain second image data corresponding to a second noisy image according to the first image data and the first predicted noise; the second noisy image is an image obtained by removing the first predicted noise from the first noisy image; Input the second image data into a preset quantum-classical hybrid diffusion model to obtain a second predicted noise corresponding to the second noisy image; Perform image generation based on the second image data and the second predicted noise.
8. An apparatus for image diffusion, characterized in that, Including: An acquisition module configured to obtain first image data corresponding to a first noisy image to be diffused; An input module configured to input the first image data into a preset quantum-classical hybrid diffusion model to obtain a first predicted noise corresponding to the first noisy image; The quantum-classical hybrid diffusion model includes a downsampling network, an upsampling network, and a quantum network layer; the downsampling network and the upsampling network are connected through the quantum network; A generation module configured to generate an image based on the first image data and the first predicted noise.
9. An electronic device, characterized in that, Comprising: One or more processors; A storage device for storing one or more programs, which when executed by the one or more processors, cause the electronic device to implement the method for image diffusion according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Computer-readable instructions are stored thereon, which when executed by a processor of a computer, cause the computer to execute the method for image diffusion according to any one of claims 1 to 7.