Artificial intelligence model training method and device, electronic equipment and storage medium
By introducing a quantum processor to encode the loss function as a quantum state and measure the gradient in the classical AI model training method, the problem of high hardware requirements for training quantized AI models is solved, and efficient AI model training is achieved.
Patent Information
- Application Number
- CN202510838832.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-10-21
AI Technical Summary
When existing technologies reconstruct AI models into quantum states for training, they have high hardware requirements and high migration costs, cannot effectively utilize traditional AI model architectures, and have low training efficiency.
The central processing unit preprocesses the raw training data, the graphics processing unit inputs the training set, the quantum processor encodes the loss function into quantum states and constructs quantum circuits, measures the quantum gradient, combines classical AI model training methods, updates parameters, and reduces the dependence on quantum hardware resources.
While ensuring training accuracy, it improves the training efficiency of AI models, reduces dependence on quantum hardware resources, and lowers costs.
Smart Images

Figure CN120822633A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of quantum computing technology, and in particular to an artificial intelligence model training method, device, electronic device and storage medium. Background Art
[0002] Artificial Intelligence (AI) models have become the core engine driving innovation across various fields. The core of building AI models lies in model training. The current training process includes three key steps: data preprocessing (such as cleaning and partitioning the original dataset), model building (such as selecting an appropriate algorithm architecture and initializing model parameters), and model training (such as using a loss function to update parameters).
[0003] However, with the explosive growth of model complexity and data sets, traditional AI model training is facing increasingly severe challenges. For example, training complex models often requires massive computing resources and takes days or even weeks, resulting in high costs and energy consumption.
[0004] Therefore, existing techniques use quantum technology to completely reconstruct AI models into quantum states and train these AI models using quantized training sets. This approach leverages quantum technology's ability to more efficiently process mathematical problems compared to traditional computing methods to apply to AI model training. However, this method involves topological quantum computers, and the related hardware technology is not yet mature, resulting in high hardware requirements for the overall solution. Furthermore, since this method reconstructs the entire AI model into a quantum state, it places high demands on physical resources and cannot utilize existing traditional AI model architectures, resulting in extremely high migration costs.
[0005] Therefore, how to reduce the cost of applying quantum technology to AI model training and improve the training efficiency of AI models is an urgent problem to be solved. Summary of the Invention
[0006] The embodiments of the present application provide an artificial intelligence model training method, device, electronic device, and storage medium to reduce the cost of applying quantum technology to AI model training and improve the training efficiency of AI models.
[0007] In a first aspect, an embodiment of the present application provides an artificial intelligence model training method, comprising:
[0008] The central processing unit performs data preprocessing on the original training data obtained from the memory; the original training data includes a training set;
[0009] The following iterative training steps are performed on the AI model to be trained until the preset iteration stop condition is reached:
[0010] The graphics processor inputs the training set into the AI model to be trained, and obtains a loss function corresponding to the current parameters of the AI model to be trained;
[0011] The quantum processor encodes the loss function into a quantum state;
[0012] The quantum processor constructs a quantum circuit based on the loss function of the quantum state, and measures the quantum gradient of the loss function through the quantum circuit;
[0013] The quantum processor updates the current parameter based on the quantum gradient and a preset regularization gradient.
[0014] In a second aspect, an embodiment of the present application provides an artificial intelligence model training device, comprising:
[0015] An acquisition unit, configured for the central processing unit to perform data preprocessing on the original training data acquired from the memory; the original training data includes a training set;
[0016] The training unit is used to perform the following iterative training steps on the AI model to be trained until the preset iteration stop condition is reached:
[0017] The graphics processor inputs the training set into the AI model to be trained, and obtains a loss function corresponding to the current parameters of the AI model to be trained;
[0018] The quantum processor encodes the loss function into a quantum state;
[0019] The quantum processor constructs a quantum circuit based on the loss function of the quantum state, and measures the quantum gradient of the loss function through the quantum circuit;
[0020] The quantum processor updates the current parameter based on the quantum gradient and a preset regularization gradient.
[0021] In some embodiments, the training unit is specifically configured to:
[0022] Measuring the quantum gradient of the loss function by the quantum circuit includes:
[0023] The quantum processor constructs a preset number of quantum bits into an initial state;
[0024] The quantum processor adds a preset perturbation variable to the current parameters in the quantum circuit to obtain a perturbation circuit;
[0025] After the quantum processor inputs the initial state into the perturbation circuit, a positive perturbation output state and a negative perturbation output state are obtained;
[0026] The quantum processor determines a first expected value of an observation operator based on the positive perturbation output state, and calculates a second expected value of the observation operator based on the positive perturbation output state;
[0027] The quantum processor determines a quantum gradient of the loss function based on a difference between the first expected value and the second expected value and a current gradient calculation accuracy.
[0028] In some embodiments, after measuring the quantum gradient of the loss function through the quantum circuit, the training unit is further configured to:
[0029] The quantum processor monitors the decoherence time of the quantum bits in real time;
[0030] The quantum processor updates the current gradient calculation accuracy according to the relationship between the quantum bit decoherence time and a preset threshold value.
[0031] In some embodiments, the training unit is specifically configured to:
[0032] The quantum processor uses the product of a preset quantum gradient weighting factor and the quantum gradient as a first product;
[0033] The quantum processor uses the product of a preset regularization coefficient and the regularization gradient as a second product; the regularization gradient is a correction amount used to suppress overfitting;
[0034] The quantum processor uses the sum of the first product and the second product as a combined gradient;
[0035] The quantum processor updates the current parameter based on the combined gradient and a preset learning rate; the learning rate is used to control the update amplitude of the current parameter during iterative training.
[0036] In some embodiments, the original training data further includes a validation set; the validation set is used to verify the performance of the AI model to be trained; the validation set includes multiple validation samples; and the iteration stopping condition includes at least one of the following:
[0037] The decrease in validation set loss is less than a first threshold; the validation set loss is determined by inputting the validation set into the AI model to be trained, based on the loss function of each validation sample and the number of validation samples;
[0038] The increase in the accuracy of the verification set is less than a second threshold; the accuracy is determined by inputting the verification set into the AI model to be trained and based on the model output of each verification sample;
[0039] The number of iterations reaches the preset number.
[0040] In some embodiments, after the quantum processor encodes the loss function into a quantum state, the training unit is further configured to:
[0041] The quantum processor performs error detection on the encoding operation by performing stable sub-measurement and obtains the detection result;
[0042] If a quantum bit with an encoding error is detected, the quantum processor performs an error correction gate operation corresponding to the detection result on the quantum bit with the encoding error to correct the encoding operation error.
[0043] In a third aspect, an embodiment of the present application provides an electronic device, including:
[0044] a memory for storing program instructions;
[0045] The processor is used to call the program instructions stored in the memory and execute the above-mentioned artificial intelligence model training method according to the obtained program instructions.
[0046] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned artificial intelligence model training method is implemented.
[0047] In the fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which is stored in a computer-readable storage medium; when the processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the electronic device performs the above-mentioned artificial intelligence model training method.
[0048] The embodiments of the present application provide an artificial intelligence model training method, device, electronic device, and storage medium. First, the central processing unit performs data preprocessing on the original training data, and the graphics processing unit inputs the training set divided from the original training data into the AI model to be trained. During the iterative training process of the AI model to be trained using the training set, based on the classical AI model training method, the quantum processor quantizes and encodes the loss function of the current parameters of the AI model to be trained into a quantum state. The quantum processor then measures the quantum gradient through a quantum circuit and uses a gradient estimation algorithm based on quantum amplitude to update the current parameters. This accelerates the AI model training while ensuring training accuracy and reduces dependence on quantum hardware resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 A schematic diagram of an application scenario of an artificial intelligence model training method provided in an embodiment of the present application;
[0050] Figure 2 A schematic diagram of a module for an artificial intelligence model training method provided in an embodiment of the present application;
[0051] Figure 3 A flowchart of an artificial intelligence model training method provided in an embodiment of the present application;
[0052] Figure 4 A flowchart of another artificial intelligence model training method provided in an embodiment of the present application;
[0053] Figure 5 A schematic diagram of the structure of an electronic device for an artificial intelligence model training method in an embodiment of the present application;
[0054] Figure 6 A schematic diagram of the hardware structure of an electronic device to which an embodiment of the present application is applied;
[0055] Figure 7 A schematic diagram of the hardware structure of a computing device using an embodiment of the present application. DETAILED DESCRIPTION
[0056] To make the objectives, technical solutions, and advantages of this application more clear, this application will be further described in detail below with reference to the accompanying drawings. It should be understood that the embodiments described herein are only some of the embodiments of this application, and not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of this application without inventive effort are intended to fall within the scope of protection of this application.
[0057] The following is an introduction to some concepts involved in the embodiments of this application.
[0058] 1. AI: refers to the technology and science of simulating human intelligence by computer systems or machines, enabling them to perform tasks that usually require human cognitive abilities.
[0059] 2. Model training: refers to the process of adjusting the internal parameters of the model through algorithms and data so that it can learn patterns from the input data and complete specific tasks.
[0060] 3. Gradient estimation: This refers to the method of approximating the gradient of the objective function through numerical or algorithmic means. Its core goal is to efficiently obtain gradient information to support parameter optimization.
[0061] The following is a brief introduction to the design concept of the embodiment of this application:
[0062] With the explosive growth of model complexity and data volume, traditional AI model training is facing increasingly severe challenges. For example, training complex models often requires massive computing resources and takes days or even weeks, resulting in high costs and energy consumption.
[0063] Therefore, in an existing technology, an AI model based on topological quantum computing is designed: the data to be trained is obtained, and a preset algorithm is used to divide the data to be trained into a training data set and a test data set; the training data set and the test data set are mapped into quantum mechanical states through quantum random access memory; the topological quantum deep complex network corresponding to the AI model is trained using the quantized training set to obtain the quantum state of the optimal weight value of the AI model; the topological quantum deep complex network and the AI model have a corresponding relationship and are a quantized mapping of the AI model; the optimal weight value is mapped into classical data to obtain a trained AI model.
[0064] However, this method involves topological quantum computers, and the hardware technology for this method is not yet mature, and the overall solution has high hardware requirements; mapping classical data into quantum states through quantum random access memory requires higher physical resources (such as low-temperature environment and high-precision control); this method requires reconstructing the entire AI model into a "topological quantum deep complex network", which makes it impossible to directly migrate traditional models (such as transformer model (Transformer) and convolutional neural network (CNN)); the quantum state mapping / measurement process will introduce errors, requiring quantum error correction, and increasing computational overhead; the training process is highly complex, and the "quantum state mapping → quantum training → classical decoding" cycle must be repeated, and the quantum-classical interface becomes a performance bottleneck.
[0065] In another prior art, a training method for a hybrid quantum-classical generative adversarial network is designed: an input image is obtained, and the input image is input into the generator to obtain an output image; an image set including the output image is input into the discriminator to obtain a discrimination result of the image set; a first loss function value of the generator and a second loss function value of the discriminator are calculated based on the discrimination result and the label data of the image set; the parameters of the generator are updated based on the first loss function value, and the parameters of the discriminator are updated based on the second loss function value to train the generative adversarial network.
[0066] However, the quantum generator / discriminator involved in this patent requires a specific quantum circuit design. Current quantum hardware (such as superconducting / ion traps) is difficult to support high-fidelity quantization of high-dimensional images; the cost of implementation in the short term is high. This architecture is designed for Generative Adversarial Network (GAN) tasks, and the cost of migrating to other AI models (such as Transformer and CNN classifiers) is high. The quantum generator involved in this method requires a large number of quantum bits to process image data, which is more than the current noisy medium-scale quantum device capabilities (usually ≤100 quantum bits and noisy).
[0067] In view of this, the embodiments of the present application provide an artificial intelligence model training method, device, electronic device, and storage medium. First, the central processing unit performs data preprocessing on the original training data, and the graphics processing unit inputs the training set divided from the original training data into the AI model to be trained. During the iterative training process of the AI model to be trained using the training set, based on the classical AI model training method, the quantum processor quantizes and encodes the loss function of the current parameters of the AI model to be trained into a quantum state. The quantum processor then measures the quantum gradient through a quantum circuit and uses a gradient estimation algorithm based on quantum amplitude to update the current parameters, thereby accelerating the AI model training while ensuring training accuracy and reducing dependence on quantum hardware resources.
[0068] It should be noted that the application scenarios described in the following embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Ordinary technicians in this field can know that with the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0069] The following first briefly introduces the application scenarios to which the technical solutions of the embodiments of the present application can be applied. It should be noted that the application scenarios introduced below are only used to illustrate the embodiments of the present application and are not limiting. In specific implementation, the technical solutions provided by the embodiments of the present application can be flexibly applied according to actual needs.
[0070] like Figure 1 As shown, it is a schematic diagram of an application scenario of an artificial intelligence model training method provided by an embodiment of the present application. The application scenario diagram includes a server 110, a central processing unit 1101, a graphics processing unit 1102 and a quantum processor 1103.
[0071] It should be noted that the artificial intelligence model training method in each embodiment of the present application can be jointly executed by the central processing unit 1101, the graphics processing unit 1102 and the quantum processor 1103 in the server 110. The central processing unit 1101 performs data preprocessing on the original training data obtained from the memory; the training set divided from the preprocessed original training data is sent to the graphics processing unit 1102, and the graphics processing unit 1102 inputs the training set into the AI model to be trained to obtain the loss function corresponding to the current parameters of the AI model to be trained; the quantum processor 1103 encodes the loss function into a quantum state; the quantum processor 1103 constructs a quantum circuit based on the loss function of the quantum state, and measures the quantum gradient of the loss function through the quantum circuit; the quantum processor 1103 updates the current parameter based on the quantum gradient and the preset regularization gradient, and repeats the current parameter update step until the preset iteration stop condition is reached.
[0072] In an optional embodiment, the central processing unit 1101, the graphics processing unit 1102 and the quantum processor 1103 can communicate with each other through a communication network.
[0073] In an optional implementation, the communication network is a wired network or a wireless network.
[0074] It should be noted that Figure 1 The examples shown are just for illustration. In fact, the number of terminal devices and servers is not limited and is not specifically limited in the embodiments of this application.
[0075] like Figure 2 As shown, compared with the classic AI model training method, the artificial intelligence model training method provided by this application involves the following modules:
[0076] Input data: The CPU loads the raw training data.
[0077] Data preprocessing: The central processing unit divides the original training data into training set, test set, and validation set, and performs preprocessing operations such as standardization, feature extraction, and null value processing.
[0078] Model initialization and forward propagation: The GPU randomly initializes the model parameters and calculates the loss function for the current parameters.
[0079] Quantum gradient estimation: During the iterative training of the AI model to be trained, the quantum processor encodes the loss function into quantum states, constructs parameterized quantum circuits for gradient estimation, and uses noise suppression to protect key parameters.
[0080] Dynamic task allocation: Based on the quantum hardware state (such as decoherence time or gate fidelity), the quantum processor dynamically allocates high-precision quantum computing or classical cluster tasks.
[0081] Hybrid optimization: The quantum processor uses an optimizer to combine quantum gradient weighting factors with regularized gradients to update the current parameters and determine whether the validation set loss has converged. If not, the model initialization and forward propagation, quantum gradient estimation, and dynamic task allocation steps are repeated until convergence.
[0082] Model output: Output the final AI model after training.
[0083] The artificial intelligence training method provided in this application is suitable for the training of large-scale neural network models (such as Generative Pre-trained Transformer 4 (GPT-4), Bidirectional Encoder Representations from Transformers (BERT), Residual Network (ResNet), etc.), dynamic structure models (such as Neural Architecture Search (NAS) models, etc.) and multimodal fusion models (such as Contrastive Language-Image Pretraining (CLIP), Flamingo, BERT-3, etc.).
[0084] It performs outstandingly in scenarios such as medical image analysis, autonomous driving decision-making, and financial risk prediction.
[0085] The following uses the image classification model as an example to introduce the artificial intelligence model training method in this application.
[0086] The central processing unit performs data preprocessing on the original image data obtained from the memory; the original image data includes a training set; the training set includes the image data and the category labels corresponding to the image data;
[0087] The following iterative training steps are performed on the AI model to be trained until the preset iteration stop condition is reached:
[0088] The graphics processor inputs the training set into the AI model to be trained, and obtains a loss function corresponding to the current parameters of the AI model to be trained; wherein the AI model to be trained is an image classification model to be trained;
[0089] The quantum processor encodes the loss function into a quantum state;
[0090] The quantum processor constructs a quantum circuit based on the loss function of the quantum state and measures the quantum gradient of the loss function through the quantum circuit;
[0091] The quantum processor updates the current parameters based on the quantum gradient and the preset regularization gradient.
[0092] Obtain the target image classification model after iterative training to use the target image classification model for image classification tasks.
[0093] The following combination Figure 3 This paper introduces each module and the artificial intelligence model training method provided by this application.
[0094] Figure 3 A flowchart of an artificial intelligence model training method provided in an embodiment of the present application is shown. Figure 3 As shown, the method may include the following steps S31-S32:
[0095] S31: The central processing unit performs data preprocessing on the original training data obtained from the memory.
[0096] The original training data includes the training set.
[0097] In an optional embodiment, S31 can be executed by a data preprocessing module, which divides the original training data into a training set, a test set, and a validation set, and performs preprocessing operations such as standardization, feature extraction, and null value processing on each data set.
[0098] S32: Execute the following iterative training steps S3201 to S3204 on the artificial intelligence AI model to be trained until the preset iteration stop condition is reached.
[0099] From the above content, we can see that the original training data is also divided into a validation set; the validation set is used to: verify the performance of the AI model to be trained; the validation set includes multiple validation samples.
[0100] In an optional embodiment, the present application determines the training result of the AI model to be trained by the convergence of the validation set loss:
[0101] (1) The decrease in the loss of the validation set is less than the first threshold.
[0102] The validation set loss is: the validation set is input into the AI model to be trained, and is determined based on the loss function of each validation sample and the number of validation samples.
[0103] Specifically, the loss function of each verification sample after inputting into the AI model to be trained is calculated, and the average value of these loss functions can be used as the verification set loss.
[0104] When the loss of the validation set decreases very little, it indicates that the optimization space of the AI model to be trained is already very small. Therefore, a first threshold can be set to stop iterative training when it is determined that the optimization space of the AI model to be trained is already small enough.
[0105] The first threshold value may be set to 0.1% to 0.5% of the initial loss, or may be an empirical value set, such as 0.001 to 0.005.
[0106] (2) The increase in the accuracy of the validation set is less than the second threshold.
[0107] The accuracy is determined by inputting the validation set into the AI model to be trained and based on the model output of each validation sample.
[0108] Specifically, for each verification sample, if the output of the AI model to be trained is consistent with the actual output of the verification sample, it is considered that the prediction of the AI model to be trained for the verification sample is correct.
[0109] The validation set accuracy is the percentage of correctly predicted validation samples to the total validation samples.
[0110] When the validation set accuracy rate increases minimally, it indicates that the performance improvement of the AI model to be trained has slowed down. Therefore, a second threshold can be set to stop iterative training when it is determined that the performance of the AI model to be trained has almost stopped improving.
[0111] The second threshold can be dynamically adjusted based on the use case and task difficulty of the AI model being trained. For example, the second threshold can be increased when facing high-noise data, and can be decreased when facing high-precision tasks. Possible values for the second threshold can range from 0.0005 to 0.002, etc.
[0112] In another optional embodiment, the present application determines the timing of stopping iteration by limiting the number of iterative training: the number of iterations reaches a preset number.
[0113] This application prevents the AI model to be trained from falling into infinite training or overfitting by presetting the maximum number of iterations as a hard constraint condition for iteration termination.
[0114] The preset number of times can be dynamically adjusted according to the difficulty of the task and the amount of data that the AI model to be trained needs to perform. For example, if the data set is a large data set, the preset number of times can be 100 to 300. If the task is a simple task, the preset number of times can be 10 to 20, etc.
[0115] S3201: The graphics processor inputs the training set into the AI model to be trained and obtains the loss function corresponding to the current parameters of the AI model to be trained.
[0116] In an optional embodiment, the loss function corresponding to the current parameters is obtained by model initialization and forward propagation.
[0117] In this application, the current parameter is recorded as θ, and the loss function of the current parameter is recorded as L(θ).
[0118] The loss function can be mean square error, cross entropy loss, etc.
[0119] S3202: The quantum processor encodes the loss function into a quantum state.
[0120] In an alternative embodiment, the quantum processor encodes the loss function L(θ) as a quantum state
[0121] and will Normalized to the interval [0, 1].
[0122] S3203: The quantum processor constructs a quantum circuit based on the loss function of the quantum state and measures the quantum gradient of the loss function through the quantum circuit.
[0123] During the iterative training process of this application, the quantum processor encodes the loss function into quantum states, constructs parameterized quantum circuits for gradient estimation, and uses noise suppression to protect key parameters.
[0124] The following first introduces the noise suppression steps:
[0125] In an optional embodiment, the operation suppression step is performed by a quantum gradient estimation module, and the specific steps are as follows:
[0126] The quantum processor detects errors in the coding operation by performing stabilizer measurements and obtains the detection results. If a quantum bit with a coding error is detected, the quantum processor performs an error correction gate operation corresponding to the detection result on the quantum bit with the coding error to correct the coding operation error.
[0127] Specifically, the noise suppression step is used to protect key parameters and reduce the impact of noise on algorithm performance. After encoding the single logical quantum bit information to be encoded in the loss function onto N physical quantum bits, a set of stabilizer measurements is performed to detect whether any coding errors occur. Once an error is detected, the corresponding error correction module (i.e., error correction gate) is inserted based on the detection result to correct the coding error.
[0128] In the above implementation, key parameters are protected through the noise suppression step, thereby improving the noise robustness of the AI model training process.
[0129] Then the gradient estimation steps are introduced:
[0130] In an optional embodiment, the gradient estimation step is performed by a quantum gradient estimation module, and the specific steps are as follows:
[0131] The quantum processor constructs a preset number of quantum bits as an initial state; the quantum processor adds a preset perturbation variable to the current parameters in the quantum circuit to obtain a perturbation circuit; after inputting the initial state into the perturbation circuit, the quantum processor obtains a positive perturbation output state and a negative perturbation output state; the quantum processor determines a first expected value of the observation operator based on the positive perturbation output state, and calculates a second expected value of the observation operator based on the positive perturbation output state; the quantum processor determines the quantum gradient of the loss function based on the difference between the first expected value and the second expected value and the current gradient calculation accuracy.
[0132] It is divided into the following four steps:
[0133] (1) Construct a parameterized quantum circuit U(θ), allocate n quantum bits, and construct the initial state
[0134] (2) Apply a preset small perturbation ±ε to the parameter θ of the quantum circuit U(θ) to obtain the perturbation circuit U(θ±ε);
[0135] The initial state Input the perturbation circuit U(θ±ε) and obtain the output state shown in the following formula 1:
[0136]
[0137] Among them, |ψ(θ+ε)〉 is the positive perturbation output state, and |ψ(θ-ε)〉 is the negative perturbation output state.
[0138] (3) Define the observation operator H and calculate the first expected value<U(θ+ε)|H|U(θ+ε)> and the second expected value<U(θ-ε)|H|U(θ-ε)> .
[0139] (4) Based on the current gradient calculation accuracy, the quantum gradient of the loss function L(θ) is measured through the quantum circuit, as shown in the following formula 2:
[0140]
[0141] in, is the quantum gradient.
[0142] In the above implementation, the loss function is encoded as a quantum state, and quantum amplitude estimation is combined to accelerate gradient calculation and maintain estimation accuracy. Compared with the classic AI model training method, this application only quantizes the gradient calculation link, without the need to reconstruct the generator or discriminator, effectively saving deployment costs.
[0143] Furthermore, the method for determining the accuracy of gradient calculation when calculating quantum gradients is introduced.
[0144] In this application, the accuracy of gradient calculation is adjusted according to the state of quantum hardware. The quantum hardware state can be reflected by parameters such as quantum bit decoherence time and gate fidelity. The following uses the quantum bit decoherence time as an example to introduce the method for determining the accuracy of gradient calculation in this application.
[0145] In an optional embodiment, the step of determining the gradient calculation accuracy is performed by the dynamic task allocation module, and the specific steps are as follows:
[0146] The quantum processor monitors the decoherence time of the quantum bit in real time; the quantum processor updates the current gradient calculation accuracy based on the relationship between the quantum bit decoherence time and the preset threshold value.
[0147] Specifically, the quantum processor monitors the decoherence time (T1) of quantum bits in real time and dynamically allocates quantum gradient calculation tasks to improve computing efficiency:
[0148] (1) When T1 ≥ threshold value, high-precision gradient calculation is performed.
[0149] (2) When T1 < threshold, the classical cluster performs gradient calculation.
[0150] Among them, the threshold value can be set according to the characteristics of quantum hardware or the task requirements of the AI model. For example, the threshold value can be set to 100us, 120us, etc.
[0151] In the above embodiment, a dynamic task allocation control module is added to the training process of the AI model to be trained. This module perceives the quantum hardware status such as the decoherence time of the quantum bit in real time, intelligently allocates computing tasks, and improves the utilization rate of quantum hardware resources.
[0152] S3204: The quantum processor updates the current parameters based on the quantum gradient and the preset regularization gradient.
[0153] In this application, the current parameters are updated by combining the quantum gradient weighting factor and the regularization gradient through the optimizer.
[0154] In an optional implementation, the current parameters are updated through a hybrid optimization module, and the specific implementation steps are as follows:
[0155] (1) The quantum processor adds the preset quantum gradient weighting factor Q and the quantum gradient The product of is the first product;
[0156] (2) The quantum processor uses the preset regularization coefficient λ and regularization gradient The product of is used as the second product; the regularization gradient is the correction amount used to suppress overfitting;
[0157] (3) The quantum processor takes the sum of the first product and the second product as the combined gradient
[0158] (4) The quantum processor updates the current parameters based on the combined gradient and the preset learning rate η; the learning rate is used to control the update amplitude of the current parameters during the iterative training process.
[0159] The current parameter before updating in this iterative training is recorded as θ t , the updated current parameter is recorded as θ t+1 , then θ t+1 Calculated by the following formula 3:
[0160]
[0161] like Figure 4 As shown, it is a flowchart of another artificial intelligence model training method provided in an embodiment of the present application.
[0162] S401: data preprocessing;
[0163] S402: Randomly initialize model parameters
[0164] S403: Calculate loss function;
[0165] S404: Quantum coding;
[0166] S405: Quantum circuit construction;
[0167] S406: Quantum gradient estimation;
[0168] S407: Quantum coding error correction;
[0169] S408: Quantum coding error correction;
[0170] S409: Allocating tasks according to decoherence time;
[0171] Specifically, if the decoherence time is greater than or equal to the threshold value, jump to S410; otherwise, jump to S411;
[0172] S410: high-precision gradient calculation;
[0173] S411: Classic cluster processing;
[0174] S412: Current parameters are updated;
[0175] S413: Convergence judgment;
[0176] S414: Save the model and perform reasoning and application.
[0177] The following is an introduction to the artificial intelligence model training method in this application with reference to the embodiments:
[0178] The AI model to be trained is the ResNet-50 image classification model.
[0179] (1) Input data: The CPU loads the ImageNet dataset as the original training data.
[0180] (2) Data preprocessing: The central processing unit divides the original training data into training set, test set, and validation set, and performs preprocessing operations such as standardization, feature extraction, and null value processing.
[0181] (3) Model initialization and forward propagation: The GPU randomly initializes the model parameters θ and calculates the loss function L(θ) for the current parameters θ.
[0182] For example, an 8-qubit Quantum Random Access Memory (QRAM) module is used to encode the loss function L(θ) into a quantum state And normalized to the interval [0, 1].
[0183] (4) Quantum gradient estimation: During the iterative training of the ResNet-50 image classification model, the quantum processor encodes the loss function into quantum states, constructs a parameterized quantum circuit for gradient estimation, and uses noise suppression to protect key parameters.
[0184] (5) Dynamic task allocation: Based on the quantum hardware state (such as decoherence time or gate fidelity), the quantum processor dynamically allocates high-precision quantum computing or classical cluster tasks.
[0185] For example, when T1 ≥ a threshold value (such as 120 μs), high-precision gradient calculation is performed; when T1 < a threshold value (such as 120 μs), the classic cluster performs gradient calculation.
[0186] (6) Hybrid optimization: The quantum processor updates the current parameters through the optimizer by combining the quantum gradient weighting factor and the regularized gradient, and determines whether the validation set loss converges. If not, the model initialization and forward propagation, quantum gradient estimation and dynamic task allocation steps are repeated until convergence.
[0187] (7) Model output: Output the trained ResNet-50 image classification model
[0188] In addition, the gradient estimation algorithm in this application can also be expanded to introduce distributed technology, with one quantum computing center as the master node and N edge quantum devices (8-16 quantum bits) as slave nodes, using a variational quantum compression algorithm to improve computing performance and accuracy.
[0189] Based on the same inventive concept, the present application also provides a training device for an artificial intelligence model, such as Figure 5 As shown, the training device 5000 of the artificial intelligence model includes:
[0190] The acquisition unit 5001 is used for the central processing unit to perform data preprocessing on the original training data acquired from the memory; the original training data includes a training set;
[0191] The training unit 5002 is configured to perform the following iterative training steps on the AI model to be trained until a preset iteration stop condition is reached:
[0192] The GPU inputs the training set into the AI model to be trained and obtains the loss function corresponding to the current parameters of the AI model to be trained;
[0193] The quantum processor encodes the loss function into a quantum state;
[0194] The quantum processor constructs a quantum circuit based on the loss function of the quantum state and measures the quantum gradient of the loss function through the quantum circuit;
[0195] The quantum processor updates the current parameters based on the quantum gradient and the preset regularization gradient.
[0196] In some embodiments, the training unit 5002 is specifically configured to:
[0197] Measuring the quantum gradient of the loss function through a quantum circuit includes:
[0198] The quantum processor constructs a preset number of quantum bits into an initial state;
[0199] The quantum processor adds a preset perturbation variable to the current parameters in the quantum circuit to obtain a perturbation circuit;
[0200] After the quantum processor inputs the initial state into the perturbation circuit, it obtains a positive perturbation output state and a negative perturbation output state;
[0201] The quantum processor determines a first expected value of the observation operator based on the positive perturbation output state, and calculates a second expected value of the observation operator based on the positive perturbation output state;
[0202] The quantum processor determines a quantum gradient of the loss function based on a difference between the first expected value and the second expected value and a current gradient calculation accuracy.
[0203] In some embodiments, after measuring the quantum gradient of the loss function through the quantum circuit, the training unit 5002 is further configured to:
[0204] The quantum processor monitors the decoherence time of the quantum bits in real time;
[0205] The quantum processor updates the current gradient calculation accuracy based on the relationship between the quantum bit decoherence time and the preset threshold value.
[0206] In some embodiments, the training unit 5002 is specifically configured to:
[0207] The quantum processor uses the product of the preset quantum gradient weighting factor and the quantum gradient as the first product;
[0208] The quantum processor uses the product of the preset regularization coefficient and the regularization gradient as the second product; the regularization gradient is a correction used to suppress overfitting;
[0209] The quantum processor takes the sum of the first product and the second product as the combined gradient;
[0210] The quantum processor updates the current parameters based on the combined gradient and the preset learning rate; the learning rate is used to control the update amplitude of the current parameters during the iterative training process.
[0211] In some embodiments, the original training data also includes a validation set; the validation set is used to verify the performance of the AI model to be trained; the validation set includes multiple validation samples; and the iteration stopping condition includes at least one of the following:
[0212] The decrease in validation set loss is less than a first threshold; validation set loss is determined by the loss function for each validation sample and the number of validation samples when the validation set is input into the AI model to be trained;
[0213] The increase in the accuracy of the validation set is less than a second threshold; the accuracy is determined by inputting the validation set into the AI model to be trained and based on the model output of each validation sample;
[0214] The number of iterations reaches the preset number.
[0215] In some embodiments, after the quantum processor encodes the loss function into a quantum state, the training unit 5002 is further configured to:
[0216] The quantum processor performs error detection on the encoding operation by performing stable quantum measurement and obtains the detection result;
[0217] If a quantum bit with a coding error is detected, the quantum processor performs an error correction gate operation corresponding to the detection result on the quantum bit with the coding error to correct the coding operation error.
[0218] Based on the same inventive concept, an electronic device is also provided in the embodiment of the present application. In this embodiment, the structure of the electronic device can be as follows Figure 6 As shown, it includes a memory 601 , a communication module 603 and one or more processors 602 .
[0219] Memory 601 is used to store computer programs executed by processor 602. Memory 601 may mainly include a program storage area and a data storage area. The program storage area may store an operating system and programs required for running instant messaging functions, while the data storage area may store various instant messaging messages and operating instruction sets.
[0220] Memory 601 may be a volatile memory, such as random-access memory (RAM); a non-volatile memory, such as read-only memory, flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or any other medium capable of carrying or storing a desired computer program in the form of instructions or data structures and accessible by a computer, but is not limited thereto. Memory 601 may be a combination of the above memories.
[0221] The processor 602 may include one or more central processing units (CPUs) or digital processing units, etc. The processor 602 is configured to implement the aforementioned artificial intelligence model training method when calling the computer program stored in the memory 601 .
[0222] The communication module 603 is used to communicate with terminal devices and other servers.
[0223] The specific connection medium between the memory 601, the communication module 603 and the processor 602 is not limited in the embodiment of the present application. Figure 6 In the embodiment, the memory 601 and the processor 602 are connected via a bus 604. Figure 6 The connections between the other components are shown in bold lines, which are only for illustration and are not intended to be limiting. The bus 604 can be divided into an address bus, a data bus, a control bus, etc. For ease of description, Figure 6 The diagram shows a single thick line, but this does not indicate that there is only one bus or one type of bus.
[0224] The memory 601 stores a computer storage medium, and the computer storage medium stores computer executable instructions, which are used to implement the training method of the artificial intelligence model of the embodiment of the present application. The processor 602 is used to execute the training method of the above-mentioned artificial intelligence model. Based on the same inventive concept, the embodiment of the present application provides a computer-readable storage medium, and the computer program product includes: computer program code, when the computer program code is run on a computer, it enables the computer to execute any of the training methods of the artificial intelligence model discussed above. Since the principle of solving the problem by the above-mentioned computer-readable storage medium is similar to the training method of the artificial intelligence model, the implementation of the above-mentioned computer-readable storage medium can refer to the implementation of the method, and the repeated parts will not be repeated.
[0225] Refer to the following Figure 7 hereinafter, a computing device 700 according to this embodiment of the present application is described. Figure 7 The computing device 700 is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0226] like Figure 7 The computing device 700 is implemented as a general-purpose computing device. Components of the computing device 700 may include, but are not limited to, the at least one processing unit 701 described above, the at least one storage unit 702 described above, and a bus 703 connecting various system components (including the storage unit 702 and the processing unit 701).
[0227] Bus 703 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, and a processor or local bus using any of a variety of bus architectures.
[0228] The storage unit 702 may include a readable medium in the form of a volatile memory, such as a random access memory (RAM) 721 and / or a cache memory 722 , and may further include a read-only memory (ROM) 723 .
[0229] The storage unit 702 may also include a program / utility 725 having a set (at least one) of program modules 724, such program modules 724 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0230] The computing device 700 may also communicate with one or more external devices 704 (e.g., a keyboard, a pointing device, etc.), one or more devices that enable a user to interact with the computing device 700, and / or any device that enables the computing device 700 to communicate with one or more other computing devices (e.g., a router, a modem, etc.). Such communication may be performed via an input / output (I / O) interface 705. Furthermore, the computing device 700 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 706. Figure 7 As shown, network adapter 706 communicates with other modules used in computing device 700 via bus 703. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with computing device 700, including but not limited to microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0231] The present application also provides a computer program product. The methods described herein can be implemented in whole or in part using software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described herein are performed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device, a core network device, an OAM, or other programmable device.
[0232] A computer-readable storage medium can be implemented as a computer program product, that is, an embodiment of the present application also provides a computer-readable storage medium, which includes a computer program, and when the computer program is executed by a processor, it implements a training method for any of the above-mentioned artificial intelligence models.
[0233] The computer program or instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program or instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired or wireless method. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; an optical medium, such as a digital video disk; or a semiconductor medium, such as a solid-state drive. The computer-readable storage medium may be a volatile or non-volatile storage medium, or may include both volatile and non-volatile types of storage media.
[0234] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0235] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0236] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0237] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0238] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.
Claims
1. An artificial intelligence model training method, characterized in that: The method comprises: The central processing unit performs data preprocessing on the original training data obtained from the memory; the original training data includes a training set; The following iterative training steps are performed on the AI model to be trained until the preset iteration stop condition is reached: The graphics processor inputs the training set into the AI model to be trained, and obtains a loss function corresponding to the current parameters of the AI model to be trained; The quantum processor encodes the loss function into a quantum state; The quantum processor constructs a quantum circuit based on the loss function of the quantum state, and measures the quantum gradient of the loss function through the quantum circuit; The quantum processor updates the current parameter based on the quantum gradient and a preset regularization gradient.
2. The method according to claim 1, wherein Measuring the quantum gradient of the loss function by the quantum circuit includes: The quantum processor constructs a preset number of quantum bits into an initial state; The quantum processor adds a preset perturbation variable to the current parameters in the quantum circuit to obtain a perturbation circuit; After the quantum processor inputs the initial state into the perturbation circuit, a positive perturbation output state and a negative perturbation output state are obtained; The quantum processor determines a first expected value of an observation operator based on the positive perturbation output state, and calculates a second expected value of the observation operator based on the positive perturbation output state; The quantum processor determines a quantum gradient of the loss function based on a difference between the first expected value and the second expected value and a current gradient calculation accuracy.
3. The method according to claim 2, wherein After measuring the quantum gradient of the loss function by the quantum circuit, the method includes: The quantum processor monitors the decoherence time of the quantum bits in real time; The quantum processor updates the current gradient calculation accuracy according to the relationship between the quantum bit decoherence time and a preset threshold value.
4. The method according to claim 1, wherein The quantum processor updates the current parameter based on the quantum gradient and a preset regularization gradient, including: The quantum processor uses the product of a preset quantum gradient weighting factor and the quantum gradient as a first product; The quantum processor uses the product of a preset regularization coefficient and the regularization gradient as a second product; the regularization gradient is a correction amount used to suppress overfitting; The quantum processor uses the sum of the first product and the second product as a combined gradient; The quantum processor updates the current parameter based on the combined gradient and a preset learning rate; the learning rate is used to control the update amplitude of the current parameter during iterative training.
5. The method according to claim 1, wherein The original training data also includes a validation set; the validation set is used to verify the performance of the AI model to be trained; the validation set includes multiple validation samples; and the iteration stopping condition includes at least one of the following: The decrease in validation set loss is less than a first threshold; the validation set loss is determined by inputting the validation set into the AI model to be trained, based on the loss function of each validation sample and the number of validation samples; The increase in the accuracy of the verification set is less than a second threshold; the accuracy is determined by inputting the verification set into the AI model to be trained and based on the model output of each verification sample; The number of iterations reaches the preset number.
6. The method according to claim 1, wherein After the quantum processor encodes the loss function into a quantum state, the method further includes: The quantum processor performs error detection on the encoding operation by performing stable sub-measurement and obtains the detection result; If a quantum bit with an encoding error is detected, the quantum processor performs an error correction gate operation corresponding to the detection result on the quantum bit with the encoding error to correct the encoding operation error.
7. An artificial intelligence model training device, characterized in that: include: An acquisition unit, configured for the central processing unit to perform data preprocessing on the original training data acquired from the memory; The original training data includes a training set; The training unit is used to perform the following iterative training steps on the AI model to be trained until the preset iteration stop condition is reached: The graphics processor inputs the training set into the AI model to be trained, and obtains a loss function corresponding to the current parameters of the AI model to be trained; The quantum processor encodes the loss function into a quantum state; The quantum processor constructs a quantum circuit based on the loss function of the quantum state, and measures the quantum gradient of the loss function through the quantum circuit; The quantum processor updates the current parameter based on the quantum gradient and a preset regularization gradient.
8. An electronic device, characterized in that: include: a memory for storing program instructions; A processor is configured to call the program instructions stored in the memory and execute the steps of the method according to any one of claims 1 to 6 according to the obtained program instructions.
9. A computer-readable storage medium storing a computer program, wherein: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
10. A computer program product, characterized in that The method comprises a computer program stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the electronic device performs the steps of the method according to any one of claims 1 to 6.
Citation Information
Cited By
Non-orthogonal quantum state distinguishing method based on quantum machine learning
CN122198173A