LoRA-based large model fine tuning method, apparatus and device, and medium

Through tensor decomposition technology and quantum neural network update the core tensor of the low-rank matrix in the LoRA fine-tuning method, the problem of high computational cost in fine-tuning of hyperscale language models is solved, and more efficient model fine-tuning effect is achieved.

CN120216852APending Publication Date: 2025-06-27ORIGIN QUANTUM COMPUTING TECH (HEFEI) CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510285286.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When dealing with hyperscale language models, the existing LoRA fine-tuning methods still require high computational costs. How to further optimize the fine-tuning methods to handle hyperscale language models more efficiently has become a hot research direction in the academic and industrial circles.

Method used

By utilizing tensor decomposition technology and quantum neural networks, the target data is decomposed to obtain the results of characterizing the data through the low-rank matrix, and the core tensors of the low-rank matrix are updated through the data processed by multiple decomposition and quantum neural networks to achieve more refined data processing and feature extraction.

Benefits of technology

This method can take into account the characteristics and changes of the data more comprehensively, making the update of core tensors more targeted and effective, so as to converge in a better direction during the model fine-tuning process and achieve more efficient model fine-tuning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216852A_ABST
    Figure CN120216852A_ABST
Patent Text Reader

Abstract

The invention discloses a large model fine tuning method and device based on LoRA, equipment and a medium, and belongs to the technical field of quantum computing, the method comprises the following steps: decomposing target data by using a tensor decomposition technology to obtain a result representing that the target data passes through a first low-rank matrix as first data, the target data being training data of a target large model, and the training data being training data of the target large model; the first low-rank matrix is represented by a core tensor used for decomposing the target data; based on the first data, obtaining second data through a quantum neural network used for extracting high-dimensional features; based on the second data, a tensor decomposition technology is utilized to obtain a result representing the target data passing through a first low-rank matrix and a second low-rank matrix, the result serves as third data, and the second low-rank matrix is represented by a core tensor utilized by current decomposition; and respectively updating the core tensor representing the first low-rank matrix and the core tensor representing the second low-rank matrix by using the third data. By applying the embodiment of the invention, more efficient model fine tuning can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of quantum computing technology, and in particular to a large model fine-tuning method, device, equipment and medium based on LoRA. Background Art

[0002] Pre-trained Large Language Models (LLMs) are undoubtedly one of the most eye-catching technological breakthroughs in the field of Natural Language Processing (NLP) in recent years. These models can accurately capture the common features and intrinsic patterns of language by performing unsupervised learning on massive text data, and thus show excellent performance in many language tasks. Pre-trained LLMs can not only generate natural, fluent and logical text, but also flexibly adapt to various specific tasks such as text classification, question-answering systems, machine translation, etc. through fine-tuning. However, the performance of LLMs is often limited by the number of its parameters - generally speaking, the more parameters, the richer the knowledge the model can master and the better the performance. But at the same time, training these models requires massive computing resources and data, which makes them complicated and time-consuming in practical applications.

[0003] To meet this challenge, the current mainstream fine-tuning method is Parameter-EfficientFine-Tuning. This method only trains a small part of the model's parameters, thus alleviating the problem of resource consumption to a certain extent. In recent years, inspired by the inherent low-rank characteristics of large language models, with the continuous emergence of new deep learning technologies, the fine-tuning technology of large language models has made breakthrough progress. Among them, the most representative method is LoRA (Low-Rank Adaption, low-rank adapter, or low-rank matrix). Since fine-tuning is only performed on a few trainable parameters, LoRA greatly reduces the number of training parameters required for downstream tasks, thereby significantly reducing the GPU memory and time costs required for fine-tuning using large language models, bringing new possibilities for fine-tuning large models.

[0004] Even with an efficient fine-tuning method like LoRA, it still requires considerable computational costs to process very large models with tens of billions of parameters. Therefore, how to further optimize the fine-tuning method to more efficiently process very large-scale language models is still a hot research direction in academia and industry, and a difficult problem that needs to be solved urgently. Summary of the invention

[0005] The purpose of this application is to provide a large model fine-tuning method, device, equipment and medium based on LoRA, aiming to achieve efficient parameter fine-tuning.

[0006] An embodiment of the present application provides a method for fine-tuning a large model based on LoRA, and the method includes:

[0007] Using tensor decomposition technology, decompose the target data to obtain the result of the target data after the first low-rank matrix as the first data, where the target data is the training data of the target large model, and the first low-rank matrix is characterized by the core tensor used for decomposing the target data;

[0008] Based on the first data, obtain the second data through a quantum neural network for extracting high-dimensional features;

[0009] Based on the second data, use tensor decomposition technology to obtain the result of the target data after the first low-rank matrix and the second low-rank matrix as the third data, where the second low-rank matrix is characterized by the core tensor used for the current decomposition;

[0010] Use the third data to update the core tensor representing the first low-rank matrix and the core tensor representing the second low-rank matrix respectively.

[0011] Optionally, the tensor decomposition technology is the matrix product operator technology.

[0012] Optionally, the step of using tensor decomposition technology to decompose the target data to obtain the result of the target data after the first low-rank matrix as the first data includes:

[0013] Use matrix reset operation to process the target data to obtain the first matrix;

[0014] Perform matrix product operation on the first matrix and the corresponding core tensor to obtain the second matrix;

[0015] By performing matrix reset operation on the second matrix, obtain the result of the target data after the first low-rank matrix as the first data.

[0016] Optionally, the step of obtaining the result of the target data after the first low-rank matrix as the first data by performing matrix reset operation on the second matrix includes:

[0017] Perform matrix reset operation on the second matrix, and use the second matrix after matrix reset operation as the new target data, and return to perform the step of using matrix reset operation to process the target data to obtain the first matrix until all core vectors have performed matrix product operation, and obtain the result of performing matrix reset operation on the current second matrix as the first data.

[0018] Optionally, the step of obtaining the second data through a quantum neural network for extracting high-dimensional features based on the first data includes:

[0019] Perform dimensionality reduction on the first data to obtain fourth data;

[0020] Extract high-dimensional features of the fourth data through a quantum neural network to obtain second data.

[0021] Optionally, the extracting high-dimensional features of the fourth data through a quantum neural network to obtain second data includes:

[0022] Extract high-dimensional features of the fourth data through a quantum neural network to obtain fifth data;

[0023] Fuse the fourth data and the fifth data to obtain second data.

[0024] Optionally, the fusing the fourth data and the fifth data to obtain second data includes:

[0025] Perform an element-wise multiplication operation on the fourth data and the fifth data to obtain a multiplication result;

[0026] Fuse the multiplication result with the fourth data to obtain second data.

[0027] Optionally, the fusing the fourth data and the fifth data to obtain second data includes:

[0028] Perform a linear weighted combination operation on the fourth data and the fifth data, and obtain the added result as the second data.

[0029] Optionally, the obtaining the result representing the target data after the first low-rank matrix and the second low-rank matrix based on the second data by using tensor decomposition technology includes:

[0030] Perform dimensionality increase on the second data;

[0031] Use tensor decomposition technology to perform tensor decomposition on the second data after dimensionality increase, and obtain the result representing the target data after the first low-rank matrix and the second low-rank matrix.

[0032] Another embodiment of the present application provides a large model fine-tuning device based on LoRA, and the device includes:

[0033] A first decomposition module, configured to decompose target data by using tensor decomposition technology to obtain a result representing the target data after the first low-rank matrix as first data, where the target data is training data of a target large model, and the first low-rank matrix is represented by a core tensor used for decomposing the target data;

[0034] A feature extraction module, configured to obtain second data based on the first data through a quantum neural network for extracting high-dimensional features;

[0035] A second decomposition module, configured to obtain, based on the second data and by using a tensor decomposition technique, a result of the target data after a first low-rank matrix and a second low-rank matrix as third data, where the second low-rank matrix is characterized by a core tensor used in the current decomposition;

[0036] A replacement module, configured to update, by using the third data, a core tensor characterizing the first low-rank matrix and a core tensor characterizing the second low-rank matrix respectively.

[0037] Another embodiment of the present application provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the method for fine-tuning a large model based on LoRA in any of the above embodiments is implemented.

[0038] Another embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a computer, the computer executes the method for fine-tuning a large model based on LoRA in any of the above embodiments.

[0039] Traditional LoRA generally updates parameter matrices based on a fixed low-rank decomposition strategy to fine-tune a large model, while the present application updates the core tensors characterizing the first low-rank matrix and the second low-rank matrix respectively by using the third data. This way of updating the core tensors based on the data processed by multiple decompositions and a quantum neural network performs more refined processing and feature extraction on the data. From the original data to the third data finally used for updating, a more complex and in-depth feature mining process is experienced, which can more comprehensively consider the features and changes of the data, make the update of the core tensors more targeted and effective, and is more conducive to the model converging in a better direction during the fine-tuning process, thereby achieving more efficient model fine-tuning. Description of the Drawings

[0040] Figure 1 It is a schematic block diagram of an example system for the method for fine-tuning a large model based on LoRA provided by an embodiment of the present application;

[0041] Figure 2 It is a flowchart of a method for fine-tuning a large model based on LoRA provided by an embodiment of the present application;

[0042] Figure 3 It is a schematic diagram of a quantum circuit of a quantum neural network provided by an embodiment of the present application;

[0043] Figure 4 It is a schematic diagram of processing data by using MPO provided by an embodiment of the present application;

[0044] Figure 5 The structural diagram of a large model fine-tuning device based on LoRA provided by an embodiment of the present application;

[0045] Figure 6 The structural schematic diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0046] The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and should not be construed as a limitation to the present application.

[0047] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0048] Classical computers use transistors to encode information in binary data, such as bits, where each bit can represent a value of 1 or 0. These 1s and 0s are used as switches to drive the functions of classical computers. If there are n bits of data, there are 2n possible classical states, and only one state is represented at a time.

[0049] Quantum computers use quantum processors that operate on data represented by qubits, also known as quantum bits. A qubit can represent the classical binary states "0", "1", or a superposition state of "0" and "1". Due to the ability to represent a superposition of "0" and "1", a qubit can represent both the "0" and "1" states simultaneously. For example, if there are n bits of data, then quantum states can be represented simultaneously. Further, the qubits in the superposition can be correlated with each other, called entanglement, where the state of one qubit (whether it is 1 or 0 or both) can depend on the state of another qubit, and more information can be encoded within two entangled qubits. Based on the principles of superposition and entanglement, qubits can enable quantum computers to perform functions that may be relatively complex and time-consuming for classical computers.

[0050] Please refer to Figure 1 , which shows an example system block diagram of a method for fine-tuning a large model based on LoRA provided by an embodiment of the present application. The system 100 can be a hybrid computing system including a combination of one or more quantum computers, quantum systems, and / or classical computers. As Figure 1In the example shown, system 100 may include a quantum system 110 and a classical computer 120. In one embodiment, the quantum system 110 and the classical computer 120 may be configured to communicate via one or more of a wired connection and / or a wireless connection (e.g., a wireless network). The quantum system 110 may include a quantum chipset composed of one or more quantum chips, and the quantum chipset includes various hardware components for processing data encoded in qubits. The quantum chipset may be a quantum computing core surrounded by infrastructure to protect the quantum chips from sources of electromagnetic noise, mechanical vibration, heat, and other noise sources that would degrade the performance of the quantum chips. The classical computer 120 may be electronically integrated with the quantum system 110 via any suitable wired and / or wireless electronic connection.

[0051] In Figure 1 the example shown, the quantum system 110 may be any suitable set of components capable of performing quantum operations on a physical system. Quantum operations, such as quantum logic gate operations, can manipulate the quantum states of qubits to evolve and / or entangle. In Figure 1 the example embodiment shown, the quantum system 110 may include a measurement and control integrated machine 111, an interface 112, and a quantum chip 113. In some embodiments, all or part of each of the measurement and control integrated machine 111, the interface 112, and the quantum chip 113 may be located in a cryogenic environment to assist in performing quantum operations. The quantum chip 113 may be any hardware capable of processing information using quantum states. This hardware may include a plurality of qubits and means for coupling or entangling the qubits in order to process information using quantum states. Qubits may include, but are not limited to, charge qubits, flux qubits, phase qubits, spin qubits, and ion qubits. The quantum chip may include a set of quantum logic gates configured to perform quantum logic operations on the qubits stored in a quantum register. Quantum gates may include one or more single-qubit gates, two-qubit gates, and / or other multi-qubit gates.

[0052] The measurement and control integrated machine 111 can be any combination of digital computing devices capable of performing quantum computing (e.g., performing quantum circuits) in combination with the interface 112. The digital computing device can include a digital processor and a memory for storing and executing quantum instructions using the interface 112. The digital computing device can also include a communication protocol device having a communication protocol for receiving instructions and sending the results of the performed quantum computation to a classical computer. Additionally, the digital computing device can also include a communication interface having the interface 112. In one embodiment, the measurement and control integrated machine 111 can be configured to receive classical instructions (e.g., from the classical computer 120) and convert the classical instructions into measurement and control instructions for the interface 112. The measurement and control instructions provided by the measurement and control integrated machine 111 to the interface 112 can be, for example, digital signals indicating which quantum gates in the quantum gates need to act on the qubits to perform a specific function. The interface 112 can be configured to convert these digital signals into analog signals (e.g., analog pulses of microwave pulses), and the analog signals can be used to apply quantum gates on the qubits to manipulate the interaction between the qubits.

[0053] The interface 112 can be a classical-quantum interface, including a combination of devices capable of receiving instructions from the measurement and control integrated machine 111 and converting the instructions into devices for implementing quantum operations. In one embodiment, the interface 112 can turn the instructions from the measurement and control integrated machine 111 into drive signals that can drive or manipulate the qubits, and / or act quantum gates on the qubits. Additionally, the interface 112 can be configured to convert the signals received from the quantum chip 113 into digital signals that can be processed and transmitted by the measurement and control integrated machine 111. The devices included in the interface 112 can include, but are not limited to, digital-to-analog converters, analog-to-digital converters, waveform generators, attenuators, amplifiers, optical fibers, lasers, and filters. The interface 112 can further include circuit components configured to measure multiple qubits after the quantum gate acts, where the measurement can produce results represented in classical bits. Each measurement performed by the interface 112 can be read out to a device connected to the quantum system 110, such as the classical computer 120. The multiple measurement results provided by the interface 112 can represent probability results.

[0054] The classical computer 120 can include hardware components such as processors and storage devices (e.g., including memory devices and classical registers) for processing data encoded in classical bits. In one embodiment, the classical computer 120 can be configured to provide various control signals, instructions, and data encoded in classical bits to the quantum subsystem 110. Further, the quantum states measured by the quantum system 110 can be read out by the classical computer 120, and the classical computer 120 can store the measured quantum states as classical bits in classical registers. In one embodiment, the classical computer 120 can be any suitable combination of computer-executable hardware and / or computer-executable software capable of executing the preparation module 121 to perform quantum computing using the data stored in the data storage module 122 as part of the construction and calculation. The data storage module 122 can be a repository for the data to be analyzed using quantum computing algorithms and the results of such analysis. The preparation module 121 can be a program or module capable of preparing classical data from the data storage module 122 as part of the quantum circuit implementation. The preparation module 121 can be instantiated as part of a larger algorithm, such as a function call of an application programming interface (API), or by resolving hybrid classical-quantum computing into quantum and classical aspects of the computing. For example, the preparation module 121 can generate instructions for creating a quantum circuit using quantum gates. In an embodiment, such instructions can be stored by the measurement and control integrated machine 111 and can instantiate components of the interface 112 to execute such that the quantum operations of the quantum gates can be performed on the quantum chip 113.

[0055] The classical computer 120 can be a laptop computer, a desktop computer, a vehicle-integrated computer, a smart mobile device, a tablet device, and / or any other suitable classical computing device. Additionally or alternatively, the classical computer 120 can also operate as part of a cloud computing service model, such as software as a service (SaaS), platform as a service (PaaS), or infrastructure as a service (IaaS). The classical computer 120 can also be located in a cloud computing deployment model, such as a private cloud, a community cloud, a public cloud, or a hybrid cloud.

[0056] The following provides a detailed introduction to LoRA involved in this application.

[0057] The core idea of LoRA is to reduce the number of parameters required for fine-tuning by introducing low-rank matrices on the basis of a pre-trained model, reduce the video memory occupancy, and when updating the parameters, only update the two low-rank matrices, reduce the computational overhead, thereby improving the training efficiency and avoiding overfitting. For LLMs models, the core approach of LoRA fine-tuning is to freeze the weights of the pre-trained model and only train the two low-rank matrices A and B.

[0058] The mathematical expression of LoRA fine-tuning is:

[0059] h = W0x + ΔWx = W0x + BAx

[0060] where x is the input vector, i.e., the training data, W0 is the pre-trained weight matrix, W0 ∈ R d*d , ΔW is the incremental parameter matrix, and A and B are the weight matrices for low-rank adaptation, i.e., low-rank matrices, B ∈ R d*r , A ∈ R r*d , r << d. It can be found that the number of parameters to be fine-tuned changes from the original d * d to 2 * d * r, significantly reducing the number of parameters and thus significantly reducing the video memory occupancy.

[0061] For LoRA fine-tuning, during training, the pre-trained weight matrix W0 is frozen, which means that during forward and backward propagation, the corresponding gradients are not calculated and its parameters are not updated, and only A and B are trained. The solution of this application realizes LoRA fine-tuning by fine-tuning the core tensors representing the low-rank matrices.

[0062] See Figure 2 , Figure 2 A large model fine-tuning method based on LoRA provided by an embodiment of this application, the method includes the following steps:

[0063] S201: Using tensor decomposition technology, decompose the target data to obtain the result after the target data passes through the first low-rank matrix as the first data, where the target data is the training data of the target large model, and the first low-rank matrix is represented by the core tensor used to decompose the target data.

[0064] The target large model is a pre-trained large language model. A tensor is a multi-dimensional array and can be regarded as a high-dimensional generalization of a vector (first-order tensor) and a matrix (second-order tensor). Tensor decomposition is the process of representing a high-order tensor as a combination of several low-order tensors (usually matrices or vectors). In the embodiment of this application, the data after tensor decomposition technology decomposition can be the first data, and the first data contains the core tensor, and the core tensor represents one of the low-rank matrices in LoRA, that is, the matrix A. When there is more than one core tensor used for decomposition, these core tensors jointly represent the first low-rank matrix. The core tensor can be initialized using kaiming initialization or Xavier initialization, etc. Through tensor decomposition technology, the high-dimensional data is converted into a low-rank matrix representation, reducing the data dimension and computational complexity.

[0065] S202: Based on the first data, obtain the second data through a quantum neural network for extracting high-dimensional features.

[0066] The core of a quantum neural network lies in its utilization of the superposition and entanglement properties of quantum states, enabling it to efficiently represent high-dimensional features in Hilbert space. Assume the input data x ∈ R n is encoded into the quantum state |ψ(x)>. After a series of quantum gate operations, the system evolves into a new quantum state |φ(x)>. By measuring |φ(x)>, a classical output y ∈ R m can be obtained. This process can be described by the following formula:

[0067] y q ∈ M(|φ(x)>)

[0068] where M(·) represents the measurement operation. Although y q is classical data, its generation process involves non-linear transformations of quantum states and can capture features that are difficult to extract by classical methods.

[0069] In some embodiments of this application, the quantum neural network may include an encoding layer, a variational layer, and a measurement layer, specifically as Figure 3 shown. In the figure, dotted lines are used to distinguish different layers. The encoding layer consists of RY gates applied to each qubit. The encoding layer encodes the input data into the quantum circuit using angle encoding. The variational layer, through quantum state transformation, utilizes the superposition and entanglement properties of quantum states to efficiently represent high-dimensional features in Hilbert space. It mainly includes a first variational sub-layer and a second variational sub-layer. The first variational sub-layer successively includes RY gates applied to each qubit and n - 1 controlled RZ gates. The control qubits of the n - 1 controlled RZ gates are the qubits from the highest bit to the lowest bit in sequence, where n is the number of qubits in the quantum neural network. The second variational sub-layer includes RY gates applied to each qubit, an RZ gate with the control qubit being the highest bit qubit and the target qubit being the qubit adjacent to the lower bit of the control qubit, an RZ gate with the control qubit being the lowest bit qubit and the target qubit being the qubit adjacent to the lower bit of the control qubit, and n - 2 controlled RZ gates. The target qubits of the n - 1 controlled RZ gates are the qubits from the lowest bit to the (n - 2)-th bit in sequence. The RY gates and controlled RZ gates in the variational layer all contain variational parameters. After the quantum state passes through the variational layer, the quantum state changes. By measuring the Pauli Z of each qubit through the measurement layer, the output of the quantum neural network is obtained.

[0070] S203: Based on the second data, using tensor decomposition technology, obtain the result representing the target data after the first low-rank matrix and the second low-rank matrix as the third data, where the second low-rank matrix is characterized by the core tensor utilized in the current decomposition.

[0071] S203 and S201 utilize the same tensor decomposition technique, except for the different input data. The input of S203 is the dimensionality-reduced data processed by S201, and the third data obtained is ΔWx, that is, BAx. There is a difference in the initialization of the second low-rank matrix and the first low-rank matrix. When the second low-rank matrix is represented by a single core tensor, all elements of this core tensor are initialized to 0 when it is initialized. When the second low-rank matrix is represented by multiple core tensors, the initialization methods of other core tensors are the same as those of the first low-rank matrix, and the last core tensor is initialized to 0. Such initialization enables the core tensor to represent the low-rank matrix B in LoRA fine-tuning.

[0072] S204: Use the third data to update the core tensors representing the first low-rank matrix and the core tensors representing the second low-rank matrix respectively.

[0073] Through optimization algorithms such as the gradient descent method, the corresponding gradients can be calculated using the third data, and then the corresponding core tensors can be updated. Adjust the parameters of the core tensors to make the decomposed data better fit the original training data, thereby optimizing the model performance.

[0074] The target large model can be composed of multiple layers. When fine-tuning the model based on LoRA, each layer corresponds to a LoRA layer. The third data in this application is the output of a LoRA layer, and one third data can be used as the input of the next LoRA layer. The direction of data transmission is the same as that of the target large model data transmission.

[0075] Traditional LoRA generally updates the parameter matrix based on a fixed low-rank decomposition strategy to fine-tune the large model. However, this application uses the third data to update the core tensors representing the first low-rank matrix and the second low-rank matrix respectively. This way of updating the core tensors based on the data processed by multiple decompositions and quantum neural networks performs more refined processing and feature extraction on the data. From the original data to the final third data used for updating, it has experienced a more complex and in-depth feature mining process, which can consider the features and changes of the data more comprehensively, making the update of the core tensors more targeted and effective, and more conducive to the model converging in a better direction during the fine-tuning process, thereby achieving more efficient model fine-tuning.

[0076] In this application, a quantum large model fine-tuning framework is constructed through tensor decomposition technology and quantum neural networks. The core idea of this framework is to re-parameterize the LoRA layer of the pre-trained large model into a quantum tensor hybrid network, and use the quantum neural network to capture complex transformations, so as to achieve parameter-efficient fine-tuning beyond LoRA. Specifically, by utilizing the learning capabilities of quantum neural networks and tensor networks for data synthesis, parameter-efficient fine-tuning is achieved on the basis of parameter-efficient fine-tuning (PEFT), and it can run on real quantum computers. Due to the reduction in the number of parameters, the model is easier to train. In some cases, from the perspective of the step size, the model convergence can be accelerated by 20% (the change rate of the loss value per unit time), achieving a better fitting effect with a shorter training step size, reducing the risk of overfitting, and having the potential to reduce training costs. In addition, the tensor network emphasizes the local correlation of input signals, which helps to prevent the training data from falling into certain local minima. The solution of this application has higher training efficiency in the low-rank space and can improve the expressive ability of the current low-rank fine-tuning network.

[0077] In some embodiments of this application, the tensor decomposition technology can be the matrix product operator technology or the Tucker decomposition technology. Tucker decomposition is a mathematical method that decomposes a high-dimensional tensor into a core tensor and multiple factor matrices, efficiently representing the high-dimensional tensor. Through the combination of the core tensor and factor matrices, the number of parameters can be significantly reduced, making it more efficient in storage and calculation. The matrix product operator (MPO) is a tensor network structure used to represent high-dimensional operators. Specifically, it represents a high-dimensional operator as a product of a series of low-dimensional matrices.

[0078] The principle of MPO is introduced below to elaborate on the reduction of the number of parameters. In LoRA, the dimension of the weight matrix must match the dimension of the training data to ensure that the model can correctly process the input data. When processing the target data, since the number of parameters of the weight matrix is reduced, the number of fine-tuned parameters is also reduced. In this application, the implementation of MPO is to rearrange the elements of the original weight matrix W into a higher-dimensional tensor. Define the dimensions of the input space X and the output space Y as N x and N y , specifically, through matrix resetting, the original weight matrix is mapped to a high-order tensor:

[0079]

[0080] Satisfying the following dimension constraint conditions:

[0081]

[0082] Reset the input vector x ∈ X and the output vector y ∈ Y to multi-dimensional tensors (i1, i2, …, i n ), (j1, j2, …, j n ), and then decompose the weight matrix W into multiple tensor components w (k) by tensor decomposition, which can be specifically expressed as:

[0083]

[0084] where the tensor component satisfies D k-1 → D k , D k is the k-th value in the key-value vector, just in a different representation. The key-value parameter D = max{D k} controls the expressive power of the large model, and its value is positively correlated with the quantum entanglement entropy. Establishing a controllable balance mechanism between model complexity and expressive power provides a new technical path for the lightweight design of large-scale neural networks.

[0085] In terms of parameter optimization, the total number of trainable parameters satisfies:

[0086]

[0087] When using the uniform key-value D k ≡ D, it simplifies to:

[0088]

[0089] Compared with the traditional number of parameters, exponential compression can be achieved when n ≥ 3.

[0090] In some possible implementation manners of the present application, by using the tensor decomposition technology, decompose the target data to obtain the result representing the target data after the first low-rank matrix as the first data, where the target data is the training data of the target large model, including:

[0091] Use matrix reset operation to process the target data to obtain the first matrix;

[0092] Perform matrix multiplication operation on the first matrix and the corresponding core tensor to obtain the second matrix;

[0093] By performing matrix reset operation on the second matrix, obtain the result representing the target data after the first low-rank matrix as the first data.

[0094] When performing tensor decomposition on target data using MPO, by resetting the target data, the target data is mapped to a first matrix. The dimension of the first matrix is determined by the input dimension decomposition parameter, the key-value vector, and the dimension of the target data. The dimension of the first matrix is: [mat_ranks[i] * inp_modes[i], the volume of the target data / [mat_ranks[i] * inp_modes[i]]]. The dimension of the core tensor is [out_modes[i] * mat_ranks[i + 1], mat_ranks[i] * inp_modes[i]]. out_modes[i] represents the output dimension decomposition parameter i, and mat_ranks[i] is the i-th value in the key-value vector. Exemplarily, the dimension of the target data is [2, 8, 16], the input dimension decomposition parameter: inp_modes = [4, 4], the key-value vector mat_ranks = [1, 8, 1], the output dimension decomposition parameter: out_modes = [4, 2]. The number of input dimension decomposition parameters is determined by the number of core tensors. When i = 0, the dimension of the first matrix is [1 * 4, 2 * 8 * 16 / 4], specifically [4, 64]. The dimension of the core tensor is mat_cores = [4 * 8, 1 * 4] = [32, 4]. After multiplying the first matrix by the corresponding core tensor, the dimension of the second matrix obtained is [32, 64]. At this time, the second matrix needs to be reset to [4, 512]. Because the 0-th output dimension decomposition parameter is 4, the first dimension is 4, and the second dimension is 32 * 64 / 4 = 512.

[0095] When there is only one core tensor, the second matrix represents the result of the target data after the first low-rank matrix. When there are multiple core tensors, perform a matrix reset operation on the second matrix, and use the second matrix after the matrix reset operation as the new target data. Return to execute the step of processing the target data using the matrix reset operation to obtain the first matrix until all core vectors have performed matrix multiplication operations, and obtain the result of performing the matrix reset operation on the current second matrix as the first data. Continuing the above example, for an input data, the processing flow using MPO is as Figure 4 shown. The dimension of the finally obtained first data is [2, 8, 8]. Because the input data is [2, 8, 16], the first dimension 2 usually represents the batch size of the data, the second dimension 8 usually represents the number of data features or sequence length, and the third dimension 8 usually represents the dimension of each feature. Therefore, the dimension of the first data is [2, 8, 8]. The method of obtaining the third data using MPO is the same. The difference is that with different parameter settings, the dimension of the third data needs to match the output dimension of the target large model.

[0096] In some embodiments of the present application, obtaining the second data based on the first data by using a quantum neural network for extracting high-dimensional features includes:

[0097] Performing dimensionality reduction processing on the first data to obtain fourth data;

[0098] The high-dimensional features of the fourth data are extracted through a quantum neural network to obtain the second data.

[0099] Reduce high-dimensional data to a lower dimension for more efficient processing and analysis. Common dimensionality reduction techniques include principal component analysis (PCA), linear discriminant analysis (LDA), t-SNE, and UMAP. These methods project data into a low-dimensional space through different mathematical transformations while retaining the main features and structure of the data as much as possible. Dimensionality reduction is done to reduce quantum resource usage on the one hand, and to reduce information loss when a high-dimensional matrix is ​​directly reduced to a low-rank matrix on the other.

[0100] In some embodiments of the present application, extracting the high-dimensional features of the fourth data through a quantum neural network to obtain the second data includes:

[0101] Extracting high-dimensional features of the fourth data through a quantum neural network to obtain fifth data;

[0102] The fourth data and the fifth data are fused to obtain the second data.

[0103] Tensor decomposition technology can be implemented using classical neural networks. Classical neural networks can generate non-harmonic functions, while QNN is good at fitting truncated Fourier series. In order to make up for the insufficient expression ability of classical neural networks in low-rank space, it is necessary to fuse the data of classical neural networks and QNN, which can not only learn smooth features, but also fill the non-harmonic gaps by processing data in a classical way, thereby significantly reducing training errors and improving generalization ability.

[0104] In some embodiments of the present application, the fourth data and the fifth data can be fused to obtain the second data in two ways. The first is to perform an element-by-element multiplication operation on the fourth data and the fifth data to obtain the multiplication result; and fuse the multiplication result with the fourth data to obtain the second data. The other is to perform a linear weighted combination operation on the fourth data and the fifth data to obtain the addition result as the second data, and the weighted weight is adjusted as the large model is fine-tuned.

[0105] To make up for the insufficient expressive power of tensor decomposition technology in the low-rank space, a fusion operation is performed on the fourth data and the fifth data. Both the multiplication result and the addition result incorporate the high-dimensional features extracted by the QNN while maintaining the low-rank structure, thereby enhancing the model's ability to model complex non-linear relationships. Through element-wise multiplication operations or linear weighted combination operations, the output fifth data of the QNN provides additional non-linear features for the fourth data. The elements in the fifth data contain high-order feature interaction information related to the input data, which is ignored during the low-rank approximation process. Therefore, the expressive power of the multiplication result or the addition result can theoretically be improved. When using the element-wise multiplication method, after obtaining the multiplication result, the multiplication result can be added to the fourth data to obtain the second data.

[0106] In some embodiments of the present application, obtaining the result representing the target data after the first low-rank matrix and the second low-rank matrix based on the second data includes:

[0107] Performing dimensionality increase processing on the second data;

[0108] Using tensor decomposition technology, performing tensor decomposition processing on the second data after dimensionality increase processing to obtain the result representing the target data after the first low-rank matrix and the second low-rank matrix.

[0109] Dimensionality increase can increase the dimension of the data, map the data to a higher space, making it easier to capture non-linear relationships in the data and enabling the data to express more complex features and patterns. When using tensor decomposition technology to process the second data, the core tensor used represents the second low-rank matrix, and the second low-rank matrix is responsible for performing dimensionality increase processing on the data after dimensionality reduction processing to restore it to the original high-dimensional space. This combination not only retains the main features of the data but also can extract more representative high-dimensional features, thereby improving the expressive power and performance of the model. In this way, the output can be compared and combined with the output of the target large model. Although dimensionality increase increases the dimension of the data, tensor decomposition, through core tensor decomposition, constrains the update of the original weight matrix to the product of low-rank matrices, thus significantly reducing the number of parameters to be learned and the computational complexity.

[0110] In the present application, both dimensionality reduction operations and dimensionality increase operations can be implemented using fully connected layers. When fine-tuning the large model through the update of the core tensor, the variational parameters of the quantum neural network are also updated, and the optimization parameters in the fully connected layer are also updated. When fusing data using linear weighted combination, the weighted weights are also updated. Since the ranks of the first low-rank matrix and the second low-rank matrix are much smaller than the rank of the weight matrix, whether it is the quantum neural network or the fully connected layer, the number of parameters to be updated is relatively small. Adding these parameters, the total number of parameters to be updated is also much smaller than the number of parameters that need to be updated for the weight matrix.

[0111] See Figure 5 , Figure 5 A large model fine-tuning device provided by an embodiment of the present application. The device may include:

[0112] A first decomposition module 501, configured to decompose target data by using tensor decomposition technology to obtain a result representing the target data after a first low-rank matrix as first data, where the target data is training data of a target large model, and the first low-rank matrix is represented by a core tensor used for decomposing the target data;

[0113] A feature extraction module 502, configured to obtain second data based on the first data through a quantum neural network for extracting high-dimensional features;

[0114] A second decomposition module 503, configured to obtain a result representing the target data after a first low-rank matrix and a second low-rank matrix as third data based on the second data by using tensor decomposition technology, where the second low-rank matrix is represented by a core tensor used for the current decomposition;

[0115] A replacement module 504, configured to update the core tensor representing the first low-rank matrix and the core tensor representing the second low-rank matrix respectively by using the third data.

[0116] Regarding the specific functions and effects achieved by the above-mentioned large model fine-tuning device based on LoRA, reference may be made to other embodiments of the present application for explanation, which will not be elaborated here. Each sub-circuit in the large model fine-tuning device based on LoRA can be implemented in whole or in part by software, hardware, and their combination. Each sub-circuit can be embedded in or independent of a processor in a computer device in hardware form, or stored in a memory in a computer device in software form, so that the processor can call and execute operations corresponding to each of the above sub-circuits.

[0117] Please refer to Figure 6 , an embodiment of the present application also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the large model fine-tuning method based on LoRA in any of the above embodiments. Please refer to Figure 6 , the computer device may be a classical computer. The computer device may also be a quantum computer.

[0118] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a computer, the computer executes the large model fine-tuning method based on LoRA in any of the above embodiments.

[0119] The embodiments of the present application also provide a computer program product containing instructions, which when executed by a computer cause the computer to execute the LoRA-based large model fine-tuning method in any of the above embodiments.

[0120] It can be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the various processes do not mean the order of execution, and the order of execution of the various processes should be determined by their functions and internal logics, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0121] It can be understood that the various embodiments described in the present application can be implemented alone or in combination, and the embodiments of the present application do not limit this.

[0122] Unless otherwise specified, all technical and scientific terms used in the embodiments of the present application have the same meaning as commonly understood by those skilled in the technical field of the present application. The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the scope of the present application. The term "and / or" used in the embodiments of the present application and the appended claims includes any and all combinations of one or more of the related listed items. The singular forms "a", "above", and "the" used in the embodiments of the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0123] It can be understood that the processor in the embodiments of the present application can be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method embodiments can be completed by the integrated logic circuit in the hardware of the processor or the instructions in software form. The above-mentioned processor can be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by the hardware decoding processor, or executed by the combination of the hardware and software sub-circuits in the decoding processor. The software sub-circuit can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0124] It can be understood that the memory in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM). It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0125] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0126] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0127] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0128] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0129] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0130] If the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

Claims

1. A large model fine-tuning method based on LoRA, characterized in that: The method comprises: Decomposing the target data by using tensor decomposition technology to obtain a result representing the target data after passing through a first low-rank matrix as the first data, wherein the target data is training data of a target large model, and the first low-rank matrix is ​​represented by a core tensor used to decompose the target data; Based on the first data, obtaining second data by using a quantum neural network for extracting high-dimensional features; Based on the second data, using a tensor decomposition technique, a result representing the target data after being processed by a first low-rank matrix and a second low-rank matrix is ​​obtained as third data, wherein the second low-rank matrix is ​​represented by a core tensor used in the current decomposition; The core tensor representing the first low-rank matrix and the core tensor representing the second low-rank matrix are updated respectively using the third data.

2. The method according to claim 1, characterized in that The tensor decomposition technique is a matrix product operator technique.

3. The method according to claim 2, characterized in that The method of decomposing the target data by using the tensor decomposition technology to obtain a result representing the target data after passing through the first low-rank matrix as the first data includes: Using a matrix reset operation, the target data is processed to obtain a first matrix; Performing a matrix product operation on the first matrix and the corresponding core tensor to obtain a second matrix; By performing a matrix resetting operation on the second matrix, a result representing the target data after passing through the first low-rank matrix is ​​obtained as the first data.

4. The method according to claim 3, characterized in that The performing of the matrix resetting operation on the second matrix to obtain a result representing the target data after the first low-rank matrix is ​​passed through as the first data includes: A matrix reset operation is performed on the second matrix, and the second matrix after the matrix reset operation is used as the new target data, and the step of using the matrix reset operation to process the target data to obtain the first matrix is ​​returned, until all core vectors have performed the matrix product operation, and the result of performing the matrix reset operation on the current second matrix is ​​obtained as the first data.

5. The method according to any one of claims 1 to 4, characterized in that: The method of obtaining second data based on the first data by using a quantum neural network for extracting high-dimensional features comprises: Performing dimensionality reduction processing on the first data to obtain fourth data; The high-dimensional features of the fourth data are extracted through a quantum neural network to obtain the second data.

6. The method according to claim 5, characterized in that The step of extracting the high-dimensional features of the fourth data through a quantum neural network to obtain the second data includes: Extracting high-dimensional features of the fourth data through a quantum neural network to obtain fifth data; The fourth data and the fifth data are fused to obtain the second data.

7. The method according to claim 6, characterized in that The fusing the fourth data and the fifth data to obtain the second data includes: Performing an element-by-element multiplication operation on the fourth data and the fifth data to obtain a multiplication result; The multiplication result and the fourth data are combined to obtain second data.

8. The method according to claim 6, characterized in that The fusing the fourth data and the fifth data to obtain the second data includes: A linear weighted combination operation is performed on the fourth data and the fifth data to obtain an addition result as the second data.

9. The method according to claim 5, characterized in that The method of obtaining, based on the second data, a result representing the target data after being processed by the first low-rank matrix and the second low-rank matrix by using a tensor decomposition technique comprises: Performing dimensionality upgrading processing on the second data; The second data after the dimensionality increase processing is subjected to tensor decomposition processing by using the tensor decomposition technology to obtain the result of characterizing the target data after the first low-rank matrix and the second low-rank matrix.

10. A large model fine-tuning device based on LoRA, characterized in that: The device comprises: A first decomposition module is used to decompose the target data by using a tensor decomposition technology to obtain a result representing the target data after passing through a first low-rank matrix as the first data, wherein the target data is training data of a target large model, and the first low-rank matrix is ​​represented by a core tensor used to decompose the target data; A feature extraction module, configured to obtain second data based on the first data by using a quantum neural network for extracting high-dimensional features; A second decomposition module is used to obtain, based on the second data, a result of representing the target data after passing through a first low-rank matrix and a second low-rank matrix by using a tensor decomposition technology, as third data, wherein the second low-rank matrix is ​​represented by a core tensor used in the current decomposition; The replacement module is used to use the third data to update the core tensor representing the first low-rank matrix and the core tensor representing the second low-rank matrix respectively.

11. A computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the LoRA-based large model fine-tuning method according to any one of claims 1 to 9 is implemented.

12. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a computer, the computer executes the LoRA-based large model fine-tuning method according to any one of claims 1 to 9.

Citation Information

Cited By

  • Model fine tuning method and related product

    CN120494023A

  • Prediction model implementation method and device, prediction model application method and device, equipment, medium and product

    CN120880925A

  • Response text generation method and device

    CN121189294A